A&A INSIGHTS
How Much Customer Data to Send to an AI: Scoping Input in Contract Work
Split customer data into three lanes before starting. Removing information the judgment does not use raises accuracy, and custody narrows only when access is revoked.
日本語で読む
THE STARTING POINT
Data you take custody of from a customer should be split into three lanes — always passed to the model, referenced only when needed, and never passed — and agreed before the project starts. Passing everything just in case is not the safe choice: removing information the judgment does not use is also an accuracy requirement, and what you hold does not narrow by keeping data out of the model — it narrows only when the access permission is removed.
The answer: split customer data into always-passed, referenced-on-demand, and never-passed
A&A perspective
When a customer hands you a whole shared folder, the first decision is not which model to use. It is to divide the data into three lanes. First, what is always passed to the model: the limited set of information the task needs for every single judgment. Second, what is referenced only when needed: data for which you keep only the location and the way to look it up, retrieving the part you need partway through processing. Third, what is never passed: data you have decided this project's judgments will not use, and which therefore never enters the model. Put that three-way split in writing before you start, and get the customer's agreement to it. For anything in the never-passed lane, settle access in the same conversation: decline the grant, have the permission removed, or write down why it has to remain and until when. Not passing data to the model and not holding it are different things. Alongside that, decide who signs the contract with any third-party service the data passes through, and where you will record which data went through which service and when.
A&A perspective
Passing everything just in case looks like the safe side because of an intuition that too much is less dangerous than too little. In contract work it cuts the other way, and one distinction is needed to see why. What you hold is set by the access the customer granted, not by the volume you send to the model. The moment you are given permission over the whole shared drive, the scope of what you must account for in a leak, and of what you must confirm deleted when the project ends, covers that whole drive. Deciding not to put something into the model does not narrow that scope. Narrowing it means declining the grant or having the permission removed. That is why the responsibility grows while the estimate does not.
A&A perspective
There is a second, technical reason. More information does not reliably mean more accuracy. The volume of context and the correctness of the model's judgment are not in simple proportion, as the next sections confirm against the sources. To state the conclusion first: narrowing the data is not a concession made for the lawyers. It points in the same direction as designing for accuracy. That is why the two concerns can be settled as one decision. Joining them into one decision is, however, A&A's own framing, not a claim either source makes.
Anthropic: context is a finite resource with diminishing marginal returns
From the sources
Anthropic's write-up on context engineering, published on 29 September 2025, describes context rot: as the number of tokens in the context window increases, the model's ability to accurately recall information from that context decreases. Some models degrade more gently than others, but the page states that "this characteristic emerges across all models." From that it concludes: "Context, therefore, must be treated as a finite resource with diminishing marginal returns."
From the sources
The same page defines the goal directly: good context engineering means "finding the smallest possible set of high-signal tokens that maximize the likelihood of some desired outcome."
From the sources
The page explains why. LLMs have an "attention budget" they draw on when parsing large volumes of context, and every new token depletes that budget by some amount. Because the transformer architecture lets every token attend to every other token, n tokens produce n-squared pairwise relationships, so as context length grows the model's ability to capture those relationships gets stretched thin — a natural tension between context size and attention focus. The page is careful to call this "a performance gradient rather than a hard cliff": models remain highly capable at longer contexts, but may show reduced precision for information retrieval and long-range reasoning compared with shorter contexts.
A&A perspective
Translated into contract work, dropping in an entire shared folder is not "more material for the judgment." It is spending the attention budget on information the judgment does not use. One caution matters here: the source is describing model performance. It says nothing about confidentiality or legal liability. The only thing this section carries forward is that narrowing what you pass has a reason on the accuracy side as well.
| Data lane | How it reaches the model | Decide before starting |
|---|---|---|
| Always passed (needed for every judgment, limited in volume) | Kept in context | The use the customer agreed to, and the scope at column granularity |
| Referenced on demand (large, only a sliver needed per case) | Keep location and lookup only; fetch at runtime | Whether the customer has access control and a record of who saw what, when |
| Never passed (decided out of this project's judgments) | Not loaded | Why it is out of scope, and whether access is removed or retained (with an end date) |
| Undecided (lane not yet assigned) | Treated as never-passed until decided | When and with whom it gets decided, and whether permission is removed meanwhile |
| The part that passes through a third party | List every provider it transits | Contracting party, where the record lives, and end-of-project handling |
Rather than pre-processing everything, keep identifiers and load on demand
From the sources
The same page contrasts pre-processing all relevant data up front with a "just in time" context strategy. Agents built that way "maintain lightweight identifiers (file paths, stored queries, web links, etc.)" and use those references to dynamically load data into context at runtime using tools.
From the sources
As an example, the page describes Anthropic's own agentic coding tool performing complex data analysis over large databases: the model writes targeted queries, stores results, and uses Bash commands like head and tail to analyze large volumes of data without ever loading the full data objects into context. The page likens this to human cognition — we do not memorize entire corpuses, but build external organization and indexing systems such as file systems, inboxes and bookmarks to retrieve what is relevant on demand.
From the sources
The page also credits autonomous navigation with enabling progressive disclosure: each interaction yields context that informs the next decision, with file sizes suggesting complexity, naming conventions hinting at purpose, and timestamps acting as a proxy for relevance. The effect, it argues, is that the agent keeps its working context focused on relevant subsets rather than drowning in exhaustive but potentially irrelevant information.
A&A perspective
What contract work can take from this is the distinction between having access to data and having passed the data. If the customer's data can be retrieved when it is needed, it does not have to sit in context all the time. That distinction is the substance of the middle lane — referenced only when needed. Note again that the source frames this as a question of accuracy and efficiency, not of custody. The responsibility side of the argument comes from a different source, in the next section.
NIST AI RMF: third-party data and resources need documented, followed procedures
From the sources
The Core of the NIST AI Risk Management Framework 1.0 (2023) states, as Map 4, that "risks and benefits are mapped for all components of the AI system including third-party software and data." Beneath it, Map 4.1 requires that approaches for mapping the technology and legal risks of those components — "including the use of third-party data or software" — "are in place, followed, and documented," and it brings risks of infringing a third party's intellectual property or other rights into the same scope.
From the sources
In the same Core, the MANAGE function states as Manage 3 that "AI risks and benefits from third-party entities are managed," and Manage 3.1 requires that "AI risks and benefits from third-party resources are regularly monitored, and risk controls are applied and documented."
From the sources
Map 1.1 requires that "intended purposes, potentially beneficial uses, context-specific laws, norms and expectations, and prospective settings in which the AI system will be deployed are understood and documented." Its listed considerations include the specific set or types of users along with their expectations, the potential positive and negative impacts of system uses on individuals, communities, organizations, society and the planet, and the assumptions and related limitations about the system's purposes, uses and risks.
A&A perspective
What a contract shop can read from this is that the requirement is not "choose a safe model." The requirement is a state of affairs: for the part that travels through a third party, there is a procedure, it is followed, and there is a record. Separately from the quality of the deliverable, you need to be able to explain after the fact which data went through which third party. And because "regularly monitored" is written explicitly, the obligation does not end when the decision is made — it continues for as long as the system is in operation. In an estimate, this is the part that quietly belongs to maintenance.
Folding the legal question and the accuracy question into one design decision
A&A perspective
The two sources so far say different things. The context-engineering write-up is about model accuracy; the NIST Core is about documenting risk that travels through third parties. Neither tells you to narrow customer data. What follows is A&A's own framing.
A&A perspective
The framing is this. "Passing less raises accuracy" and "narrowing what you hold reduces responsibility" are two different causal claims. The first is conditional: it holds only when what you remove is information the judgment does not use. Cut the relevant information and accuracy falls. The second comes from the structure of contract work — the range you hold is the range you must account for and confirm deleted. They are different pieces of work, but they are settled on the same single sheet, because deciding which data the judgments will not use is what identifies which permissions can be removed. So they do not need to be thought through twice.
A&A perspective
Where this framing earns its keep is an internal disagreement. The engineering side believes more context means more accuracy; the governance side says reduce what you send. While those look like opposing positions, the resolution tends to be "pass everything for now and think about it later." Once it is clear that removing information the judgment does not use also favours accuracy, the two stop being opposed and become the same piece of work.
A&A perspective
The opposite misreading has to be avoided too: it is not "the less the better." The source itself warns that overly aggressive compaction "can result in the loss of subtle but critical context whose importance only becomes apparent later." The purpose of the three-way split is not to minimize volume. It is to decide in advance which data belongs in which lane, and to write that down.
Defining the three lanes, and the fourth row

A&A perspective
The always-passed lane is for information the task needs for every judgment, limited in volume, and either free of personally identifying data or covered by the customer's agreement to use it that way. A glossary of internal terms, a summary of past decisions, the list of products in scope: things that are small and change slowly. Designing this lane means keeping "just in case" out of it.
A&A perspective
The referenced-on-demand lane holds data for which you keep only the location and the way to look it up, retrieving what you need partway through processing. Operational records, transaction detail, the bodies of past inquiries: things that are large, and of which any one judgment needs only a sliver. Whether data can live in this lane depends on whether the customer already has access control and a record of who looked at what and when. What to do when they do not is covered further below.
A&A perspective
The never-passed lane is data you have decided this project's judgments will not use. The distinction that matters is between data you cannot use and data you have decided not to use. Write one line giving the reason, and when someone later asks why the accuracy is what it is, you can answer that the data was placed out of scope. Leave the reason out and the same question comes back as a design failure instead. Then settle access on the same row: decline it, have the permission removed, or write why it must remain and until when. Keeping data out of the model is input design; narrowing what you hold is permission design, and they are two separate pieces of work.
A&A perspective
The third lane is, in our editorial view, the one most often skipped — that is a judgement about how scope gets decided, not a count we have taken. Decide only what is always passed and what is referenced on demand, leave the rest sitting in the shared folder, and you effectively have two lanes. What remains is data you hold without having decided what it is for — responsibility without purpose. Deciding not to pass something is not a passive choice, but it does not by itself close the boundary either. The boundary closes when the permission is removed; the never-passed decision is the precondition for doing that.
A&A perspective
Data you cannot yet classify gets an explicit fourth row: undecided. The point is to stop undecided data from drifting into the always-passed lane, so treat it as never-passed until someone decides, and settle it with the customer when it actually comes up. Write on that row when, and with whom, it will be decided. Without that, undecided rows do not resolve; they survive to the end of the project.
Six columns to fill in for each kind of data
A&A perspective
Once the lanes are set, fill in six columns for each kind of data. More columns is not better. These six were chosen as the items that, while blank, mean you are not ready to start.
A&A perspective
Column one is the lane and the reason: which of always-passed, on-demand, never-passed or undecided, and one line on why it sits there. Column two is the handling of access: declined, permission removed, or retained with a reason and an end date. This is the only column that actually narrows what you hold. Column three is the name of any third-party service the data passes through, and whether you or the customer is the contracting party with that provider. Leave this undecided and run it on your own account, and the customer's data travels under your contractual terms. In most cases that is not the design anyone intended.
A&A perspective
Column four is the access work needed on the customer's side for on-demand retrieval: is there permission management and an audit record, and if not, who builds it and when. Column five is where the record lives — where, and at what granularity, you log which data went through which third party and when. Column six is what happens when the project ends: deleted, returned to the customer, or retained, and if retained, where and for how long. Any extract or index you built yourself belongs in this column too.
A&A perspective
Of the six, only columns one and five can be filled in from your side alone. Columns two, three, four and six need the customer's decision. So the sheet is built as material for the meeting before work starts. Try to fill it in afterwards and implementation proceeds with column four blank, until it stalls waiting for permissions that never arrive.
From the sources
Columns three and five correspond to the places where the NIST AI RMF Core requires that approaches covering the use of third-party data or software "are in place, followed, and documented," and that "AI risks and benefits from third-party resources are regularly monitored, and risk controls are applied and documented." What the Core sets out, though, are outcomes to achieve, not a document format; the same page notes that a separate companion Playbook offers suggested tactical actions organizations can apply within their own contexts. The design of these six columns is A&A's.
Three things this split adds to the estimate
From the sources
The source is explicit about the cost of retrieving on demand: "Of course, there's a trade-off: runtime exploration is slower than retrieving pre-computed data." It adds that without opinionated and thoughtful engineering to give the model the right tools and heuristics, an agent "can waste context by misusing tools, chasing dead-ends, or failing to identify key information."
A&A perspective
So what the estimate gains is not the work of connecting a model. It gains three things: the access work on the customer's side — both granting what on-demand retrieval needs and asking for permission to be removed on what you decided not to use; the mechanism that logs which data went where; and the check on whether this task can tolerate the added latency. The third is a requirements check rather than a volume of work, but skipping it means redoing the design after implementation.
A&A perspective
The access setup on the customer's side cannot be finished from your side alone. It needs time from whoever runs their IT, or from whoever is doing that job alongside another one. So do not fold it into "our work": carve it out as separately estimated work. Fold it in and the wait that follows looks like your delay. Carve it out and the location of the delay is shared from the start.
A&A perspective
How much latency is acceptable depends on the task. A first reply to an inbound inquiry that must land in seconds and a monthly reconciliation that can take minutes do not permit the same amount of on-demand retrieval. The three-way split is not a fixed right answer; it is decided together with the response-time requirement. Where that requirement is tight, rebuilding the always-passed lane smaller is sometimes more realistic than moving more data to on-demand.
Putting it on one sheet: first-line inquiry triage for a retail chain (hypothetical)
Hypothetical example
What follows is hypothetical. It is not an A&A engagement; it is a setting constructed to show the three-way split concretely. Suppose a ten-store retailer asks you to take on first-line triage of the inquiry email arriving at each store: classify it, and check whether the same case has come up before. The customer hands over access to a shared drive in one go. Inside are a product master, daily inventory files, three years of inquiry email, an employee roster, and per-store sales detail.
Hypothetical example
The always-passed lane gets three columns from the product master — name, model number, warranty period — plus store locations and opening hours, and the list of inquiry categories. Small, slow-changing, and not personally identifying. Taking three columns rather than the whole product master is how this lane is built.
Hypothetical example
The on-demand lane gets the bodies of three years of inquiry email. Triaging one message needs only whether a handful of past exchanges exist for the same model number; there is no reason for three years to sit in context. Keep the location and the search condition, and fetch the matching few at runtime.
Hypothetical example
The never-passed lane gets the employee roster and the per-store sales detail: triage does not use them. Each gets one line of reason — "not needed for classification or past-case lookup." The access column says read permission on those two folders is to be removed, and it actually is removed: not putting them into the model leaves the whole shared drive in your custody. The daily inventory files go to undecided, because whether a first reply should state stock availability has not been settled; until it is, they are treated as never-passed and the permission on the inventory folder is removed for now, and the row records with whom and when that will be decided.
Hypothetical example
The third-party column names one model provider, and the contracting party — you or the customer — is settled there. The record column says that each triage logs one line for which past messages it referenced. The end-of-project column says the retrieval index is deleted and the three-column extract built for the always-passed lane is returned to the customer.
A&A perspective
In this example, what sat outside the estimate was the customer-side decision needed to resolve the inventory files, the decision about which account gets read access to past email, and the act of removing permission on the roster and the sales detail. None of them can be settled from your side. Produce the sheet before starting and they go across together as three requests; notice them after starting and you begin by explaining why implementation has stopped.
When this split does not fit, and what is not being claimed
A&A perspective
There are conditions under which this split is unnecessary or a poor fit. If you take no custody of customer data at all and the processing stays inside the customer's own environment, columns three onward are largely empty. Building the three-way split as a formality there serves no purpose.
From the sources
The source does not argue that everything should be on demand. It says that "in certain settings, the most effective agents might employ a hybrid strategy, retrieving some data up front for speed, and pursuing further autonomous exploration at its discretion," and that the right level of autonomy depends on the task. It also suggests the hybrid "might be better suited for contexts with less dynamic content, such as legal or finance work."
A&A perspective
The case that bites, in our editorial view rather than by any count we have taken, is the reverse. When the customer has neither access management nor an audit record, pushing work toward on-demand retrieval means you end up building the permission machinery for them, and complexity rises. The alternative there is to empty the on-demand lane and build the always-passed lane as a copy containing only the minimum columns the task needs. You take on the work of making the copy and the problem of the copy going stale, and the copy itself enters your own custody — where it sits and when it is deleted must go in the end-of-project column. In exchange you can start without waiting for the customer's setup. Record that you chose this alternative, why, and when the copy gets rebuilt.
From the sources
The NIST AI RMF is a voluntary framework. The same page states of the accompanying Playbook that "Like the AI RMF, the Playbook is voluntary and organizations can utilize the suggestions according to their needs and interests." The same page also carries a notice at the top: "The AI RMF 1.0 is being updated. A revised version is in progress." What this article quotes is an excerpt from the 2023 version 1.0.
A&A perspective
What is not being claimed, stated plainly. This article does not determine compliance with Japan's Act on the Protection of Personal Information or any other law. Questions such as whether a given transfer counts as provision to a third party or as delegated processing are for the reader's own counsel. Having records aligned with the NIST Core is not proof of legal compliance. "Passing less raises accuracy" is a conditional claim, holding only where what you removed was information the judgment does not use; cut relevant information and accuracy falls. No A&A customer results, reduction rates or accuracy improvements are presented. The three-way split and the six columns are A&A's design proposal, not a format supplied by either source.
Next step: on the project you are unsure about, write only the never-passed rows
A&A perspective
If there is a project you are currently unsure about, do not try to fill in all three lanes. Write only the never-passed rows. List the data you hold, mark what this project's judgments will not use, and add one line of reason. That much needs no agreement from the customer and, depending on how long the list is, takes very little time. It is enough to show how far the boundary of what you hold had spread. Then ask the customer to remove permission on the marked rows where that is possible, and the boundary actually narrows.
A&A perspective
Then decide, for what is left, whether it is always passed or referenced on demand. Rows you still cannot classify become undecided, treated as never-passed, and go on the agenda for the meeting with the customer. Input scope and write scope are decided on different axes. Input scope turns on whether you can decide not to use something; write scope turns on who restores the original state when an action fails. Even when one project needs both decided, they cannot be decided by the same test.
A&A perspective
The write side of the boundary is covered in "顧客システムへの書き込みまで請けるか:最初のAI納品で引く線" (Deciding whether to take on writes into customer systems). The whole picture from acquisition through retention is in "AIネイティブGTMとは?一人・少人数で顧客獲得から継続まで回す実践ガイド" (the AI-native GTM guide). If the target workflow is already settled and you want to fix the implementation scope, a development conversation is the right next step. If which workflow to target is still open, that comes first, and the input-scope sheet comes after it.
The scope of customer data is not something to discuss after starting; it is something to split three ways and agree before starting. Take the three lanes — always passed, referenced on demand, never passed — add a row for what is still undecided, decide on each row whether access is removed, and settle who contracts with each third party the data transits and where the record lives. The reason to narrow what you pass is not only confidentiality. Removing information the judgment does not use is also an accuracy requirement. That is why passing everything just in case loses on both responsibility and quality. Note that the responsibility side only actually lightens when the permission is removed, not at the moment you decide to keep the data out of the model. But estimate the cost of narrowing on the same sheet: on-demand retrieval is slower than pre-fetching, and it presupposes access control on the customer's side.
Sources & editorial note
Primary pages read for this article. Publication dates below belong to the sources; access dates record our research.
- Effective context engineering for AI agents
Anthropic · 2025-09-29
Accessed 2026-10-03 - AI RMF Core
NIST AI Resource Center · excerpt from NIST AI RMF 1.0 (2023); page shows no separate date
Accessed 2026-10-03
AI-assisted editorial production
A&A uses AI for research, writing, translation and editorial checks. Source facts, our analysis and hypothetical examples are labeled separately.
Editorial check: 2026-10-03