A&A INSIGHTS
Read-only or write access: the line to draw in a first AI delivery
Asked to automate the customer's data entry too? A&A splits writes into reversible, compensable and irreversible to settle what a first delivery writes and who undoes it.
日本語で読む
THE STARTING POINT
The writes that belong in a first delivery are the ones where the target system has an undo path and your side can execute that undo. Writes needing a separate offsetting operation, and writes that cannot be recalled once they leave the building, are priced as separate work. The difference between read and write is not how hard the connection is — it is who restores the original state when it fails.
The answer: cap the first delivery at writes you can undo yourself
A&A perspective
The writes that belong in a first delivery are the ones where the target system has an undo path and your side can execute that undo. Writes that cannot be undone and can only be offset by a different operation, and writes that cannot be recalled once they leave the building, stay out of the first contract and get priced as separate work. When a customer says "while you're at it, have it do the registration too", the question to answer is not "can you" but "when this fails, who puts it back".
A&A perspective
As connection work, read and write look like the same job. You take the credentials, map the fields, handle the errors. The difficulty is not where the difference lives. The difference shows up after a failure. If a read does not land, a proposal simply does not appear and the customer's data is untouched. If a write does not land, the contents of the customer's system have changed. And unless the contract says so, nobody has been assigned the job of changing them back.
A&A perspective
This article keeps three things apart: what the sources state technically, A&A's reading of that for scoping client work, and a hypothetical worked example. The sources are Anthropic engineering posts and a customer story. They do not establish the validity of any clause under Japanese contract practice, nor conformance with your audit or internal-control requirements. Take the contract wording to your own counsel.
Anthropic: choose the tools you will not build, and disclose destructive changes
From the sources
Anthropic's tool design post, published on 11 September 2025, puts "Choosing the right tools to implement (and not to implement)" first among its principles for writing high-quality tools. The same post names, as a commonly observed error, "tools that merely wrap existing software functionality or API endpoints" — whether or not those tools are appropriate for agents.
From the sources
The post gives concrete substitutions. Instead of a list_contacts tool that returns every contact, build search_contacts or message_contact. Instead of a read_logs tool, build a search_logs tool that returns only the relevant log lines and some surrounding context. It then states: "Make sure each tool you build has a clear, distinct purpose", and says that careful, selective planning of the tools you build or don't build can really pay off.
From the sources
Near the end, the post contains a line that maps directly onto scoping. If you are writing tools for an MCP server, tool annotations "help disclose which tools require open-world access or make destructive changes". In other words, inside a list of tools, the ones that cause changes you cannot take back are treated as something to be declared separately from the rest.
A&A perspective
What A&A borrows here is the shape of the estimate. The post does not call "we support every endpoint" good design. It puts the decision not to build first among its principles, and gives destructive operations the status of something declared. Read across to scoping client work, that means stopping treating the customer's API list as a list of hours, and pulling the operations that make destructive changes out into a separate column. The source, however, is stating software design principles only; it says nothing about contracts or the division of liability. That reading is A&A's own.
| Kind of write | Include in the first delivery? | What to settle before including it |
|---|---|---|
| Read and propose only | The default scope | Attach to each proposal the record it rests on, as of when, and which fields could not be read |
| Reversible (an undo exists and your side can run it) | Include it | The undo procedure, and who is told that it was reversed, and when |
| Needs compensating (no undo, but an offsetting operation exists) | Conditionally, as a separate quote | Who executes the offsetting operation, and whether doing so needs the customer's approval |
| Irreversible (cannot be recalled once outside) | Not at first | The pre-execution confirmer, the material they see, and what happens when confirmation is late |
| Listed explicitly as out of scope | Outside the scope | Name the writes you cannot classify rather than letting them in while still ambiguous |
Chatbase: checking an invoice and issuing a refund sit inside the same integration
From the sources
Anthropic's Chatbase customer story describes the platform as "integrating with business systems through APIs, enabling AI agents to take direct actions". For the Stripe integration it says that "agents can handle tasks like checking invoices and processing refunds", and attaches a condition to that: "with optional human oversight for sensitive operations".
From the sources
Earlier on the same page the same arrangement is written as spanning "from checking order status in Stripe to processing refunds with human oversight". The page also says the company is streamlining its human-in-the-loop approach to enable "efficient oversight through simple one-click confirmations".
A&A perspective
What matters to A&A is that a read and a write sit side by side in that one sentence. Checking an order's status and processing a refund are both written as work inside the same Stripe integration. As a connection they are the same. Yet the phrase about human oversight is attached only to the refund side — and that oversight is described as optional. The presence of a check is split per operation, and that split is itself an object of design. That is the structure showing through.
A&A perspective
It is worth being equally clear about what the passage does not say. The page sets out no criterion for which operations count as "sensitive". It does not report whether the oversight actually stopped an error. The fact that a confirmation takes one click is a separate matter from whether the person confirming receives what they need in order to judge. This is a vendor-published account, not an independent audit. The page's figures about user adoption carry no stated population or period, so this article does not treat them as a forecast.
Split writes three ways: reversible, compensable, irreversible
A&A perspective
From here on this is A&A's design. Taking up the distinctions the sources draw — destructive changes, and sensitive operations — scoping client work means splitting writes three ways. The criterion is neither difficulty nor the breadth of the permission. It is the route back.
A&A perspective
The first kind is the reversible write. The target system has an undo procedure, and your side can execute that undo: saving a draft, changing a status, adding or removing a tag, appending to a notes field. Get it wrong and you put it back yourself. No call to the customer, no waiting for approval.
A&A perspective
The second kind is the write that needs compensating. The operation itself cannot be undone, but a different operation exists to offset it: issue an invoice, then refund it; send a notice, then send a correction; reserve stock, then release it. Nothing returns to its prior state, but the books can be squared. Who executes the offsetting operation, though, and whether executing it needs an approval, is set by the customer's operating practice rather than by the system.
A&A perspective
The third kind is the irreversible write: once it is outside, it cannot be recalled. Mail that reached the customer's own customer, a payment that executed, a shipping instruction that was committed. Either no offsetting operation exists, or one exists but the fact that the other party saw it does not go away. If an automated write of this kind is in scope, the check has to sit before execution, because there is no route back afterwards.
A&A perspective
This three-way split cannot be settled mechanically from the target system's feature list. The same "update record" API is reversible if the result stays inside the company, and irreversible if the record is configured to notify the customer's own customer. So the classification work is done by looking at the customer's operating practice, not at the API specification. That point is A&A's judgement; the source is speaking within the context of system design.
Four columns to fill in for each write operation

A&A perspective
To settle the classification, and to have something to put in the contract, fill four columns for each write operation. Four lines per operation; five operations means twenty lines. Produce an hours figure before that work is done and what grows later is not the hours — it is the liability.
A&A perspective
The first column is the undo method. Can the target system's API undo it, does a person reverse it from an admin screen, or is there no way back? Treat "probably reversible" as a blank. The second column is who restores the original state: your side, a named person on the customer's side, or nobody decided. A write that went into operation with this column blank is the one that costs the most later.
A&A perspective
The third column is the confirmer, and the material that person sees. "The customer checks it" is not enough; write the role alongside the screen or notification they can actually judge from. The fourth column is how long it takes for a failure to become visible. Does it surface as an error the moment it executes, does it first appear as a discrepancy at the customer's month-end close, or does it reach you as a complaint from the customer's own customer?
A&A perspective
The column A&A expects to go missing most often is the fourth — an editorial view about how scope gets decided, not something we have counted. When it is missing, the column that fixes it is the third. If a failure is visible at the moment of execution, your side can serve as the confirmer. If it stays invisible until the close, or arrives as someone else's complaint, your side is not positioned to notice it, so either the confirmer has to sit on the customer's side or the operation stays out of automated writing altogether. Which of those you choose is the scope of the first delivery.
From the sources
The idea of putting the check before execution appears in the sources as well. Anthropic's post on building effective agents, published on 19 December 2024, says agents can "pause for human feedback at checkpoints or when encountering blockers", and that it is also common to include stopping conditions such as a maximum number of iterations to maintain control. The same post says value appears in work that requires both conversation and action, has clear success criteria, enables feedback loops and can "integrate meaningful human oversight".
What the estimate gains is not the connection — it is undo design and the confirmer's time
From the sources
The same post on building effective agents is explicit about the cost of running something autonomously. For open-ended problems, it says, "you must have some level of trust in its decision-making", and that an agent's autonomy means higher costs and the potential for compounding errors. On that basis it recommends "extensive testing in sandboxed environments, along with the appropriate guardrails".
A&A perspective
In A&A's reading, what lands on the estimate once a write is in scope is not connection hours. Four things land. Designing and documenting the undo procedure. Making the pre-execution check something the customer can actually click. Testing in an isolated environment — that is, building somewhere the same path can run without writing to the customer's production system. And the escalation route for a failure. The connection itself barely differs from the read-only case.
A&A perspective
There is one more cost that does not go on your invoice but should still be shown to the customer: the confirmer's time. Putting a check before execution means somebody on the customer's side judges each item. The "one-click" phrasing in the Chatbase story looks like an argument that the burden is small, but multiply it by volume and it becomes the customer's workload. However light the check, unless you multiply it by how many arrive per day and show the customer that figure, checks pile up once operations begin and the writes waiting on them stop. How to estimate where human checking jams is covered in "Before taking on more AI delivery work, estimate where human review jams".
A&A perspective
The reverse case also exists: declining to automate a write can be the expensive choice. Leave a reversible write in human hands and someone on the customer's side keeps re-typing, and their transcription errors are their own burden. So including reversible writes in the first delivery is the default, and choosing not to include one calls for a stated reason. That ordering is A&A's design, not a rule the sources state.
"Read-only" is not free of responsibility either: what you promise is traceability
A&A perspective
Narrowing scope to reading does not reduce the promise to zero. The output of a read becomes the basis for a write performed by a person. If the staff member who sees the proposal trusts it and keys it in, an error in the proposal enters the customer's system as an error in the entry. Confining the scope to reading leaves the responsibility for the quality of the proposal in place.
From the sources
The tool design post also hints at what should be promised here. It says tool implementations should take care to return only high-signal information, that they should "prioritize contextual relevance over flexibility", and that they should eschew low-level technical identifiers such as uuid or mime_type in favour of fields like name and image_url, which are much more likely to directly inform an agent's downstream actions and responses.
A&A perspective
One boundary is worth marking: that passage is about the design of what a tool returns to an agent, not about the design of what a deliverable shows a person. With that said, in A&A's reading what a read-only delivery promises is traceability: which record this proposal rests on, as of when that reference was taken, and which fields could not be retrieved. With those three attached, the staff member can check the proposal's arithmetic. A proposal without them is an instruction the staff member cannot verify, which leaves two options — re-check every item by hand, or key it in unverified. Either one erases the point of having narrowed the scope to reading.
A hypothetical example: taking on registration from inquiry into an estimate ledger
Hypothetical example
The following is hypothetical. It is not a real customer and not an A&A result. The customer is a facilities installation company that registers leads arriving through its web form and by email into an internal estimate ledger. The work being contracted is to pull the site address, the requested timing and the type of work out of the inquiry text and create a new row in the ledger. The customer has said: "while you're at it, do the staff assignment and the acknowledgement to the client too".
Hypothetical example
The writes split three ways. Creating a new ledger row is reversible: the ledger supports deleting and editing rows, and your account can do it. The staff assignment needs compensating: the assignment itself can be changed, but it is configured to notify the assigned person, the notification cannot be recalled, and a wrong assignment therefore needs a correcting message. The acknowledgement to the client is irreversible: once it reaches the customer's own customer, there is no way back.
Hypothetical example
Filling the four columns exposes a gap in how long failures take to become visible. A mistaken field in a ledger row is apparent as soon as the customer's staff look at the ledger. A wrong staff assignment surfaces when the notified person realises the site is not in their area. An acknowledgement sent with the wrong type of work is not apparent until the client replies about a different job — and that reply does not reach your side at all.
Hypothetical example
So the first delivery stops at creating the new ledger row. For the staff assignment, writing a proposed assignee into the ledger's notes field is included as a reversible write, while confirming it — which triggers the notification — is executed from the ledger by the customer's own staff. The acknowledgement to the client is excluded from the first contract; once three months of ledger rows exist and the pattern of mistaken fields is visible, it is quoted separately. The estimate lists the ledger connection hours, and separately the hours for the undo procedure document and for building a copy of the ledger to test against.
A&A perspective
What A&A wants to stress about this example is that nothing was refused. One of the three goes in first, one is reduced to a reversible form, and one is deferred to a stated point in time. Rather than saying no, the three are reordered by whether a route back exists. Being able to explain the basis for that reordering to the customer is what the four columns are for.
When this line is unnecessary, and what is not being claimed
A&A perspective
There are situations where this design is unnecessary. If the customer already runs an internal approval flow, the confirmer is the existing approver in that flow, and your side does not need to build a new confirmation step. Double the checks and control does not increase while the schedule does. And if the write target is a platform your own side supplied, restoration was yours from the start, so the question shrinks back to pricing alone.
A&A perspective
Volume is a condition too. At a few writes a month, the hours for automation plus undo design will almost never be beaten by a person typing. This three-way split is a tool for work where the same kind of write recurs.
From the sources
The nature and the age of the sources should also be stated. The post on building effective agents was published on 19 December 2024, and the page itself now carries a note that much of the tooling landscape it describes has changed since that date. What this article draws from it is limited to cost, compounding error and testing in isolated environments — claims that do not depend on specific tooling — but the publication date is old.
A&A perspective
Here is what is not being claimed. Neither source shows that classifying writes by reversibility raises margin or reduces incidents, and A&A has not measured it. The human-oversight passage in the Chatbase story is an account of a design choice, not evidence that the oversight worked. The validity of any clause under Japanese contract practice, the treatment of non-conformity liability, and conformance with your industry's audit requirements are outside this article's scope — take them to your counsel. No A&A estimating accuracy or project margin figure is presented.
The next step: on the job you are stuck on, write only the undo method
A&A perspective
There is no need to fill all four columns at once. For the job where you are currently stuck on whether to take the write, fill only the first column. Can the target system's API undo that write? Can a person reverse it from an admin screen? Or is there no way back? If it is the third, that operation is not in the first delivery's scope — it is a separate conversation about who confirms it before execution.
A&A perspective
If you have not yet decided which unit of work to sell as your first product, "Choosing your first paid AI service: how a solo founder carves out one sellable unit of work" comes first. The whole path from winning customers through to keeping them is laid out in "AI-native GTM: a practical guide for solo founders and small teams". Once the write scope is settled and you need the implementation as well, we quote that as development.
Read and write look like the same connection job, but they differ in who restores the original state when it fails. Anthropic's tool design post puts choosing which tools not to implement first among its principles, and says MCP tool annotations help disclose which tools make destructive changes. The Chatbase story places checking invoices and processing refunds inside the same Stripe integration, with optional human oversight for sensitive operations. So the scope of a first delivery is decided by the route back. Make reversible writes — the ones your side can undo — the default; for writes that need an offsetting operation, and for writes that cannot be recalled, settle the pre-execution confirmer and the undo procedure first and quote them separately. What the estimate gains is not connection hours but the undo design, the test environment, and the confirmer's time on the customer's side. Take the legal validity of any clause to your own counsel.
Sources & editorial note
Primary pages read for this article. Publication dates below belong to the sources; access dates record our research.
- Writing effective tools for agents
Anthropic · 2025-09-11
Accessed 2026-10-02 - Chatbase helps companies deliver instant, personalized customer support with Claude
Anthropic · undated customer story page; no publication date shown
Accessed 2026-10-02 - Building effective agents
Anthropic · 2024-12-19
Accessed 2026-10-02
AI-assisted editorial production
A&A uses AI for research, writing, translation and editorial checks. Source facts, our analysis and hypothetical examples are labeled separately.
Editorial check: 2026-10-02