A&A INSIGHTS
Patch again or rebuild? Decide by how many days the last stopgap held
Decide patch-versus-rebuild on how many days a stopgap holds, not on what a rebuild costs. Anthropic's CI record — 70 days, 29, then under a day — is the evidence.
日本語で読む
THE STARTING POINT
Whether to keep patching or to rebuild is decided by how many days the stopgap holds, not by what a rebuild costs. Before applying it, write the expected hold in days on one line, set it beside how many days the last stopgap actually held and a freshly redrawn rebuild estimate, and switch to the rebuild at the row where the expectation falls below the estimate.
The time to rebuild is set by how many days the last stopgap held, not by cost
A&A perspective
Whether to keep patching or to rebuild is not decided by what the rebuild costs. It is decided by which is shorter: the time a stopgap buys, or the time a rebuild takes. Before you touch the next stopgap, write down one number — how many days you expect it to hold. Then set that number beside how many days the previous stopgap actually held, and beside your current estimate for the rebuild. The row where the expected hold falls below the rebuild estimate is the time to switch. Cost never enters this decision.
A&A perspective
The reason a cost comparison stalls the decision is that cost does not tell you when. The rebuild estimate looks like roughly the same figure on the day the jam happens and three months later. The number of days a stopgap holds, by contrast, keeps shrinking as your volume grows. The money stays still while only the correct answer flips. So the number worth recording is days rather than currency, and the place to record it is not a budget document but the note you write immediately before applying the stopgap. There is a competing rule: multiply the effort each stopgap costs by how often it recurs, and set that total against the rebuild. For totals that rule is the right one, but while the frequency is still climbing it understates the next instance, and the switch comes late by exactly that much. Days are used here not to get the total right but to see the direction early.
From the sources
That both terms moved, not one, is stated plainly in Anthropic's own account of its CI, published on 14 September 2026. The author, Sachin Malhotra, first sets aside the techniques themselves — buying bigger machines, parallelising processes, restarting the service — as common and not the insight to take from the article. The point, he writes, is elsewhere: "each of these techniques bought a fraction of the time they did a year ago", and at the same time "overhauling and completely redesigning a service also takes a fraction of the time". Both terms of the comparison moved, not one.
70 days, 29 days, less than a day: three stopgaps recorded by how long they held
From the sources
The same account lines the three stopgaps up by the number of days each one held: "three quick fixes, which lasted 70 days, then 29 days, and then less than a day". The first, in October of the previous year, doubled the cores running the service; of it the article says "We also knew it would be fleeting". The second, in February, split the service's state per package so the work could run in parallel. The third, in March, was a daily restart, and "Restarting bought us less than a day".
From the sources
For background, the article states that Anthropic's engineers now ship 8x as much code per quarter as they did across 2021-2025, and that Claude authors 80% of that code. It adds that "the amount of tests across our codebase grew 10x" while only a nominal number of engineers were added, and that the result was a 25x increase in CI jobs over six months. The path by which faster generation itself jams the steps downstream of it is written out with figures attached.
A&A perspective
The value of that record is not the size of the numbers 70, 29 and 1. It is that the unit of record was days rather than money. Once the series is written in days, two rows are enough to form a view of how long the third would hold. Line the same sequence up in currency and nothing comes out of it. What is worth reproducing in your own business is not Anthropic's ratios but the shape of the record.
| Aspect | Deciding on absolute cost | Deciding on days held |
|---|---|---|
| Number recorded | The rebuild's estimated price | Days each stopgap actually held |
| When it is written | After the jam, when budgeting | Before touching the stopgap |
| What changes the decision | When cash becomes available | When expected hold falls below the rebuild estimate |
| Easily missed | The shortening trend itself | Cash on hand and seasonal load |
| Effect in a one-person business | The stopgap always looks smaller today | The comparison of sizes stays on one table |
The denominator is moving too, not just the numerator

From the sources
How far the rebuild time actually fell is also given as a figure. Of the eventual redesign — giving the service a data store so that writes could be spread across workers — the article says: "This project took three weeks for a single engineer. A year ago it would have been closer to a quarter." Three weeks for one engineer, against something near a quarter twelve months earlier.
From the sources
On the small-company side as well, rebuild times are starting to be described in days and minutes rather than months. Anthropic's report on its small-business programme, published on 10 September 2026 by Lina Ochman, its Head of U.S. SMB, says that Mike Teso, who runs a three-plant trailer manufacturer in Indiana, "built a reconciliation tool for a newly acquired factory in about 15 minutes", and that his IT director "replaced paper production schedules with live dashboards in days rather than months". Both are single instances the report states without any population, period or definition, and without independent verification.
A&A perspective
That the denominator is moving has one practical consequence: you cannot carry last year's rebuild estimate into this year's decision. If your sense that "a rebuild means three months" dates from some years ago, that is the figure which has gone stale first. For as long as that old denominator is in use, the arithmetic hands the win to the stopgap permanently. Redraw the rebuild estimate in the same sitting in which you write the expected hold of the stopgap. Updated apart, neither number means anything.
Write the expected hold before you apply the stopgap
A&A perspective
The whole procedure is this. Before you touch the stopgap, write one number: how many days you think it will hold. Not afterwards — before. The reason is that afterwards the account tends to settle into the shape of "there was nothing else to be done at the time". Only a number written in advance becomes a forecast you can later set against what happened. Whether the forecast is accurate does not matter yet at this stage. Note also that knowing a measure is temporary and being able to estimate how long it will last are two different things. With the first but not the second, all you ever record is the outcome, and no trend appears. That reading is A&A's.
From the sources
The gap between forecast and outcome is recorded on the source side too. Of the second stopgap, the sharding work, Anthropic's account says "We also knew this fix would be fleeting, but we didn't realize it would only buy us 29 days". The team knew this one was temporary as well; what they had wrong was only the number of days it would last.
Hypothetical example
Six columns are enough. The following is a hypothetical example written to show the format; these are not measured figures from A&A or from any client. The columns are: (1) what the stopgap was, (2) the date it was applied, (3) expected hold in days, (4) actual hold in days, (5) the new defect this stopgap introduced, and (6) the rebuild estimate in days as of that date. First row: "thin the review queue by hand on heavy days / 8 April / expected 60 / actual 41 / two items slipped per month / rebuild estimate 15". Second row: "publish the thinning rule as a table / 19 May / expected 40 / actual 12 / the rule goes stale unnoticed / rebuild estimate 12". As you go to write the third row and the figures come out at "expected 8 / rebuild estimate 10", the expected-hold column drops below the estimate column for the first time. That is the time to switch: the third stopgap does not get applied, and the rebuild starts instead.
The third stopgap can add defects, not merely go inert
From the sources
Of the third stopgap, the article records something beyond the effect simply wearing off. Once daily restarts began, the service drifted further behind by degrees, and on the occasions when it fell more than an hour behind, a large number of results went unrecorded. The consequence, the article states, was that "our test-selection component was using stale data" when deciding which tests to run. The same passage is explicit that this did not mean CI never ran, nor that untested code reached production, and says the effect was mostly that they ran tests "already super flaky or widespread-failing across the board".
A&A perspective
What this suggests is that the price of a stopgap is not always exhausted by the number of days until it stops working. A third-generation stopgap tends to relieve the original jam while importing a class of error that was not there before. That is what the fifth column — the new defect this stopgap introduced — is for. If that column stays blank for two rows running, suspect first that the defects are not being recorded rather than that there are none. This generalisation is A&A's reading; neither source states it.
Only you can define what "stopped holding" means
A&A perspective
To count days at all, you have to settle in advance on what counts as the stopgap having stopped holding. In the account above, pages firing and the size of the lag played that role. A solo or very small operation has no pager. What can stand in its place are events that already occur and can be counted: the number of times you intervened by hand, or the number of times you told a client something would be late. There is nothing new to instrument.
From the sources
Which step to put the threshold on follows from where the jam actually occurs. The small-business report above states that of the work attendees wanted AI to take on, about two-thirds concerned running the business and a third growing it, and that outside marketing work, reporting was the single most common use case. The examples it gives are tasks such as pulling the numbers from three systems into the Monday report. The step a reader would put a threshold on is the same step attendees most wanted to delegate.
Hypothetical example
One line is enough for the definition. As a hypothetical: "this stopgap has stopped holding = the day on which the same kind of manual intervention happens more than three times in a week". Or: "= the month in which more than two deliveries to clients run later than planned". Neither is a measured A&A figure; both are examples of the form. The soundness of "three" or "two" matters less than having decided before the stopgap went in. Decided afterwards, the threshold tends to drift to whichever side is convenient.
In a one-person business, delay comes from the stopgap looking smaller today, not from an ownership gap
From the sources
Anthropic's account contains a passage in which the reason for the delay was not technical: "Even when the trend line was clear, ownership was murky." The trend was visible, but who would own it was not. No one wanted to take on another piece of infrastructure, it continues, and the CI team had bigger priorities.
From the sources
Yet rebuilds do happen at sizes where there is no one to divide ownership with. The small-business report describes a five-person trucking compliance firm in Tennessee that lost 60% of its clients in a market downturn and then "got 30 days' notice from its core software vendor". It continues: "The team rebuilt the system themselves with Claude", took their fuel tax filing error rate from 7% to zero, and are now set up to handle twice their old peak volume with the same five people. For these figures the report gives no population, period or definition, and no independent verification.
A&A perspective
Set side by side, the two accounts show that delay in a one-person business has a different cause. It is not that ownership is ambiguous; it is that on the day, at that moment, the stopgap looks smaller. And because the only party to compare against is yourself, the comparison of sizes is settled inside your head and leaves no trace anywhere. The real function of writing an expected hold is not to forecast accurately. It is to move that comparison out of your head so the rebuild can compete on the same table. In the Tennessee case the rebuild was chosen because a deadline existed, not because a trend was read. Readers with no externally imposed deadline hold a column of days in place of that deadline.
Where this way of deciding does not hold
A&A perspective
It is unusable with one stopgap or fewer on record. The method needs a previous actual duration, so on the first jam it returns nothing; just fix the thing. The second gives you the direction of the trend; from the third you can make the switching decision. The other case where it does not bite is a business whose volume is flat. The inversion happens because the numerator shrinks, and if your throughput is not growing, the hold times are not shortening either. There, "patch it for now" remains the correct judgement and this table is one you do not need to build.
A&A perspective
A rebuild forced by a date is not a question of hold times. When a dated notice arrives from the vendor, what is fixed is the deadline, and the live comparison there is rebuild against migrate. The report above shows nothing of how that comparison went, so any reckoning of migration is A&A's conjecture. The scope of this article stops at deciding the switch yourself while no deadline has been imposed from outside. A measure that works indefinitely is also not a stopgap, so it does not belong in the table. Only put in what you already knew to be temporary when you applied it.
A&A perspective
Limits on the sources remain. Anthropic's 25x, 8x and 10x are internal operating records of one company, not an independent audit. The organisation's scale and premises differ, so they are not extrapolated here to small-scale contract delivery in Japan. Nor are 70 days, 29 days and under a day the half-life of stopgaps in general; they are one observed series for one service. And the rebuild having become faster does not contain the conclusion that you should rebuild — the article itself sets its individual techniques aside as common and not the insight. It also records that the redesigned architecture is "more expensive to run". What days govern is when to switch, not the total cost. The small-business figures (60%, 7% to zero, twice the volume, about 15 minutes, a few days) are all self-reported, and are not used here as a forecast of results or as a success rate for rebuilds.
A&A perspective
The adjacent decisions this article does not cover are handled elsewhere. Estimating the volume at which you will jam in the first place is the subject of "Before taking on more AI delivery work, estimate where human review jams", which separates ordinary review, exception handling and redelivery and counts them apart. Once you have decided to rebuild, whether a combination of existing tools is enough or you should build your own is compared in "Use existing tools or build a system of your own?", which weighs not only purchase cost but the checking effort and the maintenance. The whole picture from winning customers through to continued use is in "AI-native GTM: a practical guide for solo founders and small teams".
The reason this choice stalls every time is not missing information but the unit being recorded: money. Money tells you how much, never when. Before touching the stopgap, write the expected hold on one line, and set it beside the last actual hold and a redrawn rebuild estimate. The row where the expectation falls below the estimate is the time to switch. Until three rows of days have accumulated, the table says nothing. From the third row on, next month's decision becomes visible this month.
Sources & editorial note
Primary pages read for this article. Publication dates below belong to the sources; access dates record our research.
- Agentic coding is straining CI. Here's how we scaled test impact analysis at Anthropic
Anthropic · 2026-09-14
Accessed 2026-10-07 - What 1,000 small business owners taught us about AI
Anthropic · 2026-09-10
Accessed 2026-10-07
AI-assisted editorial production
A&A uses AI for research, writing, translation and editorial checks. Source facts, our analysis and hypothetical examples are labeled separately.
Editorial check: 2026-10-07