A&A INSIGHTS
Build now or wait for a better model: split the work into bets and assets
Sort what you build into bets a better model erases and assets that survive, then attach detection to the bets only — for solo and small teams, from primary sources.
日本語で読む
THE STARTING POINT
Whether to build now or wait for a better model comes down to whether you will notice when the capability gap your scaffolding fills has closed. Sort each feature into a bet that a better model makes unnecessary and an asset that survives either way, then write one line — the condition that says the gap closed, and the action you will take — for the bets only, before you start.
The question is not build-or-wait. It is whether you can tell when the gap closes
A&A perspective
Whether to build now or wait for the next model is not decided by how much scaffolding you have written, nor by how fast models are improving. It is decided by one thing: when the capability gap your scaffolding fills is finally closed, will you notice? There are three steps. First, take each feature you are about to build and sort it into a bet that becomes unnecessary if the model improves, or an asset that remains either way. Second, for the items you labelled bets only, write down one condition — before you start, in prose — that will tell you the gap has closed. Third, write down at the same time what you will do when that condition is met: discard the work, thin it out, or leave it in place. That is this article's answer.
A&A perspective
The reason to sort this way is that a model update does not destroy your scaffolding. It closes the capability gap your scaffolding was filling. Once the gap closes, the pre-processing and post-processing sitting on top of it have finished their job. If the gap never closes, that work stays necessary for years. So the real loss is not in having built something, but in carrying it after it has finished its job, without noticing. The only scaffolding that becomes debt is scaffolding with no means of detection. Put the other way round: write one detection condition and the loss from building on the bet side acquires a ceiling. What follows is the evidence for this judgement, and the one-page sheet to fill in before you start.
A small company's founder says plainly that the application layer is his and the models are the vendor's
From the sources
On the Anthropic case-study page for Lex, a writing platform, founder Nathan Baschez states his forward plan in these words: “We’re focused on improving our application layer, while Anthropic is focused on developing even better models for us to use”. That is a division of labour — his team concentrates on improving the application layer, while the model vendor concentrates on developing better models. The same page records the company's size as "Company size: Small". On how the model was chosen, it says: "After experimenting with various large language models, Lex chose Claude as their primary AI model for its tone, cost, and quality of output" — several models were tried, and the choice was made on tone, cost and output quality.
A&A perspective
What transfers here is the fact that the founder of a small team states publicly, and without hedging, where his own build work belongs. A policy of staying in the application layer only holds together if you assume the models will get better. When they do, part of the application layer becomes unnecessary — but the application layer can also be rebuilt on top of the better model. So "when models improve, my work disappears" is a less accurate reading than "when models improve, the contents of the application layer get swapped out." That is A&A's reading, not the page's claim. The same page names what the team is investing in next: "creating an interface for writers to compare versions, collaborate with AI, and navigate complex writing projects with ease." In this article’s terms those sit on the asset side — work the vendor will not do for you however much the models improve. That one sentence is where the policy of staying in the application layer becomes concrete about what you keep holding yourself.
A&A perspective
The page does not, however, prove this policy correct. It does not say which parts of Lex's application layer were expected to survive model improvement. The three numbers it carries — signups, churn and cost — come with no population, period, definition or independent verification, and the page itself notes “This case study was written with help from Lex!” These are self-reported figures on the vendor's own site, and no causal link is shown between the division of labour and those numbers. This article does not use any of them as evidence. It uses exactly one thing: that a named founder of a small company states this division of labour openly.
| Scaffolding | Capability gap it fills | Bet or asset / detection condition |
|---|---|---|
| Splitting long attachments before passing them to the model | Cannot handle a long input in one pass (the model's shortfall) | Bet. Condition: pass the longest attachment through unsplit and the amounts and line-item count match the original document |
| Post-processing that normalises how amounts are written | Output format is not stable (the model's shortfall) | Bet. Condition: remove the post-processing and the output is in a format the customer's accounting software ingests as-is |
| A per-customer mapping of account codes | Does not know the customer's own account names (the customer's circumstances) | Asset. No detection condition. The vendor has no motive to close this however much the model improves |
| Rules that decide who the approver is | Does not know the internal approval rules (the customer's circumstances) | Asset. No detection condition. It does not exist in public information |
| A dictionary replacing industry-specific phrasing | Vocabulary mapping (a boundary: general vocabulary is the model's side, in-house terms are not) | Write them separately. Replacing general terms is a bet; in-house terms are an asset. A condition written over both at once will never be satisfied |
Building a feature that works "well enough" as a bet on a model a few months out
From the sources
Anthropic's engineering post on agent evaluation, published Jan 09, 2026, describes its own practice this way: "Internally, we often build features that work “well enough” today but are bets on what models can do in a few months. Capability evals that start at a low pass rate make this visible. When a new model drops, running the suite quickly reveals which bets paid off." In other words, features that only work well enough today are built as bets on what models will be able to do in a few months; capability evals that begin at a low pass rate make the bet visible; and when a new model ships, running the suite shows immediately which bets paid off. The post gives the name "eval-driven development" to the eval-first practice itself, recommending you "build evals to define planned capabilities before agents can fulfill them" — write the evaluation first, to define a capability the agent cannot yet deliver. The passage about building features as bets is placed as an illustration of that practice.
A&A perspective
The important thing in that passage is not whether to bet, but the ordering: the bet and its detection are built at the same time. The evaluation is written first and the feature catches up later. The bet is not something you wait on until it pays; it is something you place after attaching an instrument that stops you from holding it once it has lost. What a solo or two-person team cannot copy is the scale of the evaluation infrastructure. The ordering does not depend on scale. Writing one line of condition before you start works in a business with no evaluation infrastructure at all.
A&A perspective
There is an asymmetry not to overlook. The party making that statement is Anthropic, which develops the models. Being able to hold "models will get better in a few months" as a premise is not the same position as running a small business on top of those models. The vendor holds both sides of the bet; a small operator holds one. The same sentence therefore carries a different expected value depending on who reads it. What this article recommends is not placing more bets, but attaching an exit condition to the ones you place.
Three questions that separate a bet from an asset

A&A perspective
Three questions are enough. First: is this scaffolding filling a shortfall in the model's capability, or something else? Second: is that shortfall the kind of thing the model vendor wants to close? Third: if it does close, can you remove the scaffolding, or has it grown into the customer's operations in a way that cannot be unpicked? If the answer to the first and second is "the model's side" in both cases, it is a bet. If either answer is "something other than the model", it is an asset. The third question checks whether an item you labelled a bet can later be removed safely; if it cannot, decide how you will remove it before you start. The bet-and-asset split, the content of these three questions and the asymmetry of attaching a condition only to the bets are all A&A design proposals. Neither of the sources cited below states them.
A&A perspective
The second question carries the weight. The gaps a model vendor wants to close are general capabilities that pay off across many users: holding long context, following instructions, keeping output format stable, reasoning further. Pre-processing and post-processing that exist to patch those are bets. Gaps like where a customer's data lives, an internal approval rule, the layout of a counterparty's form, an industry turn of phrase, a permission boundary — the vendor has no motive to close those. Scaffolding that fills them survives the model getting better. Those are the assets.
A&A perspective
This distinction has an explicit failure condition. Where the vendor has no reason to close the gap — a customer's private data, their approval rules, a counterparty's form layout — the work is an asset, and attaching a detection condition to it is pure overhead. The more instruments you add that never fire, the harder it becomes to know which instrument to read. Attach conditions only to the items you labelled bets, and attach none to the assets. Holding that asymmetry is what makes a single-page sheet work.
Hypothetical example
The following is a hypothetical example. It is not a real customer and not an A&A measurement. Suppose a solo operator delivers invoice extraction and has added four pieces of scaffolding to hold the current quality: (1) splitting long attachments before passing them to the model, (2) post-processing that normalises how amounts are written, (3) a per-customer mapping of account codes, and (4) rules that decide who the approver is. Run the first two questions over them: (1) and (2) sit on "the model's shortfall" and are bets; (3) and (4) sit on "the customer's circumstances" and are assets. Only the two labelled bets get the detection condition built in the next section.
Write the condition as a state the customer receives, not as an impression of the output
A&A perspective
One line per bet is enough, but there is a constraint on how to write it. "When the output gets better" cannot be judged. What can be judged is whether, with that scaffolding removed, the result the customer was supposed to receive still holds. In the example above, the condition for (1), the splitting step, is: pass the longest attachment through without splitting it, and see whether the resulting invoice data matches the original document on both the amounts and the number of line items. For (2), the normalisation step: remove the post-processing and see whether the output is in a format the customer's accounting software ingests as-is. Neither is an opinion about quality. Both are written as the presence or absence of a state left on the customer’s side. How to write a pass criterion as a state the customer receives is covered in detail in “When to switch models in a small AI SaaS: customer failures, not benchmarks”, so it is not repeated here. What is new here is only that the same way of writing is used for a bet’s exit condition rather than for a pass criterion.
From the sources
Separating how the output looks from what the result is, is stated outright in the same evaluation post. Using a flight-booking example, it says: "A flight-booking agent might say “Your flight has been booked” at the end of the transcript, but the outcome is whether a reservation exists in the environment’s SQL database." The agent may close by saying it booked the flight, but the outcome is whether a reservation exists in the database. The same post also treats ambiguity in task specifications as noise in metrics, stating that "Everything the grader checks should be clear from the task description."
A&A perspective
In a one-person business you need no evaluation infrastructure to run that line. What you need is to keep one or two inputs that actually broke in the past. Run those inputs once through a path with the scaffolding removed, and check by eye whether the state you wrote down holds. Three bets means three lines, and each condition should be scoped so that checking it is one run-and-eyeball. It is at the point where you start building an "evaluation suite" that you run into the cost problem covered further down.
The moment a bet turns into an asset can be observed, not predicted
From the sources
The same post distinguishes capability evals from regression evals and then describes the transition between them. Capability evals ask "What can this agent do well?" and "They should start at a low pass rate". Regression evals ask "Does the agent still handle all the tasks it used to?" and should sit at a nearly 100% pass rate. Then: "After an agent is launched and optimized, capability evals with high pass rates can “graduate” to become a regression suite that is run continuously to catch any drift. Tasks that once measured “Can we do this at all?” then measure “Can we still do this reliably?”" A capability eval whose pass rate has risen graduates into a regression suite, and a task that once measured whether something was possible at all thereafter measures whether it is still reliably done.
A&A perspective
That transition carries straight over to this decision. The day the detection condition attached to a bet is met for the first time is the day that gap closed. From then on, the same condition changes role: it stops asking "can the model do this yet?" and starts asking "is the model still doing this?" The point of the structure is that the moment a bet turns into an asset arrives as an observation rather than a forecast. You do not need to guess which scaffolding will become unnecessary. You only need a mechanism that tells you when a guess came in.
A&A perspective
In practice it runs like this. When the condition is met, carry out the action you wrote down before starting: remove or thin the scaffolding on the bet side. Then re-point the same condition at confirming that the path without it has not broken. What disappeared is the scaffolding, not the knowledge the scaffolding produced. What remains is a single line — "if this input does not produce this result, something is wrong" — which is cheaper than the original build and stays usable when the model changes again.
If the condition never fires, suspect the condition before you suspect the model
From the sources
The same post carries a caution about how to read a check that never passes: "With frontier models, a 0% pass rate across many trials (i.e. 0% pass@100) is most often a signal of a broken task, not an incapable agent, and a sign to double-check your task specification and graders." With frontier models, a 0% pass rate across many trials is usually a signal that the task definition is broken rather than that the agent lacks the capability, and a signal to re-examine the task description and the graders.
A&A perspective
Translated into a one-person business, that becomes a concrete inspection step. When the model has gone through two generations and your detection condition has still never been met, the conclusion is not "the gap has not closed yet." Suspect the condition itself first. Has a requirement unrelated to the bet crept into it? If the condition for removing the splitting step also requires mapping the customer's own account codes, then a piece of the asset side has been mixed into a bet's condition, and no amount of model improvement will satisfy it. A mixed condition locks in the loss on that bet permanently.
A&A perspective
The inspection is simple. Re-read the condition and sort every requirement in that sentence into "the model's shortfall" or "the customer's circumstances." If even one requirement is on the customer's side, take that part out of the condition. If what remains is only things that should be satisfied once the model improves, the condition is usable. This inspection is work to finish before you add another bet.
Where this way of deciding does not hold
A&A perspective
First, where the number of bets grows until maintaining detection is heavy. The method works while the sheet fits on one page; once there are ten or twenty conditions, the checking itself is a fixed cost. The moment you start thinking about automating the evaluation, suspect that the cost of detection may have overtaken the losses it prevents.
From the sources
The source is candid about the shape of that cost. On the value of evaluation it says: "Their compounding value is easy to miss given that costs are visible upfront while benefits accumulate later." Costs show up first and benefits accumulate afterwards, so the compounding value is easy to miss. The post makes that point to argue evaluation is underrated, however, not that detection can cost more than it saves; the judgement in the paragraph above is A&A’s, not the source’s. The same post also contrasts teams with and without evaluations: "teams without evals face weeks of testing while competitors with evals can quickly determine the model’s strengths, tune their prompts, and upgrade in days." No population, sample or measurement method is given for that contrast, however. It should be read as a general assertion by Anthropic, not as a forecast of how many days this would take in a reader's business.
A&A perspective
Second, where you want to predict which scaffolding will become unnecessary. This method offers no prediction. Neither source gives a way to tell which parts of an application layer will be erased by model improvement. What can be designed is detection, and nothing here estimates the hit rate of bets or the development cost saved. If the decision needs a forecast number, this method is not enough.
A&A perspective
Third, where every gap is on the customer's side. If all of the scaffolding fills customer-specific data, rules and formats, there are no bets at all, and writing detection conditions is wasted work outright. The correct decision in that case is to build now, and the hesitation is coming from somewhere else.
A&A perspective
Once the bet-and-asset split is in place, the remaining work is turning the build scope into an estimate. Carving out the first unit of work to sell is covered in "Choosing your first paid AI service: how a solo founder carves out one sellable unit of work", and the whole path from winning customers to retaining them in "AI-native GTM: a practical guide for solo founders and small teams". Deciding whether to actually move to a new model once a condition has been met is a separate problem: "When to switch models in a small AI SaaS: customer failures, not benchmarks" covers building representative tasks out of customer failures. This article stops one step before that, at whether to build now at all.
Getting stuck between building now and waiting is not caused by missing information. It is caused by not having, to hand, a condition that says a piece of work has finished its job. Sort the features you are about to build into bets and assets, and for the bets only, write one line before you start — "if this input produces this result, this scaffolding is unnecessary" — along with the action you will take. The day that condition is first met is the day the bet turned into an asset, and the same line moves over to watching whether the capability is still reliably there. No prediction is required. All that is required is a means of noticing when a guess came in. Once the bets outgrow a single page, first suspect that the cost of detection has overtaken the losses it prevents.
Sources & editorial note
Primary pages read for this article. Publication dates below belong to the sources; access dates record our research.
- Lex streamlines the writing process with Claude
Anthropic · no publication date shown on page; page states Company size: Small
Accessed 2026-10-06 - Demystifying evals for AI agents
Anthropic · Published Jan 09, 2026
Accessed 2026-10-06
AI-assisted editorial production
A&A uses AI for research, writing, translation and editorial checks. Source facts, our analysis and hypothetical examples are labeled separately.
Editorial check: 2026-10-06