Skip to main content
Navigation

A&A INSIGHTS

Business & AI strategyFor business owners

Model selector or fixed default: decide before you ship the setting

Decide before launch whether your AI service exposes model settings to customers or fixes a default from your own evaluation, using the Lindy and Genspark cases.

AI servicesmodel selectionoffer scopeevaluationsolo founder
日本語で読む
Scattered documents become an organized comparison and a decision
A conceptual illustration of gathering information, organizing it, comparing conditions and making a decision. Illustration generated with AI

THE STARTING POINT

Whether you expose the model or detailed settings to customers is not decided by how rich the feature set is. It is decided by whether you can say which option is better from your own evaluation. If you can, fix it as the default; if you cannot, build the evaluation before shipping the choice. Putting options on a settings screen is not the provision of freedom — it hands the evaluation work to the customer, and only the decision transfers. The consequence of output that does not fit the work comes back to you.

Whether to expose the setting depends on whether you can judge it from your own evaluation

A&A perspective

Whether you expose the model or the detailed settings to the customer is not decided by how rich the feature set is. It is decided by one thing: can you say which option is better from your own evaluation? If you can, fix it as the default. If you cannot, build the evaluation before you ship the choice. That is the order, because putting options on a settings screen is not the provision of freedom; it hands the work of evaluating to the customer.

A&A perspective

Evaluation here does not mean a position in a public benchmark. It means a list of somewhere between a dozen and a few dozen jobs copied from the work you actually sell, each with one sentence saying what counts as a pass. One question is enough to decide: can you explain to a customer, in a single paragraph, why this setting is the default right now? If you cannot, that setting is not yet in a state you can expose.

Shipping a choice hands the evaluation work to the customer

A five-row, three-column comparison table headed "The choice moves. The consequence does not". The columns are the aspect, fixing a default, and shipping a choice. Row one, "Who says which is better": fixing a default puts it with the provider's own evals; shipping a choice puts it with the customer, who has no comparison. Row two, "Output that does not fit": both columns read "comes back to you". That row is the point of the figure, showing the asymmetry that shipping a choice moves the decision to the customer while the consequence of output that does not fit the work does not move and returns to the provider. Row three, "What you must reproduce": one default path, against a combination per customer. Row four, "What you publish": the default and its date, against a list of options. Row five, "When it breaks down": when you cannot sustain evals, against when procurement names the model. That Lindy chose its default model after extensive evaluation, and that customers can select alternative models but rarely override it, and that Genspark's Super Agent is described as model-agnostic by design with which model takes which job presented as part of the product's design, all come from the company case studies published by Anthropic. The choice of aspects, the framing of the consequence as the thing that does not move, and the contrast of what is published and when each path breaks down are A&A's design proposal. This figure shows no effect measured by A&A, and neither page states that offering a choice is harmful.

A&A perspective

The moment you put model A and model B on a settings screen, you have handed the customer a question: which of these is better for your work? Answering it requires running both against their own jobs, comparing the output, and deciding what passes. Customers do not do that work. They pick by name recognition, by price, or by the order the options appear on the screen.

A&A perspective

Only the decision transfers; the consequence comes back. When the output of the model a customer picked does not fit their work, what arrives is not a report that says this model is inaccurate. It is a ticket that says this service does not work. Having let them choose is not a defence. So what shipping the choice reduces is only your burden of deciding; the duty to explain quality and the work of investigating both remain — and they grow, because behaviour now differs per customer.

Hypothetical example

Take a hypothetical. You run a service, alone, that summarises inbound email and routes it to the right person. You have ten customers, and you put three models on the settings screen. When someone asks you to reproduce a problem, the combinations you might have to check come to thirty. That is not a number one person can hold. Fix a single default and the set drops back to ten, and the quality explanation only has to be written once to serve every customer. This is an illustrative count showing how the arithmetic works, not a measured figure.

The customer's questions, re-sorted by what you need in order to answer them
The customer's questionWhat you need in order to answerWhat happens if you ship a choice instead
"Which model do you use?"The default model's name, and as of whenThe answer is unchanged. Disclosure is enough
"Can we choose it on our side?"A pass/fail comparison on your representative requestsCustomers with no comparison pick by name or screen order
"Wouldn't another model be more accurate?"Results compared under your work's conditions, with those conditions statedBehaviour diverges per customer; reproduction paths multiply
"Our procurement specifies a model."Pass/fail on the specified model, and the boundary you cannot meetA requirements problem. An evaluation cannot substitute for it
"Will you switch when a new model ships?"Whether the update criteria are written down in advanceSwitch timing diverges customer by customer

Source: Lindy lets customers switch, yet decides the default by its own evaluation

From the sources

Lindy sells an automation platform on which businesses build AI agents to take over repetitive work. The company case study published by Anthropic describes how the default model was chosen: after extensive evaluation, Lindy chose Claude as their default model, and while customers can select alternative LLMs, they rarely override Claude. The order is that the design allows switching, and the default is still decided by the company's own evaluation. The page records the company size as Small.

Anthropic ↗

From the sources

The same page says where the grounds sit. Quoting Luiz Scheidegger, Head of Engineering: Claude models simply performed better in our evals and use cases. On how often the default is actually replaced, the page states almost no one overrides the default LLM, and then that the vast majority of all model calls on the platform are for Claude.

Anthropic ↗

A&A perspective

What this case shows is not whether a choice exists but whose evaluation decided the default. Having a switchable design is not what makes the product work. The order runs the other way: because the evaluation is held in-house, the default can be named; because the default is good enough, customers do not override it. Note also that the page never explains why the selector was kept. On how rarely the default is overridden, no population, time window or measurement method appears on the page either, so we read it as a stated tendency rather than a measured rate. Whether the selector exists for procurement reasons or for migration is not stated, so that stays blank here rather than guessed.

Anthropic ↗

The evaluation axis is the work you sell, not general intelligence

From the sources

The same page is specific about what the evaluation measured. Quoting Scheidegger, the models were able to better navigate ambiguity in large context windows, and to make better tool calls when dealing with complex workflows such as calendar management, natural language conditions. Calendar handling and natural-language conditions are the areas the page calls two areas where other models struggled. It also mentions handling deeply nested data structures.

Anthropic ↗

A&A perspective

What is listed there is not a measure of general intelligence; it is the set of conditions in the work that company sells. That is the part that matters for a solo or very small company. Building an evaluation sounds like building an exhaustive test suite, but what you need is a narrow list copied from your main kinds of request. Put the other way round: until you can write down the conditions of your own work, no evaluation can be built. The first thing you write is not a test — it is an inventory of the work you sell.

A&A perspective

Once the evaluation exists, the procedure for deciding whether to switch when a new model ships is treated in a separate article: "When to switch models in a small AI SaaS: customer failures, not benchmarks". What this article handles is the decision one step earlier — whether you may expose a choice before you hold that evaluation at all.

Source: Genspark's "model-agnostic" is not a customer-facing choice

From the sources

The second case is Genspark. Describing the company's Super Agent, the page says it is model-agnostic by design, with different frontier models doing different jobs. What matters is who does the assigning. As the page describes it, deciding which tool to call next at each step, when to gather more information, when to backtrack and when to summarise and return is the model's role. Which model takes which job is described on the page as part of the product's design.

Anthropic ↗

From the sources

As a concrete instance of that assignment, for the most demanding planning problems the company fans out to multiple frontier models in parallel and uses Claude to reconcile their plans into one. Several models are queried in parallel and a model is then given the job of consolidating the results. The page records the company's scale as roughly 50 engineers and more than 150 tools orchestrated.

Anthropic ↗

A&A perspective

The phrase model-agnostic is used here to mean that the provider can assign per job, not that the customer can choose. The distinction is not wordplay. The first is a design that hands evaluation to the customer; the second is a design that keeps it in-house. When you write "because we are model-agnostic" as the reason for putting options on your own settings screen, that phrase is pointing at something different from what the source means by it.

Evaluation is recurring work, not a one-time task

From the sources

The same page also shows that evaluation does not finish once. Kay Zhu, co-founder and CTO, is described as having tested the approach in which the model decides each next step against every new frontier model that came out, roughly every three months, for two years. The failures are recorded too: the model fell into infinite loops, hit the same error repeatedly, or failed to recognize when it had enough information to stop.

Anthropic ↗

From the sources

The architecture before that was a directed graph of predefined workflow steps. Zhu's assessment of it was that it was too rigid, and that it often broke on edge cases. The page states that when the same approach was tried again in early 2025 against the newest model of the time, the outcome changed. It is a record of the answer to one design proposal changing with the model generation.

Anthropic ↗

A&A perspective

What you should be costing from this is not one round of building an evaluation but the number of times you rebuild it. Deciding to fix a default includes a promise to run the same evaluation every time a model is updated. So the conclusion of this article is not "always fix the default". If you cannot sustain the evaluation, saying so is the more accurate position. But what you ship in that case is still not a choice. You fix one default and publish how far your evaluation reaches, because a choice is not a substitute for not having an evaluation. We give no figure for the time or money an evaluation costs: neither of the two pages cited states one.

What to publish in place of the setting once the default is fixed

A&A perspective

When you fix the default, what you publish in place of the setting is three lines. First line: the model you run as the default, and as of when that judgement was made. Second line: on what work you judged pass and fail — not an exhaustive test suite, but the kinds of request you mainly take on. Third line: what you can and cannot do if the customer needs to specify a model. Those three lines are shorter than a settings screen, and they answer the customer's question first.

Hypothetical example

Here is template wording, shown as a shape a reader's own company fills in with its own facts. "As of October 2026 this service fixes a specific model as the default for the summarising and routing steps. The selection was made by comparing pass and fail on a check list built from the inbound-summarisation requests we have actually taken on. If you need to specify a model on your side, we will quote the verification on that model separately. Where a requested model cannot be supported, we will state that boundary up front." This is a template — not A&A's terms of service, and not a result we measured.

Three cases where exposing a choice is still right

A&A perspective

There are cases where shipping the choice is the right call. They narrow to three. First, when procurement requirements or a contract demand a named model or serving region. That is a requirements problem, not an evaluation problem, and it is not yours to settle on your own terms. Second, when the customer's data-handling policy differs by provider. What you are letting them choose there is the data path, not model accuracy. Third, when you genuinely have two defaults for two kinds of work and can state the routing rule in your own words. That is not a choice; it is disclosure of an assignment.

A&A perspective

On the third case: the design question of which model is assigned to which step is outside this article. That is an internal design matter, separate from whether anything is exposed to the customer. All that can be said here is that if you cannot explain the routing rule in your own words, you are not at the stage of putting that assignment on a settings screen.

When this rule does not apply

A&A perspective

The limits first. Both pages cited are self-description by the subject company and the provider, not independent audits. In particular, the statement that almost no one overrides the default and the statement that most model calls go to Claude are internal observations whose population, period and measurement method are not on the page. The growth and reduction figures on the Lindy page, and the revenue figure on the Genspark page, are not used here either, because no definition or population is given. Neither page says that offering a choice is harmful, nor that fixing a default caused the growth. We do not claim that causation either.

Anthropic ↗Anthropic ↗

A&A perspective

The difference in scale is worth keeping as well. Genspark orchestrates more than 150 tools with roughly 50 engineers and, for hard planning problems, queries several models in parallel and consolidates the results. That operation, and the cost of the evaluation behind it, does not transfer to a one-person company as it stands. What transfers is only the direction of the design: the provider decides the assignment.

Anthropic ↗

A&A perspective

Finally, what this article does not measure. We have not measured the effect of fixing a default on churn, win rate or ticket volume. This is a claim about ordering, not a prediction of outcomes. And in a market where deals frequently arrive with a procurement-specified model, this decision leans towards the requirements side from the start. In that case, counting how often model specification actually appears in your own deals pays off before any of this does.

Whether you expose the setting to customers is not a question of how much product to build. It is one question: can you say which option is better from your own evaluation? If you can, fix it as the default and publish the reason in three lines. If you cannot, build the evaluation before shipping the choice. What the Lindy case showed was not the value of a switchable design but the order in which the default was decided by the company's own evaluation. Genspark's model-agnostic likewise pointed at the provider's assignment, not the customer's freedom. Options laid out without an evaluation behind them only hand the decision over; the consequence still comes back to you.

Sources & editorial note

Primary pages read for this article. Publication dates below belong to the sources; access dates record our research.

  1. Lindy empowers teams to scale with AI Agents powered by Claude

    Anthropic · no date shown on page

    Accessed 2026-10-05
  2. Genspark's Super Agent orchestrates 150+ tools with Claude

    Anthropic · no date shown on page

    Accessed 2026-10-05

AI-assisted editorial production

A&A uses AI for research, writing, translation and editorial checks. Source facts, our analysis and hypothetical examples are labeled separately.

Editorial check: 2026-10-05

← All articles