Skip to main content
Navigation

A&A INSIGHTS

Business & AI strategyFor business owners

What a small team should research with AI before meeting its first customers

AI customer research can only produce a hypothesis about whom to meet and what to ask. Using the Lovable and Lex primary case studies, this article separates what desk research settles from what only contact and actual use settle, and gives small teams a record to prepare first.

customer discoveryfirst customersAI-native GTMsmall teamscustomer research
日本語で読む
Scattered documents become an organized comparison and a decision
A conceptual illustration of gathering information, organizing it, comparing conditions and making a decision. Illustration generated with AI

THE STARTING POINT

For small founding teams that can build but have not settled whom to sell to. In the Lovable and Lex case studies published by Anthropic, the gap between assumed users and the users who actually appeared, and the movement of what needed fixing, became visible through release and observation rather than research. This article separates what AI research can settle from what only meeting people and watching them use the thing can settle, and shows a record that keeps the candidate hypothesis, the questions, the observed behaviour and the offer change apart, using a hypothetical example. Source descriptions, A&A interpretation and the hypothetical example are kept distinct.

Before researching more, decide who you will meet next

A&A perspective

AI-assisted customer research can take you as far as a list of people worth meeting and a hypothesis about the questions worth asking. It cannot tell you whether any of them will pay. So at the stage of finding your first customers there are only two things to decide: the names of the next three people you will meet, and which of their behaviours would break your hypothesis if you observed it. Whether to keep researching or stop is answered by whether you can write those two things down.

From the sources

According to the Lovable case study published by Anthropic, founder Anton Osika built a weekend side project in early 2023 to help developers move faster with AI and posted it to GitHub as an open-source experiment. The case study says that it went viral, but the people showing up to use it were not the developers he had built it for.

Anthropic

A&A perspective

A&A treats that ordering as the important part. The most basic assumption of all, who would use the thing, was overturned not while researching but once it was released and people turned up. It fits practice better to treat research output as an order of visits rather than as a conclusion. Once you have a list of candidates, pick one assumption you want to test, and write down in advance the observation that would break it. If you have also decided whether a break would change the target or the scope, the research you do before meeting anyone can be short.

Lovable: the assumed users and the users who actually appeared were different

From the sources

The same case study reports that Osika kept having the same conversation beforehand: many people came to him saying they would love to build something, that they had a great idea they wanted to turn into reality, but that they could not code. Even so, the early-2023 weekend project was built for developers, and posted publicly to GitHub.

Anthropic

From the sources

On what followed the release, the case study records that the experiment spread widely while the people showing up to use it were not the developers he had built it for. A few days later Osika is described as becoming certain that the world is not going into a future where only engineers look at code to build software, and that it would be a completely new type of interface; that morning he and co-founder Fabian Hedin mapped out what would soon become Lovable.

Anthropic

A&A perspective

What A&A notes is that what was overturned was the target, not the solution. The conversations had been arriving for some time. The first thing actually built and published was still aimed at developers. What people tell you and who turns up at the thing you built and left in public are two separate pieces of information. The longer a small team spends on research, the later it obtains the second one. This account is a customer story published by Anthropic, not an independent verification.

Separating what research settles, what meeting settles, and what actual use settles (A&A's own organisation)
What you want to establishWhat AI research can coverWhat only meeting or actual use reveals
Who uses itCandidate generation from industry, role and public informationWhether the person who actually does the work is the person you assumed
What the difficulty isSituations likely to be occurring, drawn from public documents and job postingsWhether that difficulty ranks high enough to receive this period's budget
How they want it solvedA list of current methods and alternativesWhether the scope you propose is sufficient, and whether the exclusions are acceptable
Who paysA rough sense of the usual price rangeWhose approval releases the payment, and when it moves
Whether use continuesAn estimate of how often work of this kind arisesWhether they return to it unprompted a second time

Lex: the object of improvement moved from interface friction to help with writing judgement

From the sources

According to the Lex case study published by Anthropic, founder Nathan Baschez observed that the tools writers and editors use in their workflow are not nearly as good as those used by programmers and designers. The case study records that the team initially focused on improving the user interface for writers, addressing pain points such as cumbersome track changes.

Anthropic

From the sources

The case study goes on to say that they soon realized the potential of generative AI to enhance the writing process, offering benefits such as instant feedback, idea generation and personalized learning.

Anthropic

From the sources

The company also states that it saw 25,000 new signups within 24 hours of launching its AI-powered features. That is a figure the company itself reports. The case study page does not state the population, the counting method or the comparison basis for it.

Anthropic

A&A perspective

A&A reads Lex's first hypothesis as not wrong but too coarse. Saying that the tools for writing work are poor was a reasonable observation, but that phrasing does not decide whether the thing to fix is the interface or the judgement involved in writing. What you need to establish when meeting first customers is not whether a difficulty exists, but which layer of that difficulty the other party already spends time and money on. Numbers such as signup counts appear after release; they do not come out of the research you do beforehand.

What AI research can settle, and what it cannot

A&A perspective

AI-assisted research can produce a list of candidate companies, the vocabulary used in that industry, the situations likely to be occurring, the questions worth asking, and a falsification condition stating what you would have to see for the hypothesis to break. What it cannot produce is whether that party will pay out of this period's budget, and whose approval releases the payment. The table in this article separates that boundary across five things you might want to establish. The table is A&A's own organisation of the problem, not a claim made by any source.

From the sources

In a July 2013 essay, Y Combinator's Paul Graham writes that the most common unscalable thing founders have to do at the start is to recruit users manually. He also states that you cannot wait for users to come to you, and that you have to go out and get them.

Paul Graham

A&A perspective

That essay addresses early-stage startups and is not an endorsement of unlimited bespoke work. A&A reads it as follows: AI research narrows where you go and reduces travel and preparation time, but it is not a substitute for going. The stopping rule can be equally simple. When you can write down the names of three people to meet, and one assumption per person that could break, the research is finished.

The candidate record: keep hypothesis, questions, observed behaviour and offer changes apart

A&A perspective

Split the record into four columns. First, the candidate hypothesis: who, when, in what situation, and how they currently cope. Second, the questions you want to ask, each annotated with the part of the hypothesis it tests. Third, the behaviour you observed, meaning what the other party actually did, not what they said. Fourth, the change to the offer: which of target, scope or price you changed in response, and how. The reason for keeping them apart is so that you can later trace why the offer changed. Mixed into one column, the other party's words and your own interpretation become indistinguishable within days.

Hypothetical example

The following is a hypothetical example, not a real customer. Suppose a two-person team is trying to sell quotation-drafting support to a building-materials distributor with somewhat over ten employees. The hypothesis is that a sales representative spends more than an hour each time drafting a quotation while searching for comparable past jobs. The question is whether they will show you the three most recent quotations and the order in which each was assembled. The observed behaviour was that two of the three had been produced by back-office staff rather than the sales representative, and that in one case the president had rewritten the unit price at the end. The change to the offer is to move the target from the sales representative to the back-office staff, to take unit-price decisions out of scope for automation, and to leave a final confirmation step for the president. That single meeting reveals both that the buyer and the user differ and which step must not be automated.

A&A perspective

Next week's work can then be written down as: meet three people and record one observation per person that breaks the hypothesis. It is not another round of research. If all three match the hypothesis, suspect that the questions were not shaped to test it before concluding that the target was correct. A testing question is one where a yes and a no lead to different next actions.

Which observations come closest to intent to buy

A&A perspective

Statements are not observations. “That sounds useful” tells you nothing about where this sits in their budget priorities. What A&A treats as an observation is evidence that the other party spent effort: they showed you an internal document, they fixed the next meeting date on the spot, they forwarded it to someone in another department, they already pay for their current alternative, they volunteered which parts should be out of scope. In each case their time or their authority moved. Conversely, being told a document will be sent and then not receiving it is an equally useful observation.

From the sources

As an example of continuing to observe after release, the Lovable case study explains that every new Claude release goes through the evaluation the company has run from the start, measuring how often the system hits a wall and produces an app that is broken or is not what the user asked for. The case study gives as the reason that gate matters that Lovable's users often cannot read the code themselves, so they are trusting the output to work.

Anthropic

A&A perspective

A&A reads this as part of customer understanding too. A mechanism that counts what failed, and how often, after people have used the thing picks up dissatisfaction the other party never puts into words. A small team that prepares this column before its first customer is settled will make faster proposals to the second and third. The more useful thing to measure is not output accuracy in the abstract but the number of times the result differed from what was asked for.

Where this reading does not apply, and the next step

A&A perspective

Both companies are product companies that could release publicly and observe many users at once. A contracting or small service business meets customers one at a time, so the same method will not produce a statistical tendency. As an A&A hypothesis, we suggest replacing the unit of observation with the kinds of exception encountered rather than a count. If the same exception appears twice across three meetings, that is a reason to change the definition of the target. If all three exceptions differ, the target is still too broad.

A&A perspective

The limits of the evidence should be stated explicitly. The Lovable and Lex accounts are customer stories published by Anthropic; both are the companies' own descriptions, not independent audits. Neither company's growth can be treated as an effect of customer interviews. Lex's signup figure is a self-reported number with no stated population or definition and cannot be used as an expectation for another company. Paul Graham's essay dates from 2013 and addresses early-stage startups. The table and the candidate record in this article are provisional instruments made by A&A, not measured methods.

A&A perspective

There are situations where the opposing position is partly right. If the buyer population is small and can be enumerated from a public register, removing candidates by research before meeting them is legitimate. Even then, research does not decide whether the remaining parties will pay. To see the whole workflow first, read “AI-native GTM: a practical guide for solo founders and small teams”; once you have moved on to organising the statements you collected, read “Customer interviews do not become decisions: a Quillit-inspired way to keep evidence usable”. If how to choose a target cannot be settled from your own circumstances, an initial free consultation can start by filling in the first candidate record together.

AI research is material for narrowing whom to meet and preparing what to ask. Intent to buy only becomes observable when the other party's time or authority moves. Keep the candidate hypothesis, the questions, the observed behaviour and the offer change in separate columns, and stop researching once you can write down three names and, for each, the assumption that could break.

Sources & editorial note

Primary pages read for this article. Publication dates below belong to the sources; access dates record our research.

  1. Lovable helps anyone create software 20x faster with Claude

    Anthropic · n.d. (no date shown on page)

    Accessed 2026-09-22
  2. Lex streamlines the writing process with Claude

    Anthropic · n.d. (no date shown on page)

    Accessed 2026-09-22
  3. Do Things that Don't Scale

    Paul Graham · 2013-07

    Accessed 2026-09-22

AI-assisted editorial production

A&A uses AI for research, writing, translation and editorial checks. Source facts, our analysis and hypothetical examples are labeled separately.

Editorial check: 2026-09-22

All articles