A&A INSIGHTS
Usage-Based Billing: Who Fixes a Missed or Double-Counted Event?
When a usage-based invoice is wrong, who investigates and who repairs it? Missed events, double counts and late arrivals need three separate rows in your estimate.
日本語で読む
THE STARTING POINT
Do not accept missed events, double counts and post-cut-off arrivals under a single clause. A billing platform reports only invalid events, and an event you never sent or sent twice looks perfectly normal from its side, so apart from reconciling against a count you keep yourself there is no way to notice either one.
The answer: the three failures differ in who can even notice them
A&A perspective
Do not accept missed events, double-counted events and events arriving after the billing cut-off under a single clause that says you will handle billing discrepancies. The billing platform reports only invalid events. An event you never sent and an event you sent twice both look perfectly normal from the platform's side. So apart from reconciling against a count you keep yourself, there is no way to notice either of them.
A&A perspective
When the party capable of noticing changes, the place where repair responsibility sits has to change with it. A failure the platform flags as invalid is finished once you build a handler that receives the notification and resends. A failure the platform says nothing about leaves your client's complaint as the only detector, unless you build something that keeps counting. And a failure whose detector is the client's complaint always comes back as a question of trust.
A&A perspective
That is why the estimate needs three rows, not one. For each kind of failure, write down how it is detected, by when it must be corrected, and whose hours the correction is performed on. What follows is the content of those three rows, and how to recognise the case where you do not need three.
Lateness is not a defect; it is written down as the specification
From the sources
Stripe's documentation on recording usage states plainly that aggregation is asynchronous: "Stripe processes meter events asynchronously, so aggregated usage in meter event summaries and on upcoming invoices might not immediately reflect recently received meter events." The delay is documented behaviour, not an incident.
From the sources
The same page treats the recording cadence as the implementer's decision: "You can decide how often you record usage in Stripe, for example as it occurs or in batches." Whether you send each event as it happens or send them in batches is therefore not a constraint the platform imposes on you. It is a design choice you make.
A&A perspective
Put those two statements side by side and it follows that an invoiced amount is only ever the amount that has arrived by that moment. When your client says that today's usage is not on the invoice, that is not a bug; it is the behaviour the documentation describes. But the documentation saying so and your client understanding it are two different things. Producing that understanding is the work of explaining the specification.
A&A perspective
The thing to be careful about here is that explaining lateness as a specification does not, on its own, solve anything. Once you have said that usage arrives late, you still need a line: how much lateness is treated as normal, and at what point it becomes something to be corrected. Splitting the failures into three is how you draw that line.
| Kind of failure | Signal from the platform | Detection you supply, and where the correction sits |
|---|---|---|
| Missed (never sent) | None. The platform does not know how many events should have arrived. | Compare, daily, the total your application intended to send against the platform's aggregated usage. Inside 35 calendar days, resend through the same path to correct. |
| Double-counted (sent twice) | None. A duplicate is not invalid, so it enters the aggregate as a valid event. | No detection after the fact. The only control is deciding the idempotency key before you implement. Route any excess that offsetting cannot absorb to an invoice-side correction. |
| Late (arrived after the cut-off) | Anything beyond 35 calendar days is reported as invalid. Lateness inside the window has no signal. | Catch in-window lateness with the reconciliation. Write the deadline as a date: the shorter of your client's book-amendment cut-off and the 35 calendar days. |
The platform reports only invalid events
From the sources
The same documentation explains that the platform emits its own events when something is wrong with a meter event. One is described as "This event occurs when a meter has invalid usage events." The other is described as "This event occurs when usage events have missing or invalid meter IDs." So a notification path does exist.
From the sources
The page also lists the reasons an event is judged invalid. A timestamp too far in the past, a timestamp in the future, a customer that cannot be found, a missing value and an invalid value are each listed separately. On the page these appear as underscore-joined identifiers, so only their meaning is rendered here.
From the sources
And the documentation assigns the follow-up work explicitly to the implementer: "Correct and resend invalid events for re-processing." The division of labour is settled right there — whoever receives the notification is the one who repairs the data.
A&A perspective
This is the starting point for the whole article. For events that arrived but were invalid, the platform tells you, and it even states that repairing them is the implementer's job. Turn that around and you have the limit of what the platform can tell you. About an event that never arrived, the platform has nothing to report from. That last step is not something the documentation states; it is A&A's reading of what the documentation does not cover.
Double counting can only be prevented at design time
From the sources
The documentation puts duplicate prevention on the implementer: "Use idempotency keys to prevent reporting usage for each event more than one time because of latency or other issues." Retries caused by latency are named as the reason you would need the key at all.
From the sources
It then describes what happens if you skip the key: "If you don’t specify an identifier, we auto-generate one for you." In other words, not designing an idempotency key is a permitted choice. Nothing refuses your integration for omitting it.
A&A perspective
The documentation states that "Every meter event corresponds to an identifier that you can specify in your request." and then instructs you to use idempotency keys to prevent duplicates. A&A reads that pairing in the obvious way: if leaving the identifier to auto-generation prevented duplicates, the instruction would be unnecessary. So when you do leave it, reporting the same real-world usage twice gives the platform two distinct events. Nothing about them is invalid, which means the notification from the previous section never fires. The duplicate enters the aggregate as an ordinary, valid event.
A&A perspective
That makes double counting a single decision at design time rather than a detection problem. Decide before you implement what counts as the same piece of usage — which business-level unit becomes the key — and it does not happen. Skip that decision and the platform side can no longer notice it afterwards. So the double-counting row of your estimate should not say that you will investigate. It should state the key definition itself.
A&A perspective
This is also a place where you have to agree on vocabulary with the client. What "the same operation" means is settled in the language of their business, not yours. Whether a retry of the same request counts as one or two is decided by how you construct the key, so do not let implementation convenience decide it. Match the way your client counts.
A missed event produces no external signal
A&A perspective
A missed measurement is a failure in which your code never called the platform. The sender was down, an exception skipped it, the retries ran out, the queue backed up. All of those happen outside the platform, so the platform has nothing to report from.
From the sources
Given that the documentation leaves the recording cadence to the implementer — "You can decide how often you record usage in Stripe, for example as it occurs or in batches." — the platform does not know how many events should have arrived. A party with no expected count cannot report a shortfall against it.
A&A perspective
So the only detector available for a missed measurement is a second count that you hold yourself. Keep a running total on your application side of what should have been sent, and compare it daily against the platform's aggregate. The platform does hold one — the documentation refers to "aggregated usage in meter event summaries" — so what is missing is not the platform's number but something to compare it against. That comparison is partly an implementation task and partly an ongoing operational one.
A&A perspective
If that operational half is not in the estimate, it becomes unpaid work. There are two honest ways to put it in. Either include the effort to build the comparison in the build fee and carve out the daily checking and discrepancy investigation as a separately paid operations scope, or hand the comparison view over as the client's own routine and accept only the investigation on a time-and-materials basis when a discrepancy appears. Both are defensible. Not deciding is not.
A&A perspective
It is also worth writing down why the client's complaint must not be the detector. The client notices when they look at the invoice, and that is after the close. A missed measurement discovered after the close runs straight into the correction limits in the next section.
Late arrival is arithmetic, not judgement
From the sources
The documentation sets a valid range for timestamps: "Make sure the timestamp is within the past 35 calendar days and isn’t more than 5 minutes in the future." It explains the forward allowance as well: "The 5-minute window is for clock drift between your server and Stripe systems."
From the sources
Events outside that range fall into the invalid list described earlier. Among the error reasons the page lists, a timestamp too far in the past and a timestamp in the future appear as two separate cases.
A&A perspective
So late arrival splits in two. Lateness that still falls inside the 35 calendar days is accepted with its original timestamp. Lateness beyond 35 calendar days cannot be sent through this path at all and is reported as invalid. Correction works for the first case. For the second, the correction route itself does not exist.
A&A perspective
What goes into the estimate is therefore not a judgement rule about how late is too late. It is date arithmetic. If the client closes at month end, until what day of the following month will you accept corrections? Does that deadline sit inside the 35 calendar days? If it does not, the thing to redesign first is the route, not the deadline.
A&A perspective
This is where you have to mesh with your client's accounting close. The deadline by which you can technically send data and the deadline by which they can still amend their books are different numbers, and the shorter one is the real deadline. Writing "we can correct anything within 35 days" without checking which is shorter is how you make a promise you cannot keep.
Correcting an over-count and correcting an under-count are different jobs

From the sources
The documentation is explicit about negative usage: "If the overall cycle usage is negative, Stripe reports the invoice line item usage quantity as 0." It also notes that the value itself can be fractional: "The numerical usage value in the payload accepts decimal values."
A&A perspective
Where that sentence bites is at the edge of what offsetting can absorb. Sending offsetting events repairs an over-count only as far as the cycle total stays positive. Push past that, so the overall cycle usage goes negative, and the reported invoice quantity becomes 0. Zero is not a refund.
A&A perspective
So an over-count and an under-count must not live in the same clause. An under-count, inside the 35 calendar days, is repaired by resending through the same usage path. An over-count reverses through that path only as far as it can be offset within the same cycle; anything beyond that, and anything in a cycle already invoiced, does not come back. So an over-count needs a second route arranged in advance — a correction on the invoice side. That is a different system, and usually a different person.
A&A perspective
The effect on the estimate is concrete. Correcting an under-count can be completed inside your own code and your own route, so it fits within the build and operations scope. Correcting an over-count that offsetting cannot absorb reaches into the client's billing operation and their accounting, so what you can accept is identifying which events were excessive and handing that list over. Writing "we will fix billing discrepancies" without drawing that line means you have accepted their bookkeeping too.
A&A perspective
To be clear, handling the over-count on the invoice side is A&A's proposal. The source only describes the behaviour — that when overall cycle usage is negative the invoice line item quantity is reported as 0 — and says nothing at all about how a correction should then be made.
What the published case shows is that defining it is the fast part
From the sources
Stripe's published page about Lovable states that, for the launch of Lovable Cloud and Lovable AI, the company used usage-based billing features and within two weeks got as far as "to define meters and rate cards" and on to "automatically bill customers for actual consumption". The statistics panel on the page reads "2 weeks to implement usage based billing".
From the sources
The same page also indicates scale, stating "4.6 million credits being granted every month". This is Stripe describing its own customer on its own marketing page. It is vendor testimony, not an independent audit.
A&A perspective
You cannot use that as a basis for estimating contract work. What you can read from it is which part was fast: the defining, and getting automatic billing switched on. Settle the meters, settle the rate cards, make invoices come out in proportion to consumption. That is the part the platform carries for you.
A&A perspective
The allocation of the three failures this article has been working through does not appear on that case page. Its absence is not a flaw in the page. The point is that the part the platform makes fast and the part nobody decides unless the contractor decides it are two different parts — and the second is the one that gets cut from an estimate.
A&A perspective
This is exactly where Lovable's position, building its own service, parts company with a contractor implementing inside somebody else's service. In-house, investigating a billing discrepancy is your own work and the cost comes out of the same pocket as the revenue. Under contract, unless you decide in advance whose hours that investigation runs on, it shows up after handover for free.
What goes in the three rows you add to the estimate
A&A perspective
All three rows carry the same three items: how the failure is detected, by when it must be corrected, and whose hours the correction runs on. Keeping the items aligned also means your client can read the rows comparatively.
A&A perspective
The missed-measurement row. Detection is the reconciliation you build: compare, daily, the total your application intended to send against the platform's aggregated usage. The deadline is whichever is shorter, your client's book-amendment cut-off or the 35 calendar days. Ownership splits as building the reconciliation into the build fee, and the daily check plus discrepancy investigation into the operations fee.
A&A perspective
The double-counting row. Here you write a definition, not a detection method: the key that decides what counts as the same piece of usage, and the business reason behind it. State alongside it that no platform-side detection exists after the fact. Ownership is the key design, in the build fee. If the definition later has to change, treat that as a specification change and price it separately.
A&A perspective
The late-arrival row. Detection has two routes: anything beyond 35 calendar days comes back as an invalid-event notification from the platform, and anything late inside that window is caught by the reconciliation. Write the deadline as a date, not as a judgement rule. Ownership is the receive-and-resend handler in the build fee, while the treatment of anything past the deadline is your client's accounting decision.
A&A perspective
Then add one more line underneath the three: the over-count line. Because an over-count does not reverse through the usage path, your scope ends at identifying the excessive events and handing that list over, and the invoice-side correction stays with your client. Without that line, the three rows above will be read as a promise to fix everything.
Hypothetical example: splitting "bill them for what they use" into three rows
Hypothetical example
What follows is a hypothetical design example. It is not an A&A engagement, and it is not evidence of any rate of occurrence or any outcome. Suppose the client sells automatic image correction as a service, and you have accepted the build for a system that bills one unit per image put through correction.
Hypothetical example
Start with the double-counting row. Define one piece of usage as one correction-job ID, and use that ID as the idempotency key. What you then have to confirm with the client is how a user's retry of the same image is counted. If the client says retries are not billable, the key stops being the job ID and becomes the pair of image ID and version. Skip that confirmation and implementation convenience decides the counting rule for them.
Hypothetical example
Fill the missed-measurement row with the reconciliation. Keep one row per completed correction in your own database, and each morning compare that day's total against the platform's aggregated usage. When they differ, resend the rows that carry no sent marker. Write that morning comparison and resend out explicitly as something the operations fee covers.
Hypothetical example
Fill the late-arrival row with dates. If the client closes at month end and accepts book amendments until the fifth of the following month, your correction deadline is the fifth of the following month. The 35 calendar days is the longer of the two, so the real deadline is set by their accounting. Agree with their finance side in advance whether anything found after the fifth is treated as the following month's usage or handled on the invoice.
Hypothetical example
Finally, the over-count line. If changing the retry counting rule later reveals that past periods were over-counted, you accept the work as far as producing the list of job IDs that were excessive. Whether a credit is issued against that list or the adjustment is made on the next invoice stays with the client's billing operation.
A&A perspective
What matters in this example is that two of the three rows are filled in by asking, not by building. The double-counting key and the correction cut-off both require the client to decide — and once asked, both are settled. Work that finishes in a single pre-estimate conversation becomes unpaid post-handover investigation if it is left undecided.
Where this does not apply, and what is not being claimed
A&A perspective
Start with the counterexample. When the metered unit is created inside the system you are building, and is already persisted transactionally, these three rows are over-engineering.
A&A perspective
If every completed job leaves one row in a database you control, then the missed-measurement risk shrinks to a single question: did the sender process run? One daily total comparison is enough. Designing and maintaining the three-row arrangement would cost more than the problems it prevents. In that case the honest thing to write is that you will add one daily total comparison, and nothing more.
A&A perspective
The three rows earn their cost when the event the measurement rests on happens outside your control. The client's existing system is the trigger; you are counting webhooks from a third party; you are counting actions on the user's own device. In those cases the truth cannot be reconstructed afterwards, so you need something that captures it at the moment it happens.
From the sources
Next, the conditions in the source. The 35 calendar days, the 5 minutes into the future, and the per-customer combination ceiling — the documentation states "For each customer on a meter, Stripe accepts up to 100 unique combinations across all events." and, about events arriving after a limit is reached, "Events that introduce a new dimension combination after either limit is reached are invalid" — are all as printed on the page when it was read on 3 October 2026.
From the sources
The same documentation also carries a comparison between basic usage-based billing and Metronome, and describes a separate API v2 meter event stream path for sending up to 10,000 events per second. Which route you take changes the constraints. Re-read the current page immediately before you implement.
A&A perspective
Now what is not being claimed. The three-row allocation, the placement of detection and the proposal to handle over-counting on the invoice side are all A&A's design. Neither source says anything about contracts, estimates, acceptance testing, correction deadlines, or who bears the cost of investigation time.
A&A perspective
Correction deadlines and liability depend on your client's accounting process and on the contract your client holds with their own users. Whether a refund or a re-issued invoice is legally available is not something A&A determines; it is something the reader confirms with their own adviser. Lovable's two weeks is a statement about that company's own setup and is not a basis for estimating contract effort. A&A gives no guidance on elapsed days.
A&A perspective
No A&A figures are offered for how often billing discrepancies occur, how many investigations they required or how many corrections were made. No price level is recommended. The numbers in the example above belong to a hypothetical scenario.
Next step: add the three rows to one estimate you are working on now
A&A perspective
Open one usage-based billing engagement you have in front of you and add three rows to the estimate. In each row, write the detection method, the correction deadline and the owner. Any cell you cannot fill is an item you have not yet asked your client about.
A&A perspective
Of those unfillable cells, the double-counting key definition and the correction cut-off are both settled in a single conversation with the client. The earlier you fill them, the cheaper they are; filled after handover, they are expensive.
A&A perspective
We take on service builds that include metering design as a development enquiry. For adjacent decisions, "Outcome pricing for AI agents: how do you count one completion?" covers the definition of counting itself, "How to explain credits in an AI SaaS plan: show the usable amount and the extra spend" covers disclosure to the end customer, and "After an AI workflow launches, who restores the interrupted work?" covers the post-handover operational split. The overall picture is in "AI-native GTM: a practical guide for solo founders and small teams".
A billing discrepancy is three separate failures. Separate the ones the platform reports as invalid from the ones only your own running count can catch, write detection, correction deadline and owner into three distinct rows, and quote from those. One asymmetry stays: beyond what offsetting absorbs inside the same cycle, an over-count does not reverse through your usage path, so arrange the invoice-side route in advance.
Sources & editorial note
Primary pages read for this article. Publication dates below belong to the sources; access dates record our research.
- Record usage for billing with the API
Stripe · undated documentation page
Accessed 2026-10-03 - Riding the AI boom: How Lovable grew into a vibe-coding juggernaut with Stripe
Stripe · undated customer story
Accessed 2026-10-03
AI-assisted editorial production
A&A uses AI for research, writing, translation and editorial checks. Source facts, our analysis and hypothetical examples are labeled separately.
Editorial check: 2026-10-03