How do we meter AI usage without charging customers for our failed retries?

Build an AI usage ledger that separates provider costs from customer outcomes, prevents duplicate retry charges and supports invoice checks and recovery.

Saas Development
6 October 2026Updated 06 Oct 202610 min readBukhosi Moyo

Quick Answer

For a plan that charges per delivered AI outcome, keep provider attempts and customer billable events separate. Create one durable, unique billable outcome only after validation, persistence and delivery. Reconcile uncertain retries and billing submissions; a failed customer job may still incur provider costs.

Key Takeaways

  • Keep provider consumption and customer invoice units in separate linked records.
  • Define usable, persisted and delivered before counting an outcome as billable.
  • Use durable action IDs and atomic uniqueness to prevent two workers billing the same action.
  • Reconcile unknown results and lost billing acknowledgements before replaying submissions.
  • Use Batch only for deferred work; it does not guarantee immediate replies or free cancellation.

Want the full breakdown? Scroll below.

People reviewing work together at a desk with laptops
On this pageJump to a section
  1. 1Charge for the customer outcome you promised
  2. 2Separate provider attempts from billable outcomes
  3. 3Define a precise delivery boundary
  4. 4Practical output: a defensible AI usage billing ledger
  5. 5Create one outcome despite retries and crashes
  6. 6Choose the billing integration before naming its API
  7. 7Batch processing fits deferred work
  8. 8Worked hypothetical case: deferred summaries
  9. 9Acceptance, failure and recovery tests
  10. 10Practical worksheet for the pricing decision
  11. 11Frequently asked questions
  12. 12Sources

Share this article

Bukhosi Moyo

Growth Partner

Need help growing your company?

We build SEO-first websites and growth systems for South African businesses.

Get Started

Charge for the customer outcome you promised

For a plan that charges per delivered AI outcome, an internal retry should not create a second customer unit. Define what the customer bought before choosing a meter. A completed document summary, an accepted classification and a delivered chat reply are different outcomes, even if each uses the same model API.

The policy in this article is proposed outcome billing: one eligible delivered result per customer action. Other plans can use other disclosed units. The important boundary is that the provider's consumption record measures your costs; it does not automatically define the customer's invoice. Your business may pay for several attempts while charging for one result, or pay for work that never produces a billable outcome.

A boolean named billable is useful for reporting, but it is not enough to enforce the promise. Two workers can both observe an empty record, finish the same job and set their own rows to true. The ledger needs stable identities, atomic uniqueness and recovery rules that survive delayed results and process restarts.

Separate provider attempts from billable outcomes

Create a stable request_id for the customer action before starting work. A retry keeps that action ID and receives its own attempt_id. A customer intentionally requesting fresh work receives a new action under the agreed product rules. Never generate a new action ID merely to recover from a timeout.

Record the tenant or workspace, provider request identifier, attempt timestamps, status and any reported token usage. A timeout is an unknown result, not proof that the provider performed no work. Store unknown usage as unknown rather than inventing a zero-token failure. Reconcile the result before deciding whether another attempt is appropriate.

Keep a separate outcome record with the charging unit and the policy version that applied when the customer requested the work. For this proposed plan, a model success becomes eligible only after product validation, durable persistence and delivery to the entitled customer. A response that failed validation or was never made available contributes no customer unit under this policy.

Define a precise delivery boundary

Write down what makes the result usable. For a document summary, that might mean it covers the submitted document, passes required structural checks and is stored in the customer's workspace. For a classification service, an unsupported or invalid label may fail the product contract even if the model returned HTTP success.

Define delivery just as clearly. Persisting an output and making it available in the authenticated account is one possible boundary; an email receipt is another. Do not claim confirmed customer delivery solely because the provider returned text. Decide how customer cancellation, late completion and results arriving after a deadline affect eligibility, and present those rules before accepting the work.

A proposed retry deadline can limit latency and internal spend. It must not become the deduplication window: a result arriving after five minutes should not acquire a fresh charging identity. Retain the stable action key throughout reconciliation, correction and the invoice dispute period defined by your policy.

Practical output: a defensible AI usage billing ledger

Use linked attempt and outcome records. The fictional rows below demonstrate the distinction; they are not measured usage or provider pricing examples.

Record Customer action Attempt/outcome ID Result Provider consumption Customer units
Attempt req-1001 attempt-1 Client timeout; result unknown Unknown until reconciled 0
Attempt req-1001 attempt-2 Response returned Reported usage if available 0
Outcome req-1001 outcome-1001 Validated, persisted and delivered Links to both attempts 1
Attempt req-1002 attempt-1 Error; no delivered output Verified usage or unknown 0

Include these fields in the worksheet for your actual integration:

Field Purpose Verification or recovery rule
tenant_id, request_id Identify the customer action Derive workspace scope from trusted account context
attempt_id, provider_request_id Trace internal attempts Preserve unknown results for reconciliation
outcome_id, unit_type Identify the chargeable product result Enforce one eligible unit per action and unit type for this plan
policy_version Explain the rule applied Preserve the accepted version rather than silently replacing it
occurred_at, delivered_at Establish timing Store the actual outcome time and delivery evidence
billing_event_id, submission status Trace the external meter submission Keep a stable identity across uncertain responses
Correction reference and reason Explain an adjustment Link the adjustment to the original outcome and invoice

The internal cost report sums available provider consumption. The customer invoice report sums eligible outcome units. These totals need not match. Show customers the result, unit and charging rule relevant to their workspace; do not expose other customers' records, raw credentials or sensitive prompt contents merely to make the ledger look transparent.

Create one outcome despite retries and crashes

Persist the result, create the unique billable outcome and enqueue its billing-submission job in a transaction where those records share a database. A database uniqueness constraint on the stable action and unit type prevents concurrent workers creating two units. If multiple outputs arrive, preserve their attempt evidence but select only the eligible outcome allowed by the product policy.

Use a durable submission job or outbox for the external billing operation. The external meter is not part of your database transaction. A worker can crash after the provider accepted an event but before your application recorded the acknowledgement. Reconcile the original identifier, or retry using the chosen provider's documented idempotency mechanism, before treating that submission as new.

Track pending, accepted, aggregated and invoiced observations separately. Avoid silently deleting a historical charge when correcting it. An adjustment linked to the original outcome explains the change to support, finance and the customer. A failed submission also should not cause the product to rerun the AI job; result production and billing delivery are separate recovery tasks.

Choose the billing integration before naming its API

Stripe's usage-based overview presents Metronome for new use cases and documents Billing Meters for existing integrations. Verify the selected product, merchant eligibility and its event contract before building an adapter. A Billing Meters API example is not a Metronome API contract, and the overview does not establish that this particular South African merchant can use either product. Stripe usage-based billing

The ledger remains useful with another verified billing system. Your adapter maps the approved customer unit to that system's supported event or invoice procedure. It must handle stable event identifiers, eligible timestamps, corrections and submission failures according to the actual provider contract.

Stripe's Billing Meters reporting guide describes asynchronous usage processing. Submission acceptance and an updated aggregated total are separate observations. Verify the reporting delay and reconciliation procedure for the chosen product rather than expecting a meter to update synchronously. Stripe usage event reporting

Before invoicing, reconcile the outcome ledger to accepted meter events and invoice units. Reconcile provider bills to internal consumption records separately. Investigate unresolved results, missing acknowledgements and unexpected adjustments rather than declaring the two ledgers equal by design.

Batch processing fits deferred work

OpenAI Batch processes asynchronous groups with a 24-hour completion window. Input requests require unique custom_id values, and result/error records can be retrieved for reconciliation. Cancellation takes time and does not make completed work free. This is appropriate for deferred jobs rather than an immediate live-chat response. OpenAI Batch API

In your design, map an attempt-specific custom_id and batch ID to the stable product action. Submitting work every ten minutes is a scheduling decision, not a ten-minute provider turnaround guarantee. Reconcile results by identifier, including unfinished or uncertain requests, before deciding which failures to retry. The same validation and delivery checks apply before creating a customer outcome.

Batching may simplify processing a deferred workload. It does not by itself prevent duplicate billing, eliminate retries or establish that a model output is usable. Those remain product and ledger responsibilities.

Worked hypothetical case: deferred summaries

Imagine ReportFlow SA, a fictional document-summary SaaS. The proposed plan charges one unit for a validated summary made available in the customer's account. The example permits two retry attempts, with a customer deadline that accommodates deferred processing. These settings need owner approval and workload tests.

  1. The application stores the job, workspace and policy version before submission.
  2. Each attempt receives an ID linked to the same customer action.
  3. Returned outputs are validated and stored. Failed or uncertain attempts remain in the internal ledger.
  4. One atomic transaction creates the eligible outcome and billing job. A second successful worker cannot create a second unit.
  5. The billing worker submits that event through the verified adapter and records the provider response.
  6. Reconciliation checks the original event before replaying an ambiguous submission. Finance checks invoice units before finalising the billing run.

In a fictional month, 1,000 eligible summaries and 200 retry attempts yield 1,000 customer units. The retry attempts remain visible in the internal cost ledger. This arithmetic illustrates the proposed policy; it is not a measured saving or evidence of provider charges.

Acceptance, failure and recovery tests

Test Expected result Recovery evidence
Two workers return valid outputs for one action One outcome and one intended meter event Unique-action constraint and both attempt records
Model returns unusable output No outcome unit under this policy Validation reason and notification path
Client times out after provider completion Result stays unresolved until reconciled Provider identifier and retained original action
Worker crashes after result persistence Recovery reuses the existing result No second AI job solely to repair billing
Provider accepts a meter event but acknowledgement is lost No new charging identity Original identifier lookup or documented idempotent replay
Batch result file is missing or corrupt Affected jobs remain unresolved Retrieval/review path; no assertion that work was free
Customer requests intentionally fresh work New action follows disclosed rules Separate customer authorisation and action ID
Incorrect event already invoiced Linked, explained adjustment Original outcome, invoice and correction record

Run a known small fixture through the actual billing integration and read back the invoice result. A local ledger test proves only the local rule; it does not prove provider processing or a correct customer invoice. Choose reconciliation frequency according to volume, invoice timing and recovery limits, with alerts before unresolved events become charges.

Practical worksheet for the pricing decision

Have product and finance owners approve the outcome definition, retry limit, late-result policy, correction rules and customer usage view. Have engineering verify durable identifiers, tenant scope, atomic outcome creation, adapter recovery and invoice readback. Record the evidence for each decision and the person responsible for resolving exceptions.

The CMS vs custom development guide helps assess which usage rules require custom behaviour. Plan ongoing reconciliation and support with the website maintenance cost guide, and map the charging boundary against the user journey.

If your business needs help defining AI outcome billing, explore our SaaS pricing models service and SaaS development services. To turn the ledger and recovery tests into a reviewable implementation brief, get in touch.

Frequently asked questions

Can a failed customer job still cost us money?

Yes. Check provider usage records and pricing terms. Excluding an attempt from the customer invoice does not waive provider charges; monitor that internal cost and margin separately.

Should a partial output be billed?

Under this proposed completed-outcome plan, an unusable partial output creates no unit. A different plan needs a disclosed and approved charging definition, rather than a rule changed after the failure.

Does a successful meter submission prove the invoice is correct?

No. Read back aggregation and the invoice under the actual billing product. Reconcile units, timing and any adjustments to the approved customer outcome records.

Sources

Share this article

Bukhosi Moyo

Written by

Bukhosi Moyo

CEO & Founder

Bukhosi is the founder and lead SEO strategist at Symaxx. He architects search-first digital systems for South African businesses, combining technical engineering with commercial strategy to build long-term organic assets.

Feedback

Was this helpful?

Tell us how this article felt in one click.

Back to Insights

Need help executing this strategy?

Scope the product journey, technical requirements and next delivery milestone for your software.