A research task can expand as each source raises another question. An agent may open more documents, ask additional tools or delegate work while the original objective stays vague. A useful budget policy must decide when the next action is permitted and what the user receives if the task ends before every question is answered.
This article proposes a bounded research-job policy. Its currency values and source examples are fictional. It does not change provider billing, assert a native arbitrary-spend cap or claim that a model can reliably account for its own usage from conversation alone.
Define the task and allowed expansion
Write the decision the research should support, permitted sources and excluded actions. A task to compare documented product features should not quietly expand into contacting vendors, purchasing subscriptions or collecting unrelated customer information. New scope requires the responsible owner's decision.
Record the required output and minimum evidence. A useful partial report can contain verified findings, unresolved questions and source coverage. It should not declare a winner when the criteria or sources remain incomplete.
OpenAI's 10 September 2026 Agents API announcement describes cloud agents using a managed Codex harness, with long-running sessions and tool capabilities. That provides context for sustained work; it does not establish this article's proposed monetary admission policy or prove that an account has a particular budget-control feature. Source: Introducing the Agents API
Separate observable usage from inferred spending
Choose the authoritative usage and charge evidence available in the selected implementation. Record model usage, tool charges, storage or environment costs where applicable, with source and reconciliation status. A model-generated statement that the job cost little is not a billing record.
Some costs are known only after execution. Use verified upper bounds where the system can enforce or reliably reserve them. Where those bounds are unavailable, describe the allowance as a planning estimate and use the available request, output or step controls without claiming a guaranteed monetary ceiling.
Also account for in-flight work. A stop decision prevents new admissions under the proposed policy; it does not erase calls already running or charges already incurred. Reconcile their outcomes before reporting final cost or releasing reserved allowance.
Bounded research-job budget policy
Use this complete proposed policy for fictional task TEST-RESEARCH-A. The owner approves a R50 planning envelope and a separate eight-call limit for the scenario. A hard money ceiling is claimed only if the selected implementation has verified bounds for all relevant charges; otherwise the report labels it an estimate.
| Policy area | Proposed rule | Evidence or stop condition |
|---|---|---|
| Scope | Compare approved public product documentation against the owner's criteria | Hold newly requested sources or actions outside scope |
| Usage ledger | Store confirmed charges, estimated charges, reserved work and their references | Missing usage evidence stops further admission pending reconciliation |
| Next-call admission | Confirm remaining allowance and a verified bound for the proposed operation | Do not admit a call with unknown required cost under a claimed hard ceiling |
| In-flight work | Reserve its approved bound before admitting another call | Count reservations against remaining allowance until settled |
| Step limit | No more than eight approved calls in this fictional task | Stop when the configured enforceable limit is reached, even if money remains |
| Scope or authority change | Recheck current permissions and objective before continuing | Old approval does not authorise a changed research task |
| Stop output | Return verified findings, sources, gaps, current accounting and reason | Do not turn incomplete coverage into a completed recommendation |
Fictional admission arithmetic: confirmed spending is R18 and current reservations are R12, leaving R50 − R18 − R12 = R20 uncommitted. A next operation with a verified R8 bound can be reserved, leaving R12. A further operation requiring a R15 reservation is not admitted because it exceeds that remaining R12. These amounts are fixture inputs, not provider prices or observed costs.
Ledger fields: task reference; owner; scope version; operation and attempt reference; admitted bound; current state; observable usage; confirmed or estimated charge; reservation; source of cost evidence; prior-outcome check; and stop reason. Parallel work shares this ledger so two workers cannot each spend the same apparent remaining allowance.
Acceptance cases: a bounded normal call; an unknown-cost tool; two concurrent admissions; missing usage; a failed call with unresolved charge; a duplicate operation; and the step limit reached before all questions are answered. Inspect the actual enforcement boundary, not only the assistant's final promise to stop.
The policy is accepted when its claims match implemented controls and every admitted operation has the required authority and accounting evidence. If hard bounds cannot be verified, retain the useful planning and stop controls while explicitly declining to call them a guaranteed spend cap.
Partial-evidence output when the budget ends
Use this complete fictional stop report for TEST-RESEARCH-A:
Status: partial research; no final product recommendation. Reason: the next required operation exceeds the remaining approved reservation allowance under the fixture policy.
Verified findings: fixture source SOURCE-A supports feature X for Product A, and fixture source SOURCE-B supports feature Y for Product B. Each statement retains the exact source reference and reviewed version. These are invented example findings, not claims about real products.
Unresolved questions: Product A's current account eligibility has not been verified; Product B's relevant retention configuration remains unknown; the owner has not approved how those gaps affect selection. No missing answer is inferred from a vendor summary.
Accounting: R18 confirmed, R12 previously reserved and R8 newly reserved in the fictional ledger, leaving R12 uncommitted. The R15 next operation was not admitted. Final spending remains pending reconciliation of reserved operations; the report does not claim all reservations became actual charges.
Next authorised decision: the owner can review the partial evidence, narrow the remaining question, authorise a revised bounded plan or end the task. The agent does not increase its own envelope or start an alternative paid tool to bypass the stop.
This report is usable because it preserves what was established and what was not. It is not a complete answer dressed as a success message.
Enforce tool admission outside model assertions
OpenAI's function-calling guide separates model requests from application execution and returned output. In this proposed design, the trusted application checks scope and the admission record before executing an eligible requested tool. A request alone is not authorisation to spend or act. Source: OpenAI function calling
Model processing itself also consumes resources. Limiting application tools does not automatically bound all internal reasoning, generated output, hosted execution or delegated work. Verify the actual provider and environment controls for those components before claiming an end-to-end hard ceiling.
Use explicit stopped, pending and unresolved states. If a request times out after a possible action, reconcile its operation and charge before retrying. A repeated research call may add cost, while a repeated external action can create a second effect.
Keep report structure separate from evidence truth
OpenAI's Structured Outputs guide supports constrained response formats and refusals. A schema can require findings, citations, gaps and accounting status, but does not prove source support or confirm charges. Populate authoritative usage from the trusted ledger and verify findings against sources. Source: OpenAI structured outputs
In a hypothetical normal call, the source supports the stated feature and the ledger records its settled cost. In a missing-source case, the report preserves the gap. In a duplicate case, the application identifies the prior operation and reconciles it rather than spending again because the model forgot the earlier result.
FAQ about research-job spending ceilings
Can the assistant enforce a budget just by promising to stop?
No. Use implemented admission and usage controls, and verify their scope. A model instruction is not an authoritative cost ledger or provider billing limit.
What if the next tool's cost cannot be bounded?
Do not claim a guaranteed monetary ceiling for that operation. Hold it for the owner's decision or use an explicitly approved planning allowance with accurately described limits.
Is a partial report still useful when no recommendation is possible?
Yes. Preserve sourced findings, coverage, unresolved questions and the stop reason. The owner can decide the next bounded step without losing the evidence already gathered.
If your business needs help defining this process, explore Custom AI agents, the wider AI automation services, and our custom-agent workflow guide. The agents and automation comparison and custom AI agent glossary explain the terms. To discuss your records and approval rules, get in touch.

