Why Cheap Tokens Still Make AI Agents Expensive to Run

Suppose your AI agent completes the same number of tasks this month as last month.

Model prices have fallen. You expect the bill to fall too.

Instead, the bill rises.

The first instinct is to blame the model price.

But production builders keep running into a different problem: the expensive part may be the workflow around the model.

An agent can read the same instructions again, carry old tool results forward, call another tool, retry a failed step, ask a second model to check the result, and then repeat the cycle.

The user still sees one task.

The billing system sees a sequence.

Cheap tokens do not guarantee a cheap agent because an agent can buy the same context, tools and decisions many times inside one task.

The goal of this article is to make that bill easier to diagnose. By the end, you should be able to separate model price from context rereading, retries, tool use, multi-agent handoffs and success rate.

Quick Answer

An AI agent costs more than a chatbot when one user request turns into a loop:

Goal
→ model
→ tool
→ result
→ model again
→ retry / branch / verify
→ repeat until done

Each new model call can include not only the new question but also old instructions, tool definitions, conversation history and previous tool results.

So agent economics have four major layers:

  1. Model inference — input, cached input, reasoning and output tokens.
  2. Tool cost — search, computer use, external APIs or other metered services.
  3. Loop overhead — retries, branches, verification and repeated context.
  4. Operational overhead — evaluation runs, logging and human review.

The right question is not “Which model has the lowest token price?”

It is:

How much does a correctly completed task cost?

The Hidden Cost Is Often Re-Reading

A normal chatbot conversation makes the loop visible.

You ask. The model answers. You decide whether to continue.

An agent can make that decision itself.

That matters because later calls often need context from earlier calls.

Context is the information sent to the model for the current request: system instructions, user messages, tool descriptions, retrieved documents, previous decisions and other state.

If the workflow keeps resending most of that information, one token can be billed more than once across the run.

This is the part that is easy to miss when you look only at the final answer.

Original Asset 1: The Context Reread Multiplier

A useful diagnostic is:

Context Reread Multiplier
=
total billed input tokens ÷ unique useful context introduced

This is not a provider billing metric. It is a diagnostic concept for understanding your own workflow.

If a task introduces roughly 50,000 useful tokens of instructions, documents and tool results but the full run bills 500,000 input tokens, the workflow has effectively paid for about ten times as much input processing as the unique information it introduced.

That does not automatically mean waste.

The agent may genuinely need some context repeatedly.

But a high reread multiplier tells you where to look.

  • Are all tool schemas sent on every step?
  • Are old search results still in context after they are no longer needed?
  • Is raw HTML or JSON being carried forward?
  • Does every sub-agent receive the entire history?
  • Could stable prefixes be cached?
  • Could completed work be summarized or stored outside the active prompt?

Original Asset 2: Why Long Agent Loops Can Grow Faster Than You Expect

Consider a simplified workflow.

It starts with a base context of 20,000 tokens. Each step adds another 4,000 tokens of tool results, decisions or intermediate work. Assume each later model call receives the full accumulated history.

After 10 calls, total billed input is approximately:

10 × 20,000
+ 4,000 × (0 + 1 + ... + 9)
= 380,000 input tokens

Now extend the same workflow to 20 calls:

20 × 20,000
+ 4,000 × (0 + 1 + ... + 19)
= 1,160,000 input tokens

The number of calls doubled.

In this simplified example, billed input more than tripled.

The general form is:

Total input ≈ nB + d·n(n−1)/2

where:

  • B = base context,
  • d = new history added per step,
  • n = number of model calls.

This is not a universal law of agent cost. Caching, pruning, summarization, state management and different architectures change the result.

But it explains why “twice as many steps” can become much more than twice the input bill when history keeps growing.

Prompt Caching Helps—But It Does Not Make Context Free

Prompt caching lets a provider reuse previously processed prompt prefixes instead of charging the full uncached-input rate again.

OpenAI says supported requests can receive discounted cached-input pricing when they share reusable prefixes. It also notes that keeping a session alive does not by itself guarantee a cache hit.[1]

That makes prompt design an economic decision.

Stable instructions and tool definitions are easier to cache when they remain in a consistent prefix.

But changing context still has to be processed, and a long cached prompt is still not the same as zero-cost input.

The Average Run Can Hide the Expensive Tail

Another repeated production question is why a system looks affordable on average but still blows through its budget.

The answer can sit in the tail.

P95 cost means the cost level below which 95% of runs fall. The remaining 5% are more expensive.

For agents, those expensive runs often have more:

  • tool calls,
  • failed tool calls,
  • retries,
  • branches,
  • longer context,
  • verification passes.

A useful dashboard therefore includes:

Median cost
+ P90 / P95 cost
+ maximum cost
+ retries
+ tool calls

If you monitor only average cost, you can miss the workflows that create the operational risk.

Original Asset 3: The Retry Tax

Retries are useful.

A temporary API failure should not destroy a task.

But retries need a stopping rule.

An agent does not become tired or financially cautious on its own. If software keeps telling it to try again, it can keep paying for another model call and another tool call.

So the cost of a retry is not just:

one more API request

It may be:

reloaded context
+ another model decision
+ another tool call
+ another result
+ another verification step

A production agent should therefore have explicit limits such as maximum steps, maximum retries, maximum tool calls or a maximum spend per task.

Multi-Agent Systems Add a Coordination Tax

Adding more agents can improve specialization.

It can also duplicate work.

Imagine a planner, researcher, writer and reviewer.

At every handoff, the next agent may need:

  • the original task,
  • important constraints,
  • tool definitions,
  • previous findings,
  • the current draft,
  • feedback from another agent.

If all of that is copied into every handoff, coordination becomes a token cost.

This is why “more agents” is not automatically “more intelligence per dollar.”

The better question is:

Does the extra agent improve the success rate enough to pay for its context, calls and coordination?

Cheaper Models Can Be a False Economy

Model routing can create enormous savings because current API prices span a wide range.

As of October 4, 2026, OpenAI lists GPT-6.1 Sol at $2 per million input tokens and $10 per million output tokens, while GPT-6 Luna is $0.10 per million input tokens and $0.50 per million output tokens.[2][3]

That is a large price difference.

But the cheapest model is not automatically the cheapest workflow.

If a cheaper model:

  • chooses the wrong tool more often,
  • needs more retries,
  • creates more human review,
  • fails more tasks,

then lower cost per call can become higher cost per completed job.

Routing works best when cheap models handle steps they can perform reliably, while expensive capability is reserved for steps where it changes the outcome.

Original Asset 4: Cost per Successful Task

This is still the most important metric from the original article.

Cost per successful task
=
total workflow cost ÷ correctly completed tasks

For a very simple repeated-attempt example:

Agent A costs $0.08 per attempt and succeeds 50% of the time.

Agent B costs $0.12 per attempt and succeeds 95% of the time.

Ignoring other costs:

Agent A ≈ $0.16 per success
Agent B ≈ $0.126 per success

The more expensive attempt is the cheaper successful workflow.

In real deployments, the denominator matters just as much as the numerator.

The Full Cost Receipt

Model tokens are only one layer.

A useful per-task receipt looks like this:

Cost layerWhat to record
Modelinput, cached input, reasoning/output, model used
Toolssearch, computer use, external API or database charges
Loopmodel calls, tool calls, retries, branches, sub-agents
Outcomesuccess/failure, human correction, business value
Tailmedian, P90/P95, maximum cost and duration

Google's current Gemini pricing makes the same architecture visible from another direction: agent usage includes underlying inference across agentic loops, while some tools have their own pricing rules.[4]

The important management lesson is simple.

You need a receipt for the whole workflow, not only a provider invoice at the end of the month.

Five Ways to Cut Cost Without Breaking the Agent

1. Simplify the architecture first

Anthropic recommends starting with the simplest solution that works because agentic systems trade additional cost and latency for flexibility and performance.[5]

2. Reduce rereading

Expire tool results that are no longer needed, retrieve only relevant documents, and summarize completed work when the summary is cheaper than repeatedly carrying the full history.

3. Route by capability

Use cheaper models for steps where they meet the quality threshold. Reserve stronger models for decisions where capability changes success.

4. Budget loops

Set hard limits on steps, retries, tool calls or dollars. Escalate uncertain cases instead of allowing open-ended attempts.

5. Optimize the outcome, not the token count

A shorter prompt that causes a failure and a retry can cost more than a slightly larger prompt that succeeds once.

Why This Matters More as Model Prices Fall

Falling model prices can make agents cheaper.

They can also make teams comfortable adding more steps.

More tools. More verification. More sub-agents. More context. More autonomous work.

McKinsey's 2026 analysis makes a similar point: token prices alone do not describe agentic unit economics because repeated inference and workflow design determine the cost of the business outcome.[6]

So cheaper intelligence does not remove the need for cost engineering.

It changes where the optimization moves.

What to Watch Next

  1. Context efficiency. How many billed input tokens are spent per unit of genuinely new information?
  2. Tail behavior. Do P95 runs stay bounded as the agent gets more capable?
  3. Routing quality. Does using cheaper models preserve success rate?
  4. Multi-agent overhead. Do additional agents increase completion quality more than coordination cost?
  5. Tool economics. As model tokens get cheaper, do search, browser, database or external API fees become a larger share?
  6. Outcome attribution. Can a team trace every expensive run to a task, branch and result?

The Simple Idea to Remember

An agent is not one answer.

It is a workflow whose meter can run every time the system reads context, calls a model, uses a tool, retries, branches or verifies.

The cheapest token is not the goal. The cheapest reliable completion is.

If agent bills surprise you, do not start by asking which model is expensive.

Start by asking what the workflow keeps paying to do again.

Key Vocabulary

context
The instructions, messages, documents, tool definitions and other state sent to a model for a request.

prompt caching
A provider mechanism that can reuse previously processed prompt prefixes at a lower input cost.

model routing
Sending different steps to different models based on the capability, cost and latency required.

P95 cost
A tail metric: 95% of runs cost at or below this level, while 5% cost more.

context reread multiplier
A diagnostic concept used in this article: total billed input divided by the unique useful context introduced during a workflow.

cost per successful task
Total workflow cost divided by the number of tasks completed correctly.

Related Articles

Sources

  1. OpenAI — Prompt caching, checked October 4, 2026.
  2. OpenAI API — GPT-6.1 Sol model and pricing, checked October 4, 2026.
  3. OpenAI API — GPT-6 Luna model and pricing, checked October 4, 2026.
  4. Google AI for Developers — Gemini API pricing, including tools and agents, checked October 4, 2026.
  5. Anthropic — Building Effective Agents.
  6. McKinsey — Where AI agents pay off: A practical guide to the economics of agentic workflows, August 24, 2026.

Update History

October 4, 2026 — Updated with current model pricing, traceable production-cost questions, and new context-reread and tail-cost frameworks.

Sources checked through October 4, 2026. Provider pricing changes frequently. The worked examples illustrate workflow mechanics and are not forecasts of any company's agent bill. Community discussions were used to identify reader questions, not as factual benchmarks.