A chat message costs one model call. An agent task costs five to fifty. That multiplier — not the per-token price — is what decides whether your AI agent costs cents or dollars per task. This guide does the math at October 2026 prices and shows the exact knobs that keep agentic workloads inside a budget.
Last verified: October 1, 2026. Prices from the providers' published pages — see the full price table for the complete comparison.
Why agents burn more tokens than chat
Four multipliers stack on top of the base per-token price:
- The loop. An agent reasons, calls a tool, reads the result, reasons again. Each cycle re-sends the growing conversation. A 10-step task sends ~10 model calls, and call #10 carries the context of the previous nine.
- Quadratic-ish context. As the transcript grows, every subsequent call re-pays for earlier tokens. A task that ends with 50K tokens of history can bill 200K+ cumulative tokens across the run.
- Reasoning tokens. Reasoning models (GPT-6 Astra/Sol, Claude Opus 5.5 with extended thinking) bill internal chain-of-thought as output tokens — at output prices, 3–5x input.
- Retries and verification. Well-built agents verify their work — which means extra calls. Undisciplined agents retry failures without backoff and multiply the bill further.
Worked example: one research task, three models
A representative research-and-summarize agent task: 15 tool-calling rounds, ending with ~60K tokens of cumulative context and ~8K tokens of generated output (including reasoning):
| Model | Cumulative input | Output (incl. reasoning) | Task cost |
|---|---|---|---|
| GPT-6 Luna (0.50 per MTok) | ~60K → $0.006 | 8K → $0.004 | ~$0.01 |
| DeepSeek V4 (~1.75 per MTok) | ~60K → $0.045 | 8K → $0.014 | ~$0.06 |
| GPT-6.1 Sol (10 per MTok) | ~60K → $0.12 | 8K → $0.08 | ~$0.20 |
| Claude Opus 5.5 (~20 per MTok) | ~60K → $0.24 | 8K → $0.16 | ~$0.40 |
| GPT-6 Astra (50 per MTok) | ~60K → $0.60 | 8K → $0.40 | ~$1.00 |
The same task, executed faithfully by the same agent logic, spans two orders of magnitude: one cent to one dollar. Multiply by volume — 1,000 tasks/month is 1,000 on the flagship — and the routing decision becomes the budget.
The five budget knobs
- Route by task tier. Cheap models for exploration and routine steps; escalate to flagships only for the hard reasoning step. This is the single biggest lever — the table above is the size of the prize.
- Cap max output tokens. Reasoning tokens bill at output prices; an uncapped reasoning model can spend more thinking than working. Set the ceiling in your global params.
- Compress context aggressively. Truncate tool outputs, summarize old turns, drop what the loop no longer needs. Halving cumulative input halves the dominant cost term.
- Bound the loop. Max iterations with hard stops — a looping agent is a burning budget. Every serious agent runtime exposes this; use it.
- Cache what's cacheable. Prompt caching (where the provider supports it) discounts re-sent context steeply — on long agent transcripts this is often a 30-60% input saving.
Why BYOK makes agent budgets transparent
On a token-resale platform, agent costs arrive as abstract credits with a margin baked in — you can't tell model price from platform markup, and you can't route to the cheap tier if the platform doesn't offer it. With BYOK, every call's cost is the provider's list price on your own bill: agent telemetry maps 1:1 to provider costs, you choose the model per task tier, and the platform never marks anything up. SynthHires' usage dashboard shows tokens and cost per task for exactly this reason — see Usage, health & teams.
FAQ
How much does it cost to run an AI agent per month?
A light personal agent (a few tasks a day on efficient models) runs 20–100/month. Heavy automation on flagship models can run $500+/month. The dominant variables are task volume, loop length and model tier — in that order.
Why is my AI agent so expensive?
Almost always one of: an unbounded loop (capped by max iterations), reasoning tokens left uncapped (cap max output), flagship models doing exploration work (route by tier), or uncontrolled context growth (truncate and summarize). Audit in that order.
Do reasoning models cost more for agents?
Yes — their internal chain-of-thought bills as output tokens at 3–5x input price. They earn it on hard steps (planning, debugging, verification) and waste it on routine ones. The efficient pattern: cheap fast models in the loop, one reasoning-model call for the hard decision.
How do I cap an agent's spend?
Stack four limits: max iterations per task, max output tokens per call, a provider-side monthly spend cap (every billing console supports one — set it the day you create the key), and model routing so expensive tiers only see hard steps. SynthHires' governance dashboard surfaces cost per task so the caps are observable.
What's the cheapest model that can still run a tool-using agent loop?
The efficient tiers — GPT-6 Luna, DeepSeek V4, Grok 4.1 Fast — handle structured tool-calling loops well at 1.00/MTok input. The October 2026 price table ranks them; grab keys via the Get API keys guide and test your workload on two tiers before committing.
Set your caps, grab your keys, and check the math against your own dashboard — the Usage & governance doc shows what SynthHires tracks per task.