All guides
CostProduction·Intermediate9 min

How to Cut Prompt Token Cost Without Losing Quality

Cut the cost per successful output, not the length of the prompt: output and hidden reasoning tokens usually cost several times as much as input, a cached prefix costs a fraction of it, and a short prompt that needs three attempts costs more than a longer one that works the first time.

A prompt optimizer that makes prompts longer, writing about cutting token cost, owes you an explanation. Here it is: the prompt is rarely where the money goes. Output costs more than input, reasoning you never see is billed as output, and every retry pays for the whole exchange again. This guide covers where LLM spend actually comes from and the levers that reduce it without making the answers worse.

By Andrei Bădulescu, Founder at VantagePrompt·Updated

What should you measure instead of prompt length?

Cost per successful output: everything you paid, divided by the answers you could actually use. A prompt is a means to an output, and an output you throw away still cost money — you paid for its input, its output, and any reasoning behind it. Shortening a prompt saves input tokens on every call; if it also makes the model miss more often, the retries eat the saving and more.

Short promptLonger prompt
Input tokens per attempt2001,200
Output tokens per attempt800800
Cost per attempt (output at 5× input)4,200 units5,200 units
Attempts until a usable answer31
Cost per successful output12,600 units5,200 units
Illustrative numbers, not a measurement: one unit = the price of one input token. The longer prompt costs 24% more per attempt and less than half as much per usable answer.

That is the honest case for a longer prompt, and its limit. Words that add a decision — a format, an exclusion, a check — pay for themselves when they prevent a retry. Words that add nothing are pure input cost. Prompt Optimization covers telling the two apart.

Where does the money actually go?

To the output side. On every current model below, an output token costs five to eight times as much as an input token:

ModelInput, per 1M tokensOutput, per 1M tokensRatio
Claude Sonnet 5.5$2.00$10.005×
Claude Haiku 4.5$1.00$5.005×
gpt-6-luna$0.10$0.505×
Gemini 2.5 Flash$0.30$2.50≈8×
Standard list prices from the Anthropic, OpenAI and Google pricing pages on October 3, 2026. Prices change; the ratio has been stable for longer than the numbers.

Then there are tokens you never see. Reasoning models think before they answer, and that thinking is billed as output whether or not it is returned to you. OpenAI and Google say the same in their own docs; Anthropic puts it plainly:

Thinking has a cost: the tokens Claude spends reasoning are billed as output tokens, even when the thinking text isn't returned to you, and they count toward max_tokens alongside the response text.
— Anthropic, Thinking

So the cheapest tokens in a request are the ones in your prompt, and the most expensive are the ones you cannot read. Most of the levers below work on the output side.

How much can turning reasoning off save?

Many current models are hybrids: they reason by default unless told not to. For a task that does not need it — rewriting, formatting, extraction, a structured expansion — that reasoning is paid output that never reaches you. We measured it in VantagePrompt’s own pipeline. The same prompt on the same economy model, xiaomi/mimo-v2.5, before and after sending the provider’s reasoning-off switch:

Reasoning on (default)Reasoning off
Wall time60.6 s23.0 s
Total tokens5,7894,054
Credits charged32
Quality score100100
One before/after pair on the same prompt, recorded on August 23, 2026 — a single measurement, not a benchmark. Recorded completion tokens fell from 2,875 to 1,096 while the visible output stayed the same length.

The output got slightly longer and was not truncated; the time and the bill fell because the hidden thinking was gone. Since then VantagePrompt sends the reasoning-off switch on free-tier optimizer runs unless the user picks a reasoning effort — an explicit effort always wins. The rule generalises: turn reasoning down or off for tasks that are transformations of text you already have, and keep it for problems with steps. Chain-of-Thought covers which is which.

Hiding reasoning is not the same as skipping it. An option that only excludes the reasoning text from the response still bills for it — you need the setting that turns reasoning off, or a lower reasoning effort.

When does prompt caching pay off?

When many requests share a long, identical beginning — a system prompt, a document, a block of examples. The provider stores that prefix after the first request and charges a fraction of the input price to reuse it. Order the prompt for it: everything stable first, everything that varies last, because a cache matches on an exact prefix and one changed character ends the match.

ProviderHow it turns onCache readCache writeLifetime
AnthropicAdd cache_control — automatically, or at explicit breakpoints0.1× base input (lower on some models)1.25× for 5 minutes, 2× for 1 hour5 minutes by default, refreshed on each hit
OpenAIAutomatic on supported models, from 1,024 tokens on GPT-5.6 and laterGPT-5.6+: 0.1× input on most models (0.05× on GPT-6.1 Sol)GPT-5.6+: 1.25× input; free on earlier modelsGPT-5.6+: at least 30 minutes after the last use
Google GeminiImplicit caching on by default (2.5 and newer); explicit caches on requestDiscounted; guaranteed only with explicit cachingExplicit caches also bill storage per hour1 hour by default for explicit caches
From each provider’s prompt-caching documentation, October 3, 2026. Minimum cacheable lengths vary by model; Anthropic notes that a prompt below the minimum is processed without caching and returns no error.

Caching does nothing for a prompt that changes from the first token, and a write costs more than a plain read — so a prefix used once is more expensive cached than not. It pays from the first reuse at a 1.25× write and from the second at a 2× write, as long as the reuse lands within the cache lifetime. VantagePrompt marks its system prompt as cacheable on Anthropic and recent Gemini models when it passes a length threshold, and records cached tokens per run.

Does every request need the same model?

No, and this is usually the largest lever. Prices between models differ by an order of magnitude or more, and most traffic — classification, extraction, rewriting, short structured answers — does not need the strongest model available. Send that traffic to a cheap model, keep the expensive one for the requests that fail without it, and cap what any single request may cost. Model Selection & Smart Routing covers routing by price, fallback chains and hard price caps.

In VantagePrompt the economy band — models priced at or under $0.15 per million input tokens and $0.60 per million output tokens — is offered on every plan. One credit is $0.0005 of upstream model cost, rounded up, with a one-credit minimum per run; the credit calculator prices the same reference run on every model in the curated list.

Is batching worth the wait?

If the answer is not needed now, usually yes. OpenAI, Anthropic and Google all offer a batch interface at 50% of the standard price, in exchange for asynchronous processing with a turnaround of up to 24 hours. Anthropic notes that its discounts stack, so cached input inside a batch is cheaper still. For work that a person is waiting on, the delay rules it out.

VantagePrompt’s own batch runner optimizes 2 to 20 prompts in one submission and uses the reduced Flex inference rate where the provider accepts it. It does not use half-price overnight batch models, because the optimizer is built to return within minutes.

Which cost cuts end up costing more?

  • Cutting context until quality drops. Removing the examples or the source document saves input tokens and buys retries, which cost input and output again.
  • Capping output below what the answer needs. A low output limit truncates the answer; a truncated answer is usually a full retry. Set the cap to bound runaway generation, not to squeeze a complete answer.
  • The cheapest model for a task that needs reasoning. A model that cannot hold the steps fails more often, and each failure costs a full call.
  • Hiding reasoning instead of disabling it. Excluding the reasoning text from the response saves nothing; it is still billed.
  • Caching a prefix that is used once. The write costs more than a normal read; caching pays only on reuse.

What does a token-cost checklist look like?

  • Track cost per usable output, not tokens per prompt.
  • Look at output and reasoning tokens first — they are the expensive side.
  • Turn reasoning off or down for text transformations; keep it for multi-step problems.
  • Put the stable part of the prompt first and cache it if it repeats.
  • Route routine traffic to a cheaper model and cap the price per request.
  • Batch anything that can wait a day.
  • Before cutting words from a prompt, check whether they prevent a retry.

Frequently asked questions

How do I reduce the token cost of my prompts?
Measure cost per usable output rather than prompt length, then work on the expensive side first: disable or reduce reasoning for tasks that do not need it, cache long prompt prefixes that repeat, route routine requests to cheaper models, and batch work that can wait. Shorten a prompt only where the words removed were not preventing failures.
Are output tokens more expensive than input tokens?
Yes. On current Anthropic, OpenAI and Google models an output token usually costs about five to eight times as much as an input token. Reasoning tokens from thinking models are billed as output too, even when the reasoning text is not returned.
Does a longer prompt cost more?
Per call, slightly — input tokens are the cheap side of the bill. Per usable answer it can cost less, if the added instructions prevent retries. A prompt that works the first time is cheaper than a shorter one that needs three attempts.
How much does prompt caching save?
A cache read is billed at a fraction of the normal input price — 0.1 times the base input price on most current Anthropic and OpenAI models — but writing the cache costs more than a normal read. It pays off from the first reuse when the write costs 1.25 times the input price, and from the second when it costs twice, provided the reuse comes within the cache lifetime.
Should I turn off reasoning to save money?
For tasks that transform text you already have — rewriting, formatting, extraction — usually yes. In one measured VantagePrompt run, turning reasoning off on the same prompt cut wall time from 60.6 to 23.0 seconds and the charge from 3 credits to 2 with the same quality score. Keep reasoning for problems that genuinely have steps.

Sources

Put it into practice.

Run this technique in the optimizer.

Open the optimizer

Keep reading