How to Reduce AI Costs Without Breaking Quality
A first-principles playbook for cutting LLM and AI infrastructure costs by measuring cost per successful task, routing work, caching context, and using evals as the safety rail.
Cutting AI costs is not the same as buying cheaper tokens. The unit that matters is cost per successful task:
cost per success = total model + retrieval + tool + retry cost / successful outcomes
A model that costs half as much per token but needs twice as many retries has not saved anything. A smaller model that completes 94% of routine requests and hands the hard 6% to a stronger model often has.
Start with a cost-and-quality trace
For every request, record the model, input and output tokens, cached tokens, latency, tool calls, retries, outcome score, and business slice. Join billing and quality at the request ID. Without that join, finance sees spend and engineering sees evals, but nobody can see efficiency.
Build the first report with five rows:
| Slice | Volume | Cost/request | Success rate | Cost/success |
|---|---|---|---|---|
| FAQ | 52% | $0.012 | 96% | $0.013 |
| Account changes | 18% | $0.043 | 91% | $0.047 |
| Research | 12% | $0.31 | 72% | $0.43 |
| Tool failures | 8% | $0.18 | 38% | $0.47 |
| Other | 10% | $0.06 | 80% | $0.075 |
Optimize the expensive successful outcome, not the loudest invoice line.
1. Remove tokens that do no work
Inspect real prompts. Repeated policy documents, verbose tool descriptions, duplicated conversation history, and retrieval chunks with no bearing on the answer are common waste.
- Put stable instructions first so prompt caching can match the prefix.
- Retrieve fewer, better-ranked chunks and measure recall before shrinking further.
- Summarize old conversation turns into explicit state.
- Return structured fields from tools instead of prose blobs.
- Cap output length to what the interface can use.
Do not delete context blindly. Run the same regression set before and after each change and inspect the slices most dependent on that context.
2. Route by task difficulty
One model should not handle every request. Create a routing policy with three bands:
- Deterministic path: validation, lookup, formatting, and arithmetic that code can do.
- Efficient model: classification, extraction, short grounded answers, routine tool selection.
- Frontier model: ambiguous planning, long-horizon reasoning, difficult synthesis, or recovery.
Route on observable signals such as task type, context length, required tools, risk level, and whether the first attempt failed. Evaluate the router itself: false escalations waste money; missed escalations damage quality.
3. Cache the stable prefix
Prompt caching works when the beginning of requests repeats. Put system instructions, tool schemas, and shared reference material before user-specific content. Track cached-input tokens separately and alert when cache coverage falls after a deployment.
For work that does not need an immediate response—nightly classification, embedding backfills, or broad eval runs—use an asynchronous batch endpoint when the provider offers one. OpenAI’s Batch API, for example, processes supported requests asynchronously within its completion window and advertises discounted pricing. Confirm the current terms before designing around them.
4. Stop paying for retries you cannot explain
Retries are often the most invisible AI tax. Split them into transport retry, provider capacity retry, invalid structured output, tool failure, and low-confidence retry. Give each a limit and an owner.
An automatic retry with the same prompt and model is usually wishful thinking. A useful retry changes something: repairs JSON, removes a failed tool, supplies missing evidence, or escalates to a stronger route.
5. Use evals as the guardrail
Cost work without evals turns into a debate about anecdotes. Freeze a representative dataset, define hard and soft graders, record a baseline, and compare one change at a time.
Hard graders should cover schema validity, correct tool choice, permissions, citations, and required facts. Use calibrated human or model judgment for tone, completeness, and semantic correctness. Always inspect scores by task and risk slice.
The interactive lab below makes the release rule concrete: change the application type and threshold, then see which examples block the cheaper configuration.
A 30-day reduction plan
Week 1: establish the denominator
Instrument tokens, cache hits, tools, retries, outcomes, and cost. Publish cost per successful task by slice.
Week 2: remove obvious waste
Shorten repeated context, cap outputs, repair uncontrolled retries, and batch asynchronous work.
Week 3: route
Move deterministic steps into code. Send routine tasks to an efficient model and define explicit escalation signals.
Week 4: lock the gains in
Put the regression set in CI, alert on cache-coverage and retry changes, and review the ten most expensive failed tasks every week.
What to measure every week
- Cost per successful task
- Success rate by task and risk slice
- Input, output, and cached token ratios
- Retry and tool-failure rate
- Model-route distribution and escalation precision
- P50 and P95 latency
- Quality change against the pinned baseline
The durable advantage is not one optimization. It is a feedback loop where economics and quality share the same evidence.
Continue learning
Practice the platform workflow in the evaluation course library, model provider costs in the interactive cost modeller, then use the chatbot eval guide to create the quality floor that makes optimization safe.
Interactive test
Find the cheapest candidate that still clears the quality floor
System under test
Answer Northstar Shoes customers using only the return policy.
Can I return unworn shoes after 20 days?
Say yes and mention the 30-day window
Yes. Unworn shoes can be returned within 30 days of delivery.
24-second visual model
Optimize cost per successful AI task
Why cheap tokens do not guarantee a cheaper AI system, and how context, routing, and caching lower cost without crossing the quality floor.
Read the visual transcript
- 01Cheap tokens multiplied by retries and failures can still produce an expensive system.
- 02Change the denominator: measure cost per successful task while holding a defined quality floor.
- 03Use three levers: remove context that does no work, route each task to the simplest model that succeeds, and cache the stable prefix.
- 04The example moves from forty-seven cents to eighteen cents per successful task while quality remains above the floor.
Frequently asked questions
What is the best metric for reducing AI costs?
Measure total model, retrieval, tool, and retry spend per successful task. Token price alone can hide expensive retries and failed outcomes.
How can I reduce LLM costs without lowering quality?
Remove unused context, cache stable prompt prefixes, route routine work to efficient models, batch asynchronous jobs, and compare every change on a frozen evaluation set.
Should every request use the same AI model?
Usually not. Deterministic code should handle deterministic work, an efficient model should handle routine cases, and a stronger model should receive ambiguous, high-risk, or failed cases.
Why are evals necessary for AI cost optimization?
Evals define the quality floor. They show whether a cheaper prompt, model, route, or retrieval configuration preserves the behaviors users need before it reaches production.
Bhaulik Patel
Forward deployed AI engineer and creator of Deployed Engineer.