Google AI ModeSep 26, 2026
Data as of Sep 29, 2026 · Based on 53 AI responses from ChatGPT Search and Google AI Mode · See how Parse measures this
OpenAI Platform is named in 60% of the latest answers, most often first. Anthropic follows, in 40%.
Brands AI recommends here
Mentioned inRecommended in · Sep 6 – Sep 26, 2026
Question: How do I cut my model bill with caching and smart routing between cheap and premium models?
Google AI ModeSep 26, 2026
Question: How do I cut my model bill with caching and smart routing between cheap and premium models?
ChatGPT SearchSep 26, 2026
The biggest LLM cost reductions usually come from not calling the expensive model, not resending the same tokens, and only escalating hard cases. A practical architecture is:
Question: How do I cut my model bill with caching and smart routing between cheap and premium models?
ChatGPT SearchSep 22, 2026
Question: How do I cut my model bill with caching and smart routing between cheap and premium models?
Google AI ModeSep 18, 2026
Question: How do I cut my model bill with caching and smart routing between cheap and premium models?
ChatGPT SearchSep 18, 2026
Question: How do I cut my model bill with caching and smart routing between cheap and premium models?
ChatGPT SearchSep 14, 2026
Question: How do I cut my model bill with caching and smart routing between cheap and premium models?
ChatGPT SearchSep 10, 2026
Yes. The strongest pattern is cache aggressively, route conservatively, and escalate only when the cheap model shows signs it needs help.
Question: How do I cut my model bill with caching and smart routing between cheap and premium models?
Google AI ModeSep 10, 2026
39% of citations to these sources link to brands' own websites.
pub.towardsai.net
openai.com
platform.claude.com
Cutting your LLM bill by combining caching and smart routing requires a layered architecture. By intercepting repetitive queries and sending simple tasks to cheap models while reserving frontier models for complex reasoning, you can slash costs by 50–80% without hurting output quality.
Here is the blueprint to set this up effectively:
Caching has the highest ROI because 0 computation beats a cheap model every time. Stack your caches in order of precedence:
Don't use a frontier model (GPT-4o, Claude 3.5 Sonnet) for tasks a smaller model (GPT-4o mini, Claude 3 Haiku, or an open-source Llama 3 derivative) can handle. Route requests based on complexity:
A router is only as good as the evaluation harness backing it. If your cheap model starts failing at a task it shouldn't handle, your routing threshold is wrong.
If you want to dive deeper, let me know:
I can give you a more specific architectural pattern or code snippet to match your stack.
These approaches are commonly combined in production systems because they attack different parts of the bill.
Use for deterministic or repeated queries:
Examples:
Flow:
User request
|
v
Hash(normalized prompt + settings)
|
+--> Cache hit --> return response
|
+--> Cache miss --> call model --> store resultCache key should include:
hash(
system_prompt,
user_prompt,
model,
temperature,
tools,
output_format
)Avoid caching:
Exact matching misses when wording changes:
"How do I reset my password?" "I forgot my password, how can I get back in?" Use embeddings:
User query
|
Embedding
|
Vector similarity search
|
Similarity > threshold?
|
+--> Return cached answer
|
+--> Call LLMTypical pattern:
Semantic caching can eliminate entire model calls rather than merely making them cheaper.
For long repeated prefixes:
SYSTEM PROMPT (stable)
TOOLS (stable)
COMPANY DOCS (stable)
---------------------
USER MESSAGE (changes)Cache the stable section.
Good candidates:
Keep cached prefixes identical; small changes can invalidate the cache.
Do not choose the model in application code:
Bad:
model="premium-model"Better:
model = router.choose(request)Example routing policy:
| Request type | Model |
|---|---|
| Classification | Cheap |
| FAQ lookup | Cheap |
| Extraction to JSON | Cheap |
| Simple summaries | Cheap |
| Coding help | Mid |
| Complex reasoning | Premium |
| Customer escalation | Premium |
A router can be rule-based first; you do not need ML immediately.
Example:
if task == "classify":
return "cheap-model"
if tokens < 1000 and difficulty < 0.3:
return "cheap-model"
if requires_reasoning:
return "premium-model"A very effective pattern:
Request
|
v
Cheap model
|
confidence check
/ \
good bad
| |
v v
return Premium modelExamples of escalation signals:
This works especially well for:
It works less well for:
Before calling a model, score the request.
Signals:
difficulty =
+ long context
+ many constraints
+ coding request
+ math/reasoning
+ multiple documents
- FAQ match
- known intent
- structured outputExample:
{
"difficulty": 0.82,
"route": "premium"
}You can generate this score with:
Routing a smaller model helps, but unnecessary tokens hurt every model.
Common wins:
Instead of:
20,000 token conversation
+
new questionsend:
2,000 token summary
+
new questionBad:
Retrieve 20 documents
Send allBetter:
Retrieve 20
rerank
send top 3-5Set:
{
"max_tokens": 500
}instead of letting responses grow indefinitely.
Context trimming and token hygiene are frequently among the easiest cost reductions.
Architecture:
App
|
v
LLM Gateway Layer
|
+------------+-------------+
| | |
Cache Router Budget
| | |
v v v
Cheap Model Premium Model LimitsThe gateway handles:
Do not only track token cost.
Track:
| Metric | Why |
|---|---|
| Cost/request | Basic economics |
| Cost/resolved task | Real business metric |
| Cache hit rate | Shows avoided spend |
| Cheap-model success rate | Determines routing quality |
| Escalation rate | Shows router behavior |
| User correction rate | Quality signal |
| Latency | User experience |
A cheap model that causes retries may cost more than a premium model that succeeds once.
A well-tuned system often ends up with a distribution like:
70-90% cheap model
10-25% mid model
1-10% premium modelThe exact split depends on your workload and quality requirements; the key is measuring it rather than assuming every request needs the largest model.
If you tell me your stack (OpenAI API, Anthropic, Azure, self-hosted, LangChain/LlamaIndex, etc.) and workload (chatbot, RAG, coding agent, extraction), I can sketch a concrete router + cache design.
70-90% cheap model
10-25% mid model
1-10% premium modelThe exact split depends on your workload and quality requirements; the key is measuring it rather than assuming every request needs the largest model.
Architecture:
Yes. The biggest savings usually come from (1) caching repeated context, (2) routing easy requests to cheap models, and (3) only escalating when the cheap model is likely to fail.
Put the stable material first:
[SYSTEM INSTRUCTIONS — stable]
[FEW-SHOT EXAMPLES — stable]
[REFERENCE DOCUMENTS — stable]
[USER REQUEST — changes every call]Don't put timestamps, user IDs, conversation-specific text, or other changing material in the middle of the reusable prefix. For providers that use prefix caching, even small changes before the variable portion can destroy the cache hit. OpenAI's current documentation describes automatic prefix caching, while Anthropic supports automatic or explicit cache breakpoints.
Track at least:
cache_read_tokens / input_tokens
cache_write_tokens
cache_hit_rateCaching economics matter too. For example, Anthropic currently charges 1.25× base input price for a 5-minute cache write and 0.1× for cache reads; the 1-hour write is 2×. So frequent reuse can make caching extremely attractive.
Instead of:
every request → premium modeluse:
┌─ cheap model ──→ answer
request → classifier ────┤
└─ premium ──────→ answer
↑
uncertainty/failureA practical policy:
| Request | Default |
|---|---|
| Classification / extraction | Cheap |
| Simple transformation | Cheap |
| FAQ / straightforward lookup | Cheap |
| Simple coding change | Cheap |
| Summarization | Cheap |
| Ambiguous reasoning | Premium |
| Complex coding/debugging | Premium |
| Multi-step planning | Premium |
The important trick is not to send the entire prompt to a premium model just because the request might be difficult.
Have the cheap model return structured metadata:
{
"answer": "...",
"confidence": 0.91,
"needs_premium": false,
"reason": "straightforward extraction"
}Then:
if result.needs_premium or result.confidence < 0.80:
return premium(request)
return result.answerEven better, use task-specific escalation rules rather than one universal confidence threshold.
For example:
coding:
escalate if tests fail
escalate if required files aren't identified
escalate if confidence < .85
customer_support:
escalate if refund/legal/security issue
escalate if confidence < .90
summarization:
don't escalate unless required fields are missingThis prevents your router from becoming an expensive second LLM.
Don't think of "caching" as just prompt caching.
Use:
L1 — application cache
Exact same request → return previous result.
hash(model + prompt + relevant parameters)
↓
RedisL2 — semantic cache
Similar questions with sufficiently similar answers → reuse an existing response.
Be conservative here; semantic caching is dangerous when answers depend on freshness, permissions, or user-specific information.
L3 — provider prompt cache
Reuse your giant system prompt, tools, examples, documentation, etc.
The provider-level cache is particularly valuable for agents because the same instructions/tools/context can be sent repeatedly. Anthropic's own cost guidance identifies prompt caching as one of its largest cost levers for agent workloads.
Common mistakes:
Today is 2026-09-22...
Request #918273...
Random ID: abc123...
Current user message...
[huge static prompt]Instead:
[huge static prompt]
[stable tools]
[stable documentation]
[dynamic request]Also keep tool definitions stable. Anthropic specifically notes that changes to tool use settings can invalidate the cache.
People often obsess over input-token costs while ignoring output.
Set an appropriate output ceiling:
max_output_tokens = {
"classification": 100,
"extraction": 500,
"normal_chat": 1000,
"complex_reasoning": 4000,
}And tell the model what format you actually need.
If you need:
{"sentiment":"positive"}don't ask for a 500-word explanation.
A useful mental model is:
Expected cost = cheap-model cost + escalation probability × premium-model cost
Suppose:
cheap = $0.10/request
premium = $1.00/requestIf the cheap model solves 90% of requests:
$0.10 + (10% × $1.00)
= $0.20/requestversus:
premium every time = $1.00/requestThat's an 80% reduction before even considering caching.
The exact economics will depend on your token mix and provider pricing; current provider pricing differs substantially by model and cached-vs-uncached input.
I'd implement this:
┌───────────────┐
│ Exact cache? │
└───────┬───────┘
│ miss
▼
┌───────────────┐
│ Cheap router │
└───────┬───────┘
│
┌─────────┴─────────┐
│ │
easy/confident difficult/
│ uncertain
▼ ▼
CHEAP MODEL PREMIUM MODEL
│ │
└─────────┬─────────┘
▼
┌───────────────┐
│ Validate │
└───────┬───────┘
│
failure/uncertain
▼
PREMIUM RETRYThen measure cost per successful task, not simply cost per API call.
That's crucial: if the cheap model costs 5× less but requires lots of retries, the apparent savings can disappear.
If you tell me which API/provider you're using (OpenAI, Anthropic, Gemini, Bedrock, etc.) and roughly how many requests/tokens you process per day, I can sketch a concrete router + caching architecture and estimate the potential monthly savings.
Trimming your LLM bill down without hurting output quality comes down to a two-punch combo: prompt caching to slash repeated input costs, and smart routing to make sure simple tasks go to cheap models instead of burning cash on frontier giants.
Using frontier models for every single request is the fastest way to hemorrhage budget. Here is a practical playbook to get your unit economics under control.
Prompt caching saves up to 50% to 90% on input token costs for repeating prefixes. It doesn't change what the model outputs; it just stops you from re-paying to process the same block of text over and over.
Prompt caching is the highest-leverage, lowest-risk cost reduction available to production LLM teams. The saving is structural, not a quality trade-off.
Smart routing means putting a fast, cheap model (like GPT-4o mini, Claude 3 Haiku, or Llama 3.8B/8B variants) at the front door. A heavy frontier model is only called when the cheap model signals it's out of its depth.
Using frontier models for every request is the single biggest source of wasted LLM spend. Smaller models handle most classification, summarisation, and routing tasks just as well.
If you'd like, tell me:
I can give you a more specific routing strategy or architecture layout.
Yes. The biggest savings usually come from (1) caching repeated context, (2) routing easy requests to a cheaper model, and (3) only escalating when the cheap model is uncertain or the task is genuinely hard.
Put stable material first:
SYSTEM / STATIC
- system instructions
- policies
- tool definitions
- schemas
- few-shot examples
- long reference material
DYNAMIC
- user request
- conversation-specific data
- retrieved documentsPrompt caching generally works on exact prefix matches, so changing something near the beginning can invalidate the useful cached prefix. OpenAI's documentation specifically recommends putting static content first and variable content last.
For an agent, this can be a huge deal because every iteration may resend the system prompt, tools, and accumulated context. Anthropic reports that caching can reduce agent-loop costs substantially in its benchmarks.
Track:
cache_hit_rate = cached_input_tokens / total_input_tokensI'd aim for 80%+ on workloads with lots of repeated context and investigate anything substantially lower.
Don't route based merely on "short prompt = cheap model." Route based on task difficulty and required reliability.
A practical policy:
| Request | Route |
|---|---|
| Classification, extraction, formatting | Cheap |
| Simple Q&A / summarization | Cheap |
| Routine tool calls | Cheap |
| Coding/debugging with obvious solution | Cheap/medium |
| Ambiguous reasoning | Medium |
| Complex coding | Premium |
The important pattern is:
┌── cheap ── success ──> answer
request ──> router
└── difficult/uncertain ──> premiumRather than:
request ──> premiumThis is essentially a selective escalation architecture.
For example, have the inexpensive model return structured metadata:
{
"answer": "...",
"confidence": 0.91,
"needs_escalation": false,
"reason": "straightforward extraction"
}Then escalate if:
if result.needs_escalation or result.confidence < 0.75:
return premium_model(request)
return result.answerBut don't blindly trust a model's self-reported confidence. Calibrate the threshold against a labeled evaluation set.
A stronger approach is to train/measure a router on your actual traffic:
features:
task type
input length
requested output length
tool requirements
previous failure
cheap-model confidence
complexity indicators
↓
routing decision
cheap / medium / premiumMeasure quality per dollar, not just dollars per request.
Caching has economics.
If a prefix is rarely reused, paying to create/store it may provide little benefit. Some providers also have different cache-write/read economics. For example, Google's Vertex AI documentation for Claude describes cache reads as substantially cheaper than normal input while cache writes carry additional cost.
So separate your prompt into:
HIGH REUSE
↓
cache
LOW REUSE
↓
normal inputA 20k-token company policy used 100 times is an excellent cache candidate.
A 20k-token document used once isn't.
People often focus on input caching while ignoring output.
If your model generates 2,000 tokens when 300 would do, no amount of input optimization fixes that.
Use:
max_output_tokensOpenAI's current guidance notes that actual cost depends on input, cached input, output, and—where applicable—reasoning tokens, and that a cheaper per-token model doesn't necessarily mean a cheaper completed task.
This is the metric I'd build around:
cost_per_successful_tasknot:
cost_per_API_callFor example:
Cheap model:
$0.002/request
92% success
→ roughly $0.00217 per successful request
Premium:
$0.015/request
99% success
→ roughly $0.01515 per successful requestIf the cheap model can handle most requests and premium handles the difficult tail, you can get much closer to the first number without accepting the cheap model's failure rate across everything.
Also measure:
% requests routed cheap
% routed premium
cache hit rate
input tokens
cached tokens
output tokens
reasoning tokens
retry rate
escalation rate
successful-task rate
cost / successful taskI'd start with:
┌───────────────┐
│ Request │
└───────┬───────┘
↓
┌───────────────┐
│ Cache lookup │
└───────┬───────┘
↓
┌───────────────┐
│ Cheap router │
└───────┬───────┘
easy │ hard
┌───────┘ └────────┐
↓ ↓
Cheap model Premium model
│ │
└────────┬─────────┘
↓
quality / validation
↓
retry/escalateThe validation step is particularly useful: if you can cheaply detect bad answers, you don't have to pay premium-model prices for every request.
Don't create separate prompt templates for cheap and premium models if doing so destroys your cache prefix. Keep the large shared prefix identical, and vary only the model-specific instructions near the end.
That lets you combine:
stable prefix + caching + cheap-first routing + selective escalation.
As a rough strategy, I'd implement caching first, then measure your traffic and introduce routing based on the actual distribution of task difficulty. Current model pricing can differ dramatically between ordinary, cached, and output tokens, so calculate using your provider's current rates rather than assuming a fixed percentage saving.
If you tell me which API/provider you're using (OpenAI, Anthropic, Gemini, or a multi-provider setup) and roughly how many requests/tokens you process per month, I can sketch a concrete router and estimate the savings.
The biggest savings usually come from stacking two ideas:
Prompt caching can reduce input-token costs dramatically when the prefix repeats; OpenAI says its caching can reduce input costs by up to 90% on supported workloads.
Think of every request as:
[STABLE PREFIX]
system instructions
tool definitions
output schema
few-shot examples
stable reference material
[VARIABLE SUFFIX]
conversation-specific context
retrieved documents
user requestKeep the stable portion byte-for-byte identical and at the beginning. Put changing material at the end. Prefix-based caching works on exact matches, so even small changes in the cached prefix can hurt your hit rate.
For an agent, this is especially valuable because every loop may resend the system prompt, tools and previous context. Anthropic's current cost guidance likewise identifies prompt caching as one of the biggest levers for agent workloads.
Don't assume caching is working. Log something like:
request_id
model
input_tokens
cached_input_tokens
output_tokens
cache_hit_rate
latency
cost
task_type
quality_scoreA useful metric is:
cache_hit_rate = cached_input_tokens / total_input_tokensIf your agent has a huge stable prefix but a low hit rate, investigate what is changing between requests.
Don't ask, "Which model is best?"
Ask:
"What's the cheapest model that has a high probability of succeeding on this particular request?" A practical architecture is:
┌── cheap model ── success ──> return
Incoming request ──┤
└── uncertain/failure ──> premium model ──> returnFor example:
def answer(request):
classification = cheap_model.classify(request)
if classification == "simple":
return cheap_model.answer(request)
if classification == "complex":
return premium_model.answer(request)
result = cheap_model.answer(request)
if not passes_quality_check(result):
result = premium_model.answer(request)
return resultThe important part is not simply routing by request length. A 500-token question can require difficult reasoning, while a 20,000-token document may only need straightforward extraction.
Good routing signals include:
A simple policy might be:
| Task | Default |
|---|---|
| Classification | Cheap |
| Simple extraction | Cheap |
| Rewrite/summarize | Cheap |
| FAQ | Cheap |
| Basic coding | Cheap |
| Ambiguous reasoning | Premium |
| Difficult debugging |
The exact models depend on your provider and benchmark results.
This is where routing becomes much more powerful.
Instead of:
20% premium / 80% cheapuse:
100% cheap initially
↓
confidence / validator
↓
~10–30% escalatedFor example, have the inexpensive model return:
{
"answer": "...",
"confidence": 0.91,
"needs_escalation": false,
"reason": "Straightforward extraction"
}Then escalate when:
confidence < 0.80
OR
validator_failed
OR
required_tool_failed
OR
task_type == "hard_reasoning"Be careful with self-reported confidence, though. A cheap model saying "I'm 99% confident" isn't itself evidence. A separate deterministic validator, test suite, schema checker, retrieval check, or small judge is generally stronger.
This is an easy trap.
Your bill is roughly:
cost =
uncached_input_tokens × input_price
+ cached_input_tokens × cached_price
+ output/reasoning_tokens × output_priceOpenAI's current documentation explicitly separates input, cached-input, and output tokens, and reasoning tokens can contribute to output-token usage.
So if your premium model generates huge amounts of reasoning/output, routing more traffic away from it can matter much more than shaving a few cents from input.
For example, suppose:
Cheap model:
$0.50 / 1M input
$2 / 1M output
Premium:
$5 / 1M input
$30 / 1M outputIf the premium model produces 10× as many output tokens, its effective cost difference can be much larger than the advertised input-price difference.
There's an important interaction:
Don't blindly route every request between providers if doing so destroys cache locality.
A useful architecture is:
session ID
│
▼
┌───────────────┐
│ Sticky router │
└───────┬───────┘
│
┌──────────┴──────────┐
▼ ▼
Cheap model Premium model
│ │
└──── same provider/session
│
▼
warm cacheSticky routing can preserve a provider's warm cache across turns; this is particularly useful for multi-turn agents with large repeated prefixes.
So your router should optimize cost × quality × cache locality, not just model price.
You can potentially have three different caches:
Reuse the stable prompt prefix.
system + tools + schemas + stable instructionsIf users ask essentially the same question, reuse the completed answer.
"What's our refund policy?"
"What is your refund policy?"could map to the same cached response if freshness requirements permit.
Cache expensive retrieval operations:
query → retrieved document IDsrather than repeatedly searching the same corpus.
This can save both LLM tokens and tool/API costs.
Give each request a budget:
budget = {
"max_total_cost": 0.03,
"max_premium_calls": 1,
"max_iterations": 4,
"max_output_tokens": 1000,
}Then enforce:
cheap → validate → premium → stoprather than allowing an agent to loop indefinitely.
Also set per-customer/project limits and monitor cost per successful outcome, not just cost per request. Current industry guidance increasingly treats AI spend as a usage/FinOps problem rather than merely a model-selection problem.
If I were building this from scratch, I'd start with:
Request
│
▼
┌────────────────┐
│ Request router │
└───────┬────────┘
│
┌─────────┴─────────┐
│ │
simple/high-cache difficult
│ │
▼ ▼
cheap model premium model
│ │
└─────────┬─────────┘
▼
validator
│
┌─────┴─────┐
│ │
pass fail
│ │
▼ ▼
return premiumAnd I'd optimize in this order:
That combination can be substantially more effective than simply switching everything to a cheaper model—and it avoids sacrificing quality where the expensive model actually earns its keep. Anthropic's current optimization guidance similarly recommends caching/token hygiene first, then model selection and multi-model architectures.
If you tell me which API/provider you're using (OpenAI, Anthropic, Gemini, OpenRouter, etc.) and roughly how your workload looks, I can lay out a concrete routing + caching architecture with example code and a break-even calculation.
Instead of:
20% premium / 80% cheapuse:
100% cheap initially
↓
confidence / validator
↓
~10–30% escalatedFor example, have the inexpensive model return:
{
"answer": "...",
"confidence": 0.91,
"needs_escalation": false,
"reason": "Straightforward extraction"
}Then escalate when:
confidence < 0.80
OR
validator_failed
OR
required_tool_failed
OR
task_type == "hard_reasoning"Be careful with self-reported confidence, though. A cheap model saying "I'm 99% confident" isn't itself evidence. A separate deterministic validator, test suite, schema checker, retrieval check, or small judge is generally stronger.
This is an easy trap.
Your bill is roughly:
cost =
uncached_input_tokens × input_price
+ cached_input_tokens × cached_price
+ output/reasoning_tokens × output_priceOpenAI's current documentation explicitly separates input, cached-input, and output tokens, and reasoning tokens can contribute to output-token usage.
So if your premium model generates huge amounts of reasoning/output, routing more traffic away from it can matter much more than shaving a few cents from input.
For example, suppose:
Cutting your AI inference bill by combining prompt caching and smart routing can slash production costs by 60–80%.
Instead of sending every request to an expensive frontier model (like Claude 3.5 Sonnet or GPT-4o), build or use a routing layer that assesses query complexity.
Prompt caching stores the precomputed key-value states of static context (system prompts, few-shot examples, large reference documents) in GPU memory.
Make your routing layer cache-aware so that savings compound:
If you'd like to dive deeper, let me know:
I can provide a code snippet or architectural layout tailored to your stack.
Structure requests like:
[STATIC]
system instructions
tool definitions
schemas
few-shot examples
stable business rules
[DYNAMIC]
user request
retrieved documents
conversation-specific dataPut stable content first and changing content last. Prefix caching generally works on exact repeated prefixes; OpenAI's current documentation says cached input can be substantially cheaper, with discounts varying by model.
Also:
cached_tokens rather than assuming caching is working.Important: caching and routing interact. If you bounce every request among providers, you can lose the warm cache. Sticky/session-based routing can preserve it.
Don't ask "Which model is best?" Ask:
What's the cheapest model that passes my quality bar for this request? A practical cascade:
┌─ cheap model ──> confidence OK ──> return
request ──> router ─┤
└─ hard/uncertain ──> premium model ──> returnFor example:
| Request | Default |
|---|---|
| Classification | Cheap |
| Extraction / structured JSON | Cheap |
| Summarization | Cheap |
| Simple Q&A | Cheap |
| Routine coding | Cheap/mid |
| Ambiguous reasoning | Premium |
| Difficult debugging | Premium |
| High-stakes customer response | Premium |
| Complex multi-step planning | Premium |
Don't route solely from the user's words. Route based on measured task difficulty.
Have the cheap model return something like:
{
"answer": "...",
"confidence": 0.91,
"needs_escalation": false
}Then escalate when:
if (
confidence < 0.80
or missing_required_fields
or failed_validation
or tool_call_is_invalid
or task_is_known_hard
):
use_premium()
else:
return_answer()Even better, add deterministic validators before paying for another model:
cheap model
↓
JSON/schema validation
↓
business-rule validation
↓
confidence check
↓
PASS → return
FAIL → premiumThis prevents spending premium-model tokens on failures that a $0 validator can catch.
A $0.20/M model isn't necessarily cheaper if it succeeds only 80% of the time and 20% of requests need a second attempt.
Track:
cost/request
cache-hit %
cheap-model %
premium-escalation %
failure/retry %
tokens/request
latency
quality/pass rate
cost per successful taskThe last metric is the one management actually cares about.
OpenAI explicitly cautions that lower per-token pricing doesn't necessarily mean lower total cost because models can tokenize and generate differently.
Suppose your current workload looks like:
100 requests
× 20k input tokens
× premium modelInstead, aim for:
100 requests
├─ 80 cheap-model requests
│ └─ 70% cache hits
└─ 20 premium escalations
└─ 70% cache hitsYou are attacking three different cost centers:
That combination is much more powerful than simply switching to a cheaper model.
I'd implement this as four components:
┌──────────────┐
│ Cost / cache │
│ telemetry │
└──────┬───────┘
│
Request → Cache/session router
│
┌──────▼───────┐
│ Cheap model │
└──────┬───────┘
│
┌──────────▼──────────┐
│ Validator + quality │
│ classifier │
└──────┬────────┬─────┘
│ │
PASS FAIL
│ │
▼ ▼
Return Premium
│
▼
ReturnOne particularly important optimization: don't send the entire conversation history to the premium model if you don't have to. Have the cheap model produce a compact task state/summary and give the premium model only the information it actually needs.
And if you're using OpenAI specifically, current pricing illustrates why this matters: cached input can be dramatically cheaper than uncached input, while premium output tokens can be many times more expensive than input tokens.
If you do only one thing, don't start by changing models. First measure your cache-hit rate and what percentage of requests genuinely require the premium model. That's usually where the biggest savings become visible.
┌──────────────┐
│ Cost / cache │
│ telemetry │
└──────┬───────┘
│
Request → Cache/session router
│
┌──────▼───────┐
│ Cheap model │
└──────┬───────┘
│
┌──────────▼──────────┐
│ Validator + quality │
│ classifier │
└──────┬────────┬─────┘
│ │
PASS FAIL
│ │
▼ ▼
Return Premium
│
▼
ReturnOne particularly important optimization: don't send the entire conversation history to the premium model if you don't have to. Have the cheap model produce a compact task state/summary and give the premium model only the information it actually needs.
And if you're using OpenAI specifically, current pricing illustrates why this matters: cached input can be dramatically cheaper than uncached input, while premium output tokens can be many times more expensive than input tokens.
Combining prompt caching and smart model routing is the single fastest way to slash your AI inference bill by 60$-80% without sacrificing output quality.
If your application uses massive system prompts, static few-shot examples, or long conversation histories, you are paying to re-process identical input tokens on every single API call. Caching stores the "prefill" phase so you pay a fraction of the cost for repeated prefixes.
cache_control breakpoints, while others like OpenAI, Gemini, or DeepSeek handle prefix caching automatically once you cross specific token thresholds like 1 024 tokens).Not every query requires a frontier model (GPT$-4o or Claude 3.5 Sonnet). Up to 70% of standard production tasks—like simple classification, basic extraction, formatting, or straightforward chit-chat—can be handled identically by cheap mini/flash models.
When you route a repetitive, heavy-context task to a cheap model and cache that context, your marginal cost drops near zero.
max_tokens limit on your API calls to prevent runaway generation loops on both cheap and expensive tiers.If you want, let me know:
I can help you outline a more specific routing architecture or caching strategy for your stack.
For example, OpenAI's current token pricing shows how large the spread can be: GPT-5.6 Terra is $2/M input and $12/M output, while GPT-5.6 Sol is $4/M input and $20/M output; cached input is much cheaper still at $0.20/M and $0.40/M respectively.
Think of your prompt as:
[STABLE PREFIX]
system instructions
tool definitions
company policy
few-shot examples
RAG/static reference material
[DYNAMIC SUFFIX]
user's current question
current retrieved documents
conversation-specific informationPut stable material first and changing material last.
Provider prompt caching generally works on repeated prefixes. OpenAI, for example, automatically caches supported prompts once they exceed 1,024 tokens, with longer matching prefixes receiving the cached-input rate.
So instead of:
System: You are...
User: Here's today's customer data...
System: Company policy...
User: Answer this...prefer:
System:
stable instructions
company policy
tools
examples
User:
today's customer data
current questionThat makes the expensive part reusable.
Track cached_tokens in your telemetry. Otherwise you won't know whether your caching strategy is actually working.
Build a simple router:
┌─> cheap model ──> answer
Request ─> classifier
└─> premium model ─> answerStart with rules before building an elaborate ML router.
For example:
| Request | Model |
|---|---|
| Classification | Cheap |
| Extraction / JSON | Cheap |
| Summarization | Cheap |
| Simple Q&A | Cheap |
| Simple coding | Mid-tier |
| Complex debugging | Premium |
| Difficult reasoning | Premium |
| High-value customer response | Premium |
| Failed cheap-model attempt | Premium |
A good rule is:
Use the cheapest model that passes your quality threshold, not the cheapest model available. For instance, GPT-5.6 Terra is explicitly positioned as balancing intelligence and cost, while Sol targets higher-end capability.
Instead of deciding perfectly before generation:
cheap model
↓
confidence / quality check
↓
┌───────┴───────┐
pass fail
↓ ↓
return premium modelThis is powerful because most traffic is often relatively easy.
Your cheap model could return:
{
"answer": "...",
"confidence": 0.91,
"needs_escalation": false
}Then escalate when:
Don't blindly trust a model's self-reported confidence, though. Ideally calibrate the router against a labeled evaluation set.
There are actually two different caches:
The model still runs, but repeated input tokens are cheaper.
same long system prompt
↓
model runs
↓
cached input discountYou avoid the model entirely when a sufficiently similar question already has a good answer.
"What's your refund policy?"
↓
embedding / similarity search
↓
existing answer?
↓
YES → return cached answerThis can produce much larger savings for repetitive workloads.
Use semantic caching carefully for things like FAQs and documentation. Avoid it for queries where freshness, personalization, permissions, or rapidly changing data matter.
Don't just cache the LLM.
For RAG systems, consider:
User query
↓
query normalization
↓
semantic cache
↓ miss
retrieval cache
↓
LLM prompt cache
↓
cheap model
↓
premium escalation if necessaryIf you're repeatedly searching the same documents, doing the same database queries, or constructing the same context, those operations can become a hidden part of your bill and latency.
Give every request a cost/quality budget:
if task.type in SIMPLE_TASKS:
model = "cheap"
elif task.risk == "high":
model = "premium"
elif task.complexity > 0.8:
model = "premium"
else:
model = "cheap"Then add per-customer or per-workflow limits:
monthly budget
↓
remaining budget?
├── plenty → normal routing
├── moderate → prefer cheap
└── nearly exhausted → cheap / cached onlyThis prevents a small number of pathological agent runs from consuming your entire budget.
Don't optimize solely for $ / 1M tokens.
Track:
That last metric is particularly important.
A model that costs $0.01 but succeeds 70% of the time can be more expensive than one costing $0.02 that succeeds 98% of the time.
OpenAI likewise notes that comparing models purely by token price can be misleading because different models may tokenize inputs differently and generate different amounts of reasoning/output.
I'd implement this first:
┌──────────────────┐
│ Semantic cache │
└────────┬─────────┘
│ miss
▼
┌──────────────────┐
│ Request router │
└────────┬─────────┘
│
┌───────────────┼────────────────┐
▼ ▼ ▼
cheap medium premium
│ │ │
└───────────────┼────────────────┘
▼
┌──────────────────┐
│ Quality checker │
└────────┬─────────┘
│ fail
▼
premium retry
All requests:
└── stable prompt prefix → provider prompt cache1. Instrument first. Get per-request token/model/cost/quality data.
2. Fix prompt layout. Move stable context to the beginning and dynamic content to the end.
3. Enable/verify provider caching. Measure your actual cache-hit percentage.
4. Introduce cheap-model routing. Start with obvious low-complexity tasks.
5. Add escalation. Let failures move upward to the premium model.
6. Add semantic caching. Only after you understand which requests repeat.
7. Optimize prompts/context/output. Reducing tokens multiplies the savings from everything above.
As a rough mental model, your final bill becomes:
Total cost ≈ cache misses × cheap-model cost + escalations × premium-model cost + uncached context + non-LLM infrastructure
The sweet spot is therefore high cache reuse + high cheap-model coverage + low escalation rate, while maintaining your required quality. Current industry guidance similarly treats caching, routing/cascading, and context reduction as complementary rather than competing optimizations.
If you tell me which providers/models you're using and roughly how many requests/tokens you process per month, I can sketch a concrete routing policy and calculate the potential monthly savings.
| High-value customer interaction | Premium |
| Cheap-model retry/failure | Premium |
| High-stakes reasoning or difficult planning |
| Premium |
| Cheap model expresses uncertainty | Escalate |
| Premium |
| Complex planning | Premium |
| High-stakes answer | Premium + verification |