ToolsExplainer
Semantic caching for LLMs: how it works and where it goes wrong
Semantic caching replays a stored LLM answer when a new prompt means the same thing. How thresholds, vector stores and cache keys work, how it differs from prompt caching, and how to measure savings.

Semantic caching is a technique that stores LLM responses against a vector embedding of the prompt and replays a stored response when a new prompt is close enough in meaning to an old one. It targets the cost of repeated questions: when a thousand users ask how to reset a password in a thousand different phrasings, a semantic cache can answer most of them without calling the model. Bifrost, an open-source AI gateway written in Go by Maxim AI, implements it as a dual-layer cache at the gateway, and serves as the worked example throughout this explainer.
The technique is simple to describe and easy to get wrong. A similarity score is not a correctness check, a shared cache is a data-leakage path, and a semantic lookup makes every cache miss slower. This piece covers how a semantic cache LLM layer works, how it differs from exact-match caching and from the prompt caching Anthropic and OpenAI run by default, where it fails, what the research says, and how to measure savings honestly.
What semantic caching is, and what it is not
A semantic cache is a lookup table whose keys are meanings rather than strings. Each stored entry holds three things: the embedding vector of the original prompt, the response the model gave, and metadata such as an expiry time and a partition key. A new prompt is embedded with the same model, the nearest stored vector is found, and if the cosine similarity between the two clears a configured threshold, the stored response is returned instead of calling the provider.
That makes semantic caching approximate reuse: the replayed response was generated for a different prompt that the cache judged equivalent. Everything that goes wrong follows from that judgement being a number rather than a guarantee.
Three kinds of LLM caching are often lumped together, and they behave differently:
- Exact-match caching hashes the full request (model, parameters, messages) and replays a stored response only when the hash matches. It never serves a wrong answer for a different question, but it misses every paraphrase.
- Semantic caching replays a stored response for a prompt that is similar in meaning. It catches paraphrases and typos, at the cost of occasionally serving an answer to a question that was not asked.
- Provider prompt caching is run by Anthropic, OpenAI, Google and others on their own infrastructure. It reuses the processed form of an identical prompt prefix. The model still generates a fresh response, and output is billed in full.
Only the first two avoid a model call, and some gateways run them together as a dual-layer cache that tries the exact path first.
How a semantic cache works
A semantic cache sits in front of the model call, in the application or a gateway. Each request passes up to five steps: an eligibility check, an exact-match lookup, an embedding, a nearest-neighbour search, and either a replay or a provider call whose result is stored.
Embedding the prompt
The prompt is converted into a vector by an embedding model such as OpenAI’s text-embedding-3-small, which produces 1,536 dimensions and costs $0.02 per million tokens on OpenAI’s pricing page as of September 2026. The embedding model matters: prompts a human would treat as identical can score below the threshold, and prompts with opposite meanings can score above it, depending on how the model represents negation, numbers and names.
What to embed is a design choice. Embedding a long shared system prompt makes every request look alike; embedding only the latest message ignores context that changes the answer. Most implementations can exclude the system prompt and cap conversation history.
The similarity threshold
The threshold is the cosine similarity above which two prompts are treated as the same question. Both GPTCache, the open-source semantic cache from Zilliz, and Bifrost ship with a default of 0.8. A higher threshold serves fewer, safer hits; a lower one serves more hits and more wrong answers. There is no threshold that is correct for all workloads, a point the research section returns to.
The vector store
Stored vectors live in a database with approximate nearest-neighbour search. Redis or Valkey keeps the lookup in memory; Qdrant, Weaviate and Pinecone scale to larger caches. The store must support per-entry expiry and filtering by partition key, or the cache cannot enforce freshness or isolation.
TTL and invalidation
Every entry needs a time to live. A cached refund-policy answer is correct until the policy changes; an order-status answer is wrong within minutes. A production cache also needs on-demand deletion, by entry ID or by partition, for the day a cached answer turns out to be wrong.
Cache keys and partitions
The cache key decides which requests are allowed to share answers. A good implementation scopes each lookup by the model and provider (a GPT answer should not be replayed for a Claude request), by generation parameters where they change the output, and by a partition key chosen by the operator: tenant, user, session or feature. The partition key is the difference between a cache that saves money and one that leaks data.

Figure 1: A dual-layer cache pays for an embedding only when the cheap exact-match lookup misses. Every semantic miss costs an embedding call on top of the full model call.
Prompt caching vs semantic caching vs exact-match caching
Prompt caching and semantic caching solve different problems. Prompt caching cuts the cost of a long, repeated prefix (system prompt, tool schemas, reference documents) on requests that each need a fresh answer. Semantic caching cuts the cost of the whole request when the answer can be reused. Their savings stack, and a fair evaluation of either assumes the other is already on.
The provider mechanics, read from each vendor’s documentation in September 2026:
- Anthropic. Claude prompt caching bills cache reads at 0.1 times the base input price on most models, 0.05 times on Claude Opus 5.5 and 0.025 times on Claude Fable 5.1. Writing to the default 5-minute cache costs 1.25 times the input price; the 1-hour cache costs 2 times. The minimum cacheable prefix is 512 to 4,096 tokens depending on the model, and a shorter prompt is processed without caching and without an error. Hits require “100% identical prompt segments”, and caches are isolated per organisation, and per workspace on the Claude API.
- OpenAI. OpenAI prompt caching is on by default for supported models and discounts cached input “up to 90%”. For GPT-5.6 and later the minimum prefix is 1,024 tokens, cache writes cost 1.25 times the uncached rate, reads cost 0.1 times, and entries live at least 30 minutes after the latest write or reuse. Earlier models use a
prompt_cache_retentionsetting ofin_memory(typically 5 to 10 minutes of inactivity) or24h. Caches “are not shared across organizations”.
| Exact-match cache | Semantic cache | Provider prompt cache | |
|---|---|---|---|
| Who runs it | Application or gateway | Application or gateway | Model provider |
| What is reused | The whole response | The whole response | Processed prompt prefix |
| Match rule | Identical request hash | Cosine similarity ≥ threshold | Identical prefix, above a minimum length |
| Model call on hit | None | None | Yes, output generated fresh |
| Cost on hit | Storage lookup | Embedding call plus lookup | Cached input at about 0.1x; output at full price |
| Can serve a wrong answer | No, only a stale one | Yes | No |
| Isolation | Whatever the key includes | Whatever the key includes | Organisation or workspace |
| Typical lifetime | Operator-set TTL | Operator-set TTL | 5 minutes to 24 hours, provider-set |
The cost gap on a hit is large. A semantic hit costs an embedding call, $0.00003 for a 1,500-token prompt on text-embedding-3-small. A prompt-cache hit still pays for output tokens, which on most models cost five times the input rate, and output usually dominates the bill. The inference price analysis on this site walks through why caching and batch discounts now move a bill more than headline price cuts.
The gateway layer can manage both. Bifrost’s semantic cache replays responses, and its auto prompt caching setting injects provider cache breakpoints for clients that send none, such as agentic tools that never mark a cache_control block. The two settings are independent and can both be on.

Figure 2: Exact and semantic caches skip the model; the provider’s prompt cache only makes the input cheaper. The semantic lane is the only one that can return an answer to a different question.
Where semantic caching goes wrong
Semantic caching fails in predictable ways. Most mitigations reduce the hit rate, which is why hit-rate claims should be read alongside the error rate they imply.
False hits
A false hit is a cached answer served for a prompt that needed a different answer. Embedding models are trained to place related text close together, and related is not the same as equivalent. An open issue on the GPTCache repository, filed in August 2026, reports a cosine similarity of 0.96 between “Withhold the study drug from any participant who reports chest tightness” and “Administer the study drug to any participant who reports chest tightness”, comfortably above the 0.8 default. The reporter also notes that genuinely equivalent pairs with different wording can fall below 0.8. Negation, quantities, dates, product names and units are the usual culprits.
Threshold tuning is workload-specific
A threshold tuned on customer-support questions will be wrong for code questions. The category-aware caching paper observes that code queries cluster densely in embedding space while conversational ones spread out, so a fixed threshold produces false positives in dense regions and misses valid paraphrases in sparse ones. Tuning needs labelled pairs from real traffic: pairs that should share an answer and pairs that should not, scored with the production embedding model.
Personalised and time-sensitive answers
Any response that depends on who is asking, or when, is unsafe to share. “What is my balance?” embeds almost identically for every user; “What is the outage status?” changes hourly; answers built from retrieved private documents carry their reader’s permissions. These need a per-user partition, a TTL in minutes, or no cache.
Multi-turn conversations and agents
“And for the enterprise plan?” means nothing without the previous turn, and an agent’s tenth tool call depends on the nine before it. Caches either embed more history, which makes matches rarer, or refuse to cache past a set conversation length. The second is usually the right default.
Data leakage across tenants
A semantic cache without a partition key is a shared memory across every caller. The provider-side analogue has been measured: Gu and colleagues at Stanford used timing differences to detect prompt-cache sharing across users in seven API providers, including OpenAI, at the time of their early-2025 audit; both Anthropic and OpenAI now document organisation-level isolation. Timing leaks apply to semantic caches too, since a hit returns far faster than a miss. InputSnatch demonstrated reconstructing other users’ inputs from response-time differences across several cache mechanisms.
Streaming and latency on misses
A semantic lookup cannot start searching until the prompt is embedded, which is typically tens to a few hundred milliseconds. On a hit that is still far faster than a multi-second generation. On a miss it is pure overhead added before the first token of the real response, so a cache with a low hit rate makes the median request slower. Replaying a cached answer as a stream is straightforward; hiding the embedding latency on misses is not.
Adversarial collisions
Semantic caches were not designed to resist attackers. Zhang and colleagues, in a paper accepted to ICML 2026, built a black-box attack called CacheAttack that crafts prompts to collide with a target cached entry, and report an 86 percent hit rate in hijacking LLM responses, including in a financial-agent workflow. Their argument is structural: the locality that makes similar prompts collide on purpose is the opposite of the avalanche property that makes cryptographic hashes collision-resistant. Per-tenant partitions, which some gateways enforce by making the cache key a required field, limit the blast radius; they do not remove the attack within a tenant.

Figure 3: Most semantic caching incidents come from caching a request that should never have been eligible, not from a badly chosen threshold.
What the research says about semantic caching
The academic record on semantic caching is short and consistent. Early systems showed the speed-up; later work showed that static thresholds are the weak point and that the cache is an attack surface.
| Work | Year | Finding |
|---|---|---|
| GPTCache (Bang, NLP-OSS workshop) | 2023 | Open-source semantic cache; reports 2 to 10 times faster responses on a cache hit |
| MeanCache (Gill et al., IPDPS 2025) | 2024 | Repeated queries are “about 31%” of the total; federated per-user cache gives 17% higher F-score, 20% higher precision, 83% less storage |
| vCache (Schroeder et al.) | 2025 | Static thresholds give no correctness guarantee; per-prompt learned thresholds reach up to 12.5x higher hit rate and 26x lower error |
| Category-aware caching (Wang et al.) | 2025 | Hit rates of 40 to 60% for high-repetition categories and 5 to 15% for low ones; search latency cut from 30 ms to 2 ms with in-memory HNSW |
| CacheAttack (Zhang et al., ICML 2026) | 2026 | Black-box collision attack hijacks cached responses 86% of the time |
| Prompt-cache audit (Gu et al., ICML 2025) | 2025 | Timing audits found cross-user cache sharing in seven API providers |
Three conclusions follow. The 31 percent repeat rate MeanCache cites is unsourced in the abstract and an upper bound on what any cache can catch, so it is not a planning figure. vCache’s result makes a single global threshold the weakest part of most deployments; per-category or per-entry thresholds are where the research is heading, and gateways such as Bifrost already accept a per-request threshold override. And the attack papers put the cache key in the threat model, not just the cost model. On hit rates, 5 to 15 percent is a sensible floor for open-ended assistants; a vendor figure quoted without its workload is not evidence.
Semantic caching at the gateway: Bifrost as a worked example
A gateway is the natural place for a semantic cache: it sees every request, knows the model and provider, and holds the credentials for an embedding call. The explainer on AI gateways covers the broader role.
Bifrost implements semantic caching as a plugin named semantic_cache with two lookup paths. Direct mode hashes the normalised input, parameters and stream flag and replays only exact matches, with no embedding provider needed. Direct plus semantic mode adds embedding similarity on a direct miss. Running direct-only is a supported configuration, which makes it possible to start with zero false-hit risk and add the semantic layer later.
The configuration surface
The plugin’s fields map closely to the design choices above:
{
"name": "semantic_cache",
"enabled": true,
"config": {
"provider": "openai",
"embedding_model": "text-embedding-3-small",
"dimension": 1536,
"ttl": "5m",
"threshold": 0.8,
"conversation_history_threshold": 3,
"exclude_system_prompt": false,
"cache_by_model": true,
"cache_by_provider": true,
"vector_store_namespace": "BifrostSemanticCachePlugin",
"default_cache_key": ""
}
}
The defaults are a 5-minute TTL and a 0.8 threshold. conversation_history_threshold skips caching for conversations with more than three messages by default, which handles the multi-turn problem by declining it. exclude_system_prompt keeps a long shared system prompt from dominating the key. cache_by_model and cache_by_provider are on by default, so a response from one model is never replayed for another.
Storage runs on the Bifrost vector store layer, which supports Redis and Valkey, Weaviate, Qdrant and Pinecone. Redis or Valkey is the recommended store for direct-only mode, because the other three require a vector for every entry.
Cache keys and per-request headers
Caching in Bifrost is opt-in per request. A request is cached only if it carries an x-bf-cache-key header or the operator sets a default_cache_key; without either, the request bypasses the cache. Entries are matched only within the same key, so a request under tenant-A is never served a response cached under tenant-B, even for an identical prompt. That turns the partition decision from an afterthought into a required field.
Per-request headers override the defaults:
| Header | Effect |
|---|---|
x-bf-cache-key | Partition for lookup and storage; required unless a default key is set |
x-bf-cache-ttl | TTL for this request, as a duration or seconds |
x-bf-cache-threshold | Similarity threshold for this request, clamped to 0 to 1 |
x-bf-cache-type | Restrict lookup to direct or semantic |
x-bf-cache-no-store | Serve hits but do not store this response |
The threshold header makes careful rollout practical: run most traffic at a strict 0.95, route a labelled sample at 0.85, and compare error rates before changing the default. A time-sensitive endpoint can send x-bf-cache-ttl: 30s; a personalised one can use a per-user key.
Observability and invalidation
Every cache-checked response carries a cache_debug object with cache_hit, hit_type (direct or semantic), the cache_id, and on semantic hits the threshold used, the actual similarity score, and the embedding tokens consumed. Logs tag hits with a Direct Cache or Semantic Cache badge, and the Prometheus exporter publishes bifrost_cache_hits_total by hit type alongside bifrost_cost_total. Entries can be deleted by cache_id or by cache key through the API. Streamed responses are cached and replayed chunk by chunk, with the debug payload on the final chunk.
The latency trade-off is stated plainly in the configuration guidance: a semantic hit costs roughly an embedding round trip, and a semantic miss pays the embedding call on top of the full LLM call, which makes it slower than running without the cache. Writes happen asynchronously and add no latency. That warning is the right starting assumption for any semantic cache, whichever implementation is chosen.
Why the gateway layer is the author’s pick
Of the ways to add a semantic cache (a library such as GPTCache in each service, a managed cache from a vector database vendor, or a gateway plugin), the gateway is the better default, and Bifrost is the strongest implementation reviewed here: the dual-layer design allows an exact-only start, the mandatory cache key makes tenant isolation the default, per-request headers vary thresholds and TTLs by endpoint without redeploying, the debug fields expose similarity scores for tuning, and the same gateway handles prompt caching, failover and budgets. For readers comparing gateways more broadly, a comparison of production LLM gateways and a survey of open-source LLM gateways you can self-host both cover caching alongside routing and governance, as does this site’s five-gateway comparison.
Caching sits inside a wider control plane. Beyond routing and caching, Bifrost applies governance and security controls centrally (virtual keys, budgets, guardrails and audit logs), and Bifrost Edge, currently in alpha, extends that same governance and security to AI traffic on employee machines, with endpoint enforcement of those guardrails on each device.
How to measure hit rate and savings honestly
LLM cost optimisation claims for semantic caching are easy to inflate. The honest version measures correct hits against a realistic baseline and charges the cache for everything it costs.
Five rules for an honest number
- Count correct hits, not hits. Sample served semantic hits, regenerate a fresh answer for the same prompt, and have a reviewer or a judge model decide whether the cached answer was acceptable. Report the correct-hit rate and the error rate together.
- Split direct and semantic hits. Exact-match hits carry no correctness risk and often come from retries, double-submits and load tests. Semantic hits are where both the savings and the errors are. Bifrost’s
hit_typefield and per-type metric make the split straightforward. - Charge the cache for its costs. Every direct miss triggers an embedding call. Every semantic miss adds latency. Subtract embedding spend from savings and report p50 and p95 latency for misses, not just hits.
- Use a prompt-cached baseline. Comparing a semantic cache against uncached traffic overstates savings on any workload with a long stable prefix, because the provider would have discounted that prefix anyway.
- Weight by cost, not by request count. A cache that catches short, cheap questions and misses long, expensive ones can report a 30 percent hit rate and save 10 percent of spend.
A worked example
Consider a customer-support assistant on Claude Sonnet 5, priced by Anthropic in September 2026 at $2 per million input tokens, $10 per million output tokens and $0.20 per million cached-input tokens. It handles 100,000 requests a day. Each request sends 1,500 input tokens, of which 1,200 are a stable system prompt, and receives 400 output tokens. Embeddings use text-embedding-3-small at $0.02 per million tokens, embedding the full 1,500 tokens on every request as a conservative assumption.
| Setup (Sep 2026 prices) | Model cost per request | Embedding per day | Per day | Per 30 days |
|---|---|---|---|---|
| No caching | $0.00700 | none | $700.00 | $21,000 |
| Provider prompt caching only | $0.00484 | none | $484.00 | $14,520 |
| Prompt caching + semantic cache, 20% correct hits | $0.00484 on 80,000 misses | $3.00 | $390.20 | $11,706 |
| Prompt caching + semantic cache, 40% correct hits | $0.00484 on 60,000 misses | $3.00 | $293.40 | $8,802 |
The prompt-cached rows assume the 1,200-token prefix is read at $0.20 per million on nearly every request; at over one request a second the 5-minute cache stays warm and writes cost under a dollar a day.
Three things stand out. Provider prompt caching alone saves $6,480 a month and carries no correctness risk, so it comes first. The semantic cache’s incremental saving over that baseline is $2,814 a month at a 20 percent correct-hit rate and $5,718 at 40 percent, which is 19 and 39 percent of the prompt-cached bill. The embedding line is negligible at $90 a month; the real cost is elsewhere. At a 20 percent hit rate, if one semantic hit in thirty is wrong, the assistant serves about 670 wrong answers a day. Whether $94 a day is worth that depends on what a wrong answer costs, which the breakdown of what an AI product really costs on this site treats as a line item, not a rounding error.
A rollout plan for semantic caching
Add semantic caching once provider prompt caching is in place and traffic shows repeated intent.
- Turn on provider prompt caching first. Put the stable content at the front of every request and check the provider’s minimum prefix length. It is the cheapest saving available and cannot return a wrong answer.
- Start with exact-match caching. Enable a direct-only cache with short TTLs. It catches retries and duplicate submissions and gives a baseline hit rate with no false hits.
- Classify endpoints before enabling semantic lookup. Mark each route as shareable, per-user, time-sensitive or uncacheable. Only the first category gets a shared semantic cache key.
- Log similarity scores before serving semantic hits. Collect nearest-neighbour scores on real traffic and build a labelled set of pairs that should and should not share an answer. Choose the threshold from that set, not from a default.
- Launch strict, then relax. Start at a high threshold such as 0.95 on a small share of traffic, measure the correct-hit rate, and lower it only while the error rate stays inside an agreed budget.
- Treat the cache key as a security control. Partition by tenant at minimum, include the model and provider, and review the partition scheme when a new endpoint is added.
- Report savings the honest way. Correct hits, cost-weighted, net of embeddings, against a prompt-cached baseline, with miss latency alongside.
Teams evaluating gateway-layer semantic caching can request a Bifrost demo or review the Bifrost source on GitHub and the leading AI gateways for production side by side. Those running their own infrastructure may prefer the self-hosted AI gateway options roundup. Either way, the threshold is a product decision about acceptable error, made with labelled data rather than a default.
Feature claims are drawn from each vendor’s public documentation as of September 2026.
Sources
- Bifrost semantic caching Maxim AI
- Bifrost auto prompt caching Maxim AI
- Bifrost vector store Maxim AI
- Bifrost Prometheus metrics Maxim AI
- Bifrost Edge overview Maxim AI
- Prompt caching (Claude API) Anthropic
- Prompt caching (OpenAI API) OpenAI
- OpenAI API pricing OpenAI
- GPTCache: An Open-Source Semantic Cache for LLM Applications Enabling Faster Answers and Cost Savings NLP-OSS 2023 (ACL Anthology)
- GPTCache repository Zilliz
- MeanCache: User-Centric Semantic Caching for LLM Web Services arXiv (IPDPS 2025)
- vCache: Verified Semantic Prompt Caching arXiv
- Category-Aware Semantic Caching for Heterogeneous LLM Workloads arXiv
- From Similarity to Vulnerability: Key Collision Attack on LLM Semantic Caching arXiv (ICML 2026)
- Auditing Prompt Caching in Language Model APIs arXiv (ICML 2025)
- InputSnatch: Stealing Input in LLM Services via Timing Side-Channel Attacks arXiv
Frequently asked questions
What is semantic caching?
Semantic caching stores LLM responses alongside a vector embedding of the prompt that produced them. When a new prompt arrives, the cache embeds it, finds the most similar stored prompt, and returns that prompt's response if the similarity score clears a threshold. A hit skips the model call entirely, so it costs no model tokens and returns in milliseconds rather than seconds.
What is the difference between prompt caching and semantic caching?
Prompt caching is run by the model provider: an identical prompt prefix is reused, cached input is billed at a fraction of the normal rate, but the model still runs and output is billed in full. Semantic caching runs in the application or gateway and replays a previously generated answer for a similar prompt, so the provider is never called. The two are independent and can run together.
What similarity threshold should a semantic cache use?
There is no universal value. Common defaults sit around 0.8 cosine similarity, which is what Bifrost and GPTCache ship with, but the right number depends on the embedding model and the workload. Tune it on a labelled set of prompt pairs from real traffic, start stricter than the default, and track how often served hits are actually wrong.
Is semantic caching safe for multi-tenant applications?
Only if cache entries are partitioned by tenant. A cache keyed on prompt similarity alone will replay one customer's answer to another customer who asks something similar. Scope the cache key to the tenant, or to the user for personalised answers, and never cache responses that include account data, retrieved private documents or permission-dependent content under a shared key.
Does semantic caching work with streaming responses?
It can. A cached response can be replayed to the client as a stream, which Bifrost does chunk by chunk. The catch is latency on misses: a semantic lookup must embed the prompt before searching, so a miss pays an embedding round trip before the first streamed token of the real model call arrives.
How much can semantic caching reduce LLM costs?
It depends almost entirely on how often users repeat the same intent. Research on categorised workloads reports hit rates of 40 to 60 percent for high-repetition query types and 5 to 15 percent for low-repetition ones. Against a baseline already using provider prompt caching, a 20 percent correct-hit rate on a typical support workload cuts the model bill by roughly 19 percent.
Which vector database is best for semantic caching?
Any store with fast approximate nearest-neighbour search and per-entry expiry will work. Redis or Valkey is the usual choice when latency matters and the cache is small enough to hold in memory; Qdrant, Weaviate and Pinecone suit larger caches. Bifrost supports all four, and recommends Redis or Valkey when only exact-match caching is enabled.


