Prompt caching
Prompt caching lets you reuse the large, static sections of your prompts - system prompts, schemas, tool definitions, context from documents or codebases - across many requests, and pay a discounted rate for the reused part. It's most valuable when many calls share a big chunk of context: a fleet of agents working from the same instructions, a batch job running one prompt template over thousands of inputs, or a conversation that grows turn by turn.
Doubleword offers caching in two modes:
- Implicit caching is on by default for every cache-supported model and needs no code changes at all. When the infrastructure serving your request reuses your prompt's prefix, the reused tokens are automatically billed at the model's discounted cache-read rate, and the discount shows up in the response's
usagefields. It's best-effort meaning available savings depend on how your requests are served and how long a prefix stays warm. However, it is does not require any work to activate and continues to work from clients, even if you cannot modify them. - Explicit caching is the opt-in upgrade: you place
cache_controlmarkers on the parts of the prompt you want cached, and in exchange you get guaranteed, predictable caching - reads that keep working regardless of how your requests are served, sliding TTLs that keep an actively-used prefix warm, layered breakpoints for content that changes at different rates, and cache sharing across your organization's keys. If you can add markers, explicit caching is the way to get the most out of caching; switching it on takes only 3–4 extra lines of code.
A request with markers is priced by the explicit system, so your bill is always explained by your own breakpoints; a request without markers gets implicit caching. One practical rule of thumb: explicit caching has a minimum cacheable prefix length (1,024 tokens on most models), so if your stable prefix is usually shorter than that, skip the markers and let implicit caching do its thing.
Availability
Not every model supports caching. A model supports it (both modes) when its pricing metadata reports cache_pricing.enabled: true - check the model catalog for the selected model rather than assuming caching is universally available. On models without caching support, cache_control markers are simply ignored and requests bill at standard rates, so they're always safe to send.
Endpoint support
| Endpoint | Implicit | Explicit (cache_control markers) |
|---|---|---|
/v1/chat/completions | ✅ | ✅ block markers + automatic placement |
/v1/messages (Anthropic-compatible) | ✅ | ✅ block markers + automatic placement |
/v1/responses | ✅ | Automatic placement only - the Responses input shape has no per-block markers |
/v1/completions | ✅ | - the prompt is a single string, so there's nothing to mark; implicit covers this endpoint |
Usage fields
Both modes report caching through the same usage fields on every response, so this is the first place to look when adopting either one - send a repeat request and check them. The shape depends on which API standard you call: toggle the provider dropdown on the block below to switch between the OpenAI-style shape (returned by /v1/chat/completions and /v1/completions) and the Anthropic shape (returned by /v1/messages). The totals differ - OpenAI's prompt_tokens includes the cached tokens, Anthropic's input_tokens reports the cached prefix separately - but the cache fields share the same names on both:
prompt_tokens total input tokens (includes the cached ones)
completion_tokens tokens generated by the model
cache_read_input_tokens tokens served from cache - discounted read rate
prompt_tokens_details.cached_tokens same count as cache_read_input_tokens (OpenAI-standard field)
cache_creation_input_tokens tokens written to the cache this request - write rateinput_tokens new (non-cached) input tokens - the cached prefix is reported separately
output_tokens tokens generated by the model
cache_read_input_tokens tokens served from cache - discounted read rate (omitted when zero)
cache_creation_input_tokens tokens written to the cache this request - write rate (omitted when zero)cache_read_input_tokens is your discount, whichever mode produced it. cache_creation_input_tokens is only ever non-zero under explicit caching - implicit caching never writes. On streaming responses the cache fields arrive in the final usage chunk.
Pricing
Cache prices are multipliers of the selected model's standard input price. The model catalog shows model-specific cache-read prices.
| Operation | Current multiplier on supported models |
|---|---|
| Cache read | r × standard input price1 |
| Cache write - 5m | 1.25× standard input price |
| Cache write - 1h | 2× standard input price |
| Uncached input | Model's standard input price |
| Output | Model's standard output price |
These are the current multipliers for cache-enabled models in the public catalog. Availability, minimum prefix length, and multipliers are model-specific and may change.
Both modes bill cache reads at the same discounted read rate. Implicit caching never bills writes - tokens that weren't reused are simply standard input. Explicit cache writes carry the small premium above, which is what buys the guarantee: once a prefix is written, every later read costs only the discounted read rate - no further write charges - for as long as the TTL keeps sliding. The precise break-even point depends on the model, TTL, and active multipliers.
Implicit caching
There is nothing to switch on. Send requests as normal; when the serving infrastructure reuses a prefix of your prompt, the response reports it in the usage fields (cache_read_input_tokens, mirrored in prompt_tokens_details.cached_tokens) and the bill reflects it.
A few things to know:
- Reuse is best-effort. Whether a repeat prompt gets a discount depends on how your requests are served and how recently the prefix was used. Identical prompts sent close together do well; sporadic traffic may not. There are no write charges and no penalty when nothing is reused - you simply pay the standard rate.
- Prefix reuse is left-anchored: keep the stable part of your prompt at the front and per-request content (questions, timestamps, IDs) at the end, and you'll maximise your hit rate in either mode.
- When it's the right choice: clients you can't modify (off-the-shelf tools, proxied traffic), stable prefixes under the explicit minimum (1,024 tokens on most models), or any workload where zero-effort savings beat guaranteed ones.
If your workload has a big, genuinely stable prefix and you control the client, keep reading - explicit caching turns those best-effort savings into guaranteed ones.
Explicit caching
Explicit caching puts you in control: cache_control markers say exactly which prefix to cache and for how long, and the platform guarantees the result - the first request writes the prefix, and every later request carrying the same marker reads it at the discounted rate for the TTL, regardless of how your requests are served - and at no extra cost: a written prefix can be read back indefinitely without ever paying the write premium again, for as long as the TTL keeps sliding. Reads slide the expiry forward, so an actively-reused prefix stays warm indefinitely; breakpoints can be layered so stable content (tools, system prompts) is written once while faster-moving content caches on its own cadence; and a cached prefix is shared across your organization's API keys, so a prefix written by one service is readable by all of them (personal keys keep caching private to you).
How it works
A breakpoint (cache_control) marks the end of a cacheable prefix: everything from the start of the prompt up to and including that block.
- The first request with a given prefix writes it - a cache creation.
- Later requests that begin with the same prefix read it - a cache read, billed at the discounted rate.
Explicit caching is driven by the markers: every request that should read the cache carries the same cache_control marker, and your cache activity is always explained by the breakpoints you placed. (A request with no markers at all is handled by implicit caching instead.)
Matching is prefix-based and left-anchored: a request reuses the cache up to the longest leading run of content that is byte-for-byte identical to something cached before. As soon as content diverges, everything after it is treated as new.
Request A (writes): [ tools + system ]●[ user question 1 ]
Request B (reads): [ tools + system ]●[ user question 2 ]
└── identical ──┘ └── different ──┘
cache READ full price(● = breakpoint.)
TTL - each breakpoint lasts "5m" (the default) or "1h". The TTL is a sliding window: every read refreshes the expiry, so an actively-reused prefix stays warm and only expires after a full TTL with no reuse. Use 5m for bursty, back-to-back calls; 1h for interactive sessions.
Privacy & sharing - a cached prefix is scoped to the API key's owner and is never shared across customers. This gives you a choice of scope: use a key from your personal account to keep caching to yourself, or use one of your organization's keys to share prefix caching across your whole org - any org key can read a prefix that another org key wrote.
Quick start
Put the reusable content in an array-style content block and add the marker. Send the request twice - the second call reads the prefix from cache.
POST https://api.doubleword.ai/v1/chat/completions
{
"model": "{{selectedModel.id}}",
"messages": [
{ "role": "system", "content": [
{ "type": "text",
"text": "<~2,000 tokens of stable instructions>",
"cache_control": { "type": "ephemeral", "ttl": "1h" } } ] },
{ "role": "user", "content": "How do I reset my password?" }
]
}POST https://api.doubleword.ai/v1/messages
{
"model": "{{selectedModel.id}}",
"max_tokens": 256,
"system": [
{ "type": "text",
"text": "<~2,000 tokens of stable instructions>",
"cache_control": { "type": "ephemeral", "ttl": "1h" } } ],
"messages": [ { "role": "user", "content": "How do I reset my password?" } ]
}Read the result back from usage:
// 1st call - writes the prefix
"usage": { "prompt_tokens": 2088, "cache_creation_input_tokens": 2048, "cache_read_input_tokens": 0 }
// 2nd call - reads it; only the new user turn is billed at full price
"usage": { "prompt_tokens": 2096, "cache_read_input_tokens": 2048, "cache_creation_input_tokens": 0,
"prompt_tokens_details": { "cached_tokens": 2048 } }// 1st call - writes the prefix (zero-valued cache fields are omitted)
"usage": { "input_tokens": 2088, "output_tokens": 24, "cache_creation_input_tokens": 2048 }
// 2nd call - reads it; input_tokens counts only the new turn - the cached prefix is in cache_read_input_tokens
"usage": { "input_tokens": 48, "output_tokens": 24, "cache_read_input_tokens": 2048 }Streaming works the same way - the cache fields arrive in the final usage chunk.
Multiple breakpoints
You can place up to 4 breakpoints in one request, and using more than one pays off when your prompt has segments that change at different frequencies.
Take a common case: your tools never change, but you have a small set of system prompts - say one per task or agent mode - and pick one for each request. With a single breakpoint at the end of the system prompt, each system prompt caches tools + system together as one entry - so the identical tools get cached again under every variant, and you pay to write the tools once for each system prompt.
Put two breakpoints instead - one after the tools, one after the system prompt - and each segment is cached on its own cadence: the tools are written once and read by every system-prompt variant, and each system prompt is cached on top of them.
[ tools ]① ← cached once; read by every system-prompt variant
[ system prompt ]② ← cached per distinct system prompt
[ user message ]The rule: put a breakpoint at the end of each segment that changes on its own cadence - most-stable first (tools), then less-stable (system prompt, retrieved documents), then the conversation. Each layer is written once and reused until it changes, so you never re-pay to cache the stable parts underneath.
Adding a breakpoint never costs anything on its own - you're billed only for the tokens actually written to or read from the cache, never per breakpoint. So there's no harm in placing one wherever the prompt might start to change further down: an extra breakpoint can only ever help.
Automatic marker placement
If you'd rather not manage per-block markers, add a single top-level cache_control to the request and the platform places the breakpoint for you - at the last viable point in the prompt, so the cached prefix covers as much of the request as possible and moves forward automatically as a conversation grows:
{
"model": "{{selectedModel.id}}",
"cache_control": { "type": "ephemeral", "ttl": "1h" },
"messages": [
{ "role": "system", "content": "<agent instructions>" },
{ "role": "user", "content": "How do I reset my password?" }
]
}This follows the same semantics as Anthropic's automatic caching: it consumes one of the 4 breakpoint slots, combines freely with explicit block markers (the automatic breakpoint simply lands after them), and bills identically to a marker you placed yourself. It's the easiest way to get guaranteed caching on a growing conversation - and on the /v1/responses endpoint it's the way to use explicit caching, since that input shape has no per-block markers.
Patterns
System prompt & tools
Tool definitions and a long system prompt are identical on every call - cache them as the stable prefix and vary only the user turn. Tools sit first in the prefix, so a cache_control on the last tool caches every tool before it; a marker on the system block caches the tools and the system prompt.
{
"model": "{{selectedModel.id}}",
"tools": [
{ "type": "function", "function": { "name": "search", "parameters": {} } },
{ "type": "function", "function": { "name": "book_flight", "parameters": {} },
"cache_control": { "type": "ephemeral", "ttl": "1h" } }
],
"messages": [
{ "role": "system", "content": [
{ "type": "text", "text": "<agent instructions>",
"cache_control": { "type": "ephemeral", "ttl": "1h" } } ] },
{ "role": "user", "content": "Book me a flight to Tokyo next Friday." }
]
}{
"model": "{{selectedModel.id}}",
"max_tokens": 512,
"tools": [
{ "name": "search", "input_schema": { "type": "object" } },
{ "name": "book_flight", "input_schema": { "type": "object" },
"cache_control": { "type": "ephemeral", "ttl": "1h" } }
],
"system": [ { "type": "text", "text": "<agent instructions>",
"cache_control": { "type": "ephemeral", "ttl": "1h" } } ],
"messages": [ { "role": "user", "content": "Book me a flight to Tokyo next Friday." } ]
}RAG / document Q&A
Cache a large retrieved document once, then ask many questions about it - each question only pays full price for itself.
{
"model": "{{selectedModel.id}}",
"messages": [
{ "role": "system", "content": [ { "type": "text", "text": "Answer only from the document." } ] },
{ "role": "user", "content": [
{ "type": "text", "text": "<the 50-page contract>",
"cache_control": { "type": "ephemeral", "ttl": "1h" } },
{ "type": "text", "text": "What is the termination notice period?" }
] }
]
}The next question about the same document reads the whole contract from cache; only the new question is billed at full rate.
Growing conversation
Multi-turn chat is layered caching with a moving target: the system and tools stay fixed, any retrieved docs are per-session, and the conversation grows a message at a time. Put a breakpoint on each fixed layer, and put a cache_control marker on the latest message in every request (or let automatic placement do exactly that for you). Each turn then reads the whole conversation so far from cache at the discounted rate, and the only tokens that aren't already cached - what you added since your last request (the previous reply plus your new message) - are written at the write rate, ready to read cheaply next turn. You never re-write earlier turns, so a long conversation stays cheap: a discounted read of the history plus a small write on just the new tail.
Turn 1: [sys+tools]①[docs]②[user₁ ●] → mark user₁; writes through user₁
Turn 2: [sys+tools]①[docs]②[user₁][asst₁][user₂ ●] → mark user₂; reads through asst₁, writes the new turn
Turn 3: [sys+tools]①[docs]②…[asst₂][user₃ ●] → mark user₃; reads through asst₂, writes the new turnBecause matching is longest-common-prefix, a follow-up that shares only breakpoint ① still reads that layer and re-caches from there - you don't have to plan reuse perfectly.
What you can cache
cache_control works on content blocks throughout the request, in the tools → system → messages order:
- Tool definitions - objects in the
toolsarray. - System content blocks.
- User and assistant text content blocks.
- Tool results - a
role: "tool"message (Chat Completions) or atool_resultblock (Messages API).
A few things worth knowing:
- Images & documents still benefit from caching when they sit inside a cached prefix (and changing one invalidates everything after it), but the image/document tokens themselves aren't discounted - only the text tokens in the prefix are.
- Empty text blocks and sub-content blocks (like citations) can't be marked directly - cache the top-level block instead.
- Confirm it landed by reading the usage fields. On a model that doesn't support caching, markers are simply ignored and you're billed at standard rates - so they're always safe to include.
Getting the best out of explicit caching
Golden rule: stable content first, volatile content last, breakpoint at the boundary.
[ STABLE PREFIX: tools, system prompt, instructions, examples, documents ] ● breakpoint
[ VOLATILE SUFFIX: this request's actual question ]- Order matters. Everything before the breakpoint must be identical across requests to get a hit. Keep per-request content - the question, timestamps, IDs - after the cached blocks, and send the cached blocks in a consistent order (the same content in a different order is a cache miss).
- Keep prefixes byte-identical. A single changed character before a breakpoint - a different date, a reordered sentence - breaks the match for everything after it.
- Cache the big stuff. Minimum cacheable prefix length is model-specific - most supported models currently require 1,024 tokens. If your stable prefix usually sits under the minimum, leave the markers off and let implicit caching handle it.
- Don't cache one-offs. A write carries a small premium, so only cache prefixes you'll reuse.
- Mark the latest message in a conversation (or use automatic placement) so each turn reads the whole history and writes only the new tail.
Reference
Marker - add to any content block, or to a tool object in the tools array:
"cache_control": { "type": "ephemeral", "ttl": "5m" }typemust be"ephemeral";ttlis"5m"(default) or"1h".- Up to 4 breakpoints per request.
- Minimum prefix tokens are model-specific; check the selected model's pricing information.
- Markers change the bill, never the output.
The usage fields that report reads and writes are shared with implicit caching and documented above.
Further information
Doubleword's explicit caching follows the same model and semantics as Anthropic's prompt caching - markers, TTLs, breakpoints, and automatic placement all carry over directly. For deeper background - cache-aware prompt structure, what invalidates a cache, and more worked examples - see Anthropic's prompt caching guide. The available TTLs are 5m and 1h; consult the model catalog for the selected model's authoritative pricing information.
Footnotes
-
r is the selected model's current cache-read multiplier. See the model pricing table for the authoritative rate. ↩