langchain · difficulty ◆◆
openai.completions: Explicit Prompt Caching with cache_control
Cut cost and latency on long-context calls you control
Your long system prompts were being re-sent and re-billed on every single call.
$ pip install -U langchain-openai>=1.3.5What it does
langchain-openai 1.3.5 ships explicit prompt caching, a new integration with OpenAI\u2019s cache_control parameter. You mark portions of a conversation as cacheable by attaching cache_control={"type": "cache", "status": "active"} to the extra fields of a SystemMessage or HumanMessage, and LangChain forwards it to the OpenAI API correctly. Unlike OpenAI\u2019s automatic detection, the explicit approach gives you fine-grained control over exactly which blocks get cached.
Why it matters
Prompt caching slashes cost and latency for long-context applications. A 100k-token system prompt repeated across thousands of requests costs up to 10x less when cached. The explicit version wins on predictability: you know precisely what is cached, cold starts are faster, and it applies to any model that supports the feature. RAG pipelines, agentic workflows with fixed system prompts, and batch processing all benefit.
Example
$ Wrap a stable code-review system prompt in explicit cache_control and invoke a reviewer.Summary: The function eval_user_expr contains a critical security vulnerability.\n\nIssues:\n1. [Critical] The use of eval() on user input (expr) allows arbitrary code execution.\n\nSuggestions:\n1. Replace eval() with a safe parser such as ast.literal_eval, or sandbox it.Cached tokens show up in prompt_tokens_details.cached_tokens on repeated calls.
Common flags
- ChatOpenAI.invoke()
- Invoke the model with messages plus caching metadata
- cache_control
- Extra field {"type": "cache", "status": "active"} on any message
- additional_kwargs
- Message field used to attach cache metadata
- PromptCache
- langchain-core base utilities for prompt caching
History
Origin
OpenAI caching was previously automatic and opaque, forcing developers to rely on the model detecting repeated content and discounting it behind the scenes.
The fix
The 1.3.5 release wired cache_control through LangChain messages so developers could explicitly declare cacheable blocks and get predictable pricing.
Fun facts
Pros & cons
pros
- + Predictable cost modelling
- + Faster repeated-call latency
- + Works with any supporting model
cons
- − Requires a langchain-openai upgrade
- − Only helps when context repeats
Takeaways
- 1Attach cache_control to stable blocks
- 2Inspect cached_tokens on repeated calls
- 3Pair with a fixed system prompt for best savings