kmail.at
← learning

langchain · difficulty ◆◆

openai.completions: Explicit Prompt Caching with cache_control

Cut cost and latency on long-context calls you control

Your long system prompts were being re-sent and re-billed on every single call.

2026-07-13 · 7 min read

$ pip install -U langchain-openai>=1.3.5

What it does

langchain-openai 1.3.5 ships explicit prompt caching, a new integration with OpenAI\u2019s cache_control parameter. You mark portions of a conversation as cacheable by attaching cache_control={"type": "cache", "status": "active"} to the extra fields of a SystemMessage or HumanMessage, and LangChain forwards it to the OpenAI API correctly. Unlike OpenAI\u2019s automatic detection, the explicit approach gives you fine-grained control over exactly which blocks get cached.

Why it matters

Prompt caching slashes cost and latency for long-context applications. A 100k-token system prompt repeated across thousands of requests costs up to 10x less when cached. The explicit version wins on predictability: you know precisely what is cached, cold starts are faster, and it applies to any model that supports the feature. RAG pipelines, agentic workflows with fixed system prompts, and batch processing all benefit.

Example

$ Wrap a stable code-review system prompt in explicit cache_control and invoke a reviewer.
Summary: The function eval_user_expr contains a critical security vulnerability.\n\nIssues:\n1. [Critical] The use of eval() on user input (expr) allows arbitrary code execution.\n\nSuggestions:\n1. Replace eval() with a safe parser such as ast.literal_eval, or sandbox it.

Cached tokens show up in prompt_tokens_details.cached_tokens on repeated calls.

Common flags

ChatOpenAI.invoke()
Invoke the model with messages plus caching metadata
cache_control
Extra field {"type": "cache", "status": "active"} on any message
additional_kwargs
Message field used to attach cache metadata
PromptCache
langchain-core base utilities for prompt caching

History

Origin

OpenAI caching was previously automatic and opaque, forcing developers to rely on the model detecting repeated content and discounting it behind the scenes.

The fix

The 1.3.5 release wired cache_control through LangChain messages so developers could explicitly declare cacheable blocks and get predictable pricing.

Fun facts

Pros & cons

pros

  • + Predictable cost modelling
  • + Faster repeated-call latency
  • + Works with any supporting model

cons

  • − Requires a langchain-openai upgrade
  • − Only helps when context repeats

Takeaways

  1. 1Attach cache_control to stable blocks
  2. 2Inspect cached_tokens on repeated calls
  3. 3Pair with a fixed system prompt for best savings

Related commands

← all learning