langchain · difficulty ◆◆
ChatOpenAI: Explicit Prompt Caching to Cut Latency and Cost
Pay for that system prompt once per cache window
You were re-paying to re-tokenize a giant system prompt on every single turn.
$ pip install -U langchain-openai>=1.3.5What it does
langchain-openai 1.3.5 lets you explicitly request prompt caching by passing cache_implicit=True, or a dict with TTL and max-age hints, when calling the model. LangChain passes cache_implicit to the ChatCompletion request and surfaces cache hits in AIMessage.additional_kwargs["cached_content"].
Why it matters
Prompt caching is one of the most effective ways to cut latency and cost on long conversations. In RAG pipelines and agentic loops, the same large system prompt is sent every turn. Explicit caching lets you pay for it once per cache window and control TTLs from your code.
Example
$ Run three queries against a heavy system prompt with explicit caching.[LGTM!] One-line: Doubles input. Issues: None. Corrected: def foo(x): return x * 2
[CACHED] Code review: LGTM! ...
[CACHED] Code review: LGTM! ...After the first call the system prompt is cached; later calls report cached_content.
Common flags
- ChatOpenAI.invoke()
- Core invocation, now accepts cache_implicit
- ChatOpenAI._generate()
- Internal generator that checks cache_implicit before building the request
- HumanMessage
- User turn message type
- SystemMessage
- System-prompt message type
History
Origin
Prompt caching was automatic or inaccessible via the LangChain API before this release.
The fix
PR #38762 added the cache_implicit parameter, giving fine-grained control over which messages get cached.
Fun facts
Pros & cons
pros
- + Explicit control over caching
- + Surfaces cache hits in kwargs
- + Directly cuts latency and cost
cons
- − Requires an OpenAI account with caching support
- − OpenAI-specific parameter
Takeaways
- 1Request caching with cache_implicit
- 2Check cached_content in kwargs
- 3Set TTLs based on session age