langchain · difficulty ◆◆
cache_weight: Explicit Prompt Caching in langchain-openai
A single knob for cost and latency on repeat context
You wanted cached, predictable context without hand-building OpenAI headers.
$ pip install -U langchain-openai langchain-coreWhat it does
LangChain now lets you explicitly pass cache_weight and related parameters when calling OpenAI chat models, giving direct control over which parts of a conversation are cached. Previously caching was automatic or required manual API-level configuration; now ChatOpenAI accepts explicit caching hints per call without you touching raw OpenAI options.
Why it matters
Prompt caching is one of the most effective ways to cut latency and cost on long conversation chains, especially system prompts that repeat across every request. Before this, you had to build cache_control headers by hand. Now you write ChatOpenAI(model="gpt-4o", cache_weight=0.5) and LangChain translates it into the correct API structure.
Example
$ Initialize ChatOpenAI with a cache_weight hint and invoke the same system plus user message twice.The function below uses Python's `re` module to validate...\nadditional_kwargs={"‘cached’": True} # on the second callcache_weight=0.5 means roughly half of this input may be cached server-side.
Common flags
- ChatOpenAI
- OpenAI chat model client; now accepts cache_weight
- cache_weight
- Constructor hint controlling how much input may be cached
- BaseChatModel._generate
- Internal hook where cache metadata is attached to the response
- langchain-fireworks
- Sibling package fix reporting cached token usage correctly
History
Origin
Caching required manually constructing additional_controllers or cache_control headers on the raw OpenAI API.
The change
The 1.3.5 release (bundled in meta-release langchain 1.3.13) promoted caching to a first-class constructor hint on ChatOpenAI.
Fun facts
Pros & cons
pros
- + Clean constructor-level control
- + No header boilerplate
- + Ideal for repeated system prompts
cons
- − Saving depends on repeated context
- − Semantics vary across providers
Takeaways
- 1Use cache_weight for predictable caching
- 2Re-run to see the cached marker
- 3Trade memory vs cost per pipeline