langchain · difficulty ◆◆◆
CacheConfig: TTL-Aware Explicit Prompt Caching
Only the dynamic user turn travels the wire after the first call
You could not see or control what the provider cached.
$ pip install -U langchain-openai "langchain-core>=0.3.0"What it does
OpenAI prompt caching lets you mark specific conversation parts (system message, tools, few-shot examples) as cached so they are not re-sent every call. In langchain-openai>=1.3.5 you explicitly request caching by passing a cache parameter to ChatOpenAI or init_chat_model, and can configure TTL and staleness via CacheConfig. Before this, caching was implicit and opaque.
Why it matters
Long system prompts and large tool definitions are common in production: RAG pipelines with 20 context documents, multi-step agents with bulky tool schemas, or few-shot learners with 50 example pairs. With explicit caching only the dynamic user turn travels the wire after the first call, cutting latency by 30-70% and reducing billed tokens for repeated-same-context workloads.
Example
$ Build a reviewer model with a cache TTL and invoke the same system prompt twice.<issues>No issues found.</issues> Score: 10/10\n<issues>No issues found.</issues> Score: 10/10\n# second call returned in ~30% of first-call timeInspect the cache via langchain_core.cache.InMemoryCache.lookup().
Common flags
- init_chat_model
- Model-agnostic chat model initialiser with a cache kwarg
- CacheConfig
- Configures cache TTL and staleness for prompt blocks
- langchain_core.caches
- Built-in InMemoryCache, SQLCache, and RedisCache
- ChatOpenAI
- OpenAI Chat API client with explicit cache parameter
History
Origin
Caching was implicit: you could not control what stuck in cache or for how long.
The change
The 1.3.5 release (via meta-release langchain 1.3.13) added explicit cache control and a CacheConfig TTL so developers decide the cache lifetime.
Fun facts
Pros & cons
pros
- + Fine-grained TTL control
- + Big latency wins on repeated context
- + Visible cache state
cons
- − Cache semantics are provider-specific
- − Needs langchain-core>=0.3.0
Takeaways
- 1Set an explicit cache TTL
- 2Time both calls to observe the speedup
- 3Inspect cache state directly