langchain · difficulty ◆◆
Tool Schema Cache for Throughput
Warm cache, cooler bills, faster turns.
Token overhead shrank the moment the schema hit a warm cache.
$ get_weather.tool_call_schemaWhat it does
In langchain-core 1.5.1 the token-counting path for BaseTool was refactored to cache the tool_call_schema internally. When count_tokens_approximately() runs on a BaseTool, the schema is now looked up from a warm cache rather than recomputed or reconstructed on every call. Token counting runs on every LLM call that involves tools — often thousands of times in a production chain — so any redundant schema processing adds measurable overhead.
Why it matters
If you have traced LangChain token usage in a tool-calling chain and noticed unexpectedly high counts, part of the cause may have been repeated re-serialization of tool schemas. The fix adds a caching layer: the schema is computed once and reused via the tool_call_schema key inside BaseTool. The impact is clearest in high-throughput scenarios where agents call many tools or pipelines invoke the same tool repeatedly across turns — both latency and token billing improve.
Example
$ @tool
def get_weather(city: str) -> str:
"""Return the weather for a given city."""
return f"The weather in {city} is sunny."
llm = FakeListLLM(responses=["It is sunny in Tokyo."])
print("First call:", get_weather.invoke({"city": "Tokyo"}))
print("Second call:", get_weather.invoke({"city": "Paris"}))First call: The weather in Tokyo is sunny.
Second call: The weather in Paris is sunny.After the first call the schema is cached; later invocations skip re-computation.
Common flags
- count_tokens_approximately()
- Estimates token count; now uses the tool_call_schema cache
- BaseTool.invoke()
- Synchronous tool invocation that triggers token counting
- tool_call_schema
- Property storing the cached schema, populated lazily
History
A caching layer for billing
langchain-core 1.5.1 shipped the fix (PR #39020) alongside LangSmith gateway env-var support. High-throughput teams saw both latency and token-overhead drop for repeated tool-call turns.
Fun facts
Pros & cons
pros
- + Lower latency per turn
- + Smaller token overhead
- + Predictable counts
cons
- − Perf improvement only
- − Needs 1.5.1+
Takeaways
- 1Schemas are computed once and reused.
- 2High-throughput loops benefit most.
- 3Track prompt_tokens to see the savings.