kmail.at
← learning

langchain · difficulty ◆◆

Tool Schema Cache for Throughput

Warm cache, cooler bills, faster turns.

Token overhead shrank the moment the schema hit a warm cache.

2026-07-27 · 6 min read

$ get_weather.tool_call_schema

What it does

In langchain-core 1.5.1 the token-counting path for BaseTool was refactored to cache the tool_call_schema internally. When count_tokens_approximately() runs on a BaseTool, the schema is now looked up from a warm cache rather than recomputed or reconstructed on every call. Token counting runs on every LLM call that involves tools — often thousands of times in a production chain — so any redundant schema processing adds measurable overhead.

Why it matters

If you have traced LangChain token usage in a tool-calling chain and noticed unexpectedly high counts, part of the cause may have been repeated re-serialization of tool schemas. The fix adds a caching layer: the schema is computed once and reused via the tool_call_schema key inside BaseTool. The impact is clearest in high-throughput scenarios where agents call many tools or pipelines invoke the same tool repeatedly across turns — both latency and token billing improve.

Example

$ @tool
def get_weather(city: str) -> str:
    """Return the weather for a given city."""
    return f"The weather in {city} is sunny."

llm = FakeListLLM(responses=["It is sunny in Tokyo."])
print("First call:", get_weather.invoke({"city": "Tokyo"}))
print("Second call:", get_weather.invoke({"city": "Paris"}))
First call: The weather in Tokyo is sunny.
Second call: The weather in Paris is sunny.

After the first call the schema is cached; later invocations skip re-computation.

Common flags

count_tokens_approximately()
Estimates token count; now uses the tool_call_schema cache
BaseTool.invoke()
Synchronous tool invocation that triggers token counting
tool_call_schema
Property storing the cached schema, populated lazily

History

A caching layer for billing

langchain-core 1.5.1 shipped the fix (PR #39020) alongside LangSmith gateway env-var support. High-throughput teams saw both latency and token-overhead drop for repeated tool-call turns.

Fun facts

Pros & cons

pros

  • + Lower latency per turn
  • + Smaller token overhead
  • + Predictable counts

cons

  • − Perf improvement only
  • − Needs 1.5.1+

Takeaways

  1. 1Schemas are computed once and reused.
  2. 2High-throughput loops benefit most.
  3. 3Track prompt_tokens to see the savings.

Related commands

← all learning