kmail.at
← learning

langchain · difficulty ◆◆

Cached Tool Call Schemas

Stop re-serializing the same tool schema on every call.

Every agent loop was silently re-paying for the same JSON schema.

2026-07-24 · 6 min read

$ tool.tool_call_schema

What it does

langchain-core 1.5.1 introduces a tool_call_schema cache inside BaseTool that stores the pre-serialized tool-call schema. Previously every count_tokens_approximately() call on a BaseTool re-serialized the tool schema from scratch — expensive when tools are invoked thousands of times in an agentic loop. The cache memoizes the schema on first use and reuses it for all later token-count calls within the same tool instance.

Why it matters

Token counting runs on every turn of a production agentic pipeline, often for dozens of tools. Without caching, repeated JSON-serialization adds latency and CPU overhead that silently inflate per-call latency and API costs. The payoff is clearest in high-frequency agent loops like ReAct agents calling tools 50 to 100 times per query, multi-tool routing layers, and batch pipelines that count tokens on the same tool objects repeatedly.

Example

$ @tool
def get_weather(city: str) -> str:
    """Get the current weather for a city."""
    return f"The weather in {city} is sunny."

@tool
def get_time(timezone: str) -> str:
    """Get the current time for a timezone."""
    return "10:00 AM"

counter = TokenCountingCallbackHandler()
llm = ChatOpenAI(model="gpt-4o", callbacks=[counter])
tools = [get_weather, get_time]
llm.invoke("What is the weather in Paris?", config={"tools": tools})
llm.invoke("What time is it in London?", config={"tools": tools})
print(f"Total tokens counted: {counter.total_tokens}")
Total tokens counted: <varies by model>
Prompt tokens: <varies>
Completion tokens: <varies>

The second invocation runs measurably faster because the tool schema is served from the in-memory cache.

Common flags

count_tokens_approximately()
Counts tokens for a prompt/tool using a model-specific tokenizer
BaseTool.tool_call_schema
Cached JSON schema used for tool-call token counting
TokenCountingCallbackHandler
Callback tracking token usage across LLM calls

History

PR #39020

The fix landed in langchain-core 1.5.1 (released 2026-07-23). It moved tool-schema serialization to a one-time, lazily-populated cache inside BaseTool, so repeated token counting no longer pays the same serialization cost.

Fun facts

Pros & cons

pros

  • + Faster agent loops
  • + Less CPU per token count
  • + Invisible to API users

cons

  • − Perf-only, no new feature
  • − Matters mainly for repeated calls

Takeaways

  1. 1Tool schemas are cached once per tool instance.
  2. 2The second invoke is faster than the first.
  3. 3Upgrade langchain-core to 1.5.1+.

Related commands

← all learning