langchain · difficulty ◆◆
Cached Tool Call Schemas
Stop re-serializing the same tool schema on every call.
Every agent loop was silently re-paying for the same JSON schema.
$ tool.tool_call_schemaWhat it does
langchain-core 1.5.1 introduces a tool_call_schema cache inside BaseTool that stores the pre-serialized tool-call schema. Previously every count_tokens_approximately() call on a BaseTool re-serialized the tool schema from scratch — expensive when tools are invoked thousands of times in an agentic loop. The cache memoizes the schema on first use and reuses it for all later token-count calls within the same tool instance.
Why it matters
Token counting runs on every turn of a production agentic pipeline, often for dozens of tools. Without caching, repeated JSON-serialization adds latency and CPU overhead that silently inflate per-call latency and API costs. The payoff is clearest in high-frequency agent loops like ReAct agents calling tools 50 to 100 times per query, multi-tool routing layers, and batch pipelines that count tokens on the same tool objects repeatedly.
Example
$ @tool
def get_weather(city: str) -> str:
"""Get the current weather for a city."""
return f"The weather in {city} is sunny."
@tool
def get_time(timezone: str) -> str:
"""Get the current time for a timezone."""
return "10:00 AM"
counter = TokenCountingCallbackHandler()
llm = ChatOpenAI(model="gpt-4o", callbacks=[counter])
tools = [get_weather, get_time]
llm.invoke("What is the weather in Paris?", config={"tools": tools})
llm.invoke("What time is it in London?", config={"tools": tools})
print(f"Total tokens counted: {counter.total_tokens}")Total tokens counted: <varies by model>
Prompt tokens: <varies>
Completion tokens: <varies>The second invocation runs measurably faster because the tool schema is served from the in-memory cache.
Common flags
- count_tokens_approximately()
- Counts tokens for a prompt/tool using a model-specific tokenizer
- BaseTool.tool_call_schema
- Cached JSON schema used for tool-call token counting
- TokenCountingCallbackHandler
- Callback tracking token usage across LLM calls
History
PR #39020
The fix landed in langchain-core 1.5.1 (released 2026-07-23). It moved tool-schema serialization to a one-time, lazily-populated cache inside BaseTool, so repeated token counting no longer pays the same serialization cost.
Fun facts
Pros & cons
pros
- + Faster agent loops
- + Less CPU per token count
- + Invisible to API users
cons
- − Perf-only, no new feature
- − Matters mainly for repeated calls
Takeaways
- 1Tool schemas are cached once per tool instance.
- 2The second invoke is faster than the first.
- 3Upgrade langchain-core to 1.5.1+.