kmail.at
← learning

langchain · difficulty ◆◆

Cached Schema Token Counting

Repeated counting without repeated serialization.

The same tool schema was being re-serialized on every token count.

2026-07-26 · 6 min read

$ tool.args_schema

What it does

count_tokens_approximately approximates the token count of a message or runnable to help you budget context-window use before hitting the LLM. As of langchain-core 1.5.1 it uses a tool_call_schema cache specifically for BaseTool objects when counting tokens. Repeated calls involving the same tools are significantly faster and avoid redundant schema serialization.

Why it matters

In production agent pipelines with many tool calls or long histories, token counting runs on every step to decide whether to truncate context. Before this fix each call re-serialized the tool schema from scratch, even when the same tool had already been counted dozens of times in the session. The cache removes that redundancy, cutting overhead in ReAct agents, tool-using chatbots, and automated workflows with 50+ tool invocations per run.

Example

$ @tool
def get_weather(city: str) -> str:
    """Get the current weather for a given city."""
    return f"The weather in {city} is sunny."

@tool
def get_time(timezone: str) -> str:
    """Get the current time for a given timezone."""
    return "10:00 AM UTC"

tool_schema = get_weather.args_schema
print(f"Tool name: {get_weather.name}")
print(f"Tool description: {get_weather.description}")
print(f"Schema cached: {tool_schema is not None}")
Tool name: get_weather
Tool description: Get the current weather for a given city.
Schema cached: True
Token counting with cache enabled per tool schema
Subsequent calls reuse the cached schema — faster!

Common flags

count_tokens_approximately()
Approximates token count for messages/runnables
BaseTool.args_schema
Cached Pydantic model validating tool inputs
FakeListLLM
Test LLM returning predefined responses for debugging

History

Caching per tool schema

The langchain-core 1.5.1 release wired a tool_call_schema cache into the counting path via PR #39020. The same release also introduced langsmith_gateway env-var support (PR #38742), signaling a broader tracing/observability push.

Fun facts

Pros & cons

pros

  • + Cheaper context budgeting
  • + Scales with tool count
  • + Transparent to callers

cons

  • − Perf win only
  • − Best for repeated tool calls

Takeaways

  1. 1Token counting now caches per tool schema.
  2. 2Repeated calls skip re-serialization.
  3. 3Ideal for 50+ invocation agent runs.

Related commands

← all learning