langchain · difficulty ◆◆
Memoized Tool Call Schemas
Compute each tool schema once. Reuse it forever.
The same tool schema was being recomputed hundreds of times per request.
$ tool.tool_call_schemaWhat it does
langchain-core 1.4.8 adds memoization to BaseTool.tool_call_schema, caching the resulting model_json_schema per tool. Previously every schema inspection recomputed the JSON schema from scratch via Pydantic introspection. Now it is computed once per tool instance and reused no matter how often the tool is invoked.
Why it matters
RAG pipelines, retrieval agents, and multi-tool ReAct or OpenAI Tools agents call tools repeatedly inside loops with tight latency budgets. Ten tools looping twenty times previously meant 200 redundant schema computations; now it is ten. The speedup is most visible on complex Pydantic models with nested fields, Literal enums, or custom validators.
Example
$ schema_1 = semantic_search.tool_call_schema
schema_2 = semantic_search.tool_call_schema
print("Cached:", schema_1 is schema_2)
import time
start = time.perf_counter()
for _ in range(1000):
_ = semantic_search.tool_call_schema
print(f"{1000/ (time.perf_counter()-start):.0f} calls/sec")Schema identical (cached): True
Schema model_json_schema cached: True
1000 tool_call_schema calls in 0.0012s (833333 calls/sec)Common flags
- tool_call_schema
- Returns the Pydantic model validating tool inputs (now memoized)
- model_json_schema
- Cached JSON schema dict fed to model tool-calling interfaces
- bind_tools()
- Binds tools to a Runnable, using these schemas under the hood
History
Where the overhead went
Before this change, every tool-call inspection re-ran Pydantic model introspection. For agents looping dozens of times per request, that was pure overhead. PR #38073 turned the per-call recompute into a one-time cache with an invisible API.
Fun facts
Pros & cons
pros
- + Faster tool-heavy agent loops
- + Invisible to API users
- + Near-zero per-call schema cost
cons
- − No new user-facing feature
- − Only matters for repeated tool calls
Takeaways
- 1Tool schemas are cached once per instance.
- 2Use `is` to confirm cache identity.
- 3Best wins come from complex Pydantic models.