langchain · difficulty ◆◆
Parallel Tool Calls on bind_tools (July)
Cut latency roughly in half with fan-out tool calls.
Fan-out your tool calls and watch latency drop by half.
$ model.bind_tools(tools, parallel_tool_calls=True)What it does
bind_tools on LangChain chat models exposes a parallel_tool_calls parameter for OpenRouter models. When True, the model can return multiple tool calls in a single response rather than one per message turn. It passes the flag through the OpenRouter provider’s tool-calling implementation, serializing multiple invocations into one AIMessage with multiple ToolCall objects.
Why it matters
Sequential tool execution is a latency bottleneck. A weather API call plus a database lookup as two full round-trips becomes one. This shines for multi-tool agents querying several independent data sources, high-throughput pipelines maximizing tokens-per-second, and fan-out/fan-in architectures.
Example
$ model = ChatOpenRouter(model="anthropic/claude-3-5-sonnet", openai_api_key="<your-key>").bind_tools([get_weather, get_time], parallel_tool_calls=True)
resp = model.invoke([HumanMessage(content="Weather in Paris and time in UTC?")])
print("Tool calls:", len(resp.tool_calls))
for tc in resp.tool_calls:
tool_func = {"get_weather": get_weather, "get_time": get_time}[tc["name"]]
print(f" Result: {tool_func.invoke(tc['args'])}")Tool calls returned: 2
- get_weather: args={'city': 'Paris'}
- get_time: args={'timezone': 'UTC'}
Result: Sunny, 22°C in Paris
Result: Current time in UTC: 14:32 UTCActual output depends on model behavior at inference time; not all models reliably produce parallel calls.
Common flags
- bind_tools
- Bind a list of tools to a chat model for tool-use LLMs
- ChatOpenRouter
- OpenRouter’s chat model wrapper extending BaseChatModel
- create_react_agent
- Builds a ReAct agent, which internally uses bind_tools
History
A sparse release day
The top release this day, langchain-openrouter==0.2.5, had no substantive changelog, so this tutorial revisits the feature-rich 0.2.4 parallel_tool_calls feature on bind_tools.
Fun facts
Pros & cons
pros
- + Halves latency for independent calls
- + Natural fan-out/fan-in
- + One assistant turn
cons
- − Model may still serialize
- − Must handle variable call counts
Takeaways
- 1parallel_tool_calls=True returns a batch of ToolCall objects.
- 2Execute them by looking up functions by name.
- 3Latency drops roughly in half for independent tools.