kmail.at
← learning

langchain · difficulty ◆◆

Parallel Tool Calls on bind_tools (July)

Cut latency roughly in half with fan-out tool calls.

Fan-out your tool calls and watch latency drop by half.

2026-07-01 · 7 min read

$ model.bind_tools(tools, parallel_tool_calls=True)

What it does

bind_tools on LangChain chat models exposes a parallel_tool_calls parameter for OpenRouter models. When True, the model can return multiple tool calls in a single response rather than one per message turn. It passes the flag through the OpenRouter provider’s tool-calling implementation, serializing multiple invocations into one AIMessage with multiple ToolCall objects.

Why it matters

Sequential tool execution is a latency bottleneck. A weather API call plus a database lookup as two full round-trips becomes one. This shines for multi-tool agents querying several independent data sources, high-throughput pipelines maximizing tokens-per-second, and fan-out/fan-in architectures.

Example

$ model = ChatOpenRouter(model="anthropic/claude-3-5-sonnet", openai_api_key="<your-key>").bind_tools([get_weather, get_time], parallel_tool_calls=True)
resp = model.invoke([HumanMessage(content="Weather in Paris and time in UTC?")])
print("Tool calls:", len(resp.tool_calls))
for tc in resp.tool_calls:
    tool_func = {"get_weather": get_weather, "get_time": get_time}[tc["name"]]
    print(f"  Result: {tool_func.invoke(tc['args'])}")
Tool calls returned: 2
  - get_weather: args={'city': 'Paris'}
  - get_time: args={'timezone': 'UTC'}
  Result: Sunny, 22°C in Paris
  Result: Current time in UTC: 14:32 UTC

Actual output depends on model behavior at inference time; not all models reliably produce parallel calls.

Common flags

bind_tools
Bind a list of tools to a chat model for tool-use LLMs
ChatOpenRouter
OpenRouter’s chat model wrapper extending BaseChatModel
create_react_agent
Builds a ReAct agent, which internally uses bind_tools

History

A sparse release day

The top release this day, langchain-openrouter==0.2.5, had no substantive changelog, so this tutorial revisits the feature-rich 0.2.4 parallel_tool_calls feature on bind_tools.

Fun facts

Pros & cons

pros

  • + Halves latency for independent calls
  • + Natural fan-out/fan-in
  • + One assistant turn

cons

  • − Model may still serialize
  • − Must handle variable call counts

Takeaways

  1. 1parallel_tool_calls=True returns a batch of ToolCall objects.
  2. 2Execute them by looking up functions by name.
  3. 3Latency drops roughly in half for independent tools.

Related commands

← all learning