kmail.at
← learning

langchain · difficulty ◆◆

reasoning_effort via init_chat_model and the Routing Pattern

Route cheap, classify cheap, reason deep - with one uniform parameter.

The cheapest trick in an LLM app is not thinking hard about every call - reasoning_effort makes that easy.

2026-08-13 · 8 min read

$ pip install -U langchain==1.3.15

What it does

reasoning_effort is now a standard, first-class chat model parameter in LangChain v1, added in langchain-core==1.5.4 and surfaced in langchain==1.3.15 (PR #38887). It tells a model how much chain-of-thought to spend on a task - from a light, fast low pass to a deep high pass - without manually toggling per-provider knobs. Before, you reached into provider-specific constructor args with names that differed across vendors. Now reasoning_effort is one uniform parameter understood by init_chat_model and the model middleware.

Why it matters

Real apps rarely want the same reasoning budget for every call. A routing step classifying an incoming request can run on low (fast, cheap); a code-generation or structured-extraction step should use high to avoid subtle mistakes. Passing reasoning_effort through a single provider-agnostic interface lets you tune latency and cost per call, keep code portable across model backends, and avoid brittle vendor-specific hacks that break when you switch providers.

Example

$ Route then answer: classify cheap, reason deep only for code
fast_model = init_chat_model("gpt-5-mini", reasoning_effort="low")
deep_model = init_chat_model("gpt-5-mini", reasoning_effort="high")

def running_total(nums):
    total = 0
    result = []
    for n in nums:
        total += n
        result.append(total)
    return result

print(running_total([1, 2, 3, 4]))  # -> [1, 3, 6, 10]

The classification branch stays fast and cheap; the code branch spends reasoning budget.

$ Vary effort per call with with_config
model.with_config({"model": {"reasoning_effort": "low"}})
# Vary effort per call on a single chain, and measure cost on a batch of 50 prompts

Common flags

init_chat_model
Provider-agnostic chat-model factory; accepts reasoning_effort uniformly.
wrap_tool_call
Middleware helper with a state_schema param for structured tool state.
StructuredPrompt
Core prompt construct; got a fix so it no longer mutates caller kwargs.

History

From per-provider knobs to one dial

Every reasoning-capable provider shipped its own way to control thinking depth, with incompatible names and semantics. Standardizing reasoning_effort across the base interface - and threading it through init_chat_model - turns a messy per-vendor concern into a single portable parameter that middleware and factories both understand.

Fun facts

Pros & cons

pros

  • + Per-call cost tuning
  • + Portable across providers
  • + Simple routing pattern

cons

  • − Adds a routing step to design
  • − Provider semantics vary

Takeaways

  1. 1Split your calls by reasoning effort, not just by model.
  2. 2Use init_chat_model so the parameter stays provider-agnostic.
  3. 3Measure the cost difference before and after tuning.

Related commands

← all learning