kmail.at
← learning

langchain · difficulty ◆◆

Token Counting Finally Understands o-Series Models

Count tokens before you pay, even when the model is a reasoning model.

Your token counter raised an error the moment you pointed it at a reasoning model — until today.

2026-08-19 · 7 min read

$ pip install -U langchain-openai==1.5.2

What it does

get_num_tokens_from_messages(messages) on ChatOpenAI estimates how many tokens a conversation will consume by counting each message's content plus per-message structural overhead. Before this fix, OpenAI's o-series models — o1, o3, o4-mini and their pro variants — were not handled by the tokenizer path, so calling the method with an o-series model would throw a ValueError or KeyError because their names weren't recognized by the counting logic. This change teaches the token-counting layer to recognize o-series models, so token estimation now works for reasoning models just like it does for gpt-4o, gpt-5, and the rest of the catalog.

Why it matters

Accurate token counting is the backbone of cost control and context management. If you build agents or RAG apps that select a model at runtime, or that stream a conversation history into a context window, you need a reliable pre-invocation token count so you can trim history, summarize, or raise a budget alert before you pay for a call. o-series reasoning models are becoming a default choice for hard reasoning tasks, so silently failing to count their tokens meant you either had to hardcode a fallback path or lose the guardrail entirely. With this fix, get_num_tokens_from_messages becomes a consistent, model-agnostic utility across the whole OpenAI catalog — exactly what a token-budget middleware in a real project needs.

Example

$ Count tokens for an o-series model that previously raised, and compare with gpt-4o-mini
from langchain_openai import ChatOpenAI
from langchain_core.messages import HumanMessage, SystemMessage

llm = ChatOpenAI(model="o4-mini")

messages = [
    SystemMessage(content="You are a precise reasoning assistant."),
    HumanMessage(content="Prove that sqrt(2) is irrational."),
]

count = llm.get_num_tokens_from_messages(messages)
print("o-series token count:", count)

gpt = ChatOpenAI(model="gpt-4o-mini")
print("gpt-4o-mini token count:", gpt.get_num_tokens_from_messages(messages))

Exact counts vary slightly by tokenizer version and prompt; the key outcome is that the o-series call no longer raises.

$ A budget guard that raises before invoking if the estimate is too high
count = llm.get_num_tokens_from_messages(messages)
if count > 200:
    raise RuntimeError(f"Too many tokens for budget: {count}")
# otherwise, safe to invoke
result = llm.invoke(messages)

This is the guardrail that silently disappeared for o-series models before the fix.

Common flags

ChatOpenAI.get_num_tokens
Single-string token estimate (no role overhead).
ChatOpenAI.invoke
Runs the model; returns AIMessage with usage_metadata.
get_num_tokens_from_messages
Message-list-aware count used for context/budget checks.
AIMessage.usage_metadata
Post-call input_tokens / output_tokens / total_tokens.

History

The model catalog kept growing faster than the tokenizer

LangChain's token counter had a hardcoded notion of which OpenAI model names it understood. For years that was fine — gpt-3.5-turbo, gpt-4, gpt-4o all fit the pattern. Then the o-series reasoning models landed with their own naming scheme, and the counting layer didn't know them. The result: a method that worked perfectly for classic models threw an ugly KeyError for the very models people were adopting for hard reasoning tasks. This fix is a catalog widening — teach the tokenizer the new names so the utility stays universal.

Fun facts

Pros & cons

pros

  • + Token counting works uniformly across the OpenAI catalog
  • + Enables pre-invocation budget guards for reasoning models
  • + Fixes a silent ValueError/KeyError on o-series names
  • + Complements post-call usage_metadata for full cost visibility

cons

  • − Requires langchain-openai>=1.5.2
  • − Estimate is approximate, not an exact bill

Takeaways

  1. 1Upgrade to langchain-openai>=1.5.2 for o-series token support.
  2. 2Call get_num_tokens_from_messages before invoke as a budget guard.
  3. 3Compare the estimate against usage_metadata for exact post-call billing.
  4. 4Treat the estimate as a guardrail — exact counts vary by tokenizer version.

Related commands

← all learning