langchain · difficulty ◆◆
Token Counting Finally Understands o-Series Models
Count tokens before you pay, even when the model is a reasoning model.
Your token counter raised an error the moment you pointed it at a reasoning model — until today.
$ pip install -U langchain-openai==1.5.2What it does
get_num_tokens_from_messages(messages) on ChatOpenAI estimates how many tokens a conversation will consume by counting each message's content plus per-message structural overhead. Before this fix, OpenAI's o-series models — o1, o3, o4-mini and their pro variants — were not handled by the tokenizer path, so calling the method with an o-series model would throw a ValueError or KeyError because their names weren't recognized by the counting logic. This change teaches the token-counting layer to recognize o-series models, so token estimation now works for reasoning models just like it does for gpt-4o, gpt-5, and the rest of the catalog.
Why it matters
Accurate token counting is the backbone of cost control and context management. If you build agents or RAG apps that select a model at runtime, or that stream a conversation history into a context window, you need a reliable pre-invocation token count so you can trim history, summarize, or raise a budget alert before you pay for a call. o-series reasoning models are becoming a default choice for hard reasoning tasks, so silently failing to count their tokens meant you either had to hardcode a fallback path or lose the guardrail entirely. With this fix, get_num_tokens_from_messages becomes a consistent, model-agnostic utility across the whole OpenAI catalog — exactly what a token-budget middleware in a real project needs.
Example
$ Count tokens for an o-series model that previously raised, and compare with gpt-4o-minifrom langchain_openai import ChatOpenAI
from langchain_core.messages import HumanMessage, SystemMessage
llm = ChatOpenAI(model="o4-mini")
messages = [
SystemMessage(content="You are a precise reasoning assistant."),
HumanMessage(content="Prove that sqrt(2) is irrational."),
]
count = llm.get_num_tokens_from_messages(messages)
print("o-series token count:", count)
gpt = ChatOpenAI(model="gpt-4o-mini")
print("gpt-4o-mini token count:", gpt.get_num_tokens_from_messages(messages))Exact counts vary slightly by tokenizer version and prompt; the key outcome is that the o-series call no longer raises.
$ A budget guard that raises before invoking if the estimate is too highcount = llm.get_num_tokens_from_messages(messages)
if count > 200:
raise RuntimeError(f"Too many tokens for budget: {count}")
# otherwise, safe to invoke
result = llm.invoke(messages)This is the guardrail that silently disappeared for o-series models before the fix.
Common flags
- ChatOpenAI.get_num_tokens
- Single-string token estimate (no role overhead).
- ChatOpenAI.invoke
- Runs the model; returns AIMessage with usage_metadata.
- get_num_tokens_from_messages
- Message-list-aware count used for context/budget checks.
- AIMessage.usage_metadata
- Post-call input_tokens / output_tokens / total_tokens.
History
The model catalog kept growing faster than the tokenizer
LangChain's token counter had a hardcoded notion of which OpenAI model names it understood. For years that was fine — gpt-3.5-turbo, gpt-4, gpt-4o all fit the pattern. Then the o-series reasoning models landed with their own naming scheme, and the counting layer didn't know them. The result: a method that worked perfectly for classic models threw an ugly KeyError for the very models people were adopting for hard reasoning tasks. This fix is a catalog widening — teach the tokenizer the new names so the utility stays universal.
Fun facts
Pros & cons
pros
- + Token counting works uniformly across the OpenAI catalog
- + Enables pre-invocation budget guards for reasoning models
- + Fixes a silent ValueError/KeyError on o-series names
- + Complements post-call usage_metadata for full cost visibility
cons
- − Requires langchain-openai>=1.5.2
- − Estimate is approximate, not an exact bill
Takeaways
- 1Upgrade to langchain-openai>=1.5.2 for o-series token support.
- 2Call get_num_tokens_from_messages before invoke as a budget guard.
- 3Compare the estimate against usage_metadata for exact post-call billing.
- 4Treat the estimate as a guardrail — exact counts vary by tokenizer version.