langchain · difficulty ◆◆
Preserve Usage Tokens in v3 Streaming
Stream chunks without losing your token meter.
Your streamed tokens were hiding their cost. Until now.
$ llm.stream(..., version="v3")What it does
With langchain-core >= 1.4.8 (shipped the same day as langchain==1.3.10), token-usage metadata that previously vanished during streaming is now preserved in the final event payload. When you consume a stream chunk by chunk through the v3 streaming protocol, the aggregate usage counts are still available at the end.
Why it matters
In production LLM apps, token usage drives cost tracking, rate-limit management, and debugging. Before this fix, streaming paths gave you no reliable way to read the final token count without switching to the non-streaming invoke() path. Now you get both: a low-latency streaming UX and accurate usage reporting for billing, observability, and optimization.
Example
$ for event in llm.stream("Explain token usage in one sentence."):
events.append(event)
print(event.content, end="", flush=True)
final = events[-1]
print(final.usage)LangChain is a framework for building applications powered by large language models.
— Usage: prompt_tokens=12, completion_tokens=15, total_tokens=27$ for event in chat.stream_events([HumanMessage(content="What is LangChain?")], version="v3"):
events.append(event)
if event.event == "on_chat_model_stream":
print(event.data.chunk.content, end="", flush=True)LangChain is a framework for building context-aware applications...
— Final usage: prompt_tokens=9, completion_tokens=18, total_tokens=27Common flags
- stream_events(version="v3")
- Yield v3 streaming events from a chain or LLM with usage preserved
- on_chat_model_stream
- Event type emitted per streaming chunk in the v3 protocol
- final_event.usage
- Usage metadata now present on the last event after streaming
History
The v3 streaming protocol
LangChain evolved from simple invoke() calls to rich event streams. The v3 protocol introduced structured events, but early versions dropped usage metadata mid-stream. PR #38021 restored that missing link so stream() is as informative as invoke() for metering.
Fun facts
Pros & cons
pros
- + Accurate cost tracking on streaming paths
- + Low latency UX plus usage reporting
- + Same result via stream() and invoke()
cons
- − Requires upgrading to langchain-core 1.4.8+
- − Usage only guaranteed on the final event
Takeaways
- 1Streaming no longer hides your token counts.
- 2Check final_event.usage, not the earlier chunks.
- 3Upgrade langchain-core to at least 1.4.8 to get the fix.