langchain · difficulty ◆◆
Usage on Every Streaming Chunk
Track token spend while tokens are still arriving.
A live token meter while the answer is still typing itself.
$ chunk.usage_metadataWhat it does
With the v3 streaming fix in langchain-core 1.4.8, AIMessageChunk objects yielded during streaming now carry forward usage data as each chunk arrives. Instead of losing it until the final AIMessage, the cumulative counts live on AIMessageChunk.usage_metadata per chunk.
Why it matters
Real apps need per-token cost tracking and live token meters. Before the fix, streaming responses were a black box for usage. Now you can build live cost estimators that update spend as tokens arrive, streaming-aware budgets that halt generation mid-stream, and debugging tools that correlate usage with response segments.
Example
$ total_tokens = 0
chunks_with_usage = 0
for chunk in llm.stream("Explain why the sky is blue in 3 sentences."):
if chunk.usage_metadata:
total_tokens += chunk.usage_metadata.get("output_tokens", 0)
chunks_with_usage += 1
if chunk.content:
print(chunk.content, end="", flush=True)Explain why the sky is blue in 3 sentences.
Chunk usage: {'input_tokens': 8, 'output_tokens': 3, 'total_tokens': 11}
Chunk usage: {'input_tokens': 8, 'output_tokens': 7, 'total_tokens': 15}
Chunk usage: {'input_tokens': 8, 'output_tokens': 14, 'total_tokens': 22}
Total output tokens (from chunks): 21
Chunks with usage metadata: 4input_tokens repeats per chunk as the running context total; output_tokens is cumulative.
Common flags
- ChatOpenAI.stream()
- Streams AIMessageChunk objects with usage metadata per chunk
- usage_metadata
- Dict with input_tokens, output_tokens, total_tokens
- llm.invoke()
- Non-streaming call returning a single AIMessage with full usage
History
Closing the streaming gap
invoke() always reported full usage, but streaming could not. This release restores the missing link, making stream() as informative as invoke() for metering purposes.
Fun facts
Pros & cons
pros
- + Live cost estimators
- + Streaming-aware budgets
- + Debug usage per segment
cons
- − input_tokens is the running total, not per-chunk cost
- − Only guaranteed on chunk objects in 1.4.8+
Takeaways
- 1Read usage_metadata on every chunk, not just the last message.
- 2output_tokens is cumulative; input_tokens is the running total.
- 3Live cost meters are now trivial to build.