kmail.at
← learning

langchain · difficulty ◆◆

Usage on Every Streaming Chunk

Track token spend while tokens are still arriving.

A live token meter while the answer is still typing itself.

2026-06-21 · 6 min read

$ chunk.usage_metadata

What it does

With the v3 streaming fix in langchain-core 1.4.8, AIMessageChunk objects yielded during streaming now carry forward usage data as each chunk arrives. Instead of losing it until the final AIMessage, the cumulative counts live on AIMessageChunk.usage_metadata per chunk.

Why it matters

Real apps need per-token cost tracking and live token meters. Before the fix, streaming responses were a black box for usage. Now you can build live cost estimators that update spend as tokens arrive, streaming-aware budgets that halt generation mid-stream, and debugging tools that correlate usage with response segments.

Example

$ total_tokens = 0
chunks_with_usage = 0
for chunk in llm.stream("Explain why the sky is blue in 3 sentences."):
    if chunk.usage_metadata:
        total_tokens += chunk.usage_metadata.get("output_tokens", 0)
        chunks_with_usage += 1
    if chunk.content:
        print(chunk.content, end="", flush=True)
Explain why the sky is blue in 3 sentences.
  Chunk usage: {'input_tokens': 8, 'output_tokens': 3, 'total_tokens': 11}
  Chunk usage: {'input_tokens': 8, 'output_tokens': 7, 'total_tokens': 15}
  Chunk usage: {'input_tokens': 8, 'output_tokens': 14, 'total_tokens': 22}

Total output tokens (from chunks): 21
Chunks with usage metadata: 4

input_tokens repeats per chunk as the running context total; output_tokens is cumulative.

Common flags

ChatOpenAI.stream()
Streams AIMessageChunk objects with usage metadata per chunk
usage_metadata
Dict with input_tokens, output_tokens, total_tokens
llm.invoke()
Non-streaming call returning a single AIMessage with full usage

History

Closing the streaming gap

invoke() always reported full usage, but streaming could not. This release restores the missing link, making stream() as informative as invoke() for metering purposes.

Fun facts

Pros & cons

pros

  • + Live cost estimators
  • + Streaming-aware budgets
  • + Debug usage per segment

cons

  • − input_tokens is the running total, not per-chunk cost
  • − Only guaranteed on chunk objects in 1.4.8+

Takeaways

  1. 1Read usage_metadata on every chunk, not just the last message.
  2. 2output_tokens is cumulative; input_tokens is the running total.
  3. 3Live cost meters are now trivial to build.

Related commands

← all learning