langchain · difficulty ◆◆
ChatOpenAI Now Surfaces Gateway Metadata
Find out whether your LLM call was served from cache, and which region handled it, straight off the AIMessage.
Your LLM gateway knew whether it served from cache — now your AIMessage finally knows too.
$ pip install -U langchain-openai==1.5.2a1 langchain-core==1.5.6What it does
When your OpenAI calls run through a LangSmith gateway — an LLM proxy that sits between your app and the provider for caching, routing, and tracing — the gateway embeds useful context into the response HTTP headers: cache hit or miss, the underlying provider or region that served the request, gateway latency, and trace identifiers. In langchain-openai 1.5.2a1, ChatOpenAI parses those headers after each response and attaches the parsed values to the result's response_metadata. You read them with result.response_metadata["gateway"], without writing your own header-parsing code. When no gateway is present, the field is simply absent — a strict superset of normal usage.
Why it matters
In production, teams increasingly route LLM traffic through a gateway to enforce budgets, cache expensive calls, and trace requests across providers. Today you can finally observe what the gateway actually did from inside LangChain — was this answer served from cache? which upstream region handled my request? — and log it alongside the response. That turns a black-box proxy into a measurable part of your pipeline: you can build cost dashboards, detect cache thrash, and verify gateway routing from a single line of code. Combined with the core change in langchain-core 1.5.6 that pushes the same metadata into LangSmith traces, you get end-to-end visibility without hand-rolling HTTP header parsing.
Example
$ Point ChatOpenAI at a LangSmith gateway and read the extracted gateway metadata off the responsefrom langchain_openai import ChatOpenAI
llm = ChatOpenAI(
model="gpt-4o-mini",
base_url="http://localhost:8080", # your gateway endpoint
api_key="your-key",
temperature=0,
)
result = llm.invoke("What is the capital of Austria?")
print("Answer:", result.content)
meta = result.response_metadata
print("Model:", meta.get("model"))
print("Gateway metadata:", meta.get("gateway", "no gateway header present"))
gw = meta.get("gateway") or {}
print("Cache hit?", gw.get("cache"))
print("Upstream region:", gw.get("region"))
print("Gateway latency ms:", gw.get("latency_ms"))The gateway field mirrors whatever the gateway emits in headers.
$ Same call with no gateway — the feature is a no-opAnswer: Vienna
Model: gpt-4o-mini
Gateway metadata: no gateway header presentWithout a base_url gateway, response_metadata.get("gateway") returns None. Everything else still works.
Common flags
- ChatOpenAI.invoke
- Sends a prompt to the OpenAI-compatible endpoint and returns an AIMessage carrying response_metadata.
- AIMessage.response_metadata
- Dict where the gateway-extracted header metadata now lands.
- get_num_tokens_from_messages
- Fixed in this release to correctly support OpenAI o-series reasoning models.
- with_structured_output
- Builds typed, structured responses that also carry gateway metadata.
History
From black box to measurable proxy
LLM gateways grew up fast — what began as a simple reverse proxy for a single provider became a routing, caching, and tracing layer that teams depend on for cost control. But the app calling the gateway had no idea what the gateway actually did: was that answer cheap because it came from cache, or expensive because it hit a distant region? The gateway already knew, but only in HTTP headers nobody bothered to read. This change wires that knowledge back into the LangChain objects you already hold — a small patch release that makes the proxy legible.
Fun facts
Pros & cons
pros
- + Read gateway behavior with zero header-parsing code
- + Strictly additive — no-op without a gateway
- + Pairs with core change for full LangSmith trace visibility
- + Enables cost dashboards and cache-thrash detection
cons
- − Key shape mirrors the gateway, so it varies by gateway
- − Only meaningful when a LangSmith gateway sits in front
Takeaways
- 1Upgrade to langchain-openai 1.5.2a1 + langchain-core 1.5.6.
- 2Point ChatOpenAI at your gateway and read response_metadata["gateway"].
- 3Turn on LANGCHAIN_TRACING_V2 to see the same metadata in traces.
- 4Log cache status and region per call to start building cost observability.