LiteLLM Integration
Overview
LiteLLM provides a unified interface for calling 100+ LLM providers. Fiddler supports two integration modes:
Both modes work by routing OpenTelemetry traces to Fiddler’s OTLP ingestion endpoint using standard environment variables.
LiteLLM SDK Integration
Overview
LiteLLM includes a built-in OpenTelemetry integration. When you enable it and point the OTLP exporter at Fiddler, every LLM call is automatically traced — with no Fiddler-specific package required. Fiddler natively ingests LiteLLM SDK-generated OTel traces and maps them to the Fiddler schema, giving you full observability over prompts, responses, and token usage across all LLM providers. The following SDK functions are supported:Notes on the table above:
completionoperation name: LiteLLM versions before1.82.1(released January 2026) emitgen_ai.operation.name = "completion"literally forlitellm.completion()calls. Newer versions rewrite it to"chat". Both are classified identically asllm.- Non-text APIs classified as
chain: Fiddler’s LLM observability currently focuses on text-based generative completions. Image, audio, embedding, moderation, ranking, and OCR operations are classified aschainso they remain visible in traces without being treated as LLM completions.
Architecture
Prerequisites
- Fiddler account with a GenAI application already created
pip install litellm(oruv add litellm)- A valid LLM provider API key (e.g.
OPENAI_API_KEYfor OpenAI models)
Quick Start
Step 1: Set environment variables
Set these before starting your application:Step 2: Enable the built-in OTel callback
Add one line to your application startup:Step 3: Make completions as normal
No other code changes are required:Step 4: Verify traces are arriving
Open the Fiddler UI and navigate to your application’s Explorer. You should see the trace within a few seconds of making your first completion call.What Gets Captured
Message Content
Token Usage
Model Information
gen_ai.system and gen_ai.request.model are SDK first-class LLM attributes. They are stored at their unprefixed keys and resolved at query time by the Fiddler backend’s field registry, making them queryable via SpanAttribute::gen_ai.system and SpanAttribute::gen_ai.request.model.
Supported Features
Troubleshooting
Traces not appearing in Fiddler Check that all three environment variables are set correctly:litellm.callbacks = ["otel"] is set before your first litellm.completion() call.
Check the fiddler-application-id header and application.id resource attribute are both set
Both are required. fiddler-application-id must be a valid UUID for an existing Fiddler application, otherwise spans will be dropped during ingestion.
LiteLLM Proxy Integration
Overview
LiteLLM is an OpenAI-compatible proxy gateway that lets you call 100+ LLM providers through a single API. When LiteLLM proxy is configured to emit OpenTelemetry traces, Fiddler automatically detects and ingests them — no additional SDK or code changes required. Fiddler includes a purpose-built mapper for LiteLLM proxy traces that handles the proxy’s specific span format, attribute layout, and operation naming conventions. This gives you full observability over every LLM call routed through your proxy: prompts, responses, token usage, cost metadata, and latency — across all models and providers in one place.Architecture
When to Use This Integration
Use the LiteLLM proxy integration when:- You are already running LiteLLM proxy as your LLM gateway
- You want to monitor all LLM traffic centrally regardless of underlying provider (OpenAI, Anthropic, Bedrock, etc.)
- You want cost attribution and latency tracking without instrumenting individual applications
Quick Start
Step 1: Configure LiteLLM proxy to emit OpenTelemetry
Set the following environment variables before starting the proxy:config.yaml:
Step 2: Set your Fiddler application ID
Two environment variables carry your application ID and both are required:OTEL_RESOURCE_ATTRIBUTES— setsapplication.idon every OTel resource, which Fiddler uses to route traces to the correct applicationOTEL_EXPORTER_OTLP_HEADERS— includesfiddler-application-idas an HTTP header for authentication and routing at the ingestion endpoint
Step 3: Verify traces are arriving
Make a test request through your proxy:Using with OpenAI Codex CLI
OpenAI Codex CLI does not need its own Fiddler integration. Codex’s native OpenTelemetry traces carry only token counts — no prompt, response, or tool content. (That content exists only in Codex’s OpenTelemetry logs, which Fiddler does not ingest.) So the recommended way to observe Codex in Fiddler is to point it at your LiteLLM proxy. The proxy then emits fullllm spans (prompt + response + token usage) that Fiddler ingests through the integration described above, with no Fiddler-side code.
Step 1: Use LiteLLM 1.88+ with a Codex-compatible model
Codex uses the OpenAI Responses API, and its streamed output is only captured by LiteLLM 1.88.0 or later. Ensure the proxy is on a supported version with the OTel exporter SDK installed:config.yaml, map an alias (the name Codex will request) to a Codex-capable model. Use a current Codex-family model from your OpenAI account for the upstream model — for example gpt-5.2-codex. General chat models such as gpt-4o-mini will not work, because they reject Codex’s tool calls.
Step 2: Point Codex at the proxy
Add a custom model provider to~/.codex/config.toml so Codex sends requests to your LiteLLM proxy instead of OpenAI directly:
Step 3: Use Codex normally
llm span with the full user prompt, assistant response, and token usage.
Codex uses the OpenAI Responses API. On the proxy path with LiteLLM ≥ 1.88.0, LiteLLM captures both the input and output messages for Responses-API calls, so prompts and responses appear in full. The
/v1/responses content gap noted in Known LiteLLM Upstream Caveats applies only to earlier LiteLLM versions and does not affect this setup.What Gets Captured
Span Types
LiteLLM proxy emits several span types per request. Fiddler classifies them based on thegen_ai.operation.name attribute:
LLM endpoints — classified as llm (generative text completions):
Non-LLM endpoints — classified as
chain (not generative text completions):
Infrastructure spans — classified as
chain:
Each proxy request typically produces up to 3 spans:
Captured Attributes
Message Content LiteLLM writes full conversation history as JSON on the span (not as span events). Fiddler extracts:If you have disabled message logging in LiteLLM (
turn_off_message_logging: true), the message content fields will be absent from traces. Token counts and cost metadata are still captured.
Model Information
gen_ai.system and gen_ai.request.model are SDK first-class LLM attributes. They are stored at their unprefixed keys and resolved at query time by the Fiddler backend’s field registry, making them queryable via SpanAttribute::gen_ai.system and SpanAttribute::gen_ai.request.model.
Cost Metadata (stored as
fiddler.span.user.*)
LiteLLM emits cost fields under gen_ai.cost.*. These are preserved in Fiddler as user-visible span attributes:
Proxy Metadata (stored as
fiddler.span.user.*)
LiteLLM proxy emits metadata.* attributes containing API key, team, user, and routing information. These are preserved as user-visible span attributes for auditing and cost attribution.
Supported Features
Endpoint Coverage
Platform Features
Troubleshooting
Traces not appearing in Fiddler Check that OTel is enabled in LiteLLM:fiddler-application-id header and application.id resource attribute are both set:
Both are required. fiddler-application-id must be a valid UUID for an existing Fiddler application, otherwise spans will be dropped during ingestion.
Check service.name is "litellm"
Fiddler detects LiteLLM proxy spans by service.name. LiteLLM proxy sets this to "litellm" by default. If you have overridden OTEL_SERVICE_NAME, ensure it is set to "litellm" or "litellm-proxy":
false to re-enable message capture.
Spans classified as chain instead of llm
This happens for internal LiteLLM infrastructure spans (self, router, proxy_pre_call) and for non-completion operations (embeddings, image generation, speech, etc.). This is expected behavior — only completion-generating endpoints (/chat/completions, /completions, /v1/responses) are classified as llm spans.
/v1/responses spans are missing message content
See the Known LiteLLM Upstream Caveats section below for details. Token counts, costs, and span-type classification are unaffected.
Known LiteLLM Upstream Caveats
While integrating with LiteLLM, several gaps were identified in LiteLLM’s own OpenTelemetry callback (litellm/integrations/opentelemetry.py). These are not Fiddler issues — they affect every downstream OTel consumer (Datadog, Honeycomb, Phoenix, etc.). Fiddler classifies the spans correctly and surfaces every attribute that LiteLLM does emit, but the gaps below mean some content is simply absent from the trace at the source.
/v1/responses and /responses — input and output messages both missing
LiteLLM’s OTel callback reads input messages from kwargs["messages"] and output messages from response["choices"] — both shapes specific to /chat/completions. The Responses API uses kwargs["input"] and response["output"] instead, so neither extraction block runs.
Where the data does exist: the
raw_gen_ai_request child span (a sibling of the parent litellm_request span) carries both the request and response under llm.<provider>.input / llm.<provider>.output. It is currently surfaced as a chain span without content extraction.
Tracking: BerriAI/litellm#25840
/v1/responses, /responses, /v1/messages, /v1beta/...:generateContent — system prompt missing
LiteLLM’s OTel callback writes gen_ai.system_instructions only when the kwarg name is exactly system_instructions. Other endpoints use different field names for the same concept:
The system prompt does reach LiteLLM and is included in the actual LLM request — it just never lands on
gen_ai.system_instructions in the OTel trace. As with the output-text gap, the data is visible on the raw_gen_ai_request child span (llm.<provider>.instructions / llm.<provider>.system / llm.<provider>.systemInstruction).
Tracking: BerriAI/litellm#25840 (follow-up comment)
Non-chat-completion endpoints — gen_ai.system empty and llm.None.* attribute prefix
For every endpoint family except /chat/completions, LiteLLM’s custom_llm_provider is not propagated into the OTel callback’s view of litellm_params. This causes two visible symptoms:
gen_ai.systemis set to an empty string instead of the provider (e.g."openai","vertex_ai","anthropic").- Raw provider attributes on the
raw_gen_ai_requestchild span use allm.None.*prefix (e.g.llm.None.output,llm.None.model) instead ofllm.openai.*orllm.vertex_ai.*.
/v1/messages and Gemini may need follow-up after merge).
Gemini streaming variant — gen_ai.response.model not set
/v1beta/models/{model}:streamGenerateContent does not emit gen_ai.response.model on the parent span, even though the non-streaming :generateContent variant does. Likely lives in LiteLLM’s Gemini streaming aggregation path. Low impact; not yet filed upstream.
Summary — what works and what doesn’t, by endpoint
These caveats will resolve as the upstream LiteLLM PRs land. Fiddler will pick up the improvements automatically — no Fiddler-side changes will be needed when LiteLLM fixes ship.
Guardrails
Fiddler implements the LiteLLM Generic Guardrail API spec, letting you plug Fiddler’s real-time guardrails directly into the LiteLLM proxy gateway. Once configured, every request routed through the proxy is checked for secrets, PII, and unsafe content before it reaches the model. See LiteLLM Guardrails for setup instructions, check behavior, and the full API reference.Related Documentation
- OpenTelemetry Integration — Manual OTel instrumentation for custom frameworks
- Strands Agents SDK — Native monitoring for Strands agent applications
- LangGraph SDK — Auto-instrumentation for LangGraph applications
- LiteLLM OTel documentation — LiteLLM’s official OpenTelemetry setup guide