Skip to content

feat(observability): fix duplicate LiteLLM/LangChain traces in Langfuse - #3

Merged
Hilalalpak merged 6 commits into
mainfrom
feature/litellm-langfuse-tracing
Aug 4, 2026
Merged

Hilalalpak merged 6 commits into
mainfrom
feature/litellm-langfuse-tracing

Conversation

@Hilalalpak

Copy link
Copy Markdown
Owner

Spent a while chasing why every LLM call showed up twice in Langfuse —
once from LangChain and once from LiteLLM — while the LiteLLM records
were appearing outside the main application trace.

This branch fixes that, along with a few issues I found on the way.

What's in here:

  • Wired up LiteLLM's langfuse_otel callback correctly after the first
    configuration attempt was silently ignored
  • Fixed a health probe that was making real model calls and quietly
    consuming Groq's daily quota
  • Added explicit model priority: Groq 70B → Groq 8B → local Ollama
  • Disabled unnecessary application-level retries so LiteLLM owns the
    retry and fallback behaviour
  • Injected the W3C trace context into LiteLLM requests, allowing its
    provider spans to attach to the existing expedition_pipeline trace
    instead of creating separate traces
  • Added a small LangChain callback override so it no longer creates a
    duplicate ChatOpenAI generation; agents, chains, tools and retrievers
    are still traced normally

End result: one trace per request, one owner for model-generation data,
and the real provider, model, token usage and cost visible in Langfuse.

The 70B → 8B fallback was verified locally by deliberately breaking the
primary deployment. The Ollama fallback and Kubernetes deployment still
need separate verification.

I haven't deployed the Kubernetes side yet. The previous GKE node pool
couldn't fit LiteLLM reliably, so that remains a separate infrastructure
problem.

… Langfuse

- Update litellm config to use langfuse_otel callbacks
- Inject Langfuse OTel environment variables to LiteLLM proxy in docker-compose and Kubernetes
…outing

- swapped /health with /health/liveliness in docker and k8s probes to stop groq quota drain
- added 'order' to models in litellm config (70B -> 8B -> Ollama)
- killed useless retries on 429s. router now instantly falls back to the secondary model with a 60s cooldown for the primary.
…rings

- llm_provider now dynamically injects the active opentelemetry context (traceparent) into default_headers. this forces litellm to stitch its proxy spans under our existing langfuse trace instead of spawning orphaned roots.
- rewrote docstrings in a pragmatic, straight-to-the-point style.
- implemented AgentOnlyLangfuseHandler to override on_llm_start and related methods.
- this stops langchain from logging redundant ChatOpenAI generations, ensuring litellm is the sole source of truth for model calls and costs.
- the agent tree (langgraph, chains, tools) remains fully visible and intact.
- opted into gen_ai_latest_experimental semconv in litellm to drop noisy raw_gen_ai_request spans and adopt standard chat span naming.
- added proper OTEL_SERVICE_NAME tags for api and litellm gateway.
- cleanly separated environment tags: 'local' for docker-compose, 'production' for k8s.
- fixed litellm kubernetes readiness probe endpoint.
@Hilalalpak
Hilalalpak merged commit 66d61ca into main Aug 4, 2026
1 check passed
@Hilalalpak
Hilalalpak deleted the feature/litellm-langfuse-tracing branch August 4, 2026 23:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant