Repository navigation
feat(observability): fix duplicate LiteLLM/LangChain traces in Langfuse - #3
Merged
Merged
Conversation
… Langfuse - Update litellm config to use langfuse_otel callbacks - Inject Langfuse OTel environment variables to LiteLLM proxy in docker-compose and Kubernetes
…outing - swapped /health with /health/liveliness in docker and k8s probes to stop groq quota drain - added 'order' to models in litellm config (70B -> 8B -> Ollama) - killed useless retries on 429s. router now instantly falls back to the secondary model with a 60s cooldown for the primary.
…rings - llm_provider now dynamically injects the active opentelemetry context (traceparent) into default_headers. this forces litellm to stitch its proxy spans under our existing langfuse trace instead of spawning orphaned roots. - rewrote docstrings in a pragmatic, straight-to-the-point style.
- implemented AgentOnlyLangfuseHandler to override on_llm_start and related methods. - this stops langchain from logging redundant ChatOpenAI generations, ensuring litellm is the sole source of truth for model calls and costs. - the agent tree (langgraph, chains, tools) remains fully visible and intact.
- opted into gen_ai_latest_experimental semconv in litellm to drop noisy raw_gen_ai_request spans and adopt standard chat span naming. - added proper OTEL_SERVICE_NAME tags for api and litellm gateway. - cleanly separated environment tags: 'local' for docker-compose, 'production' for k8s. - fixed litellm kubernetes readiness probe endpoint.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Spent a while chasing why every LLM call showed up twice in Langfuse —
once from LangChain and once from LiteLLM — while the LiteLLM records
were appearing outside the main application trace.
This branch fixes that, along with a few issues I found on the way.
What's in here:
configuration attempt was silently ignored
consuming Groq's daily quota
retry and fallback behaviour
provider spans to attach to the existing expedition_pipeline trace
instead of creating separate traces
duplicate ChatOpenAI generation; agents, chains, tools and retrievers
are still traced normally
End result: one trace per request, one owner for model-generation data,
and the real provider, model, token usage and cost visible in Langfuse.
The 70B → 8B fallback was verified locally by deliberately breaking the
primary deployment. The Ollama fallback and Kubernetes deployment still
need separate verification.
I haven't deployed the Kubernetes side yet. The previous GKE node pool
couldn't fit LiteLLM reliably, so that remains a separate infrastructure
problem.