This is a known limitation that is already warned about — the issue is that the warning under-states what actually happens, and the consequence is a vendor misroute of cluster data.
Stating that up front because the code is not unaware: config.py does warn.
if self.LLM_PROVIDER == "anthropic":
if not self.CORTEX_V4_ENABLED:
logging.warning(
"LLM_PROVIDER=anthropic is only used by the V4 cortex — "
"set CORTEX_V4_ENABLED=true (the V2 graph supports azure/openai only)."
)
So this is not "nobody noticed". It is a judgement about severity and wording.
What actually happens
LLM_PROVIDER=anthropic on a default install (CORTEX_V4_ENABLED=false), with a valid Anthropic key present:
LLM_PROVIDER : anthropic
ANTHROPIC_API_KEY set : True
ANTHROPIC_LARGE_MODEL : claude-sonnet-4-6
coordinator:
client class : langchain_openai.chat_models.base.ChatOpenAI
model : gpt-4o
base_url : https://api.openai.com/v1 (default)
api key used : sk-openai-LEFTOVER-KEY...
subagent:
client class : langchain_openai.chat_models.base.ChatOpenAI
model : gpt-4o-mini
base_url : https://api.openai.com/v1 (default)
api key used : sk-openai-LEFTOVER-KEY...
The run proceeds normally. Every prompt — pod specs, events, log excerpts, namespace and resource names — goes to OpenAI, using OPENAI_API_KEY, for a user who explicitly selected Anthropic.
Why the warning is not sufficient here
Two specific gaps, both fixable in a few lines:
1. It describes the cause, not the consequence. "is only used by the V4 cortex" reads naturally as "Anthropic will not be used" — i.e. it will fail, or do nothing. It does not say "your cluster data will instead be sent to OpenAI using OPENAI_API_KEY." Those are very different things to a reader deciding whether to act on a log line.
2. It warns where it arguably should refuse. Compare the neighbouring behaviour: an unknown LLM_PROVIDER raises ValueError and stops startup. A known provider that silently routes to a different vendor only logs. The more consequential of the two is the one that continues.
And it only surfaces at all when OPENAI_API_KEY happens to be unset. When it is set — extremely common, whether left over from an earlier config or present for another tool — everything simply works, correctly-looking, against the wrong vendor.
Why this matters for this project specifically
Discussion #83 is "What cluster data is sent to the LLM provider, and can I keep it on-premises?", and v4/docs/data-handling.md exists precisely to answer it. Provider choice here is frequently a compliance decision — a data-processing agreement with one vendor and not another. Silently satisfying it with a different vendor is the failure mode those documents exist to prevent.
It is the same shape as safety invariant #1, moved from the diagnosis layer to the routing layer: the system presents a result (a working session on "anthropic") that does not correspond to what actually happened.
Scope — the Cortex path is correct, do not change it
app/cortex/models.py::_tier() handles this properly and is the reference implementation:
if provider == "anthropic":
model = settings.ANTHROPIC_SMALL_MODEL if small else settings.ANTHROPIC_LARGE_MODEL
return _anthropic(model, streaming)
including a clear ImportError → RuntimeError message when langchain-anthropic is absent (it is an optional extra and not a workspace dependency).
The gap is only in app/core/llm.py, which the default V2 graph uses:
def _coordinator_llm() -> BaseChatModel:
if settings.LLM_PROVIDER == "azure":
return _make_azure(...)
return _make_openai(...) # anthropic and qwen both land here
llm.py contains no occurrence of the string anthropic at all.
Possible directions — the call is the maintainer's
Not obvious which is right; please argue it rather than assume:
- Fail closed. Raise on
LLM_PROVIDER=anthropic without CORTEX_V4_ENABLED, the way an unknown provider already raises. Safest, and a one-line change — but it turns a working (if surprising) configuration into a hard startup failure for anyone relying on it today.
- Implement it in
core/llm.py, mirroring cortex/models.py. Makes the setting mean what it says everywhere. Largest change, and pulls in the optional langchain-anthropic dependency question.
- Keep warning, but say the consequence — name OpenAI and
OPENAI_API_KEY explicitly in the message. Cheapest, and strictly better than today even if (1) or (2) lands later.
(3) is worth doing regardless of which of (1)/(2) is chosen.
Acceptance criteria
No cluster or API key needed — the reproduction only inspects the constructed client object; nothing is ever called.
Task metadata
|
|
| Difficulty |
Intermediate — small diff, the decision is the work |
| Skills |
Python. No Kubernetes, no API key. |
| Mentor |
@MSKazemi — ask right here |
| Status |
🟢 Unclaimed — comment "I'd like to take this" and it is yours |
Found while checking whether #17 was stale (it is not — that issue is accurately scoped).
This is a known limitation that is already warned about — the issue is that the warning under-states what actually happens, and the consequence is a vendor misroute of cluster data.
Stating that up front because the code is not unaware:
config.pydoes warn.So this is not "nobody noticed". It is a judgement about severity and wording.
What actually happens
LLM_PROVIDER=anthropicon a default install (CORTEX_V4_ENABLED=false), with a valid Anthropic key present:The run proceeds normally. Every prompt — pod specs, events, log excerpts, namespace and resource names — goes to OpenAI, using
OPENAI_API_KEY, for a user who explicitly selected Anthropic.Why the warning is not sufficient here
Two specific gaps, both fixable in a few lines:
1. It describes the cause, not the consequence. "is only used by the V4 cortex" reads naturally as "Anthropic will not be used" — i.e. it will fail, or do nothing. It does not say "your cluster data will instead be sent to OpenAI using
OPENAI_API_KEY." Those are very different things to a reader deciding whether to act on a log line.2. It warns where it arguably should refuse. Compare the neighbouring behaviour: an unknown
LLM_PROVIDERraisesValueErrorand stops startup. A known provider that silently routes to a different vendor only logs. The more consequential of the two is the one that continues.And it only surfaces at all when
OPENAI_API_KEYhappens to be unset. When it is set — extremely common, whether left over from an earlier config or present for another tool — everything simply works, correctly-looking, against the wrong vendor.Why this matters for this project specifically
Discussion #83 is "What cluster data is sent to the LLM provider, and can I keep it on-premises?", and
v4/docs/data-handling.mdexists precisely to answer it. Provider choice here is frequently a compliance decision — a data-processing agreement with one vendor and not another. Silently satisfying it with a different vendor is the failure mode those documents exist to prevent.It is the same shape as safety invariant #1, moved from the diagnosis layer to the routing layer: the system presents a result (a working session on "anthropic") that does not correspond to what actually happened.
Scope — the Cortex path is correct, do not change it
app/cortex/models.py::_tier()handles this properly and is the reference implementation:including a clear
ImportError→RuntimeErrormessage whenlangchain-anthropicis absent (it is an optional extra and not a workspace dependency).The gap is only in
app/core/llm.py, which the default V2 graph uses:llm.pycontains no occurrence of the stringanthropicat all.Possible directions — the call is the maintainer's
Not obvious which is right; please argue it rather than assume:
LLM_PROVIDER=anthropicwithoutCORTEX_V4_ENABLED, the way an unknown provider already raises. Safest, and a one-line change — but it turns a working (if surprising) configuration into a hard startup failure for anyone relying on it today.core/llm.py, mirroringcortex/models.py. Makes the setting mean what it says everywhere. Largest change, and pulls in the optionallangchain-anthropicdependency question.OPENAI_API_KEYexplicitly in the message. Cheapest, and strictly better than today even if (1) or (2) lands later.(3) is worth doing regardless of which of (1)/(2) is chosen.
Acceptance criteria
anthropicon a default install cannot end up sending cluster data to OpenAI without that being unmissableenv -iwithLLM_PROVIDER=anthropic+ both keys, then assert on the constructed client's class andbase_url)v4/docs/configuration.mdstates which providers each graph supportsv4/tests/suite greenNo cluster or API key needed — the reproduction only inspects the constructed client object; nothing is ever called.
Task metadata
Found while checking whether #17 was stale (it is not — that issue is accurately scoped).