Summary
A real kq session hit a deterministic planning failure while trying to solve several cluster problems. The coordinator emitted an unresolved metavariable as an actual tool argument:
You: solve ths issues
⚙ run_kubectl → Running: kubectl describe pod ml-scorer-5bc9565b96-5tskr -n shop
⚙ run_kubectl → Running: kubectl describe pod payments-api-56f659864c-28s2w -n shop
⚙ run_kubectl → Running: kubectl get nodes -o wide
⚙ run_kubectl → Running: kubectl describe node <node-name>
...
Error: LLM error: Command contains disallowed shell characters ...
run_kubectl is correct to reject < and >: they are shell metacharacters and the runner intentionally blocks them. The bug is upstream: a placeholder that exists only as planning notation reached the execution surface.
Why this is a coordinator bug, not a sanitizer bug
The v4 coordinator prompt already says dependent calls must be sequential, but its parallel-execution guidance includes a describe node example without establishing a concrete node first. In this trace the model correctly knew it needed kubectl get nodes, but incorrectly scheduled kubectl describe node <node-name> in the same batch instead of waiting for the node name.
Do not fix this by weakening _SHELL_METACHAR in app/tools/kubectl_tool.py. The safety gate should continue rejecting < / >.
Expected behavior
If a command needs an identifier that is not already present in the cluster snapshot or conversation state, the coordinator should:
- fetch the identifier first (
kubectl get nodes -o wide),
- wait for that result,
- emit a second tool call using the concrete node name.
No metavariable such as <node-name>, <pod>, <namespace>, $NODE, {node}, or similar planning placeholder should ever be passed to run_kubectl.
Suggested scope
- Tighten the coordinator system prompt around dependent calls and replace the ambiguous parallel
describe node example.
- Add a pre-execution/model-output guard that detects unresolved metavariables in tool arguments and routes them back for replanning rather than invoking the tool.
- Prefer resolving names from the existing snapshot before doing another discovery call when possible.
Acceptance criteria
Reproduction context
The user first asked do you see any issues in my cluster; KubeIntellect identified ImagePullBackOff, CrashLoopBackOff, and an unschedulable Pending pod in namespace shop. The next request, solve ths issues, produced the batch above and failed on the placeholder command.
Summary
A real
kqsession hit a deterministic planning failure while trying to solve several cluster problems. The coordinator emitted an unresolved metavariable as an actual tool argument:run_kubectlis correct to reject<and>: they are shell metacharacters and the runner intentionally blocks them. The bug is upstream: a placeholder that exists only as planning notation reached the execution surface.Why this is a coordinator bug, not a sanitizer bug
The v4 coordinator prompt already says dependent calls must be sequential, but its parallel-execution guidance includes a
describe nodeexample without establishing a concrete node first. In this trace the model correctly knew it neededkubectl get nodes, but incorrectly scheduledkubectl describe node <node-name>in the same batch instead of waiting for the node name.Do not fix this by weakening
_SHELL_METACHARinapp/tools/kubectl_tool.py. The safety gate should continue rejecting</>.Expected behavior
If a command needs an identifier that is not already present in the cluster snapshot or conversation state, the coordinator should:
kubectl get nodes -o wide),No metavariable such as
<node-name>,<pod>,<namespace>,$NODE,{node}, or similar planning placeholder should ever be passed torun_kubectl.Suggested scope
describe nodeexample.Acceptance criteria
run_kubectlcontinues to reject<and>as shell metacharacters.run_kubectlwith an unresolved placeholder/metavariable.get nodesfirst anddescribe node <concrete-name>only after the result is available.get nodes+describe node <node-name>failure mode.Reproduction context
The user first asked
do you see any issues in my cluster; KubeIntellect identified ImagePullBackOff, CrashLoopBackOff, and an unschedulable Pending pod in namespaceshop. The next request,solve ths issues, produced the batch above and failed on the placeholder command.