Skip to content

Coordinator emits unresolved kubectl placeholders that are guaranteed to fail the command safety gate #173

Description

@MSKazemi

Summary

A real kq session hit a deterministic planning failure while trying to solve several cluster problems. The coordinator emitted an unresolved metavariable as an actual tool argument:

You: solve ths issues
⚙ run_kubectl → Running: kubectl describe pod ml-scorer-5bc9565b96-5tskr -n shop
⚙ run_kubectl → Running: kubectl describe pod payments-api-56f659864c-28s2w -n shop
⚙ run_kubectl → Running: kubectl get nodes -o wide
⚙ run_kubectl → Running: kubectl describe node <node-name>
...
Error: LLM error: Command contains disallowed shell characters ...

run_kubectl is correct to reject < and >: they are shell metacharacters and the runner intentionally blocks them. The bug is upstream: a placeholder that exists only as planning notation reached the execution surface.

Why this is a coordinator bug, not a sanitizer bug

The v4 coordinator prompt already says dependent calls must be sequential, but its parallel-execution guidance includes a describe node example without establishing a concrete node first. In this trace the model correctly knew it needed kubectl get nodes, but incorrectly scheduled kubectl describe node <node-name> in the same batch instead of waiting for the node name.

Do not fix this by weakening _SHELL_METACHAR in app/tools/kubectl_tool.py. The safety gate should continue rejecting < / >.

Expected behavior

If a command needs an identifier that is not already present in the cluster snapshot or conversation state, the coordinator should:

  1. fetch the identifier first (kubectl get nodes -o wide),
  2. wait for that result,
  3. emit a second tool call using the concrete node name.

No metavariable such as <node-name>, <pod>, <namespace>, $NODE, {node}, or similar planning placeholder should ever be passed to run_kubectl.

Suggested scope

  • Tighten the coordinator system prompt around dependent calls and replace the ambiguous parallel describe node example.
  • Add a pre-execution/model-output guard that detects unresolved metavariables in tool arguments and routes them back for replanning rather than invoking the tool.
  • Prefer resolving names from the existing snapshot before doing another discovery call when possible.

Acceptance criteria

  • run_kubectl continues to reject < and > as shell metacharacters.
  • The coordinator never invokes run_kubectl with an unresolved placeholder/metavariable.
  • A node investigation where the node name is initially unknown emits get nodes first and describe node <concrete-name> only after the result is available.
  • Add a regression test reproducing the get nodes + describe node <node-name> failure mode.
  • The parallel-tool prompt example cannot be interpreted as permission to parallelize a call whose resource name is still unknown.

Reproduction context

The user first asked do you see any issues in my cluster; KubeIntellect identified ImagePullBackOff, CrashLoopBackOff, and an unschedulable Pending pod in namespace shop. The next request, solve ths issues, produced the batch above and failed on the placeholder command.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/serverkubeintellect-serverbugSomething isn't workinghelp wantedExtra attention is needed

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions