Skip to content

Latest commit

 

History

History
48 lines (35 loc) · 1.82 KB

File metadata and controls

48 lines (35 loc) · 1.82 KB

NanoCode Evaluation

This lightweight evaluation checks NanoCode's agent mechanics and safety boundaries. It does not try to benchmark the intelligence of the selected model.

Evaluated on 2026-08-22 with Node.js 24.3.0. Automated checks use deterministic fake model responses and consume no API credits.

Automated checks

Check Expected behavior Result
Agent tool loop Execute Grep, return its result, then accept the final answer Pass
Structured answer content Render text parts instead of [object Object] Pass
Denied edit Return a structured denial and leave the file unchanged Pass
Outside-project read Reject a path that escapes the project root Pass
Exact edit Replace one unique occurrence Pass
Ambiguous edit Reject multiple matching occurrences without changing the file Pass
JSONL resume Save a session and load its messages again Pass
Context compaction Summarize older messages and preserve recent messages Pass

Run the suite with:

npm test

Terminal smoke checks

The real TUI was run in a pseudo-terminal and checked for:

  • Neon header, project path, input editor, and status footer
  • /help, /save, --continue, and /exit
  • Approval overlay with Deny selected by default
  • Clean terminal restoration after exit

Syntax checks and git diff --check also pass.

What this does not prove

  • Answer quality across models or a larger prompt dataset
  • Security isolation for an approved Bash command
  • Exact token counts, because context size is estimated
  • Reliability under very large repositories or long-running sessions
  • Cross-platform rendering outside the tested macOS terminal environment

These are honest scope boundaries for a small educational coding agent rather than missing claims of production equivalence.