This lightweight evaluation checks NanoCode's agent mechanics and safety boundaries. It does not try to benchmark the intelligence of the selected model.
Evaluated on 2026-08-22 with Node.js 24.3.0. Automated checks use deterministic fake model responses and consume no API credits.
| Check | Expected behavior | Result |
|---|---|---|
| Agent tool loop | Execute Grep, return its result, then accept the final answer |
Pass |
| Structured answer content | Render text parts instead of [object Object] |
Pass |
| Denied edit | Return a structured denial and leave the file unchanged | Pass |
| Outside-project read | Reject a path that escapes the project root | Pass |
| Exact edit | Replace one unique occurrence | Pass |
| Ambiguous edit | Reject multiple matching occurrences without changing the file | Pass |
| JSONL resume | Save a session and load its messages again | Pass |
| Context compaction | Summarize older messages and preserve recent messages | Pass |
Run the suite with:
npm testThe real TUI was run in a pseudo-terminal and checked for:
- Neon header, project path, input editor, and status footer
/help,/save,--continue, and/exit- Approval overlay with Deny selected by default
- Clean terminal restoration after exit
Syntax checks and git diff --check also pass.
- Answer quality across models or a larger prompt dataset
- Security isolation for an approved Bash command
- Exact token counts, because context size is estimated
- Reliability under very large repositories or long-running sessions
- Cross-platform rendering outside the tested macOS terminal environment
These are honest scope boundaries for a small educational coding agent rather than missing claims of production equivalence.