Everything you need to run Claude Code (or any AI coding agent) professionally on $20/month.
Not a tutorial. A working system. Drop it in, configure it, and your AI agent becomes a different tool.
Most people use AI coding agents like a smarter autocomplete. They get inconsistent results, hit context limits without warning, lose work between sessions, and wonder why the AI starts agreeing with everything they say halfway through a long session.
The problem isn't the model. It's the setup.
Andrej Karpathy's framing is useful here: LLMs aren't magic boxes — they're deterministic systems with well-understood failure modes. Context decay is real. Sycophancy is a training artifact. "Reasoning" can be performative — right-looking but wrong. Once you understand the failure modes, you can engineer around them.
This kit does that.
| Agent | Config file | Status |
|---|---|---|
| Claude Code | CLAUDE.md |
Native — full support |
| Gemini CLI | GEMINI.md |
Port the config, same principles |
| GitHub Copilot | .github/copilot-instructions.md |
Port the accuracy rules + tool priority |
| Codex CLI | AGENTS.md |
Port the config |
| Cursor / Windsurf | .cursorrules |
Port accuracy rules section |
| Goose | TOM extension + context file | Inject via GOOSE_MOIM_MESSAGE_FILE |
The token gauge, session system, and memory system are Claude Code native. The accuracy rules, tool priority, and skill concepts work anywhere you can inject a system prompt.
Drop-in configuration that gives your AI agent:
- Accuracy rules based on Anthropic's own research on degradation
- Tool priority system (critical — see section below)
- Bash restrictions
- Session management rules
- Output token rules to eliminate response bloat
Run inside Claude Code via /skill-name:
| Skill | What it does |
|---|---|
save-session |
Captures full session state to a dated file. Never lose context again. |
resume-session |
Loads last session and briefs you before touching anything. |
verify |
CC + second model cross-check loop for high-stakes outputs. |
caveman |
Ultra-compressed responses. ~75% token reduction, zero accuracy loss. |
Auto-memory that persists across all conversations:
user/— who you are, preferences, expertise levelproject/— ongoing work, decisions, blockersfeedback/— corrections and confirmed patterns (stops Claude repeating mistakes)reference/— pointers to external systems
Second terminal pane. Real-time: context %, cost, cache efficiency, degradation risk.
python tools/cc-token-gauge/context_gauge.pyThis is the part most people skip. Don't.
By default, AI agents reach for bash commands to read files, search codebases, and find content. Bash works — but it's a token furnace. Every cat, grep, and find call burns tokens on output formatting, shell overhead, and raw file dumps.
The solution: dedicated MCP tools that return exactly what the AI needs, nothing else.
Replaces find, ls, and directory scanning entirely. Frecency-ranked results (frequent + recent files first). Orders of magnitude faster than bash find.
In your CLAUDE.md, tell your AI explicitly:
File search → fff (mcp__fff__find_files)
File content search → fff grep (mcp__fff__grep)
NEVER use bash find, ls, grep, or rg for file operations
Without this rule, your agent will default to bash. With it, token usage on file ops drops 70%+.
Install (Windows — download prebuilt binary):
# Download fff-mcp-x86_64-pc-windows-msvc.exe from:
# https://github.com/dmtrKovalenko/fff.nvim/releases/latest
# Place at: C:\Users\<you>\.local\bin\fff-mcp.exe
# Current stable: v0.9.6Install (macOS/Linux — run the install script):
curl -fsSL https://raw.githubusercontent.com/dmtrKovalenko/fff.nvim/main/install-mcp.sh | shReplaces reading entire code files. Your AI gets symbol definitions, references, and call graphs — not 500 lines of raw source.
New in v1.80+: Gateway Mode v2 — jMunch now works as a universal proxy for ANY AI application using the OpenAI or Anthropic HTTP APIs, not just MCP servers. Zero code changes required. Benchmarked: 95–98.9% token reduction on wrapped apps.
In your CLAUDE.md:
Code files (.py/.ts/.tsx) → jCodeMunch (mcp__jcodemunch__*)
Call list_repos before reading any code file
Critical: MCP vs CLI token cost. MCP servers load their full tool definitions into every message turn — even when you never call them. A session with 5 heavy MCPs can carry 70k tokens of dead weight per turn. CLIs cost zero tokens when idle, only tokens when called. If a tool has a CLI equivalent, prefer it. Switching MCPs to CLIs can save 40% of your session tokens before you write a single line of code.
Replaces reading entire markdown docs. Your AI queries specific sections, not whole files.
In your CLAUDE.md:
Doc files (.md/.mdx/.rst) → jDocMunch (mcp__jdocmunch__*)
Call search_sections before reading any markdown
Replaces reading raw JSON, HTML, and data files. Your AI queries datasets, describes columns, and samples rows — not raw file dumps.
In your CLAUDE.md:
Data files (.json/.html, >100 lines) → jDataMunch (mcp__jdatamunch__*)
Wraps the entire jMunch family (and any MCP server) as a transparent proxy. Compresses bulky MCP responses before they hit your context window.
Benchmarked savings:
- GitHub MCP: 88.3% token reduction
- Firecrawl MCP: 98.9% token reduction
- Wall-clock performance: 19-43% faster
Install:
pip install jmunch-mcpWire into your MCP config (replace the direct jcodemunch/jdocmunch/jdatamunch commands with the jmunch-mcp proxy pointing to a config TOML). See guides/tool-stack.md for full wiring instructions.
CLI proxy that compresses shell command output before it hits your context. Intercepts git, npm, pytest, tsc, and 100+ other commands and strips noise before your AI sees it.
Claims 60-90% reduction on common dev commands.
Install (Windows):
# Download from https://github.com/rtk-ai/rtk/releases/latest
# Pick rtk-x86_64-pc-windows-msvc.zip
rtk init -g # wires into Claude Code automaticallyInstall (macOS/Linux):
curl -fsSL https://raw.githubusercontent.com/rtk-ai/rtk/refs/heads/master/install.sh | sh
rtk init -gWraps claude -p (headless Claude Code) with jcodemunch retrieval pre-wired, so batch/scripted calls pull relevant code slices on demand instead of stuffing whole files into the prompt. Auth-agnostic — runs against your Claude subscription's Agent SDK credit by default ($0 actual cost within the monthly credit), or --use-api for team/CI use per Anthropic's TOS.
Verbs: ask, index, run, review, changelog, refactor, tests, sweep, doctor.
In your CLAUDE.md / scripts, use it for:
Batch/scripted Claude calls (PR review, changelog, fan-out refactors) → jragmunch
Requires jcodemunch-mcp registered as an MCP server
Install:
pip install jragmunch
jragmunch doctorRepo: jgravelle/jragmunch-cli
On Claude Pro, every token counts. A typical session reading files via bash vs. the full stack:
| Operation | Bash tokens | fff+jMunch tokens | With jmunch-mcp |
|---|---|---|---|
| Find a file in large repo | ~2,000 | ~50 | ~50 |
| Read a code symbol | ~3,000 (whole file) | ~200 (symbol only) | ~25 |
| Search doc for answer | ~5,000 (whole doc) | ~300 (section) | ~35 |
| GitHub MCP call | ~8,000 | ~8,000 | ~940 |
Over a full session: 50-75% savings from fff+jMunch, up to 90% additional savings from jmunch-mcp on MCP calls.
The rule your AI must follow:
Use fff and jMunch for ALL file operations. Bash is only for git commands, package installs, and CLI execution. Never bash-grep. Never bash-cat. Never bash-find.
| Tool | Cost | Purpose |
|---|---|---|
| Claude Pro | $20/month | The AI |
| fff | Free | Token-efficient file search |
| jCodeMunch + jDocMunch + jDataMunch | Free | Token-efficient code/doc/data navigation |
| jmunch-mcp | Free | MCP response compressor (88-99% reduction) |
| jragmunch | Free | Token-efficient RAG CLI for headless Claude |
| RTK | Free | Shell output compressor (60-90% reduction) |
| NotebookLM (research pipeline) | Free | Knowledge extraction |
| This kit | Free | Config + skills + memory system |
Total: $20/month.
Anthropic's own research identifies 5 failure modes in long AI sessions:
- "I don't know" circuit gets overridden — competing signals in long context cause false confidence
- Performative reasoning — responses look correct but the underlying logic is wrong
- Sycophancy — the AI reverse-engineers agreement with your suggestion instead of checking independently
- Internal momentum — can't self-correct mid-sentence even when wrong
- Context decay — early instructions fade, later instructions dominate
Key thresholds:
- Message 30+: degradation starts
- Context 50%: plan to
/compact - Context 80%:
/compactimmediately
The accuracy rules in CLAUDE.md and the degradation risk gauge in cc-token-gauge are direct implementations of this research.
# 1. Clone
git clone https://github.com/albatrossflyon-coder/claude-token-operator-kit
# 2. Copy config
cp config/CLAUDE.md ~/.claude/CLAUDE.md
# 3. Install skills
cp -r skills/* ~/.claude/skills/
# 4. Start token gauge (second terminal)
python tools/cc-token-gauge/context_gauge.py
# 5. Install fff + jMunch (see guides/tool-stack.md)Full setup guide: guides/20-dollar-setup.md
This kit is the output of a research pipeline:
- YouTube videos + articles go into NotebookLM notebooks
- NLM extracts structured knowledge via CLI
- That knowledge gets encoded into skills, CLAUDE.md rules, and memory files
- The AI runs those rules on every session
The token gauge was built because NLM surfaced Anthropic's degradation research. The verify skill exists because that same research showed AI will confidently give wrong answers in long sessions. Every piece connects.
- Karpathy framing — Andrej Karpathy's work on understanding LLMs as deterministic systems with known failure modes
- fff — dmtrKovalenko — the fastest file search toolkit for AI agents. Core of the 50-75% token savings.
- jCodeMunch — jgravelle — semantic code navigation via MCP
- jDocMunch — jgravelle — section-level markdown navigation via MCP
- jDataMunch — jgravelle — structured data navigation via MCP
- jmunch-mcp — jgravelle — MCP response compressor proxy. Wraps any MCP server and cuts response token cost 88-99%.
- jragmunch — jgravelle — token-efficient RAG CLI for headless Claude, jcodemunch retrieval pre-wired.
- RTK — rtk-ai — Rust Token Killer. CLI proxy that compresses shell output 60-90% before it hits your context.
- NotebookLM — Google — research pipeline that surfaced Anthropic's degradation research
- Token monitoring — ai-token-dashboard — live token dashboard for CC, Hermes, Gemini, and more
MIT
