Skip to content

About

Research and monitoring CLIs that write inspectable artifact bundles instead of chat transcripts.

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Nano Research

CI Python 3.11+ License: MIT

Research and monitoring CLIs that write inspectable artifact bundles instead of chat transcripts.

Most LLM research tools hand you a conversation. When the tab closes, the reasoning, sources, and intermediate notes go with it. nano-research and nano-monitor take the opposite approach.

Every run writes a self-contained folder of typed JSON, raw notes, provenance tables, and canonical Markdown that you can read, diff, and replay without re-running the model. Schemas are validated with pydantic, JSON and Markdown are canonicalized, and behavior is checked against a fixture corpus.

Built for analysts, researchers, and engineers who need to re-ask the same question over weeks and see exactly what changed.

The design emphasis on provenance, auditability, and typed deltas is motivated by research contexts where the same question must be re-asked over time against evolving evidence — tracking model capabilities as they ship, monitoring regulatory or policy developments, or observing how a body of sources changes across runs. The tools are general-purpose, but these are the use cases the evaluation discipline is tuned for.

What's in the box

Tool Status One-line description Rendered example
nano-research stable (schemas frozen) One-shot structured research that writes a reusable evidence bundle instead of a transcript. Orderbook resilience summary
nano-monitor beta Recurring stateful monitoring that writes typed deltas between runs. AI policy monitor delta

What a research bundle looks like

.tmp/orderbook-resilience/
├── run.json          # canonical, sorted-keys JSON of the full run
├── sources.json      # deduplicated provenance with normalized URLs
├── summary.md        # final summary (ATX headings, LF endings)
└── raw_notes/
    ├── web_lane_1.json
    ├── web_lane_2.json
    └── web_followup_1.json

Every downstream workflow — reports, diffs, audits, re-ingestion — just reads files.

Installation

Create a virtual environment, activate it in your shell, and install the package:

python -m venv .venv
python -m pip install --upgrade pip
python -m pip install -e .

If you want the test tooling as well:

python -m pip install -e .[dev]

Live runs also require provider clients and API keys:

python -m pip install openai google-genai

Set these environment variables for live mode:

  • OPENAI_API_KEY
  • GEMINI_API_KEY
  • TAVILY_API_KEY

Quickstart — live run

Set your provider keys, then run a research brief:

export OPENAI_API_KEY=...
export GEMINI_API_KEY=...
export TAVILY_API_KEY=...

nano-research "How do crypto exchanges recover from orderbook snapshot gaps?"

That writes a fresh bundle under research/<slug>/ (or docs/research/<slug>/ if a docs/ directory already exists in the repo root).

Quickstart — fixture replay (no API keys needed)

To see the artifact shape without spending tokens, replay the bundled fixture:

nano-research --title "Orderbook resilience" --prompt-file tests/fixtures/nano_research/brief.md --fixture-dir tests/fixtures/nano_research --output-dir .tmp/orderbook-resilience

That writes run.json, sources.json, summary.md, and the supporting raw_notes/ bundle under .tmp/orderbook-resilience/.

Replay the monitor fixture the same way:

nano-monitor --repo-root .tmp/policy-demo --topic-file tests/fixtures/nano_monitor/policy_topic.json --fixture-run-dir tests/fixtures/nano_monitor/policy_baseline
nano-monitor --repo-root .tmp/policy-demo --topic-file tests/fixtures/nano_monitor/policy_topic.json --fixture-run-dir tests/fixtures/nano_monitor/policy_material_change

That writes watch/ai-policy-monitor/topic.json, state.json, and per-run summary.md, delta.json, delta.md, and next_agent.md artifacts under .tmp/policy-demo/.

Rendered examples

How to verify this works

Run the test suite:

pytest -q

Run the replay gate script:

python scripts/monitor_replay_eval.py

Regenerate the committed examples:

python scripts/eval_examples.py

After python scripts/eval_examples.py, the examples/ tree should remain byte-identical to what is committed.

CI additionally asserts byte-exact example regeneration via git diff --exit-code -- examples/ — any change to the Python pipeline that would affect committed example outputs blocks merge.

nano-monitor ship gates

  • no_change_precision >= 0.90
  • provenance_completeness == 1.00
  • duplicate_suppression_rate >= 0.95
  • wmdf1 must not regress vs tests/fixtures/nano_monitor/replay_baseline.json

Known limitations

  • No live API eval is included in CI. This avoids cost and nondeterminism, but prompt-quality drift is not automatically scored here.
  • CLI orchestration is checked through fixture replay and bounded runtime tests, not a fully mocked live-provider end-to-end harness.
  • Coverage is intentionally concentrated on research provenance and deterministic monitor deltas; behavior outside the fixture corpus has a narrower confidence band.
  • URL normalization is intentionally conservative and may need refinement if future sources rely on path case sensitivity.
  • The replay corpus is still small. It is useful as a release gate, but not yet broad enough to stand in for a true benchmark suite.
  • CLI errors currently surface as raw Python tracebacks. A friendlier error-reporting layer is tracked as a post-0.1.0 improvement.

Read next

About

Research and monitoring CLIs that write inspectable artifact bundles instead of chat transcripts.

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages