A multi-agent AI research assistant built with Pydantic AI, using the Programmatic Agent Hand-off pattern — a chain of specialized agents where each stage's output feeds the next.
Give it a topic. It searches arXiv for real papers, downloads and reads the PDFs, extracts structured findings from each one, compares them side by side, writes summaries at multiple reading levels, builds a concept mind map, and generates a full Markdown report — all from a single query, with no manual query-writing required.
Full original spec: SUMMARY.md. Phase-by-phase build plan and progress: TASKS.md.
- One input, one output. Type a topic, get a complete research report — plan, papers, comparison table, summary, mind map, and downloadable Markdown.
- Grounded, not hallucinated. Every paper comes from a real arXiv API lookup; every extracted fact (model, dataset, metrics, limitations) is pulled from the actual downloaded PDF text, not guessed from the topic.
- Two frontends. A live dashboard (FastAPI + Tailwind) that shows exactly which agent is running, and a Streamlit app for quick iteration.
- Free to run. Uses the arXiv API (no key required); the only paid dependency is your LLM provider.
Topic
│
▼
Planner Agent → objectives, keywords, arXiv search queries
│
▼
Paper Search Agent → real papers from the arXiv API (free, no key)
│
▼
PDF Download Agent → downloads PDFs to papers/, dedupes by arxiv_id
│
▼
PDF Reader + Knowledge → extracts full text (PyMuPDF), then per-paper:
Extraction Agent objective, methodology, dataset, model, metrics,
results, contributions, limitations, future work
│
▼
Paper Comparison Agent → one row per paper: model, dataset, key result, limitation
│
▼
Summary Agent → executive / beginner / technical summaries + key insights
│
▼
Mind Map Agent → concept graph (nodes + edges)
│
▼
Report Generator → assembled Markdown report, saved to reports/
Two frontends sit in front of the same backend pipeline:
- Dashboard (
GET /) callsPOST /api/research/start, then pollsGET /api/research/status/{job_id}to show a live, per-agent progress tracker before rendering the final result. - Streamlit (
streamlit_app.py) callsPOST /api/research, which runs the whole pipeline synchronously and returns the full result in one response.
Both routes call the same run_research_pipeline() in src/services/orchestrator.py — the dashboard's version passes an on_stage callback so progress can be observed; the Streamlit route doesn't.
| Layer | Choice |
|---|---|
| Agent framework | Pydantic AI (Programmatic Agent Hand-off) |
| Backend | FastAPI |
| Dashboard frontend | Jinja2 templates + Tailwind CSS (CDN) + vanilla JS + Mermaid.js |
| Secondary frontend | Streamlit |
| Paper search | arXiv API (free, no key) |
| PDF parsing | PyMuPDF (fitz) |
| XML parsing | defusedxml (safe parsing of arXiv's Atom feed) |
| LLM | OpenAI, via Pydantic AI (model configurable through MODEL_NAME) |
| Job tracking | In-memory (src/services/job_manager.py) — no persistence yet |
| Database | SQLite (planned — see TASKS.md Phase 9, not wired up yet) |
├── main.py # FastAPI entrypoint — mounts routes + /static
├── streamlit_app.py # Streamlit UI
├── requirements.txt
├── .env # OPENAI_API_KEY, MODEL_NAME, DATABASE_URL, API_BASE_URL
├── SUMMARY.md # original project spec
├── TASKS.md # phase-by-phase build plan (gitignored, local only)
│
├── src/
│ ├── agents/ # one Pydantic AI Agent (or plain function) per pipeline stage
│ │ ├── planner.py # ResearchPlan
│ │ ├── search.py # PaperSearchResult (arXiv)
│ │ ├── extractor.py # ExtractedKnowledge (per paper)
│ │ ├── comparator.py # ComparisonTable
│ │ ├── summarizer.py # ResearchSummary
│ │ ├── mindmap.py # MindMapGraph
│ │ └── report.py # ResearchReport (deterministic, no LLM)
│ │
│ ├── schemas/ # Pydantic models for each agent's structured output
│ ├── services/
│ │ ├── orchestrator.py # chains every agent together, exposes on_stage callback
│ │ ├── job_manager.py # in-memory Job (per-stage status tracking)
│ │ └── job_runner.py # runs the pipeline as a background asyncio task
│ │
│ ├── tools/ # deterministic helpers (no LLM)
│ │ ├── arxiv_client.py # arXiv API client
│ │ ├── pdf_downloader.py # downloads + dedupes PDFs into papers/
│ │ └── pdf_parser.py # PyMuPDF text extraction
│ │
│ ├── api/routes/ # FastAPI route handlers
│ │ ├── plan.py # POST /api/plan
│ │ ├── research.py # POST /api/research, /research/start, GET /research/status/{id}
│ │ └── web.py # GET / (dashboard)
│ │
│ ├── web/ # dashboard frontend
│ │ ├── templates/index.html # single-page dashboard (Jinja2)
│ │ └── static/{app.js,styles.css} # polling + rendering logic, monochrome theme
│ │
│ ├── database/ # reserved for SQLAlchemy models (not implemented yet)
│ └── utils/config.py # loads and validates .env
│
├── papers/ # downloaded PDFs (gitignored)
└── reports/ # generated Markdown reports (gitignored)
Requires Python 3.11+ and an OpenAI API key.
pip install -r requirements.txtCreate a .env file in the project root:
OPENAI_API_KEY="sk-..."
MODEL_NAME="openai:gpt-4.1-mini"
DATABASE_URL="sqlite:///./research.db"
API_BASE_URL="http://localhost:8000"
MODEL_NAME accepts any Pydantic AI model identifier (e.g. openai:gpt-4.1, openai:gpt-4o-mini). API_BASE_URL is only used by the Streamlit app to know where the FastAPI backend is running.
Start the FastAPI backend — this also serves the dashboard:
uvicorn main:app --reload --port 8000Open http://localhost:8000. Enter a research topic (e.g. "GraphRAG vs Traditional RAG") and click Run. A vertical stage tracker shows exactly which agent is running — Planner → Search → PDF Download → PDF Reader/Extraction → Comparison → Summary → Mind Map → Report — and the full result (summary, comparison table, mind map, papers, report download) renders inline once it finishes. A full run typically takes a few minutes, since it's downloading and analyzing several real papers with real LLM calls at each stage.
To run the Streamlit app instead (or alongside, in a second terminal):
streamlit run streamlit_app.py| Method | Path | Description |
|---|---|---|
GET |
/ |
Dashboard (HTML) |
POST |
/api/plan |
Runs only the Planner agent. Body: {"topic": "..."} → ResearchPlan |
POST |
/api/research |
Runs the full pipeline synchronously; blocks until done. Used by Streamlit. |
POST |
/api/research/start |
Starts the full pipeline in the background. Returns {"job_id": "..."} immediately. |
GET |
/api/research/status/{job_id} |
Live per-stage status (pending/running/done/failed) and the final result once done. |
GET |
/health |
Health check |
Interactive OpenAPI docs at http://localhost:8000/docs.
Example:
curl -X POST http://localhost:8000/api/research/start \
-H 'Content-Type: application/json' \
-d '{"topic": "Mixture of experts language models"}'
# => {"job_id": "..."}
curl http://localhost:8000/api/research/status/<job_id>- No persistence. Jobs and results live in memory only — they're lost on server restart, and there's no history of past research sessions. SQLite persistence is planned (see TASKS.md) but not implemented.
- Synchronous run is slow and blocking.
/api/research(used by Streamlit) blocks the request for the full multi-minute pipeline duration; there's no timeout/cancellation. - Markdown only. The Report Generator produces Markdown; PDF/HTML export isn't implemented yet.
- No token budget enforcement. Pydantic AI's
UsageLimits/ctx.usagearen't wired in yet, so there's no hard cap on combined LLM spend per run. - Capped paper count. Only the top
MAX_PAPERS_TO_PROCESS = 5search results are downloaded and analyzed per run (seesrc/services/orchestrator.py), to bound cost and latency. - arXiv only. Semantic Scholar and Crossref (mentioned in SUMMARY.md) are deferred.
See TASKS.md for the full phase-by-phase build plan. Everything through the Report Generator is implemented and has been verified end-to-end against real arXiv papers and a real LLM. Next up: database persistence and non-Markdown report export.