Skip to content

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

AI Research Agent

A multi-agent AI research assistant built with Pydantic AI, using the Programmatic Agent Hand-off pattern — a chain of specialized agents where each stage's output feeds the next.

Give it a topic. It searches arXiv for real papers, downloads and reads the PDFs, extracts structured findings from each one, compares them side by side, writes summaries at multiple reading levels, builds a concept mind map, and generates a full Markdown report — all from a single query, with no manual query-writing required.

Full original spec: SUMMARY.md. Phase-by-phase build plan and progress: TASKS.md.

Features

  • One input, one output. Type a topic, get a complete research report — plan, papers, comparison table, summary, mind map, and downloadable Markdown.
  • Grounded, not hallucinated. Every paper comes from a real arXiv API lookup; every extracted fact (model, dataset, metrics, limitations) is pulled from the actual downloaded PDF text, not guessed from the topic.
  • Two frontends. A live dashboard (FastAPI + Tailwind) that shows exactly which agent is running, and a Streamlit app for quick iteration.
  • Free to run. Uses the arXiv API (no key required); the only paid dependency is your LLM provider.

Architecture

Topic
  │
  ▼
Planner Agent               → objectives, keywords, arXiv search queries
  │
  ▼
Paper Search Agent           → real papers from the arXiv API (free, no key)
  │
  ▼
PDF Download Agent            → downloads PDFs to papers/, dedupes by arxiv_id
  │
  ▼
PDF Reader + Knowledge         → extracts full text (PyMuPDF), then per-paper:
Extraction Agent                 objective, methodology, dataset, model, metrics,
                                  results, contributions, limitations, future work
  │
  ▼
Paper Comparison Agent         → one row per paper: model, dataset, key result, limitation
  │
  ▼
Summary Agent                  → executive / beginner / technical summaries + key insights
  │
  ▼
Mind Map Agent                 → concept graph (nodes + edges)
  │
  ▼
Report Generator                → assembled Markdown report, saved to reports/

Two frontends sit in front of the same backend pipeline:

  • Dashboard (GET /) calls POST /api/research/start, then polls GET /api/research/status/{job_id} to show a live, per-agent progress tracker before rendering the final result.
  • Streamlit (streamlit_app.py) calls POST /api/research, which runs the whole pipeline synchronously and returns the full result in one response.

Both routes call the same run_research_pipeline() in src/services/orchestrator.py — the dashboard's version passes an on_stage callback so progress can be observed; the Streamlit route doesn't.

Tech Stack

Layer Choice
Agent framework Pydantic AI (Programmatic Agent Hand-off)
Backend FastAPI
Dashboard frontend Jinja2 templates + Tailwind CSS (CDN) + vanilla JS + Mermaid.js
Secondary frontend Streamlit
Paper search arXiv API (free, no key)
PDF parsing PyMuPDF (fitz)
XML parsing defusedxml (safe parsing of arXiv's Atom feed)
LLM OpenAI, via Pydantic AI (model configurable through MODEL_NAME)
Job tracking In-memory (src/services/job_manager.py) — no persistence yet
Database SQLite (planned — see TASKS.md Phase 9, not wired up yet)

Project Structure

├── main.py                    # FastAPI entrypoint — mounts routes + /static
├── streamlit_app.py            # Streamlit UI
├── requirements.txt
├── .env                        # OPENAI_API_KEY, MODEL_NAME, DATABASE_URL, API_BASE_URL
├── SUMMARY.md                  # original project spec
├── TASKS.md                    # phase-by-phase build plan (gitignored, local only)
│
├── src/
│   ├── agents/                   # one Pydantic AI Agent (or plain function) per pipeline stage
│   │   ├── planner.py               # ResearchPlan
│   │   ├── search.py                 # PaperSearchResult (arXiv)
│   │   ├── extractor.py               # ExtractedKnowledge (per paper)
│   │   ├── comparator.py               # ComparisonTable
│   │   ├── summarizer.py                # ResearchSummary
│   │   ├── mindmap.py                    # MindMapGraph
│   │   └── report.py                      # ResearchReport (deterministic, no LLM)
│   │
│   ├── schemas/                  # Pydantic models for each agent's structured output
│   ├── services/
│   │   ├── orchestrator.py          # chains every agent together, exposes on_stage callback
│   │   ├── job_manager.py            # in-memory Job (per-stage status tracking)
│   │   └── job_runner.py              # runs the pipeline as a background asyncio task
│   │
│   ├── tools/                    # deterministic helpers (no LLM)
│   │   ├── arxiv_client.py          # arXiv API client
│   │   ├── pdf_downloader.py         # downloads + dedupes PDFs into papers/
│   │   └── pdf_parser.py              # PyMuPDF text extraction
│   │
│   ├── api/routes/                # FastAPI route handlers
│   │   ├── plan.py                   # POST /api/plan
│   │   ├── research.py                # POST /api/research, /research/start, GET /research/status/{id}
│   │   └── web.py                      # GET / (dashboard)
│   │
│   ├── web/                       # dashboard frontend
│   │   ├── templates/index.html      # single-page dashboard (Jinja2)
│   │   └── static/{app.js,styles.css} # polling + rendering logic, monochrome theme
│   │
│   ├── database/                  # reserved for SQLAlchemy models (not implemented yet)
│   └── utils/config.py            # loads and validates .env
│
├── papers/                     # downloaded PDFs (gitignored)
└── reports/                    # generated Markdown reports (gitignored)

Setup

Requires Python 3.11+ and an OpenAI API key.

pip install -r requirements.txt

Create a .env file in the project root:

OPENAI_API_KEY="sk-..."
MODEL_NAME="openai:gpt-4.1-mini"
DATABASE_URL="sqlite:///./research.db"
API_BASE_URL="http://localhost:8000"

MODEL_NAME accepts any Pydantic AI model identifier (e.g. openai:gpt-4.1, openai:gpt-4o-mini). API_BASE_URL is only used by the Streamlit app to know where the FastAPI backend is running.

Running

Start the FastAPI backend — this also serves the dashboard:

uvicorn main:app --reload --port 8000

Open http://localhost:8000. Enter a research topic (e.g. "GraphRAG vs Traditional RAG") and click Run. A vertical stage tracker shows exactly which agent is running — Planner → Search → PDF Download → PDF Reader/Extraction → Comparison → Summary → Mind Map → Report — and the full result (summary, comparison table, mind map, papers, report download) renders inline once it finishes. A full run typically takes a few minutes, since it's downloading and analyzing several real papers with real LLM calls at each stage.

To run the Streamlit app instead (or alongside, in a second terminal):

streamlit run streamlit_app.py

API Reference

Method Path Description
GET / Dashboard (HTML)
POST /api/plan Runs only the Planner agent. Body: {"topic": "..."} → ResearchPlan
POST /api/research Runs the full pipeline synchronously; blocks until done. Used by Streamlit.
POST /api/research/start Starts the full pipeline in the background. Returns {"job_id": "..."} immediately.
GET /api/research/status/{job_id} Live per-stage status (pending/running/done/failed) and the final result once done.
GET /health Health check

Interactive OpenAPI docs at http://localhost:8000/docs.

Example:

curl -X POST http://localhost:8000/api/research/start \
  -H 'Content-Type: application/json' \
  -d '{"topic": "Mixture of experts language models"}'
# => {"job_id": "..."}

curl http://localhost:8000/api/research/status/<job_id>

Known Limitations

  • No persistence. Jobs and results live in memory only — they're lost on server restart, and there's no history of past research sessions. SQLite persistence is planned (see TASKS.md) but not implemented.
  • Synchronous run is slow and blocking. /api/research (used by Streamlit) blocks the request for the full multi-minute pipeline duration; there's no timeout/cancellation.
  • Markdown only. The Report Generator produces Markdown; PDF/HTML export isn't implemented yet.
  • No token budget enforcement. Pydantic AI's UsageLimits/ctx.usage aren't wired in yet, so there's no hard cap on combined LLM spend per run.
  • Capped paper count. Only the top MAX_PAPERS_TO_PROCESS = 5 search results are downloaded and analyzed per run (see src/services/orchestrator.py), to bound cost and latency.
  • arXiv only. Semantic Scholar and Crossref (mentioned in SUMMARY.md) are deferred.

Status

See TASKS.md for the full phase-by-phase build plan. Everything through the Report Generator is implemented and has been verified end-to-end against real arXiv papers and a real LLM. Next up: database persistence and non-Markdown report export.

About

A multi-agent AI research assistant built with Pydantic AI, using the Programmatic Agent Hand-off pattern — a chain of specialized agents where each stage's output feeds the next.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages