A tool-using AI research agent that reasons step-by-step (ReAct-style) and autonomously calls external tools — web search, Wikipedia, arXiv, and a calculator — to answer research questions with sourced, structured responses.
Unlike a standard RAG chatbot that just retrieves and summarizes, this agent reasons about what it doesn't know, decides which tool can fill that gap, calls it, observes the result, and iterates — the same reasoning loop that powers modern production AI agents.
Ask it something that needs multiple kinds of information in one go, e.g.:
"What's the latest news about AI, and what is 15% of 340?"
The agent will search the web for current news, use the calculator for the math, and combine both into one coherent answer — while showing exactly which tools it called and why, in an expandable trace.
User Question
│
▼
┌─────────────┐ ┌──────────────────┐
│ Agent Core │◄────►│ LLM (Groq API) │
│ (ReAct loop)│ │ openai/gpt-oss │
└──────┬──────┘ └──────────────────┘
│ selects & calls
▼
┌──────────────────────────────────────────┐
│ Tools: Web Search │ Wikipedia │ arXiv │ Calculator │
└──────────────────────────────────────────┘
│
▼
Sourced, structured answer
The agent loop (app/agent.py) sends the conversation plus tool schemas to the LLM. If the model requests a tool call, the loop executes it via a central registry (app/tools/registry.py) and feeds the result back — repeating until the model has enough information to answer, or hits a step limit as a safety net.
- Backend: FastAPI, Python
- LLM: Groq API (
openai/gpt-oss-120b) - Frontend: Streamlit
- Deployment: Docker
- Testing: pytest (16 tests, run automatically via CI)
- CI: GitHub Actions
- 4 tools, each independently tested: web search (DuckDuckGo), Wikipedia (direct REST API), arXiv paper search, and a safe AST-based calculator (no
eval()— rejects code injection) - Retry logic for transient model/API errors (malformed tool calls, output-parsing failures)
- Context-window management: tool results are truncated before being fed back into the conversation to stay under API rate limits, while the UI still shows full results
- Citation cleanup: strips formatting artifacts some models leave behind from their own training
- Full tool-call trace shown in the UI, not just the final answer — see exactly what the agent searched for and what it found
research-agent/
├── app/
│ ├── main.py # FastAPI entrypoint
│ ├── agent.py # ReAct reasoning loop
│ ├── llm_client.py # Groq API wrapper with retry logic
│ ├── config.py # Environment-based settings
│ └── tools/
│ ├── registry.py # Central tool registry
│ ├── web_search.py
│ ├── wikipedia_tool.py
│ ├── arxiv_tool.py
│ └── calculator.py
├── tests/ # pytest unit tests
├── streamlit_app.py # Chat UI
├── Dockerfile
└── .github/workflows/ # CI (runs tests on every push)
git clone https://github.com/zain-cs/research-agent.git
cd research-agent
python -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
pip install -r requirements.txtGet a free API key from console.groq.com/keys, then:
cp .env.example .envEdit .env and add your key:
GROQ_API_KEY=your_key_here
uvicorn app.main:app --reloadAPI docs available at http://127.0.0.1:8000/docs.
streamlit run streamlit_app.pyOpens at http://localhost:8501.
docker build -t research-agent .
docker run -p 8000:8000 --env-file .env research-agentpytest tests/ -vTests also run automatically on every push via GitHub Actions.
POST /research
{"question": "What is the capital of France, and what is 15% of 340?"}Returns the final answer plus a full trace of every tool call made:
{
"answer": "...",
"trace": [
{"type": "tool_call", "content": "Calling calculate(...)", "tool": "calculate"},
{"type": "tool_result", "content": "51.0", "tool": "calculate"},
{"type": "final_answer", "content": "..."}
],
"steps_used": 2
}- Persistent conversation memory across sessions
- A results-caching layer to reduce redundant tool calls
- Streaming responses instead of waiting for the full agent loop to finish
- Swappable LLM backends (currently Groq-only)
MIT — see LICENSE