Upload a CSV, ask questions in plain English, and get answers, tables, and charts back. DataChat is a small but complete LLM agent: Google Gemini decides which tool to call, the tools run sandboxed pandas and matplotlib, and the results are fed back until the model produces a grounded, natural-language answer.
Built as an original portfolio project to demonstrate agent design, safe code execution, and pandas/data-analysis engineering — not a wrapper around a one-line framework helper.
Ask a question in plain English — the agent computes the answer from the data and can chart it:
Ask for a visualization directly:
| Piece | What it shows |
|---|---|
Explicit tool-calling loop (agent.py) |
Function calling with Gemini, driven by a hand-written loop — no hidden framework magic, so the control flow is testable and easy to explain. |
AST-allowlist sandbox (sandbox.py) |
Model-generated pandas code is parsed to an AST and rejected unless every node is on an allowlist (no imports, no dunder access, restricted builtins). Honest about its limits. |
Structured chart tool (charts.py) |
Charts come from validated parameters, not free-form plotting code — safer and predictable. |
| Grounding guardrails | The system prompt forbids inventing numbers: every figure must come from a tool result. |
Test suite (tests/) |
Deterministic tests for the sandbox (including sandbox-escape attempts) and the chart builder — the parts that don't need an API key. |
┌──────────────┐ question ┌───────────────────┐
Streamlit ──►│ ask(q) │──────────────►│ Gemini (function │
(app.py) │ loop │◄──────────────│ calling) │
└──────┬───────┘ tool call └───────────────────┘
│ dispatch
┌────────────┴────────────┐
▼ ▼
run_pandas (sandbox) create_chart (matplotlib)
│ │
└────► result text ◄──────┘ (fed back to the model)
# 1. install
python -m venv .venv && source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
# 2. add your (free) Gemini key — https://aistudio.google.com/apikey
cp .env.example .env # then edit .env
# 3. (optional) regenerate the sample dataset
python sample_data/generate.py
# 4. run the app
streamlit run app.pyThen upload any CSV — or click Load sample dataset — and ask things like:
- "Which region has the highest total revenue?"
- "What's the average order value by customer segment?"
- "Plot monthly revenue as a line chart."
- "Show the top 5 products by units sold."
pytest -qThe tests cover the sandbox and chart engine (no API key or network needed), including attempts to break out of the sandbox via imports and dunder access.
datachat-agent/
├── app.py # Streamlit UI (upload, chat, charts, tool-call trace)
├── agent.py # Gemini function-calling loop + tool declarations
├── sandbox.py # AST-allowlist execution of pandas code
├── charts.py # structured chart builder
├── sample_data/ # generator + a demo sales.csv
└── tests/ # pytest suite for sandbox + charts
The sandbox is a best-effort allowlist that raises the bar for running model-generated code (blocks imports, dunder access, arbitrary builtins). It is not a hardened boundary for untrusted, multi-tenant use. For that, run the code in a real isolate (container, gVisor/seccomp, or a subprocess with resource limits). This project is meant for local, single-user analysis of data you trust.
MIT — see LICENSE.

