You ask once. Three AI models argue it out, a fourth one judges, and you get an answer with the sources it's standing on.
Screenshots · Why · How it works · Quick start · Architecture · API · Roadmap
Built during the IBM Bob Dev Day Hackathon with the IBM Bob IDE.
Starting a new debate — pick a topic, choose three debater models and a judge, then hit Start research.
The pipeline running live — research, then debate rounds, then judging and synthesis. Each stage shows progress and how many sources or messages it has produced.
I kept hitting the same wall whenever I needed a real answer to a tricky question. Asking one chatbot is fast, but you never quite trust it — it sounds confident either way. Doing the research by hand takes the whole afternoon: a dozen tabs, copy-pasted quotes, citation chases, asking another model to double-check. You finish tired and still not sure.
Disputatio is my fix for that loop. You type in a topic, pick three AI models to argue it, and pick a fourth one as the judge. The app runs the rest on its own:
- It searches the web through Exa and pulls back real sources with citations.
- The three models debate — initial positions, cross-examination, rebuttals, closing arguments.
- The judge reads everything and scores it on accuracy, logic, evidence, and rebuttal handling.
- You get the final answer in the shape you want — long write-up, short summary, or plain-language version.
It's meant for solo devs, students, analysts, writers — anyone who needs a real answer fast and doesn't want to take a single model on faith. No account; runs locally.
What makes it different from typical "AI research" tools. Most either give you one model's opinion or quietly average everything together so it sounds tidy. Disputatio does the opposite: when models disagree, you see exactly where and why. When something is uncertain, it stays uncertain. Every claim has a source attached.
┌──────────┐ ┌──────────────────────────────────────────┐ ┌──────────┐ ┌──────────────┐
│ Research │ ──▶ │ Initial → Cross-exam → Rebuttal → Final │ ──▶ │ Judge │ ──▶ │ Synthesis │
│ (Exa) │ │ 3 debater models in parallel │ │ (1 model)│ │ (3 formats) │
└──────────┘ └──────────────────────────────────────────┘ └──────────┘ └──────────────┘
| Stage | What happens | Provider |
|---|---|---|
| 🔍 Research | Multi-query Exa search; pulls structured sources with citations | Exa.ai |
| 🗣 Initial | Each debater takes a position grounded in the research | OpenRouter |
| 🎯 Cross-exam | Debaters challenge each other's claims | OpenRouter |
| 🛡 Rebuttal | Each defends and refines its position | OpenRouter |
| 🎤 Final | Closing arguments | OpenRouter |
| ⚖️ Judging | Scored on accuracy (40%), logic (30%), evidence (20%), rebuttals (10%) | OpenRouter |
| 📝 Synthesis | Final answer in Explanatory, Scientific, or Simplified form | OpenRouter |
Each stage talks to the next through validated JSON, so the pipeline stays intact even when a model gets creative with formatting.
- Node.js 18+
- An OpenRouter API key — openrouter.ai
- An Exa.ai API key — exa.ai
git clone https://github.com/lev1nson/Disputatio.git
cd Disputatio
npm installCreate .env.local:
OPENROUTER_API_KEY=sk-or-...
EXA_API_KEY=...
NEXT_PUBLIC_APP_URL=http://localhost:3000Run it:
npm run dev # http://localhost:3000
npm run typecheck # tsc --noEmit
npm run lint # eslint
npm test # vitest- Pick three debater models and one judge. GPT-4 Turbo, Claude 3 Opus, Gemini Pro, Llama 3 70B all work — anything OpenRouter supports.
- Type your topic (10+ characters). E.g. "What is the impact of AI on employment?"
- Hit Start research. Watch the pipeline strip light up stage by stage.
- Read the verdict in your preferred format.
| Layer | Tech |
|---|---|
| Framework | Next.js 16 (App Router), React 19, TypeScript 5 |
| UI | Tailwind CSS 4, shadcn/ui, Radix UI, lucide-react |
| State | Zustand |
| Validation | Zod 4 |
| HTTP | Axios |
| Tests | Vitest 4, Testing Library, happy-dom |
| LLM gateway | OpenRouter (multi-provider model orchestration) |
| Web research | Exa.ai (neural search + content extraction) |
ai-debate-platform/
├── app/
│ ├── api/
│ │ ├── research/route.ts # Exa-powered research
│ │ ├── debate/route.ts # Debate round orchestration
│ │ ├── judge/route.ts # Judge evaluation
│ │ └── synthesize/route.ts # Final answer synthesis
│ ├── layout.tsx # Root layout
│ ├── page.tsx # Main pipeline UI
│ └── globals.css
├── components/
│ ├── disputatio/ # Run shell (sidebar, header, footer, stages)
│ │ ├── PipelineRail.tsx
│ │ ├── RunningContent.tsx
│ │ ├── SynthesisContent.tsx
│ │ ├── VerdictBlock.tsx
│ │ ├── TranscriptMessage.tsx
│ │ ├── SourceList.tsx
│ │ └── …
│ ├── pipeline/ # Pipeline strip and stage panels
│ │ ├── PipelineStrip.tsx
│ │ ├── StageCell.tsx
│ │ ├── HeaderBar.tsx
│ │ ├── stageMeta.ts
│ │ └── panels/
│ ├── ui/ # shadcn primitives (button, card, tabs, …)
│ ├── DebateMessage.tsx
│ ├── ResearchDisplay.tsx
│ ├── JudgeDecisionDisplay.tsx
│ ├── FinalAnswerDisplay.tsx
│ └── ModelSelector.tsx
├── lib/
│ ├── api/{openrouter,exa}.ts # Provider clients
│ ├── store/debateStore.ts # Zustand state
│ ├── types/models.ts # Shared types
│ └── utils/ # Prompts, round summary, helpers
├── docs/ # Hackathon docs + screenshots
├── .planning/ # GSD planning artefacts (PROJECT, ROADMAP, STATE, phases)
└── package.json
All routes live under app/api/* and accept JSON.
Request: { topic: string }
Response: ResearchResultRuns an Exa-powered multi-query search and returns structured sources with citations.
Request: {
topic: string
research: ResearchResult
models: { debater1: string; debater2: string; debater3: string }
phase: 'initial' | 'cross-exam' | 'rebuttal' | 'final'
debateHistory?: DebateMessage[]
}
Response: { messages: DebateMessage[] }Orchestrates one round of the debate for the three debater models.
Request: {
topic: string
research: ResearchResult
judgeModel: string
debateHistory: DebateMessage[]
agents: Agent[]
}
Response: JudgeDecisionScores the debate on accuracy / logic / evidence / rebuttals.
Request: {
topic: string
research: ResearchResult
debateHistory: DebateMessage[]
judgeDecision: JudgeDecision
synthesisModel: string
}
Response: FinalAnswer // explanatory | scientific | simplifiedGenerates the final answer in three voice variants.
100+ models via OpenRouter, including:
- OpenAI — GPT-4 Turbo, GPT-3.5 Turbo
- Anthropic — Claude 3 Opus / Sonnet / Haiku
- Google — Gemini Pro
- Meta — Llama 3 70B
- Mistral — Mixtral 8x7B
The whole thing was built inside the IBM Bob IDE. Bob was basically a second pair of hands the entire way:
- Planning the program. Before I wrote a single line, Bob read the prototype, helped lay out the architecture, broke the work into milestones, and worked through where the pipeline should split into stages.
- Writing the code. Most of the route handlers, the Exa and OpenRouter clients, the state store, and the UI panels were written together with Bob.
- Adding features. Deepening the research step, adding a jury alongside the judge, letting users pick a report format, cleaning up citations — Bob handled the multi-file edits and walked me through the diffs.
- Tests. Stage metadata, citation parsing, prompt templates — Bob wrote tests I could actually trust, and I tightened them where needed.
I didn't end up using watsonx.ai or watsonx Orchestrate in this version — models go through OpenRouter so you can swap whatever you want. Plugging watsonx-hosted models in as extra debaters or judges is the obvious next step.
The project plans live in .planning/. Current milestone is v1.0.
| # | Phase | Status |
|---|---|---|
| 1 | Backend Logic Baseline | ✅ Complete |
| 1.5 | Pipeline Structure UI | ✅ Complete |
| 1.6 | Readable Debates | ◆ Active |
| 2 | Deep Exa Research Pipeline | ⏳ Pending |
| 3 | Model Assessment + Judge / Jury | ⏳ Pending |
| 4 | Report Generation | ⏳ Pending |
| 5 | Dashboard UI System | ⏳ Pending |
| 6 | Landing Page Output | ⏳ Pending |
See .planning/ROADMAP.md for the requirement-level breakdown.
MIT — use it for anything.
Made with ☕ and four arguing AIs · Report an issue


