Live
AI Debate Arena
Pick a motion and watch three LLM agents argue it live: a Pro and a Con agent trade rebuttals while a Judge scores every round, all streamed token-by-token.
- Role
- Design & engineering
- Timeline
- [n] weeks
- Year
- [2025]
- Stack
- LangGraph, Python, Next.js
(01)The problem
Single-model chat answers hide the reasoning trade-offs behind a confident tone.
Running several agents in real time is hard: turn-taking, shared context, and streaming three voices to one UI without it feeling chaotic.
(02)My role
- Designed the agent graph and prompts for the Pro, Con and Judge roles, including structured scoring rubrics.
- Built a streaming transport that multiplexes several agent token streams over one Server-Sent Events channel.
- Built the debate UI with per-agent lanes, live scoring and shareable transcripts.
(03)Architecture
A LangGraph state machine owns the debate: each node is an agent, edges encode turn order, and a shared state object carries the transcript. Tokens from every node are tagged with the agent id and streamed to the client over SSE.
- 01
Motion
User topic, normalised into a debatable claim
- 02
Pro agent
Opening argument, then rebuttals
- 03
Con agent
Counter-argument with cited reasoning
- 04
Judge agent
Rubric scoring per round, JSON output
- 05
Stream mux
Agent-tagged tokens → SSE → React lanes
Graph, not a prompt chain
LangGraph made turn order, retries and early termination explicit and testable instead of buried in prompt glue.
Structured judge output
The Judge returns schema-validated JSON, so scores render as UI rather than prose — and malformed output is retried automatically.
(04)Tech
AI
- LangGraph
- LangChain
- OpenAI
- Pydantic
Backend
- Python
- FastAPI
- Server-Sent Events
Frontend
- Next.js
- TypeScript
- Motion
(05)Outcome
[What users did with it, what surprised you, what's next.]
- cooperating agents
- 3
- ms to first token
- <400
- debates run
- 1,000+
Next case study
Orbit
An AI study assistant that turns notes into practice