consensus
Ask Claude, OpenAI and Gemini the same question, make them argue over their answers for several rounds, and let one of them judge — on your own subscriptions, no API credits.
- Python
- LLM
- Claude
- OpenAI
- Gemini
- Tkinter
Three models, one argument, one ruling. Consensus asks two or three AI models the same question, then makes them check each other’s work. Each model answers on its own first. In every revision round it reads the others’ answers, verifies their claims, and revises its own, under instructions not to blindly follow and not to stubbornly ignore. At the end a judge reads every final answer and rules on each disputed point.
You get one ranked answer, a table showing which model held which item at what value, and a plain-language story of the whole argument: what each model listed, what it added, dropped or re-ranked each round and why in its own words, and how the judge settled what was left.
Why
It’s for questions where the answer is a shortlist and being wrong matters: which stocks to buy, which cities to move to, which practices to learn. Two models with different training and different search habits make different mistakes. In one run, Claude caught a goodwill impairment and a founder-CEO departure the other model had missed; the other flagged a going-concern warning Claude had overlooked. Forcing them to argue surfaces errors a single chat session never would.
How it works
| Step | What happens |
|---|---|
| plan | a quick call decides the answer template: what one item is, how to key it, which attributes and units, so answers can be compared |
| round 0 | every model answers independently, with web search |
| revision rounds | each model reads the others’ latest answers, verifies, revises, and says what it changed and why |
| judge | one model reads all final answers and the agreement summary and rules on every argued item |
| story | a narrative of the run is generated from the data, not from a model |
Agreement is measured precisely: items match on normalized keys, numbers must be identical to count as agreed, and the ranking order is compared once per round. Row colors show what everyone held, what a majority held, and what’s argued.
Runs on what you already pay for
Claude and OpenAI run through their official command-line tools signed in to a
normal subscription; Gemini uses a free API key. Nothing is hosted and no API
credits are needed. Prompts are saved with fill-in wildcards ({{AMOUNT}},
{{RISK|low|medium|high}}), every run is kept with all reports, and you can
pause between rounds to type a note that steers the next one.