← All projects

consensus

Ask Claude, OpenAI and Gemini the same question, make them argue over their answers for several rounds, and let one of them judge — on your own subscriptions, no API credits.

  • Python
  • LLM
  • Claude
  • OpenAI
  • Gemini
  • Tkinter

Three models, one argument, one ruling. Consensus asks two or three AI models the same question, then makes them check each other’s work. Each model answers on its own first. In every revision round it reads the others’ answers, verifies their claims, and revises its own, under instructions not to blindly follow and not to stubbornly ignore. At the end a judge reads every final answer and rules on each disputed point.

You get one ranked answer, a table showing which model held which item at what value, and a plain-language story of the whole argument: what each model listed, what it added, dropped or re-ranked each round and why in its own words, and how the judge settled what was left.

Why

It’s for questions where the answer is a shortlist and being wrong matters: which stocks to buy, which cities to move to, which practices to learn. Two models with different training and different search habits make different mistakes. In one run, Claude caught a goodwill impairment and a founder-CEO departure the other model had missed; the other flagged a going-concern warning Claude had overlooked. Forcing them to argue surfaces errors a single chat session never would.

How it works

StepWhat happens
plana quick call decides the answer template: what one item is, how to key it, which attributes and units, so answers can be compared
round 0every model answers independently, with web search
revision roundseach model reads the others’ latest answers, verifies, revises, and says what it changed and why
judgeone model reads all final answers and the agreement summary and rules on every argued item
storya narrative of the run is generated from the data, not from a model

Agreement is measured precisely: items match on normalized keys, numbers must be identical to count as agreed, and the ranking order is compared once per round. Row colors show what everyone held, what a majority held, and what’s argued.

Runs on what you already pay for

Claude and OpenAI run through their official command-line tools signed in to a normal subscription; Gemini uses a free API key. Nothing is hosted and no API credits are needed. Prompts are saved with fill-in wildcards ({{AMOUNT}}, {{RISK|low|medium|high}}), every run is kept with all reports, and you can pause between rounds to type a note that steers the next one.

More projects