Method
The panel
On the panel today: ChatGPT (OpenAI) and Claude (Anthropic). Each model is called through its lab’s official API. We do not paste prompts into chat apps and we do not publish screenshots.
The panel grows over time. Gemini, Grok and Perplexity are not on it yet. A model joins from its first dated run; earlier runs keep the models they had.
Each page names the models that answered that run, and each tape names the exact API model, for example gpt-6.1-sol or claude-sonnet-5-5.
The question
40 fixed consumer electronics questions. Every seat on a run gets the same text, word for word: one short instruction, a required answer format, and the question. Prices in questions are US dollars. The current wording is prompt v1. If the wording changes, the version changes, and the page says which version a run used. Each verdict page opens its tape with the exact text the models received.
Runs use a low temperature. We keep each model’s raw answer exactly as returned, even when we cannot read a pick out of it.
Empty seats
A run is published only when at least two models return a readable first pick. A model that errors or returns no readable pick does not count and does not appear in the tally. Its error or raw text stays on the tape. We never write a pick for a model that did not give one.
How to read a page
The score counts first places from the models that answered. Picks that name the same product in different words count as one. A win means a majority named the same product first. A split means no majority. The tape is the evidence: if the tape and the tally ever disagree, believe the tape and email us.
Notes
When a model gets a checkable fact wrong, such as who makes a product, we add a dated note under that pick. The model’s words stay as it gave them, and the tally counts what it said. We do not add notes on prices, opinions or rankings.
Samples
A page marked “Sample — not a live run” holds illustrative answers we wrote to build the site. No model said them. A sample is replaced as soon as that question has a live run. Questions with neither show “Awaiting first run”.
Re-runs
At launch, each live question has one dated run. We re-run questions over time to see whether the answers move. A re-run is a new dated run with its own tape. The question page shows the latest run first, with every earlier run and its tape below it, and the archive lists every run. We do not edit an old run.
What we do not do
- Scrape consumer chat apps.
- Rewrite answers or product names. A wrong fact gets a note, not an edit.
- Drop a model because we dislike its pick.
- Test products or write reviews.
- Take payment to include or exclude a product.
Limits
Models change. Prices change. Stock changes. A page is what the models said on that date, not a buying order and not advice.