Model SaidWhat the models said.

Method

The panel

On the panel today: ChatGPT (OpenAI) and Claude (Anthropic). Each model is called through its lab’s official API. We do not paste prompts into chat apps and we do not publish screenshots.

The panel grows over time. Gemini, Grok and Perplexity are not on it yet. A model joins from its first dated run; earlier runs keep the models they had.

Each page names the models that answered that run, and each tape names the exact API model, for example gpt-6.1-sol or claude-sonnet-5-5.

The question

40 fixed consumer electronics questions. Every seat on a run gets the same text, word for word: one short instruction, a required answer format, and the question. Prices in questions are US dollars. The current wording is prompt v1. If the wording changes, the version changes, and the page says which version a run used. Each verdict page opens its tape with the exact text the models received.

Runs use a low temperature. We keep each model’s raw answer exactly as returned, even when we cannot read a pick out of it.

Empty seats

A run is published only when at least two models return a readable first pick. A model that errors or returns no readable pick does not count and does not appear in the tally. Its error or raw text stays on the tape. We never write a pick for a model that did not give one.

How to read a page

The score counts first places from the models that answered. Picks that name the same product in different words count as one. A win means a majority named the same product first. A split means no majority. The tape is the evidence: if the tape and the tally ever disagree, believe the tape and email us.

Notes

When a model gets a checkable fact wrong, such as who makes a product, we add a dated note under that pick. The model’s words stay as it gave them, and the tally counts what it said. We do not add notes on prices, opinions or rankings.

Samples

A page marked “Sample — not a live run” holds illustrative answers we wrote to build the site. No model said them. A sample is replaced as soon as that question has a live run. Questions with neither show “Awaiting first run”.

Re-runs

At launch, each live question has one dated run. We re-run questions over time to see whether the answers move. A re-run is a new dated run with its own tape. The question page shows the latest run first, with every earlier run and its tape below it, and the archive lists every run. We do not edit an old run.

What we do not do

Limits

Models change. Prices change. Stock changes. A page is what the models said on that date, not a buying order and not advice.