Reproducibility
How much does a repeat measurement agree with the last one?
AI answers are not identical each time. The question is how different they are. This page shows the observed results from the fixed questions we send every day — favourable or not. Where the numbers are not there yet, we say so.
What is measured
We send the same question wording to ChatGPT (with web search) on different days and take the brands its answer names. Each pair of consecutive measurements is compared, and we count how far the top 3 names agree. Only our own fixed questions are used — questions typed by visitors are never included here.
When we publish rates
Agreement rates (%) are published once we have at least 5 questions and 30 consecutive pairs. Below that, one extra pair moves the figure by a whole digit, so it cannot support a claim about accuracy. The moment the threshold is met, the numbers appear here automatically — nobody decides whether to show them.
Current sample and results
- Questions with a repeat measurement
- 3
- Measurements in total
- 46
- Consecutive pairs (repeat measurements)
- 41
Period covered: 2026-07-18 – 2026-09-05 (JST)
Agreement rates are not published yet
The sample above has not reached the publishing threshold (5 questions, 30 consecutive pairs). We will not publish a rate below it. Instead, the raw observations — what we measured, how many times, and how many pairs agreed — are shown below.
Observations by question
“Exact” counts pairs where the top 3 names were identical to the previous measurement; “partial” counts pairs where at least one name carried over. These are counts, not rates.
| Question | Measurements | Pairs | Exact | Partial |
|---|---|---|---|---|
| 「渋谷でおすすめの社労士事務所を教えて」2026-07-18 – 2026-09-05 | 21 | 20 | 1 / 20 | 14 / 20 |
| 「渋谷でおすすめの税理士事務所を教えて」2026-08-10 – 2026-09-05 | 13 | 12 | 2 / 12 | 8 / 12 |
| 「渋谷でおすすめの司法書士・弁護士事務所を教えて」2026-08-10 – 2026-09-05 | 10 | 9 | 0 / 9 | 3 / 9 |
| 「小ロット対応の金属加工メーカーを教えて」2026-07-19 – 2026-07-19 | 1 | 0 | — | — |
| 「表参道で評判の良い美容室を教えて」2026-08-26 – 2026-08-26 | 1 | 0 | — | — |
What we have not measured
- Miss rate for brand detection (a name is in the answer but we did not pick it up)
- False-positive rate (something that is not a brand was picked up as one)
- Alias resolution rate (treating “XYZ Inc.” and “XYZ” as the same company)
Each of these needs a hand-labelled ground-truth set, which we have not built. So we publish no estimate for them. When we can measure them, they will appear on this page.
Check it with your own question
The fastest way to judge the variation is to run your own question and see it. No account, no credit card.