Researchers scraped 2,781 pages from gov.uk and generated 22,066 question-and-answer pairs from them, nine per page, written to sound like real people asking about benefits, tax or immigration. They then ran 11 models, including Claude, Gemini, ChatGPT, Llama, Qwen and Kimi, against 7,355 of those pairs, scoring each answer by splitting it into individual factual claims and checking every claim against the page.
Most answers were good. The misses are the finding: scores swing widely from question to question, and a small number of bad answers drag every model’s average below its typical performance. Claude 4.5 Haiku scored highest without worked examples and also wrote the most. Open-weight models such as Qwen kept up with the commercial ones. Every model volunteered more information than the page held, and giving them worked examples did not fix that, sometimes lowering accuracy instead.
Almost no model ever said “I don’t know”, and Qwen3-32B never refused once. The test could hardly have shown anything else. Every question was generated from a page that answers it, and the authors’ own refusal detector errs about once in 150 answers, more than nearly every refusal rate it recorded. What a model does when no page covers the person asking is outside what this test can see.
The authors name the other limit themselves: the reference answers come from gov.uk as it stood, and they start going stale as soon as the rules change.