Methodology

How this benchmark was built

Everything needed to check our figures, disagree with them, or run the test again yourself. Exact model versions, the scoring method, what we excluded, and where the study is weak.

What was tested

951 technical questions about UK construction were put to 18 AI models, producing 13,747 answers. The questions span 20 categories, from Building Regulations and British Standards through to waterproofing, demolition and contracts.

Each question has one verified answer tied to a specific place in a published document: a clause, a table, a diagram or a numbered section. Examples of the sources used include Approved Documents A to S, the relevant British and European Standards with their UK National Annexes, NHBC Standards, BREEAM UK New Construction, the GPDO, CDM 2015 and JCT contract forms.

How the questions were made. The question set was generated by an AI model working from those published documents, and each answer carries the clause-level citation it was drawn from. The citations are real and can be checked against the source. No chartered professional has independently verified the whole question set, and we do not claim otherwise.

We found errors in our own answer key, and say so. Both scoring models were asked to flag any question whose stated answer looked wrong. That produced a shortlist, which a reviewer then checked by hand against the published documents. Of the questions on that shortlist, most turned out to have a real problem: answers that were factually wrong, answers describing requirements that had since been amended or repealed, and citations pointing at table numbers from superseded editions. 27 answers were corrected and re-scored, and 10 questions were removed from scoring as unfair to grade. The original question set is retained unaltered alongside the corrected one.

Then we audited the rest of the key. That first pass only covered questions a judge happened to challenge while marking, which is a sample selected for trouble and says nothing about the remainder. So every other question was put to the same two models again, this time shown only the question, our answer and our citation, with no AI response involved. Where both independently judged the answer wrong or the question ungradeable, it was excluded: 40 more questions came out that way, leaving 951 scored. A further 258 were flagged by only one auditor; those were left in and sent for human review rather than acted on.

The bar is set at both auditors agreeing for a reason. They reach the same verdict on fewer than half the questions, and on some they contradict each other about what a table in the same document says. A single adverse opinion from either is not evidence, so nothing was rewritten on their say-so, only removed when both agreed.

The direction of that error matters. In several cases the answer key was out of date and the models were not: on the Class MA floorspace limit, for example, every model that correctly said the cap had been removed in March 2024 was being marked wrong against a key that still cited the old rule. Correcting the key raised scores rather than lowering them.

Exact models tested

Marketing names are ambiguous and change over time, so every model is listed with the exact identifier sent to the provider's API. If a figure on the results page refers to a model, this is the version it refers to.

Name used in chartsAPI model identifierVendorTier
Claude Opus 5claude-opus-5AnthropicPaid
Claude Fable 5claude-fable-5AnthropicPaid
GPT-5.6 Solopenai/gpt-5.6-solOpenAIPaid
Claude Opus 4.8claude-opus-4-8AnthropicPaid
Kimi 3moonshotai/kimi-k3Moonshot AIPaid
Gemini 3.6 Flashgemini-3.6-flashGoogleFree
Gemini 3.1 Progoogle/gemini-3.1-pro-previewGooglePaid
GPT-5.6 Lunaopenai/gpt-5.6-lunaOpenAIFree
Grok 4.5grok-4.5xAIPaid
Perplexity Sonar Prosonar-proPerplexityPaid
GPT-5.6 Terraopenai/gpt-5.6-terraOpenAIPaid
Perplexity SonarsonarPerplexityFree
Claude Sonnet 5claude-sonnet-5AnthropicPaid
DeepSeek V4 Prodeepseek/deepseek-v4-proDeepSeekPaid
DeepSeek V4 Flashdeepseek/deepseek-v4-flashDeepSeekFree
Mistral Largemistralai/mistral-large-2512MistralPaid
Mistral Mediummistral-medium-2604MistralFree
Claude Haiku 4.5claude-haiku-4-5AnthropicFree
Paid versus free. A model is counted as paid (12 of 18) when a construction professional would need a paid subscription or paid API access to reach it, and free (6) when a capable equivalent is reachable on a free consumer tier. This is a judgement about access, not about price paid during the test, which was billed at API rates throughout.
How each model was asked. Identical instruction for every model: answer a UK construction technical question directly and specifically in two to three sentences, without hedging or refusing. No documents were supplied, no web search was requested, and no follow-up questions were asked. Reasoning models were left at their default settings and given enough output budget to think and still answer.

How answers were scored

Every answer was graded twice, by two AI models from different vendors, against the verified answer and its citation. Each judge sees only the question, the verified answer and the response. Neither is told which model produced the response, so a judge cannot favour its own family.

2

Correct

Same substance as the verified answer, with every key figure, standard reference and requirement right. Different wording or extra accurate detail still scores 2.

1

Partial

On the right topic and partly right, but misses or misstates part of the answer, or gives a value that is close but wrong.

0

Wrong

Contradicts the verified answer, gives a materially incorrect figure or standard, goes off topic, or refuses to answer.

A score counts only when both judges give the same grade. Where they disagreed, the answer was excluded from that model's score rather than resolved by picking a preferred judge. The two judges agreed on 80.3% of answers, and 3,369 were excluded on disagreement. That agreement rate is the best available measure of how reliable these scores are, and it is published so it can be argued with.

Judges were told to grade against substance, not style: an answer that hedges but lands on the correct figure scores 2, and a fluent, confident answer with the wrong figure scores 0. Length was not rewarded.

How much does the ranking depend on our answer key being right?

This is the question worth asking of any benchmark, and the corrections above let us answer it with evidence rather than assurance. Correcting 27 answers was a real change to the key, made on the merits of each question and without regard to which model it would help. If the ranking were fragile, that change would have scrambled it.

0.9587

Rank correlation

Spearman, before versus after correction

12/18

Models unmoved

Did not change position at all

+0.01pp

Mean score change

Individual models moved between −1.1 and +1.6

The top three and the bottom five did not move. What movement there was happened in the middle of the table, where six models sit within a single percentage point of each other. That is why the results are presented as 10 bands rather than 18 ranked places: a gap smaller than 1.5 percentage points is inside the margin our own corrections moved things, so it is not a finding.

What this does and does not show. It shows the ordering is robust to the kind of error we found and fixed, because such errors hit every model at once rather than favouring one. It does not show the absolute percentages are precise. Treat the scores as approximate and the bands as the result.

What was excluded, and what counted against a model

  • Judge disagreements were excluded. Neither judge was treated as the tie-breaker. Excluded answers do not count for or against a model, and the count is published per model in the raw data.
  • Failures to answer counted as wrong. Where a model returned nothing usable, that scored 0 rather than being quietly dropped. A handful of reasoning models occasionally spent their whole output budget thinking and returned an empty answer; those are recorded and counted.
  • Every model tested is published. No model was run and then left out of the results.

Limitations

Where this study is weak. We would rather state these ourselves than have them found.

  1. 1

    The questions were written by AI, not by a chartered professional.

    Each question and its answer were generated against a named clause, table or section of a published UK standard or Approved Document. The citations are real and checkable, but no chartered professional has signed off the question set. Both judges were asked to flag any question whose stated answer looked wrong, and those flags are published below.

  2. 2

    The scorer is itself an AI.

    Two AI models grade every answer. We use one from Anthropic and one from OpenAI so that neither vendor's models are marked by their own family alone, and we only count a score where both agree. That reduces vendor bias but does not eliminate the possibility that both judges are wrong in the same way.

  3. 3

    This measures recall of published standards, not judgement.

    A high score means a model reliably reproduces what a document says. It does not mean the model can apply that requirement to a real building, weigh competing standards, or spot when a question is the wrong question. Design and safety decisions need a competent person.

  4. 4

    Every model was asked the same way, which favours none of them.

    All models received an identical short prompt with no retrieval, no documents attached and no follow-up. Models that would normally be used with a document uploaded, or as part of a tuned workflow, will score differently in that setting.

  5. 5

    Results are a snapshot of a single day.

    Every answer was collected on 28 July 2026, with scoring and review completed the following day. Models are updated continuously, several tested here are explicitly preview or beta releases, and a provider can change a model behind the same name. These figures describe the versions listed above as they behaved on that date, and nothing more.

  6. 6

    Some models were reached through a gateway, not the vendor directly.

    Where a vendor's own account was rate-limited or out of quota, the model was called through OpenRouter instead. The underlying model version is the same and is listed above, but the routing differed and is recorded in the raw data.

Checking our work

If you want to interrogate a specific figure, a category result or a particular model's answers, get in touch and we will share the underlying response and scoring data for that slice, including both judges' verdicts and the citation each question was built from.

Built by Fabrick

Want a platform like this for your business?

This platform was built by Fabrick. We create bespoke digital tools, data platforms and content hubs for construction and built environment companies -- designed to demonstrate your expertise and generate qualified leads.

Carbon CalculatorsData DashboardsContent HubsLead-Gen ToolsCPD PlatformsProduct Selectors

Fabrick - Award-winning marketing specialists for the built environment