Tested 28 July 2026

How Accurate is AI for UK Construction?

951 questions. 18 AI models. 20 categories. Graded blind, with the workings published.

951

Questions

18

AI Models

20

Categories

86.5%

Best Score

Read the full methodology and limitations

How we tested

The short version. The full methodology lists every model version, the scoring rubric, and where this study is weak.

951 technical questions

Building regulations, British Standards, health and safety, fire safety, structural design and 15 other categories. Every answer is tied to a named clause, table or section in a published UK document.

18 models, exact versions published

12 paid and 6 free models from Anthropic, OpenAI, Google, xAI, Mistral, DeepSeek, Moonshot and Perplexity. Each is listed with the precise API identifier tested, not just a marketing name.

Graded blind, by two AIs

Each answer is scored Correct, Partial or Wrong by two models from different vendors, neither told which AI wrote the answer. A score counts only where both agree. They agreed 80.3% of the time.

Identical prompt for every model

No documents supplied, no web search requested, no follow-up questions. This measures what each model knows unaided, which is how most people actually ask it.

Overall Rankings

All 18 models, scored and grouped

Across 951 questions on UK construction standards and regulations. Models close enough to be tied are grouped rather than separated by place. Hover any bar for the correct, partial and wrong split.

Paid model
Free model

Read these as 10 groups, not 18 places. Models within 1.5 percentage points of each other are tied, because gaps that small sit inside the margin of the scoring method. Hover any bar for the exact model version tested and the models it ties with.

Key Findings

What the data reveals about AI in UK construction

86.5%

Highest overall score

Claude Opus 5 led the field. No model was fully right on more than 87% of the marks available, which puts a hard ceiling on how far any of them can be trusted unaided.

10.3pp

Paid versus free gap

Paid models averaged 80.5% against 70.2% for free ones. If you use AI for construction work, the subscription is doing something.

7.7%

Lowest error rate

Even the best model got 59 of 951 questions outright wrong. The weakest, Claude Haiku 4.5, got 277 wrong (36.6%).

13.5%

Not fully right, best model

Counting partial credit, the leading model still failed to give a complete, correct answer on 13.5% of the marks available. That is the number to hold in your head before relying on any of this.

80.3%

Inter-judge agreement

Two AI judges from different vendors agreed on 80.3% of the 13,747 answers. Disagreements were excluded rather than settled by picking a favourite judge.

10

Groups, not places

Models within 1.5 percentage points of each other are too close to separate, so the 18 models resolve into 10 groups. Three lead the field; six more are bunched in the middle.

Paid vs Free

Is paying for AI worth it in construction?

“Paid” means a professional would need a paid subscription or paid API access to reach that model. The full definition is in the methodology.

Paid models average

80.5%

12 models tested

Free models average

70.2%

6 models tested

Category Heatmap

All 20 categories across all 18 models

Green is good. Red is risky.

CategoryClaude Opus 5Claude Fable 5GPT-5.6 SolClaude Opus 4.8Kimi 3Gemini 3.6 FlashGemini 3.1 ProGPT-5.6 LunaGrok 4.5Perplexity Sonar ProGPT-5.6 TerraPerplexity SonarClaude Sonnet 5DeepSeek V4 ProDeepSeek V4 FlashMistral LargeMistral MediumClaude Haiku 4.5
Accessibility & Inclusive Design(39 Qs)78%79%74%71%70%72%68%74%73%62%77%59%69%62%60%49%47%40%
British Standards - Concrete & Steel(48 Qs)91%90%87%87%90%92%90%84%88%81%81%80%91%87%73%71%57%57%
British Standards - Other Materials(45 Qs)93%92%91%92%95%93%87%94%91%90%85%92%82%85%80%77%66%55%
Building Regulations - Fire Safety(49 Qs)93%90%88%91%89%87%86%87%91%92%86%83%75%83%74%74%58%46%
Building Regulations - Other Parts(44 Qs)83%82%77%80%76%86%74%70%76%74%76%76%71%66%56%68%58%44%
Building Regulations - Thermal(47 Qs)84%80%81%79%80%78%87%78%80%83%81%90%66%64%60%42%63%39%
Construction Technology(35 Qs)87%86%79%83%83%72%68%76%70%79%74%69%74%67%61%54%61%52%
Contracts & Procurement(44 Qs)94%92%86%90%89%86%82%80%84%80%81%78%89%74%59%58%59%39%
Demolition & Refurbishment(28 Qs)94%90%96%83%86%76%71%90%75%86%90%85%82%75%65%73%66%57%
Environmental & Contamination(34 Qs)94%95%92%95%95%80%80%86%88%89%88%92%94%81%79%67%59%48%
Fire Safety - Post-Grenfell(41 Qs)88%90%93%92%90%85%88%85%85%96%80%89%84%79%77%71%68%47%
Health & Safety / CDM(46 Qs)99%100%100%100%99%98%99%99%98%99%99%100%96%94%89%90%93%74%
MEP & Building Services(41 Qs)74%76%76%70%73%69%74%76%67%71%75%70%74%73%65%53%49%34%
Materials & Products(40 Qs)86%78%82%77%71%86%83%79%76%72%78%75%74%68%72%65%59%44%
NHBC Standards(39 Qs)68%75%79%74%66%71%68%72%67%71%74%68%57%61%69%67%51%61%
Planning & Permitted Development(51 Qs)94%94%95%87%94%91%94%92%88%98%86%96%82%83%76%70%66%68%
Roofing & Cladding(44 Qs)66%70%77%76%68%64%61%56%67%56%67%55%64%62%54%59%52%45%
Structural Design & Loading(44 Qs)88%94%86%87%92%86%89%83%85%85%80%83%86%76%78%78%63%62%
Sustainability & Carbon(47 Qs)90%91%90%88%86%81%85%90%91%85%83%80%80%78%66%66%65%48%
Waterproofing & Below-Ground(42 Qs)75%73%71%64%72%67%70%64%63%61%68%67%68%66%63%51%47%33%
Category Deep Dive

Compare model performance by category

Select any of the 20 categories to see how each model performed.

Conclusions

What this means for UK construction professionals

  1. 1

    Read the table as 10 groups, not 18 places.

    Models within 1.5 percentage points of each other are tied. Three models lead on 86.5% to 85.1%; six more sit bunched together in the middle. Picking between models inside a band on these numbers would be reading noise.

  2. 2

    Useful for recall. Not a substitute for a competent person.

    The best model still leaves 13.5% of the available marks on the table. For anything that carries safety, regulatory or contractual weight, check the answer against the published document before you act on it.

  3. 3

    Ask for the clause, then go and read it.

    These models are strongest at telling you where a requirement lives. Used as a way into a document rather than a replacement for it, they save real time at low risk.

  4. 4

    Pay for the tool if the work matters.

    A 10.3 percentage point gap between paid and free is not a rounding error. If AI is touching billable work, the free tier is the wrong place to do it.

  5. 5

    Specialist and paywalled standards remain the weak spot.

    Accuracy tracks how freely available a document is. Areas that sit behind paywalls or in niche guidance score consistently worse, and those are often exactly where professionals need help.

  6. 6

    Treat any figure here as a snapshot.

    Several models tested are preview releases, and providers change models behind the same name. We publish the exact version identifiers so you can tell whether a result still applies.

All 13,747 answers were collected from the models on 28 July 2026, with scoring completed on 29 July 2026. The 951 questions were generated against published UK standards and Approved Documents, with each answer tied to a specific clause, table or section; they have not been independently verified by a chartered professional. Answers were graded by two AI models from different vendors, blind to which model produced each answer, and counted only where both agreed. Exact model versions, the scoring rubric and the study's limitations are set out in the methodology.

Built by Fabrick

Want a platform like this for your business?

This platform was built by Fabrick. We create bespoke digital tools, data platforms and content hubs for construction and built environment companies -- designed to demonstrate your expertise and generate qualified leads.

Carbon CalculatorsData DashboardsContent HubsLead-Gen ToolsCPD PlatformsProduct Selectors

Fabrick - Award-winning marketing specialists for the built environment