How Accurate is AI for UK Construction?
951 questions. 18 AI models. 20 categories. Graded blind, with the workings published.
951
Questions
18
AI Models
20
Categories
86.5%
Best Score
How we tested
The short version. The full methodology lists every model version, the scoring rubric, and where this study is weak.
951 technical questions
Building regulations, British Standards, health and safety, fire safety, structural design and 15 other categories. Every answer is tied to a named clause, table or section in a published UK document.
18 models, exact versions published
12 paid and 6 free models from Anthropic, OpenAI, Google, xAI, Mistral, DeepSeek, Moonshot and Perplexity. Each is listed with the precise API identifier tested, not just a marketing name.
Graded blind, by two AIs
Each answer is scored Correct, Partial or Wrong by two models from different vendors, neither told which AI wrote the answer. A score counts only where both agree. They agreed 80.3% of the time.
Identical prompt for every model
No documents supplied, no web search requested, no follow-up questions. This measures what each model knows unaided, which is how most people actually ask it.
All 18 models, scored and grouped
Across 951 questions on UK construction standards and regulations. Models close enough to be tied are grouped rather than separated by place. Hover any bar for the correct, partial and wrong split.
Read these as 10 groups, not 18 places. Models within 1.5 percentage points of each other are tied, because gaps that small sit inside the margin of the scoring method. Hover any bar for the exact model version tested and the models it ties with.
What the data reveals about AI in UK construction
86.5%
Highest overall score
Claude Opus 5 led the field. No model was fully right on more than 87% of the marks available, which puts a hard ceiling on how far any of them can be trusted unaided.
10.3pp
Paid versus free gap
Paid models averaged 80.5% against 70.2% for free ones. If you use AI for construction work, the subscription is doing something.
7.7%
Lowest error rate
Even the best model got 59 of 951 questions outright wrong. The weakest, Claude Haiku 4.5, got 277 wrong (36.6%).
13.5%
Not fully right, best model
Counting partial credit, the leading model still failed to give a complete, correct answer on 13.5% of the marks available. That is the number to hold in your head before relying on any of this.
80.3%
Inter-judge agreement
Two AI judges from different vendors agreed on 80.3% of the 13,747 answers. Disagreements were excluded rather than settled by picking a favourite judge.
10
Groups, not places
Models within 1.5 percentage points of each other are too close to separate, so the 18 models resolve into 10 groups. Three lead the field; six more are bunched in the middle.
Is paying for AI worth it in construction?
“Paid” means a professional would need a paid subscription or paid API access to reach that model. The full definition is in the methodology.
Paid models average
80.5%
12 models tested
Free models average
70.2%
6 models tested
All 20 categories across all 18 models
Green is good. Red is risky.
| Category | Claude Opus 5 | Claude Fable 5 | GPT-5.6 Sol | Claude Opus 4.8 | Kimi 3 | Gemini 3.6 Flash | Gemini 3.1 Pro | GPT-5.6 Luna | Grok 4.5 | Perplexity Sonar Pro | GPT-5.6 Terra | Perplexity Sonar | Claude Sonnet 5 | DeepSeek V4 Pro | DeepSeek V4 Flash | Mistral Large | Mistral Medium | Claude Haiku 4.5 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Accessibility & Inclusive Design(39 Qs) | 78% | 79% | 74% | 71% | 70% | 72% | 68% | 74% | 73% | 62% | 77% | 59% | 69% | 62% | 60% | 49% | 47% | 40% |
| British Standards - Concrete & Steel(48 Qs) | 91% | 90% | 87% | 87% | 90% | 92% | 90% | 84% | 88% | 81% | 81% | 80% | 91% | 87% | 73% | 71% | 57% | 57% |
| British Standards - Other Materials(45 Qs) | 93% | 92% | 91% | 92% | 95% | 93% | 87% | 94% | 91% | 90% | 85% | 92% | 82% | 85% | 80% | 77% | 66% | 55% |
| Building Regulations - Fire Safety(49 Qs) | 93% | 90% | 88% | 91% | 89% | 87% | 86% | 87% | 91% | 92% | 86% | 83% | 75% | 83% | 74% | 74% | 58% | 46% |
| Building Regulations - Other Parts(44 Qs) | 83% | 82% | 77% | 80% | 76% | 86% | 74% | 70% | 76% | 74% | 76% | 76% | 71% | 66% | 56% | 68% | 58% | 44% |
| Building Regulations - Thermal(47 Qs) | 84% | 80% | 81% | 79% | 80% | 78% | 87% | 78% | 80% | 83% | 81% | 90% | 66% | 64% | 60% | 42% | 63% | 39% |
| Construction Technology(35 Qs) | 87% | 86% | 79% | 83% | 83% | 72% | 68% | 76% | 70% | 79% | 74% | 69% | 74% | 67% | 61% | 54% | 61% | 52% |
| Contracts & Procurement(44 Qs) | 94% | 92% | 86% | 90% | 89% | 86% | 82% | 80% | 84% | 80% | 81% | 78% | 89% | 74% | 59% | 58% | 59% | 39% |
| Demolition & Refurbishment(28 Qs) | 94% | 90% | 96% | 83% | 86% | 76% | 71% | 90% | 75% | 86% | 90% | 85% | 82% | 75% | 65% | 73% | 66% | 57% |
| Environmental & Contamination(34 Qs) | 94% | 95% | 92% | 95% | 95% | 80% | 80% | 86% | 88% | 89% | 88% | 92% | 94% | 81% | 79% | 67% | 59% | 48% |
| Fire Safety - Post-Grenfell(41 Qs) | 88% | 90% | 93% | 92% | 90% | 85% | 88% | 85% | 85% | 96% | 80% | 89% | 84% | 79% | 77% | 71% | 68% | 47% |
| Health & Safety / CDM(46 Qs) | 99% | 100% | 100% | 100% | 99% | 98% | 99% | 99% | 98% | 99% | 99% | 100% | 96% | 94% | 89% | 90% | 93% | 74% |
| MEP & Building Services(41 Qs) | 74% | 76% | 76% | 70% | 73% | 69% | 74% | 76% | 67% | 71% | 75% | 70% | 74% | 73% | 65% | 53% | 49% | 34% |
| Materials & Products(40 Qs) | 86% | 78% | 82% | 77% | 71% | 86% | 83% | 79% | 76% | 72% | 78% | 75% | 74% | 68% | 72% | 65% | 59% | 44% |
| NHBC Standards(39 Qs) | 68% | 75% | 79% | 74% | 66% | 71% | 68% | 72% | 67% | 71% | 74% | 68% | 57% | 61% | 69% | 67% | 51% | 61% |
| Planning & Permitted Development(51 Qs) | 94% | 94% | 95% | 87% | 94% | 91% | 94% | 92% | 88% | 98% | 86% | 96% | 82% | 83% | 76% | 70% | 66% | 68% |
| Roofing & Cladding(44 Qs) | 66% | 70% | 77% | 76% | 68% | 64% | 61% | 56% | 67% | 56% | 67% | 55% | 64% | 62% | 54% | 59% | 52% | 45% |
| Structural Design & Loading(44 Qs) | 88% | 94% | 86% | 87% | 92% | 86% | 89% | 83% | 85% | 85% | 80% | 83% | 86% | 76% | 78% | 78% | 63% | 62% |
| Sustainability & Carbon(47 Qs) | 90% | 91% | 90% | 88% | 86% | 81% | 85% | 90% | 91% | 85% | 83% | 80% | 80% | 78% | 66% | 66% | 65% | 48% |
| Waterproofing & Below-Ground(42 Qs) | 75% | 73% | 71% | 64% | 72% | 67% | 70% | 64% | 63% | 61% | 68% | 67% | 68% | 66% | 63% | 51% | 47% | 33% |
Compare model performance by category
Select any of the 20 categories to see how each model performed.
What this means for UK construction professionals
- 1
Read the table as 10 groups, not 18 places.
Models within 1.5 percentage points of each other are tied. Three models lead on 86.5% to 85.1%; six more sit bunched together in the middle. Picking between models inside a band on these numbers would be reading noise.
- 2
Useful for recall. Not a substitute for a competent person.
The best model still leaves 13.5% of the available marks on the table. For anything that carries safety, regulatory or contractual weight, check the answer against the published document before you act on it.
- 3
Ask for the clause, then go and read it.
These models are strongest at telling you where a requirement lives. Used as a way into a document rather than a replacement for it, they save real time at low risk.
- 4
Pay for the tool if the work matters.
A 10.3 percentage point gap between paid and free is not a rounding error. If AI is touching billable work, the free tier is the wrong place to do it.
- 5
Specialist and paywalled standards remain the weak spot.
Accuracy tracks how freely available a document is. Areas that sit behind paywalls or in niche guidance score consistently worse, and those are often exactly where professionals need help.
- 6
Treat any figure here as a snapshot.
Several models tested are preview releases, and providers change models behind the same name. We publish the exact version identifiers so you can tell whether a result still applies.
All 13,747 answers were collected from the models on 28 July 2026, with scoring completed on 29 July 2026. The 951 questions were generated against published UK standards and Approved Documents, with each answer tied to a specific clause, table or section; they have not been independently verified by a chartered professional. Answers were graded by two AI models from different vendors, blind to which model produced each answer, and counted only where both agreed. Exact model versions, the scoring rubric and the study's limitations are set out in the methodology.
Want a platform like this for your business?
This platform was built by Fabrick. We create bespoke digital tools, data platforms and content hubs for construction and built environment companies -- designed to demonstrate your expertise and generate qualified leads.
Fabrick - Award-winning marketing specialists for the built environment