AI makes things up: construction needs an ITP for it
How should construction combat AI hallucinations? With an inspection and test plan (ITP) is Ian Yeo’s answer, born of experience.

During testing, our AI reviewed a RAMS document that it never received. The submission had come in through a subcontractor submission page. The review ran, and back it came with a full verdict: non-compliant, with all 19 assessment criteria flagged as critical gaps and references to UK safety legislation. It read like a thorough, rejected safety audit.
But the AI had never seen the document.
A fault earlier in the process meant the content had never reached the model. Faced with an empty submission, it did not stop and say there was nothing to assess. It produced a plausible-looking answer anyway.
A safety audit of nothing, delivered with complete confidence.
We caught the failure in testing and fixed the process that allowed it to happen. More importantly, it changed the way we thought about AI reliability.
The question is not whether AI can hallucinate. It can. The useful question for construction is: what stops that hallucination from becoming part of a permanent project record?
This is not a model problem you can wait out
There is a comfortable assumption that hallucination is mainly a sign of immature AI. Models improve quickly, so perhaps we should just wait for better ones. That is not a sensible control strategy.
Models are getting better, and in many cases they make fewer mistakes. But “fewer” is not the same as “none”, and the remaining errors can be extremely convincing.
That matters in construction. A fluent but invented paragraph in a casual chatbot conversation is one thing. The same behaviour inside a RAMS review, drawing comparison, tender assessment or progress report is quite another.
So, we stopped treating reliability as something the model alone should solve. Instead, we started treating it as a process problem.
In construction, we already know how to do this. We do not assume that our people will never make mistakes. We use permits, inspections, witness points, hold points and inspection and test plans so that errors are prevented where possible – otherwise they are caught before they become consequential.
AI needs the same treatment.
An ITP for AI
Every AI task in PlanOps runs as a structured sequence, not an open-ended conversation.
The system fetches information, checks it, asks the model to perform a tightly controlled task, validates the result and decides whether the process can continue.
From failures we found in testing, five controls became particularly important.
1. No information, no judgement
The RAMS failure had a simple cause: nothing checked that the document content actually existed before the assessment started.
We now put a gate before any analysis step. The process checks that the evidence is present and readable. If it’s not, the task stops and someone is informed. No verdict is produced.
We found the same pattern elsewhere. A drawing comparison should not compare a first issue against an earlier one. A tender review shouldn’t continue if a submission hasn’t been received.
An assessment of nothing should not be unlikely. It should be impossible.
2. Do not let the model invent references
Construction systems are full of references and identifiers: document numbers, works package IDs, classifications and asset IDs. Language models are very good at producing text that looks exactly like those things.
If an AI needs to choose an identifier, we first provide the real available options. The model chooses from that list, and its choice is checked again.
We applied the same principle to Uniclass codes after finding that a model could generate a code that looked perfectly credible but did not exist.
A nearly-right classification can be worse than no classification because incorrect information can sit neatly without attracting attention.
3. Use systems for facts and AI for judgement
Some hallucinations happen because we ask the model to do the wrong job. We once asked AI to provide links to company logos. It supplied plausible URLs, many of which were wrong.
That was not really a model problem. It was a wrong task problem. The answer was not a better prompt. We replaced that step with a conventional narrow software service.
The same rule applies to analysis. If you ask an AI to compare this month with last month, you need to provide the data. If information is not available, we should not invite models to fill gaps.
4. Show the evidence boundary
There is another failure mode that receives less attention than hallucination: omission. AI can produce a perfectly reasonable assessment from an incomplete set of information. Unless you know what the model saw, the result can look indistinguishable from one based on the complete record.
That means the output should show the extent of the information used, the evidence boundary. Which documents did it read? Which revisions? What relevant information was not available or not included?
If a reviewer can see that the structural steelwork specification wasn’t read, they can immediately judge the result accordingly.
For someone reviewing an AI-generated conclusion, visibility of the source material is more valuable than a model accuracy claim.
5. Put people at the consequential hold points
The first four controls can largely be automated. The final one is deliberately human.
Where an AI output feeds an important decision or becomes part of a permanent record, a responsible person should review it. That review should sit at a defined point in the workflow, not exist as a vague expectation that somebody will probably check the result.
The important thing is that the hold points, which require a person, are placed where the consequences are.
What these controls cannot do
None of this makes AI infallible. Even when a model receives the correct document, uses valid references and reports on the information used, it can still reach the wrong conclusion.
It can misunderstand wording, overlook an implication or give too much importance to a single item. The risk is managed and reduced, but not eliminated.
That’s another reason that “ITP for AI” is a useful mental model. An inspection regime doesn’t claim that defects can never occur. It’s a way of making outcomes more repeatable, identifying failures and stopping them from travelling further than they should.
The questions construction should ask
When someone offers AI that will generate or assess information that could end up as part of a project, “does it hallucinate?” is not a particularly useful question.
Instead ask:
- What happens if information is missing?
- Can the model create references, IDs or codes that do not exist?
- Can I see exactly what evidence it used?
- Can I see what relevant evidence it did not use?
- What happens when an automated check fails?
- Where does a person approve the result before it becomes a record?
Specific answers to those questions tell you far more than a general claim that an AI system is accurate.
Construction has spent decades building controls around people, processes and physical work because competence alone is not a sufficient safeguard.
AI shouldn’t get a special exemption just because it works quickly and writes confidently. Treat it like any other part of a controlled process: define what it is allowed to do, verify what goes in, check what comes out and put hold points where failure matters.
Keep up to date with DC+: sign up for the midweek newsletter.