It's a question we hear a lot, and it's a good one. ChatGPT, Claude and Copilot are remarkable. Anyone in your business can paste in a specification, ask for a summary, and get something useful back in seconds. So when someone proposes AI tooling for construction planning, a reasonable person asks: why can't we just use the chatbot we already have?
I'm not going to pretend the answer is "because chatbots are rubbish". They aren't. I use them daily, and I'd guess you do too - around 75% of construction project professionals now report using AI in some form, up from 15% two years ago. The chat interface is the reason that happened. It made AI accessible to people who'd never have touched it otherwise, and that's been a genuinely good thing for the industry.
But there is a difference between a tool an individual finds useful and a capability a business can rely on. And that difference is worth exploring carefully, because it's exactly where governance and compliance live. We end up with a set of choices, each with consequences.
What chat is genuinely good at
Let's give credit first, because the case for structure only makes sense once you're clear about where chat shines.
Chat is superb for exploration. Thinking through a problem, testing an idea, asking "what am I missing here?", drafting a first version of a letter you'll rewrite anyway. It's brilliant for one-off questions with low stakes - what does this clause typically mean, how is this usually sequenced, explain this regulation in plain English. And it's a personal productivity multiplier: each individual, working in their own way, gets faster.
If that's all your organisation needs from AI, a chatbot subscription could be enough. Plenty of value has been created that way already and will continue to do so.
The problems start when the output stops being a private draft and starts being a business record. When it feeds a decision, a report to a client, a compliance submission, or the golden thread. At that point the question changes from "is this useful?" to "can I rely on it, repeat it, and prove it?" And those are questions chat wasn't designed to answer.
Consistency: will I get the same answer twice?
Ask a chatbot the same question on Monday and Thursday and you'll get different answers. Not wildly different - but different in structure, emphasis, and sometimes substance. The models are probabilistic by design. Add to that the human variable: your best planner writes a careful, detailed prompt; a colleague under pressure writes two lines. Same tool, same project, very different outputs.
For an individual, this barely matters. You read the answer, apply judgement, move on. For an organisation, it matters enormously. If your progress reports, risk reviews or compliance checks are produced through ad-hoc chat, their quality depends on who asked, how they asked, and which day it was. Between projects, the variation compounds - every project team develops its own habits, its own prompts, its own gaps.
Compare that with a defined workflow: the same steps, the same inputs, the same checks, every time, regardless of who runs it. That's not an AI concept. It's just what a business process is. We'd never accept a quality inspection regime that changed shape depending on who happened to do it and how they felt that morning. But that's what ad-hoc chat gives you.
Consistency is also what makes improvement possible. If every run is different, you can't tell whether a bad output was the process or the person. If every run is the same, you can find the weakness, fix it once, iterate and drive continuous improvement.
Sources: where did this answer come from?
When a chatbot gives you an answer, ask yourself a simple question: which documents did it actually use?
Sometimes you pasted them in, so you know. Often, though, the answer is a blend - some of what you provided, some of the model's general training, some inference filling the gaps. The result reads confidently either way. And that's the danger. A fluent answer with no traceable sources isn't information; it's little more than opinion.
In most industries that's a nuisance. In construction it's a liability. The Building Safety Act's golden thread requirements are explicit: information about a building must be accurate, current, and traceable - you need to be able to show what changed, when, and why. ISO 19650 information management has the same principle. If an AI-assisted output ends up in your project record, "the chatbot said so" is not a source. An auditor, a regulator, or legal counsel in a dispute will ask: what information was this based on, and was it the current revision? You need an answer.
Structured AI workflows can be built to fill this gap - every output carrying references to the specific documents, versions and data it drew from. Not because it's using advanced AI, although it is, but because the workflow controlled what went in. Which brings us to the next point.
Completeness: did it see the whole picture?
A chatbot knows exactly what you gave it, and nothing else about your project. That's easy to forget, because it talks as though it knows everything.
So the summary of the specification is a summary of the sections you pasted. The risk review covers the risks in the documents you remembered to attach - not the ones in the subcontractor's programme you didn't think to include, or the site diary entry from three weeks ago, or the revised drawing issued on Friday. The model can't tell you what's missing, because it doesn't know what exists.
This is probably the least discussed weakness of the chat approach, and I'd argue it's the most dangerous one, because the failure is silent. An incomplete answer looks identical to a complete one.
The alternative is to invert the relationship: instead of a person deciding what the AI sees, the system assembles the full project picture - documents, registers, programmes, correspondence, current revisions - and the AI works from that, every time. The person's judgement is then spent where it belongs, on the output, rather than on the thankless task of remembering which of forty documents matter today.
Guardrails: what stops it going wrong?
A chat conversation is open-ended. That's its charm, and its risk. Nothing constrains what's asked, what's assumed, or what's done with the answer. There's no natural point where a check happens.
Step-by-step workflows provide something the chat window structurally can't: guardrails. Each step has a defined purpose, defined inputs, and a defined output. The AI does the specific thing the step calls for - extract these items, check against this standard, flag these discrepancies - rather than freewheeling. Human review sits where it's needed, at the decision points, not sprinkled across a long conversation. If a step fails or a check doesn't pass, the process is aware of the failure, clearly identified and actionable.
This mirrors how the industry already thinks about safety-critical work. We don't rely on competent people alone; we use permits, hold points, inspection and test plans. Nobody considers an ITP an insult to the tradesperson's skill. It's how organisations make quality outcomes repeatable and bad ones catchable. Applying the same thinking to AI isn't caution for its own sake - it's the only approach consistent with how we manage every other risk on a project.
There's a quieter guardrail benefit too: scope. A well-designed workflow only touches the data it needs, with the permissions of the person running it. An open chat window into which staff paste whatever seems relevant - including, inevitably, commercially sensitive or personal information - is a data governance headache that most organisations haven't yet fully confronted.
Versions: what did we know, and when did we know it?
Construction disputes are archaeology. Months or years after the event, I know as I spent nearly a year working full-time on a dispute, which was a massive eye opener and had a lot of interesting learnings, doing forensic research into what was known at the time a decision was made. Which drawing revision. Which programme version. Which conditions.
Chat conversations are miserable evidence. They live in personal accounts, scroll away into history, get deleted, and record neither which document versions were discussed nor how the output evolved. If a report went through six chat iterations before someone pasted the final version into Word, the lineage is gone. You have an output with no provenance.
A structured system versions three things a chat can't:
- the inputs, which revision of each document was used;
- the outputs, each generated report or check, retained and dated; and
- the process itself, meaning which version of the workflow, with which steps and checks, produced it.
When someone asks in two years why the March report said what it said, the answer exists. That's the golden thread principle applied to AI-assisted work - and it's worth noting the golden thread guidance explicitly calls for version control and an audit trail of changes. Regulators have already told us what good looks like. Ad-hoc chat can't provide it.
The benefits nobody puts on the list
Beyond the five above, a few other advantages are worth naming.
The capability belongs to the organisation, not the individual. Prompting skill is real, and in a chat-only world it lives in people's heads and personal chat histories. When your best "AI person" leaves, their prompts, habits and hard-won lessons leave with them. A workflow is organisational knowledge - documented, shared, improvable, and always there when you need it.
New people are productive immediately. Running a well-built workflow doesn't require prompt craft. That matters for an industry with the skills pressures ours has - you want your tools to reduce the experience needed to do the job well, not add a new specialism.
Costs become predictable. Ad-hoc chat usage is invisible until the invoice arrives. Defined workflows have a clear cost footprint, which makes AI a line item you can plan rather than a habit you discover.
And the process itself becomes auditable. When a regulator, client or insurer asks "how do you use AI in your business?" - and they are starting to ask - "our staff use chatbots as they see fit" is an uncomfortable answer. "These defined workflows, with these checks, these human review points, and this audit trail" is a governance position.
So which should you choose?
For me it's clear that both are important. This isn't a case where one approach kills the other.
Chat is the right tool for thinking: exploring, drafting, learning, asking. Structure is the right tool for delivering: any output that feeds a decision, enters a project record, goes to a client, or touches compliance. The mistake isn't using chatbots. The mistake is not noticing the moment a chat output crosses the line from someone's draft working into the official record of a project. At that point, the defined and repeatable process is required.
Ask yourself these five questions of any AI-assisted output your business relies on:
- Would we get materially the same result if someone else ran it tomorrow?
- Can we show exactly which information, at which revision, it was based on?
- Do we know it considered all relevant items, not just what someone remembered to include?
- Was there a defined check before it was acted on?
- Could we reconstruct this in two years, in front of someone sceptical?
If the answer to all five is yes, it doesn't much matter what tool produced it - you have a working governed process. If the answer is no, you've got a well-written piece of opinion sitting in your project record.
Full disclosure: we're building PlanOps precisely because we think construction deserves AI in an easy-to-use package that answers yes to these questions. But whatever tools you choose - ours, someone else's, or workflows you build yourself - these questions are massively important. And they will only get more so in our increasingly AI-dominated world. They're the same questions the industry already asks of every drawing, every test certificate and every inspection record. AI shouldn't get an easy pass just because it's new.
Sources and further reading
- BSR Compliance - The Golden Thread of Information
- ICE - The golden thread running through the new building safety regime
- RICS - How BIM helps manage building safety for the golden thread
- RICS - ISO 19650 Part 6: standard enhances golden thread and risk management
- Asite - Golden thread reporting: how to demonstrate control under the Building Safety Act
- ONCREO - Golden thread of information in construction: complete guide
