Ask an AI model a question and it will give you an answer. It will be well written. It will sound certain. And you will have no idea whether it is true.
For a birthday message, that hardly matters. For a policy brief, a patient record or a decision about a pipeline, it matters a great deal.
We have built AI tools for exactly that kind of work. E4CInsights turns years of sustainable development research into policy briefs. PipelineGPT lets engineers put questions to their own inspection and incident records. The same lesson came out of both: in serious work, an answer is only as good as your ability to check it.
Fluent is not the same as correct
These models are trained to produce language that reads well. That is a different skill from being right, and from the outside the two look the same.
The danger is not that the model is often wrong. It is that when it is wrong, it is wrong in the same calm, confident voice it uses when it is right.
Three rules we build by
- Every claim points to its source. An answer comes with the document, section and page it was drawn from, so the reader can check it in a minute.
- The AI answers from your documents, not from memory. If the material does not contain the answer, the right response is to say so.
- A person signs off where it counts. Anything that could lead to an operational decision is held for a qualified reviewer before anyone acts on it.
None of this is exotic. It is how any careful organisation already treats a junior analyst's work: show your sources, stick to the evidence, and have someone senior review it before it goes out.
Why the human stays
Human review is sometimes described as a temporary measure, something to remove once the models improve. We see it the other way round. It is what makes the tool usable in the first place.
An engineer will not act on a recommendation they cannot trace. A regulator will not accept "the system said so". Keeping a person in the loop, and keeping a record of what they decided, is what allows an organisation to adopt AI without lowering its standards.
What this looks like in a real product
Principles are easy to state, so here is how they show up in something we built.
In PipelineGPT, an engineer types a question in ordinary language. The system searches the operator's own inspection reports and incident records, and the answer comes back with a reference for each claim: which document, which section and, for a PDF, which page.
If the answer contains something that could lead to action in the field, such as a repair or a pressure reduction, it is not shown to the operator straight away. It goes to a qualified engineer, who can approve it, correct it or reject it. Every one of those decisions is logged.
The documents themselves stay on the operator's own servers. You can read more in the PipelineGPT case study, and about the same approach applied to policy research in the E4C case study.
Where AI helps most
None of this is an argument against AI. It is very good at certain things, and the caution above is what lets you use it for them.
- Finding the relevant passage in thousands of pages in seconds.
- Summarising long documents, as long as the summary links back to the source.
- Drafting, where a person will read and edit the result anyway.
- Answering routine questions from an agreed set of material.
It is weakest where there is no source to check against and nobody reviewing the output. Unfortunately, that is where it is often used first.
The useful question is not how clever the AI is. It is how quickly a person can tell when it is wrong.
If you are considering AI for your organisation
Ask whoever is selling it three things. Where does each answer come from? What happens when it does not know? Who checks the output before it is used?
If the answers are clear, you are looking at a tool. If they are vague, you are looking at a demo.



