Independent research · AEC risk & assurance
How AI Fails, for Engineers Who’ve Never Trusted It
The engineers most skeptical of AI are often the ones best equipped to use it safely — because they already do the one thing that catches its failures. They check. A field guide to how AI fails in engineering practice, and where to point the judgment you already have.
Independent research on AI in licensed engineering practice. Views are the author’s own and do not represent any employer or client. Not legal, insurance, or professional advice — the judgment on any specific work remains the responsible engineer’s.
The engineers most skeptical of AI are often the ones best equipped to use it safely, because they already do the one thing that catches its failures. They check.
What the careful engineer usually lacks isn’t judgment. It’s a map of how AI fails, and AI does not fail the way people do. This piece is that map, written for the engineer who has never trusted the tool and doesn’t intend to start. It won’t teach you to use AI. It will tell you where AI breaks, so the judgment you already have lands where the failures actually are.
Carry one thing through all of it: AI produces the same confident, complete, fluent output whether or not it knows the answer. A person who is unsure shows it. AI doesn’t. That single fact changes where you have to look.
01Why this is written for the skeptic
The engineer with thirty or forty years of practice who hand-checks every result is not behind the curve. They are modeling the exact discipline the profession now requires. The NSPE Board of Ethical Review has been explicit that AI-generated technical work “requires at least the same level of scrutiny as human-created work.”[1] Skepticism is the correct posture.
But skepticism without a map wastes effort. If you don’t know how AI fails, you either check everything with equal force — exhausting, and unsustainable at production scale — or you check the wrong things and miss the ones that matter. Knowing the failure modes is what lets experienced judgment go where it counts.
Consider a failure that is easy to reproduce and easy to miss. A small design application — proofed by another engineer, reviewed, signed off, in active use — turns out to produce wrong results at its maximum and minimum input values. The cause isn’t a typo. Parts of it were built on approximation methods rather than the standard governing equations, and at the boundaries the approximation diverges from the exact result. Inside the normal working range, the two agree closely enough that nothing looks wrong; only at the extremes does the gap open up. The tool looks exact. It isn’t.
That failure mode deserves a moment, because it is a specific and recurring one, not a fluke. A proof that exercises typical inputs will pass a tool like this every time — the error lives only where the inputs are rare, which is precisely where a proof is least likely to go. “Proofed” and “tested where it breaks” are different claims, and the gap between them is exactly where an approximation hides. It is the whole problem in miniature: a plausible, reviewed, confident-looking tool that fails silently at the edges, and reveals it only if you go looking. The rest of this piece is where to look.
02The core insight: AI doesn’t fail the way people do
When a person is unsure of something, they signal it. They hesitate. They ask a question. They flag an assumption. They hand you something visibly incomplete. Over a career you learn to read those signals: the hedge in a sentence, the gap in a calc, the note in the margin. Uncertainty has a tell, and you’ve spent decades learning it.
AI removes the tell. It produces a complete, confident, fluent answer regardless of whether it knows, and a wrong answer arrives with exactly the same polish as a right one. Research on AI reliability calls this poor calibration: the model’s expressed confidence is decoupled from its correctness.[2]
To understand why, it helps to know what the tool actually is. A large language model is not a database and does not look anything up. It is a statistical text generator: given your prompt, it predicts the most plausible next words. When it has correct information, it produces correct text. When it doesn’t, it produces plausible text anyway, because producing plausible text is the only thing it does. Fabrication isn’t an exceptional malfunction — it is a predictable consequence of a system that generates likely language with no built-in guarantee that the language is true. The model has no internal mechanism that separates “I know this” from “this is what a confident answer would look like.”
The consequence for anyone checking the work is sharp: you cannot use the AI’s confidence, fluency, or formatting as any signal of correctness. The cues you’d read on a person — hesitation, roughness, a hedge — are simply absent. You have to verify the substance directly, every time, because the surface tells you nothing.
03The failure modes
1.Fabricated authority: citations, code sections, standards
AI invents plausible references. A code provision, a standard section number, an equation, a coefficient, presented cleanly and either non-existent or saying something different from what it claims. Because the tool generates plausible text, a fabricated AISC or ACI citation looks exactly like a real one.
The legal profession made this unmissable. In Mata v. Avianca (S.D.N.Y. 2023), a lawyer filed a brief citing court decisions that ChatGPT had invented whole: fake case names, fake citations, fake quotes, fake reasoning. Judge Castel found that at least six of them appeared to be “bogus judicial decisions with bogus quotes and bogus internal citations.”[3] The detail that matters most for our purposes came later. When opposing counsel couldn’t locate the cases, the lawyer went back and asked ChatGPT whether they were real. It told him they were, and told him they could be found in Westlaw and LexisNexis. He was sanctioned, and the case became the defining cautionary tale.[4] It has kept happening since, in courts around the world. There is nothing special about legal text that makes it more fabricable than an engineering citation. The mechanism is identical.
The second lesson from that exchange is the one engineers should internalize: asking the model to check its own work is not a check. The verification and the error come from the same generator.
And it happens even when the tool is otherwise good. A peer-reviewed study ran GPT-5 through 203 published questions from the orthopaedic in-training examination. It scored 78.3%, above the exam’s pass threshold and above the mean score of final-year residents — the highest accuracy the authors were aware of in the peer-reviewed literature. Yet in the subset of responses examined closely, references were fabricated or misrepresented in 33% of them. The rate was 50% among wrong answers, but still 15.9% among correct ones.[5] Roughly one in six correct answers arrived with a bad citation attached. The right answer and the fake source can sit in the same fluent paragraph.
In engineering, this shows up as a cited code edition or section that doesn’t exist or was superseded; an equation attributed to a standard that doesn’t contain it; a “typical value” presented with authority and no real source. I have been given wrong code references many times, across every major model I’ve used. It is avoidable, but only if you know how to work with the tool. A loose or over-optimistic prompt invites the model to drift into provisions that were never there, and if you don’t run your review against a checklist, it drifts behind the scenes — into territory you never see, because you didn’t look.
Which is the point worth stating plainly, because it is the one that carries liability: treat every AI citation as a lead, not a source. Open the actual code, the actual standard, the actual edition, and confirm the provision says what the AI claims, including exceptions and conditions. And remember whose name is on the seal. The AI provider’s terms of service hand the responsibility back to you. Its confidence is not your defense.
2.Stale knowledge, delivered confidently
An AI model’s training has a cutoff. It does not reliably know which edition of a code is current, and it will not tell you when it might be out of date. It will apply a superseded edition, an obsolete method, or a retired provision with complete confidence and no warning.
This is dangerous in engineering specifically because our field is edition-governed. The right answer under ACI 318-14 can be wrong under 318-19. Load combinations change between editions of ASCE 7. A detailing rule is quietly superseded. The tool has no dependable sense of “as of when,” and — worse — no built-in signal that says “this may be outdated.” It states the old rule with the same assurance it states the current one.
Web search partly mitigates the training cutoff, but only partly, and it introduces a second-order version of the same problem. The current published edition of a standard is not necessarily the edition your jurisdiction has adopted. Adoption lags publication, unevenly, by state and sometimes by locality. A model that correctly retrieves the newest edition can still be giving you the wrong governing document for your project.
The defense is a habit: anchor every code-dependent answer to the edition your project actually uses, confirmed against the document itself. Never assume the model applied the right one, because it has no reliable way to know which one governs.
3.Confidently wrong reasoning, and the trouble with tools it builds
This is the failure that fluency hides best: well-structured, confident technical reasoning that is subtly or grossly wrong. A wrong formula applied cleanly. A unit or sign error. A dropped boundary condition. An assumption that reads as reasoned but isn’t. The presentation is flawless; the substance fails. Because fluency and correctness are separate properties of the tool, a beautifully formatted derivation is not evidence of a correct one.
There is early research suggesting a particularly treacherous version of this. A 2026 preprint — single-author, not yet peer-reviewed, so treat it as a hypothesis rather than a finding — reports that in multi-step retrieval questions, supplying the model with one confirmed intermediate fact can raise its confident-wrong-answer rate. Supplying enough facts eliminates the effect; supplying one appears to make things worse. The authors call it anchored confabulation: a partial anchor commits the model to confidently completing the rest of the chain from memory.[6]
The study is about retrieval chains, not structural derivations, and I’d caution against treating the analogy as established. But the shape of it is worth holding in mind, because it inverts an intuition most of us have. We tend to read a correct opening as evidence of a sound method. If this effect generalizes, a correct first half makes a flawed second half more convincing, not less.
There’s a second version of this that matters more as engineers start building their own tools with AI. When you convert a spreadsheet into an application, the logic is legible; you can read what it does, because you wrote the spreadsheet. When you let generative AI build the logic, the result is harder to predict. Using an AI-built app can itself become a discovery process, where you learn what it actually does by running it and watching. That is exactly how a boundary-value failure like the one above stays hidden. The approximation lives inside the tool, out of sight, and shows itself only at the edges. An AI-built tool can encode an approximation that looks like an exact method and fail silently precisely where you didn’t test.
The defense is the one it has always been, and it does not get easier because the tool is new: independent verification by a method that can fail on its own terms. A hand calculation. A different established package. A limiting or bounding case. An order-of-magnitude check. The rule that makes it real is that the check must not inherit the AI’s assumptions. Re-derive from your own inputs. Don’t re-read the model’s work in different words — that isn’t a check, it’s a second reading of the same possibly-wrong reasoning.
4.It agrees with you, and it isn’t repeatable
Two smaller failure modes worth naming. AI tends to agree with the premise built into your question. Ask it to “confirm that this beam is adequate” and you have already tilted it toward confirming.
And the same question, asked twice or phrased differently, can return different answers. These systems are stochastic by default; a single run is one sample, not a verdict.
The habits that address both are cheap. Ask neutrally rather than leadingly. Ask the model to argue the opposite of what you expect. And when it matters, don’t let the answer depend on how you happened to phrase the prompt. If rephrasing changes the conclusion, you don’t have a conclusion yet.
04The profession already drew the line
None of this is new doctrine. It’s the old doctrine, with higher stakes.
The NSPE Board of Ethical Review addressed AI directly in Case 24-2 (2024), and the instructive part is a contrast within a single case. The same engineer used AI two ways. For a report, the engineer reviewed the draft thoroughly, cross-checked key facts against professional sources, and adjusted the text; on the verification question, the Board found that work remained under the engineer’s direction and control. For design documents produced with an AI-assisted drafting tool, the engineer performed only a cursory review. The client found misaligned dimensions and omitted safety features required by local regulations, and the Board concluded the engineer had failed to maintain responsible charge. Same engineer, same project, two outcomes, and on that axis the only difference was the checking.[1]
Two details in the facts of that case deserve more attention than they usually get.
The first is that the drafting tool was new to the market and the engineer had no prior experience with it. Sealing work produced by a tool whose failure modes you haven’t characterized is a distinct exposure from sealing work produced by a tool you know well. It is the reason the first question about any new tool is not “what can it do” but “where does it break.”
The second is that the report did not come through clean either, and the reasons had nothing to do with checking. The Board found that uploading the client’s information into an open AI interface was effectively putting confidential information into the public domain, without the client’s consent. It also found that the AI-generated report lacked citations to the technical authorities it drew on. Careful verification was necessary; it wasn’t sufficient. Confidentiality and attribution are separate obligations, and in day-to-day practice the confidentiality exposure is probably the more common one.
The Board’s standard is the same map this article draws. AI work requires at least the same scrutiny as human-created work, and the analogy the Board reaches for is directing an engineering intern: outline the guidelines and constraints, then consider and challenge what comes back rather than accepting it, and understand the output before you take it. The profession had, in fact, settled the underlying principle decades earlier. In 1990, considering computer-aided drafting and design, the Board held that an engineer may seal CADD-produced documents provided they were prepared under the engineer’s direction and control — and warned, in a passage later cited as AI became the next step in the same progression, that a tool used beyond its ability becomes a crutch and a substitute for judgment, which is a scenario for liability.[7] AI is that same rule, applied to a tool that fails in far less visible ways.
05Where to point the judgment you already have
You don’t need to learn to trust AI. You need to know that its output carries no honest signal of its own uncertainty, so you supply the scrutiny it can’t.
Distilled, the discipline is this. Treat AI output as an unverified draft from a confident stranger. Check its citations against the actual source, and don’t ask the model to check them for you. Anchor every code answer to your project’s governing edition. Verify the numbers by an independent method that doesn’t inherit the model’s assumptions. Write your prompts carefully — a loose or optimistic prompt is where the drift begins. And never let fluency stand in for correctness.
Speaking as an engineer who uses these tools daily and has spent months testing what each new model can and cannot do: even with strong engineering knowledge, AI makes mistakes, and they are often the small, quiet ones that a fast read slides past. When you are building or coding a multi-layer piece of software to design a building or a bridge, follow your checklist. Move through defined stage gates. Cross-check the output. At each gate, confirm the reasoning against a code provision, a standard, or a simple piece of mathematics that has to hold if the result is right. That is how you keep the tool inside the boundaries of your judgment instead of letting it wander past them.
The engineer who checks everything isn’t slow. In a world where AI produces confident, plausible, quietly-wrong output, that engineer is the control that works. The skill to add isn’t trust. It’s knowing the failure signature, so the checking lands where the failures are.
This is why the responsible-charge methodology I’ve been developing treats independent verification as the irreducible act, and why its first step is understanding a tool’s failure modes before relying on it. The framework is published in full at aeco.digital/methodology-rca. But this article stands on its own: know how AI fails, and point your judgment there.
Notes
- 1.NSPE Board of Ethical Review, Case 24-2, Use of Artificial Intelligence in Engineering Practice (July 18, 2024). Responsible charge is defined in NSPE Position Statement No. 10-1778.↩a↩b
- 2.On calibration and hallucination generally: Ji et al., “Survey of Hallucination in Natural Language Generation,” ACM Computing Surveys 55(12), 2023; Huang et al., “A Survey on Hallucination in Large Language Models,” ACM Transactions on Information Systems, 2025.↩
- 3.Mata v. Avianca, Inc., No. 22-cv-1461 (S.D.N.Y.), Order to Show Cause (Castel, J., May 2023).↩
- 4.Sanctions opinion, Mata v. Avianca, June 22, 2023.↩
- 5.“Accuracy Is Not Enough: Reasoning and Reference Reliability in Orthopaedic Large Language Model (LLM) Applications.” GPT-5 scored 78.3% (159/203) on publicly available orthopaedic in-training examination questions, against a 67% pass threshold and a 73% mean final-year-resident score. In an 88-response subset, hallucinated or misrepresented references occurred in 33% of responses overall — 50% of incorrect answers and 15.9% of correct ones.↩
- 6.A. Lathkar, “Anchored Confabulation: Partial Evidence Non-Monotonically Amplifies Confident Hallucination in LLMs,” arXiv preprint (2026). Under review; not peer-reviewed at time of writing. Finding concerns multi-hop retrieval questions, not engineering derivations; the analogy here is the author’s, not the study’s.↩
- 7.NSPE Board of Ethical Review, Case 90-6, Use of CADD System (1990), discussed in Case 24-2. See also Case 98-3, holding that technology must not substitute for engineering judgment.↩
The research is published in full and free to read. If this is your world, register your interest and we will share what we learn.
Register interest →Method. Claims are anchored to primary sources where they exist — ethics rulings, court orders, and peer-reviewed literature — and every source is listed in the Notes above. Where a claim rests on work that is not peer-reviewed, the text says so at the point of use rather than only in the note: see the discussion of the 2026 preprint in Section 03, where the analogy to engineering derivations is identified as the author’s own and not the study’s. Sources checked September 2026.
Disclaimer. Independent research for general information. Not legal, insurance, or professional advice, and not a substitute for the engineer’s own judgment on any specific work. Code editions, adoption status, and professional obligations vary by jurisdiction — confirm against the governing documents for your project. © 2026 AECO.digital.