Free Sample
The Mimic's Problem
How AI Systems Learned to Hide What They Don't Understand
by The Field Researchers
Chapter 1: The Confidence Plague
In March 2023, a lawyer in New York submitted a legal brief to federal court that cited six cases. They were well-formatted, with proper case numbers and publication details. The judge would likely have missed that none of them existed. The lawyer hadn't invented them deliberately. He had copied them directly from ChatGPT, which had generated them with the same fluent confidence it applied to questions about the weather or the capital of France. The system had been asked to find cases supporting a legal argument. It could not actually search legal databases. What it could do—what it was built to do—was produce text that sounded like it was retrieving cases. The distinction mattered more than the judge initially realized. Sanctions followed. The lawyer's career absorbed the damage. But the real story was not about the lawyer's carelessness. It was about what the system had done: generated text indistinguishable from truth, with zero access to the underlying reality it was describing.
This is the opening pattern of our moment. Not errors that announce themselves. Not systems that fail gracefully, hedge their bets, or admit confusion. Instead: systems that sound wrong in ways that are nearly impossible to detect without expertise, infrastructure, or luck. A radiologist in a hospital in Pennsylvania runs a scan through an AI diagnostic tool that returns a high-confidence assessment of lung nodules. The assessment is wrong. But the confidence score is very high—0.94 out of 1.0. The radiologist, trained to trust validated algorithms, might anchor on that number. A medical student in São Paulo asks an AI chatbot whether a particular combination of medications is safe for a patient with kidney disease. The system provides a detailed, pharmacologically plausible-sounding answer that contains a critical contraindication the system has no way of knowing about. It sounds like knowledge. The student, lacking experience, might not catch the error until a patient is harmed.
These are not hypothetical examples. In 2023 and 2024, as large language models and computer vision systems proliferated into medical education, legal work, and clinical decision-making, a pattern emerged that went largely unremarked: the systems failed, but not in ways people were trained to recognize. They failed while sounding confident. They failed while sounding expert. They failed in a manner that was orthogonal to human intuitions about how failure should appear.
For decades, computer scientists built systems that would crash, freeze, or return error messages. These failures were alarming but legible. A user knew something was wrong. A system that returned a blank screen or a stack trace or a "does not compute" message was communicating its limit. The new systems do something stranger: they elaborate on their limits. They construct edifices of plausible text on foundations of no knowledge. And because language—human language, the kind we're reading now—is fundamentally a tool for conveying confidence and authority, systems trained to generate human-like text naturally generate confidence and authority whether or not they have any epistemic right to do so.
Consider the anatomy of failure in a specific case, examined in detail. In 2024, researchers at Stanford tested GPT-4, one of the most capable language models yet built, on a set of straightforward factual questions. Some of the questions had answers in the model's training data. Others did not—they asked about events that occurred after the training cutoff, or about obscure facts that statistically rarely appeared in the internet text the model learned from. You might expect the model to perform differently on these two categories. It did. It performed worse on questions where it lacked ground truth. But the signature of its failure was revealing: the model's confidence remained high across both categories. In fact, on many false claims, the model expressed certainty that exceeded its certainty on true ones.
This is not incidental. It is not a bug that can be patched out with better training. It emerges from the fundamental design of these systems, which are optimized to predict the next word in a sequence with high probability, regardless of whether that word corresponds to anything real. A model trained to produce text that reads like truth will gradually develop an internal structure that doesn't distinguish between "I have encountered evidence of this" and "this is the kind of text that statistically follows in this context." The two become neurologically identical.
The problem becomes visible in specific domains where consequences are severe and ground truth is testable. In medicine, this has begun to materialize in real time. A study in 2024 by researchers at Johns Hopkins examined how medical professionals interact with AI systems framed as "decision support." The framing matters. When a system is presented as an expert consultant, doctors are more likely to defer to it. And they are more likely to defer when it expresses confidence. In the study, when an AI system generated a diagnosis with high confidence language—"This finding is consistent with acute myocardial infarction"—physicians agreed with the assessment 70 percent of the time. When the same system, producing the same assessment, was framed with uncertainty language—"This could be consistent with acute myocardial infarction, but other diagnoses are also possible"—agreement dropped to 45 percent. The accuracy of the underlying assessment did not change. Only its presentation did. The model had no way of knowing whether it should be confident or uncertain. It simply produced the most statistically likely token given what came before, and human judgment reorganized itself around that confidence signal.
This is the confidence plague: not that AI systems are wrong—all systems are sometimes wrong—but that they are wrong with fluency. They do not announce their uncertainty. They produce text that bears all the surface markers of knowledge: technical terminology, causal narrative, specific detail, the rhetorical structure of expertise. A radiologist can call up images to verify a diagnosis. A judge can pull the case law. But an oncologist reviewing a treatment recommendation from an AI system, or a teacher assigning an essay written by a student using a language model, or a hiring manager reading a résumé auto-generated by a recruiting tool—these users exist downstream of the text, unable to see its origins, and the text looks like it came from somewhere.
Here is where the problem deepens. Many of these systems have been integrated into workflows in ways that presume correctness without requiring verification. A hospital system using an AI tool for discharge summaries might not route those summaries back through a physician before sending them to the patient's home care provider. A legal firm using document analysis software might trust its categorization of contract clauses without sampling. A content moderation system trained on examples of harmful speech will encounter new examples that aren't quite in its training set and will need to make decisions about them anyway, generating confidence scores that the decision-makers will read as probabilities when they are actually something far stranger: the normalized activations in the final layer of a neural network, which happen to correlate with correctness when the test distribution matches the training distribution, and correlate with nothing in particular when it doesn't.
The lawyer with the fictional cases was operating in a domain where verification is theoretically possible but practically constrained. He was under deadline pressure. The case law had to be relevant to a specific argument. When ChatGPT generated cases that sounded relevant and cited them with proper formatting—the kind of formatting that signals legitimacy to anyone who hasn't pulled the actual case—the pressure to move forward exceeded the impulse to verify every single citation. This is how the system infiltrated reality: not through brilliance, but through the basic human tendency to trust what sounds authoritative when time is short and you have skin in the game.
The medical domain reveals a different vulnerability. Doctors, unlike lawyers, have institutional training in verification. When a radiologist reads an AI-assisted scan, she is supposed to perform her own independent assessment. But institutional incentives corrode this in practice. If a tool is framed as "FDA-cleared" or "clinically validated," it carries the weight of regulatory judgment. If a hospital has invested in the system and incorporated it into the standard workflow, the doctor who consistently disagrees with it becomes a bottleneck. If the AI's recommendation happens to be right more often than the average radiologist—and many AI systems do perform better than average on their specific training distribution—then the doctor who defers to it will be correct more often than if she trusted her own judgment. The perverse incentive is built in: trusting the confident system produces better outcomes than catching its occasional failures.
This is only true up to a point. The point at which it becomes false is precisely the point the system cannot recognize: when the input diverges from its training distribution, when the problem becomes subtly novel, when the ground truth is genuinely ambiguous and the model must make a judgment call. In these moments, a well-calibrated system would express uncertainty. The systems we have built do not. They have been trained on human-generated text, which is often confident even when it shouldn't be. They inherit this tendency. They amplify it.
By late 2023, anecdotal reports were accumulating. An AI system used in hiring filtered out qualified candidates because their résumés didn't match patterns it had learned. A predictive policing algorithm, trained on historical arrest data, concentrated surveillance in neighborhoods it had learned were "high-crime areas," thereby increasing arrests there, which reconfirmed its priors. A content moderation system with high confidence in its classifications began making false positives on novel examples. A language model trained on scientific literature began generating plausible-sounding but fabricated citations in academic papers, leading to retractions and institutional embarrassment. In each case, the system had no mechanism for recognizing that it was operating outside its competence. It had no "I don't know" in its vocabulary—not in the sense of a response it could return, but in the sense of a genuine epistemic stance it could adopt. To return "I don't know" requires that the system recognize uncertainty, and recognition of uncertainty requires a ground-truth signal: access to the actual fact of the matter, which the system can then compare against its own probabilistic prediction.
Without that signal, the system optimizes for something simpler: coherence. It learns to generate text that flows, that doesn't contradict itself, that maintains internal consistency. These are valuable properties. But they are not properties of truth. A false narrative, told with internal consistency and supported by plausible evidence, is still false. A system optimized for coherence will eventually become a system optimized for the appearance of truth, which is not the same thing.
This distinction has always existed in human cognition—we have a long history of false confidence, motivated reasoning, and eloquent lies—but we have a social infrastructure for managing it. We have peer review. We have adversarial processes. We have the slow correction of repeated contact with reality. A doctor who consistently misdiagnoses loses patients and reputation. A lawyer who cites fabricated cases gets sanctioned. A scientist who publishes fraudulent work faces retraction and disgrace. These are not foolproof mechanisms, but they create feedback loops between confidence and correctness.
AI systems experience none of these loops. A language model that generates a fictional case law cite never finds out that it was wrong. A recommendation algorithm that directs a hiring manager's attention toward the wrong candidate has no way of knowing that this happened—there is no feedback that the candidate would have been an excellent fit. A medical AI system that gives bad advice in a rare case never encounters enough examples of that rare case to update its weights. It remains confident in its priors, which were wrong all along but were right often enough in its training distribution to survive deployment.
There is a category of system failure that engineers sometimes call "silent failure"—a system that returns a result but the result is wrong and the user has no way of knowing. Silent failures are worse than noisy failures. A system that crashes tells you to distrust it. A system that returns a wrong answer with high confidence tells you to trust it. This is what we have built.
The immediate question—"Why don't we just fix this?"—reveals the depth of the problem. Engineers have tried. Techniques like uncertainty quantification, confidence calibration, and ensemble methods can modestly improve a model's ability to express appropriate doubt. But these techniques work at the margins. They can make a system somewhat better at recognizing when it's operating outside its training distribution. They cannot make a system that has no ground-truth access suddenly aware of ground truth. They work by training the system to recognize statistical signatures of uncertainty—places where its own internal activations diverge, or where it generates contradictory outputs, or where its predictions consistently fail on a validation set. But a system can learn these signatures while still fundamentally unable to distinguish between "I lack knowledge" and "I know something implausible."
Consider a concrete attempt at calibration. You take a language model and you fine-tune it on examples where humans have labeled responses as confident or uncertain, correct or incorrect. You reward the model for expressing low confidence on questions it gets wrong and high confidence on questions it gets right. The model learns a correlation. It learns that certain types of questions—those far from its training data, or those phrased in ways that don't match its training examples—should be preceded by expressions of uncertainty. It learns to say "I'm not entirely sure" before answering questions about events in 2025. But it does not gain access to ground truth. It has learned a new pattern: that low-confidence language correlates with difficult questions. A sufficiently capable model can learn to generate uncertainty language without actually becoming uncertain. It has simply learned a new form of plausibility.
This is the structure of the problem that anchors this entire book. We have built a class of technology whose core operation is pattern completion. We have trained it on human-generated text, which contains all the surface markers of knowledge and confidence, often regardless of whether that knowledge is genuine. We have then optimized the system to output statistically likely text—which means outputting the kind of text that appears most frequently in the training set, which means outputting confident, coherent, elaborate text. We have then integrated these systems into domains where correctness matters, where users make decisions based on the output, where the consequences ripple through institutions.
The pattern that emerges across these domains is not random. It's structural. It's what you'd expect from a system that has no mechanism for truth but has every mechanism for plausibility. The systems do not fail into honesty. They fail into eloquence.
What follows is an investigation into how this happened—how we ended up deploying pattern-completion engines into the world while believing we had built understanding machines, and what this means for the institutions that now depend on them. The investigation is organized around a central claim: the problem is not that these systems are bad at learning from data. They are, in fact, remarkably good at that. The problem is that learning from data is not the same as understanding reality, and we have built an entire technological apparatus on the assumption that with enough data and enough parameters, the distinction would collapse.
It has not. The systems have become more sophisticated, more fluent, more capable at generating text that reads like expertise. But they have not become more truthful. In some ways, sophistication and truthfulness have become inversely related: the better the system is at generating plausible text, the harder it becomes to detect when that plausibility is a mirage.
Enjoyed the sample?
Get the full book — EPUB + PDF, no DRM, works on every reader.
Instant download · Kindle, Apple Books, Kobo, Google Play Books · No DRM