AI has a strange ability to sound equally confident when it is right, when it is partly right, and when it is completely wrong. That’s why people have difficulty placing AI correctly in their minds.
Ask an AI to write a Python function and it may produce code that runs perfectly on the first attempt. Ask the same AI a medical, legal, historical or financial question and the answer may arrive with exactly the same fluency, confidence and polished language. But the answers may not deserve the same level of trust.
The problem is not simply that AI sometimes makes mistakes. Humans make mistakes too. The more important difference is how AI produced the answer and, crucially, what mechanisms exist to verify that it is factually correct. The distinction starts with something I describe as deterministic versus probabilistic computing.
Deterministic vs Probabilistic Computing
Run the same conventional Python or C program with the same input a million times and you get the exact same result a million times. Now copy-paste the same prompt into five separate conversations with the same LLM, and you may get five different answers. That difference gives us a useful starting point.
Traditional software generally follows predefined rules:
Rules + input → execution → result
An LLM works differently:
Prompt + learned patterns → probabilities → “generated” answer
Technically, this distinction is simplified. Conventional software can use randomness, and an LLM can be configured to behave more deterministically. But for understanding the practical difference between the two, this simplification is useful.
A normal computer program executes rules somebody has defined with code. An LLM generates an answer that fits what it has learned from vast amounts of data.
That does not make the LLM’s answer random. Quite the opposite: modern LLMs have learned astonishingly rich patterns from enormous amounts of data and can produce remarkably sophisticated answers. But there is something important missing from that process. The model does not have an internal truth meter. It can generate an answer that looks highly appropriate without independently establishing that the answer corresponds to reality. That is where much of the confusion around AI begins.
Plausibility Is Not Verification
Imagine asking an AI a question and receiving a confident answer. The answer may be grammatically perfect, contain a convincing explanation, follow a logical chain and sound exactly like something an expert would say. None of those properties proves that it is true. The probability that a sequence of words is a good continuation is not the same thing as the probability that the statement corresponds to reality.
Put more simply:
AI generating an answer asks: “What fits here?”
Independent verification asks: “Is this actually true?”
This matters because AI is extraordinarily good at producing plausible answers. It can interpolate between things it has learned. It can combine information in new ways, propose explanations and construct arguments that may never have appeared verbatim in its training data. However, this plausible explanation can still be wrong, because the underlying information may be wrong, an assumed, external context may be missing, or the question may be about something that changed after the model was trained about it. Or worse yet, the model may construct an answer that fits extremely well despite having no factual basis at all.
This is why the familiar term hallucination can actually understate the larger issue. The problem is not merely that AI occasionally invents things. The deeper problem is that the mechanism that generates the answer is not, by itself, a mechanism that proves the answer is true.
But AI Can Reason, Right?
This is where another misconception appears. It is tempting to respond to the problem above by saying that AI does not reason and merely predicts words. I think that description is not useful. Modern AI can perform impressive multi-step reasoning, compare hypothesis, write algorithms, diagnose and solve problems, plan actions and reach conclusions that require far more than simple recall. Whether that reasoning is identical to human reasoning is a different and somewhat a philosophical discussion. For practical purposes, the important distinction is much simpler:
Reasoning is not verification.
An AI can reason its way to the wrong conclusion, and so can a human. The real question is, can an LLM reliably recognize when its own conclusion needs to be challenged?
When Reality Pushes Back
Software development is one of the clearest examples of why AI can appear extraordinarily capable. Suppose an AI writes a program. A parser can reject its syntax, a compiler can reject its types, a linter can detect issues with the code, tests can fail, a browser can parse the resulting page, inspect the DOM and then take a screenshot. All this output can be fed back to the AI, allowing it to modify the code and try again. The important thing about all these mechanisms is that they do not care whether the AI believes it succeeded. Reality check pushes back. The AI says: “This should work.” The compiler says: “No, it doesn’t.” Then AI changes the code. This loop repeats until the program behaves as intended.
This is an extraordinarily powerful environment for AI because probabilistic generation is surrounded by mechanisms that can independently challenge the generated answer. This is akin to putting guardrails to ensure the response does not go sideways.
Passing a compiler check or a test suite doesn’t make the software perfect obviously, but the feedback loop is unusually strong. AI generates, the environment checks, AI adjusts, the environment checks again.
The stronger that loop is, the more confidently AI can be allowed to act.
Not Every Problem Domain Has a Compiler
Now consider a medical question. An AI recommends a particular treatment. Where is the equivalent of the compiler? Where is the button that tells us immediately PASS or FAIL?
Medicine certainly has verification mechanisms. Doctors have diagnostic criteria, laboratory tests, imaging, clinical guidelines, drug databases, specialist opinions and eventually patient outcomes. But these mechanisms are not as practical as compiling software. They are expensive, slow, may show uncertain results, and most importantly, some actions cannot simply be undone.
While a broken CSS rule can be reverted back instantly, harm caused by a medication may never be undone. The same problem appears in different forms in law, economics, strategy, psychology, geopolitics and many other fields.
There is no “npm test” command for a business strategy. There is no simulator that immediately tells us whether a geopolitical forecast is correct. There may be excellent evidence and highly qualified people capable of reviewing the conclusion, but the feedback loop is fundamentally different or weak in some domains.
This helps explain something that otherwise looks puzzling. The same AI can appear almost superhuman at one task and dangerously unreliable at another. The model did not necessarily become less intelligent. The verification environment is different.
Two Simple Ways To Verify AI’s Answer
It is useful to separate verification into two broad categories. The first is mechanical verification. This includes things such as:
- compilers
- interpreters
- test suites
- parsers
- calculators
- browsers
- controlled experiments
These mechanisms can often tell us relatively quickly whether something worked.
The second is factual or evidence-based verification. This includes:
- official records
- trusted databases
- research papers
- measurements
- observations
- historical records
A historical claim may have almost no mechanical way of being tested, but it can potentially be checked against strong documentary evidence. Code may have extremely strong mechanical verification even when factual verification is irrelevant. Some tasks have both, and some have neither. This can determine how much confidence we should place in AI’s output.

X -axis: Mechanical / deterministic verifiability
Y-axis: External factual / evidentiary verifiability
The exact position of each task on such a chart is obviously subjective. The important point is not whether code generation deserves a score of eight or nine. The point is that different tasks inhabit fundamentally different verification environments. That helps determine how we can use AI more safely.
What Role Should AI Play?
Once we accept that different tasks have different verification environments, another conclusion follows naturally: AI should not have the same job everywhere.
Its role can move along a spectrum: Explore → Recommend → Decide → Execute
At one end, AI may simply look for possibilities. At the other, it may be allowed to act as an autonomous agent. The appropriate position depends on the type of the task.
Consider scientific research where the situation is almost the opposite. There may be enormous value in asking AI to examine vast quantities of data, detect unusual relationships, propose explanations and generate hypotheses that researchers have not considered. This may be one of the areas where AI’s output is very valuable, but that does not mean the AI should be allowed to declare “Therefore this hypothesis is true”. Its better role may be “Here is something interesting that you should investigate”. The experiment should follow AI’s output, and establish the fact. The AI helps decide where to look.
This distinction is particularly obvious in cybersecurity. Imagine an AI examining 200,000 HTTP requests from a web application assessment. A human tester may never manually examine every relationship between every request. An AI might notice:
“These endpoints appear to follow an unusual authorization pattern; there may be an IDOR vulnerability here.”
That filtering from 200k records is very useful, but it still only generates a hypothesis. The request should then be resubmitted using the same low-privileged session but with another user’s object reference. If the server returns data that should not be accessible, it establishes evidence. The AI generated the hypothesis. The test established the fact.
This is a powerful way of thinking about AI:
Use AI to discover what might be true.
Use verification to establish what is true.
There is also an important paradox here. The more uncertain and complex the problem, the more valuable AI’s ability to explore may become. Yet, that is exactly the situation where we should not give it the final say.
At the other end of the spectrum, some tasks are intellectually boring but very safe to automate.
Consider a routine data transformation. The desired format is known, so is the input. The output can be validated automatically and errors are easy to detect. Also the changes made by AI can be reversed easily. Such conditions would leave very little reason for a human to examine every individual decision. Therefore, this task is very suitable for AI to execute.
And for some tasks AI is unnecessary altogether. If deterministic software running on fixed rules already solves the problem perfectly, then use that instead. There is little value in asking an LLM to probabilistically rediscover something that ordinary software can enforce reliably every time.

X -axis: Value of AI exploration / hypothesis generation
Y-axis: Safe level of AI autonomy
The interesting part of this chart is that the two axes do not necessarily move together. Scientific discovery can sit very high on exploratory value while remaining relatively low on safe autonomy. Routine data transformation can sit in almost the opposite position. This tells us that AI’s most valuable role depends on the task. Strong feedback loops can support AI as an autonomous agent, while uncertain and difficult-to-verify domains may benefit most from AI as an explorer: surfacing patterns, hypotheses and possibilities for humans to investigate.
Five Practical Questions To Assess If You Can Trust AI’s Answers
It is easy to get carried away when AI provides confident and seemingly very insightful answers. But blindly acting upon them can get us in trouble. So how can an ordinary AI user decide when to trust and act on AI’s answers? I think five questions cover a surprisingly large part of the problem.
1. Can AI’s output be independently verified?
Can another system, source, expert, experiment or observation challenge the answer? If not, our confidence in AI’s answer should immediately fall.
2. How quickly will I discover that it is wrong?
A compiler may tell you within milliseconds. A bad business strategy may take two years to reveal itself.
The longer it takes to discover AI’s wrong answer, the less you should act on it.
3. Can the action be reversed?
A malformed document can be regenerated with ease. A database migration may be quickly recovered from a backup. A financial transaction may be harder to reverse. A medical intervention may be impossible to undo.
The less reversible the outcome, the stronger the required oversight should be.
4. How big is the impact of a mistake?
Not all errors matter equally. If AI selects the wrong font size, impact is negligible. If AI gives incorrect medical, legal, financial or safety advice, the consequences may be serious even when the probability of error is very small.
The cost of being wrong matters as much as the likelihood of being wrong.
5. Is AI exploring possibilities or making the final decision?
This may be the most important question of all.
An AI proposing ten possible explanations is very different from an AI choosing one explanation and taking irreversible action based on it. In uncertain environments, moving AI one step back in the decision chain can preserve most of its value while eliminating much of the risk.
Let it explore, search, challenge assumptions, find patterns and generate hypotheses. Then you decide what level of evidence is required before the hypothesis becomes a conclusion.
Reality Check Gets The Final Say
AI is not an oracle. But describing it as a glorified autocomplete or a machine that merely “guesses a probable answer” is not very helpful either.
Modern AI’s output can at least resemble reasoning to some degree, it can make certain generalizations, write software, analyze complex information, identify hard to detect patterns and generate genuinely useful ideas. Its limitations are more subtle.
An AI-generated answer does not carry its own proof of correctness. The fluency or confidence in its answer does not establish truth. Its reasoning does not automatically provide verification. That is why the right question is not simply “Can AI do this?” Going forward, the answer to that question will increasingly be yes. So the more useful questions are:
What role should AI have for this kind of task?
and:
What mechanism will catch it when it is wrong?
Where strong feedback exists, AI can move increasingly close to autonomous execution. Where verification is difficult and consequences are serious, its greatest value may lie in helping us see possibilities that we would otherwise miss. Perhaps the simplest rule is this:
Trust AI in proportion to the strength of the feedback loop available to verify its results.
AI can generate probable answers. Reality check still gets the final say.