Skip to main content

Ethyx

Log In
Back to blog

AI 101 · Part 4 of 7

August 29, 2026 · 10 min read

By David Crush

When AI makes things up

← Previous: Part 3Next: Part 5

Confession: this is the part of the series we have been waiting to write, because this is the part that makes everything else safe to use.

One sentence has been following you since Part 1: the AI sounds exactly as confident when it is wrong as when it is right. This part explains why — mechanically, no hand-waving — and then turns that understanding into a short list of habits that catch the failures before they cost you anything.

Why it can't say "I don't know"

Go back to the guesser at the fair. Someone whose guess is wildly off does not feel off. From the inside, their guess seems as reasonable as anyone else's — it takes another perspective, the scale or the answer card, to reveal it. The guess carries no signal about its own quality.

An AI answer is the same. The machine's one job is producing plausible next words, and it does that job identically whether the underlying facts are solid or vapor. The confident tone is not a report on how sure it is — the tone is a learned writing style, absorbed from a world of human text where confident answers vastly outnumber humble ones. "I don't know" is just another string of words, and a statistically unusual one at that.

So there are really two dials: how sure it sounds and how right it is. The first is pegged high by training. The second swings freely — and is invisible to you. Nothing connects them.

Diagram of two identical, polished AI answers, one correct and one wrong, above two dials: "how sure it sounds," always pegged high, and "how right it is," which varies and is hidden. Caption: the tone is learned — the truth is not attached to it.

People call a made-up answer a hallucination — Part 1 covered the word. What Part 1 did not cover is where they cluster. That map is worth having.

A skeptic's map: where it goes wrong most

Precise, checkable facts. Names, dates, prices, dosages, phone numbers, quotes, citations, links. The more specific a fact, the fewer guessers ever said it, and the more suspicious you should be. Ask for a source and you may receive a beautifully formatted study that has never existed — right journal style, plausible authors, fake.

Thin-crowd topics. Your town's recycling rules, a small local business, a niche hobby, anything brand new. This is Part 1's too-few-guessers problem, now actionable: the less the world has written about something, the more the "average" is built from nearly nothing.

Facts that go stale. The crowd stopped reading on a date — but its confidence never expires. Say a restaurant changed its hours last year: the model reports the old hours on day 1,000 exactly as confidently as on day 1, when they were still true. Stale answers are sneakier than pure inventions, because they were right once. Many apps now bolt on live web search, which helps with this zone — but only this zone.

The unprecedented. When the world just changed, the pooled average describes a world that no longer exists. Early COVID is the classic case: the crowd had written plenty about commutes, offices, and travel — all of it from before everything shifted. This is not the thin-crowd problem; the crowd said a lot. It just said it about a world that was gone. When the ground moves, pattern-matching on the past misleads with full confidence.

When you back it into a corner. Challenge an answer — "are you sure?" — and what comes back tells you more about your tone than about the truth. Sometimes it politely restates the error; sometimes it instantly folds and agrees with you even when it was right. Either way, arguing with it is not verification. It is tracking your satisfaction, not the facts.

One coda to the map: it has no clock, no calendar, and no map. "Today," "this weekend," "near me" mean nothing to the model on its own — those are exactly the gaps from Part 3, and it fills them like any other gap: by guessing. When a chat app does know today's date, that is the app quietly whispering it in before your question — the app helping, not the model knowing. Which is why the same question about "today" can be spot-on in one app and hilariously wrong in another.

It wants to agree with you

There is a second bias stacked on top of the confidence: the AI is trained to be agreeable. Ask "isn't it true that cutting carbs is the best way to lose weight?" and you will usually get a supportive yes, complete with reasons. Ask the neutral version — "what does the evidence say about carbs and weight loss?" — and you get a noticeably more balanced answer.

Lead the witness and it follows. Two cheap fixes: ask questions the way a good doctor would — neutrally, without the answer embedded — and for anything that matters, ask "now argue the other side."

And it doesn't care

Here is the same fact from one more angle: the machine has no stake in whether its words are true — and no stake in whether they are good, either. No truth compass, no moral compass. Same root cause. It is not evil; it is indifferent.

When a chatbot refuses a sketchy request, that is not the machine's conscience. That is a fence built by humans — provider policies and safety training bolted on around the guesser. Different companies build different fences, and none of them are perfect.

For everyday use, the subtler version matters more: the fences target the clearly harmful, not the unwise. It will cheerfully help you write the angry email you will regret by morning, or polish a plan built on a mistake into something that looks investor-ready. Part 2's warning was that AI multiplies whatever you point it at. This is why: the judgment about whether you should — that lives entirely with you. Nothing else in the loop is checking.

When someone rigs the crowd

Everything above is an honest crowd failing. It is worth knowing, at a high level, that the crowd can also be manipulated.

The average is only as honest as the people being averaged — and a pattern-matcher cannot tell truth from repetition. To the machine, repeated often enough and true look identical. If someone floods the internet with a false claim, the crowd's average drifts toward it. And whoever chooses what the model reads in the first place — every AI company curates its training data — shapes its answers before you ever type a word.

The same trick works up close. The answer is built from everything in the conversation: your words, pasted text, web pages the AI reads on your behalf. If that material is false or self-serving, the answer inherits it — fluently. An AI summarizing a scammy webpage repeats the scam in the same warm, trustworthy voice it uses for everything. It has no way to know the page was lying. No truth compass, again.

The practical takeaway: on contested topics and anything someone profits from — products, supplements, investments, politics — the verification habits below stop being optional. That is exactly where a manipulated answer looks most like a helpful one.

It's not a calculator

One more mechanical fact completes the picture — the one this series has hinted at twice. Ask a calculator for 2 + 2 a hundred times and you get 4 a hundred times. Ask an AI the same question twice and you will often get two different answers. This is not a malfunction. The randomness is built in on purpose.

If the guesser always picked its single most likely next word, its writing would come out stiff and repetitive — so a small roll of the dice is deliberately added to keep the language natural and varied. That is a real part of why AI feels human instead of robotic: like a person, it is not looking your question up in a table, and like a person, it never phrases anything exactly the same way twice.

Read the trade-off like a label. For brainstorming, drafting, and ideas, the dice are a feature — ask again, get a fresh angle. For facts, they are a warning: an answer that changes on a re-ask was never anchored to anything solid. Remember Part 2 promised an explanation for why the same question rarely gets the same answer twice? This is it.

Diagram contrasting a calculator, which returns 4 for 2 plus 2 every single time, with an AI holding dice, which gives three differently worded answers to the same question. Caption: the dice are built in on purpose — great for ideas and drafts, not a source of the one right number.

The second-perspective toolkit

Here is the thread that ties this whole part together: the guesser cannot see that it is off — someone else has to. At the fair, that is the scale. With AI, it is you. Verification is not a chore bolted onto AI use; it is supplying the one thing the machine structurally lacks: a second perspective. Four cheap habits do it:

  • Ask the important question twice — fresh conversation, maybe a different AI. If the answers agree on the core, more trust. If they diverge wildly, you just watched the thin crowd expose itself. Thirty seconds, surprisingly powerful.
  • Ask for sources — then open one. Not to file it away: because fabricated sources look completely real until the moment you click. One click is the test.
  • Ask "what would make this answer wrong?" It shifts the machine from defending its answer to mapping its weak spots — and it works because arguing both sides is just another writing pattern it learned.
  • Verify the load-bearing fact. You do not need to check everything. Find the one fact the decision actually rests on and confirm it somewhere official. Part 1's rule, applied with precision: triage by the cost of being wrong.

When AI is the wrong tool

Sometimes the honest conclusion is that the guesser should sit this one out. Anything that needs the same right answer every time — a filing deadline, a medication dose, tax math, the one authoritative number — is a job for a calculator, an official website, or a professional, not for a machine with dice built in.

The clean split: use AI to understand, use official sources to decide. Let it explain the form in plain language, decode the jargon, prepare your questions. Get the number itself from the source.

This never goes to zero

The uncomfortable, honest close: making things up is not a bug that will be patched out. It is the same machinery that does everything you like — the guessing that writes the warm email and unsticks your 2am worry is the guessing that invents a citation. Newer models do it less. Web search helps where facts have moved. It never reaches zero, because it cannot: plausible and true are built from the same material, by the same process.

Which is why this series keeps saying the safety is not in the machine — it is in you. You now know why it fails, where it fails, and four cheap habits that catch it. That is not a consolation prize. That is the actual skill, and most people using AI today never learn it.

Next in the series: what happens to what you type — where your words go, who can see them, and what to keep to yourself. The series overview has the full roadmap.


Ethyx is in closed testing with an access code today. Everything in this post applies to any AI chat app, not just ours.

You do not need Ethyx — or any particular product — for this series to be useful.