INTEEVO
Beyond the Pilot28 July 2026 · André Jacyshyn

Thirty spelling mistakes and you bin it, thirty wrong facts and you sign it

We judge documents on the errors we can see. Automated systems make only the other kind.

two apples, the perfect one with a wormhole.
“The error did not degrade as it travelled. It got polished.”

Imagine a proposal lands on your desk. The argument is elegant, the structure clean, the recommendations sensible. Scattered through it are thirty spelling mistakes, about one word in twenty.

You bin it. Not after careful deliberation, immediately, and you think slightly less of whoever sent it. You never get a second chance to make a first impression, and this one told you everything: if they could not be bothered to check the words, why would they have checked the reasoning?

Now the same document, spelled immaculately, with thirty statements of fact that are simply wrong. Not absurd, just wrong. A figure transposed, a date shifted, a conclusion that does not follow from its evidence. It sails through. It gets circulated, quoted, actioned.

That asymmetry is the entire problem with automated decision-making, and we walked into it with our eyes open.

We built error detectors for the wrong kind of error#

Human writing gave us reliable tells. Someone who is guessing hedges, or their prose goes vague at the exact moment it should be precise, or they contradict themselves two pages later. Sloppiness in one dimension predicted sloppiness in another, and we learned to read that correlation without knowing we were doing it.

Machine output breaks the correlation. The spelling is perfect whether or not the content is true. The tone is identical whether the system is reciting something well established or improvising over a gap. Every surface signal we spent decades learning to read has been decoupled from the thing it used to indicate.

So we are left with the only reliable check, which is knowing the subject well enough to spot when the answer is wrong. That works beautifully as long as you already knew the answer, which rather defeats the purpose of asking.

Where it stops being an annoyance#

One wrong fact in a document is a nuisance. Someone eventually catches it, there is a slightly awkward email, life continues.

Chain the systems together and the arithmetic changes. Agent one extracts, agent two calculates, agent three plans, agent four writes the summary that reaches a human. At no point does anything downstream interrogate what arrived from upstream, because interrogating it was never anyone's job. Agent two has no way of knowing that agent one was improvising, and even if it did, it has no basis to overrule it.

By the fourth step you have an elaborate, internally consistent, beautifully presented answer that is wrong all the way down, and the presentation quality has gone up at every stage. The error did not degrade as it travelled. It got polished.

There is a particular version of this that deserves its own name, and I think of it as the consensus illusion. Run several systems over the same flawed source, ask them to vote, and they agree enthusiastically. You read that agreement as corroboration. It is nothing of the sort. It is four people who all read the same wrong newspaper.

Confidence that means something#

The fix is not more accuracy, which is a race nobody wins. It is calibration, and the distinction matters.

An accurate system is right a lot. A calibrated system knows how often it is right. When a calibrated system says it is 70% sure, it turns out to be correct roughly seven times in ten, which means the number is doing work rather than decorating the output. Raw confidence scores from a language model, taken at face value, are closer to a Magic 8-Ball with a maths degree. They can be brought into alignment, but not by accident and not by default.

Once you have numbers that mean something, you can actually build with them. High confidence proceeds with spot checks. Middling confidence gets a second opinion from a different system or a different method. Low confidence stops and waits for a person. Nothing sophisticated about it, it is triage, the same logic any hospital applies at the door.

What that buys you is not a system that never errs. It is a system that fails in the right direction, loudly and early, at the point where a human is cheapest to involve. Expensive expertise gets spent where the uncertainty actually is, rather than spread evenly across everything as a comfort blanket.

Layered humility#

I keep coming back to a phrase for what a well-built agentic system needs, which is layered humility. Not one check at the end, which catches almost nothing and arrives too late to be useful, but doubt built in at several points: at generation, at handover between steps, against external ground truth, and at the governance level where somebody senior has actually decided what happens when the machine says it is unsure.

The vote pattern deserves particular attention, because it is the cheapest signal you will ever get and most implementations throw it away. When your systems agree instantly, that tells you little. When they scatter, something interesting has happened, and that scatter is worth escalating even when the majority answer looks fine. Disagreement is information. Most architectures discard it in favour of a tidy single answer, which is a bit like a jury foreman announcing the verdict and burning the deliberation notes.

Sceptics make better architects#

None of this is anti-automation. I have spent long enough building these systems to know what they save, and I am not about to argue for doing it all by hand out of nostalgia.

But there is a difference between a system that is right most of the time and a system you can trust, and the gap between them is entirely made of knowing when it is wrong. The organisations that get this right are not the ones with the best models. They are the ones that treated scepticism as an engineering requirement rather than a personality trait, and built for the bad day rather than the demo.

A speedometer that reads consistently high is worse than no speedometer at all, because you will act on it. That is the thing worth designing against.

Pass it on
LinkedIn X Email

Bring us the problem.

A short, no-obligation call. If we are not the right fit, we will say so and point you somewhere better.