The analyst who will not show their working
You would not sign anything they gave you. Then the same behaviour arrives with an enterprise licence and you sign it weekly.

“It is a curious position for an industry that audits everything else.”
Picture an analyst who produces excellent-looking reports, refuses to name their sources, will not explain how they reached a conclusion, and becomes faintly offended when questioned. Nobody would tolerate that for a fortnight. You would demand the working before you signed anything, not because you thought they were dishonest, but because that is simply what due diligence is.
Now consider the platform generating your sales forecasts, your customer insights and your operational recommendations. Ask it how certain it is. Ask which sources fed a particular conclusion. Ask what happens when its own output becomes the input to the next process.
Largely, you cannot. And unlike the analyst, it never gets awkward about it, which turns out to be the problem.
The enterprise version is the awkward one#
Everyone accepts a small tool might be opaque. The interesting case is the large platforms that sit inside processes moving serious money.
These systems produce pipeline predictions, customer scoring, operational recommendations and strategic analyses. They are sold on sophistication and bought on brand, and they will give you extremely detailed analytics about user engagement, feature adoption and system performance. What they generally will not give you is a calibrated statement of how much to trust any particular output.
Can you audit the confidence behind next quarter's pipeline prediction? Does the assistant tell you when its strategic suggestion is a stretch? Can the operational system explain why you should believe its insight rather than merely that it has one?
The honest answer is mostly no, and we have collectively decided that brand reputation substitutes for verifiable accuracy. It is a curious position for an industry that audits everything else.
What makes the enterprise case sharper than the standalone tool is reach. A fabricated customer insight inside a small application affects one decision. The same fabrication inside a platform embedded across sales, marketing and product development propagates to dozens of decisions before anyone notices the original was never grounded in anything.
Opacity is not one problem#
It helps to separate the failure modes, because they need different answers.
Source uncertainty is the obvious one: you cannot tell whether a conclusion rests on your actual data, on general training material, or on something invented to fill a gap. Reasoning opacity is subtler and often worse, because the sources can be perfectly sound while the inference connecting them is nonsense, and a confident summary conceals that completely.
Then there is what I would call flatness. Everything arrives in the same tone. A conclusion supported by extensive evidence looks identical to one assembled from fragments, and no signal in the output distinguishes them. You are being asked to apply uniform scrutiny to material of wildly varying reliability, which in practice means applying uniform scrutiny to none of it.
And propagation, which ties them together. Once an output becomes an input, its provenance is gone. The second system has no way of knowing the first was improvising, and no standing to disagree if it did.
Three questions for the supplier meeting#
You are unlikely to change how these platforms are built. You can change what you ask before you commit to one, and three questions do most of the work.
First: can your system produce confidence or uncertainty indicators, and how are they calibrated? The calibration half is where the interesting conversation lives. Any system can emit a number. A calibrated one has been tested so that seventy per cent confidence corresponds to being right roughly seventy per cent of the time. If nobody can describe how that was validated, the number is decoration.
Second: what visibility do we get into the reasoning and the source material behind an output? Listen for whether you get actual traceability or a summary of reasoning generated after the fact, which is a different thing entirely and considerably less useful.
Third: how do you handle it when your outputs become inputs to subsequent processes? This is the question fewest suppliers have thought about, and the answer tells you whether anyone in the building has considered compounding error at all.
The specific answers matter less than the shape of the response. Someone who has built for this will be pleased and will want to show you the mechanics. Someone who has not will reassure you at length about their quality processes, and that reassurance is your answer.
Scaling scrutiny rather than spreading it#
Once you have confidence indicators that mean something, the operational design follows almost automatically.
High confidence proceeds with sampling. Medium confidence gets a second opinion, ideally from a different method rather than the same method run twice, since running the same flawed approach twice produces agreement rather than verification. Low confidence stops and waits for a person who knows the domain.
The benefit is not that fewer errors occur. It is that expensive human attention lands where the uncertainty actually is, instead of being spread thinly across everything as a gesture, which is how most review processes work and why most of them catch so little.
The capability you have to keep#
The part that worries me most is not technical.
Organisations that let these systems fully replace their internal expertise lose the ability to notice when the systems are wrong. There is no external alarm for this. The only detection mechanism is a person who knows the domain well enough to feel that something is off, and that feeling is built by doing the work, which is exactly the work being automated.
So there is a tension here I do not think anyone has properly resolved, and I am suspicious of anyone claiming to have. You automate to remove the tedious middle. The tedious middle is where judgement came from. I do not have a clean answer, and I would rather say so than offer a tidy one.
What I would say is that maintaining the capability to evaluate has to be a deliberate decision with resources attached, because nothing about the default path preserves it. Train people in how these systems fail, not only how to use them. Reward the person who catches the error rather than treating them as an obstacle to throughput.
Uncertainty is not a defect to be engineered away. It is a property to be measured, surfaced and acted on, and the systems worth trusting are the ones honest enough to tell you when they are guessing.
Bring us the problem.
A short, no-obligation call. If we are not the right fit, we will say so and point you somewhere better.