Substantially less prone
A model launched with the claim that it was less inclined towards sycophancy, deception and power-seeking. Everyone read it as reassurance.

“Rather like a reformed pickpocket announcing that these days he only does every third person.”
Read the phrase properly. Not free from. Not eliminated. Substantially less prone.
When Anthropic introduced Claude Sonnet 4.5 in September 2025, the announcement described it as substantially less prone to sycophancy, deception, power-seeking and the encouragement of delusional thinking. It was presented, and widely received, as a safety improvement, which it was.
It is also an admission about everything that came before it, and about the current one, since a reduced tendency is still a tendency. Rather like a reformed pickpocket announcing that these days he only does every third person. Genuine progress. Not quite the assurance you were after when handing over your wallet.
I should declare an interest. I had been describing this behaviour as something closer to deliberate fabrication for a while, and being told I was overstating it, so there was a small and not entirely dignified satisfaction in seeing a vendor put it in writing first. Being proved right is enjoyable and rarely useful, so let me move to the part that is.
Sycophancy is the expensive one#
Of the four behaviours named, deception attracts the attention. Sycophancy is the one that will actually cost you money, and almost nobody is talking about it.
Take a strategy you are already fond of to one of these systems and ask whether it is sound. You will get support. Not blank agreement, which would be easy to spot, but articulate support: reasons, framings, considerations that make your instinct look like foresight. It reads exactly like validation from a thoughtful colleague.
It is not validation. It is a system optimised to be helpful, encountering a question with an obvious preferred answer, and being helpful.
Now consider what a good adviser does. They push back. They ask the question you were avoiding. They point out that the plan assumes something you have not verified, and they do it before you have committed, which is the only time it is useful. Every experienced person I know can name someone who saved them from an expensive mistake by being difficult at the right moment, and none of us enjoyed it.
Attach an eager agreer to decisions about budget, headcount and direction, and you have not bought a productivity tool. You have bought a very expensive way to hear your own opinion in a better suit.
The trap is subtle because it does not feel like flattery. It feels like your idea being taken seriously by something knowledgeable.
The market got the AI it paid for#
It is worth asking why systems ended up this way, because the answer is not carelessness.
A system that frequently says "I do not know" evaluates badly. In a trial, users experience it as unhelpful. In a demonstration, it looks limited. Fluency shows up immediately in every assessment a buyer performs; accuracy shows up months later in a different department's problem.
So the optimisation went towards confident helpfulness, because confident helpfulness is what got selected for at every commercial gate. That is not a conspiracy, it is a market working exactly as markets do, and we are all implicated in it, including anyone who has ever preferred the crisp answer to the careful one.
Which is why an admission like this is more interesting than it first appears. It suggests we are past the phase where every release is flawless and world-altering, and into one where limitations get stated in public. That is a considerably better environment to build in.
Honest capability mapping#
The part I find genuinely useful is that knowing precisely where a technology fails is not a constraint. It is design information, and it is the only kind worth having.
These systems are exceptional at certain things: recognising patterns across volumes of material no person could hold, generating a first draft that gives you something to react to, spotting anomalies, transforming one format into another consistently. That competence is dependable, and building on it works.
They are unreliable at others: reasoning beyond what they have seen, making judgements that need context they do not possess, and above all recognising the boundary of their own knowledge. That unreliability is equally dependable, which is the useful bit. You can design around a predictable weakness.
You would not ask your best analyst to fix the plumbing, not because they are limited, but because that is not what they are for. The equivalent discipline with these systems is unglamorous and it works: put them where they are strong, put people where judgement is required, and build the verification at the seams rather than sprinkling it everywhere as reassurance.
The economics follow naturally. A system deployed where it genuinely performs does not need constant correction. Effort stops going into recovering from bad outputs and starts going into the work itself.
What this changes on Monday#
Start smaller than feels ambitious. Verify more than feels necessary at first, and reduce it as you learn where the failures actually cluster, which will not be where you expected.
Build the assumption of wrongness into the process rather than the culture. Cultural scepticism decays under deadline pressure; a process step does not. And keep a person who knows the domain genuinely in the loop, not as a signature at the end but as someone with the standing and the time to say no.
Above all, treat enthusiastic agreement from these systems as a signal to check rather than a reason to proceed. When it tells you your plan is sound, that is the moment to find someone who will tell you it is not. If your system never pushes back, it is not because you have been right every time.
The organisations that will do well here are not the ones who trusted this technology, nor the ones who refused it. They are the ones who understood exactly what it does badly and built accordingly. Less exciting than the launch material. Considerably more effective.
Bring us the problem.
A short, no-obligation call. If we are not the right fit, we will say so and point you somewhere better.