A Story That Has Become Routine
A CIO of a mid-sized engineering services firm called me earlier this year. They had just shut down their second enterprise AI initiative in eighteen months. Both had been technically functional. Both had been adopted with apparent enthusiasm in the first quarter. Both had quietly stalled. He wanted to understand what was happening before authorising a third attempt.
I asked him what the two post-mortems had concluded. He said, with some embarrassment, "They both said the same thing. 'Change management issues.'"
"Change management issues" is the diagnosis enterprise organizations give themselves when an AI initiative dies and they do not know why. It is technically true and operationally useless. It does not tell you which change, in which direction, to manage.
The honest answer — supported now by an embarrassingly large body of research — is that the failure rarely lives in the model, the integration, or the data pipeline. It lives in three specific human failure modes that the technology actively accelerates and that no dashboard in the buyer's stack is designed to detect.
The Empirical Floor
MIT's State of AI in Business 2025 reported that approximately 95% of enterprise AI pilots fail to scale beyond their initial deployment. Fortune 500 companies are estimated to lose roughly $31.5 billion annually to what the same body of research calls organizational amnesia — the cost of work that was searched for, reconstructed, or duplicated because the AI sessions that produced the original work were stateless and the reasoning was never captured.
Different studies report different specific numbers — BCG and McKinsey have cited 70%-80% failure ranges on broader digital transformation; sector-specific surveys vary. The exact number is less interesting than the consistency of the underlying finding: the modal outcome of an enterprise AI initiative is to underdeliver against the business case used to fund it.
What is striking, when you read the post-mortems, is how rarely the model itself is identified as the cause. The model worked. The integration shipped. The team learned how to prompt. And then the initiative quietly stopped delivering value.
Three failure modes, in my consulting practice, account for the majority of these slow stalls. None of them is technical. All of them are predictable. None of them appears on the executive dashboard until the initiative has already failed.
Failure Mode 1: Organizational Amnesia
The first failure mode is the one that the MIT figure points at most directly. Most enterprise AI deployments are session-based: a worker opens a chat, asks a question, gets an answer, and closes the chat. The reasoning that produced the answer is discarded. The next worker, asking a related question the next week, starts from zero. The organization accumulates output — drafts, summaries, analyses, recommendations — but does not accumulate the reasoning that produced any of it.
Over twelve to eighteen months, this produces a strange organizational state. The volume of AI-assisted work rises sharply. The institutional capacity to learn from that work does not rise with it. Decisions get made and then are not retrievable; rationale is generated and then evaporates; lessons that would have been written into a colleague's head in the pre-AI era are now produced inside a chat session that no one returns to.
The cost is invisible in the short run because the immediate outputs look fine. It becomes visible in the medium run when the organization realizes it has produced eighteen months of work whose reasoning it cannot reconstruct, whose authors cannot remember how they arrived at the conclusions, and whose underlying judgments cannot be audited, challenged, or built upon.
The fix is not a technical one. Stateless tools can be replaced with stateful ones; that is the easy part. The harder part is cultural: the organization has to value the capture of reasoning as much as the production of output, and the reward systems have to reflect that. Most do not. The worker who writes down why they accepted an AI's suggestion is producing a more valuable record than the worker who simply accepts it — and is, in most organizations, paid the same.
Failure Mode 2: The Erosion of Verification Capacity
The second failure mode is the one I have written about most extensively in our coaching practice — the slow degradation of the senior expert's ability to catch the AI when it is subtly wrong. (I covered this in detail in our piece on the engineer's EQ gap, so I will keep the summary short here.)
The mechanism is well-attested: as workers delegate judgment-heavy tasks to AI, their independent capacity to perform those tasks atrophies. The endoscopist study is the cleanest illustration — adenoma detection rates of 28% with the AI, 22% without it after routine use, below the pre-AI baseline. The Anthropic developer trial found a 17% comprehension gap; the PNAS math study found students who had used ChatGPT during practice performed 17% worse than peers once the tool was removed. The pattern is now robust enough that organizations should treat it as a baseline expectation.
The failure mode this produces is one of the most expensive an enterprise AI initiative can encounter: the senior expert at the top of the chain — whose job is to catch errors that propagate downstream — silently loses the capacity to catch them. The initiative does not fail through a single visible incident. It fails through a slow rise in the rate of plausible-but-wrong outputs that nobody upstream is calibrated to detect, until a recall, an audit, or a customer-discovered defect makes the erosion visible. By then the verification capacity required to recover has thinned across the senior cadre.
This is the single highest-leverage point for a coaching intervention in an AI-augmented organization, and almost no organization is currently making the investment.
Failure Mode 3: The Quiet Atrophy of Team Fabric
The third failure mode is the one that organizations most consistently underweight, because it is the slowest to manifest and the easiest to attribute to other causes.
AI-augmented work reduces the need for synchronous human coordination. Some of this is welcome — fewer status meetings, less email, fewer "quick syncs" that consume an afternoon. But the meetings that get eliminated are not only the wasteful ones. They are also the meetings in which colleagues debated trade-offs, in which juniors absorbed how seniors thought, in which the small social work of being a team was done.
The Harvard-Microsoft study of more than 180,000 developers found that individual coding activity rose 12.4% under Copilot adoption while peer collaboration fell nearly 80%. The OECD's 2025 study of approximately 6,800 workers across seven countries found that algorithmic management was associated with reduced autonomy, lower trust, and higher stress — except in workplaces where workers had substantial voice in how the systems were configured.
The downstream consequence in our client data is consistent. Teams under heavy AI adoption become measurably more productive and less cohesive over twelve to twenty-four months. The productivity gains are real and they accrue to the organization. The cohesion losses are also real and they accrue too — they show up later, as gradual disengagement, slow attrition of the workers whose professional identity included a social-collaborative dimension, and a brittleness under stress that the organization mistakes for "we need more process."
The team is more efficient. The team is also slowly losing the social fabric that lets it recover when something non-routine happens. In most organizations, that fabric was load-bearing in ways that nobody had explicitly named — until it was gone.
What These Three Failure Modes Have in Common
The pattern across all three is the same. The technology delivers what it promised. The organization captures a real, measurable productivity gain. And then, on a timeline the buyer did not budget for, a slow human cost emerges that consumes the productivity gain and continues compounding.
The cost is structural. It is produced by the deployment, not in spite of it. Better models will not fix it. Better integrations will not fix it. The fix is in the human system the technology sits inside.
This is what "67% fail on people, not tech" actually means. It does not mean that the engineers were lazy or that the change-management plan was thin. It means that AI deployment is — empirically, predictably — not behaviourally neutral. The same platform, deployed into two organizations with different cultural conditions, produces different outcomes. The buyer who treats the deployment as a technology project gets the modal outcome. The buyer who treats it as an organizational behaviour project gets a different outcome, and a more durable one.
What a Properly Framed Deployment Actually Requires
The interventions are not exotic. They are, however, ones that almost no AI vendor will surface in their sales process, because they require the buyer to do work the vendor cannot bill for.
- Diagnose verification capacity in the senior expert layer before deployment, and track it during it. The erosion is invisible to the senior themselves — that is the defining property of the failure mode. It has to be measured externally.
- Build reasoning capture into the workflow, not as a compliance checkbox but as the substantive work the role now does. The worker's job is no longer producing output; it is producing justified output. The reward system has to follow.
- Invest deliberately in the relational rituals that the platform makes nominally unnecessary — design reviews, mentorship pairings, the small social work of being a team. The platform's efficiency must be partially spent on the cohesion it dilutes.
- Involve workers in configuring the system. The OECD finding is unambiguous on this. Worker voice is not a soft variable; it is the variable that decides whether the negative effects appear at all.
These are EQ-shaped interventions. They are coachable, diagnosable, and — in the organizations that take them seriously — they are the difference between an AI initiative that scales and one that joins the 95%.
Where to Start
The diagnostic comes before the prescription. Before you fund the next pilot, you want to know what shape your organization is currently in along the three dimensions above — because the same deployment will produce different outcomes depending on what it lands on.
Our EQ Frontier assessment is built specifically for this question. It measures the emotional and cognitive dimensions of AI readiness at the leader and team level — including the early indicators of verification erosion, attribution drift, and relational atrophy that the literature predicts and that conventional readiness audits miss.
The CIO I spoke with at the start of this piece had not been wrong to fund either of his first two initiatives. The technology had been sound. The conditions the technology landed in had not been measured. We started, this time, with the diagnostic. We will not know for another six months whether the third initiative scales — but for the first time, the organization can see what it is actually managing.
Diagnose what you are actually deploying into. Start with the EQ Frontier assessment, or read our companion piece on the engineer's EQ gap. For the full synthesis across all three organizational layers, see our pillar on the AI Anxiety Stack.