The hidden liability of generative AI: assisted, or augmented

Published on

Brian PLUS 2026-08-10 inspearit
Contents

What are people still worth if they don't think much

I support 3 Product Managers on a transformation programme. With one of them, it genuinely works. Augmented workshops, and an AI he knocks the ball around with on his features.

And I caught myself wondering, bluntly, what a PM still brings if they don't think much and aren't particularly creative. The question didn't stop there. It extended to developers. Then to me.

I have no reassuring answer to offer. What I can say is that this question is the real cost of an AI transformation, that it first arises in silence, and that it appears in no business case.

The only indicator I really watch

The signal I look for in someone I'm supporting is counter-intuitive: I want them to end their day as tired as before, and often more so. Tiredness is not a virtue, but it is the only reliable sign that they went further in their thinking rather than faster to the same place.

That is what separates, in my head, someone assisted from someone augmented: assistance lightens the day and produces more, whereas augmentation relieves nothing at all and merely moves the ceiling of what you are able to handle.

The criterion looks like nothing, and yet it turns the way you steer an AI rollout upside down. If everyone goes home lighter in the evening and the production counters are climbing, the organisation has probably started borrowing without noticing, rather than gaining productivity.

So I am going to talk about debts. The word is not an accusation. A debt finances an immediate acceleration and is repaid later, with interest. The problem is never borrowing, it is borrowing without knowing it. In the organisations where I work — large structures engaged in agile transformation at scale — I see 3 being contracted in parallel. The debt of people, the debt of the collective, the debt of the system.

A word of caution on vocabulary. AI does not invent these debts. Most of these mechanisms predate it, the first of them having been described in aircraft cockpits. What AI does is make them glaring and industrialise them. DORA puts it better than I do: "AI amplifies what already exists" [1] — a mechanism I have detailed elsewhere in AI reveals and amplifies organizational dysfunction. An organisation with solid foundations therefore borrows less than the others, which is rather a reason to invest in its fundamentals.

The bill for people

AI excels at scales and drills. The first draft, the summary, standard code, routine analysis. These are exactly the tasks through which you learn a craft, through which you keep your hand in once you have learned it, and which are tiring.

A controlled trial published by Anthropic in January 2026 gives the order of magnitude [2]. 52 Python developers, mostly juniors, discover a library they do not know. Half have an AI assistant in addition to the official documentation, the other half the documentation alone. On the comprehension quiz taken without AI after the task, the assisted group scores 50 % correct against 67 %. On time, the 2-minute gap is not significant. On comprehension, it is, and it is worth 17 points. In other words, at lower effort and equivalent output, the assisted group paid for the saving in understanding.

MIT measured the short version of the same phenomenon [3]. Participants write a text with ChatGPT, electroencephalogram on their heads, and 83 % of them prove unable, a few minutes later, to quote the content of "their" own text. The study is a preprint on a small sample and I take it for what it is, but the mechanism is the one I observe. You produce without encoding, and nothing flags it at the time.

I see it in myself first. I built a mission second brain that produces the summaries of my meetings. The temptation, every time, is to pass it on without ever redoing the exercise, even mentally. The document is defensible. I am not, if someone questions me on it the next day.

For an executive, the difficulty is that this bill arrives in 2 instalments. At 12 months it is invisible, since everyone is producing more. At 5 years, you need seniors capable of challenging the AI and you have not trained the generation that was supposed to become them.

The other half of the bill is the one that opens this article, and it is not measured in skill. Guy Champniss surveyed 1,200 employees and identified 6 effects that set in when AI arrives without support [4]. Cognitive delegation and loss of critical distance, a sense of dispossession from one's own work, doubt about one's own competence, weakening of collective exchange, fear of losing credibility with peers by the very fact of using the tool, and worry about the future value of one's profession. The point that matters for anyone steering a transformation is that these effects also hit — and sometimes first — the people who find the tool wonderful and use it every day, so that treating them as change resistance is the fastest way to increase the debt rather than reduce it.

The bill for the collective

A strategic relaunch deck, prepared for a cross-programme leadership session, reached me at an advanced stage of preparation. Sections duplicated in French and English repeating the same value proposition without adding anything. Vague titles. Polls and draft slides that should have stayed internal notes. Nothing was wrong. Everything needed reworking by the person who had to use it to run the session.

BetterUp Labs and the Stanford Social Media Lab gave this a name: "workslop" [5]. AI-generated content, presentable, empty of substance, which transfers the workload onto whoever receives it. Of 1,150 US employees surveyed, roughly 40 % received some within the month, and each occurrence costs its recipient 1 hour 56 minutes. Scaled to an organisation of 10,000 people, the authors put the loss at a little over 9 million dollars a year.

The figure is striking, but it is not what should worry an executive. Return to the criterion from the start: workslop is tiredness changing hands. The sender finished earlier, the recipient will spend 2 hours on it, and the dashboard records only the first half of the operation.

Then comes an effect I have seen measured nowhere and which I give for what it is, a field deduction. When workslop circulates, you stop knowing what someone's word is worth. Does this document reflect their analysis or a model's? Is that confidence expertise, or borrowed fluency? The cost of verification, which mutual trust existed precisely to save, reinstalls itself everywhere. The signal is easy to spot once you look for it: someone quietly redoes work already delivered rather than asking its author what they checked.

There is a third collective debt, more counter-intuitive. Doshi and Hauser showed in Science Advances that AI improves individual creativity, especially that of the least creative, while reducing the collective diversity of output [6]. For a leadership team seeking differentiation, that is a strategic problem, since value lies in the ideas others have not had rather than in the average quality of the ideas produced.

The bill for the system

The DORA 2025 report establishes that AI increases delivery throughput and product performance, while retaining a negative relationship with stability [1]. GitClear confirms it on repositories, more than 200 million lines changed between 2020 and 2024 [7], with copy-pasted code rising from 8.3 % to 12.3 % and refactoring collapsing from 24.1 % to 9.5 %. Generating code has never been faster, while understanding it, maintaining it and keeping it alive takes exactly as long as before, which widens month after month a gap nobody provisions for — the core of what I call elsewhere the generative AI plus technical debt cocktail.

I paid this bill directly. An agent I had designed to support a Product Management role, able to run scripts that create and modify backlog items in production, operated for several weeks without a single test. Neither for the scripts, nor for the agent's behaviour. The architecture review I eventually ran put it this way: "AI is probabilistic; without regression tests, every update is Russian roulette". Generating the tool had taken a few hours. Realising it was not sustainable took several weeks. The same review found an access token for a backlog management tool exposed in clear text in a configuration file, with no rule having anticipated protecting it.

That last point opens the governance debt. The MIT Media Lab and Project NANDA report notes that in more than 90 % of the companies surveyed, employees regularly use personal AI tools for their work, while only 40 % of those companies have bought an official subscription [8]. Usage moves outside the frame, client data included — which is the very definition of Shadow AI.

The reflex is to ban. The Samsung case deserves to be read to the end, though, because it says the opposite of what it is usually made to say. Samsung Device Solutions had authorised ChatGPT in March 2023, then experienced 3 leaks in 20 days, including a meeting record and source code [9]. The ban came a month later, temporary, announced for as long as it took to deploy an internal tool. Authorising is therefore not enough, and banning settles nothing. What was missing, between the two, were usage rules written before opening up and a date by which the internal alternative exists — in other words, operational AI governance.

That leaves the most discreet debt of all, the debt of measurement. Automating a flow you have not rethought freezes the dysfunction under a layer of apparent efficiency. And measuring AI by volume produced installs a perfect incentive to manufacture workslop. Marilyn Strathern condensed it in 1997, reformulating Goodhart's law: "when a measure becomes a target, it ceases to be a good measure" [10].

I fell into it. For the Product Management support agent, I settled on 2 metrics: a rework rate with a target below 20 %, and an output-template conformity rate with a target of 100 %. Demanding 100 % on form and tolerating 20 % rework on substance describes, read coldly, the exact profile of what I have just described: a deliverable that conforms to the template and is improvable on content.

What I do, concretely

Nothing above argues for slowing down. Well orchestrated, AI does exactly the opposite of what this text describes, and that is what I try to bring about.

The augmented workshop. Humans in a room, and one person holding a scribe role who feeds the AI as the discussion goes. The point is not speed: the people in the room arrive with their biases, and a well-orchestrated AI makes it possible to go beyond them and cover the blind spots the group will not see on its own. It is the only setup where I have seen AI reduce a collective bias instead of adding one.

Knocking the ball around. The PM it works best with does not ask the AI to write his features. He has it embody proto-personas and sends them to challenge what he proposes, to check that the feature makes sense and delivers value. Same thing on acceptance criteria. What comes out at the end is far beyond what was generated at the start, and that is precisely why the session is tiring.

Dose on the expertise curve, not the rollout curve. I open access in layers, and I wait to see that the person has mastered one layer before opening the next. The passing criterion is not that they can use the tool, it is that they use it to go further rather than to make life easier. I do it by feel, because I went through these stages myself while paying attention, and I have no better method to offer today. That is a real limitation of what I am describing here.

Preserve what is learned. On tasks with learning at stake, invert the order of operations: think first, generate second. Ring-fence AI-free work at the start of a skill ramp-up. Assess understanding as much as the deliverable, for instance by asking someone to defend a solution as if the AI had not written it.

Institutionalise contradiction. Ask the AI to attack its own proposal, require a structured alternative before any assisted decision, and keep one simple criterion. If nobody can explain why the AI's proposal is good, it is not accepted yet. This reflex is the best antidote to the cognitive biases that sabotage an AI transformation.

One objection remains to all of this, and it stops me. These practices ultimately rest on a human who checks, and that is exactly the faculty that degrades in contact with a machine that is often right. The report Deloitte delivered to the Australian government in 2025, riddled with invented references and a fabricated judicial citation, had passed through every level of internal review [11].

I have no complete answer. I have a distinction, and it only holds if it is measured. Nominal validation consists of reading and approving. Effective validation leaves a trace of what it changed, and its rejection rate is a number someone looks at. The question to ask in a programme review is therefore not "do you have human validation", but "how many times has it said no".

Assisted, or augmented

Generative AI is probably the best asset offered to organisations in 20 years. That is precisely why its full accounting must be kept, and not only the column that pays.

As for the question at the start — what people are still worth if they don't think much — I believe it is badly posed, and it took me a while to understand why. It assumes that thinking is a property of people. What I see in the field is that it is above all a property of the situations we put them in. A PM given a tool that produces in their place will have fewer and fewer occasions to think. The same PM, in an augmented workshop where an AI comes to challenge their feature with 3 proto-personas, will have far more than before.

Which of the two it turns out to be does not depend on the tool. It depends on how the work is organised around it, on the time someone is given to disagree with what a machine proposes, and on what we decide to look at in the evening.

Appendix: the liability in 13 items

ItemWhat is being borrowedThe warning signalThe countermeasureSource
Cognitive debtProducing without memorisingSomeone unable to defend the document they "wrote" yesterdayThink first, generate second, on tasks with learning at stake[3]
Psychological debt6 effects that set in when AI arrives without supportTeams producing more, asking for less help, and avoiding the topic in retrospectivesTreat adoption as a management subject[4]
Skill debtThe junior's learning path dries upHighly productive juniors whose questions become rareRing-fence AI-free work, assess understanding[2]
Automation biasHumans stop checking a machine that is often rightA deliverable crossing several reviews without a trace of changeInstitutionalise contradiction[12][13]
Anchoring biasThe AI's 1st proposal becomes the frame of the thinkingWorkshops where the chosen solution is always a variant of the first draftRequire a structured alternative before deciding[14]
Illusion of competenceHaving read a summary is mistaken for having understoodSomeone speaking about a file with more confidence than time spent on itMake them explain; ask for the demonstration[15][16]
SycophancyThe model agrees with the userAn AI that has refused you nothing in weeksAsk the AI to attack its own proposal[17][18]
WorkslopThe sender's tiredness moves to the recipient"It's well written but it says nothing"Sign what you pass on, in the strong sense[5]
HomogenisationEveryone improves, the whole looks alike3 proposals, 3 workshops, the same structure and the same examplesDiverge without AI before converging with it[6]
Trust debtYou no longer know what someone's word is worthSomeone quietly redoing work already deliveredDeclare usage: what was assisted, to what degree, what was checkedField deduction
Technical debtThroughput rises, stability fallsCode that works and nobody can explain any moreCode review and continuous integration, non-negotiable[1][7]
Governance debtReal usage moves outside the frameA leadership team discovering the scale of usage through an incidentEnable, with written rules and a date[8][9]
Process and measurement debtYou measure volume, so you produce volumeA dashboard where everything is up and nobody is at easeRethink the flow before augmenting it; measure value[10]

References

  1. Google Cloud / DORA, State of AI-assisted Software Development 2025. Self-reported data, roughly 5,000 respondents; the 2025 throughput result reverses that of 2024.
  2. Shen, J. H., Tamkin, A. (Anthropic), How AI Impacts Skill Formation, arXiv:2601.20245, January 2026.
  3. Kosmyna, N. et al. (MIT Media Lab), Your Brain on ChatGPT, arXiv:2506.08872, 2025. Preprint, not peer-reviewed, 54 participants.
  4. Champniss, G., The Psychological Costs of Adopting AI, Harvard Business Review, 1 May 2026.
  5. Niederhoffer, K. et al. (BetterUp Labs / Stanford Social Media Lab), AI-Generated "Workslop" Is Destroying Productivity, Harvard Business Review, 22 September 2025.
  6. Doshi, A. R., Hauser, O. P., Generative AI enhances individual creativity but reduces the collective diversity of novel content, Science Advances, 10(28), 2024.
  7. GitClear, AI Copilot Code Quality, repository data 2020-2024.
  8. MIT Media Lab / Project NANDA, The GenAI Divide: State of AI in Business 2025. Not peer-reviewed, 52 interviews and 153 executive responses.
  9. Samsung / ChatGPT incident: authorised in March 2023, 3 leaks in 20 days, temporary restriction in May 2023.
  10. Strathern, M., 'Improving ratings': audit in the British University system, European Review, 5(3), 1997, p. 308, reformulating Goodhart (1975).
  11. Deloitte Australia, Targeted Compliance Framework Assurance Review, Department of Employment and Workplace Relations, published 4 July 2025, corrected version 26 September 2025, refund of 97,587 Australian dollars.
  12. Parasuraman, R., Manzey, D. H., Complacency and Bias in Human Use of Automation, Human Factors, 52(3), 2010.
  13. Moffatt v. Air Canada, 2024 BCCRT 149, February 2024.
  14. Tversky, A., Kahneman, D., Judgment under Uncertainty: Heuristics and Biases, Science, 185(4157), 1974.
  15. Fernandes, D. P. B. et al., AI makes you smarter but none the wiser, Computers in Human Behavior, 175, 2025.
  16. Fisher, M., Goddu, M. K., Keil, F. C., Searching for explanations, Journal of Experimental Psychology: General, 144(3), 2015.
  17. Cheng, M. et al., Sycophantic AI decreases prosocial intentions and promotes dependence, Science, 391(6792), 2026.
  18. OpenAI, Expanding on what we missed with sycophancy, 2 May 2025.

About the author and declaration of interests

I am a consultant at inspearit, a firm that sells an AI transformation offering under the IAgile™ brand. The field observations come from an ongoing agile transformation at scale, anonymised.

Running an AI rollout and wondering which of these 3 debts you are currently contracting? Let's take stock in 30 minutes, no slides.

Review your AI liability →