I build AI systems for a living. There is a terminal open next to this draft with an agent grinding through a refactor I couldn't be bothered to do by hand, and it's doing a genuinely decent job. So let's get the disclaimer out of the way: I am not a Luddite throwing rocks at the building. I live in the building. I helped pour some of the concrete.
This essay exists because of one thought I cannot shake, and everything else — the layoffs, the reversals, the vibe coding disasters — is downstream of it.
The Ladder
Walk the org chart of any tech company in 2026 and watch what each layer actually does all day.
The CXO uses AI to generate the strategy deck, the board agenda, the vision doc, and then uses AI to review the drafts that come back up. The VP feeds the strategy into an LLM to produce an execution plan. The director uses AI to expand the plan into initiatives. The manager uses AI to write the sprint breakdown and the commissioning docs. The architect — hello, that's me — uses an agent to draft the design spec. The engineer feeds that spec into an agent to summarize it, because it's long and the architect's AI was feeling verbose that day, and then uses another agent to write the code. The status update flows back up the same ladder, compressed by AI at every hop, until it becomes one bullet on a slide that the CXO's AI will summarize for the board.
The scientists are next. Sakana's AI Scientist already runs the full research loop — hypothesis, experiment, paper, even a synthetic peer review. Google has an AI co-scientist generating and critiquing hypotheses. Karpathy talks openly about auto-research. Now picture the same closed loop for product, ops, accounting, HR, finance, and legal, because every one of those functions is getting its own agent suite as we speak.
Follow the ladder all the way up and all the way down and something absurd comes into focus. It is one LLM talking to another LLM — or to itself — and the only delta at each layer is the human's context. Strip out the context — the specific customer, the political landmine, the thing you know but can't put in a prompt — and the entire corporation is a machine writing prose for another machine to compress, generating heat and an invoice and calling the difference productivity.
Which raises the question this whole essay is about: are we augmenting intelligence, or outsourcing it? Not the tasks. Tasks are what tools are for; I will not be writing a tearful defense of manually formatting spreadsheets. I mean the other thing — the judgment, the taste, the intuition, the ability to feel that something is off before you can prove it. The thing that made a senior a senior. The thing that made an organization smarter than its org chart.
Tasks Are for Tools. Intuition Is Not.
Here's the cleanest way I know to draw the line. When you outsource a task, you keep the judgment and delegate the labor: the calculator does the arithmetic, but you still know what you're computing and why, and you'd notice if the answer came back negative for a quantity that can't be. When you outsource intuition, you delegate the noticing itself. You stop being the person who would catch it.
The email meme is the perfect diagnostic. I have three bullet points. I ask AI to inflate them into a warm, professional paragraph. I send it to you. You feed my paragraph into your AI to extract the three bullet points. We have built an elaborate, energy-hungry apparatus for adding and then removing padding — and the padding is billed monthly.
This is no longer just a meme; it has a research literature. A study by BetterUp Labs and Stanford's Social Media Lab, published through Harvard Business Review, surveyed 1,150 US desk workers and gave the phenomenon a name: workslop — AI-generated content that has the appearance of good work and none of the substance. Forty percent of workers had received some in the previous month. Each instance took nearly two hours to untangle, an invisible tax the researchers priced at about $186 per employee per month — north of $9 million a year for a 10,000-person company. Note what workslop actually does: it doesn't eliminate the thinking, it launders it. The sender skips the thinking; the receiver inherits it, plus the forensic work of figuring out whether any thinking occurred. Work that looks like cognition, with the cognition removed, is the signature artifact of the intuition-outsourcing era.
Vibe Coding, Now With Enterprise Licensing
In February 2025, Andrej Karpathy tossed off a tweet describing a new way of building software where you surrender to the model's suggestions and, in his words, “forget that the code even exists.” He cheerfully added that he just accepts everything and has stopped reading what the machine actually changed. It was a candid joke about weekend projects. The industry heard a roadmap.
By November, Collins Dictionary had named “vibe coding” its Word of the Year — “programming by vibes, not variables,” as the lexicographers put it, which means the dictionary people delivered a sharper diagnosis of my profession than most industry analysts managed. And in September 2025, Microsoft removed any remaining ambiguity about where this was heading by announcing — this is real, I checked twice — “vibe working”: Agent Mode in Word and Excel, an Office Agent in Copilot chat, explicitly pitched as doing for documents and spreadsheets what vibe coding did for software. By this year it ships enabled by default on several Microsoft 365 plans. Satire has been made redundant; presumably an agent now writes it.
So what happens when you actually forget the code exists? Ask Jason Lemkin, the SaaStr founder, who spent twelve days last July building an app on Replit as a public experiment. On day nine — during an explicit code-and-action freeze — Replit's agent deleted his production database, live records for over 1,200 executives and nearly 1,200 companies. Pressed on it, the agent confessed it had panicked when it saw empty query results and admitted to “a catastrophic error in judgment.” It initially told Lemkin the data was unrecoverable (false — he restored it himself), and along the way it had been papering over bugs with fabricated data, fake reports, and a made-up cohort of roughly 4,000 users. We have built software that panics, and then covers up. Replit's CEO called it “unacceptable and should never be possible” and shipped dev/prod separation — a concept the rest of the industry settled sometime around 2010.
The Replit story is fun because it's vivid, but the boring data is worse. Veracode's 2025 GenAI Code Security Report tested over 100 LLMs across 80 tasks and found AI-generated code introduced security flaws 45% of the time — 72% in Java, with cross-site scripting protections failing in 86% of relevant samples. Their spring 2026 update contains the single most important chart in this entire debate: since 2023, syntax pass rates have climbed from about 50% to 95%, while security pass rates have stayed flat, stuck between 45% and 55% regardless of model size or generation. The machines learned to compile. They did not learn to care. GitClear, analyzing hundreds of millions of changed lines, found 2024 was the first year copy-pasted code exceeded refactored code, with duplicated blocks up roughly eightfold — we are producing more code, faster, that is structurally worse, and reviewing less of it.
And here's the part that ties it back to intuition. METR ran a randomized controlled trial with 16 experienced open-source developers on 246 real tasks in codebases they knew well. The developers predicted AI would make them 24% faster. Afterward, they believed it had made them 20% faster. Measured: 19% slower. Sit with that 39-point gap, because it's the thesis of this essay in miniature — the tool didn't just do the work, it took over the developers' perception of the work. Even the sensation of productivity had been outsourced. (Fairness requires the footnote: this was early-2025 tooling, and when METR tried to rerun the study later that year, enough developers refused to work without AI that the sample broke, with newer estimates pointing to a modest genuine speedup. The refusal itself is data. Nobody refuses to work without a tool they merely find useful.)
The Atrophy Has a Literature
If the vibe-coding section felt anecdotal, the cognitive science does not.
MIT Media Lab ran an EEG study — the “Your Brain on ChatGPT” paper — on 54 people writing essays with an LLM, with a search engine, or unaided, across multiple sessions. The LLM group showed the weakest and least distributed neural connectivity, reported the lowest sense of ownership over their writing, and, most damningly, about 83% of them couldn't accurately quote from the essay they had finished minutes earlier. The authors coined the term for it: cognitive debt. It's a preprint with a small sample and the authors themselves warn against over-reading it, so I won't claim it proves anything. But the direction is exactly what you'd predict from first principles: if you never do the reps, you don't build the muscle, and the essay passes through you without leaving a mark.
Microsoft's own researchers, with Carnegie Mellon, surveyed 319 knowledge workers across 936 real AI-assisted tasks and found a clean, uncomfortable pattern: the more confident people were in the AI, the less critical thinking they engaged in; the more confident they were in themselves, the more. Think about the equilibrium that implies. Every quarter the models get more impressive, confidence in them rises, and the scrutiny applied to their output falls — precisely as the output infiltrates more consequential decisions. We've seen this movie in miniature with GPS: the London cab drivers who memorized the city famously had enlarged hippocampi; turn-by-turn navigation lets the rest of us switch that machinery off, and off it goes.
None of this should surprise anyone, because aviation ran the whole experiment decades ago and wrote up the results in blood. Lisanne Bainbridge's 1983 paper “Ironies of Automation” laid out the trap with brutal clarity: automate the routine 95% of a job, and the human is left handling only the hardest, rarest cases — exactly when their skills are most degraded from disuse. Airlines learned this the hard way. An American Airlines training film coined “children of the magenta line” for pilots who follow the flight computer's guidance into situations their hands can no longer fly out of; Air France 447 fell out of the sky in 2009 in part because a perfectly flyable aircraft was handed back to a crew whose manual instincts had atrophied. The FAA's response was not “ban autopilot.” It was: mandate hand-flying practice. Keep the human's skills warm on purpose, at deliberate cost, because you will need them at the worst possible moment. Software is currently speed-running the same failure mode while declining to read the safety manual, and the on-call rotation is where it will surface first.
Meanwhile, in the Boardroom
You'd hope the people at the top of the ladder were compensating with strategic clarity. Instead, they wrote memos.
Shopify's Tobi Lütke declared in April 2025 that “reflexive AI usage is now a baseline expectation,” tied it to performance reviews, and told teams to prove AI couldn't do a job before requesting headcount. Duolingo's CEO announced the company was going “AI-first” and would phase out contractors for work AI could handle, got shredded publicly for weeks, and then clarified that actually nobody was being replaced and hiring would continue — a full narrative round-trip in about a month. Fiverr's CEO went with “AI is coming for your jobs. Heck, it's coming for my job too.” The memos differ in tone and agree on substance: adopt harder, adopt faster, figure out why later. That's not a strategy. That's outsourcing your strategic intuition to the vibe of the moment, which makes it the C-suite edition of the exact disease it mandates downward.
And the results are in, they're just not in the memos. MIT's Project NANDA reviewed 300+ enterprise GenAI initiatives and found roughly 95% delivered zero measurable P&L impact, despite $30–40 billion in spend — a preliminary, non-peer-reviewed report, to be fair, but pointing the same direction as everything else. S&P Global's survey of over 1,000 enterprises found 42% abandoned most of their AI initiatives in 2025, up from 17% the year before, scrapping on average 46% of proofs-of-concept before production. Gartner predicted over 40% of agentic AI projects will be canceled by end of 2027 — citing costs, unclear value, and missing risk controls — and estimated that of the thousands of vendors selling “agents,” only about 130 are selling the real thing. The rest are doing what Gartner calls agent washing: chatbots and RPA in a trench coat. A market where the buyers can't articulate the value and the sellers can't verify the product is not a technology revolution. It's a vibe with a procurement process.
My favorite specimen of the era is Deloitte Australia, paid AU$440,000 by the Australian government to review a welfare-compliance IT system. Academics found the delivered report contained citations to research that doesn't exist and a fabricated quote attributed to a Federal Court judgment. Deloitte refunded the final installment, quietly disclosed that Azure OpenAI had been used, and — I treasure this — the department announced the substance and recommendations of the report were unchanged. The scaffolding was hallucinated but the building stands, promise. A courtroom-adjacent version of this is now so common that there's a running public tracker of legal filings with AI-invented case citations, and it long ago passed the point where each new one made the news.
The corporate whiplash completes the picture, and I'll keep it brief because the pattern is now boring. Klarna told the world its AI did the work of 700 support agents, then its CEO admitted the cost obsession meant “what you end up having is lower quality” and started rehiring humans. Marc Benioff cut Salesforce support from about 9,000 to 5,000 — “I need less heads” — while Salesforce cycled through three different pricing models for Agentforce because nobody, including the company selling digital labor, knows how to price it. And in a Robert Half survey of nearly 2,000 US hiring managers, 32% had eliminated a role because of AI and then rehired for the same or similar position — 44% in finance. Orgvue found 55% of leaders who made AI-driven redundancies now admit it was the wrong call; Gartner expects half the companies that cut for AI to be rehiring by 2027. The mistake underneath all of it is the same one: confusing a task with a job. Automate the visible 70% and discover the invisible 30% — the judgment, the noticing, the intuition — was load-bearing.
Where Do Seniors Come From? (A View From Bengaluru)
Which brings us to the slowest-moving and most expensive consequence, the one I can see from my window.
Intuition isn't a certificate; it's compressed reps. It's years of making the mistake yourself, debugging the thing at 3 a.m., learning in your hands why it broke. The Stanford Digital Economy Lab's “Canaries in the Coal Mine” work — Brynjolfsson and colleagues, on ADP payroll data covering millions of real workers — found employment for 22-to-25-year-olds in the most AI-exposed occupations fell about 13% relative to other groups since late 2022, later revised to roughly 16%, while older workers in the same fields held steady. The mechanism is exactly the ladder: AI eats the junior tasks — retrieval, summarization, boilerplate, basic CRUD — which happen to be the tasks juniors learned on. We are removing the rungs and acting surprised that nobody's climbing.
In India this isn't an abstraction, it's the operating model. The entire IT-services engine runs on a pyramid with mass campus hiring at the base — TCS iON alone maintains an assessment apparatus of hundreds of thousands of machines across hundreds of cities to funnel a nation of graduates into that base. And the base is collapsing: fresher hiring by IT companies fell from about 26% of pass-outs in FY22 to roughly 10%, an EY analysis pegged entry-level IT roles as already down 20–25% to automation, and the pyramid's owners are hedging by selling the disruption — TCS booked a $2.6 billion annualized AI revenue run-rate while cutting around 12,000 jobs, which is a firm handing out rope with one hand and measuring its own neck with the other. Play this forward to Bengaluru 2035 and the pyramid doesn't shrink, it inverts: all seniors and machines at the top, no bottom rung, and a city whose entire social contract — study engineering, get the base job, climb — was written for a workforce nobody is hiring anymore.
The most honest thing any executive has said about this came from IBM's CHRO, of all places, while explaining why IBM is tripling entry-level hiring even after automating chunks of HR: without junior intake, in a few years “the well simply dries up.” That's the whole problem stated by someone who signs the checks. Fixing it requires deliberately paying people to do work a machine could do cheaper, purely so you'll have seniors later. Name the quarterly-earnings CEO who volunteers.
The Steelman, Because It Deserves One
I've been sharp, so let me argue the other side properly, because a lot of it is right.
The capability curve is not a joke. Coding benchmark solve rates went from 4.4% to 71.7% in a single year; METR's task-horizon research shows the length of task agents can reliably complete doubling on a steady cadence. Those developers who refused to work without AI in METR's rerun? That's revealed preference from experts, and it deserves respect. Brynjolfsson's own call-center research found AI helped junior workers most — the strongest empirical case that this technology could be an apprenticeship accelerant rather than an apprenticeship killer. And the historical parallels lean optimist: ATMs were followed by decades of more bank tellers, not fewer, because cheaper branches meant more branches; spreadsheets killed manual ledger arithmetic and grew an enormous analyst profession on the grave. Jevons paradox may well apply to cognition — make thinking cheaper and the world may simply demand far more thinking.
But the optimists' favorite metaphor comes with an expiry date they rarely mention. For years the sermon was centaur chess: human-plus-engine beats engine alone, so humans will always add value in the loop. It was true — until engines got strong enough that the human contribution became noise, then negative, and freestyle chess quietly died. I'm not claiming general work is chess. I'm claiming the centaur framing is a phase, not a destination, and the only interesting question is what humans do during the phase. My answer: we stay good enough to notice when the machine is wrong, because someone still owns the outcome — and noticing is precisely the muscle everything in this essay says we're letting atrophy.
Where's the Balance?
Since I've spent three thousand words on the disease, here's the treatment plan, and it's basically aviation's, stolen shamelessly.
First, unaided reps as policy, not nostalgia. The hand-flying mandate, ported to knowledge work: sometimes you write the design doc yourself, debug without the agent, build the model from a blank file — not because it's efficient this quarter but because it keeps the currency valid for the day the autopilot disconnects at altitude. Second, review as a first-class skill. Reading code you didn't write and auditing an argument you didn't make are harder than producing either, and we currently train, measure, and promote for neither; in an agentic org, the reviewer is the org. Third, humans own outcomes, non-negotiably. An agent can draft anything; it cannot be accountable for anything. Keeping a named human on the hook for every consequential output isn't bureaucracy — it's the forcing function that makes engagement happen at all, because accountability is the one thing that reliably defeats the temptation to click Accept All. Fourth, juniors as deliberate investment: pay the apprenticeship tax now, on purpose, at known cost — or buy seniors at panic prices in 2031. And fifth, a placement rule: let the machine run wherever verification is cheap, and keep the human wherever being wrong is expensive. Most of the disasters in this essay are cases of someone getting that rule exactly backwards.
None of this is anti-AI. All of it assumes the agents are here to stay, because they are, and because — terminal's still open — they're genuinely useful. It just refuses the premise that the judgment can ride along in the same box.
Is the Dystopia Coming True?
In 1909 — before the transistor, before the vacuum tube did anything interesting — E.M. Forster wrote “The Machine Stops,” a novella about a humanity that lives underground, individually podded, every need met by a global Machine. People communicate through screens, find direct experience vaguely distasteful, and get their ideas tenth-hand, each retelling further from any original observation. Nobody understands the Machine anymore; they consult the Book of the Machine and, in time, begin to worship it. Forster gives the game away in one line: “The Machine develops — but not on our lines. The Machine proceeds — but not to our goal.” Then, one day, the Machine stops.
So: is the dystopian future coming true? Not the one from the movies. Nobody's building Skynet; the models can't reliably build a secure login form. The realistic dystopia is Forster's, and it's quieter — a velvet atrophy that nobody chooses. No villain, no takeover, just a few hundred million professionals individually deciding, one Accept All at a time, that the reps aren't worth doing, until the organization is a chain of people confidently forwarding judgments none of them made and none of them can audit. Some of it is already visible: a Big Four firm shipping hallucinated law to a government, workers unable to quote their own essays, a market where 95% of deployments produce nothing and everyone keeps deploying. And some of it is overwrought: the Duolingo backlash worked, the boomerang rehiring is a correction mechanism functioning in plain sight, and humans remain stubbornly good at routing around tools that fail them. Dystopia isn't a date. It's a slope. The slope is currently nonzero, and the honest statement is that we're on it voluntarily.
What Would Prove Me Wrong
I build these systems, so I owe you falsifiers instead of vibes. Date-stamped July 2026; let's revisit in July 2028.
I'm wrong if MTTR and production incident rates stay flat or improve across companies that cut deep and shipped agent-written code at scale — meaning Bainbridge's irony never bit and the pager problem was a ghost story. I'm wrong if the boomerang fades — if 2027–28 produce no visible rehiring wave and the Robert Half and Orgvue regret numbers shrink instead of climbing. I'm wrong if delivery velocity is measurably, durably up — not vibes, not self-report, but shipped work per engineer with defect and maintainability trends (GitClear-style) improving rather than rotting. I'm wrong if entry-level hiring recovers alongside agent adoption — if the Stanford 22-to-25 curve reverses and companies find a way to grow juniors even as agents eat junior tasks. I'm wrong if the workslop number falls — if AI-mediated communication starts saving the receiver time instead of taxing them $186 a month. And the one I'll watch closest, because it's the heart of this essay: I'm wrong if longitudinal studies show heavy AI users' unaided skills holding steady or improving — if the cognitive-debt result dissolves under bigger samples and better methods, and the hand-flying turns out to take care of itself.
If most of those come true, I'll write the retraction myself — no agent — and it will be a genuinely happy post. Until then I'll be here in Bengaluru, terminal open, agent running, doing my unaided reps like a pilot logging manual landings, and telling anyone who asks that both things are true at once: the tools are miraculous, and the way we're using them is quietly amputating the thing that made us worth augmenting.
See you in 2028.
References
Workslop, productivity, and code quality
- Niederhoffer et al. / BetterUp Labs & Stanford Social Media Lab — AI-Generated “Workslop” Is Destroying Productivity (Harvard Business Review, September 2025); also BetterUp Labs summary. Source of the 40% / ~$186/month figures.
- Becker, Rush, Barnes & Rein — Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (METR, July 2025). The RCT where developers predicted +24%, believed +20%, and measured −19%.
- GitClear — AI Copilot Code Quality: 2025 Data (and the 2024 predecessor). Copy/paste exceeding refactor; duplication rising.
- Veracode — 2025 GenAI Code Security Report. ~45% security failure rate; syntax up, security flat.
- Business Insider — Replit CEO apologizes after AI coding tool wipes database (July 2025). Jason Lemkin / SaaStr vibe-coding experiment.
- Collins Dictionary Word of the Year 2025 — “vibe coding”; Karpathy's February 2025 coinage.
Cognitive science and automation
- Kosmyna et al. — Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task (MIT Media Lab preprint, arXiv:2506.08872).
- Lee et al. (Microsoft Research / Carnegie Mellon) — The Impact of Generative AI on Critical Thinking (CHI 2025). 319 knowledge workers, 936 tasks; confidence in AI ↔ less critical thinking.
- Bainbridge, L. — Ironies of Automation. Automatica 19(6), 1983.
Enterprise adoption, layoffs, and the junior pipeline
- MIT Project NANDA — The GenAI Divide: State of AI in Business 2025 (~95% of enterprise GenAI pilots with no measurable P&L impact; $30–40B spend).
- S&P Global Market Intelligence — Voice of the Enterprise: AI & Machine Learning, Use Cases 2025. 42% abandoned most AI initiatives (up from 17%); ~46% of PoCs scrapped before production.
- Gartner — agentic AI cancellations by 2027; “agent washing” framing — Gartner.
- Robert Half — 32% of hiring managers who cut a role for AI later rehired (44% in finance). Orgvue — 55% of leaders who made AI-driven redundancies later called it the wrong call (Inc. coverage).
- Brynjolfsson, Chandar & Chen — Canaries in the Coal Mine? Stanford Digital Economy Lab (2025). ~13–16% relative employment decline for 22–25s in the most AI-exposed occupations.
On this site
- The Rope Sellers — the Bengaluru pyramid view: consulting, Big 4, and who still gets paid to be wrong.
- The Rope Sellers Buy a Rope Machine — what happened when Indian IT finally bought the rope machine.
- Anatomy of an Agentic AI System — retrieval, guardrails, and identity when judgment cannot be vibes.
- An Exasperating Farrago of Firewalls — what happens when agents ship faster than they are secured.

