Fats & Sugars

AI Hype & Signal

Is AI making us stupid?

Tuesday 25 August 2026

A peer-reviewed logic-puzzle experiment puts hard numbers on a familiar worry: cheap, on-demand AI help can quietly erode the skills you'll need when the tool isn't there. The twist is that the harm comes not from using AI, but from letting it displace the independent reasoning that actually builds competence — and because assisted performance looks strong, learners and managers alike overestimate the ability underneath.

In this episode:
- People who leaned on AI performed worse once it was taken away
- The damaging factor is displaced independent reasoning, not AI use itself
- Cheaper help means more help — lower cost drove more requests and more dependence
- AI-assisted performance flatters true ability, and predictions of later solo work inherit that flattery
- The task was still learnable — those who engaged improved markedly, so weaker gains reflect offloading
- The AI was deliberately perfect, so this isn't about bad answers
- The findings are early and narrow, and the authors say so

Sources:
How AI Assistance Affects Human Skill Development: A Study of Learning with Logic Puzzles
Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task

AI Hype & Signal is produced with AI, including its two synthetic hosts, and every episode is grounded in cited sources and reviewed before release. Even so, it is intended for general information and discussion, not professional advice, so please check anything important against the original sources linked above before relying on it.

Transcript

The full word-for-word transcript of this episode. Plain text version

Marcus: You know, the common story we always hear is that artificial intelligence either makes us brilliant, granting us these sort of superpowers, or, well, it makes us irreparably lazy, right? But the reality is actually entirely different. The actual risk has nothing to do with the tool itself. It is entirely about the mechanism of substitution. A single tool can either scaffold your thinking, or it can substitute for it. And it's the substitution that actively and quietly erodes the exact skill you're going to rely on when that tool is suddenly gone.

Devon: I mean, the genuinely horrifying and frankly slightly amusing edge to this is what we might call the flattery effect.

Marcus: The flattery effect.

Devon: Exactly, because frictionless, on-demand assistance makes our work look spectacular, perfectly flatters our perceived competence. I mean, if you're sitting at a desk turning in flawless reports, you look incredibly capable, right up until the exact moment the Wi-Fi drops, or, you know, you're asked to perform without the safety net.

Marcus: And that's exactly what we're getting into in this deep dive. We're examining two pivotal pieces of research to unpack exactly how that safety net alters our capabilities. We're looking at a 2026 controlled experiment out of the University of California, Irvine, by Wu and colleagues, which used logic puzzles to track skill development.

Devon: Right, logic puzzles.

Marcus: Yeah. And then we're going to look at a 2025 MIT study by Cosmina and her team, who actually strapped EEG monitors to people's heads while they wrote essays with large language models.

Devon: Oh, wow.

Marcus: So, our mission today, well, our mission for this dive is to answer one critical question for anyone listening: is on-demand assistance scaffolding our minds, or are we actively replacing our own cognitive maps?

Devon: Let's start with the logic puzzle experiment, because I think that gives us the foundational mechanism. The research team at UC Irvine, they wanted to isolate exactly what happens to human skill when an artificial assistant is introduced and then taken away.

Marcus: Yeah.

Devon: So, they brought in 124 participants and put them through three very distinct phases.

Marcus: And the design of those phases is exactly what traps the variable we're looking for here. So, phase one was a pure pre-assessment, no artificial help, just a baseline test of the human's ability to solve a complex logic puzzle.

Devon: Right, just to see where they start.

Marcus: Exactly. Then phase two was the experimental session. Here, the participants were given similar puzzles, but they had access to an artificial assistant on demand, though using it came at varying costs, meaning they would lose a fraction of their potential bonus payment if they asked for a hint.

Devon: Ah, oh.

Marcus: And finally, phase three was the post-assessment. The assistant was entirely removed, they were completely back on their own.

Devon: And I think the control they used in phase two is absolutely crucial for the listener to grasp. The simulated assistant in that middle phase was deliberately flawless.

Marcus: Yes.

Devon: So, if a participant asked for help, the machine provided the 100% mathematically correct next step, every single time. It didn't hallucinate, it didn't give bad advice.

Marcus: Which means, if we see a drop in human performance in phase three, it is absolutely not a story about the machine giving bad answers and leading the human astray.

Devon: Right, because the tool worked perfectly.

Marcus: Exactly. So, if the tool is perfect but the human gets worse, the damage has to be in the displacement of effort.

Devon: But how do we actually measure that displacement? Because, I mean, just looking at a final score doesn't tell us what happened inside the participants' heads. The paper mentions they used a Bayesian latent ability model to figure this out.

Marcus: Yes, they did.

Devon: Before we gloss over that, how do you actually mathematically separate someone's raw brainpower from the flawless help they just received? I mean, for the listener, what is that model actually doing?

Marcus: It's a great question, and it's vital to understanding the results. A Bayesian latent ability model essentially looks at the probability of a participant getting a specific question right, based on two things: their established baseline ability from phase one, and the known difficulty of the specific puzzle. So, if a participant with a low baseline suddenly solves a notoriously difficult puzzle in phase two, the model can mathematically attribute that success entirely to the machine's intervention.

Devon: Right, because they couldn't have done it otherwise.

Marcus: Precisely. It isolates the human's actual contribution from the inflated outcome.

Devon: And that isolation revealed the core finding. The determining factor in whether someone actually developed their logical reasoning skill wasn't how often they used the tool, it was what the researchers called the solo share.

Marcus: The solo share.

Devon: Yeah. The actual muscle builder was the time spent thinking independently, you know, the friction of wrestling with the problem before reaching for help.

Marcus: The displacement of that independent reasoning is what actually caused the damage. When participants preserved their solo share, their underlying skills improved. When they outsourced the initial friction, their skills stagnated or eroded.

Devon: Which really forces us to look at the human stakes of this dynamic. I mean, does this mean cheap, frictionless help is inherently dangerous?

Marcus: Yeah.

Devon: Because when the researchers lowered the cost of asking for a hint, literally just making the financial penalty smaller, people simply used the tool more and they thought less.

Marcus: Yeah.

Devon: Think about the workplace or a classroom. Who actually pays for this?

Marcus: The individual pays the price, eventually.

Devon: Exactly. The individual whose manager sees their heavily assisted performance in phase two. Because the machine is flawless, the employee's output looks flawless. So, the manager drastically overestimates the employee's baseline competence.

Marcus: Right.

Devon: The system flatters the individual, the manager buys into the flattery, and then that individual is promoted into a senior role that requires the very independent reasoning they've been quietly offloading for a year.

Marcus: It's a terrifying thought. And Wu's team actually ran an additional variation in 2026 testing the informativeness of the tool. They wanted to see what happens when the assistant gives high-information, rich hints versus low-information, weak hints.

Devon: Well, I would assume the richer the information, the better the learning.

Marcus: You would assume that, but it caused a massive, immediate divergence in skill based entirely on the user's initial baseline ability.

Devon: Oh, really?

Marcus: Yeah. High-ability users treated the high-information tool as a scaffold. They engaged with it selectively, maintained a high solo share of independent thought, and their underlying skills grew rapidly. But lower-ability users treated that exact same rich information as a complete substitute.

Devon: Oh, no.

Marcus: Yes. They asked for help earlier, relied on it heavily, and their underlying skill actively degraded.

Devon: I mean, that is a deeply uncomfortable dynamic for educational equality. The very tool we assume will level the playing field actually widens the gap if we don't account for how different people deploy their solo share.

Marcus: It becomes even more complicated. In the post-study surveys, the lower-ability users who relied heavily on the tool reported highly elevated perceptions of their own ability.

Devon: Of course they did.

Marcus: They felt incredibly confident. They genuinely believed they had developed a successful puzzle-solving strategy, when their only actual strategy was offloading the cognitive work.

Devon: It's the ultimate trap of frictionless design. It doesn't just erode your skill, it completely blinds you to the erosion. But, I have to say, solving a logic puzzle is essentially mathematical. It's binary. You follow the rules and you're either right or you're wrong.

Marcus: True.

Devon: So, I'm highly skeptical that this clean, mathematical erosion translates to complex, creative work.

Marcus: You know, solving a logic puzzle is indeed just following rules. I wasn't entirely convinced this applied to actual creative work either, until I looked at the 2025 MIT study by Cosmina and her team. They wanted to see exactly what happens in the brain during a highly complex, open-ended task.

Devon: Right, writing an essay. This is a task that demands working memory to hold a sentence in your head, semantic retrieval to find the right vocabulary, narrative planning, critical thinking, all firing simultaneously. And they actually strapped EEG monitors to the participants' heads to watch it happen in real time.

Marcus: Yeah, the physical setup is fascinating. They split the participants into three groups. One group wrote their essays using a large language model, one group wrote using a standard internet search engine, and the third group wrote using only their brains, with absolutely no external tools.

Devon: The neurological data they captured here is what takes this from a sort of behavioural theory to a biological reality. What did the brainwaves actually show?

Marcus: Well, in the brain-only group, the EEG showed massive neural connectivity across the alpha, theta, and delta bands. Let's unpack what those bands mean for a moment. When those specific bands fire together in a state called neural coupling, it indicates high executive control, intense semantic memory retrieval, and heavy working memory load.

Devon: Which makes perfect sense. I mean, if you're staring at a blank screen trying to find the perfect word, your semantic retrieval network is working overtime. You're holding your previous sentence in your working memory while plotting the next paragraph. Your brain is integrating slow, deep associative thoughts with fast, executive decision-making.

Marcus: Right. Now look at the group using the large language model. They showed the weakest neural coupling, by far. Their frontal theta connectivity, which is the network heavily associated with holding things in your working memory and executing control, was significantly lower. Their brains literally powered down.

Devon: I really want to pause on the why there, because the brain is fundamentally an energy management system. It is desperate to conserve calories. Neuroplasticity is essentially biological efficiency.

Marcus: Exactly.

Devon: So, if a machine is holding the context, maintaining the working memory, and retrieving the vocabulary, your brain registers that those networks are no longer required for the task. It conserves energy by shutting them down.

Marcus: And the behavioural fallout from that neural powering down is staggering. It resulted in massive memory impairment regarding the participants' own work. In the very first session of the study, 83% of the LLM users could not correctly quote a single line from the essay they had just finished writing minutes prior.

Devon: 83%?

Marcus: 83%.

Devon: They couldn't remember their own words.

Marcus: No.

Devon: And during the interviews, that manifested as a deeply fragmented sense of ownership over the text. Many LLM users either claimed partial ownership, saying they wrote maybe 50% of it, or explicitly stated they felt zero ownership over the final essay.

Marcus: Wow.

Devon: Meanwhile, the brain-only group claimed absolute full authorship, and the search engine group landed somewhere in the middle.

Marcus: It's the exact same mechanism as relying on a GPS navigation system. Think about it. If you drive on your route yourself, you take wrong turns, you consciously look at landmarks, and you engage your spatial memory. It requires friction.

Devon: Yeah.

Marcus: But when the GPS is on, your spatial awareness powers down. You get to the destination flawlessly, but you build absolutely no internal map of how you got there.

Devon: It's like using a calculator for basic arithmetic over a period of years. Eventually, if someone asks you to do long division on a napkin in a restaurant, you just freeze. It's not just that you forgot a rule. The neural pathways that execute that specific logical sequence have literally been pruned back, because you haven't forced the friction of doing it yourself. The mechanism of execution has atrophied.

Marcus: And the MIT researchers tested that exact atrophy. They ran a fourth session, which they called the crossover. They took the participants who had been using LLMs for the first three sessions and suddenly forced them to write completely brain-only, no tools, a cold turkey cutoff.

Devon: Ah, the moment the cognitive debt comes due.

Marcus: So, if neuroplasticity is as responsive as we think, do their brainwaves just immediately spike back up to match the original brain-only group, because they're suddenly forced to do the hard work again?

Devon: They don't. No. Instead of returning to normal brain-only connectivity levels, these users showed distinct under-engagement of their alpha and beta networks. They had structurally adapted to the machine doing the heavy lifting. They were attempting to write an essay independently, but their executive function and deep semantic retrieval networks were lagging. The skill had measurably atrophied.

Marcus: And we can actually see the evidence of that atrophy in the words they chose. Because the researchers ran natural language processing, or NLP analysis, on the text of the essays themselves.

Devon: Break that down for us. What does NLP analysis actually look for in this context?

Marcus: So, they look at named entity recognition and specific n-grams. An n-gram is essentially just a cluster of words that frequently appear together. And the analysis showed that the essays produced by the LLM groups were highly homogeneous. They were statistically similar across the board.

Devon: Right.

Marcus: They were reusing the exact same predictable phrasing, the same clusters like "before speaking", and gravitating toward the exact same generic topics.

Devon: They all sounded identical.

Marcus: Yes.

Devon: And there is a profound irony here. We seek out these generative tools to expand our capabilities, to brainstorm wildly, to push beyond our own limits. But the reality is that they actually narrow our collective ideas.

Marcus: Yeah.

Devon: When you offload the friction of ideation, you default to the statistical mean of the model. You create a monoculture of thought.

Marcus: Which, to my mind, completely proves the underlying mechanism. Offloading independent reasoning actively erodes skill, full stop. The mechanism is completely clear across both the logic puzzles and the highly creative essay tasks. Whether it's mathematical reasoning or narrative generation, the neurological and behavioural laws hold true. When you substitute the cognitive effort, you lose the cognitive capacity. I honestly believe this represents a fundamental, generalisable cognitive law.

Devon: Okay, let me stop you there, because I have to challenge that conclusion quite heavily. I am absolutely not ready to declare a fundamental cognitive law based on the scope of what we're looking at here. Let's be honest about what we are actually measuring.

Marcus: Go on.

Devon: We have 124 people doing logic puzzles on a screen for a tiny financial bonus. And we have a small group of university students wearing EEG caps, writing essays in a highly artificial, short-term lab setting.

Marcus: I hear that, but the physiological data from the EEG is entirely objective. Brainwaves are brainwaves.

Devon: The data is objective, yes. But the scale is microscopic. That controlled environment is miles away from a professional architect, or a senior software engineer, or a financial analyst using these tools in a real job over a period of six months.

Marcus: Fair enough.

Devon: But in the real world, a professional is interleaving assistance with deep domain expertise, receiving team feedback, and going through iterative design. They are not just sitting in a vacuum hitting a generate button and walking away. I think we are at serious risk of overreading small, short-term experiments to declare a sudden crisis of human competence.

Marcus: I see the logic of the real-world environment being richer, absolutely. But the mechanism of biological efficiency does not care if you're in a sterile lab or a corporate office.

Devon: Right, true.

Marcus: If your frontal theta band isn't engaging because the software is maintaining the working memory, your brain will reallocate those resources. The logic puzzle study perfectly isolated this. The people who preserved their solo share of independent thought maintained their skill, the people who outsourced it lost it. The office environment doesn't change the underlying biology.

Devon: But the office environment entirely changes the incentive to preserve that solo share. A professional whose livelihood and reputation depend on the quality of their work has a massively higher incentive to critically evaluate the output than a participant doing a low-stakes puzzle on an online platform for a $3 bonus.

Marcus: That is a fair critique of the incentive structure, I'll give you that. But look back at the MIT study. Even when those participants were explicitly tasked with creating something original, they still fell into the homogeneity trap. They still couldn't remember their own quotes mere minutes later.

Devon: I agree, the immediate memory impairment is a stark, undeniable finding. But we are fundamentally missing the longitudinal data here. What happens after a year of using these tools?

Marcus: Oh, I don't know.

Devon: Exactly. Do we develop new, higher-level executive skills that we simply haven't identified yet? Perhaps the brain powers down at semantic retrieval because it is actively reallocating that energy to synthesis, curation, and editing. Skills that the specific EEG setup might not be highlighting in a simple essay task.

Marcus: We certainly don't have the decade-long studies yet, and we have to explicitly flag the limits of this research. Both sets of authors state very clearly that their findings are early and narrow.

Devon: Yeah.

Marcus: The link in the logic study between independent reasoning and skill gain is associational. It is not a proven causal law. It desperately needs rigorous testing in real-world professional training settings.

Devon: Exactly. If this holds up in the real world, then we have a massive pedagogical problem on our hands. But we cannot state as a fact that every worker currently using an LLM is experiencing irreversible cognitive atrophy.

Marcus: Agreed. But if this does hold up, it completely changes how we must view every software update that promises to make our workflows more seamless.

Devon: Which brings us to the actual design of these systems. If we want to avoid this atrophy, we have to critique the system architecture, not the user.

Marcus: Spot on.

Devon: It is fundamentally unfair to blame a stressed student or an overworked employee for taking the easy route, when the software is engineered by thousands of highly paid developers to be as frictionless and seductive as possible. Frictionless help is explicitly designed to make you offload.

Marcus: It is the path of least resistance by design. So, how do we actually fix the system?

Devon: Well, the researchers suggest highly practical counterweights that avoid the whole doom and gloom scenario. We can fix the design by intentionally reintroducing friction. Imagine an interface that simply prompts an initial human attempt first. You literally cannot click the assist button until you've typed 50 words of your own, or spent two minutes mapping out a logic problem on the screen.

Marcus: Forcing the solo share to occur before the assistant unlocks.

Devon: Precisely. Or intentionally delaying the help. Instead of generating a thousand words in two seconds, what if the system intentionally slows down to match the speed of human thought?

Marcus: That's interesting.

Devon: Or, perhaps most effectively, offering partial hints rather than the whole answer. If you're stuck in a piece of code, the system doesn't write the script for you. It highlights the line where the logic breaks and asks you what you think is missing. That is true scaffolding. It preserves the cognitive engagement while still providing support.

Marcus: It transforms the tool from a substitute into a cognitive forcing function. But at the time of release, we don't have those intentionally designed friction interfaces at scale. We have the frictionless ones.

Devon: We do.

Marcus: So, what is the concrete takeaway for the listener who has to go to work and use these tools?

Devon: The takeaway is that you have to artificially introduce the friction yourself. Treat frictionless assistants with intense suspicion.

Marcus: The practical application of this data is simple. Try the task yourself before you hit generate. Dedicate the first 10 minutes of any complex task to a blank page and your own brain.

Devon: And, more importantly, treat the availability of easy assistance as a signal to routinely check your own competence. Ask yourself, if the Wi-Fi dropped right now, could I still do my job? Or have I completely offloaded the cognitive map?

Marcus: It is the ultimate stress test for your own mind. If our tools are gradually removing the friction that builds our cognitive maps, what happens to the next generation of professionals who never experienced the friction in the first place? Do they even know what a cognitive map feels like, or will they only ever know the destination? Make sure you can still carry your own weight.