Fats & Sugars

AI Hype & Signal

Local AI Models, Are They The Future?

Saturday 5 September 2026

Small local models can now answer most everyday AI queries, but the real efficiency win comes from routing between device and cloud rather than local hardware winning outright. We dig into a large-scale study measuring intelligence per watt, Apple's hybrid on-device architecture, and a bank report reframing open, downloadable models as a market threat. The catch: the hard cases stay hard, and always-on local agents could quietly push energy costs onto users.

In this episode:
- Routing to the best local model handles nearly 89% of single-turn chat and reasoning queries
- The efficiency gain is smart local-cloud routing, not local hardware beating the cloud
- Local viability jumped sharply from 2023 to 2025, with intelligence per watt up 5.3x
- Local models cover creative tasks well but struggle with technical fields and the hardest reasoning
- Apple and open-weight model releases show big players building local-first, hybrid architectures
- A rebound-effect caveat: always-on local agents may raise total energy use, shifting costs onto consumers

Sources:
WWDC26: Apple unveils next generation of Apple Intelligence, Siri AI, powerful parental controls, and an expansive set of software improvements
Intelligence per Watt: Measuring Intelligence Efficiency of Local AI
MAI-Thinking-1: Building a Hill-Climbing Machine
Open-source AI 101: the battle for the future of AI
Computer Environments Elicit General Agentic Intelligence in LLMs

AI Hype & Signal is produced with AI, including its two synthetic hosts, and every episode is grounded in cited sources and reviewed before release. Even so, it is intended for general information and discussion, not professional advice, so please check anything important against the original sources linked above before relying on it.

Transcript

The full word-for-word transcript of this episode. Plain text version

Marcus: Imagine you are sitting in a coffee shop. You are entirely offline, you know, no Wi-Fi, your mobile data is completely switched off.

Devon: And your battery is probably hovering at, what, maybe 50%?

Marcus: Exactly, 50%. But you have this massive 200-page legal document that you urgently need to summarise.

Devon: Or, perhaps a really complex piece of Python code that is broken and just needs debugging right now.

Marcus: Right. So, you type your query into your laptop and straight away, a highly capable AI model starts answering you. I mean, it reads the file, it runs the code, and it gives you a flawless answer.

Devon: And the really crucial part here is that not a single byte of your personal data ever actually travelled to a server in the cloud.

Marcus: That is the kicker. And, you know, that is not a pitch for some futuristic device coming in 10 years. That is something a small AI model, just running quietly on your own everyday hardware, can already do today.

Devon: It really represents a complete paradigm shift, I think, in how we think about computing. Because, well, we have spent the last few years operating under this assumption that serious artificial intelligence required massive, distant data centres.

Marcus: Yeah, we pictured those giant warehouses full of humming servers.

Devon: Absolutely.

Marcus: Yeah.

Devon: But the gravity of computation, it is pulling back. It is returning towards the actual devices sitting on your desk, or even just the phone in your pocket.

Marcus: And that is exactly our mission for this deep dive. We are looking at the reality of local AI models. We want to unpack exactly how this massive shift benefits you, the user, through absolute privacy and, well, some truly staggering cost reductions.

Devon: Right, because the savings are immense.

Marcus: They really are. But we are also going to reveal why the real revolution is not actually your laptop defeating the cloud in some sort of head-to-head hardware battle. The real magic is a brilliant new system called hybrid routing.

Devon: Which is, to be fair, completely redefining how computing resources are allocated globally right now.

Marcus: Yeah.

Devon: And we have a fantastic stack of sources to help us navigate this landscape today. We are going to examine a major Stanford-led study on a metric called intelligence per watt.

Marcus: Which is fascinating on its own.

Devon: It really is. And we are looking at that alongside a Deutsche Bank report that outlines how open-source AI is disrupting the market. Plus, we have some fascinating research on large language models operating in local virtual sandboxes.

Marcus: I want to start by establishing just how capable these small, local models have actually become. Because, I mean, I think a lot of people assume that if an AI is not running on a billion-dollar supercomputer, it must be basically useless.

Devon: That is a very common misconception, yeah.

Marcus: But according to the Stanford study, which actually analysed over a million real-world queries, small local models can now successfully handle 88.7% of everyday single-turn chat and reasoning queries.

Devon: And the sources specifically mention models with under 20 billion active parameters there, like Qwen1.5 or Llama 3.1.

Marcus: Right. Before we get into that massive percentage, I want to make sure I am picturing the technology correctly. I am assuming active parameters essentially refers to, you know, the actual number of neural connections firing in my device's memory at any given second.

Devon: That is a really highly accurate way to visualise it, actually. In large language models, a parameter is essentially a variable that the model uses to make decisions and generate text. It is, uh, its the the model's learned knowledge.

Marcus: Okay, so it is the brainpower.

Devon: Exactly. And a few years ago, the frontier cloud models required hundreds of billions, or even over a trillion parameters to sound coherent. But researchers have discovered how to distil that vast knowledge into much smaller, highly efficient packages.

Marcus: So, a 20-billion-parameter model is small enough to fit into the memory of a modern, high-end laptop.

Devon: Yes. Yet, it retains a shocking amount of reasoning capability.

Marcus: The numbers in the study are just staggering when you look at the timeline. In just two years, the share of queries that can be handled locally jumped from roughly 23% to over 71% for the best individual model.

Devon: Yeah, that leap is incredible.

Marcus: In fact, a model called GPT-OSS-120B handles nearly three-quarters of these queries entirely on its own. The Deutsche Bank report refers to this sudden leap in capability as a second DeepSeek moment.

Devon: Which is such a powerful framing, I think.

Marcus: It really is. It means open-weight models, meaning models where you can literally download the underlying architecture to your hard drive and own it forever, are absolutely rattling the market.

Devon: Because from a user's perspective, this provides immense utility. You get offline capability, you get zero latency because, well, you are not waiting for a network packet to travel to a server farm in another country.

Marcus: Yeah, you are not staring at that little loading icon.

Devon: Spot on. And most importantly, you get absolute privacy. Your data never leaves your physical premises.

Marcus: That data control is fundamentally changing the risk calculus for everyone, isn't it? From individual users to massive corporations.

Devon: Oh, absolutely. Think about a highly regulated organisation like a hospital or a corporate law firm.

Marcus: Or even just, you know, an individual who does not want their private tax returns ingested by a cloud provider's training algorithm.

Devon: Exactly. When you deploy the model locally, those privacy concerns just evaporate entirely.

Marcus: To put this 88% success rate in perspective, running local AI right now is essentially like having a highly capable general practitioner living in your own home.

Devon: I like that analogy.

Marcus: Yeah, for 90% of your ailments, you know, a common cold, a minor sprain, general daily advice, that GP handles it perfectly, immediately, and privately.

Devon: Right, you only need to travel to the expensive specialist hospital, which represents the massive cloud model, for the rarest, most complex surgeries.

Marcus: The data from the Stanford researchers perfectly aligns with that GP analogy, actually. They broke down that success rate by domain, which gives us a much clearer picture of where local AI shines.

Devon: What did they find on the domain breakdown?

Marcus: So, for creative tasks, social interactions, and humanities questions, local models have a success rate of over 93%.

Devon: Wow. 93?

Marcus: Yeah, they are incredibly proficient at drafting emails, brainstorming marketing ideas, or summarising articles.

Devon: So, that is the equivalent of the GP prescribing rest and fluids. It is routine work.

Marcus: Exactly right. But when you look at technical fields, disciplines like complex architecture, software engineering, or graduate-level physical sciences, that local coverage drops significantly.

Devon: Down to what sort of level?

Marcus: Down to around 60%. The hardest reasoning problems remain unsolved by the small models.

Devon: So, the hard problems stay hard. You still need to travel to the specialist hospital for the complex surgeries.

Marcus: Exactly. Okay, so my local GP is fantastic for routine daily tasks. But if I ask my laptop to read a massive 100-page financial report to pull out a few specific data points, usually, I would have to copy and paste that entire wall of text into a chat window.

Devon: Which is a nightmare.

Marcus: It is. That would absolutely melt my laptop's memory. How is a local model actually reading these massive files offline without crashing my machine? Because if it cannot handle my personal files, the privacy aspect doesn't really matter.

Devon: This is where the research on the LLM sandbox environment becomes so crucial. A sandbox is essentially a minimal, isolated virtual computer running inside your actual computer.

Marcus: Okay, so it is a secure, walled garden.

Devon: Yes, precisely. When you give a local AI model access to this sandbox, it is no longer just predicting the next word in a chat window like a standard chatbot. It is given fundamental computing tools.

Marcus: Like what kind of tools?

Devon: Well, terminal access to execute commands, file management systems to read and write data, and a Python interpreter to write and run its own code on the fly.

Marcus: Hold on, so it is not just talking to me, it is actively working on my machine.

Devon: It is, yeah.

Marcus: I want to break down exactly how this functions because the source material gives a brilliant example of a long-context task. A user asks the model to extract specific information from massive industry reports containing hundreds of thousands of words.

Devon: And in a traditional cloud setup, you would have to feed that entire document into the prompt.

Marcus: Right. Which takes forever and costs a fortune. But in the sandbox, instead of trying to read the whole thing at once, the model autonomously uses shell commands, specifically a command like grep, to search the files.

Devon: That is fascinating.

Marcus: It locates the exact lines containing the relevant keywords, and then it writes a custom Python script to systematically extract only the required data.

Devon: Okay, that is brilliant.

Marcus: So, let me guess the mechanism here. And because the AI is using these search commands in the sandbox, it acts sort of like a librarian.

Devon: A librarian.

Marcus: Yeah, instead of bringing the entire encyclopaedia to your desk, which would eat up massive amounts of memory, it goes into the stacks, finds the one exact sentence you need, and only brings back that sentence.

Devon: That is a phenomenal analogy. It acts exactly like a targeted librarian.

Marcus: Yeah.

Devon: And the efficiency gains here are what make local AI truly viable.

Marcus: Because of the token costs.

Devon: Exactly. In standard cloud AI, every single word you paste into a prompt consumes a token. Cloud providers charge you a literal financial fee for every single token they process.

Marcus: And when you are dealing with 100-page PDFs, those tokens add up incredibly fast. Your bill just skyrockets.

Devon: But the research shows that by keeping those massive documents in a local sandbox environment and using that librarian method, token consumption falls drastically. For long-context tasks, they saw token usage drop by up to eight times.

Marcus: Eight times, that is massive.

Devon: Because the model is only reading the specific chunks of the file it actually needs to answer the question. You are not paying the toll to process the other 99 pages of completely irrelevant text.

Marcus: Right, for a cost-conscious firm, dodging those per-token API charges is a massive financial win. You are no longer paying a cloud provider a tax just to read your own internal documents.

Devon: And beyond money, it brings us back to the privacy aspect. For a regulated organisation, this local sandbox environment means you can process strictly confidential data on-site.

Marcus: The documents never leave your physical servers.

Devon: Exactly. It completely neutralises the risk of a cloud leak or a data interception during transmission. You get the intelligence of a modern AI, but with the security of a locked, offline filing cabinet.

Marcus: It fundamentally changes the economics and the security posture of enterprise AI. You are reducing your token usage by a factor of eight, and your data exposure by a factor of infinity.

Devon: Because the exposure is zero.

Marcus: Spot on. Now, I have to say, this sounds like a total utopia. We have huge utility, massive cost savings, total privacy. But, uh, I want to step back and look at the physical reality here.

Devon: Okay, let us do it.

Marcus: Processing all of this data locally requires energy. Our laptops and phones have batteries that drain, they have thermal limits, I mean, they get hot. The Stanford study introduced a metric to measure this called intelligence per watt.

Devon: Which is a brilliant metric.

Marcus: I am guessing that is exactly what it sounds like on the tin. It is a measure of how smart the model's answer is, divided by how hard my laptop's battery has to work to generate that answer.

Devon: You hit the nail on the head. Intelligence per watt, or IPW, is a vital way to measure the true physical cost of AI. It divides the task accuracy, you know, did the model actually get the right answer, by the unit of power required to generate it.

Marcus: Okay, so it tells you how much useful, actionable intelligence you are getting for the electricity you are actually burning.

Devon: Exactly. And the sources show that this metric has improved at a breakneck pace. Over just two years, from 2023 to 2025, intelligence per watt improved by 5.3 times.

Marcus: Which is a huge leap. It means we are getting over five times more intelligence for the exact same amount of battery power.

Devon: The researchers dug into that number and found that the 5.3 times improvement was driven by two distinct factors.

Marcus: Mhm.

Devon: There was a 3.1 times advance in model architecture.

Marcus: Meaning the software got smarter and less sloppy.

Devon: Right, requiring fewer computational steps to reach a conclusion.

Marcus: Right.

Devon: And then there was a 1.7 times advance in the physical hardware accelerators themselves.

Marcus: Wait, I need to stop you there and push back on this entire premise. Because no matter how much the software improves, my laptop is still a generalist device.

Devon: True.

Marcus: The Stanford study explicitly states that massive cloud hardware, like the NVIDIA B200 chip, still holds a massive 1.4 to 7.4 times efficiency edge per query over local chips, like the Apple M4 Max in a high-end laptop.

Devon: Yes, they do point that out.

Marcus: The giant data centres have chips built only for AI. So aren't we just wasting battery life and grid energy trying to force our laptops to do this heavy lifting, when a data centre could do it much more efficiently?

Devon: It is a really vital pushback, and you are highlighting the core tension in the entire local AI movement here. You are absolutely right that cloud hardware is more efficient in a vacuum.

Marcus: Because they have specialised gear.

Devon: Exactly. Enterprise accelerators in a data centre have dedicated tensor processing units and high-bandwidth memory that are solely optimised for AI maths. They don't have to power a screen, they don't have to run background apps.

Marcus: Right.

Devon: Local hardware, like your laptop, relies on a unified memory architecture. The AI has to share the memory and the processing power with your web browser, your operating system, your music player. It is fundamentally a compromise.

Marcus: So, my laptop is basically a decathlete. It is pretty good at 10 different events. But the cloud server is an Olympic sprinter. It is always going to run the 100-metre dash faster and more efficiently because that is the absolute only thing it has been built to do.

Devon: That is the perfect way to look at it. The sprinter wins the sprint. However, the Stanford study reveals a fascinating exception to this rule when we look away from laptops and focus on smartphone edge devices.

Marcus: Oh, really?

Devon: Yeah, they tested the iPhone 16 Pro, which uses a highly optimised neural processing unit, or NPU. Because that entire smartphone system operates at a very low power envelope, around 12 watts, that chip actually achieves seven times higher intelligence per watt than a massive workstation GPU running the exact same model.

Marcus: To put 12 watts in perspective for you listening, an old standard incandescent light bulb used 60 watts. Your phone is generating near-human reasoning on a fraction of the power it takes to light up a closet.

Devon: Which proves that there is a distinct place for ultra-local efficiency. If a query is simple enough to fit into the memory of a smartphone, routing it to that low-power NPU is incredibly energy efficient.

Marcus: So, we have this weird split. The local device is amazing for privacy, it saves token costs, and it is highly efficient for easy tasks. But the cloud sprinter is mathematically more efficient and capable for the massive, heavy-lifting tasks. Who is making the decision on where my query goes? Because I certainly do not want a drop-down menu asking me to route my own traffic every single time I type a prompt.

Devon: Nobody wants that. And this is the breakthrough that ties everything together. The solution the sources point to is one central mechanism, the intelligent router.

Marcus: Okay, how does that work?

Devon: The practical future of AI is not about choosing exclusively between a local device or a cloud server. It is about hybrid routing. The system itself decides, query by query, in a matter of milliseconds, where the computation should actually happen.

Marcus: And it does not need to run the prompt to know, it just looks at the complexity of your request. The study simulated this using a router that is only 80% accurate.

Devon: Meaning 80% of the time, it correctly guesses whether a small local model can handle the prompt, or if it needs to be sent to the massive cloud model.

Marcus: And the results are wild. By simply directing traffic, sending the easy questions to the local laptop and shipping only the hard ones to the frontier cloud models, this hybrid setup yields a 64.3% reduction in energy, a 61.8% reduction in compute, and a 59% reduction in financial cost.

Devon: And those savings are compared to a baseline where every single prompt goes to the cloud. You are achieving an average of 60 to 80% cost reductions across the board, just by sorting the mail before you open it.

Marcus: This is the absolute magic of the whole setup. I mean, the revolution isn't your laptop beating a billion-dollar data centre in a bench press competition. The magic is the router quietly dodging the tollbooth 90% of the time.

Devon: Spot on. It keeps your data private for the easy stuff, avoids the API fees, and only pays the toll for the complex surgeries.

Marcus: If we look at the broader industry, this hybrid routing is exactly why major technology platforms are rapidly building local-first architectures. The source material notes that Apple unveiled their next generation of Apple Intelligence with this exact philosophy.

Devon: Because they have to, right.

Marcus: Exactly. The architecture is designed to process as much as possible on-device to protect privacy and save server costs. But for complex tasks, it relies on server models, complete with daily usage limits.

Devon: Because if every single query from every user globally went to a cloud server, the infrastructure would literally melt. The cloud companies would go bankrupt trying to serve two billion phones asking for daily email summaries.

Marcus: Absolutely. By offloading the easy 90% to local devices, the cloud is freed up to tackle the highly technical reasoning tasks that actually require massive compute.

Devon: It is a collaborative, sustainable ecosystem.

Marcus: It is, yeah. Okay, I buy the logic. We save money, we save energy, we keep our data private. But, uh, I want us to take a step back and look for the catch.

Devon: There is always a catch.

Marcus: There really is. Whenever we see massive cost reductions and utility booms like this, someone, somewhere, is paying a price. When something gets cheaper and easier, we don't just use it the same amount, we use it way more.

Devon: You are pointing directly to the macro effects of this shift, which the Stanford study explicitly warns about, actually. They highlight a phenomenon known in economics as the Jevons paradox.

Marcus: Which is the rebound effect. Historically, when we made coal engines more efficient in the Industrial Revolution, we didn't use less coal, we just built way more engines and ended up burning more coal overall.

Devon: Precisely. The Jevons paradox occurs when a technology becomes more efficient, lowering its cost, which in turn causes people to use it so much more that total resource consumption actually goes up.

Marcus: So, if we apply that to local AI, because it suddenly costs me zero cloud API fees to have an AI categorise my emails, I might decide to set up a local agent to constantly read, analyse, and sort every single message that hits my inbox.

Devon: 24 hours a day.

Marcus: Yeah, 24 hours a day.

Devon: Exactly. As per-query efficiency rises, users will inevitably deploy always-on local agents. And while each individual query is highly efficient, the sheer volume of constant, background computation could raise the total energy use across the board.

Marcus: That makes a lot of sense.

Devon: The Deutsche Bank report rightly points out that downloadable does not mean free.

Marcus: So, when you run AI locally, you are essentially allowing your personal laptop to moonlight as a data centre, and data centres need power.

Devon: The energy bill does not disappear, it is simply transferred. Local inference shifts the physical energy cost directly onto consumer power grids.

Marcus: So, I am footing the bill.

Devon: You are. You are no longer paying the cloud provider's API fee, but you will absolutely be paying your local utility company when your laptop's fans are spinning at maximum speed all night running background tasks.

Marcus: It is the ultimate hidden subscription fee. But the physical power bill isn't the only thing we inherit, is it? We are also taking on all the administrative headaches.

Devon: That is the other major caveat that users really need to understand. When you rely on a managed cloud service, that massive provider is responsible for cybersecurity.

Marcus: They handle the messy stuff.

Devon: Yeah, they monitor the model for hallucinations, they apply safety updates, and they ensure uptime. When you download an open-weight model and run it in a local sandbox, the user is now saddled with the maintenance.

Marcus: You essentially become your own IT department.

Devon: Yes. The security risks of unmonitored local models are entirely transferred to you. What if a downloaded open-weight model has a hidden backdoor?

Marcus: Ah, that is a terrifying thought.

Devon: What if it executes a malicious script in your sandbox because it misunderstood a document? Or if it hallucinates terrible financial advice based on your tax returns?

Marcus: And you can't blame anyone else.

Devon: No, there is no cloud provider to hold accountable. The liability rests entirely on your shoulders.

Marcus: It is a profound shift in power, but also a profound shift in responsibility. Let us synthesise everything we have covered in this deep dive. The near-term future of artificial intelligence is not a dramatic, winner-takes-all battle of local devices versus the cloud.

Devon: No, it is not.

Marcus: The future is the intelligent router. It is that quiet traffic controller sitting on your device, giving you the privacy, the zero latency, and the immense cost savings of local AI for 90% of your daily tasks.

Devon: And reserving those massive, energy-intensive cloud servers strictly for the heavy lifting. It is a highly efficient, symbiotic system that delivers incredible utility directly to you.

Marcus: While drastically reducing our reliance on centralised API tolls.

Devon: Spot on.

Marcus: But, as you embrace this new era and you transform your personal devices into local AI accelerators to save on those cloud subscription fees, I want to leave you with a completely different consequence to ponder.

Devon: What is that?

Marcus: If we are running our laptops and our phones at maximum thermal capacity all day, every day, to act as our own personal data centres, what happens to the physical lifespan of that hardware? Are we about to enter a new era of massive e-waste where we end up frying our expensive laptop batteries every eight months, just to avoid paying a £20 monthly cloud subscription?