Fats & Sugars

AI Hype & Signal

One number to Measure AI Models: Epoch Capabilities Index

Thursday 13 August 2026

Epoch AI's Capabilities Index promises a single headline number for ranking model progress, but that number is a scaled composite stitched together from more than 50 benchmarks. We unpack what the ECI actually measures, why its comparability is engineered rather than natural, and what it was deliberately built to hide.

In this episode:
- How the ECI collapses over 50 benchmarks into one capability scale
- Why the score rewards models for passing harder tests, not simply more of them
- How domain-specific versions reveal that benchmark choice shapes the result
- Why the scaling deliberately obscures long-term progress trends across domains
- Why the cyber index rests on different foundations and lags behind the general one
- How Google DeepMind funding sits alongside Epoch's claim of full independence

Sources:
Epoch Capabilities Index

AI Hype & Signal is produced with AI, including its two synthetic hosts, and every episode is grounded in cited sources and reviewed before release. Even so, it is intended for general information and discussion, not professional advice, so please check anything important against the original sources linked above before relying on it.

Transcript

The full word-for-word transcript of this episode. Plain text version

Marcus: Imagine staring at a massive, just utterly confusing spreadsheet of AI models. I mean, you have rows upon rows of versions, parameter counts, context windows, release dates.

Devon: Oh, it's a nightmare.

Marcus: Right. And you're desperately scrolling and you're wishing for just one single definitive headline number that proves without a shadow of a doubt that model A is smarter than model B.

Devon: Because that is the ultimate corporate fantasy, isn't it? Yeah. People just want the number so they can justify the multi-million pound procurement budget to their board and, you know, finally go to lunch.

Marcus: Exactly. We all want an IQ score for the machine, and that hunt for the magic number is exactly why we're looking at the Epoch Capabilities Index today, or the ECI.

Devon: ECI, yeah.

Marcus: So our mission for this deep dive is to unpack this single, highly sought-after capability score. We're drawing from a really fascinating stack of sources for this. We've got Epoch AI's own documentation, their interactive data explorers, the underlying methodology paper, which is titled A Rosetta Stone for AI Benchmarks, and a rather revealing discussion thread on the LessWrong forum.

Devon: That thread is a gold mine, honestly.

Marcus: It really is. So the common story out there, the way you'll hear people talk about the ECI on social media or in pitch decks, is that it's an objective, naturally discovered measure of intelligence. People treat it like a pure thermometer for the AI frontier.

Devon: Yeah, but the reality is completely different. The ECI is a scaled composite. I mean, it's stitched together from dozens of distinct benchmarks, and its comparability is heavily engineered rather than organically observed. Right. We're not discovering a fundamental law of physics here. Oh, yeah. We are constructing a highly complex statistical narrative.

Marcus: And this matters for you listening because whether you're an executive trying to pick the right tools for your organisation or you're just someone trying to navigate the endless hype cycle of tech announcements, understanding how this specific number is cooked is the only way to avoid getting burned by it.

Devon: If you don't understand the ingredients, you have no idea what you're actually buying when you look at the final score on the menu.

Marcus: Exactly. Which means we need to look under the hood at the mechanism creating those numbers. So before we can judge if this metric is useful for your tech stack, let's step into the engine room. How does the ECI actually work?

Devon: Right. Let's get into the mechanics.

Marcus: What the ECI does is collapse a massive set of distinct benchmarks into one unified scale. We're talking anywhere between 30, 9 to over 50 different exams.

Devon: And just to give you an idea what those exams look like, they include things like GSM8k, which is essentially a massive dataset of primary school maths word problems.

Marcus: Yeah, the classic maths test.

Devon: Right. Then you have HellaSwag, which tests common sense reasoning by asking the AI to finish a sentence logically.

Marcus: Okay.

Devon: And then there's MMLU, the Massive Multitask Language Understanding test, which is basically a giant multiple-choice exam covering 57 different subjects. It ranges from high school physics to professional law.

Marcus: So they're completely different types of tests measuring entirely different skills.

Devon: Completely different.

Marcus: But the ECI takes all of those and just smashes them together. And the vital part of the methodology here is how it actually scores them. It doesn't just hand out one point for every correct answer.

Devon: No, that would be too simple.

Marcus: Right. It uses a difficulty-based weighting system, which is heavily inspired by something called item response theory.

Devon: Which is a concept borrowed from human psychometrics, right?

Marcus: It is, yeah. Let's break down how that actually works in practice because it's not super intuitive at first glance.

Devon: Please do, because the maths gets thick quite quickly.

Marcus: So in a normal school test, every question is worth one mark, regardless of how hard it is. The ECI algorithm doesn't do that. It looks at the overlap of how different models perform on the same tests to figure out which tests are genuinely difficult.

Devon: Okay, so it's evaluating the test itself, not just the model.

Marcus: Precisely. If a relatively weak, older open-source model easily passes a specific maths test, the algorithm essentially says, "Ah, that test must be quite easy," and mathematically downgrades its value.

Devon: Oh, I see.

Marcus: But if a test is only successfully passed by the absolute elite, state-of-the-art models, it gets weighted as highly difficult. It rewards models for succeeding on the harder tests rather than just handing out points for clearing a massive volume of easy ones.

Devon: Which sounds incredibly clever on the surface. I mean, it does. But we have to ask, who is actually leaning on this highly abstracted number in the real world?

Marcus: Well, a lot of people are.

Devon: Sure, but let's contrast the ECI with another metric discussed heavily in our LessWrong sources. The METR time horizons benchmark.

Marcus: Ah, yeah, the METR benchmark.

Devon: METR is wonderfully concrete. It gives you a result you can actually explain to a chief financial officer. It tells you, "Look, this model can successfully complete software engineering tasks that would take a human expert two hours, with a 50% success probability."

Marcus: It's grounded in real-world labour. You can tie a financial value to it.

Devon: Completely. Now look at the ECI. It's totally abstract. The scaling is anchored quite arbitrarily. Well, at the time of the index's design, the creators decided Claude 3.5 Sonnet would sit at an arbitrary score of 130, and a hypothetical future model like GPT-5 would sit at 150.

Marcus: Right. What is 130? What is 150? It's a manufactured yardstick. It's a relative construct that tells you absolutely nothing about what the machine can actually do when you plug it into your corporate workflow.

Devon: Okay, I understand the appeal of the concrete METR score, but I'm going to push back here. Relying on one highly specific benchmark like that to judge general AI is like measuring an Olympic decathlete solely by their long jump. Well, yes, the long jump is concrete, you can measure it in metres and centimetres, but it tells you absolutely nothing about whether that athlete can throw a javelin or run the hurdles. The decathlon requires general athleticism, just like general AI capability requires broad competence.

Marcus: Okay, but that—

Devon: ECI is trying to capture that breadth.

Marcus: The long jump might not tell me about the javelin, but at least I know exactly how far the athlete jumped. With the ECI, I just know they scored a 130 on, you know, general athletics. Yeah, it's meaningless.

Devon: Fair point. But there's a much bigger catch here that we have to acknowledge with this generalist approach.

Marcus: Which is?

Devon: What happens if a highly specialised model absolutely aces its specific field? Let's say a lab builds an AI that is brilliant at complex genomic sequencing or protein folding, a true Nobel-level specialist.

Marcus: Okay.

Devon: But because it was trained purely on biology, it completely fails at writing poetry or answering historical trivia on the MMLU. Under the ECI methodology, that model gets a surprisingly low score. Are we effectively punishing brilliant specialists just because they aren't mediocre generalists?

Marcus: Well, yes. The index demands a generalist profile. Right. If a tool is perfectly built for your specific organisational need but fails at arbitrary tasks you never intend to use it for, the headline ECI score will make it look inferior. It's the classic trap of the generalist metric.

Devon: And this tension over generalist versus specialist leads us straight into how the ECI handles specific domains. Because Epoch AI knew people wanted to track specific fields.

Marcus: Of course they did.

Devon: So they created domain-specific ECIs, like a software engineering ECI or a maths ECI. But here is where the methodology gets truly bizarre.

Marcus: Bizarre is putting it mildly.

Devon: There are two very surprising design choices buried in the data explorers. First, the ranking of who is the smartest completely shifts based on exactly which benchmarks you decide to feed into the dashboard.

Marcus: Wait, so it's not a stable, absolute truth?

Devon: Not at all. If you load up the software engineering explorer and you swap out a benchmark like SWE-bench Verified, which tests how well an AI solves real-world GitHub issues, and you drop in Frontier Code instead, the entire leaderboard shuffles.

Marcus: Just from swapping one test?

Devon: Yep. The model that was crowned the top coder suddenly falls to second or third place, based entirely on the specific flavour of the test the curator selected.

Marcus: Which immediately shatters the illusion of an objective, unshakable ranking. It's highly sensitive to the test papers.

Devon: It is.

Marcus: But you mentioned two design choices, and I know the second one is the mathematical engineering that really requires a deep dive.

Devon: Yeah, this is the crux of it. The domain-specific scores, so the maths score, the coding score, they are mathematically engineered to rise at the same overall pace as the general index.

Marcus: Let's just pause and make sure we unpack the mechanism of how they actually do that, because the implications of that are absolutely staggering.

Devon: The mechanism is a process of normalisation. Imagine that AI coding capabilities are genuinely exploding. The technology is advancing rapidly, and on a raw graph, the trend line is going nearly vertical.

Marcus: Right, a massive spike in coding ability.

Devon: Exactly. The ECI methodology steps in, applies a mathematical formula, and artificially flattens that vertical line so its slope matches the smoother, slower curve of general AI progress.

Marcus: It forces it down.

Devon: It does. The ECI deliberately cannot be used to compare long-term progress trends across different domains. Because of this mathematical straitjacket, you absolutely cannot use it to prove that AI coding capabilities are advancing faster than AI maths capabilities over a span of several years.

Marcus: Which is horrifying from a strategic perspective. I mean, think about this playing out in reality.

Devon: Okay, how so?

Marcus: You have executives sitting in boardrooms, right? They're looking at these sleek dashboards quoting these exact numbers to justify massive computing budgets. They're looking at the software engineering ECI trend line and saying, "Look at the trajectory of coding AI, we must pivot our entire engineering department right now."

Devon: Right, they see the line going up.

Marcus: But the metric they are citing is explicitly mathematically designed to obscure long-term trends across domains. It forces the domain trend to match the general trend by design.

Devon: Well.

Marcus: It is borderline useless for long-term strategic forecasting and, frankly, completely misleading for the people throwing millions of pounds at it based on those graphs.

Devon: I have to completely disagree with you there. I don't think it's misleading at all. I will vehemently defend the ECI here as an honest, brilliantly constructed tool for a very specific problem.

Marcus: Brilliantly constructed to obscure the truth.

Devon: No. All of this information, this specific normalisation limit, is clearly stated in their methodology paper. They aren't hiding it. They even specify that while it can't show long-term divergence between fields, it will catch a sudden, one-off capability jump within a domain over a short period. They aren't hiding the ball.

Marcus: Being documented in a white paper doesn't excuse a metric that fundamentally fails the intuition of its users. If I give you a speedometer for your car, but I've secretly calibrated it so it always matches the speed of the car driving next to you, writing that down on page 47 of the owner's manual doesn't make it a good speedometer. People look at a trend line for coding and assume it represents coding.

Devon: But they built it to solve the saturation problem. You have to look at why they engineered it this way.

Marcus: Go on.

Devon: Natural benchmarks. The exams we actually write for these models saturate incredibly fast. A test is invented, it's considered impossibly hard, and six months later, every frontier model is scoring 99% on it.

Marcus: Sure, they ace the test.

Devon: The tests break. Epoch AI needed a way to mathematically stitch the history of old, broken tests to the new, harder tests so we can see the full timeline. The metric only becomes dangerous when someone in the hype cycle reads more into it than Epoch AI ever claimed it could do.

Marcus: Oh, I see. You can't blame the thermometer because someone is trying to use it to measure wind speed.

Devon: But they named it the Epoch Capabilities Index. They put it on a beautiful, interactive leaderboard. They know exactly how this will be consumed by the tech press and the markets.

Marcus: They're providing a service.

Devon: They know everyone is desperate for a single horse-race number. To build a metric that actively obscures the very domain trends people are desperately trying to track, and then hide behind the technical documentation when challenged, is a magnificent piece of academic sleight of hand. It's amusing in a very dark, corporate-budget sort of way.

Devon: Look, we're going to have to agree to disagree on their culpability here. I maintain the tool does exactly what it says on the tin, provided you actually bother to read the tin before spending your budget.

Marcus: If anyone reads the tin, yes. But while we're debating the honesty of how these trend lines are stitched together, is there any domain where they don't use this mathematical straitjacket?

Devon: Ah, yes. There is one notable carve-out in the documentation.

Marcus: The cyber exception.

Devon: Exactly. The cyber ECI.

Marcus: The cyber ECI is built on completely different foundations. It incorporates specific cybersecurity benchmarks, like complex capture the flag hacking exercises, that simply aren't found anywhere in the general index.

Devon: Right.

Marcus: And because the underlying data behave so differently, and because those specialised tests are run less frequently, they don't lock it to the general trend in the same way.

Devon: So if you're tracking cybersecurity AI and you're looking at the headline cyber ECI number, you have to understand that it operates on its own cadence. It is inherently less current than the general score.

Marcus: It's slower.

Devon: It's a lagging indicator, which in the security world is vital context.

Marcus: Absolutely. But this complex stitching of data brings us to an even deeper question about the foundations of the index. Looking through the Rosetta Stone methodology paper, we really need to ask, who exactly is funding the creation of these yardsticks?

Devon: This is where the human element gets genuinely uncomfortable.

Marcus: How so?

Devon: Well, we're looking at a metric that is treated as the definitive scorecard for the global AI frontier. It's the metric used to judge the most powerful, resource-intensive models on Earth. Now, look at the underlying methodology paper that birthed it. A Rosetta Stone for AI Benchmarks was funded by Google DeepMind.

Marcus: Right.

Devon: Furthermore, it was co-written with researchers from DeepMind's own AGI Safety and Alignment team.

Marcus: Let me just step in here and balance the scales carefully because we must stick strictly to the facts presented in the sources.

Devon: Of course.

Marcus: Epoch AI states very clearly that the ECI is a fully independent product they claim full rights over. Moreover, the actual code that calculates the index is entirely open source.

Devon: Yes, that is true.

Marcus: It's sitting in a public GitHub repository right now for any data scientist to audit. They aren't running this out of a hidden black box.

Devon: Look, I'm not making an accusation of fraud here. We aren't questioning the integrity of any named researcher on that paper.

Marcus: Good. Just wanted to make that clear.

Devon: But as an observer of the culture of this industry, I have to point out two facts that sit very uncomfortably side by side.

Marcus: Okay.

Devon: On one hand, we have a highly complex, statistically engineered metric used to define who is winning the AI race. On the other hand, the foundational research for that exact metric was funded by a massive corporate lab that is actively spending billions of pounds to win that exact same race.

Marcus: It's the ultimate question of structural independence in a highly concentrated industry.

Devon: Precisely that. Even if the code is public, which is highly commendable, the structural architecture matters.

Marcus: The design choices.

Devon: Exactly. The decisions about how to weight difficulty using item response theory. The choices of which benchmarks are deemed informative versus unreliable. The mechanism of normalising domain trends. All of those subjective design choices were developed with funding and collaboration from a key player in the game.

Marcus: Right.

Devon: When you engineer comparability to this degree, the engineer's perspective shapes the reality. We're criticising the setup. When a metric is this abstracted from the real world and requires this much statistical stitching to function at all, the incentives of the architects become highly relevant context for anyone using the tool.

Marcus: That's a very fair critique of the landscape. I mean, if you're building the track and setting the hurdles and you also own the star runner, it's a dynamic worth noting.

Devon: Absolutely.

Marcus: And that brings us perfectly to the takeaway for you, the listener, who actually has to navigate this landscape. Let's bring this conversation back down to Earth.

Devon: No hype, just reality.

Marcus: If you see a single capability score sitting at the top of a leaderboard, let's say you see a hypothetical GPT-5.6 Sol or Claude 5, and it has a shiny, bold score of 162 next to it.

Devon: Yeah.

Marcus: And a vendor is using that specific number of 162 to pitch you a new enterprise tool, or justify a massive departmental software budget, or crown the new smartest model in the world, stop. Ask questions.

Devon: Because the number 162 means nothing in a vacuum. It doesn't tell you if the model can write an email without hallucinating or process a spreadsheet accurately.

Marcus: It means absolutely nothing until you ask what benchmarks actually went into it. You must remember that the comparison you're looking at is statistically engineered, not naturally observed.

Devon: And perhaps most importantly, recall what that number was explicitly built not to tell you about long-term progress trends across different domains.

Marcus: Exactly. The ranking on the dashboard is the easy part to read. The methodology document is the thing actually worth reading. If you skip the methodology and just chase the highest number on the chart, you are flying completely blind. Which leaves us with a final, rather provocative thought to mull over.

Devon: We've spent this entire deep dive discussing how Epoch AI has to statistically stitch and engineer these metrics just to keep up.

Marcus: Because the natural benchmarks, the actual exams we write for these models, saturate and become too easy entirely too fast.

Devon: The models just keep breaking the tests.

Marcus: They do.

Devon: So what happens to our fundamental understanding of capability when the AI gets so advanced that it starts saturating these highly complex, engineered tests faster than human researchers can invent new math to stitch them together?

Marcus: It's a daunting thought. We started by imagining you staring at a confusing spreadsheet, desperately wishing for one definitive number. But what happens when the spreadsheet simply runs out of columns?