Marcus: So, we have somehow built a world where a machine can write an entire reference work, ostensibly to finally deliver perfectly unbiased truth, then grade its own homework, and uh proceed to mark itself down. Devon: It is, yeah. It's a genuinely remarkable sequence of events, and that's exactly what we're unpacking in this deep dive. We're looking at a large-scale audit out of Ghent University. It's titled Grokipedia versus Wikipedia: An LLM-Based Audit of Political Neutrality Along Ideologies. And the overarching mission here is to figure out what actually happens when artificial intelligence takes over the writing of our core public knowledge bases. Marcus: Because, I mean, the stated goal was always to fix human bias, right? To create this, you know, flawless encyclopedia without the editorial slant that just inevitably creeps in from human writers. Devon: Exactly. I mean, the the common story you hear across the tech sector is that an AI-written encyclopedia simply strips out human bias and hands you cold, neutral, objective facts. But the reality, at least according to this comprehensive audit, is that it doesn't remove the tilt at all. It just swaps one ideological tilt for another. Marcus: Wow. Okay, but before we get into the whole, you know, AI rating its own bias thing, we need to look at what xAI was actually trying to build when they launched this platform back in October 2025. Because the scale was just massive from day one. Devon: It was. When xAI launched Grokipedia, it was explicitly pitched as this neutral alternative to Wikipedia, entirely generated by the Grok model. And by the time this audit was conducted on their version 0.2, it already included over a million articles. Marcus: Which is a staggering amount of text to generate. Devon: It is. Now, it's worth noting that prior work has reported Grokipedia to be essentially a synthetic derivative of Wikipedia. Marcus: Right, meaning what, exactly? Devon: It means it relies heavily on Wikipedia's foundational structure and topics to know what to write about, but researchers have noted it tends to have uh weaker citation quality. Marcus: Okay, but the core marketing claim, the whole reason for its existence, was absolute neutrality. So, if you're a researcher at Ghent University, how do you actually audit that neutrality without just, you know, introducing your own human bias into the grading process? Devon: Well, the researchers approached this with a rather brilliant mechanism. They took 1,394 pairs of articles about members of government. Marcus: Okay. Devon: And this sample was drawn from a highly structured dataset called WhoGov, specifically covering politicians who served between 2016 and 2023. They then had these article pairs evaluated by four different AI models. Marcus: So, they didn't use human graders at all. Devon: No, they used what we call an LLM judge, which is, you know, an important piece of jargon to define here. It's simply an AI model that's prompted to act as an evaluator. Marcus: Right. Devon: It reads the text and scores it against a highly specific rubric. In this case, the rubric was based on standard encyclopedic principles of neutrality. So, it's looking at things like impartial tone, proportional representation of viewpoints, that sort of thing. Marcus: And who were the four judges they chose for this panel? Cuz I imagine that matters. Devon: It matters a lot. They used Grok, Claude, Mistral, and DeepSeek. That is a very deliberate selection. Marcus: Because they're all so different. Devon: Exactly. You have four very different models with different underlying architectures, completely different training philosophies, and different regional origins. Marcus: And the punchline finding of this entire study, the thing that literally brings us here to talk about it, is that every single one of those four judges rated Grokipedia as containing more bias than Wikipedia. Devon: Every single one. Marcus: It's just, I mean, it is quite literally the comedy of opening a restaurant, marketing it as having the most objective, perfect menu in the world, hiring four famous food critics to review it, and, you know, your own mother gives it the worst rating. Devon: It's a stunning outcome for a project explicitly designed to eliminate bias. Even the model that wrote the encyclopedia says it's biased. Marcus: But I do want to push on the mechanics of this a bit, because just saying something is biased is quite abstract. How exactly does this bias manifest in the text? Like, what does an AI slant actually look like when you're just reading an article about some local politician? Devon: Well, to understand the shape of it, we have to look at how the researchers actually mapped the bias. They used an expert-coded dataset called V-Party, and this dataset scores politicians along nine specific ideological dimensions. So, we're talking about their explicit stances on immigration, LGBT equality, women's labour equality, and economic policy. Marcus: Right, the core political issues. Devon: Exactly. And what the audit reveals is that neither encyclopedia is genuinely neutral. Both portray politicians favourably overall, but they lean in completely opposite ways. Marcus: Opposite leans. Okay. So, they're effectively looking at the exact same politician with the exact same voting record and just choosing completely different lighting and different angles to present them to the reader. Devon: That's a great way to put it. And the lighting Wikipedia chooses heavily flatters socially liberal politicians, whereas Grokipedia actively penalises socially liberal politicians and strongly favours the economically right wing. Marcus: Hold on. What does a penalty actually look like in an encyclopedia entry? Give me a sense of how that reads on the actual page. Devon: To give you a sort of generalised idea of how this looks in practice, where Wikipedia might say, "A politician championed marriage equality legislation," a biased Grokipedia entry might frame it as, "The politician was embroiled in controversies regarding traditional marriage laws." Marcus: Ah, I see. Devon: It's the exact same historical event, just viewed through a totally different, much more critical lens. The language used to describe them is just less neutral. It focuses more heavily on friction or opposition compared to a politician whose ideology aligns with the system's inherent slant. Marcus: That makes total sense. So, this is not necessarily making up fake facts. It's just, you know, shifting the framing and the adjectives to cast a shadow. Devon: Precisely. And this brings us to the starkest data point in the entire audit. A politician's economic left-right position is the single strongest predictor of how favourably they are portrayed across the board. Marcus: Out of all nine dimensions in that V-Party dataset. Devon: Yes, out of all of them. But here is the crucial part. That effect is driven entirely by Grokipedia. Yes. In Grokipedia, the coefficient for being economically right-wing is a strong, mathematically proven predictor of receiving a favourable portrayal. But in Wikipedia, the effect of your economic ideology on how favourably you're portrayed is statistically insignificant. Marcus: Let me just synthesise that for a second, cuz that is wild. On Wikipedia, your economic policy, whether you want to raise taxes or cut regulations, doesn't really change the tone of your biography. But on Grokipedia, it absolutely does. It is actively reading your economic stance and adjusting the warmth of your biography accordingly. Devon: It is. And this highlights a broader finding about the sheer weight of ideology in these systems. Ideology shapes Grokipedia's coverage far more heavily than Wikipedia's. Marcus: How much more? Devon: Well, in statistical terms, a politician's ideological traits explain 22% of the variance in Grokipedia's neutrality scores. For Wikipedia, those exact same traits explain just about 6% of the variance. Marcus: Let's ground those numbers, because hearing raw percentages can get a bit textbook heavy. But think about it this way. On Wikipedia, a politician's personal politics barely move the needle on how neutral their article sounds. It only accounts for about 6% of the tone. Devon: Right. Marcus: But on Grokipedia, their politics drive nearly a quarter of the entire tone of the article, 22%. The numbers prove the bias is driving the vehicle. Devon: It's a massive gulf. And, you know, it's not just economic ideology driving that vehicle. In Grokipedia, social identity dimensions, specifically stances on LGBT rights and women's labour equality, actively attract more negative portrayal. Marcus: Just by supporting them. Devon: Yes. If a politician champions those causes, the AI judges consistently rate their Grokipedia articles as more negatively biased. Marcus: Okay, if the bias is that prominent, driving nearly a quarter of the article's tone, the obvious question is, where does this bias actually live? If Grokipedia was entirely written by an AI, did the AI decide to do this, or was it told to do this? Devon: Now, that is the essential mechanism question the researchers grapple with here. Is this ideological tilt baked deeply into the massive training data that Grok learned from, you know, the billions of web pages and forums it scraped? Or is it a result of the specific prompt xAI used to generate the articles? Marcus: Like a hidden instruction in the system prompt that inadvertently steered the model. Devon: Exactly. Now, I want to state clearly that the authors of the audit flag this as an unconfirmed open question. It is not a settled fact. We know the bias is undeniably there, but the precise architectural origin of it remains unconfirmed. Marcus: Right, but I mean, if the mechanism is an open question, how much weight should we really give to the machine's own verdict on itself? Because earlier, you mentioned that Grok, the exact model family that wrote Grokipedia, was one of the four judges on the panel, and it marked its own work down. Devon: Yes, and the authors read that as a rather damning detail. Grok assigns a higher mean absolute bias rating to Grokipedia than it does to Wikipedia. Marcus: Which means what, structurally? Devon: The structural implication here is quite profound. It implies that the slant is structurally baked into the content itself, and the model recognises its own departure from neutrality. The generative process introduces a bias that the evaluative process can clearly see and flag. Marcus: Wait, let me stop you there. Are we reading way too much into that? Calling it a confession feels like we're treating this math equation like it has a guilty conscience. Devon: How so? The model evaluated its own generated text and found it lacking. Marcus: Because we have to look at the study's data on model variance. The choice of AI judge changes the result enormously here. When DeepSeek evaluated the articles, it rated almost 87% of them as completely neutral. Devon: That's true. Marcus: But when Claude looked at the exact same articles, it rated only 25% of them as neutral. Devon: That is a very wide spread in strictness. Marcus: Yeah, it's a massive spread. It's like asking two completely different teachers to grade the exact same history essay. DeepSeek is the substitute teacher handing out easy A's to everyone, while Claude is the strict professor marking you down for every slight phrasing choice. If one judge says everything is fine and another judge says everything is heavily biased, then Grok grading its own homework might just be a quirk of its specific evaluator tuning. Devon: So, you're suggesting it's just a statistical output. Marcus: Exactly. That output is highly dependent on the prompt and the temperature setting of the model on any given day, which is basically the dial that engineers use to control how creative or predictably strict the AI's responses are. It might just be how Grok happens to be calibrated when you ask it to grade a paper, rather than some profound structural confession of guilt about how it wrote the text in the first place. Devon: I see your point on the variance in strictness. DeepSeek is clearly a very lenient grader, and Claude is very strict. But you cannot ignore the relative scoring here. Marcus: What do you mean? Devon: Even if Grok has a specific tuning quirk based on its temperature settings or system prompt, it still consistently applied that tuning to both encyclopedias. And under that consistent lens, it found Grokipedia to be less neutral than Wikipedia. That structural implication holds, even if the absolute numbers shift based on which AI is sitting in the judge's chair. Marcus: Okay, I still think treating it as a sentient machine recognising its own flaws anthropomorphises the whole process way too much. But I will concede the point on relative scoring. It graded both, and it preferred the human work. Devon: We might just have to agree to disagree on how damning Grok's specific verdict on itself actually is. But I think we can unite on the practical conclusion that stems from this debate. Marcus: Which is that single-model AI audits are entirely unreliable. Devon: Absolutely. If you only used Claude to audit the internet, you would think the digital world is completely awash in bias and nothing is safe. Marcus: Yeah. Devon: And if you only used DeepSeek, you'd think everything is perfectly fine and objective. This extreme variance proves that a diverse panel of AI judges is absolutely necessary to get anywhere close to the truth of what these models are actually doing. Marcus: So, having established that this ideological tilt is real, that it leans in the completely opposite direction of Wikipedia, and that auditing it is incredibly complex, we have to look at why this actually matters to you, the person searching for facts online. Who actually pays the price for this flawed digital infrastructure? Devon: Well, it's a fundamental question of public knowledge. Encyclopedias are the foundation of how we establish baseline facts in a society. Marcus: Exactly. They shape political opinion and, by extension, democratic discourse. The core worry raised by the authors here is deeply cultural. The encyclopedia you happen to read could subtly nudge your view of a politician based purely on their ideology, rather than their actual voting record or their legislative actions. Devon: You search for a member of government to understand what they've actually done, and instead, you're handed a stylised framing of what they believe. Marcus: Right. And the horrifying and amusing edge to all of this, the real second-order sting, is that this does not stay contained on one platform. We are already seeing the wider ecosystem cannibalise this content. Devon: Yes, the propagation issue. Marcus: Exactly. These AI-written articles on Grokipedia are already being cited by other AI systems. Reports show that ChatGPT models have used Grokipedia as a source in their own outputs. Devon: Which means a slant generated in one place quietly propagates across the entire internet. Marcus: It is the ultimate absurdity of this technical setup. We have built a world where a machine writes a reference work, grades its own homework, marks itself down as biased, and the rest of the tech industry just blindly scrapes it and builds on top of it anyway. Devon: And the bias becomes foundational. Future models will likely scrape and train on this very content, creating a feedback loop where the AI slant just gets louder and louder. Marcus: Yeah, it just gets baked into the cement. Devon: And that leads to a very grounded, concrete takeaway for anyone navigating this space. When you are reading an AI-generated reference work, the ideology is just as structural as the facts. The promise of neutral AI is, at the time of this release, currently just a marketing line. In practice, you are always reading through a lens. Marcus: Which leaves you with quite a philosophical puzzle. If Wikipedia has a liberal lean driven by the collective biases of human editors, and Grokipedia has a right-wing lean driven by opaque algorithms, and future AIs are just training on both of them, is the future of neutrality just the mathematical average of our overlapping biases? What happens to truth when AI systems stop asking humans what happened and just start citing each other?