A new analysis of major AI chatbots found they frequently spread political falsehoods and backed claims with sources that don’t exist or fail to support the answers given.
According to reporting from Fox News, the study by the nonprofit research group Just Facts tested paid versions of ChatGPT, Google Gemini, Grok, and Claude on 100 carefully designed multiple-choice questions covering immigration, abortion, climate change, elections, crime, gun control, COVID-19, and related topics.
The questions were structured to probe falsehoods commonly circulated on both the left and the right, and each chatbot was required to cite sources for its responses.
Across the 400 total answers (100 questions × four systems), the chatbots collectively produced 74 incorrect responses. More striking was the quality of their citations.
The models offered 419 sources in total. Of those, 104 pointed to webpages that did not exist (including 86 with no trace in the Internet Archive or Google search results), and 77 sources did not actually address the question they were meant to support.
Additional problems included unrelated material, sources that said the opposite of the claim, and demonstrably false references.
Overall, only about 46% of the cited sources proved both real and valid. Validity rates varied by model: ChatGPT at 57%, Gemini at 49%, Claude at 44%, and Grok at 32%.
Accuracy also showed a directional pattern. ChatGPT correctly handled 94% of questions designed to surface right-leaning falsehoods but only 75% of those aimed at left-leaning ones.
Gemini scored 91% versus 76%, and Claude 91% versus 81%. Grok reversed the trend, scoring 73% on right-leaning probes and 84% on left-leaning ones.
Just Facts described ChatGPT and Gemini as performing roughly like C students on the left-leaning set of questions, Claude like a B student, while Grok showed the opposite strength and weakness.
Only ChatGPT’s left-right gap reached statistical significance at the 95% confidence level.
James D. Agresti, president of Just Facts, told Fox News Digital the work goes beyond measuring political tilt.
Prior research has repeatedly found left-leaning tendencies in major models, he noted, but bias does not automatically equal factual error.
This study specifically tested whether the systems were getting the facts wrong.
“The thing that really jumped out at me when I started going through the sources they provided is that roughly half of the sources were illegitimate,” Agresti said.
He compared the overall performance to having “a B student in your pocket” rather than the Ph.D.-level expertise some industry leaders have claimed.
The study’s authors added an important caveat: the questions were worded with high precision to eliminate ambiguity about the correct answers.
That clarity may have given the models clearer paths to accurate replies than users typically provide in everyday, more open-ended queries, where performance could be worse.
Agresti’s practical advice was straightforward: do not treat the systems as authoritative. “Don’t trust, verify.” Users should check that cited sources actually exist and support the claims made.
Fox News also reported responses from the companies. Google stated Gemini is designed for neutral, accurate answers and that it could not fully replicate the study’s results.
Anthropic emphasized its testing for political neutrality and noted that the multiple-choice format does not reflect how people normally use Claude.
OpenAI said ChatGPT is built to be objective by default and to surface sources so users can evaluate them. xAI did not respond to a request for comment.
The findings underscore a broader documented issue with large language models: their tendency to generate confident but unsupported or fabricated references, a problem previously observed in technical and biomedical contexts as well.
While the systems remain powerful tools, the Just Facts results reinforce the need for independent verification on politically charged or high-stakes topics.
