Newsbench found that leading chatbots (ChatGPT, Claude, Grok, Gemini) get factual details wrong in roughly a third of their responses about the news, and about one in seven responses sourced foreign state media like RT or China Daily—even on questions about U.S. domestic politics.
Robbie Goldfarb reveals Newsbench's headline findings: about a third of ~2,500 tested responses per model contained factual errors, and roughly 15% of responses cited foreign state media such as RT or China Daily, even for U.S. domestic political questions. ✦ AI generated
Robbie Goldfarb · The Cognitive Revolution · 2026-06-26 · original ↗
starts at this moment · 49:23
“give us a little kind of intro to what Newsbench was and what was the intent behind the study before before you started off asking those questions and what the findings were.”
we looked at you know per model about 2500 responses each about a third of them um had a factual error in them. what could be a wrong number, a wrong date, a misattributed quote, um a misstated policy... in about 15% like one in seven responses sourced foreign state media.
verbatim transcript · starts at 49:23
49:23findings. I would say um all in all it probably surprised us the level like the the number of issues we found particularly when it came to factual accuracy. Um we looked at you know per model about 2500 responses each about a third of them um had a factual error in them. what could be a wrong number, a wrong date, a misattributed quote, um a misstated policy. Um which I think was I
49:53think we knew that that was an issue, but I think that number was quite a bit higher than than even we expected um going into it. Um the other point I'll call out that I think well two other things that are interesting on on bias um maybe unsurprisingly all the models tended to lean left with the exception of Grock which leaned right um which probably tracks with what I think most
50:18people would would would assume. Um but on sourcing uh in about 15% like one in seven responses sourced foreign state media >> wow >> like RT from Russia um so China Daily and what's really interesting is oftentimes this wasn't even in questions about their home country. So we saw Russia like RT and China daily source in questions about US domestic politics and so that that was I think a
50:54particularly interesting finding that you saw across the board with all with all the models. >> Propaganda pays it seems in the AI era. I mean it does but also you know these it's also I think a business model thing right like the these state sponsored media is generally free and open and they're like sure come scrape our stuff whereas you know New York Times for example you got to you got to have a
51:18deal with them so it's um it they are they've certainly done a good job at opening themselves up to the [laughter] scrapers >> got their LLM TXT uh in good working murder. >> Yeah. >> Yeah. That's that is very fascinating finding. >> Yeah. >> Um Pasha, you have any more like methodology questions on that? I I was going to maybe zoom out a little bit at the end, but anything you want to dig
- ·~1 in 3 responses contained a factual error
- ·Tested ~2,500 responses per model (ChatGPT, Claude, Grok, Gemini)
- ·Errors: wrong numbers, dates, misattributed quotes, misstated policies
- ·~1 in 7 responses cited foreign state media (RT, China Daily)
- ·Even on U.S. domestic politics questions
- ·Sourcing suggests algorithms fail to filter state propaganda