Why You Can’t Fact-Check AI

ChatGPT Logo

Yesterday, I wrote a post analyzing how six of the most popular AI chatbot platforms handle citation.

The findings were, to put it mildly, not encouraging. The standard way that AI chatbots cite their sources mimics the way academic papers cite theirs. Using a combination of in-text citations and a reference list. 

But, while this works well for academic papers, which often exist as references to physical works, it doesn’t work well for a purely digital product. It does little to support the authors whose work was used and is less intuitive for readers.

However, as I noted then, it is likely by design. AI companies are pushing to make their products as “sticky” as possible. One of the ways to do that is by making it difficult or unappealing for users to go elsewhere. Even if that elsewhere might be better for the user.

But the original plan was to test each AI with three different prompts. The first two were prompts designed to have the chatbot return news stories, and thus generate a large number of citations. The third prompt was designed to mimic a more generic research question, something a student or curious person might ask.

However, that third prompt was not included in the analysis. The reason is simple: it didn’t generate many citations. In fact, with several chatbots, it didn’t generate a single one.

As someone who doesn’t use chatbots (other than for tests like these), this was surprising. But this raises an obvious question: How is someone supposed to fact-check a chatbot if the chatbot can’t be bothered to cite its sources?

The answer is that you can’t. At least not in a way that doesn’t eliminate any of the bot’s usefulness.

When AI Doesn’t Show its Work

In the initial test, the third prompt was the question: “How did the Enlightenment period change the way we view authorship and creativity?”

I chose this topic because it’s one that I’ve been researching off and on for a while, especially when comparing it to Romanticism. I’ve amassed a large number of sources and have been debating what to do with the information. That said, this was a topic that I was already familiar with, so it made sense to test the chatbots with it.

Of the six AI chatbots tested (Gemini, ChatGPT, Copilot, Claude, Perplexity, and DeepSeek), only two (Copilot and Perplexity) returned any citations. All of the others did not return any citations or external links.

Gemini did offer a tool to “Double Check” the response. This produced links to supporting and contradictory sources, but those results were based on Google search results and were not an indication of how the chatbot reached its conclusions.

In short, it wasn’t citing where the AI got its information from. Instead, it was simply validating the information.

This creates a very serious problem. AI, as we all know, can and does make mistakes. Five of the chatbots had some (tiny) warning about that below their responses (Perplexity didn’t).

But if AI is so flawed that we are not supposed to trust it, then how can we fact-check it if it doesn’t cite its sources?

The Importance of Citing Sources

As we discussed in 2017, citing sources is important for multiple reasons. Though giving credit where it’s due gets the lion’s share of the attention, it’s also useful for strengthening the credibility of your arguments and showing your due diligence.

As such, fact-checking a human-written work is relatively easy as long as it is properly cited. You can examine the sources, see if they are credible, and if they were accurately represented. You can then look for contradictory sources to see if the author missed something in their research.

Fact checking without sources is much more difficult. You have to examine every single claim independently and start from scratch. ChatGPT, for example, gave me a 500-word response to the query. That includes dozens of different claims, some of which are very vague. It is much easier to verify a credible source than to verify an ambiguous claim after the fact.

Even Gemini’s “Double Check” tool was not particularly useful. It only checked a fraction of the claims, and many of the sources it provided were not credible by themselves. This included encyclopedias and other secondary sources that are not allowed in academic environments.

In short, AI is a black box. We don’t know and can’t know how it reached the response it did. Since we can’t check the path that it took, we can’t know if that path was accurate and correct.

But that may be by design.

The Real Problem

In November 2023, Troy Mikanovich at the University of Southern California published an article entitled AI Writing and Attribution: AI Cannot Cite *Anything.*

The paper illustrates the problem succinctly. AI systems don’t engage directly with their sources, at least not in most cases. Instead, they are trained on a distilled version of the sources, which enables the AI to create responses that sound plausible, but are completely removed from the original context.

So, when I perform a query like my first two, the AI is, in the background, performing a web search and parsing the results. In those cases, it interacts with the source material and can cite them directly. But for a more general query, it doesn’t do that. Instead, it simply strings together the information it has into plausible-sounding responses.

This is why AI has been repeatedly likened to autocomplete (not AI autocomplete, which is a different thing). It’s guessing words, not actually processing and interacting with the information.

In my testing, the result was the same across a series of similar prompts on related and unrelated topics. The chatbots, with the notable exemption of Perplexity, didn’t cite many or any sources. However, even when it does cite sources, it’s unclear how trustworthy that citation and if it’s actually where the information came from, or if the AI is simply working backwards.

In short, AI is not regularly citing its sources, of dubious quality when it does and doesn’t cite it clearly enough to support or acknowledge the original authors. It fails at every level when it comes to citation.

Bottom Line

There are two directions to look at this problem, and neither is good. 

The first is the work of the human authors and creators who AI systems were trained on. As with Google, that relationship could be at least somewhat symbiotic if AI systems properly credited their sources. Much of the time, they don’t and they can’t. 

But, even when they do, they do it so poorly that the user is likely to be unaware it even exists. The proof is in the pudding. As more and more search has moved to AI, traffic to regular, human-written content has dropped significantly. This site is no exception.

The symbiotic web has turned fully parasitic.

That is bad enough, but things don’t get better when looking at it from how useful AI is practically. If AI is so flawed and unreliable that it requires warnings, then one would think it should provide the tools to mitigate those issues. However, it doesn’t and, in many cases, it can’t. 

AI is famous for its hallucinations and the problem is only getting worse. However, when you can’t easily fact-check it, there’s no way to mitigate the danger without spending more time and energy than AI would save. So it becomes a matter of either accepting the risk of AI’s mistakes or not using it at all.

This is a big part of why I think that, if AI is going to truly prove its worth, it’s going to be in spaces like AI autocomplete, where citation isn’t important as the human is still in control. Once the computer takes the reins, you either have to trust it completely or spend more time fixing its mistakes.

There is no middle ground that I can see when it comes to generative AI.

Want to Reuse or Republish this Content?

If you want to feature this article in your site, classroom or elsewhere, just let us know! We usually grant permission within 24 hours.

Click Here to Get Permission for Free