AI Conference Overrun with AI Peer Review

The Fourteenth International Conference on Learning Representations (ICLR) is slated to be held in April 2026 in Rio de Janeiro, Brazil. Known as one of the major AI conferences, it’s expected to attract researchers from around the world and has already received nearly 20,000 paper submissions.
However, even before registration has opened, the conference has been beset by controversy.
According to an article by Miryam Naddaf at Nature, one of the researchers who submitted a paper to the conference, Graham Neubig at Carnegie Mellon University, became suspicious that some of the peer reviews he received were AI-generated.
After publishing a call for help on social media, the AI-detection company Pangram (previous coverage) stepped up and analyzed all 19,490 studies and 75,800 peer reviews submitted to the conference. They found that 21% of the submitted peer reviews were fully AI-generated, and over half contained signs of AI use.
This was in stark contrast to the 2022 conference (before the broad availability of generative AI), where a similar test found that nearly all the peer reviews were labeled as “fully human-written.”
Fortunately, the news was much better for the paper submissions themselves. Very few papers were entirely AI-written, and only 9% had more than 50% AI content. The majority, 61% were either mostly or entirely human-written.
Still, the announcement set off alarm bells for many in the field. ICLR was forced to address the issue in a blog post, where they reiterated their policy on AI usage and stated that they would desk reject any paper with “concrete evidence” of unauthorized and undisclosed AI usage.
Unfortunately for ICLR, this wasn’t the only issue with the peer review process. Shortly after this news broke, the conference was hit by a security breach that exposed the identities of peer reviewers.
To put it mildly, the schadenfreude among AI skeptics has been immense. However, this issue isn’t unique to an AI conference and, as we discussed in October, is a growing problem in all academic fields.
To that end, the most interesting part of this story isn’t the presence of AI peer review, but the other data that Pangram dug up.
Some Surprising Statistics
With so many papers and peer reviews analyzed, Pangram took the time to examine any additional trends it found in the work. To that end, there were a few surprises.
The first was that, at least in this sample, AI peer reviews actually graded AI papers more harshly. When an AI peer review examined a fully human-written paper, it averaged a whole point higher (out of five) than fully AI-generated papers.
In fact, the drop off was almost entirely linear. The more AI that was detected in the paper, the lower the score it received. This runs counter to existing research, which has found that LLMs prefer content written by other LLMs over human-written works.

However, there is a flip side to that. The more AI was present in the peer review, the higher the average score the peer review gave. Fully human-written reviews averaged roughly 4.1 points, while the average fully-AI-generated review was 4.4.

According to Pangram, this was likely due to the sycophantic nature of AI. Chatbots routinely tell users what they want to hear rather than providing a real analysis. Humans are simply more critical than chatbots in most areas, including peer review.
Finally, the company’s analysis found that fully AI-generated peer reviews were significantly longer than human-written ones. Pangram attributed this to AI having a “low information density”, meaning that it uses a lot of words to say very little.
That said, fully human-written peer reviews had a similar word count to those edited by AI, regardless of the extent of the editing.
While all this amounts to just one sample and one datapoint in the growing world of AI research, these are still interesting observations.
What About False Positives?
As with any discussion of using AI detection tools, the issue of false positives must be addressed.
To that end, Pangram touts its false-positive rate and the multiple studies that have found it to be exceedingly accurate at spotting AI-written text.
However, what I find most compelling is the use of the 2022 dataset as a control. That dataset included over 10,000 peer reviews written before the widespread availability of generative AI. Of those, “nope” was marked as fully AI-generated, and only 12 were marked as edited by AI to any degree.
Even if we assume that all findings of AI editing were false positives, that’s still a false positive rate of roughly 0.1%. This is impressive, especially considering that the 2022 dataset will be similar in content (though not quantity) to the 2026 dataset.
While this isn’t perfect and many will not want to base decisions on individual peer reviews on this information, it does lend credence to the overall data.
Still, ICLR has said that they will not be basing any decisions solely on AI detection tools due to the risks of false positives. They are asking those who suspect they received AI-generated reviews to seek out additional evidence, such as hallucinations and other false claims.
But that raises another question. What motivation does a submitter have to report an AI peer review? Especially given that the AI peer review is more likely to be positive?
To be clear, the conference did say that it will be “leveraging recent LLM detection tools to identify papers that potentially have a significant amount of LLM-generated content. These will then be given to ACs for further checking.” However, it is unclear what that threshold is and what other evidence they will be screening for.
Still, from a broad data standpoint, there’s not much reason to mistrust this data. While there is still a possibility that a particular work could be misclassified, such cases are almost certainly the exception rather than the rule.
Bottom Line
It’s important to remember that, despite the obvious schadenfreude, this isn’t an issue unique to ICLR or AI conferences as a whole. This is a problem endemic to academic publishing broadly.
That said, there are two reasons to suspect the problem is more significant here.
First and most obviously, this is an AI conference. Its members will be more familiar with and more comfortable using AI than the general public. So, it’s natural that they would use it more.
However, the second is the workload that the peer reviewers were placed under. According to the Nature article, peer reviewers were assigned five papers each and had to review them in two weeks. That is, to put it mildly, an incredibly high load, especially since this is volunteer work and isn’t compensated.
This is something that Bharath Hariharan, the senior programme chair for ICLR 2026, acknowledged. Some reviewers might have felt they had to turn to AI to complete the task, especially while balancing their other duties.
This is likely illustrated by how much lower the AI usage rate was for papers than for peer reviews. Simply put, the papers were the work that researchers wanted to do, and they could do on their own timetable. Peer reviews were required under a tight deadline.
Balancing the peer review workload will be critical not only to defend against inappropriate AI use but also against low-quality human-written peer reviews.
Want to Reuse or Republish this Content?
If you want to feature this article in your site, classroom or elsewhere, just let us know! We usually grant permission within 24 hours.
