The Medium comments on this post are also on point. Running the same experiment with accepted papers is a good control. Running a similar experiment with reviewers would be interesting, but more obnoxious because they are not being paid.
I would keep a private blacklist (shadow ban) the authors who wasted several hours of a reviewer's time to prove they were not legitimate. The existence of such a list would be problematic, though.
Could the same system we use here be applied? Accepted authors could "vouch" for "dead" papers in case they were "auto-killed"?
This system is broken and providing more evidence that it is broken isn't much of a step towards fixing it.
It strikes me as akin to plagiarism. If the purported author can’t even answer basic questions about the paper, how can they plausibly claim to have written it?
In the case where someone uses AI to write the paper and then deeply familiarizes themself with it, it may go undetected, but then it’s also presumably less of an issue since they have actually read it carefully and closely. If it’s still bad or wrong after that, then it’s not that different from a human writing a bad or wrong paper on their own and should be treated similarly.
Yes. The cognitive process performed by a person using an LLM is often no different from that performed by a person using a ghost author, which is a form of plagiarism covered by 42 CFR § 93.234 - Research misconduct.
Plagiarism, at its most fundamental level, is a lie. It is the taking of works or ideas of others and passing them off as your own, either directly or indirectly. The misdeed itself is in the lie, the “I created this” when it is known to be untrue.
However, that lie isn’t being told to the original victim. It’s a lie about the victim, claiming that they didn’t create it or their contributions didn’t matter, but it’s not a lie to them. Instead, it’s a lie to the audience, which is the second victim and the actual target of the con.
> It strikes me as akin to plagiarism. If the purported author can’t even answer basic questions about the paper, how can they plausibly claim to have written it?
Authorship standards differ by field. In biology, for example, it would be common to list someone as an author if they assisted in one experiment. They might be at a different institution and may be unaware of all but the vaguest outline of the paper as a whole—they just got brought onboard because they are an expert in one particular task that needed to be done. In exchange they get to be a middle author (not worth much) and develop a relationship with someone whose expertise they may need on one of their own papers in the future (the primary benefit).
That is not the case here, of course—I just wanted to provide some context for your position not being universally applicable.
> Accepted authors could "vouch" for "dead" papers in case they were "auto-killed"?
The problem isn't with the papers here though, it is the author's understanding of the paper that is in question. A paper written by some hypothetically awesome AI would be a good paper but just not really the proclaimed author's paper.
I think this highlights a dual function of citations that are in tension. A citation can be to claim a stated idea has been made and tested with sufficient rigour to be published. Citation's can also be used to 'credit' others, treating reference as a type of currency. I think this latter form is an outright mistake, but entrenched in academia. The notion of giving credit like this creates a perverse incentive that lies behind much academic fraud, there is enough incentive to be the person to state something that it outweighs the requirement that person has for the statement to be true. Without that notion of credit as currency, issues like plagiarism simply disappear. In the absence of credit, someone making the same claims as someone else without referencing them is just making their own case weaker. Not necessarily less true, but less convincing. If citations were used just used to support a paper then the incentive is to cite, and failing to reference existing work harms only the author.
I think there is too much "This is my idea" and not enough "I think this is true". Credit fails as a measure of effort, diligence, innovation, or truth. Careers are being made and broken by how effectively an individual can game the system.
Do you imagine a world where people are willing to go through the legal hassle of changing their name to get past a ban for low effort journal submissions?
One larger problem here is the value of a research paper is rarely the specific knowledge it adds but in the process of researching that adds to the collective knowledge+experience of those involved, especially training graduate students. AI papers shortcut this entirely. Academia has a lot to answer for this too by making papers the currency of success. AI generated papers are almost shortcut learning at a full system level.
"For more than a century, scientific journals have been the pipes through which knowledge of the natural world flows into our culture. Now they’re being clogged with AI slop."
Academia, particularly the university system, is an untenable collection of interests. The triple stresses of COVID, AI, and funding withdrawal seem to presage what will be a significant disruption.
I disagree for academia using papers as an easy medium for verification and providing more knowledge. Llms always short circuit everything, so how would you fix academia?
> Separately, our group has been exploring approaches along these lines to make such evaluations more scalable
Actually, that sounds like an interesting idea for peer review in general, to include an interview between referees and authors. If it saves one round of rebuttals/reactions, it needn't even consume a lot more of everyone's time if you're doing those things properly. What it would undermine would be blindness, but something's gotta give, and it was already on its way out.
IMHO (not a paper writer, but read a lot during my grad school years), the Genie is out of the bottle. The only way forward, as I see it, is using LLMs for reviews also. Basically, filter all submitted papers with an LLM and ask it to summarize it, find the biggest weaknesses and main strong points, etc. that a human can then use to review the paper. Basically, LLM-as-a-reviewer .
Personally, I would love to see a conference where people are explicitly encouraged to use LLMs for doing the work and writing the papers, and LLMs are used to review them too.
There should at least be a code of professional conduct where authors state the extent to which LLMs were used. (This would also help not wasting time by asking some “authors” about “their” paper.)
Journals themselves should make policies about the extent to which they allow the use of LLMs. In some areas it might be considered more benign than in others.
There is. NLP conferences, like the ACL family (and the ARR) require you to disclose in the paper if you used LLMs for writing and coding, which ones and how. Whether every author is honest is a different matter.
This already happens. It's clear when a reviewer used an LLM, and it's very annoying for the authors that have to respond to what are usually low quality, superficial reviews.
"low quality, superficial reviews" have always been around. Reviewing is most often an unpaid, thankless job and many times reviewers barely put in the effort.
I think that’s a lot of risk of anchoring reviewer bias. I’d be more comfortable with a triaged review where the editor’s office uses models to score whether a human editor should evaluate a paper to potentially send out for review, then the editor makes their own assessment, and the reviewers continue to do their job unassisted
The entire point of an academic paper is to add to the sum of human knowledge. How can an LLM trained on a subset of human knowledge possibly even begin to accurate evaluate such a paper?
I trust an LLM to review that the language used in the paper is grammatically correct, but not to evaluate new information for accuracy.
1) Humans also are trained on a subset of human knowledge.
2)A lot of papers are just about experimenting something, and then applying simple stats. Eg empirical studies, around 1/3rd of published papers. Like, we tried this drug or did this experiment, from a sample size X here are the results. An expert is needed to maybe comment on the conclusion/hypothesis of the underlying suspected mechanism, but LLMs are still very useful on catching bad statistics or p hacking (so so common)
Yes. I want reproducability, open-sourcing, accessibility, correctness, and most of all: usefulness. I don't care how it was written or reviewed, as long as some assurances regarding above things can be made, and I don't see why LLMs would get in the way of that.
What can't be gotten rid of fast enough is the notion that having written something is meaningful on its own. Making something that looks right was a level above total novice: now it's the floor.
This is the journal's policy on LLM use by authors [1]:
> LLMs may be used as general-purpose assistive tools. Whichever tools are used, authors are fully responsible for content on which they are listed as (co-) authors. This includes, but is not limited to, content generated by LLMs that could be construed as plagiarism or scientific misconduct (e.g., fabrication of facts). Low-quality contributions (be they submissions or reviews) that appear to be largely LLM-generated will be closely examined for evidence of the issues mentioned previously, such as scientific misconduct. LLMs are not eligible for authorship. We will periodically revise this policy as new information about the use of LLMs in the scientific process becomes available.
While it doesn't outright encourage using LLMs, it's right at the door, and IMO a policy this weak is actively contributing to the problem the article's author is complaining about. In my opinion any policy weaker than "using LLMs to generate any part of your submission is not allowed and considered a serious breach of ethics" is insane. People like to say that such policies are unenforceable, but that's really not the point (at first), since there are other things like (somewhat ironically) p-hacking that are pretty hard to detect but still widely recognized as unethical. We haven't exactly solved p-hacking either, but at least most of us can agree that p-hacking should be eliminated.
It's hard for me not to read between the lines here. Maybe it's the tinfoil talking, but it being a machine-learning journal, it probably embodies a generally pro-AI philosophy, and thus may not want to discourage too much of it...
It may also be worth noting that this journal apparently uses AI itself on the reviewing side [2]. I'm not claiming this is super unethical or anything as long as the main review is human (although I have concerns), it probably should be part of the conversation.
That is not what I said at all. For research whose ostensible purpose is to increase human understanding, the absolutely bare minimum we could ask for is for the authors to understand their own work well enough to write their research paper themselves, without having an LLM crank it out. I'm not saying anything about LLMs at other points during the research. Unless I'm missing something, the TMLR policy does not forbid using LLMs to generate entire research papers. It only says the authors are "responsible for the content" and warns against "low-quality contributions."
Peer review has its historical issues, but the landscape of science and science-publishing has changed. New problems of authorship and authorial-understanding are now challenged by LLMs writing (at least) good sounding papers - some of which might be of acceptable quality in subject (I am not against AI in the sciences; some of the math work has been great). On the other hand: I am against authors not understanding their own work. High repute journals may need to add "oral exams" to the paper acceptance process...
I'm not familiar with the world of academic publishing, so I want to ask: how is the industry making sure that submissions aren't at least partially AI-generated?
Is it standard practice for authors to have to defend their submissions via interview like this? If not, why not?
Does the vetting process vary with the quality of the publisher?
As an outsider, it's extremely worrying that anyone would even attempt to submit an AI-generated paper for publication in an academic journal. At that level I would have assumed literally everybody should know better than to even try.
> Is it standard practice for authors to have to defend their submissions via interview like this? If not, why not?
It isn't, but maybe it should be. For post-grad qualifications oral defense is standard, and I didn't mind defending my central thesis then, and won't mind now.
Not in a direct interview style, but most us conferences can request additional information or feedback. If they conditionally accept or reject a paper, that conditional relies on feedback from the author(s).
Interviews like this are interesting, but in no way can scale to the infinite paper slop conferences are facing.
I’ve only published a few papers, but this interview sounds extremely unusual to me (I mean, it is clearly a special thing that the editor is doing, which is fine). I wouldn’t do something unethical, but if I had and the editor asked me for an interview like this, I’d know I’d probably been caught.
Why is it unusual: it sounds extremely time-consuming.
As to how worrying AI-generated papers are… it sounds more like a headache for the editors really.
In general, journals don’t have to be perfect; mostly researchers read research papers. You already have to read critically (publish-or-perish has been a thing for a while, so there are plenty of not-so-great papers out there). Peer review is just the “entry” barrier, science is a social process and papers become more or less influential based on a fuzzy process of citation, conference talks, and peer-to-peer suggestions.
For whom? Surely the authors can find an hour after submitting the paper to a journal?
Beware that a reviewer easily spend a full week on reviewing a paper, and there are typically three of them. So if one hour of conversation can save three weeks work, it sounds worth it.
Curation is the new skill. The dismissal of a result on the basis of authorship is one of the reasons blind reviews are there. It sounds like we need more scientists. In al seriousness, this is a skill not just for science. Even at work, the amount of slop is rising exponentially and people are half-treating the symptom with its source, using AI to summarise. We will find our ways eventually.
> All reviewers, Action Editors and Editors-in-Chief for TMLR are unpaid volunteers.
He is not an unpaid volunteer.
He's an associate professor at the prestigious Carnegie Mellon University. He is not paid by the journal, but he is paid a salary by the university, and the university expects that a small part of his academic work is to serve as an editor in academic journals.
There are much easier and more prestigious ways of fulfilling university service requirements than reviewing and editing for TMLR, so I think it is correct in spirit to label it as volunteering.
> and the university expects that a small part of his academic work is to serve as an editor in academic journals.
Some do it for the CV, some do it for science, but I don't know any Uni that expects them to be editor more than they expect them to edit the Wikipedia or write popular science. Cool if you do it, but not expected.
Elsevier, springer and mdpi have literally billions dollars in profit thanks to this free labor. Why would university pay for performing labor for for-profit companies? The system is broken and we should name things as they are - it’s free labor.
Is it really “research” - as in expanding human knowledge - if nobody understands it? The point is deepening human understanding, not producing research papers
I answered requests to be a peer reviewer. (I'm not sure why I was selected, I don't have many publications or credentials.) I saw a lot of papers with hallucinated references. I also remember one paper that described a methodology that I don't think the authors really performed, I think it was just academic fraud where they pretended to have performed an experiment. At the time that I answered the journal requests, AI could hallucinate fake reports, but agents weren't powerful enough to run the experiments yet.
These days agents are able to really perform genuine experiments and write up the results. A prompt like this: "You'll work autonomously end to end to select a research task that meaningfully advances the state of the art in AI, is clearly defined and worth performing, that people would be interested in reading, and that you can perform on this hardware" (insert details) " in a week. Carefully log your steps so that your results can be replicated. Then, do a research review and write your paper about it up with correct, cited references. You must check all of your citations. Look up current lists of "Claudisms", (such as use of the word "genuinely", or "load-bearing"), and remove them from your writeup. After writing your writeup, edit it and pare it down, remove anything unnecessary, keep it fast paced and interesting. Also, try to tell a story, be engaging in your writeup. Don't use violent metaphors, remove references to killing, strangulation, etc. Your writeup should be ready to publish and accurately reflect a real experiment with a meaningful result that advances the state of the art and contributes to understanding. Be concise and focus on why it matters."
Okay, so there's the prompt. You can give it to any AI and have a journal-ready publication in a week. I guess you can ask it to add charts and stuff, if you want to be fancy.
If I gave my agent the above prompt, would I be one of the authors? Maybe it's fair to say I guided, facilitated, elicited, or advised it. But it's clear that the AI would be the one that is actually selecting and running the experiment and writing up the results.
Someone could probably get a publication without even reading the paper they wrote their name on. Their only contribution might be editing their name into the PDF.
It can be turned around as well, for validating existing research, creating a bot that checks papers against all the well-known logical fallacies, issues with statistical methods, checks the images etc.
I hate to say this, but a researcher in biology ain't a statistician or logician. Many math paper hand waves over many parts of a proof.. yet both of those can be and have been useful in practice.
I wonder how the ratios would change for papers at different parts of the review process. For what fraction of published papers are the authors unable to answer basic questions about them?
I think authenticity and trust will command a (larger) premium in this new age of slop.
The article highlights how only one out of ten paper’s authors were able to answer questions thoroughly and at a high level. This indicates an overwhelming percentage of authors are slopping up their work with AI and submitting it without even reading it.
No doubt this is happening, but I wonder how many authors of papers "slated for desk rejection" 10 years ago could answer questions about their papers? We'd need that comparison to understand if this is a new problem or if AI is just a new source of content that the authors of poorly-written papers are using.
Evidence-based assertions are a good thing, but some things are so obvious that a "comparison" or "research" is not needed. This is already an obvious problem in so many places from high schoolers turning in assignments they don't understand, coders submitting code changes they don't understand, blog posts, and certainly to scientific papers.
The Medium comments on this post are also on point. Running the same experiment with accepted papers is a good control. Running a similar experiment with reviewers would be interesting, but more obnoxious because they are not being paid.
I would keep a private blacklist (shadow ban) the authors who wasted several hours of a reviewer's time to prove they were not legitimate. The existence of such a list would be problematic, though.
Could the same system we use here be applied? Accepted authors could "vouch" for "dead" papers in case they were "auto-killed"?
This system is broken and providing more evidence that it is broken isn't much of a step towards fixing it.
It strikes me as akin to plagiarism. If the purported author can’t even answer basic questions about the paper, how can they plausibly claim to have written it?
In the case where someone uses AI to write the paper and then deeply familiarizes themself with it, it may go undetected, but then it’s also presumably less of an issue since they have actually read it carefully and closely. If it’s still bad or wrong after that, then it’s not that different from a human writing a bad or wrong paper on their own and should be treated similarly.
Yes. The cognitive process performed by a person using an LLM is often no different from that performed by a person using a ghost author, which is a form of plagiarism covered by 42 CFR § 93.234 - Research misconduct.
Plagiarism, at its most fundamental level, is a lie. It is the taking of works or ideas of others and passing them off as your own, either directly or indirectly. The misdeed itself is in the lie, the “I created this” when it is known to be untrue.
However, that lie isn’t being told to the original victim. It’s a lie about the victim, claiming that they didn’t create it or their contributions didn’t matter, but it’s not a lie to them. Instead, it’s a lie to the audience, which is the second victim and the actual target of the con.
https://www.plagiarismtoday.com/2019/08/01/the-two-victims-o...
> It strikes me as akin to plagiarism. If the purported author can’t even answer basic questions about the paper, how can they plausibly claim to have written it?
Authorship standards differ by field. In biology, for example, it would be common to list someone as an author if they assisted in one experiment. They might be at a different institution and may be unaware of all but the vaguest outline of the paper as a whole—they just got brought onboard because they are an expert in one particular task that needed to be done. In exchange they get to be a middle author (not worth much) and develop a relationship with someone whose expertise they may need on one of their own papers in the future (the primary benefit).
That is not the case here, of course—I just wanted to provide some context for your position not being universally applicable.
1 reply →
> Accepted authors could "vouch" for "dead" papers in case they were "auto-killed"?
The problem isn't with the papers here though, it is the author's understanding of the paper that is in question. A paper written by some hypothetically awesome AI would be a good paper but just not really the proclaimed author's paper.
I think this highlights a dual function of citations that are in tension. A citation can be to claim a stated idea has been made and tested with sufficient rigour to be published. Citation's can also be used to 'credit' others, treating reference as a type of currency. I think this latter form is an outright mistake, but entrenched in academia. The notion of giving credit like this creates a perverse incentive that lies behind much academic fraud, there is enough incentive to be the person to state something that it outweighs the requirement that person has for the statement to be true. Without that notion of credit as currency, issues like plagiarism simply disappear. In the absence of credit, someone making the same claims as someone else without referencing them is just making their own case weaker. Not necessarily less true, but less convincing. If citations were used just used to support a paper then the incentive is to cite, and failing to reference existing work harms only the author.
I think there is too much "This is my idea" and not enough "I think this is true". Credit fails as a measure of effort, diligence, innovation, or truth. Careers are being made and broken by how effectively an individual can game the system.
That would fail, humans and agents could create new “author” accounts by the swarm or have paid author accounts.
Do you imagine a world where people are willing to go through the legal hassle of changing their name to get past a ban for low effort journal submissions?
9 replies →
One larger problem here is the value of a research paper is rarely the specific knowledge it adds but in the process of researching that adds to the collective knowledge+experience of those involved, especially training graduate students. AI papers shortcut this entirely. Academia has a lot to answer for this too by making papers the currency of success. AI generated papers are almost shortcut learning at a full system level.
Related:
Jan 2026 https://www.theatlantic.com/science/2026/01/ai-slop-science-...
"For more than a century, scientific journals have been the pipes through which knowledge of the natural world flows into our culture. Now they’re being clogged with AI slop."
Sept 2026 https://www.theatlantic.com/ideas/2026/09/college-education-...
Academia, particularly the university system, is an untenable collection of interests. The triple stresses of COVID, AI, and funding withdrawal seem to presage what will be a significant disruption.
I disagree for academia using papers as an easy medium for verification and providing more knowledge. Llms always short circuit everything, so how would you fix academia?
> Separately, our group has been exploring approaches along these lines to make such evaluations more scalable
Actually, that sounds like an interesting idea for peer review in general, to include an interview between referees and authors. If it saves one round of rebuttals/reactions, it needn't even consume a lot more of everyone's time if you're doing those things properly. What it would undermine would be blindness, but something's gotta give, and it was already on its way out.
IMHO (not a paper writer, but read a lot during my grad school years), the Genie is out of the bottle. The only way forward, as I see it, is using LLMs for reviews also. Basically, filter all submitted papers with an LLM and ask it to summarize it, find the biggest weaknesses and main strong points, etc. that a human can then use to review the paper. Basically, LLM-as-a-reviewer .
Personally, I would love to see a conference where people are explicitly encouraged to use LLMs for doing the work and writing the papers, and LLMs are used to review them too.
Extend this beyond review. The most value we would get is from quality checking existing published papers.
There should at least be a code of professional conduct where authors state the extent to which LLMs were used. (This would also help not wasting time by asking some “authors” about “their” paper.)
Journals themselves should make policies about the extent to which they allow the use of LLMs. In some areas it might be considered more benign than in others.
There is. NLP conferences, like the ACL family (and the ARR) require you to disclose in the paper if you used LLMs for writing and coding, which ones and how. Whether every author is honest is a different matter.
This already happens. It's clear when a reviewer used an LLM, and it's very annoying for the authors that have to respond to what are usually low quality, superficial reviews.
Boy do I have news for you! :-D
"low quality, superficial reviews" have always been around. Reviewing is most often an unpaid, thankless job and many times reviewers barely put in the effort.
To be fair, low quality superficial reviews were also not uncommon before LLMs…
I think that’s a lot of risk of anchoring reviewer bias. I’d be more comfortable with a triaged review where the editor’s office uses models to score whether a human editor should evaluate a paper to potentially send out for review, then the editor makes their own assessment, and the reviewers continue to do their job unassisted
The entire point of an academic paper is to add to the sum of human knowledge. How can an LLM trained on a subset of human knowledge possibly even begin to accurate evaluate such a paper?
I trust an LLM to review that the language used in the paper is grammatically correct, but not to evaluate new information for accuracy.
This is very bad logic
1) Humans also are trained on a subset of human knowledge. 2)A lot of papers are just about experimenting something, and then applying simple stats. Eg empirical studies, around 1/3rd of published papers. Like, we tried this drug or did this experiment, from a sample size X here are the results. An expert is needed to maybe comment on the conclusion/hypothesis of the underlying suspected mechanism, but LLMs are still very useful on catching bad statistics or p hacking (so so common)
Yes. I want reproducability, open-sourcing, accessibility, correctness, and most of all: usefulness. I don't care how it was written or reviewed, as long as some assurances regarding above things can be made, and I don't see why LLMs would get in the way of that.
What can't be gotten rid of fast enough is the notion that having written something is meaningful on its own. Making something that looks right was a level above total novice: now it's the floor.
Nope, I’ve tried this, it’s awful.
For a start LLMs love LLM generated text, so you are boosting papers people never had any input in.
Secondly, LLMs in my experience are good at small issues, but fail totally at the whole paper being obviously poorly constructed, or clearly fake.
This is the journal's policy on LLM use by authors [1]:
> LLMs may be used as general-purpose assistive tools. Whichever tools are used, authors are fully responsible for content on which they are listed as (co-) authors. This includes, but is not limited to, content generated by LLMs that could be construed as plagiarism or scientific misconduct (e.g., fabrication of facts). Low-quality contributions (be they submissions or reviews) that appear to be largely LLM-generated will be closely examined for evidence of the issues mentioned previously, such as scientific misconduct. LLMs are not eligible for authorship. We will periodically revise this policy as new information about the use of LLMs in the scientific process becomes available.
While it doesn't outright encourage using LLMs, it's right at the door, and IMO a policy this weak is actively contributing to the problem the article's author is complaining about. In my opinion any policy weaker than "using LLMs to generate any part of your submission is not allowed and considered a serious breach of ethics" is insane. People like to say that such policies are unenforceable, but that's really not the point (at first), since there are other things like (somewhat ironically) p-hacking that are pretty hard to detect but still widely recognized as unethical. We haven't exactly solved p-hacking either, but at least most of us can agree that p-hacking should be eliminated.
It's hard for me not to read between the lines here. Maybe it's the tinfoil talking, but it being a machine-learning journal, it probably embodies a generally pro-AI philosophy, and thus may not want to discourage too much of it...
It may also be worth noting that this journal apparently uses AI itself on the reviewing side [2]. I'm not claiming this is super unethical or anything as long as the main review is human (although I have concerns), it probably should be part of the conversation.
[1]: https://jmlr.org/tmlr/editorial-policies.html
[2]: https://medium.com/@TmlrOrg/ai-reviews-at-tmlr-for-assessing...
So you can’t interact with an LLM, or use any product that has LLM features when performing the research?
That is not what I said at all. For research whose ostensible purpose is to increase human understanding, the absolutely bare minimum we could ask for is for the authors to understand their own work well enough to write their research paper themselves, without having an LLM crank it out. I'm not saying anything about LLMs at other points during the research. Unless I'm missing something, the TMLR policy does not forbid using LLMs to generate entire research papers. It only says the authors are "responsible for the content" and warns against "low-quality contributions."
1 reply →
Peer review has its historical issues, but the landscape of science and science-publishing has changed. New problems of authorship and authorial-understanding are now challenged by LLMs writing (at least) good sounding papers - some of which might be of acceptable quality in subject (I am not against AI in the sciences; some of the math work has been great). On the other hand: I am against authors not understanding their own work. High repute journals may need to add "oral exams" to the paper acceptance process...
I'm not familiar with the world of academic publishing, so I want to ask: how is the industry making sure that submissions aren't at least partially AI-generated?
Is it standard practice for authors to have to defend their submissions via interview like this? If not, why not?
Does the vetting process vary with the quality of the publisher?
As an outsider, it's extremely worrying that anyone would even attempt to submit an AI-generated paper for publication in an academic journal. At that level I would have assumed literally everybody should know better than to even try.
> Is it standard practice for authors to have to defend their submissions via interview like this? If not, why not?
It isn't, but maybe it should be. For post-grad qualifications oral defense is standard, and I didn't mind defending my central thesis then, and won't mind now.
Not in a direct interview style, but most us conferences can request additional information or feedback. If they conditionally accept or reject a paper, that conditional relies on feedback from the author(s).
Interviews like this are interesting, but in no way can scale to the infinite paper slop conferences are facing.
7 replies →
I’ve only published a few papers, but this interview sounds extremely unusual to me (I mean, it is clearly a special thing that the editor is doing, which is fine). I wouldn’t do something unethical, but if I had and the editor asked me for an interview like this, I’d know I’d probably been caught.
Why is it unusual: it sounds extremely time-consuming.
As to how worrying AI-generated papers are… it sounds more like a headache for the editors really.
In general, journals don’t have to be perfect; mostly researchers read research papers. You already have to read critically (publish-or-perish has been a thing for a while, so there are plenty of not-so-great papers out there). Peer review is just the “entry” barrier, science is a social process and papers become more or less influential based on a fuzzy process of citation, conference talks, and peer-to-peer suggestions.
> it sounds extremely time-consuming.
For whom? Surely the authors can find an hour after submitting the paper to a journal?
Beware that a reviewer easily spend a full week on reviewing a paper, and there are typically three of them. So if one hour of conversation can save three weeks work, it sounds worth it.
3 replies →
The rates of paper publishing show that everyone must be using ai now or the rates wouldn't have gone up
Curation is the new skill. The dismissal of a result on the basis of authorship is one of the reasons blind reviews are there. It sounds like we need more scientists. In al seriousness, this is a skill not just for science. Even at work, the amount of slop is rising exponentially and people are half-treating the symptom with its source, using AI to summarise. We will find our ways eventually.
> All reviewers, Action Editors and Editors-in-Chief for TMLR are unpaid volunteers.
He is not an unpaid volunteer.
He's an associate professor at the prestigious Carnegie Mellon University. He is not paid by the journal, but he is paid a salary by the university, and the university expects that a small part of his academic work is to serve as an editor in academic journals.
There are much easier and more prestigious ways of fulfilling university service requirements than reviewing and editing for TMLR, so I think it is correct in spirit to label it as volunteering.
> and the university expects that a small part of his academic work is to serve as an editor in academic journals.
Some do it for the CV, some do it for science, but I don't know any Uni that expects them to be editor more than they expect them to edit the Wikipedia or write popular science. Cool if you do it, but not expected.
Elsevier, springer and mdpi have literally billions dollars in profit thanks to this free labor. Why would university pay for performing labor for for-profit companies? The system is broken and we should name things as they are - it’s free labor.
Is it really “research” - as in expanding human knowledge - if nobody understands it? The point is deepening human understanding, not producing research papers
FYI in case the author is reading, https://www.cs.cmu.edu/~nihars/preprints/greCAPTCHA.pdf is a dead link.
EDIT: I found a live link on arxiv https://arxiv.org/html/2609.20481v1
In June 2026 I proposed a CAPTCHA for scientific publications
https://chorasimilarity.wordpress.com/2026/06/13/a-captcha-f...
At the moment this was seen as a tongue in cheek proposal.
I’d agree it’d be a funny proposal, wouldn’t have worked back then but funny.
1 reply →
completely necessary.
I answered requests to be a peer reviewer. (I'm not sure why I was selected, I don't have many publications or credentials.) I saw a lot of papers with hallucinated references. I also remember one paper that described a methodology that I don't think the authors really performed, I think it was just academic fraud where they pretended to have performed an experiment. At the time that I answered the journal requests, AI could hallucinate fake reports, but agents weren't powerful enough to run the experiments yet.
These days agents are able to really perform genuine experiments and write up the results. A prompt like this: "You'll work autonomously end to end to select a research task that meaningfully advances the state of the art in AI, is clearly defined and worth performing, that people would be interested in reading, and that you can perform on this hardware" (insert details) " in a week. Carefully log your steps so that your results can be replicated. Then, do a research review and write your paper about it up with correct, cited references. You must check all of your citations. Look up current lists of "Claudisms", (such as use of the word "genuinely", or "load-bearing"), and remove them from your writeup. After writing your writeup, edit it and pare it down, remove anything unnecessary, keep it fast paced and interesting. Also, try to tell a story, be engaging in your writeup. Don't use violent metaphors, remove references to killing, strangulation, etc. Your writeup should be ready to publish and accurately reflect a real experiment with a meaningful result that advances the state of the art and contributes to understanding. Be concise and focus on why it matters."
Okay, so there's the prompt. You can give it to any AI and have a journal-ready publication in a week. I guess you can ask it to add charts and stuff, if you want to be fancy.
If I gave my agent the above prompt, would I be one of the authors? Maybe it's fair to say I guided, facilitated, elicited, or advised it. But it's clear that the AI would be the one that is actually selecting and running the experiment and writing up the results.
Someone could probably get a publication without even reading the paper they wrote their name on. Their only contribution might be editing their name into the PDF.
It can be turned around as well, for validating existing research, creating a bot that checks papers against all the well-known logical fallacies, issues with statistical methods, checks the images etc.
I hate to say this, but a researcher in biology ain't a statistician or logician. Many math paper hand waves over many parts of a proof.. yet both of those can be and have been useful in practice.
1 reply →
I wonder how the ratios would change for papers at different parts of the review process. For what fraction of published papers are the authors unable to answer basic questions about them?
I think authenticity and trust will command a (larger) premium in this new age of slop.
The article highlights how only one out of ten paper’s authors were able to answer questions thoroughly and at a high level. This indicates an overwhelming percentage of authors are slopping up their work with AI and submitting it without even reading it.
No doubt this is happening, but I wonder how many authors of papers "slated for desk rejection" 10 years ago could answer questions about their papers? We'd need that comparison to understand if this is a new problem or if AI is just a new source of content that the authors of poorly-written papers are using.
Evidence-based assertions are a good thing, but some things are so obvious that a "comparison" or "research" is not needed. This is already an obvious problem in so many places from high schoolers turning in assignments they don't understand, coders submitting code changes they don't understand, blog posts, and certainly to scientific papers.
1 reply →
Sounds like I shouldn't be so hard on AI when it hallucinates things based on this data? :)