← Back to context

Comment by jrm4

20 hours ago

The important question is:

So what?

Genuinely. I get that there may be some visceral reaction against this, but when I break it down, I mostly fail to see the problem. Seems like what is actually important is:

Compared to before, when a human reads it, do they -- or society -- get something good out of it? Is it worth it to add this to the "pantheon?"

If that's not what's happening enough, and if this doesn't describe the process -- then the problem lies elsewhere, no?

I recently desk-rejected a paper where every single citation in its Introduction was hallucinated. That means that the entire connection between what the author(s) did and how it relates to existing research was simply made up. I've never seen this happening before AI but now there's at least one paper in every cycle pulling something similar.

My problem therefore is: we are seeing more and more papers written with tools that are known to make up facts, citations, and even entire papers. And the number of papers has increased, too. I therefore see it less as "people are being more productive" and more "people are releasing bad science much faster than we can keep up with".

  • Yeah but that's an issue with the researcher putting out a bad paper, and it suggests you'll have to reject more papers. We wouldn't ban email because many of the emails are spam, it just means we need new tools to filter out junk. AI will allow researchers to be more productive all together and take less time to publish a paper, which is good.

    • The problem with spam is that it's not a technology like AI is. So I suggest taking cars instead.

      Cars have plenty of advantages, and yet no one would say "the number of pedestrians killed by cars is rising, but that's an issue with the drivers". In fact, the opposite is true: from fines and school zones to speed bumps and bollards, we have accepted that cars bring structural problems with them that cannot be solved at the driver level alone.

      > we need new tools to filter out junk

      Agreed, but if my office suddenly was flooded with garbage my first thought wouldn't be "I need more, bigger trash cans" but rather "who brought all this junk here and why?". To simply assume that the garbage is a sudden natural phenomena that I have to live with seems, at the very least, unfair.

    • > Yeah but that's an issue with the researcher putting out a bad paper, and it suggests you'll have to reject more papers. We wouldn't ban email because many of the emails are spam, it just means we need new tools to filter out junk. AI will allow researchers to be more productive all together and take less time to publish a paper, which is good.

      It's a signal:noise ratio thing. If 1 out of every 1000 AI-written papers are bad, it makes sense to put in a filter that auto-rejects any paper that has AI tells.

      After all, if that 1 researcher was any good, he wouldn't have used AI to write the thing in the first place.

      Publishing was always about getting past the filters. There's one more filter - "AI-generated content" - so do what you have to to get past it. IOW, write your own paper.

  • While working on my PhD, there was one guy in the field that would publish new, better results to Springer every time someone improved upon benchmarks. No code, no data, no conference publications, paywall restricted, unverifiable, always the best.

    I am almost certain he was "hallucinating" the results. This was in the 2010s

    There are well known issues in academic publishing, though I imagine it has become much noisier like open source

Surprisingly I agree with you. My opinion is: if it makes communicating research more effective, while not reducing the quality of the output substantially, I see no issue.

A possible conclusion for this could be: If the majority of CS papers is AI written, let's just accept this reality universally and stop worrying about it altogether.

  • I think the key question is more effective at what?

    I see plenty of anecdotal evidence that models have been trained fantastically well—and getting better—at writing to trigger the right neurons in the human population to produce “This is interesting/informative/correct” responses in bulk.

    Could their ability to produce those responses run far ahead of their ability to actually achieve the last in reality? Sure seems plausible, and then where are we?

  • > if it makes communicating research more effective, while not reducing the quality of the output substantially, I see no issue.

    That's a big "If".

    If a research is good, the author still has to clear all the hurdles in publishing. "Writing your own paper" is just one more hurdle.

    > A possible conclusion for this could be: If the majority of CS papers is AI written, let's just accept this reality universally and stop worrying about it altogether.

    That's just a different way of saying "if the majority of CS papers are crap, lets just accept this reality".

    So, go on, publish away all your AI-induced "research", but the bar is slowly going to be raised anyway to reject that. That's how science always worked - when a bar is not sufficient to exclude the crap, it is raised.

  • I reviewed for ACL (big NLP conf) and another similar conference recently. Lots of crap (though there always has been). One superficially well-written paper was likely AI-plagiarized (i.e. the AI sampled an idea from prior work and rewrote it) and got desk rejected for it.

    Reviewers were also totally unengaged. Of 20 reviews I read (from my reviewers or from reviewers on the same papers), maybe 2 were mediocre, and the rest were crap (though likely not AI).

    The notion that science will somehow benefit from this is about as stupid an idea as you can have. Science relies on skepticism. AIs are not skeptical, and many folks are submitting papers because they stand to gain something, not because they are motivated to do good research or develop new understanding. Fields are being inundated with garbage that is maximally indistinguishable from real work (that's the training objective for LLMs). This in turn maximizes the cost of identifying bad work.

    This is the same enshittification process that we see everywhere else. You get spam phone calls because there is no reason for a spammer not to call you. "Researchers" are submitting spam papers because there is no cost to doing so with some possible gain. Absent intervention, this eventually drives the community value of the network to zero (or potentially negative, if friction costs to switching are high).

    • And yet the quality of work at ACL is still higher than NeurIPS, and that's with ACL this year having over 40% of all posters looking nearly identical due to everyone claude coding their posters...

  • I mean, sure. But what an absolutely insane predicate. "Not reducing the quality of the output substantially" is (as an AI might say...) load bearing there.

    Other problems include: Signal to Noise Ratio going through the roof.

    • Is it? I guess look -- I'm an academic, I've read piles of articles such as these. Signal to noise, already not great.

      Yes, I feel like there's room to improve things, I just strongly doubt that "using AI to detect AI" is a particularly useful thing to do here.

  • > if it makes communicating research more effective, while not reducing the quality of the output substantially

    That’s a big if. We all know that’s not what’s happening.

  • > My opinion is: if it makes communicating research more effective, while not reducing the quality of the output substantially, I see no issue.

    That's a big if. ArXiv is not peer reviewed and LLMs basically interpolate and extrapolate text, which makes them essentially fluff generators. Even in the most charitable interpretation, LLMs enable those with nothing to say to say nothing while meeting surface-level style guides.

    • We can't know if real science is happening in the background but I'd wager that the majority of these papers is not complete slop but real findings with AI generated text used to communicate it. If it was just straight slop I would be really worried.

      2 replies →

If a human writes a given paragraph, I know it has attained a minimum level of awareness in some human mind at some point. If an AI writes a paper, I have no guarantee that it has reached that minimum. We also fear, with some reason, a correlation between AI writing and possibly being hallucinated or non-existent. Not just because AI is bad and people should feel bad for using them, but because given that a paper is somehow fake, it is probably more likely to be made with AI, and running that backwards it is reasonable to take closer looks at AI papers.

If fake papers weren't already a big problem before AI and the fields had already been policing themselves adequately, if this was already a functioning high-trust domain, maybe we could ignore this a bit more, but the fields already manifestly had problems. People taking advantage of that are reasonably more likely to use AI. The pressures to publish or perish provide the voltage and the AIs are a rather convenient path-to-ground.

I agree in some sense that if a truth is published, it doesn't matter if the AI or a human published it. However there are perfectly reasonable reasons to be concerned that AI usage is correlated to not publishing truths, especially in a world where merely being human-generated was already not an adequate check against that.

Even if you can reject the aesthetic argument for non-fiction works (although read some of Dijkstra's papers for a good counterargument), it is still a problem because it breaks an important quality signalling mechanism.

Pre-LLMs, a paper with no spelling or grammar errors showed that somebody had put effort into writing and editing it. If they cared about the presentation, they probably also cared about the content. LLMs routinely produce nonsense that looks superficially like high-quality work.

There are far too many papers to read all of them. LLM slop is evidence that something is probably low quality. As the saying goes, "if you can't be bothered writing it, I can't be bothered reading." The rare outliers will get enough citations and recommendations to overcome this filter.

Indeed. We'd likely find a similar pattern if we count spelling mistakes before and after autocorrect.

you just dont respect your readers, thats all

  • if 5% of the work that I'm reading to keep up is slop, that's a shame. that's time I spent puzzling over connections that weren't real, trying to impose some kind of logic on the arguments, and wondering why the data doesn't really support the conclusion.

    if 50% of the work is nonsense, then there's a serious concern that we can't move forward at all.

Another important questions is:

What about the papers that graduate to proper publication?

Arxiv is full of pre-prints that anyone can upload.

  • > Arxiv is full of pre-prints that anyone can upload.

    You now (at least for some categories) have to receive endorsement from someone who has multiple recent papers on arxiv in the same (or adjacent) category.

The problem is that humans start with credulity. AI hallucinates and makes up stuff some percentage of the time. Humans are not "default deny" when given information. Unleashing that was a disservice to mankind and has created a future filled with lies and people who are confident in them.

  • LLMs are in essence an attack on the concept of written language, harvesting and dissolving the social contract that underpins it; it is no longer safe to assume that text has intent, let alone content, simply because it is grammatically well-formed.

    Perhaps the most darkly amusing consequence of this particular mania is that by poisoning the majority of our information environment with hallucinated slop, we have likely crippled the next several generations of machine-learning techniques before they're even invented! Small, locally-hostable LLMs will rattle along spewing spam long after the broader "genai bubble" pops, and building clean training datasets will permanently be more difficult and expensive.

  • Oh, to be technically correct:

    AI hallucinates and makes up stuff 100% percent of the time. Never been a fan of that word for this.

    Again, I fail to see the problem here that isn't solved by careful reading WHICH IS WHAT PEOPLE SHOULD BE DOING ANYWAY. I would like to see room for AI disclosure, maybe a statement of "this is how much AI I used."

    But this blanket X% of this looks like AI? Again, so what?

    • If as a primary school student you had needed to reverify every fact /experiment that was presented in your textbooks, you would not have gotten very far.

      Trust is very important to human progress.

    • I find the fact that 39% of papers use a style that signals lack of effort somewhat worrisome. Even if that’s 100% wrong, and they’re all high effort papers, the fact that they give off the same aura as low effort papers is a problem for the authors.

LLM-written text tends to have the property that it gives the impression of expertise and knowledge in excess of what the text actually contains.

In other words, it "sounds smart" without necessarily having anything to back it up. In even more critical terms, it's very good at bullshitting.

Unfortunately for us, the scientific community current relies on a certain amount of trust. (To do otherwise is very expensive! see: bitcoin). When you introduce a known-bullshitter to write your papers, every human in the loop effectively has to defend against an adversarial attack. Not just the readers at home, or the peer reviewers, but even the author needs to be wary that the facts and arguments coming out of the LLM are true and meaningful.

Personally, I've been a minor contributor to several high-profile papers. I don't know how every field does it, but in my experience, the corresponding author (generally the PI or other senior scientist), is responsible for the accuracy of the paper. They ultimately have to trust the people who did the work that the facts are true. Introducing LLMs into the mix make it more difficult for them to identify and review sections they are unsure of. (An honest person will typically write at a confidence level reflecting their certainty. LLMs do not do this in any reliable way.)

I've also found that LLMs frequently use metaphors that are unhelpful, or used out of context in a field that isn't familiar with them. This makes understanding the text more work, for no good reason. Introducing terms or definitions with low relevance reads as impressive at first glance, but avoiding the standard terminology in the field just adds confusion. (As an analogy, imagine if you were reading a CS paper that, for no particular reason, devoted a section to a new data structure called an "akimbo tree," which after much untangling, you realized was a reinvention of a randomized splay tree.)

  • Agreed. I know nothing about nuclear physics. I still doubt you could pick a random person off the street and have them convincingly pose as a nuclear physicist to explain a "nuclear physics" concept to me. I doubt you could do it with a random PhD from a non-physics field. An LLM could probably convince me even if 90% of the content of the explanation is subtly or blatantly incorrect.