Now is the time to give LLMs access to the ACM digital library

10 hours ago (cacm.acm.org)

As a researcher with many articles in the ACM library, I have to say this is a masterclass in hypocrisy. Obviously, lawyers can decipher the terms of ACM publishing contracts and Creative Commons licences to determine if this will be acceptable or not. But ACM is not a company, it's a non-profit founded in 1947 to represent scientists.

I would be surprised if a majority of ACM members were to say yes should we ask them (but ACM is not known for such democracy). Along with book authors, we are one of the many people that provide the knowledge and expertise on which large tech firms train their models, and get nothing in return. Actually, life is getting worse for us: extra workload in universities with students' AI use, a completely broken peer review system, etc. Hence the irony of ACM thinking about licensing, and only licensing, at a time where this is the least of our priorities.

  • How do you square away the idea that you do science for the increase in knowledge of human kind, but then say that a particular use of that knowledge is verboten?

    I get the copyright aspect of this and I'm not arguing that here. I'm more asking about the moral / ethical idea of choosing who can benefit from your science.

    Obviously there are the moral / ethical arguments about AI in general here to weigh against - those have been hashed out significantly elsewhere, and I'm not interested in debating them. What I'm asking about here is the impact on science by sharing it with tooling that distributes it in ways not generally considered when originally written.

    A quick check of your post history suggests the frame that you work in strongly is privacy related research (observation - may be wrong). I'm curious how that impacts what you wrote here generally.

    (Just to be perfectly clear, I'm not arguing your points here, trying to understand them better)

    • Good question. Unfortunately, academic knowledge is widely ‘verboten’ already. Everything under paywall, researchers having to pay up to $10,000 to publish in open access in some venues, rare books unavailable even to top universities. Access to knowledge and information is increasingly difficult for everyone.

      That said, what matters here is the social contract, what do I bring to society and what do we get from tech companies. For most people around the world, access to the typical leading models is out of reach. Not many on this planet can pay the subscriptions (or even API keys) that offer access to the best models. So I'm not buying the argument that tech companies are broadening access. What we're creating is a increasingly discriminatory society where the few get access to information, and the many don't.

      1 reply →

  • If it was a non profit that trained the model - would that change your mind?

    • The issue here is ACM focusing on licensing. A non profit would not be able to pay ACM for access. Hence why this policy is hypocritical: it gives more power to the larger players and undermines smaller actors in the field who have fewer resources.

      5 replies →

  • You mean the knowledge you gathered with public grants, with a public paid salary, yet don’t want to make freely available to the public?

    Yeah, too bad

  • If you don't hold a patent for the use of the knowledge you published publicly, you can't prevent others from using the knowledge. You enjoy the prestige attached to the idea that you're an academic who participates in giving away their knowledge but then you play this game when that knowledge would actually be useful as opposed to being read by 3 other people in your special area who sit on your various committees in your career, now you want to forbid the use for culture war intra-elite signaling reasons.

    You don't own the knowledge you put out there unless you have a limited time valid patent. The rest is absurdity. If you want to keep your findings to yourself, keep them secret.

    • The intellectual property law that governs ACM articles is copyright law, not patents. I don’t know who controls these (the ACM or the authors) or what rights might have been granted to the public.

      The entire point of copyright law is so that people can make their writing public and still be able to control the right to make copies (for example, into your dataset for training an LLM).

      22 replies →

    • This makes no sense whatsoever. How would a philosopher of science, or a social scientist who publish in the ACM apply for a patent? Or someone who builds software (software patent not so easy to get ;)). I have applied for patents before and I'm pretty sure my patent application has been fed to countless LLMs by now.

      The issue is not who owns knowledge, it's how it benefits humanity.

      1 reply →

    • The world would be quite different if AI companies had to create the knowledge they trained on, rather than consume that knowledge freely given away. They're not known for freely giving away their produce either, I don't know why you think the ire should be pointing in this direction.

      1 reply →

    • “If you don’t lock your bike, you can’t prevent others from taking it for a ride. You enjoy the mobility attached to the idea you’re a bike rider who rides a bike but then you play this game when that bike would actually be useful as opposed to sitting in the bike rack all day”.

      1 reply →

  • As long as we are going toward a world of abundance where money doesn't mean much and the main currency is time, I can't complain. I will subsidize that with my brain power turned into ink on paper.

    • Abundance for whom, exactly? And if the answer is everybody: who in charge has an incentive to do this?

      Sorry to ruin your day, but if the people with money could have introduce equal society, you would have noticed their attempts by now.

      1 reply →

How about we give humans access

So give it for free to the open weight models, and charge the closed weight models. Easy

Blocking access only hurts people who follow the rules. Unblocking access lets them compete with those who break the rules.

I think the right choice is pretty clear...

I'm really not sure how this would work. I don't know how the ACM works, but in IEEE you would have to give them your publishing rights. However, training a LLM is not publishing by itself, it is a derivative work? Any way, at this point authors should be entitled to monetary compensation, not the publisher. The deal is totally different.

ACM has been leaning heavily into AI-generated content for their journals in the past year, and this article is no exception: it appears to be 100% AI generated and full of LLM verbiage.

There's something hilarious about that, but also, snake eating its own tail.

  • I don't think this is AI-generated. It is focused and direct. It reads like anodyne albeit totally human academic manager writing.

This reminds me of the "I drink your milkshake" scene from "There Will Be Blood."

Does the ACM really think LLMs haven't already consumed 90% of the content through other sources?

Would you prefer a parquet dump of acm articles to hugging face?

A llm emulating a person is why many of my used sites banned llms due to scraping bandwith costs

Is this about access or accessibility to claude (for example)

I am sure the entirely of human computing knowledge is not that big.

Are they in the position to do that?

What about the authors?

  • The ACM sent around a nice query to members which made it clear that they were going to do it even if 100% of the members said "No, don't do that".

    I suppose they could be sued.

  • Don't you grant ACM a right to distribute your work when publishing? So doesn't ACM already have the right to grant access to AI?

What about us average humans, or is it ONLY the corporate LLM token dealers who get access?

Either way, I'll pirate.

the digital library should have always been open access

now it will be fodder for the slop machine

(i think that LLMs are going to wreck the peer review system for all but hard-experimental papers)

  • Since starting to use AIs seriously for search in the last 3 months I have read and referenced more published papers and academic primary sources then I think I did in the previous 5 years. They're fantastic for pointing at some claim and asking for the primary source for it, and then it does the work of following things through the layers of backref to the original, assuming it's online. Opening the library up to AI access makes it far more accessible and usable then it was before.

    AIs, at least in their current form, make you more who ever you were. If you want snap, glib answers of dubious accuracy, they'll give them to you, more easily than ever before. If you want to dig back into primary sources and get the original content, they'll do that for you, more easily than ever before.

    Can't speak to how the science infrastructure is going to handle them, but if it takes down the peer review system, which I think has been worthless for probably going on two decades and has just given the entire enterprise a false sense of assurance, it'll probably be a net gain in the end. Peer review is a source of more problems then it is solving right now.

  • The real efficacy of the peer review system has always been somewhat questionable. That's a big reason why arxiv is so prominent.

  • > I think that LLMs are going to wreck the peer review system for all but hard-experimental papers

    I think I may be shadowbanned, but at what point do we start viewing LLMs as a national security threat?

ACM bureaucrats and AI shills are flagging dissent. Welcome to 2026. The entire Internet is just top down propaganda.

TLDR: “we’re going to try to get money from LLM providers for access to our back catalog without getting permission from the authors or providing them with any share of the revenue”.

  • Scholarly article authors never expected royalty or permission for their scholarly works to be reused.

I don't want to be mean but if it's that hard to even recognize that LLMs are even a valid thing that could intersect with your business that you need some kind of campaign for it..

It's almost like, "don't hurt yourself unc, we will just search arxiv".