Search This Blog

Sunday, October 04, 2026

The arXiv - where it's been and where it's going

Back in the ancient mists of time, scientists and mathematicians would circulate preprints of their articles among friends and colleagues via the postal service, as a courtesy, to get feedback and to try to make sure people in the community were aware of their forthcoming work.  As the wikipedia entry says, with the advent of widespread LaTeX and the development of the web, Paul Ginsparg (then at LANL) put together an html-based site for electronic sharing of preprints, initially at xxx.lanl.gov (back before "xxx" in URLs was the kind of thing filtered and blocked by employers).   In 2001 he moved to Cornell, and by then the lab was perhaps relieved to see the site, rebranded as the arXiv, shift to Cornell's library/repository infrastructure.   PIs my age remember (fondly? maybe?) the old days of the arXiv, when Prof. Ginsparg's rather dry sense of humor pervaded the site.  The skull-and-crossbones logo.  The "help"/FAQ pages that basically said, "if you can't figure out how to .tar.gz all of your necessary LaTeX files, and you can't figure out how to make your .eps figures small, maybe you should reconsider whether you're smart enough to be sharing your ideas here".  The arXiv was for preprints, without peer review (though interesting follow-on sites like Scirate now exist for organized commenting on the articles).

From these modest beginnings, the arXiv has grown enormously, including imitators/spin-offs such as chemrxiv and biorXiv and socarXiv.  The arXiv has recently become an independent nonprofit, hired a CEO, secured multiyear philanthropic support, and hosts over 3 million articles.  It's been interesting seeing some level of complaints online about some of these steps, but when the audience is so large, not everyone is going to be happy.  Rapid growth has been a major issue - see here for a graph of monthly submissions:
The exponential rise (except for a slight pandemic-correlated shoulder) has been problematic, especially recently.  One hallmark of the arXiv over the years has been its ability to function with comparatively minimal need for "moderation".  Early on, one consequence of the minimalistic help and moderate technical entry barrier was that it was unusual for fringe/pseudoscience to make its way onto the server.  (Hence the establishment of viXra.)  With the ease of cranking out properly formatted readable manuscripts using AI, clearly the arXiv has been struggling.  If 10-15% of submissions need some kind of human intervention or review, the support needs are rapidly outpacing the limited count of support staff.  There can be substantial backlogs.  To help deal with this, the arXiv recently updated its policies regarding AI-generated content (and AI cannot be a co-author, because the AI tools cannot take responsibility for content), and most recently has had to limit submission rates to two papers per month per submitting author.  These moves, too, have drawn some criticism (e.g. here).  Personally, I think the operators of the arXiv face an incredibly challenging environment and are doing the best they can - the idea that they are making moves because they are establishment sticks in the mud who don't understand the New Way of Doing Science is just wrong-headed.

It's completely unclear where all this is heading.  Exponential growth in nature signals instability and does not continue forever. If proponents of very heavily AI-driven research want to establish a repository specifically for that work, that's up to them.  [It is very on brand for the hard core AI advocates to argue that the arXiv is somehow morally obligated to host everything (regardless of hardware or personnel costs) so that future AI tools can read everything (a repository growing too quickly for human researchers to keep up) and summarize it.]

One overarching point that should come up in any arXiv discussion: The arXiv has become a global repository for an enormous amount of human knowledge, without charging anyone publication fees.  This should make interested parties think reallllllly hard about economic models of for-profit publishers. 


No comments: