Today’s biggest artificial intelligence (AI) startups make no shortage of bold promises. Their technologies, some boast, will revolutionize software development, drug discovery, and scientific research.
Yet a new preprint posted on 16 July on bioRxiv suggests many of these firms barely participate in one of science’s most fundamental practices: publicly documenting discoveries in scientific literature so other researchers can evaluate and build on them. More than half of AI unicorns—private companies valued at more than $1 billion—have never played a leading role in publishing a scientific paper or preprint, according to the new analysis. Collectively, they accounted for just one in every 1000 AI papers published in 2025.
“For a field that is supposedly reshaping science and is so advanced in terms of scientific potential, not having any scientific documentation seems like a very weird paradox,” says paper co-author John Ioannidis, a metascientist at Stanford University. “How can you judge that what they say is real, validated, and reproducible?” The scarcity of publications, others say, also makes it harder to assess AI’s social impacts, including energy use and safety.
But University of Alberta AI ethicist Mohamed Abdalla says the findings reflect the incentives facing commercial AI developers, rather than solely a failure to uphold scientific norms. “It’s not the company’s job to advance science, right?” he says. “The company’s job is to advance money.”
Ioannidis has long studied how unicorns, particularly in biotech, engage with the scientific literature. (In 2015, he was the first to publicly scrutinize the lack of peer-reviewed studies produced by Theranos, the blood testing startup that proved to be based on fraudulent data.) He wondered whether AI unicorns would show similar patterns.
To find out, he and his team first identified all 317 unicorn AI companies that have existed from 1998 to 2025. Then, they searched for publications affiliated with these startups—including journal articles, conference papers, reviews, and preprints. They selected those where a company researcher played a leading role as a first or last author, indicating the startup had made a substantial contribution to the work. The final data set included 2077 final publications, comprising 1389 peer-reviewed papers and 688 preprints.
More than half of the startups had never produced a single qualifying paper, the analysis revealed. Scientific influence proved even more concentrated, with the top 5% of firms accounting for greater than 90% of all citations. OpenAI alone was responsible for nearly 40% of all citations in the data set, followed by the Chinese computer vision company Megvii and the platform Hugging Face. And even at the most prolific companies, much of the output came from the same small group of repeat authors. For example, despite OpenAI employing roughly 4500 people, only eight researchers had authored five or more qualifying papers.
The findings are unsurprising to some AI researchers given how the industry is structured. For example, unlike the pharmaceutical industry, where published discoveries can be protected by patents, AI companies have learned they often gain little from publicly disclosing technical advances, says Nur Ahmed, an AI researcher at the University of Arkansas. Google’s landmark 2017 paper on the transformer—the architecture that underpins today’s large language models—has become a classic cautionary example, Abdalla adds. Although Google patented aspects of the technology, “I don’t think anybody’s paying Google for that,” he says.
Startups also operate on much faster timelines than academia, where peer review can lumber on for months or even years. That’s why many AI companies have embraced what Avijit Ghosh, an AI policy researcher at Hugging Face, calls the “blogification” of research: announcing new models and releasing code or data sets through blog posts and technical reports rather than scientific journals. The new analysis didn’t track those outputs, he points out.
For Ghosh, the debate shouldn’t center on publishing in journals versus blogs. What matters is whether companies are releasing enough code, data sets, or model weights (the numbers that determine how a model interprets and responds to a prompt) for others to independently verify and build on their work, he says.
The preprint also found that firms based in China consistently published more papers than their counterparts based in the United States. Whereas leading U.S. frontier labs have increasingly kept the details of their most capable models secret or “closed sourced,” leading Chinese companies have embraced “open-source” models. Moonshot AI, one of the Chinese startups included in the study, recently unveiled Kimi K3—one of the strongest open models to date—and publicly released its model weights through Hugging Face today.
But whether models are open or closed, the rapid pace toward increasingly powerful generalist AI worries Emma Pierson, a computer scientist at the University of California, Berkeley. She argues AI research—whether published freely or kept secret—risks accelerating models that pose serious societal and safety concerns, including supercharging cyberattacks. “If we were racing forward on cancer-curing AI, I would be like, ’Fantastic, full steam ahead,’” she says. “But that’s not what we’re racing toward, right?”
Facts Only
* More than half of AI unicorns have never played a leading role in publishing a scientific paper or preprint.
* These firms accounted for only one in every 1000 AI papers published in 2025.
* Scientific influence is concentrated, with the top 5% of firms accounting for over 90% of all citations.
* OpenAI was responsible for nearly 40% of all citations in the analyzed data set.
* Many AI companies have not produced any qualifying scientific papers.
* Google’s 2017 paper on the transformer architecture is cited as a cautionary example regarding patenting versus public disclosure.
* Some AI companies embrace "blogification" by releasing code and data via blog posts instead of journals.
* Firms based in China published more papers than their counterparts in the United States.
* Moonshot AI publicly released model weights through Hugging Face.
Executive Summary
Many artificial intelligence startups are not participating in the standard practice of publicly documenting scientific discoveries, which raises concerns about reproducibility and societal assessment. A new analysis found that more than half of AI unicorns, private companies valued over $1 billion, have never taken a leading role in publishing a scientific paper or preprint. This scarcity makes it difficult to evaluate the validity, reproducibility, and social impacts, such as energy use and safety, of AI advancements.
Some researchers attribute this pattern to commercial incentives rather than a failure of scientific norms, suggesting that the primary driver for these firms is financial advancement rather than advancing science itself. Furthermore, industry timelines are faster than academia's peer-review process, leading some companies to favor "blogification" of research over traditional journal publication. While some entities, like Chinese companies, embrace open-source models, others maintain secrecy or patenting strategies. The central debate shifts from publishing venues to whether sufficient data and code are released for independent verification.
Full Take
The discrepancy between the promise of revolutionary scientific advancement by AI and the observed pattern of limited public documentation reveals a fundamental tension between commercial incentives and epistemic responsibility. The structure of the AI industry, where financial gains are prioritized over open scientific dissemination, appears to create an environment where publishing is not incentivized or necessary for corporate success, contrasting sharply with established scientific norms. This dynamic is exacerbated by the speed of the AI development cycle, pushing some actors toward alternative forms of documentation like blog posts rather than slow peer review.
The concentration of scientific influence among a small subset of firms, exemplified by OpenAI's citation dominance, suggests that while the field is technologically advanced, the narrative surrounding its validation remains centralized and potentially opaque. The tension between closed-source development—where entities prioritize proprietary advantage—and open-source release presents a complex negotiation over knowledge distribution. Whether this results in safer or more robust AI advancement depends on whether the drive for competitive advantage can be decoupled from the need for transparent, verifiable scientific consensus.
What are the long-term systemic consequences if the mechanism for validating revolutionary science becomes siloed within commercially driven entities? If the focus remains on proprietary advantage, how will societal safety and the broad application of these powerful technologies be governed without a shared public epistemology? Does the emergence of open models simply reflect an arms race between competing economic and scientific agendas, or does it signal a necessary shift toward a different kind of accountability in rapidly evolving fields?
Sentinel — Human
The article presents a complex, well-supported argument by weaving quantitative data on AI publication practices with qualitative expert commentary on the implications for scientific validation and safety.
