Thank you for visiting nature.com. You are using a browser version with limited support for CSS. To obtain
the best experience, we recommend you use a more up to date browser (or turn off compatibility mode in
Internet Explorer). In the meantime, to ensure continued support, we are displaying the site without styles
and JavaScript.
Investors and technology-transfer offices expend enormous effort trying to spot commercially promising research before it reaches the point of patenting. Now, a machine-learning tool is aiming to speed up this process by scoring how ‘patent-like’ a scientific paper is — months or years before any deal, patent filing or spin-off company reveals its commercial potential. And it’s one of many proffering the same capability.
The tool, called the Translation Readiness Index (TRI), performs a linguistic analysis of a paper’s title and abstract. It then measures how similar a paper’s vocabulary is to publications that have previously been paired with patents. The method was developed by researchers at the data-analytics firm League of Scholars in Sydney, Australia. The work was posted as a preprint on arXiv1 and has not yet been peer reviewed.
“It’s a new way of triaging or ranking” research, says computational social scientist Paul McCarthy, co-founder of League of Scholars and a co-author on the preprint. The tool estimates the probability that a paper uses “patent-like language”, he says.
Picking winners
The researchers trained TRI on 20,610 scientific papers, including 9,431 that had been matched to patents. Titles and abstracts were fed into five classifiers, with the best-performing model having a 78% chance of ranking a patent-linked paper above an otherwise comparable paper that was not linked to a patent.
Papers that were eventually cited in patents included vocabulary such as ‘prototype’, ‘device’ and ‘design’ more often than did papers that were not cited in patents. TRI analyses only titles and abstracts, so doesn’t directly assess a paper’s underlying data or results.
To test whether TRI’s highest-ranked papers were correlated with other markers of commercial activity, such as whether co-authors have industry affiliations and if authors had previously patented research, the researchers looked at the 100 highest TRI-ranked papers by authors at the University of Western Australia (UWA) in Perth. The papers, published between 2019 and 2026, were more likely than a random sample to show those markers: 83 of the 100 papers had industry-affiliated co-authors, and 34 involved at least one UWA-affiliated author who had previously patented research, McCarthy says. As a result, the team is now testing TRI with several universities.
McCarthy doesn’t recommend basing investment decisions on the tool’s results alone because it’s a probabilistic ranking. However, he says it might help to uncover “unexpected gems”.
Ben Miles, co-founder of Empirical Ventures, an early-stage deep-tech investment firm based in London, says the tool could be useful as an external signal for academics and funders wanting to decide which ideas deserve further support from universities, governments or philanthropies before they are mature enough for investors.
But patentable technologies aren’t always commercially viable — and so any measurement of that will always be imperfect for investors’ needs.
Spin-off scouts
TRI is one of several research-scouting tools that aim to identify promising science. Some of the tools are being adopted in research institutes to help identify discoveries that could — with backing — become a viable business.
One such tool, called Haystack, was built for the technology team at Cornell University in Ithaca, New York, to scan the roughly 13,000 papers published each year that include Cornell-affiliated authors. That’s too many for Cornell’s technology-transfer team to inspect manually, says Matt Marx, who built the tool and is vice-provost for entrepreneurship, innovation and external engagement at the university.
Enjoying our latest content?
Log in or create an account to continue
Access the most recent journalism from Nature's award-winning team
Explore the latest features & opinion covering groundbreaking research
The company League of Scholars featured in this article has previously worked with the Nature Index as a data provider. This article was produced independently of those activities.
Facts Only
* League of Scholars is a data-analytics firm based in Sydney, Australia.
* The Translation Readiness Index (TRI) analyzes titles and abstracts of scientific papers.
* TRI was trained on 20,610 scientific papers, including 9,431 matched to patents.
* The best-performing model showed a 78% chance of ranking patent-linked papers above non-linked comparable papers.
* Vocabulary terms including "prototype," "device," and "design" appeared more frequently in papers cited in patents.
* Testing involved 100 high-ranked papers from the University of Western Australia published between 2019 and 2026.
* 83 of the 100 UWA papers had industry-affiliated co-authors.
* 34 of the 100 UWA papers involved authors with previous patents.
* Paul McCarthy is a co-founder of League of Scholars and co-author of the preprint.
* The research was posted as a preprint on arXiv and has not undergone peer review.
* Haystack is a separate research-scouting tool used by Cornell University.
* Cornell University publishes approximately 13,000 papers annually with affiliated authors.
Executive Summary
The Translation Readiness Index (TRI) is a machine-learning tool developed by League of Scholars in Sydney, Australia, designed to predict the commercial potential of scientific research by analyzing linguistic patterns in titles and abstracts. By comparing vocabulary to previously patented papers—specifically looking for terms like "prototype" and "device"—the tool aims to help investors and technology-transfer offices triage research before patents are filed.
The tool demonstrated a 78% probability of ranking patent-linked papers above comparable non-linked papers during training on over 20,000 documents. Testing at the University of Western Australia showed a correlation between high TRI scores and existing industry affiliations. However, the method relies on linguistic proxies rather than an assessment of underlying data or results. While proponents suggest it can uncover "unexpected gems" for funders and universities, critics and developers alike acknowledge that patentability does not guarantee commercial viability, leaving a gap between linguistic markers and actual market success.
Full Take
The strongest version of this narrative is that linguistic machine learning can democratize the discovery of "hidden" innovation, moving the identification of commercial potential from "who you know" to "what you wrote."
The current framing relies heavily on a specific persuasive loop: the developers of the tool provide the primary evidence for its efficacy. By citing their own preprint and the results of their own tests at UWA, the narrative positions the tool as a solution to a problem (the inefficiency of manual scouting) that the tool's own metrics define. This creates a circular validation where "patent-like language" is the goal, and the tool is the only measure of that language.
Patterns detected: ARC-0043 Authority Game
The underlying paradigm is the "quantification of intuition." It assumes that commercial viability leaves a detectable linguistic footprint. However, this risks creating a feedback loop in academia: if researchers know that words like "prototype" or "device" trigger investment signals, they may optimize their abstracts for the algorithm rather than the science—a form of "linguistic gaming" that could decouple a paper's score from its actual utility. This shifts the reward structure from breakthrough discovery to signal optimization.
Who benefits? Primarily technology-transfer offices and early-stage investors who can reduce labor costs. Who bears the cost? The "silent" innovator who produces transformative research using unconventional language that the TRI model deems "non-patent-like."
Bridge Questions:
1. Does the tool identify truly novel paradigms, or does it merely identify research that fits the existing linguistic mold of previous patents?
2. How would the tool's accuracy change if tested against papers that were highly cited but never patented?
Counterstrike Scan: An influence campaign would use this to push a "technocracy of innovation," claiming that human experts are obsolete and only AI can spot "true" value. The current content does not match this; it maintains necessary caveats regarding probabilistic ranking and commercial viability.
