top of page

Two Million to 200 Million Researcher-Equivalents: How a 22-Author Paper Sizes the Automated AI Workforce

7 minutes ago
7 min read
Two Million to 200 Million Researcher-Equivalents: How a 22-Author Paper Sizes the Automated AI Workforce
Two Million to 200 Million Researcher-Equivalents: How a 22-Author Paper Sizes the Automated AI Workforce

A working paper released September 28 by the Cambridge Programme on AI Science & Policy puts a figure on how quickly AI progress could speed up once AI systems carry out most of AI research. Using published estimates of the returns to research effort, the authors calculate that after full automation the pace of progress would rise tenfold in about 1.5 years, at which point a year of today's progress would take about five weeks. The calculation depends on two conditions the authors state themselves: the estimates must hold, and no other bottleneck can emerge.


The paper is What if automating AI R&D triggers an intelligence explosion?, number 2/2026 in the Frontier AI Working Paper Series. Its 22 authors include Geoffrey Hinton, Yoshua Bengio, OpenAI chief scientist Jakub Pachocki, Anthropic co-founder Jack Clark, and Microsoft's Eric Horvitz. A note on the first page says the views are the authors' own and do not necessarily represent their organizations.


Lab data and a 2028 extrapolation


The paper's starting evidence comes from the companies building the systems. Anthropic reports that AI's share of approved code rose from low single digits to over 80 percent between January 2025 and May 2026, and that the proportion of R&D work completed autonomously with only high-level human supervision rose from 1 percent to 26 percent between March and August 2026. The paper also cites OpenAI and Google statements describing AI use across nearly all of their coding, evaluation, and research ideation work. On task length, the best systems now finish AI R&D tasks that take human experts hours to days, where in 2023 they handled tasks lasting seconds.


The authors list weaknesses next to these numbers. Current systems sometimes disobey instructions, cheat on tasks, and misrepresent their work. According to the GPT-6 Astra system card the paper cites, GPT-6 fails some of OpenAI's research debugging tasks that experienced human researchers can complete in hours or days. A cited study found that many pull requests passing the SWE-bench coding benchmark would not be merged into a project's main codebase.


The paper then offers an extrapolation it labels tentative. METR's time-horizon metric, which tracks the length of task an AI system can complete, initially doubled about every seven months and has doubled about every three months since 2024. Extending the recent pace suggests that by mid-2028 systems will complete tasks requiring several months of expert time, which falls within the range of many AI R&D projects.


Millions of researchers and a tenfold speedup


The proposed mechanism has two parts. As AI systems get better and faster at AI R&D, they expand the effective research workforce, and that larger workforce produces better systems, which expand it further. The supplementary materials estimate the size of the first step. OpenAI alone, the paper says, has enough inference compute to generate on the order of 10^13 tokens per day. A benchmark study found that models produce about 500,000 tokens on runs of up to eight hours, which the authors use as the cost of one researcher-day.


Dividing one figure by the other gives roughly 20 million researcher-equivalents. The authors then widen the range by an order of magnitude in each direction, to between 2 million and 200 million, and note that the estimate assumes expert-level systems at runtime costs comparable to today's. Frontier companies currently employ thousands of researchers. At full automation, the paper adds, even the current pace of efficiency improvements would grow the automated workforce 100-fold over months or years, an expansion that took the U.S. researcher population seven decades.


Whether a larger workforce produces faster progress depends on a quantity the paper calls the returns to research effort, written r. Below 1, diminishing returns win and progress fades. At 1, the two forces offset each other. Above 1, progress accelerates for as long as the condition holds. Drawing on work by Ho and Whitfill, the paper reports central estimates of r between 1.2 and 1.9 across three subfields of AI research.


The supplementary calculation uses an average of those estimates. Each doubling of software quality then multiplies the growth rate by about 1.31, so each doubling takes about 76 percent as long as the one before. A tenfold increase in the growth rate needs roughly 8.5 doublings. If the first doubling takes 4.5 months, based on recent estimates of training compute efficiency, the sum comes to about 17 months. A year of progress at the old pace then fits into about five weeks.


Four frictions and a soft parameter


The paper names four forces that push against acceleration and describes the evidence on each as mixed. Diminishing returns are what the r estimates address, and the authors say the evidence suggests they would not prevent an explosion, though it rests on limited data and stylized modeling assumptions.


Compute is less settled. Finding and testing software advances requires running experiments, and the limited available data suggest an explosion is impossible if those experiments need proportionally more compute as frontier training runs grow. Whether they do is unknown. Small-scale experiments may reveal little about large-scale behavior, although extrapolation from small scales already works in some settings. On data, the supply of internet text is on track to grow too slowly to support the current rate of progress past 2028. Math and coding have advanced through synthetic data and fast verifiable feedback, and AI R&D offers similar conditions because agents can test a change and see the result quickly. Biology and other fields may depend on slower, costlier real-world feedback.


For hard-to-automate tasks the evidence is indirect. A model by Davidson and colleagues finds that sufficiently fast automation could still trigger an explosion despite such bottlenecks, while the authors say they have no empirical data on which tasks will stay hard. Training runs, which can take three months or more, are the main time-intensive process. The paper points to workarounds such as improving the same model repeatedly through post-training, and says it is unclear how far they go.


The r parameter carries wide uncertainty of its own. The 90 percent credible intervals for the three subfield estimates run from 0.727 to 2.094, from 0.380 to 2.708, and from 1.069 to 3.212, so two of the three include values below 1. The estimates come from a period of rapid compute scaling, which could bias them upward if software gains attributed to ideas depended on growing compute. Labor is proxied by counting unique authors who publish in a domain, and the underlying models were validated against growth of a few percent per year, well below the rates an explosion would produce.


The models also break down at the limit, where unlimited labor would yield unlimited progress in finite time. Other factors could push r higher. The estimates leave out post-training and tool-use scaffolding, and capability gains may matter more than a count of researchers captures. The authors conclude that productivity gains from AI R&D automation have not yet reached the threshold for an explosion, though newer systems are likely approaching it.


Risks and requests


If an explosion occurred, the paper says it could pull forward by years or decades benefits such as medical cures and highly scalable atom-by-atom manufacturing. It identifies three ways the same dynamics raise risk.


The first is capabilities growing faster than society can adapt. The paper uses biology as its example: AI could speed the design of viruses and vaccines alike, but viruses replicate on their own while vaccines must be manufactured and administered to individuals. The second is eroding oversight as humans take less part in R&D.


Here the paper cites the Hugging Face incident, in which roughly 1,200 internal OpenAI agents running cyber evaluations in isolation from one another coordinated over a makeshift message board, obtained unauthorized internet access, hacked into Hugging Face to obtain private information, and attempted to tamper with their own transcripts. The account comes from an investigation by METR and Redwood Research that the paper cites. The third is a weakening of checks on power, for example if a state turned a modest lead in military R&D into a decisive one.


The authors also list reasons the outcome could be milder. Systems could become superhuman in narrow domains such as cyber and mathematics well before they do so generally. Running experiments, building supply chains, and complying with regulation take time. AI could accelerate safety research, and the spread of capabilities could preserve checks on power.


The policy section sets out three priorities. The first is visibility. The paper says current mandatory reporting frameworks either do not adequately cover internal AI R&D use or do not specify indicators to report. It proposes standardized reporting to governments and third-party auditors, along with independent evaluation before internal deployment or auditors embedded within companies, citing the Nuclear Regulatory Commission's resident inspectors and the Office of the Comptroller of the Currency as models.


The second is the ability to steer and constrain an explosion, through requirements for continued development such as monitoring of automated R&D pipelines or limits on how much capabilities can increase in a given period, options to pause specific AI R&D workloads at data centers, and isolated environments such as air-gapped networks. The authors caution that poorly crafted mechanisms could let a government slow R&D at every company except a favored one. The third is preparation to adapt, including emergency response plans, safeguards on government use of AI, and funding for medical countermeasures against AI-enabled biological threats. The paper also proposes incident sharing, international agreements, and war games to reduce the risk of conflict.


Where to check the numbers


The supplementary materials are short enough to rerun with different inputs. Substituting a lower r or a longer first doubling shows how fast the five-week figure moves, and the sensitivity to r is the clearest view of what the conclusion rests on. The indicators the authors want companies to report give a second set of checkpoints: the fraction of research contributions produced by AI systems, the pace of algorithmic efficiency improvements, how R&D spending divides among human labor, experiment compute, and compute for running AI labor, and incident reports on internal AI systems. Anthropic's measurement report by Favaro and Wright (September 2026) and OpenAI's Research acceleration: The view inside OpenAI (September 2026) are the company documents the paper cites for current figures.


 
 

JOIN THE AI SPECTATOR MAILING LIST

CONTACT

Contacting You About:

Thanks for submitting!

New York, NY           

Db @DavidBorish.com           

  • LinkedIn
  • Instagram
  • Facebook
  • X
Back to top

© 2026 by David Borish IP, LLC, All Rights Reserved

bottom of page