EconomyยทCenter for Global Development
Policy brief

AI Research Trials Are Booming, But the Old Playbook Won’t Work

Economists are randomizing AI faster than they can finish the trials meant to test it, and most tools will not survive to see their own results published.

What the Study Found

  • AI-related trials now make up about one in five economics registrations, up from one in thirty just four years ago.
  • High-income countries still register more AI trials, but low- and middle-income countries are catching up at a similar pace.
  • AI trials involve national governments less often (10.7%) than the wider trial pool (14.1%), a 3.4-point gap.
  • The average AI trial takes about 11 months to run, often too slow to outlast the model it’s testing.

EVERY randomized controlled trial, or RCT, in economics starts with the same small bureaucratic act: a researcher registers the study, months before anyone is enrolled, describing an intervention nobody outside the team has run yet. Registry records now show artificial intelligence built into close to one in five of those studies, up from one in thirty just four years earlier. The tools being tested keep changing faster than the process built to test them, and the standard randomized trial, designed for something that holds still, was never built for that. That gap, more than the raw growth number, is what the rest of this is about.

Money is pouring in to keep the growth going. Anthropic’s Economic Futures Research Fund alone has committed $200 million to large scale trials, and J-PAL, the Abdul Latif Jameel Poverty Action Lab, and other funders have launched calls of their own. None of that money guarantees the studies it buys will still mean anything by the time they’re finished.

The two registries behind these numbers, the American Economic Association’s RCT Registry and the Registry for International Development Impact Evaluations, are where economists log a trial’s design before results exist, and together they hold 10,214 research efforts, the raw dataset behind everything that follows. Total registrations climbed from a little under 1,000 in 2019 to around 1,650 by 2024, held roughly steady in 2025, and are on pace for a similar count in 2026, going by the roughly 900 filed by midyear. Of the full total, 739 are tagged as AI related, using descriptions mentioning machine learning, large language models, chatbots, computer vision, or named tools like GPT or Claude. That’s the one in five now, against the one in thirty from four years back.

Substack Sign-up form screenshot

The growth isn’t spread evenly. High income countries, home to just 17% of the world’s population, still register a disproportionate share of AI trials, though the low and middle income share is climbing at roughly the same steepness. Only 66.4% of AI related records even disclose a country, so this picture has real gaps in it.

Sorted by topic, AI trials cluster hard around behavior, education and labor, together over half of them, more concentrated than development economics research generally. Firms and labor studies are especially overrepresented next to Jessica Leight’s separately compiled sample of 1,497 development articles from 2021 to 2025. Where the money and attention concentrate says as much about what economists think AI is for as any of the results will.

The Trials Run Slower Than the Technology

The profession has been burned by an uncoordinated pile-up before. So many disconnected mobile health pilots hit Uganda at once that in 2012 the government simply banned new ones, a pattern researchers now call pilotosis. AI funding is arriving faster and in bigger checks than mobile health funding ever did, which is exactly the condition that produced pilotosis the first time.

The clearest evidence of the mismatch is timing. Across 741 AI related efforts with usable planned dates, the average trial runs 11.3 months from first enrollment to final measurement, and the wider category of randomized trials can take two years, or even longer, from planning through baseline to midline and endline; either span is long enough for the underlying model to be retired, renamed or quietly upgraded before a single result reaches a paper. Cost complicates it further: AI makes a digital intervention cheap to build, but if evaluation costs don’t fall at the same rate, the evidence budget eats a bigger share of every grant. There’s a subtler problem too. Which team built the tool may matter more than which tool got tested, since a skilled team’s mediocre model can outperform a clumsy team’s good one. And what does a control group even mean once the general purpose model everyone else is using keeps quietly improving too, a version of the same instability education researchers have separately flagged in AI classroom trials?

Government involvement, the thing most likely to get a pilot scaled rather than shelved, is lower for AI trials than for the registry as a whole: 10.7% against 14.1%, a gap of 3.4 percentage points. A trial nobody in government helped design is a trial nobody in government is likely to act on.

Not every AI trial is doomed to irrelevance. One eight week classroom study, run by Fab AI with Google DeepMind and Sierra Leone’s Ministry of Education, tested a guided learning tool on over 1,700 junior secondary math students and found a gain of 0.26 standard deviations, fast enough to beat its own subject matter’s shelf life. It’s the exception the authors point to, not yet the rule.

Fixing the Plumbing, Not Just the Method

The fixes on offer are mostly about plumbing rather than method: a shared tracker that captures midline results before the paper exists, a registry for pilots that quietly died before results ever came, and review boards mixing economists with technologists and policymakers rather than academics alone. Funders, in this telling, should also reward projects with local researchers and government co-design from the start, on the theory that scale follows relationships more reliably than it follows novelty. One proposal borrows from marketing science: test not just the best model available, but a cheaper Option C version closer to what a government could actually afford to run at scale. None of it will make a fast moving technology hold still for an evaluation, but it might stop so many separately funded trials from asking the same three questions in three different countries.

The registries will keep filling up regardless. The funding is committed, the interest is real, and the next wave of forms is already being filed somewhere this week. Whether any of it produces evidence a government still trusts by the time a slower reader gets around to the paper is the open question no registry can answer on its own.

Reference

Ohlenburg, Hanney, Kazemi, Levine and Wani, AI RCTs Are Booming, To Be Useful, They Must Evolve, Center for Global Development, August 27, 2026.

  • Study type: Observational analysis of trial-registry metadata; grey literature CGD Note, not peer reviewed.
  • Sample size: 10,214 registered research efforts (739 AI-related), AEA RCT Registry and RIDIE.
  • Corpus: All AEA RCT Registry and RIDIE entries, 2019 to mid-2026 (2026 partial year).
  • Analytic frame: Keyword and metadata classification of AI-relatedness, income group and research topic from registry entries.
  • Period covered: 2019 through mid-2026 (2026 partial year).
  • Funding / conflicts of interest: Not declared for this note; standard CGD disclaimer that publications reflect the authors’ views and CGD takes no institutional position.
  • Data availability: Not stated in the note; underlying entries are the public AEA RCT Registry and RIDIE listings.
  • Main limitation: AI-relatedness is classified from self-reported keywords and metadata rather than a verified technical audit of each trial, so some misclassification is likely in either direction.

FAQ

Why does a slow-moving trial hurt an AI experiment more than an ordinary one?

A slow-moving trial hurts an AI experiment because the average AI-related trial takes about 11 months to run, and the underlying model can be retired, renamed or upgraded well before that window closes. A result about one specific version of a tool can quietly stop describing anything a user will ever touch again.

Could a shared results tracker actually cut down on wasted AI pilots?

A shared tracker could help by surfacing midline results and failed pilots before the final paper appears, letting other teams see early signals or avoid repeating a pilot that quietly died elsewhere. It would not speed up any single trial, but it might stop so many teams from separately rediscovering the same dead end.

Is government involvement really necessary for an AI pilot to scale?

Government involvement is not strictly necessary, but the note’s own numbers suggest it matters: AI trials involve a government partner less often than the registry average, 10.7% against 14.1%, and a pilot with no government co-design behind it is less likely to be picked up and scaled by one afterward.

Why are so many AI trials about education and labor rather than health or governance?

So many AI trials focus on education and labor because that is where economists appear to see the clearest productive payoff from the technology right now, judging by where registrations concentrate. Behavior, education and labor together account for over half of AI-related trials, more concentrated than development economics research generally.

What would make an AI trial’s results useful even after the model it tested is gone?

An AI trial’s results stay useful after the model disappears if they test something more durable than one product version, such as how a team’s technical skill shapes outcomes, or how a fixed intervention design performs against a cheaper fallback model. The authors call this a push for model-agnostic evidence, a goal the field has not yet solved.

Cite This Page

"AI Research Trials Are Booming, But the Old Playbook Won’t Work." ScholarPeer, 28 August 2026, scholarpeer.com/ai-research-trials-are-booming-but-the-old-playbook-wont-work-ai-research-trials/.

Download RIS · Download BibTeX