What the Study Found
- Traces of chatbot writing in federal grant abstracts held steady through 2021 and 2022, then rose sharply beginning in 2023.
- Use did not spread evenly. Grants fell into two groups: one barely touched by these tools and another where roughly one sentence in ten showed machine influence.
- At both NIH and NSF, and at both the proposal and award stages, heavier AI involvement went with abstracts that were closer to what the agency had funded the year before.
- Among confidential NIH submissions, moving from light to heavy AI involvement came with about a four percentage point higher chance of funding. At NSF, no such link appeared.
- NIH awards with heavier AI traces produced about 5 percent more papers in their first year or two, but no more highly cited papers.
- All of these findings are associations measured within individual investigators. They do not show that AI use caused any of the outcomes.
There is no register showing which federal grant applications have been written with help from a chatbot. NIH and NSF applicants have not been required to say. So a Northwestern team looked for the traces instead.
They collected grant abstracts written in 2021, before ChatGPT reached the public, and treated them as a record of how humans write when asking the government for research money. Then they asked a version of GPT-3.5 to rewrite the same abstracts. That produced a second record. By comparing the words each version favored, the researchers built a ruler they could apply to later batches of grants. A Stanford group used the same approach to estimate chatbot involvement across more than a million papers and preprints.
They applied that ruler to more than 130,000 NSF and NIH awards and about 5,700 confidential proposals from two large American research universities. That second collection is unusual because it includes proposals that lost, documents that almost never leave the agencies that reject them.
Grant Writing Did Not Drift, It Split in Two
The share of machine-flavored sentences stayed level through 2021 and 2022, then rose sharply in 2023 as ChatGPT spread. That much was expected. What surprised the researchers was the shape of the change. Scores didn’t spread smoothly across a range, as they might have if everyone had started using chatbots a little. Instead, they piled up in two places. One group sat near zero. Another clustered around roughly one sentence in ten. Some scientists appeared not to use these tools in their abstracts at all. Others leaned on them heavily. There wasn’t much of a middle.
How much AI appears in a grant application is interesting. But what the application proposes can shape which science actually gets done. So the team turned each abstract into a string of numbers designed to capture its meaning rather than its wording. They then measured how far that meaning sat from everything the same agency had funded the year before.
Across all four collections of grants, public and private, NSF and NIH, more AI involvement went with less distance. The gap wasn’t huge, but it was consistent. Moving from a light AI touch to a heavy one pushed a proposal down about four or five places out of 100 in the distinctiveness rankings. The analysis compared each investigator with their own other proposals. That matters because it makes the result less likely to reflect cautious scientists simply being more likely to use chatbots. The same person, when writing with more machine help, wrote a less distinctive grant. A similar pattern has appeared outside research funding. In five brainstorming experiments at Wharton, groups using ChatGPT produced narrower sets of ideas. In one round, nine different participants independently gave their invented toy the same name.
The Sameness Is in the Ideas, Not Just the Prose
There’s an obvious objection. Maybe these tools simply smooth out everyone’s writing, making different ideas sound more alike. The team ran a test designed to separate wording from substance. They returned to the 2021 abstracts, had a language model rewrite them, and measured them again. The science had not changed at all. Only the language had. Yet the distinctiveness scores stayed essentially the same. That suggests the convergence lies in the ideas, not just the phrasing. “Science advances by exploring ideas that don’t yet look obvious,” said Dashun Wang, who directs Northwestern’s Center for Science of Science and Innovation. “If AI increasingly learns from yesterday’s successful proposals, one of the questions we should ask is whether tomorrow’s scientific portfolio becomes less adventurous.”
At NIH the Pattern Paid Off, at NSF It Did Not
Then comes the finding that makes this more than a study of writing habits. Among the confidential NIH submissions, proposals with heavier AI involvement were more likely to be funded. Across the span from light use to heavy use, the difference was roughly four percentage points. That’s substantial given how thin the odds have become. NIH funded about one in eight research project applications in 2025, down from about one in five two years earlier. Four percentage points can represent a meaningful share of an applicant’s chances. The same analysis found no such pattern among NSF submissions.
None of this shows that using a chatbot causes a grant to get funded, and the authors are careful about that. Comparing investigators with themselves rules out many possible differences between people. But it can’no’t rule out another possibility: scientists may be more likely to use these tools when the project they are pitching is already safe, conventional, and easier to build.
More Papers, Not More Important Ones
Looking at what happened after awards were made, NIH grants with heavier AI traces produced about 5 percent more publications. But when the researchers narrowed the count to papers in the top 5 percent of citations for their field and year, the advantage disappeared. In other words, the grants produced more papers without clear evidence that they produced more influential work. NSF awards showed no difference in publication output.
Several caveats matter. The confidential proposals came from only two universities, so they can”t represent the entire American research system. The analysis also examined abstracts rather than the full research plans that reviewers actually read. And the publication counts covered only the first year or two after an award, which may be too early to judge whether a grant produced lasting discoveries.
Agencies are already trying to write rules for a practice they can’no’t easily see. NIH said in July 2025 that applications substantially developed by AI will not count as an applicant’s original ideas. The agency noted that these tools had enabled some investigators to submit more than 40 distinct applications in a single submission round. NSF asks proposers to say whether and how they used generative AI. It also forbids reviewers from uploading proposals into outside chatbots.
NIH also runs programs, including the Pioneer and Transformative grants, created specifically for ideas the agency expects would not fare well in its own ordinary review. Those programs exist because the center of the portfolio is not always where the surprises live. That makes one line in this paper’s acknowledgments especially striking: the authors used GPT-5 and Claude Opus 4 to improve the language and readability of the manuscript.
- Study type: Observational text analysis of grant abstracts, with regressions comparing each investigator with their own other proposals. The findings show associations, not cause and effect.
- Sample: About 1,600 NSF and 4,100 NIH confidential proposal submissions from two large US research universities. The sample included funded, unfunded, and pending applications with start dates from 2021 to 2025. Researchers also analyzed roughly 57,000 NSF and 74,000 NIH publicly released awards from the same period.
- Measures: Researchers estimated the share of LLM-modified sentences by comparing word frequencies in 2021 human-written abstracts with GPT-3.5 rewrites of the same text. They measured semantic distinctiveness with SPECTER2 sentence embeddings against abstracts funded by the same agency the previous year. Publications were linked to awards through Dimensions.
- Manipulation: None in the main analysis. In one controlled check, a language model rewrote 2021 abstracts while keeping the science the same. This let the researchers test whether the distinctiveness result came from writing style alone.
- Duration: Grants had start dates from 2021 to 2025. Publication outcomes covered awards starting in 2023 and 2024, with citations counted through November 2025.
- Funding and conflicts: Supported by NSF Award 2404035. The authors declare no competing interest and disclose using GPT-5 and Claude Opus 4 to improve the language, style, and readability of the manuscript.
- Data availability: The authors publicly posted deidentified data behind the main figures. Access to the confidential university proposal data requires a data use agreement. Award records came from Dimensions, a commercial database.
- Main limitation: The confidential proposals came from only two institutions, and researchers analyzed abstracts rather than full applications. The design also cannot separate the effects of AI use from the kinds of projects scientists choose to use AI on.
Reference
Qian, Y., Wen, Z., Furnas, A. C., Bai, Y., Shao, E., & Wang, D. (2026). The rise of large language models and the direction and impact of US federal research funding. Proceedings of the National Academy of Sciences, 123(33). https://doi.org/10.1073/pnas.2601439123
Frequently Asked Questions
Does this mean using a chatbot will get my grant funded?
No. The study found that NIH proposals with heavier AI traces were funded more often, but it cannot show that AI caused the difference. Scientists may simply be more likely to use these tools when a project is already conventional and well defined. Such projects may already fit comfortably within established review expectations.
How can anyone tell a chatbot helped write a grant?
Not from any single document. The method works across large collections of text. It compares how often certain words appear with patterns found in human writing and machine writing. It can estimate AI involvement across a collection or give a rough score for one abstract, but it cannot deliver a verdict about authorship. It is a measuring instrument, not an accusation.
Why would NIH and NSF come out so differently?
The authors say plainly that they do not know, and their data cannot test the reason. One possibility they raise involves what reviewers reward. As co-author Yifan Qian put it, “review norms may more strongly reward incremental, executable projects that yield multiple publications, and LLM-assisted drafting may help proposals conform to those established templates.” The agencies also differ in what counts as research output. NSF grants, for example, often produce software, data, and training rather than papers.
Is less distinctive research automatically worse?
No, and the paper does not claim that it is. Careful, incremental work is how much of science advances. The concern is about the overall mix. Agencies are also expected to fund some risky and unusual ideas. If a portfolio keeps moving toward what has already worked, fewer of those unconventional bets may remain.
Was AI use in the proposals actually disclosed?
Mostly not, which is part of why the researchers needed a statistical measure. NSF encourages disclosure but does not require it, while NIH’s policy sets a limit on originality rather than requiring applicants to report AI use. The study detected AI traces statistically. Applicants did not report them.
Cite This Page
