MindยทGeorgetown University
Journal article ยท Peer-reviewed

A Misleading AI Summary Can Warp What You Remember

Georgetown and University of Washington researchers found that AI video summaries omitted the crash itself in 95% of test cases, and one misleading detail was enough to warp what people remembered from an accident they had watched firsthand.

What the Study Found

  • Misleading AI summaries cut correct memory recall from 83.6% to 44.8% in a controlled experiment.
  • AI-generated video summaries omitted 51.6% of key details on average, and 95% skipped the crash itself.
  • Knowing a summary was AI-written, not human-written, made no difference: memory distortion held steady either way.
  • People’s trust in AI made no difference either: skeptics and enthusiasts alike misremembered the same detail.

A RED car rolls up to an intersection, turns right, and crashes into a pedestrian crossing the street. Georgetown and University of Washington researchers found that telling people an accident summary was written by AI, rather than a human, did nothing to protect their memory of what actually happened. Everyone who read a summary that got a single detail wrong ended up misremembering that detail themselves, whether they believed a machine wrote it or a colleague did. The wrongness, it turns out, came from ChatGPT and Gemini, both of which routinely bungled the same simple scene when asked to describe it, the same overconfidence other researchers have clocked in chatbots that dig in on a wrong answer the longer a conversation runs.

Summaries like these already show up in body cam reports, meeting notes, and hospital discharge letters, places where nobody replays the source video to check the machine’s homework. The researchers wanted to know what happens to a person’s memory once that homework is subtly, confidently wrong.

Psychologists have shown for decades that memory bends under the weight of misleading information fed to people after they witness something, a phenomenon known as the misinformation effect. The classic demonstration used a short film of a car accident and swapped one detail in a follow-up account, and later replications recreated it with a red car, a stop or yield sign, and a pedestrian who gets knocked down and gets back up. Mattea Sim of Georgetown’s McCourt School of Public Policy, working with Yoshi Kohno at Georgetown and Yael Eiger at the University of Washington, described the work in a peer reviewed paper accepted to an AI ethics conference, and used those same modernized videos with the human note taker swapped for an AI one. Their question was whether a summary written by ChatGPT could plant the same kind of false memory, and whether knowing the summary came from a machine would change anything at all.

Substack Sign-up form screenshot

It could, and it did, in a big way. But the results about the AI’s own writing came first, and they set up just how much room there was for things to go wrong.

The Summaries Kept Skipping the Crash

The team prompted ChatGPT and Gemini to write neutral, facts-only summaries of two 25-second videos, five times each, then had researchers code every summary against a list of 19 details that mattered to the story: who did what, when, and where. All 20 summaries contained at least one error, and the most common failure by far was leaving something out. On average, the AI omitted 51.6% of the central details, and 95% of the summaries never mentioned the single most important moment in the video at all: the car actually hitting the pedestrian.

“I was struck by how bad the summaries were, even at this stage in AI development,” said Eiger, whose research focuses on technology in the carceral system. “It worries me that police departments may be using video summarization technologies without rigorous testing and without an awareness of how incorrect AI-generated summaries could be.”

Telling People It Was AI Changed Nothing

The second half of the study, a randomized experiment with 328 people whose answers were ultimately analyzed, put a live audience in front of that flawed AI voice. Participants watched one of the two accident videos, waited a day or two, then read a 21-sentence ChatGPT-written summary that either got the traffic sign right or swapped it for the wrong one, before answering questions about what they remembered from the original video. Among that group, 83.6% correctly recalled the traffic sign after reading an accurate summary, compared with just 44.8% after reading the misleading one. Before anyone read a word of it, some participants were also told the text came from a professional AI transcription tool, and others were told a human transcriber wrote it, even though every participant actually received text generated by ChatGPT. It made no measurable difference: 46.3% of people in the AI-labeled misleading group got the traffic sign right, against 43.4% in the human-labeled misleading group, a gap small enough to have happened by chance.

People’s general trust in AI didn’t change the outcome either, despite a ten-item survey built specifically to measure it. Whether someone scored as an AI skeptic or an AI enthusiast, a wrong detail slipped into their memory just the same, which tracks with other research finding that people readily forgive an AI’s factual slips as long as everything else about it feels right.

None of this happened because participants were checked out or careless: 96.6% of them correctly remembered that the accident happened during the day, a detail the summary never mentioned at all, and they averaged 91.4% accuracy across eleven other memory questions the misleading summary never touched. The forgetting was narrow and specific, landing exactly on the one detail the AI got wrong, which is what makes it look like a genuine misinformation effect rather than a fog of general confusion.

The researchers point to policing as the setting where this matters most, since departments already use AI to turn body cam footage into written reports, and eyewitness memory has been a fault line in courtrooms for as long as psychologists have studied it. None of the errors here rose to the level of the AI-generated police report that, according to a case the authors cite, once claimed an officer had shape-shifted into a frog after picking up dialogue from a Disney movie playing in the background. The errors this study documents are smaller and easier to miss, and that, the authors argue, is exactly the danger: a subtle, plausible-sounding mistake is far more likely to slide past a tired human reviewer than a shape-shifting officer is. Keeping a human in the loop is supposed to catch the AI’s mistakes, but this experiment suggests the loop can absorb the AI’s mistakes into its own memory instead.

The team is already thinking about a harder test: real body cam audio instead of a cartoon car crash, and officers instead of strangers on a gig-work platform reading the summary. If a fake yield sign can outlast a warning label announcing that the text was AI-generated, it’s an open question what a subtler error might do inside an actual police report nobody thinks to double check.

Reference

Sim, M., Eiger, Y., & Kohno, T. (2026). AI-Enabled Human Memory Manipulation: Misleading AI-Generated Summaries Distort Human Memory (Version 1). arXiv. https://doi.org/10.48550/ARXIV.2609.28820

  • Study type: Peer reviewed human-subjects experiment; accepted to the proceedings of the Ninth AAAI/ACM Conference on AI, Ethics, and Society (October 2026), full-length version posted on arXiv
  • Sample size: 328 U.S. adults in the memory experiment (recruited via Prolific); 20 AI-generated video summaries separately analyzed for errors
  • Intervention: An AI-generated summary containing one misleading detail about the video, labeled to participants as either AI-written or human-written
  • Comparator: An otherwise identical AI-generated summary containing accurate information, under the same AI-written or human-written labeling
  • Duration: 24 to 48 hours between watching the event and reading the summary
  • Funding / conflicts of interest: Robert L. McDevitt Chair in Computer Science at Georgetown University; U.S. National Science Foundation Award #2205171. No competing interests declared
  • Data availability: Full dataset and supplementary materials posted on OSF
  • Preregistration: Not reported
  • Main limitation: The test videos were simple animated recreations of a staged accident, and every participant read an identical AI-written summary regardless of the labeled source, which limits how far the findings generalize to longer, more realistic AI summaries or real body camera footage

FAQ

Why does it matter if a police report is written by AI instead of a person?

It matters because departments are already using AI to turn body cam footage into written reports, and those reports can shape what officers, courts, and juries remember about an incident. The Georgetown and University of Washington researchers found that a misleading AI summary changed what people remembered about an event they had personally watched, so the same effect could distort memory of a real encounter that nobody but the AI and a body camera actually saw in full.

Could warning people to double-check an AI summary fix the problem?

Warning people that a summary might be wrong is not something this study tested directly, but the two labeling conditions it did test suggest a bare tag isn’t enough on its own: participants told the text came from an AI were just as likely to misremember the traffic sign as participants who thought a human wrote it. Separate meta-analysis research on eyewitness memory finds that more substantive warnings, ones that explain why a piece of information might be wrong rather than simply flagging it as suspect, can meaningfully cut the size of the misinformation effect, which suggests any real fix would need to go well beyond a one-line disclaimer.

Why did the AI summaries leave out the crash itself instead of just describing it wrong?

Omission was simply the most common way ChatGPT and Gemini failed in this study, with every one of the 20 summaries the researchers coded missing at least one important detail. The car hitting the pedestrian, the single most central event in the video, was left out of 95% of the summaries, which the researchers connect to older psychology research showing that omitted information degrades memory almost as much as incorrect information does.

Is this a problem specific to ChatGPT and Gemini, or would other AI tools make the same kind of mistake?

The study only tested ChatGPT and Gemini, so it cannot say for certain how other AI models would perform on the same videos. But both models made similar kinds of errors despite being built by different companies, and the researchers point to other studies of AI video summaries finding comparable problems, which suggests the pattern is not unique to a single AI product.

  • Ben Sullivan

    Veteran journalist, 25 years ยท Science & business reporting ยท Founded ScienceBlog.com

    Ben Sullivan is a veteran journalist with 25 years of experience reporting on science and business across the U.S. and Europe. His work has appeared in premier outlets, including The Economist, The New York Times Magazine, the Los Angeles Times, and Prognosis, an English-language newspaper published in Prague. A digital media pioneer, Ben founded ScienceBlog.comย and led it for two decades. Under his leadership, the site was named one of the best science blogs "in the known universe" by Popular Science and was featured on Nature's year-end list of top science news blogs. Sullivan has consulted for the U.S. Department of State, served on the board of directors of the Los Angeles Press Club, was awarded a National Press Foundation fellowship to study health insurance, and taught writing at Loyola Marymount University's Asia Media International program. He lives in Los Angeles.

    MuckRack โ†— ยท LinkedIn โ†— ยท Editorial Policy & Correctionsโ†—

    https://orcid.org/0009-0007-1842-5997

Cite This Page

"A Misleading AI Summary Can Warp What You Remember." ScholarPeer, 26 September 2026, scholarpeer.com/a-misleading-ai-summary-warp-what-you-remember/.

Download RIS · Download BibTeX