Fictional experts are leaving real traces online
A new preprint is raising alarms about a strange byproduct of generative AI: fictional identities that are being repeated so often by language models that they are beginning to contaminate parts of the academic and online record. According to reporting published Aug. 27, 2026, names such as Elena Vasquez and Marcus Chen have appeared across hundreds of AI-generated papers, articles and books despite not belonging to real people.
The researchers behind the paper, from Samsung and the University of Warsaw, describe these recurring identities as a kind of AI haunting. Their central claim is not merely that large language models invent names, which is expected, but that they tend to invent the same names repeatedly in the same kinds of contexts. Over time, that repetition can seed false authority into search results, repositories and publications that look legitimate on the surface.
This matters because names carry trust. In academic publishing and research communication, readers often use an author’s name, institutional role or subject expertise as a shortcut for credibility. If AI systems repeatedly generate plausible-sounding experts and those names make their way into apparently formal documents, the barrier between fabricated content and the real scholarly record becomes much thinner.
How repeated name patterns emerge
The paper focuses on what the authors call correlated name priors. In plain terms, some models do not generate fictional identities randomly. They appear biased toward certain names when prompted for particular types of experts, authors or characters. The reporting gives several examples: Elena Vasquez, Marcus Chen, Elara Voss, Aris Thorne, Lena Petrova and Elena Amara Okafor all reportedly show up with unusual regularity depending on the model and prompt context.
The phenomenon extends beyond single names. The researchers say models also generate “correlated character ensembles,” meaning some names are more likely to appear together. That suggests the models are not only repeating isolated labels but also reproducing clusters of identity patterns. Once that happens, fake authorship can become more scalable. A user generating academic-style text, books, biographies or websites may receive the same authoritative-seeming cast of invented people again and again.
What makes the issue more serious is persistence. A fake expert created in one AI output might be trivial on its own. But if that same identity reappears across repositories, websites and documents, it can accumulate a digital footprint. Search engines and citation systems may then amplify the illusion, making it harder for readers to determine whether a source is fabricated or merely obscure.
From model quirk to publishing problem
The article ties the pattern directly to academic publishing infrastructure. The researchers reportedly searched databases for names they knew models frequently produced and found those names embedded in a broad set of AI-generated materials. One of the most striking examples cited is Zenodo, the CERN-operated repository that issues real DataCite DOIs. The team says it identified 1,655 records attributed to ghost authors and linked to nonexistent journals with fabricated publication dates.
That detail matters because a DOI is one of the strongest signals of legitimacy in scholarly communication. It does not guarantee quality, but it signals that a record has entered an organized publishing and indexing ecosystem. If fabricated journals, dates and authors can ride on top of real DOI infrastructure, then trust failures are no longer confined to spam blogs or low-grade content farms. They begin to touch systems used by researchers, librarians and institutions.

The article also connects the pattern to earlier reporting on AI-generated “research” sold as if it were human-produced. In one cited case, a company offering medical research had listed a founder and lead methodologist named Elena Vasquez, who did not exist. After reporting on the matter, that name was reportedly removed from the company’s site. Another example described a false statement circulated on Facebook that was attributed to a nonexistent “executive director Dr. Elena Vasquez.”
Taken together, those cases suggest the problem is not purely academic. The same fabricated identities can migrate across health claims, social rumors, commercial services and publication platforms. That cross-domain spread is what gives the phenomenon its broader relevance.
Why the contamination risk is growing
Generative AI lowers the cost of producing convincing text at scale. That includes papers, reports, web pages and book-like documents that can mimic institutional formats. If models also tend to reuse a small pool of persuasive fictional names, bad actors do not even need to invent a fresh persona each time. The model does part of the work for them by supplying recurring identities that already sound credible.
The danger is not limited to deliberate fraud. Some users may not verify generated names closely, especially when producing draft content quickly. In automated or semi-automated publishing pipelines, a fake expert can slip into public records because no one checks whether the individual exists. Once published, the identity can be cited, scraped, indexed and reused elsewhere.
This creates a feedback loop. Repetition by models leads to repetition on the web, and repetition on the web can in turn reinforce what future models encounter. Over time, the distinction between a model artifact and a documented person can blur for anyone relying on search rather than direct verification.
An integrity challenge for research systems
The reporting stops short of presenting a complete solution, but it makes the scale of the integrity problem harder to dismiss. Academic publishing has long dealt with plagiarism, paper mills and fake peer review. AI ghost authors introduce a different type of failure: synthetic identity pollution inside systems that were built around assumptions of human attribution.
That challenge touches repositories, editors, publishers, indexing services and search platforms. Stronger author verification, better screening for fabricated journals and more scrutiny of suspiciously repetitive identity patterns may all become necessary. The larger point is that AI misuse is not only about whether text is machine-written. It is also about whether the people attached to that text exist at all.
The preprint’s findings are likely to resonate because they describe a subtle but structurally important shift. Language models are no longer just generating flawed prose or fabricated citations. They may also be standardizing fake experts and quietly inserting them into the knowledge ecosystem. If that process continues unchecked, the academic record could become harder to audit precisely because the errors look so ordinary.
This article is based on reporting by 404 Media. Read the original article.
Originally published on 404media.co




