The AI That Writes Scientific Papers — and the Scientists Who Fear It

The abrupt change in language in scientific journals is accompanied by an unnerving silence. However, if you focus on the wording, it makes a strong statement. Words like “intricate and multifaceted” or the oddly common word “delve” have become increasingly popular—not by accident, but because ChatGPT and other AI language models have become ghostwriters in labs, subtly infiltrating the academic canon.

The AI That Writes Scientific Papers — and the Scientists Who Fear It
The AI That Writes Scientific Papers — and the Scientists Who Fear It

In one unusually particular instance, the sentence “Certainly, here is a possible introduction for your topic” appeared in a research paper published in Surfaces and Interfaces by Elsevier. It was a leftover AI prompt that was inserted straight into the body of peer-reviewed paper, not a metaphorical flourish. It got published, avoided editorial examination, and quickly gained notoriety as a warning. When contacted, Elsevier stated that they were looking into how this got through. However, Elisabeth Bik, a specialist in scientific integrity, was already raising the alarm: if that kind of error was obvious, what about the more difficult-to-identify AI-generated text?

AI Writing Scientific Papers – Key Facts

Item Details
Technology Large Language Models (e.g., ChatGPT, GPT-4)
Core Issue AI-generated text appearing in scientific research papers
Risks Identified Fabricated citations, hallucinated facts, unverifiable references
Key Figures Elisabeth Bik (Integrity Expert), Andrew Gray (Research Librarian)
Estimated Impact Over 60,000 scientific papers in 2023 may contain AI-generated content
Reference Example Elsevier journal published a paper with a clear AI-generated prompt
Source Scientific American, arXiv.org, PubMed, Scopus, Google Scholar

Bik’s worry stems from the way big language models work. These tools produce confident, fluid prose. However, truth is what they are unable to produce consistently. An artificial intelligence (AI) may confidently cite studies that do not exist, complete with believable-sounding author names and publications, unlike a careful researcher who verifies every citation. This behavior is a recognized glitch and is referred to as “hallucination.” Hallucinated footnotes are not merely awkward for scientists, who value accuracy. They pose a threat.

University College London librarian Andrew Gray discovered a startlingly high rise in terms statistically preferred by AI by searching through massive scientific databases such as Dimensions and PubMed. These days, words like “commendable,” “intricate,” and “meticulous” are disproportionately used in scholarly publications. His findings indicated that more than 60,000 papers probably had AI-generated content in 2023 alone. That amounts to about 1% of all the scientific material that was published in that year. It might be closer to 17.5% in computer science.

Not only is the number impressive, but the subtlety amazing as well. Only language patterning can identify the change. Copying and pasting an AI question into an abstract isn’t always the smoking gun. It’s the accumulation of phrases that are a tad too smooth, strangely universal adjectives, and an uncanny tone consistency across dozens of fields. Only those who know where and how to look can see it, much like a fake signature hidden inside a lab notebook.

Using phrases that ChatGPT frequently uses as breadcrumbs, Gray’s team was able to track them down through millions of documents. In the literature on the biological sciences, for example, the word “delve” has become exponentially more prevalent. It made 349 appearances in 2020. That figure skyrocketed to 2,847 by 2023. It has already reached 2,630 this year, a 654% growth. That is not merely a peculiarity of language; rather, it is a mirror reflecting the sometimes imperceptible transformation of scientific writing.

The word “delve” appeared four times in one biomedical journal, each time surrounded by a careful study, which caused me to pause.

Researchers are currently faced with this dilemma. When it comes to eliminating inappropriate terminology or expediting the writing process, AI writing tools are incredibly successful. These technologies can be particularly powerful for non-native English speakers, providing scientists a more distinct voice in an international dialogue. However, they can covertly introduce flaws, such as phony references and speculative data disguised as fact, when used carelessly, particularly in fields like engineering or medicine.

The current editorial structure is one aspect of the problem. Peer reviewers aren’t taught to recognize a chatbot’s tone. Journals continue to rely on trust—trust that the results presented are legitimate, that the author wrote their own abstract, and that the references provided are real. AI challenges that trust in more ways. It avoids it.

On the other hand, tools that purport to identify AI-generated text continue to be incredibly unreliable. Language is subjective, clumsy, and situational. Legitimate writing, especially polished work by non-native speakers, is frequently flagged by AI detectors as “too perfect.” Ironically, by penalizing writers for increasing their clarity, the very instruments designed to maintain integrity may create new kinds of prejudice.

Some publishers are creating new disclosure guidelines in response. Others are attempting to track the origins of scientific texts by using distinctive information or watermarking. However, these initiatives are still disorganized and far from conventional. For the time being, readers must maintain their skepticism and researchers must be open and honest.

However, there is no denying the possibility of responsible use. When used appropriately, AI can be a very useful helper. It can assist scientists with article outlines, data translation for wider audiences, and even the discovery of neglected cross-disciplinary linkages. Researchers could increase science’s accessibility without sacrificing its credibility by carefully utilizing these instruments.

Cultural adaption is now required. Scholars need to be taught by research institutions not only how to use AI but also when not to. Journals need to reconsider peer review in order to incorporate both methodological rigor and linguistic pattern recognition. Perhaps most importantly, scientists need to continue to be careful with their own writing because the rhythm of their words now has greater significance than ever before.

How clearly and ethically AI is integrated will define the future of scientific publishing, not whether it is used at all. It’s difficult to distinguish between authorship and aid. However, if we draw it carefully, we might find a balance that maintains both digital speed and human rigor—two forces that, if in harmony, could greatly advance rather than undermine the scientific process.