#13 The Ethical Use of Large Language Models in Academic Writing: A Path Forward
The craft of writing is undoubtedly difficult, and this seems doubly true in an academic environment. Not only are researchers responsible for conducting and interpreting the resulting data, but they also must carefully construct a written report that accurately embodies the meaning of their work while remaining interpretable by the average reader. It is perfectly understandable that Large Language Models (LLMs) such as ChatGPTTM and ClaudeTM have been seen by many scholars as potential time savers. After all, if their use can enable study data to be interpreted and disseminated far faster than in the past, then how could this not accelerate scientific progress?
The truth, however, is far more nuanced. Unlike past technological advances such as the development of word processors or statistical analysis software, LLMs do not simply facilitate tasks that ultimately still depend on the intellectual activity of the user, but they can also be used in ways that replace that intellectual activity. While a word processor (for example) permits rapid drafting, editing, and spellchecking of a document in a way that the manual typewriters of old do not, LLMs can draft the text itself de novo with only minimal prompting, thus blurring the boundaries between providing assistance with an intellectual task and performing the task itself.
Going Under the Hood: How LLMs Function
Were LLMs infallible and without bias in their interpretations, this would not be an issue, but neither of these is true. While LLMs appear at a surface level to generate intelligent responses, an understanding of the mathematical principles undergirding LLM function suggests that this is just that: an appearance. At their core, LLMs consist of a series of powerful mathematical functions that decompose written text into statistically useful numeric letter sequences called “tokens,” which are then processed with reference to a training dataset via iterative statistical operations. The results of this are then modified to increase their variance in a way that mimics human speech. The LLM's final response is thus a statistical prediction of what tokens could potentially follow the input string.
The quality of the final text that emerges from the LLM is quite impressive and reads for all intents and purposes as if it were created by a human. It is important to note, however, that there is no step in this process that corresponds to word definitions or meanings, and that, from the LLM's perspective (if they can be said to have one), a successful output string is one that is statistically plausible, nothing more. It thus follows that some proportion of responses will be linguistically reasonable but false in terms of content (i.e., hallucinated) and that the overall statistical distribution of output text will be reflective of the biases present within the training dataset. These concerns have been borne out by studies assessing the accuracy of references cited by LLMs when asked to write text suitable for scholarly publication in medicine. The results of these were concerning, with complete citation fabrication 16%-50% of the time and a variety of issues within the remaining references1,2.
Given this, the potential cognitive outsourcing noted above risks compromising the transparency and trustworthiness of the research process. It was this concern that led us to publish a framework that practicing academics can use to make better ethical decisions regarding the use of LLMs in their analytical and writing processes. The article can be viewed here. As healthcare simulation researchers, this article focuses largely on this subset of the medical literature. The approach we suggest, however, is content-neutral and thus is applicable across scholarly writing.
Principles from Current Authorship Guidelines
As an initial step we felt it important to survey what current medical journals require in terms of LLM use disclosure, and so we assessed recommendations published by the International Committee of Medical Journal Editors, the Journal of the American Medical Association Network, and the World Association of Medical Editors. This review uncovered several key principles that we used to formulate our final recommendations. First, LLM use should be explicitly disclosed to readers within the Methods section of the article in a manner like other software tools, such as statistical analysis software. The way the LLM was employed must also be disclosed. Was it used primarily to edit and improve the flow of the text, or was it used in data analysis (or both)? Finally, LLMs do not have legal standing and thus cannot meet definitions of authorship. It is therefore incumbent on authors using LLMs to recognize that any inaccuracies, hallucinations, or biases that creep into the work due to LLM use are their responsibility alone.
Developing a Framework: Common Use Cases
The principles above facilitate both transparency and accountability. After considering them, however, we believed that additional guidance was needed, not least because clearly positive LLM uses also exist. For example, LLMs can easily facilitate the translation of text by authors who do not natively speak the dominant language of the journal, making their findings accessible to a wider range of readers. The question, then, is how authors can make informed, ethical choices about specific use cases. Thus, a more elaborate framework was needed.
In constructing this framework, our primary considerations were based on the underlying functionality of LLMs as described above along with the likelihood that ongoing LLM use might have a long-term deleterious effect on key skills due to repeated “outsourcing” of cognitive tasks. This latter concern is supported by a recent study that assessed electroencephalographic (EEG) patterns and subject recall in authors using LLMs to perform writing tasks as compared to those that did not3. Not only did the authors who used LLM assistance have fewer EEG correlates of interconnected neural activity, but they often could not recall the material they had prepared. These two considerations led us to develop three lists based on current use cases: ethically acceptable uses, ethically contingent uses, and ethically suspect uses.

Figure. Overview of ethical use cases (ethically acceptable, ethically contingent, ethically suspect) of AI generative tools in writing.
Ethically acceptable uses included those applications primarily focused on grammar, flow, translation, and the syntactic elements of language. Since LLMs are specifically designed to produce quality written work, these uses play to the strengths of the technology. In these situations, authors provide the LLM with previously written text containing their own thoughts, interpretations, and overall message and ask the LLM to improve its readability, grammar, and flow.
Ethically contingent uses include applications in which the LLM is either generating relatively larger amounts of text with less user input or summarizing longer bodies of text in a way that borders on interpretation. Each of these cases gives the LLM a wider field of play in which hallucinations and bias can creep in, and each potentially outsources cognition that may best be performed by the author to the LLM. While there are legitimate uses within this category, the concern is that long-term employment of LLMs for these tasks could result in the erosion of critical thinking over time.

Figure. Overview of ethically suspect scenarios.
The final category, ethically suspect uses involves the use of LLMs to develop whole sections of the final text with minimal to no cognitive input from the author, such as asking the LLM to write an article introduction about the management of hypertension in adults based on the current literature without any additional input. It also includes the use of LLMs in the development of new concepts and/or in the interpretation of raw study data prior to the author performing these tasks themselves. With the potential for LLM hallucination and bias remaining undetected by authors, and given their lack of intellectual engagement coupled with the likelihood of skill degradation from cognitive outsourcing, these uses should largely be avoided.
Extending the Framework: A Decision-making Heuristic
While the above may seem comprehensive, it is almost certain that potential applications beyond those listed will be developed. We chose to address this reality by transforming this model from a static list into a more dynamic decision-making process based on four questions. These questions are:
- Have I used generative AI in a fashion that ensures the primary ideas, insights, interpretations, and critical analysis within the manuscript are my own?
- Have I used generative AI in a fashion that ensures humans will maintain competency in core research and writing skills?
- Have I double-checked to ensure that all the content (and references) in my manuscript are accurate, reliable, and free from bias?
- Have I disclosed exactly how generative AI tools were used in writing the manuscript and which parts of the manuscript involved the use of generative AI?
If the answer to all these questions is yes, then the LLM use in question is likely ethical. If not, however, then deeper consideration (and perhaps reconsideration) of the specific use is needed. We believe that this questioning process is robust enough to provide ongoing guidance as the technical sophistication of LLMs continues to evolve.
Closing
In summary, LLMs offer the potential for clearer communication and more efficient scholarship but also carry with them the possibility of undetected hallucination, bias, and skill erosion. We offer the ethical framework above to the academic community in hope that it will assist researchers in navigating these tricky waters and prove a trustworthy resource as technology continues to evolve.
References
- Athaluri SA, Manthena SV, Kesapragada V, Yarlagadda V, Dave T, Duddumpudi RTS. Exploring the boundaries of reality: investigating the phenomenon of artificial intelligence hallucination in scientific writing through ChatGPT references. Cureus. 2023;15(4): e37432.
- Bhattacharyya M, Miller VM, Bhattacharyya D, Miller LE. High rates of fabricated and inaccurate references in ChatGPT-generated medical content. Cureus. 2023;15(5): e39238.
- Kosmyna N, Hauptmann E, Yuan Y, Situ J, Liao X, Beresnitzky A, et al. Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task. ArXiv Computer Science>Artificial Intelligence: ArXiv; 2025 [Available from: https://arxiv.org/abs/2506.08872.]
Author

— by Aaron W. Calhoun, MD, Professor, Department of Pediatrics; Division Chief, Pediatric Critical Care, University of Louisville, Norton Children's Hospital, 9/2026
Continue the conversation! Please email us your comments to post on this blog. Enter the blog post # in your email Subject.