With over a decade of experience training academic researchers and editors, Hema has supported authors in publishing across major scholarly platforms. Her work has appeared in outlets such as The Scholarly Kitchen, Springer Nature Communities, SAGE’s Social Science Space, Disability Horizons, and the European Association of Science Editors blog, alongside other international platforms. She also speaks at academic conferences and TEDx events on research, publishing, and innovation.
~~~~~~~~~~~~~
As artificial intelligence becomes increasingly woven into academic work, its influence is extending far beyond routine drafting and editorial assistance into intellectually demanding domains once considered distinctly human. Among the most challenging is peer review—a process that depends on technical precision, disciplinary expertise, interpretive judgment, and nuanced communication.
In a study I conducted last year (documented in a report that included transcripts of my conversations with GPT), I examined whether AI could make this specialized discourse more accessible by accurately simplifying reviewer feedback for people with cognitive disabilities. To test this, I used two prompting strategies: a general plain-language prompt and a prompt explicitly focused on cognitive accessibility. I applied both to ten technically complex reviewer comments, running each prompt twice to assess consistency.
The findings raised substantial concerns. For example, one reviewer comment referenced a difference-in-differences approach to explain the timing and scale of market reactions to central bank announcements—a concept rooted in causal inference. GPT-4 simplified this into language about “how fast and how strongly the market reacts,” shifting attention away from the crucial issue of causality. In another case, “endogeneity” became “hidden effects,” a phrase that may sound clearer but strips away the term’s precise econometric meaning. Endogeneity does not merely refer to unseen influences; it describes statistical bias in explanatory variables that often requires advanced correction.
Conceptual distortion also appeared in broader theoretical language. “Bounded rationality” became “limited thinking ability,” a revision that both mischaracterizes the concept and risks sounding reductive or offensive. Bounded rationality refers not to intellectual deficiency but to decision-making under constraints such as limited information or time.
Across multiple outputs, inconsistency compounded these problems. In some cases, GPT retained technical terms like “sensitivity analysis” without explanation; in others, it replaced them with vague phrases such as “double-checking your methods.”
Yet identifying weaknesses in AI’s ability to interpret reviewer feedback addresses only one side of peer review. Equally important is whether AI can evaluate author responses to that feedback. While some researchers may use AI primarily to clarify reviewer comments, others may also expect it to function as an evaluative tool—helping determine whether draft rebuttals genuinely address reviewer concerns before submission. This raises a more consequential question: if AI struggles to preserve conceptual precision when interpreting reviewer feedback, can it reliably judge whether an author’s response resolves that feedback?
To investigate this, I conducted a comparative evaluation this year of four AI systems—Perplexity, GPT, Claude, and Gemini. The study used five interdisciplinary cases based on published research abstracts in economics, sociology, computer science, security studies, and sports analytics. For each case, I created a reviewer comment, a hypothetical author response, and an independent human expert evaluation. Across the five cases, reviewer comments contained at least three major evaluative concerns each, producing 15 distinct criteria and 60 total observations across the four systems.
Significant weaknesses emerged, particularly when reviewer comments relied on rhetorical phrasing, implicit critique, or disciplinary nuance.
For example, one reviewer wrote, “I fail to understand ‘self-reported survey.’” Rather than requesting a basic definition, this phrasing likely signalled scepticism about the methodological justification for self-reporting. Yet the author response presented to AI systems interpreted the comment literally, merely explaining what self-reported surveys are. Most systems praised this response, overlooking the reviewer’s deeper concern.
Similarly, when a reviewer criticized a manuscript for offering a disorganized “laundry list” of concepts, AI systems often accepted expanded definitions as sufficient. But defining terms does not address criticism about prioritization, conceptual hierarchy, or structural coherence.
In another case, reviewers challenged the use of the asset turnover ratio because its denominator could be artificially deflated. The authors responded by substituting the working capital ratio, arguing that its numerator was more representative. However, they failed to explain how this substitution resolved the denominator problem itself. Most systems did not detect this gap.
A similar pattern appeared when reviewers questioned why banks were included in a study of transitions to knowledge economies, noting that banks were already knowledge-intensive and therefore not novel examples. The authors responded only by explaining why banks are knowledge-intensive—a point the reviewers had not disputed. Several AI systems again failed to recognize that the central issue of novelty remained unresolved.
Taken together, these findings do not suggest that AI is incapable of contributing meaningfully to scholarly communication; AI systems can provide substantial value in tasks such as drafting, summarizing surface-level content, improving readability, organizing ideas, or helping users identify obvious omissions. In these domains, speed and linguistic flexibility can enhance efficiency and broaden access.
However, for more complex tasks, current systems remain unreliable in ways that can be consequential. For technology firms, this means that refining academic AI tools should focus on domain-sensitive reasoning. For journals, peer reviewers, and authors, AI should be treated as a supplementary instrument rather than an adjudicative authority that requires oversight. For mentors, supervisors, and senior scholars, human guidance continues to be relevant. Future research should expand beyond interdisciplinary pilot cases to test performance across additional subject areas, such as medicine, law, and engineering, where forms of critique differ substantially. Studies could also examine whether specialized fine-tuning, reviewer-style datasets, or collaborative human-AI models improve evaluative reliability.
Ultimately, while AI can assist academic communication, the nuanced discernment required for the peer review process still depends fundamentally on human expertise.


Leave a Reply