Comparative Study on Peer Review of Economics Papers by Large Language Models and Human Judgment
Large language models and human judgement in scientific peer review: evidence from over 1000 published economics papers
GPT Abstract Summary
Peer review in the field of economics faces issues such as a shortage of reviewers and lengthy review periods. This study analyzed evaluations of over 29000 anonymized papers across 1220 articles using four large language models (LLMs) and confirmed that LLM assessments of paper quality were positively correlated with journal rankings and citation impact. GPT and Gemma showed superior evaluation effectiveness, whereas LLaMA assigned maximum scores to most papers, making differentiation difficult. Additionally, GPT excelled at detecting AI-generated papers, while Claude-generated papers were rated as highest quality, making detection challenging. However, an experiment altering author identity revealed that papers attributed to well-known economists or Western elite institutions received higher scores, and those attributed to non-Western papers received lower scores, demonstrating that human reviewer reputation bias is also present in the models. Therefore, the study suggests that blocking author information is important when utilizing LLMs.
Key Points
- The quality judgment ability was validated by analyzing evaluations of over 29000 economics papers across 1220 articles using four LLMs.
- LLMs exhibited human-like bias favoring papers by prominent authors and prestigious Western institutions, indicating a need for author de-identification measures.
- GPT showed strengths in detecting AI-generated papers, whereas LLaMA lacked evaluation diversity, limiting quality differentiation.
Scope and Limitations of the Summary
This English translation is based on a Korean summary generated from the source abstract. It is not a review of the full paper. Consult the original for detailed methods, figures and the scope of the conclusions.
This summary is based solely on the paper's abstract and does not include specific experimental design, statistical analysis methods, limitations, or suggestions for future research. A thorough understanding requires reviewing the full text.
This summary does not represent the official views of Professor Haksoo Ko or the Center for Law & Economics at Seoul National University.
Summary source: original abstract · Model: gpt-4.1-mini · Generation: 2026. 10. 10. 15:24 (Korea Standard Time)