방사선 의학 분야에서 대형언어모델(LLM) 연구가 폭증하고 있다. 하지만 이 연구들이 과연 투명하게 보고되고 있을까? 결론부터 말하면, 평균 준수율은 51.2%에 그쳤다.
튀르키예 바셰셰히르 도시병원 연구팀은 2025년 1월 1일부터 12월 26일까지 Web of Science 방사선 핵의학 의학영상 분야 Q1 저널에 발표된 LLM 연구를 대상으로 횡단면 감사(cross-sectional audit)를 수행했다. 평가 기준은 MI-CLEAR-LLM(의료 분야 LLM 정확도 보고를 위한 최소 보고 항목) 2025 업데이트 버전이었다.
201건의 적격 연구 중 102건이 최종 분석되었다. 전체 준수율은 평균 51.2% ± 14.7%(범위 22.2%~84.2%)로, 보통 수준에 그쳤다.
항목별로 보면 양극화가 뚜렷했다. 입력 데이터 유형(100%), 테스트 데이터 독립성(80.2%), 적응 전략(78.1%)은 비교적 잘 보고되었다. 반면 프롬프트 실행 설정(29.4%)과 확률적 관리(33.1%)는 대부분 누락되었다. 가장 적게 보고된 항목은 학습 데이터 기준일(9.8%)과 프롬프트 작성 근거(15.6%)였다.
저널 간 차이도 유의미했다(P=0.011). 특히 한국방사선의학회지(Korean Journal of Radiology, KJR)는 평균 준수율 72.8% ± 2.7%로 가장 높았다. 분석에 포함된 KJR 논문 4편 모두가 상대적으로 우수한 보고 품질을 보였다.
이 연구가 의미 있는 이유는 LLM 연구의 재현성 위기를 구체적 수치로 드러냈기 때문이다. 프롬프트를 어떻게 실행했는지, 학습 데이터의 기준일이 언제인지 명시하지 않으면, 다른 연구자가 동일한 결과를 재현할 수 없다. LLM의 출력은 프롬프트와 설정에 따라 크게 달라지므로, 이 정보의 누락은 단순한 서술상의 문제가 아니라 과학적 엄격성의 결여다.
다만 이 연구 자체도 한계가 있다. 2025년 한 해 동안의 논문만 분석했으므로, 시간에 따른 개선 추세를 파악하기는 어렵다. 또한 MI-CLEAR-LLM 가이드라인 자체가 2025년에 업데이트된 것이므로, 과거 연구들이 이를 따르지 않은 것은 일면 당연하다.
연구자와 임상의가 주목해야 할 점은 명확하다. LLM 연구를 수행할 때 프롬프트의 정확한 문구와 실행 환경, 학습 데이터 기준일, 확률적 설정(온도 매개변수 등)을 반드시 명시해야 한다. 독자가 결과를 재현할 수 있어야만, 비로소 과학적 증거로 인정받을 수 있다.
📖 *Reporting Quality of Large Language Model Studies: A Cross-Sectional Audit of High-Ranking Radiology and Medical Imaging Journals (횡단면 감사, 102편 분석)* |
논문 원문
※ 이 기사는 의학 논문을 바탕으로 작성되었습니다. 개인 건강 상태에 따라 다를 수 있으니 전문의와 상담하세요.
Large language model (LLM) research in radiology is exploding. But are these studies being reported transparently? The short answer: average compliance sits at just 51.2%.
A research team from Basaksehir Cam and Sakura City Hospital in Istanbul, Turkey, conducted a cross-sectional audit of original LLM research published between January 1 and December 26, 2025, in Q1 journals within the Web of Science "Radiology, Nuclear Medicine, and Medical Imaging" category. The evaluation tool was the MI-CLEAR-LLM (Minimum Reporting Items for Clear Evaluation of Accuracy Reports of Large Language Models in Healthcare) 2025 update.
Of 201 eligible studies, 102 were analyzed after quota-based subsampling. Overall adherence to MI-CLEAR-LLM was moderate at best: mean 51.2% plus or minus 14.7% (range 22.2%-84.2%).
Item-level analysis revealed striking polarization. Input data type was reported in 100% of studies, test-data independence in 80.2%, and adaptation strategy in 78.1%. In stark contrast, prompt execution setup appeared in only 29.4% and stochasticity management in just 33.1%. The least frequently reported items were training-data cutoff date (9.8%) and rationale for prompt wording (15.6%).
Inter-journal differences were statistically significant (P=0.011). Notably, the Korean Journal of Radiology (KJR) showed the highest mean adherence at 72.8% plus or minus 2.7%, with all four eligible KJR articles demonstrating relatively strong reporting quality.
This study matters because it quantifies a reproducibility crisis in LLM research. Without disclosing how prompts were executed, what the training data cutoff date was, or how stochasticity was managed, other researchers cannot reproduce the results. Since LLM outputs vary dramatically with prompt wording and configuration settings, this omission represents not merely a reporting gap but a deficiency in scientific rigor.
Limitations include the study's restriction to papers from 2025 only, making it impossible to assess temporal improvement trends. Additionally, the MI-CLEAR-LLM guideline itself was updated in 2025, so earlier studies naturally could not have followed it.
For researchers and clinicians, the message is clear. When conducting LLM research, prompt wording, execution environment, training-data cutoff dates, and stochasticity settings (e.g., temperature parameters) must be explicitly documented. Only when readers can reproduce results can studies be accepted as credible scientific evidence.
📖 *Reporting Quality of Large Language Model Studies: A Cross-Sectional Audit of High-Ranking Radiology and Medical Imaging Journals (Cross-sectional audit, 102 studies analyzed)* |
PubMed
※ This article is based on a medical research paper. Individual health conditions may vary; please consult a healthcare professional.