NEPHROLOGY / RESEARCH PAPER
 
KEYWORDS
TOPICS
ABSTRACT
Introduction:
Artificial intelligence (AI)-based large language models (LLMs) are promising tools for clinical decision support, but their clinical performance in specialized fields such as nephrology is still uncertain. ChatGPT and Claude represent distinct AI architectures with potentially different clinical utilities. We aimed to compare the diagnostic accuracy, treatment recommendations, and overall clinical utility of these two AI models in managing real-life difficult nephrology cases.

Material and methods:
Twenty-two real nephrology cases from a tertiary care university hospital were presented to both models, covering disorders such as glomerulonephritis, acute kidney injury, vasculitis, and transplant complications. Each model’s output was assessed for diagnostic accuracy, risk evaluation, test recommendations, and treatment planning. Three independent nephrologists evaluated the responses using the Quality Assessment of Medical Information (QAMAI) and Global Quality Score (GQS) tools. Statistical comparisons were performed using the Wilcoxon signed-rank test, with p < 0.05 considered significant.

Results:
Claude achieved higher diagnostic accuracy than ChatGPT (4.59 ±0.41 vs. 4.36 ±0.48; p = 0.048), whereas ChatGPT scored better in clarity (4.63 ±0.30 vs. 4.32 ±0.29; p = 0.002). No significant differences were found in relevance, completeness, usefulness, or source citation. Overall QAMAI scores were comparable between the two models (ChatGPT: 23.72 ±1.46; Claude: 23.39 ±1.43; p = 0.371). Inter-rater reliability ranged from moderate to good, with the highest agreement observed for ChatGPT’s GQS.

Conclusions:
Both ChatGPT and Claude demonstrate notable potential as decision-support tools in nephrology. Claude provided slightly higher diagnostic accuracy, while ChatGPT offered greater clarity. Despite these promising results, clinical judgment remains essential when interpreting LLM-generated suggestions.
REFERENCES (14)
1.
Benary M, Wang XD, Schmidt M, et al. Leveraging large language models for decision support in personalized oncology. JAMA Network Open 2023; 6: e2343689.
 
2.
Hu Y, Liu J, Jiang W. Large language models in nephrology: applications and challenges in chronic kidney disease management. Renal Fail 2025; 47: 2555686.
 
3.
Pal A, Sankarasubbu M, editors. Gemini goes to med school: exploring the capabilities of multimodal large language models on medical challenge problems & hallucinations. Proceedings of the 6th Clinical Natural Language Processing Workshop; 2024.
 
4.
Gomez-Cabello CA, Borna S, Pressman SM, Haider SA, Forte AJ. Large language models for intraoperative decision support in plastic surgery: a comparison between ChatGPT-4 and Gemini. Medicina 2024; 60: 957.
 
5.
Preiksaitis C, Ashenburg N, Bunney G, et al. The role of large language models in transforming emergency medicine: scoping review. JMIR Med Inform 2024; 12: e53787.
 
6.
Ao G, Chen M, Li J, Nie H, Zhang L, Chen Z. Comparative analysis of large language models on rare disease identification. Orphanet J Rare Dis 2025; 20: 150.
 
7.
Li Y, Dong J, Liu D, et al. Systematic benchmarking of large Language models in programmed cell death-oriented gastric cancer research: a comparative analysis of DeepSeek V3, DeepSeek R1, and Claude 3.5. Discov Oncol 2025; 16: 1227.
 
8.
Achiam J, Adler S, Agarwal S, et al. Gpt-4 technical report. arXiv preprint arXiv:230308774. 2023.
 
9.
Bai Y, Kadavath S, Kundu S, et al. Constitutional ai: harmlessness from ai feedback. arXiv preprint arXiv:221208073. 2022.
 
10.
Zeller NP, Shah AD, Van Heest AE, Bohn DC. Assessing accuracy of chat generative pre-trained transformer’s responses to common patient questions regarding congenital upper limb differences. J Hand Surg Glob Online 2025; 7: 100764.
 
11.
Vaira LA, Lechien JR, Abbate V, et al. Validation of the Quality Analysis of Medical Artificial Intelligence (QAMAI) tool: a new tool to assess the quality of health information provided by AI platforms. Eur Arch Otorhinolaryngol 2024; 281: 6123-31.
 
12.
Wu X, Huang Y, He Q. Diagnostic performance of newly developed large language models for critical illness cases: a comparative study. Int J Med Inform 2025; 204: 106088.
 
13.
Hancı V, Ergün B, Gül Ş, Uzun Ö, Erdemir İ, Hancı FB. Assessment of readability, reliability, and quality of ChatGPT®, BARD®, Gemini®, Copilot®, Perplexity® responses on palliative care. Medicine 2024; 103: e39305.
 
14.
Athaluri SA, Manthena SV, Kesapragada VKM, Yarlagadda V, Dave T, Duddumpudi RTS. Exploring the boundaries of reality: investigating the phenomenon of artificial intelligence hallucination in scientific writing through ChatGPT references. Cureus 2023; 15: e37432.
 
eISSN:1896-9151
ISSN:1734-1922
Journals System - logo
Scroll to top