Leading LLMs vary on thyroid cancer guidance
Large language models (LLMs) demonstrated widely varying performance in generating recommendations that aligned with established clinical guidelines for anaplastic thyroid cancer (ATC), according to a comparative study published in Scientific Reports. Although top-performing models produced guideline-concordant responses more consistently than others, the authors concluded that none were sufficiently reliable for independent clinical use.
The investigators evaluated five leading LLMs using 70 standardized clinical questions covering the diagnosis and management of ATC. Responses were assessed for adherence to recommendations from the National Comprehensive Cancer Network, the American Thyroid Association, and the European Society for Medical Oncology.
The study followed TRIPOD-LLM reporting guidance. Three surgical oncology experts reviewed model-generated responses for relevance, clarity, accuracy, and adequacy. The first four models were evaluated under blinded conditions, while ChatGPT 5 was assessed separately after its public release, wrote lead author Mohamed Yasser, of Mansoura University Hospitals and the Faculty of Medicine at Mansoura University in Mansoura, Egypt, and colleagues.
Gemini 2.5 Pro achieved the highest median accuracy, followed by DeepSeek R1, while ChatGPT 4.1 scored lowest across most metrics. ChatGPT 5 and Claude Sonnet 4 had intermediate performance.
Performance varied by clinical domain and question complexity. Accuracy differences were statistically significant only for surgical management questions, where Gemini 2.5 Pro had the highest median score. The clearest separation by complexity occurred for moderate-complexity questions, suggesting that top-performing models may be better at synthesizing multiple guideline elements into a coherent clinical response.
The findings remained consistent in sensitivity analyses, according to the study authors. Excluding the unblinded ChatGPT 5 model "confirmed the performance hierarchy among blinded models, with significance preserved or strengthened across all four metrics," they wrote.
Despite the performance differences, the investigators emphasized that no model consistently produced guideline-concordant recommendations across all scenarios.
"Leading LLMs show variable capacity to align with ATC clinical guidelines," the authors wrote. "While top-performing models hold promise as supportive tools, their inconsistencies across domains and complexity levels preclude autonomous clinical use."
These models should serve strictly as decision aids under expert supervision, they noted. LLMs may help clinicians rapidly access guideline-based information, but important limitations remain before the technology can be incorporated into independent clinical decision-making for patients with ATC.
Future work should focus on improving consistency, reducing variability across clinical domains, and evaluating model performance as clinical guidelines evolve, the investigators concluded. They also noted that the study relied on standardized clinical questions rather than real-world patient encounters and focused on a single cancer type, which may limit generalizability to other oncology settings.
Open access funding was provided by the Science Technology & Innovation Funding Authority in cooperation with the Egyptian Knowledge Bank. The authors declared no competing interests.
AACE Endocrine AI is published by Conexiant under a license arrangement with the American Association of Clinical Endocrinology, Inc. (AACE®). The ideas and opinions expressed in AACE Endocrine AI do not necessarily reflect those of Conexiant or AACE. For more information, see Policies.