Commentary Insights Thyroid Disease Management Predictive Risk Models

Commentary: The paradox of AI prediction

August 27, 2026 By Matthew Solan
Share Share via Email Share on Facebook Share on LinkedIn Share on Twitter

A machine learning model developed to predict occult lymph node metastasis in patients with papillary thyroid carcinoma showed moderate overall discrimination, but its sensitivity and negative predictive value may limit its ability to support treatment de-escalation, according to a commentary published in The Journal of Clinical Endocrinology & Metabolism.

David Toro-Tobon, MD
David Toro-Tobon, MD

“The current accuracy of the model reminds us that AI is a support tool, not a replacement for a clinician," wrote AACE Endocrine AI board member David Toro-Tobon, MD, of Mayo Clinic in Rochester, MN. 

Dr. Toro-Tobon examined a recent study by Zhongyu Wang, MD, and colleagues that developed a machine learning (ML) model combining clinical data, ultrasound findings, and genetic markers to predict occult lymph node metastasis in patients with papillary thyroid carcinoma (PTC).  

He described the study as methodologically rigorous, particularly in its integration of different data sources and use of model-interpretability methods, but added that the research illustrates the complex challenge of interpreting performance metrics and artificial intelligence (AI) explanations in clinical practice. 

The model achieved an area under the curve (AUC) of 0.733 in the test set. Dr. Toro-Tobon characterized this as a moderate overall performance but emphasized that AUC measures a model's ability to rank patients from low to high risk across possible thresholds rather than its safety at the specific threshold used for clinical decision-making. 

That distinction is particularly relevant when a prediction might be used to support less invasive treatment, according to Dr. Toro-Tobon. Active surveillance and thermal ablation have emerged as alternatives for select patients with low-risk PTC, but these approaches depend on determining whether cancer has spread to lymph nodes, which can be difficult to assess in the central neck before surgery.  

For a model intended to rule out occult lymph node metastasis, and potentially support nonsurgical treatment, Dr. Toro-Tobon emphasized sensitivity and negative predictive value (NPV). The model had a sensitivity of 63% and an NPV of 74%. By comparison, Dr. Toro-Tobon noted that established rule-out molecular tests in thyroid oncology have NPVs ranging from 94% to 97%. 

Specificity also was 73%, suggesting the model does not yet solve the problem of overtreatment. “This creates a risk of automation bias, where clinicians may over-rely on a computer’s precise-looking probability score, inadvertently overriding their own clinical judgment or ignoring other subtle warning signs,” wrote Dr. Toro-Tobon. 

The underlying study also used SHapley Additive exPlanations (SHAP), a post hoc interpretability method that assigns weight to individual features based on how much they move a model's prediction above or below its average prediction. Dr. Toro-Tobon cautioned that these values represent mathematical attribution of associations rather than causal biological explanations. "That distinction is critical because it can generate an illusion of understanding," he wrote.  

For example, a SHAP result attributing a low-risk prediction to small tumor size could provide a plausible explanation even when the prediction is a false negative. Dr. Toro-Tobon noted that such explanations could suppress clinical skepticism and suggested that future research prioritize models that are interpretable by design, including explainable boosting machines or generalized additive models. 

RET fusion positivity provided another example. In Dr. Wang and colleagues' study, it was associated with 3.34 times the odds of metastasis, but ranked relatively low in the model's global SHAP feature importance. “This discrepancy occurs because ML algorithms prioritize features that help them classify the majority of patients to maximize overall accuracy," wrote Dr. Toro-Tobon.  

Because RET fusions occurred in only 3% of cases, the model learned to deprioritize them in favor of more common features, such as age and tumor size. "This highlights a vital concept for AI literacy: AI feature importance is statistical, not necessarily biological," he wrote. 

Dr. Toro-Tobon suggested that future models incorporating additional genetic data or advanced image-analysis methods may eventually be accurate enough to safely rule out disease. “Until then, clinicians should use these predictions carefully, understanding that a ‘low-risk’ score is an estimate, not a guarantee.”  

Dr. Toro-Tobon reported having no relevant disclosures.

 

AACE Endocrine AI is published by Conexiant under a license arrangement with the American Association of Clinical Endocrinology, Inc. (AACE®). The ideas and opinions expressed in AACE Endocrine AI do not necessarily reflect those of Conexiant or AACE. For more information, see Policies.

Related Content