News Technology Diagnostics & Imaging Predictive Risk Models

Facial imaging shows promise for identifying steatotic liver disease

September 02, 2026 By Matthew Solan 7 min read
Share Share via Email Share on Facebook Share on LinkedIn Share on Twitter

A deep learning system that analyzes 3-dimensional facial images identified steatotic liver disease in internal and independent external test cohorts and retained discrimination across several clinically relevant subgroups, according to a study published in Cell Reports Medicine

"Our findings demonstrate the feasibility of using 3D facial imaging with AI for non-invasive detection of SLD, supported by biological correlates and interpretable features," wrote the authors. 

The deep learning system, called 3D-FAICE, was trained and tested using data from the China Consortium of 3D Facial Image Investigation. Participants aged 18 to 70 underwent 3-dimensional facial scanning, fasting blood collection, and assessment of demographic, clinical, and laboratory information. The study included a development set of 9,440 participants, an internal test set of 834, and an independent external test set of 525. An additional 39 participants comprised a self-controlled longitudinal cohort, and 618 alcohol users were included in a smartphone-based point-of-care cohort. Steatotic liver disease (SLD) was determined using ultrasonography.  

The system processed facial geometry and texture information to generate a composite facial map containing more than 100,000 featurel values. Its predicted probability of SLD was defined as the facial risk score (FRS). 

The system achieved an area under the receiver operating characteristic curve (AUROC) of 0.905 in the internal test set and 0.866 in the external test set. In subgroup analyses, AUROCs ranged from 0.888 to 0.921 for metabolic dysfunction-associated steatotic liver disease and from 0.886 to 0.893 for alcohol-associated liver disease. For early-stage SLD, AUROCs were 0.833 in the internal test set and 0.799 in the external test set. 

The authors also examined whether the model was primarily capturing information related to body size. Among participants with normal body mass index (BMI), AUROCs were 0.831 and 0.837 in the internal and external test sets, respectively. Performance ranged from 0.82 to 0.867 among participants with normal BMI who also did not have hypertension or diabetes. Across age and gender strata, AUROCs exceeded 0.8 in nearly all groups, reaching as high as 0.933 among participants younger than age 50.   

Adding FRS to the base model containing age, gender, BMI, and diabetes increased AUROC by 0.04 in the internal test set and 0.05 in the external test set. The association between FRS and SLD also remained across BMI strata. 

FRS was also compared with established clinical indices for hepatic steatosis. In the internal test set, the FRS AUROC of 0.905 was higher than 0.863 for the Hepatic Steatosis Index, 0.742 for the nonalcoholic fatty liver disease Liver Fat Score, and 0.469 for Fibrosis-4, and was comparable with the Fatty Liver Index at 0.893. 

In the external test set, FRS had an AUROC of 0.866 compared with 0.848 for the Hepatic Steatosis Index, 0.867 for the Fatty Liver Index, 0.719 for the nonalcoholic fatty liver disease Liver Fat Score, and 0.397 for Fibrosis-4. 

In the longitudinal analysis, which included 39 participants who did not have SLD at baseline but developed it during follow-up, FRS had an AUROC of 0.816. 

The authors also prospectively evaluated smartphone-based 3-dimensional facial imaging in 618 alcohol users undergoing annual health examinations, using smartphones equipped with a structured-light module. The model achieved an AUROC of 0.88. Total preprocessing and model inference time using mobile-based edge computing was approximately 10 seconds. 

To explore the biological signals underlying the facial predictions, the authors analyzed plasma metabolomic profiles from 310 participants. Of 100 metabolites associated with the facial risk score, 72 remained significantly associated after adjustment for age, gender, BMI, and diabetes, with pathway analyses continuing to identify amino acid and glycolipid metabolic pathways.  

The authors then evaluated whether combining facial and metabolomic data could improve predictive performance. Using the same 310 participants, they evaluated five experimental settings under five-fold cross-validation: facial imaging alone, metadata alone, metabolomics alone, multimodal fusion of facial and metabolomic data, and cross-modal distillation from metabolomics to facial images. The facial-only model achieved an AUROC of 0.908, compared with 0.936 for metabolomics alone and 0.839 for a model based on age, gender, and BMI. The multimodal fusion model achieved an AUROC of 0.979. 

The authors also used a cross-modal distillation approach to transfer information learned from metabolomics into the facial model during training, allowing the resulting model to operate using facial images alone. This approach increased the facial-only model's AUROC to 0.947. 

Group-averaged SHAP-based interpretability analyses indicated that 3D-FAICE consistently focused on the periorbital and cheek regions. Region-specific SHAP values were also associated with demographic characteristics, including age, gender, and BMI, and showed correlations with amino acid and glycolipid metabolic pathways. 

Because facial images contain potentially identifiable information, the authors also tested federated learning, in which model training can occur without exchanging the original facial data between sites. They first established a baseline using centralized training with standard single-site settings. To simulate multicenter federated training, they randomly divided the internal training data into two subsets, with each representing a local center. Federated learning produced performance similar to centralized training, with AUROCs of 0.900 vs 0.905 in the internal test set and 0.860 vs 0.866 in the external test set. Adding differential privacy reduced performance to AUROCs of 0.875 and 0.835 in the internal and external test sets, respectively, but made the model less susceptible to reconstruction in model-inversion attack experiments. 

The study had several limitations. Ultrasonography used as the reference standard for SLD has limited sensitivity for mild steatosis and is subject to operator dependence and subjective interpretation, particularly for early-stage disease. Early-stage SLD was also defined qualitatively rather than using quantitative thresholds from magnetic resonance imaging proton density fat fraction or controlled attenuation parameter.  

The dataset also remained relatively small for population-level applications. The smartphone cohort consisted exclusively of alcohol users, limiting its generalizability. The authors also noted that chronic alcohol use may produce facial features that could act as confounders and potentially inflate specificity estimates. Real-world image acquisition may be affected by lighting, facial expressions, and occlusions. Images in the study underwent standardized preprocessing and quality control, and participants were instructed to maintain a neutral expression and remove glasses or masks.  

The multimodal and cross-modal distillation analyses presented an additional limitation. They were based on only 310 participants and evaluated using cross-validation without an independent external test cohort. The authors therefore characterized these findings as proof-of-concept rather than clinically validated performance. 

For future research, the authors called for studies using quantitative reference standards and validation in unselected general populations and populations of different ethnic origins. They also recommended external validation of the multimodal and distillation models in larger independent cohorts with paired facial and metabolomic data.  

"With continued validation and technical refinement, such approaches may complement existing diagnostic strategies and expand opportunities for scalable risk assessment," the authors wrote. They reported no competing interests.  

AACE Endocrine AI is published by Conexiant under a license arrangement with the American Association of Clinical Endocrinology, Inc. (AACE®). The ideas and opinions expressed in AACE Endocrine AI do not necessarily reflect those of Conexiant or AACE. For more information, see Policies.

Related Content