OPEN ACCESS International Journal of Advances in Artificial Intelligence and Machine Learning Vol 2 No 3 . https://doi. org/10. 58723/ijaaiml. Research Article Predicting Thyroid Cancer Recurrence Using Machine Learning: An Artificial Intelligence Approach to Clinical Oncology Received: July 01, 2025 Revised: September 20, 2025 Accepted: October 11, 2025 Publish: October 21, 2025 Joy Aifuobhokhan*. Ahmad Khalid Hussain. Chijioke Cyriacus Ekechi. Aisha Olasunbo Olanrewaju. Emmanuel Afuadajo. Deborah Adetola Bowale. Oluwadare Marvellous Inioluwa Abstract: Background of study: Differentiated thyroid cancer (DTC) accounts for most thyroid malignancies and has favorable survival outcomes, yet up to 30% of patients experience recurrence, placing strain on follow-up systems in resourcelimited settings. Conventional staging tools offer limited predictive precision. With increasing interest in machine learning (ML) for precision oncology, there is a need for interpretable, deployable models suitable for low-resource Aims and scope of paper: To develop and validate an interpretable machine learning model for predicting thyroid cancer recurrence and assess its feasibility for deployment in constrained clinical settings, including African oncology Methods: A retrospective dataset of 383 DTC patients with at least 10-year follow-up was sourced from the UCI Machine Learning Repository. Thirteen demographic, clinical, and treatment-related predictors were included. Data preprocessing involved encoding, scaling, and class balancing using SMOTE. Logistic Regression. Random Forest. K-Nearest Neighbors, and Extreme Gradient Boosting (XGBoos. were trained with hyperparameter tuning via grid search and cross-validation. Performance was evaluated using accuracy, precision, recall. F1 score, and AUC-ROC. Result: XGBoost achieved the best performance with 97% accuracy, 95% recall, 94% precision, and an AUC-ROC The most influential predictors were age, smoking status. T and M staging. ATA risk category, and The final model was deployed as a browser-based decision support tool to enable real-time recurrence risk estimation. Conclusion: This study presents a high-performing and interpretable ML model for predicting DTC recurrence, demonstrating feasibility for use in low-resource oncology settings. External validation with African clinical datasets and integration into electronic health systems is recommended to enhance equity and clinical uptake. Keywords: Artificial Intelligence in Healthcare. Machine Learning. Predictive Modelling. Thyroid Cancer Recurrence. XGBoost Classifier. INTRODUCTION prognosis, recurrence remains a significant concern. to 30% of patients experience disease recurrence within 5 to 10 years after treatment, which may manifest locally in cervical lymph nodes or as distant metastases (Gordon et al. , 2. These recurrences necessitate prolonged monitoring, repeat interventions, and have substantial implications for quality of life and healthcare costs (Halder et al. , 2. Thyroid cancer is one of the fastest-rising malignancies globally, particularly affecting women and individuals in mid-life (Kim et al. , 2. Differentiated thyroid cancer (DTC), which includes papillary and follicular subtypes, accounts for over 90% of all cases (Haugen et , 2. While DTC typically carries an excellent Publisher Note: CV Media Inti Teknologi stays neutral with regard to jurisdictional claims in published maps and institutional Accurate risk stratification is essential for guiding clinical decisions, yet traditional tools such as the American Thyroid Association (ATA) risk classification and TNM staging offer static, population-based estimates (Gordon et al. , 2. These models often fail to account for the nuanced interplay of patient-specific factors like thyroid function, prior treatments, and comorbidities (Ahmad & Haddad, 2. As a result, clinicians face challenges in aligning standardized guidelines with individual patient trajectories, especially in heterogeneous populations (Wang et al. , 2. Copyright A20xx by the author. Licensee CV Media Inti Teknologi. Indonesia. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution-ShareAlike (CCBY-SA)license . ttps://creativecommons. org/licenses/bysa/4. 0/). International Journal of Advances in Artificial Intelligence and Machine Learning Recent advances in artificial intelligence (AI), particularly machine learning (ML), present a powerful opportunity to address these limitations. ML algorithms can process complex, high-dimensional clinical data and reveal patterns that conventional statistical models may overlook (Chen & Guestrin, 2. Their growing role in oncology includes applications in imaging analysis, treatment response prediction, and recurrence modelling (Sarker, 2. Studies have shown promising results using ML models such as Random Forest. Support Vector Machines, and Gradient Boosting for recurrence prediction in thyroid and other cancers (Borzooei et al. Recent advances in machine learning (ML) have been applied to predict thyroid cancer recurrence with promising results. For example, one study developed a deep learning model using ultrasound imaging features to stratify recurrence risk (Alawiyah et al. , 2. , while prediction (Habchi et al. , 2. More recently, ensemble methods applied to multi-institutional datasets demonstrated improved accuracy over traditional risk stratification systems (Borzooei et al. , 2. However, most of these models were developed and validated exclusively in high-resource settings, limiting their generalizability to diverse populations and low-resource of multiple ML algorithms in forecasting thyroid cancer . to implement the best-performing model as an accessible, clinician-friendly application. to explore feasibility in African healthcare contexts by simulating data constraints, prioritizing model interpretability, and deploying the final model as a lightweight, web-based application suitable for lowresource environments. This study advances the field by offering a contextsensitive, interpretable ML tool for thyroid cancer It emphasizes not only technical accuracy but also clinical relevance, accessibility, and regional equity, key pillars for responsible AI integration in global health (Park & Lee, 2. MATERIAL AND METHOD This study employed a supervised machine learning approach to predict the recurrence of differentiated thyroid cancer (DTC) using a retrospective clinical The dataset was sourced from the University of California. IrvineAos Machine Learning Repository and comprises anonymized medical records of 383 patients diagnosed with well-differentiated thyroid cancer (Borzooei et al. , 2. Each patient was followed longitudinally for a minimum of ten years to determine recurrence status. The dataset contained no missing values, allowing full utilization of all entries without the need for imputation or data exclusion (Park & Lee. What distinguishes the present study is not only its focus on geographic equity but also its emphasis on feature interpretability and clinical usability. Unlike prior work that often privileges accuracy over transparency, we incorporated established prognostic indicators . ATA risk classification. TNM staging, thyroid functio. alongside modern ML techniques and deployed the resulting model as a web-based application accessible in resource-constrained healthcare systems. This dual focus on interpretability and deployment feasibility aims to bridge the gap between research prototypes and tools that can realistically augment clinical decision-making in diverse settings. The dataset included a diverse range of demographic, clinical, and treatment-related features (Table . Variables captured include patient age, gender, current and past smoking status, previous exposure to radiotherapy, and thyroid function status . uthyroid, hypothyroid, hyperthyroid, and subclinical variant. (Borzooei et al. , 2. (Park & Lee, 2. (Sankar & Sathyalakshmi, 2. Other fields included staging data according to TNM (Tumor. Node. Metastasi. ATA (American Thyroid Associatio. risk classification, presence of adenopathy, tumor focality, histological subtype, and treatment response classification as defined by ATA guidelines (Borzooei et al. , 2. The outcome variable, recurrence, was recorded as a binary indicator . for no recurrence, 1 for To bridge this gap, the current study develops an MLbased recurrence prediction model using a 15-year dataset of 383 patients with well-differentiated thyroid This dataset was sourced from the University of California at Irvine (UCI) Machine Learning Repository (Borzooei et al. , 2. The dataset includes 13 key clinical features ranging from demographics and tumor staging to thyroid function and treatment history (Borzooei et al. , 2. This research addresses three main goals: . to compare the predictive performance Table 1. Clinical and Pathological Features Used in the Study (N = . Feature Name Data Type Brief Description Example Values Age Gender Smoking Hx Smoking Hx Radiotherapy Thyroid Function Continuous PatientAos age at diagnosis . Categorical Biological sex of patient Categorical Current smoking status Categorical History of smoking Categorical History of neck or chest radiotherapy Categorical Thyroid functional status at diagnosis 27, 34, 62 Male. Female Yes. No Yes. No Yes. No Euthyroid. Hyperthyroid. Hypothyroid International Journal of Advances in Artificial Intelligence and Machine Learning Feature Name Physical Examination Adenopathy Pathology Focality Risk T (Tumo. N (Nod. M (Metastasi. Stage Response Data Type Brief Description Example Values Categorical Clinical examination findings Single nodular goiter. Multinodular goiter Categorical Presence of cervical lymphadenopathy Yes. No Categorical Histological subtype of thyroid carcinoma Papillary. Micropapillary. Follicular Categorical Tumor focality Uni-focal. Multi-focal Categorical ATA recurrence risk classification Low. Intermediate. High Tumor size and local invasion (TNM Categorical T1a. T2. T3. T4 Regional lymph node involvement (TNM Categorical N0. N1a. N1b Categorical Distant metastasis (TNM stagin. M0. M1 Categorical Overall AJCC staging I. II, i. IV Excellent. Indeterminate. Biochemical Categorical Response to initial therapy All data processing and modeling tasks were performed in Python . , using libraries such as pandas, scikit-learn. XGBoost, and imbalanced-learn (Chen & Guestrin, 2. (Gomes Mantovani et al. , 2. prepare the data for analysis, categorical variables were encoded using a combination of Label Encoding and One-Hot Encoding (Yousefi et al. , 2. (Gomes Mantovani et al. , 2. Label Encoding was used for ordinal features, such as risk classification, while OneHot Encoding was applied to nominal features like gender and thyroid function to retain interpretability. Continuous variables, including patient age and TNM StandardScaler function, normalizing the values to a mean of zero and unit variance. This ensured uniform feature scaling and supported effective model performance, particularly for distance-sensitive algorithms like K-Nearest Neighbours (KNN) (Lickert et al. , 2. inclusion was further guided by clinical relevance based on established guidelines and literature (Chu et al. Importantly, none of the 13 candidate features were excluded: each demonstrated either significant or near-significant statistical association with recurrence or was retained due to its established prognostic relevance in differentiated thyroid cancer . ATA risk. TNM staging, treatment respons. Thus, the final models were trained on all 13 features age, gender, smoking status, history of radiotherapy. TNM staging components, thyroid function subtype. ATA risk classification, treatment response, pathology subtype, and tumor focality, as summarized in Table 1. To evaluate predictive performance, four machine learning models were implemented: Logistic Regression. Random Forest. K-Nearest Neighbors (KNN), and Extreme Gradient Boosting (XGBoos. (Yang & Shami, 2. These models were chosen for their established effectiveness in classification tasks and their differing algorithmic architectures. Logistic Regression served as a linear baseline, while Random Forest provided a tree-based ensemble model capable of handling non-linearity and feature interactions. KNN, an instance-based learner, was included for its simplicity and sensitivity to feature scaling. XGBoost, a highperformance gradient boosting method, was included for its superior performance in structured data prediction tasks (Probst et al. , 2. The dataset was divided into an 80:20 training-test split using stratified sampling to preserve the class distribution of the recurrence variable across both Given the presence of moderate class imbalance, 108 patients with recurrence versus 275 without, the Synthetic Minority Oversampling Technique (SMOTE) was employed to upsample the minority class in the training set. SMOTE works by interpolating between instances of the minority class to create synthetic data points, helping to reduce model bias and improve recall without compromising realworld distribution in the test set (Lundberg & Lee. Each model underwent hyperparameter optimization using grid search with five-fold cross-validation (Table For Logistic Regression, the regularization strength (C) and solver were tuned. KNN was optimized for the number of neighbors and the weighting scheme. Random ForestAos hyperparameters included the number of estimators, maximum tree depth, and split criterion. For XGBoost, parameters such as learning rate, tree depth, and number of boosting rounds were fine-tuned. This process ensured that each model was trained under Feature selection was conducted through a hybrid approach that combined statistical testing with domain Categorical variables were evaluated using chi-square tests to determine associations with recurrence, while correlation matrices were generated to examine collinearity among variables. Final feature International Journal of Advances in Artificial Intelligence and Machine Learning optimal conditions, minimizing overfitting improving generalizability (R & E, 2. 1 Best parameters: learning rate = 0. n_estimators = 100, max_depth = 5 Ie highest recall and AUC-ROC. K-Nearest Neighbors (KNN): n_neighbors . , 5, 7, 9, . , weights (AouniformAo. AodistanceA. , metric (AoeuclideanAo. AomanhattanA. 1 Best parameters: n_neighbors = 5, weights = distance, metric = euclidean Ie balanced accuracy 0. Logistic Regression: penalty (Aol1Ao. Aol2A. C . 01, 0. 1, . , solver (AoliblinearAo. AosagaA. 1 Best parameters: penalty = l2. C = 1, solver = liblinear Ie mean CV accuracy 0. To optimize performance, hyperparameter tuning was performed using grid search with five-fold crossvalidation. Each model was tuned within a defined parameter space as follows: Random Forest: n_estimators . , 100, . , max_depth (None, 10, 20, . , min_samples_split . , 5, . 1 Best parameters: n_estimators = 100, max_depth = None, min_samples_split = 2 Ie mean CV accuracy = 0. XGBoost: learning rate . 01, 0. 05, 0. n_estimators . , 100, . , max_depth . , 5, . Table 2. Hyperparameter Tuning Ranges and Optimal Parameters for Machine Learning Models Model Random Forest XGBoost KNN Logistic Regression Optimal Parameters (Best CV Resul. Parameter Grid Tested Mean CV Accuracy n_estimators=100, n_estimators: . , 100, . max_depth: [None, max_depth=None, 10, 20, . min_samples_split: . , 5, . min_samples_split=2 learning_rate: . 01, 0. 05, 0. n_estimators: . , learning_rate=0. 100, . max_depth: . , 5, . n_estimators=100, max_depth=5 n_neighbors: . , 5, 7, 9, . weights: [AouniformAo, n_neighbors=5, weights=distance. AodistanceA. metric: [AoeuclideanAo. AomanhattanA. metric=euclidean penalty: [Aol1Ao. Aol2A. C: . 01, 0. 1, 1, . penalty=l2. C=1, solver=liblinear [AoliblinearAo. AosagaA. To assess the modelsAo effectiveness, we calculated standard classification metrics: accuracy, precision, recall . F1 score, and the area under the receiver operating characteristic curve (AUC-ROC). Special emphasis was placed on recall, given the high clinical cost of false negatives in recurrence prediction (Habchi et al. , 2. Confusion matrices and ROC curves were also generated to visualize performance and evaluate calibration. would be implemented in low-resource clinical environments (Tang et al. , 2. To explore feasibility in African healthcare contexts, we incorporated methodological steps beyond standard performance evaluation. Data limitations common in African oncology systems, such as incomplete biomarker availability, heterogeneous follow-up schedules, and pronounced class imbalance, were simulated during preprocessing. To support clinician trust, model interpretability was emphasized by incorporating established prognostic features . ATA risk stratification. TNM staging, thyroid functio. 1 Results This methodology ensures both robustness and transparency, enabling replication of the study and application of similar techniques in other clinical prediction contexts. RESULT AND DISCUSSION This study presents a machine learning-based approach for predicting the recurrence of differentiated thyroid cancer (DTC), applying a comprehensive methodology from data preparation to model deployment. The results discussed herein highlight key clinical trends in the dataset and the modelAos ability to predict recurrence with high sensitivity and specificity. The discussion contextualizes these findings within the broader landscape of thyroid cancer management and artificial intelligence in healthcare, particularly in low-resource Finally, the best-performing XGBoost model was deployed as a lightweight, browser-based application, enabling risk estimation without reliance on highperformance computing infrastructure or continuous broadband connectivity. These steps were designed to approximate the conditions under which such a system The dataset comprised 383 patient records spanning a 15-year observation period, with each patient monitored for at least a decade following initial treatment. The average age at diagnosis was 41. 25 years (SD = 15. ranging from 15 to 82 years. A majority of the cohort International Journal of Advances in Artificial Intelligence and Machine Learning . %) fell between the ages of 30 and 52, suggesting a middle-aged profile typical of DTC epidemiology. Gender distribution was heavily skewed, with 82. female and 17. 5% male, mirroring global patterns in thyroid cancer prevalence. This imbalance reflects the clinical reality of DTC and underscores the need for modeling strategies that prioritize sensitivity to rare outcomes. To address this imbalance, the Synthetic Minority Oversampling Technique (SMOTE) was used to augment the recurrence class in the training set, improving the modelAos capacity to learn from limited positive Exploratory data analysis (EDA) revealed a strong class imbalance in the outcome variable: 17. 5% of patients experienced recurrence, while 82. 5% did not (Figure . Figure 1. Distribution of Recurrence Status Further descriptive analysis revealed clinical features with significant predictive relevance (Figure 2-. Only 11% of the patients were current smokers, while the remaining 89% were non-smokers. Approximately 86% of patients were euthyroid at the time of diagnosis, and subclinical thyroid disorders were observed in a small Regarding risk stratification, 87% of the cohort was categorized as low-risk according to ATA criteria, while intermediate and high-risk classifications accounted for the remainder. International Journal of Advances in Artificial Intelligence and Machine Learning . Figure 2. Categorical Feature Distributions. Distribution of Gender. Distribution of Smoking. Distribution of Thyroid FunctionAeEuthyroid. Distribution of Thyroid Function-Clinical Hypothyroidism. Distribution of Thyroid Function-Clinical Hyperthyroidism. Distribution of Thyroid Function-Subclinical Hyperthyroidism. Distribution of Thyroid Function-Subclinical Hypothyroidism. Distribution of Thyroid Function-Risk AoLowAo. Distribution of Thyroid Function-Risk AoIntermediateAo. Distribution of Thyroid Function-Risk AoHighAo. Pearson correlation analysis identified strong positive relationships between age and smoking history, tumor size (T), and metastatic spread (M). Recurrent status showed the highest correlation with ATA risk level and tumor progression metrics. Figure 3. Pearson Correlation Heatmap This heatmap illustrates pairwise Pearson correlation coefficients among all demographic, clinical, and treatment-related variables in the differentiated thyroid cancer dataset. Variables are listed along both axes, with coefficients ranging from -1 . erfect negative correlation, dark blu. to 1 . erfect positive correlation, dark re. , and the diagonal representing self-correlation. Annotated values highlight the strength and direction of relationships, supporting the identification of multicollinearity and feature Notably, tumor size (T stag. correlates strongly with ATA risk classification, metastatic spread (M stag. , and adenopathy, while age shows moderate positive correlation with smoking history and tumor stage. Thyroid function status demonstrates generally weak correlations with staging These patterns provide clinical context and inform feature selection for machine learning modeling. Network-based correlation visualizations of categorical variables highlighted interconnected nodes, such as age. T staging, adenopathy, and ATA risk level, as central to recurrence prediction. These nodes were critical in shaping model input features. International Journal of Advances in Artificial Intelligence and Machine Learning Figure 4. Categorical Correlation Network Chi-square analysis confirmed statistically significant associations . O 0. between recurrence and several categorical variables. Key associations included smoking history . = 2. ATA risk classification . = 4. N, and M stages . < 1e-. , as well as adenopathy and treatment response . < 1e-. These findings validated the clinical relevance of the chosen Boxplot visualizations stratified by recurrence status demonstrated that patients who experienced recurrence were, on average, older and had a wider age distribution (Figure . In contrast, non-recurrent patients were more tightly clustered around a median age of 38. Figure 5. Boxplot - Age vs. Recurrence Violin plots added nuance by illustrating distribution Age Distribution by Gender and Recurrence Figure . : Males with recurrence skewed older with tight age ranges. females showed wider age Age Distribution by Smoking and Recurrence Status Figure 6. : Recurrent smokers tended to be older. non-recurrent smokers were younger. Age Distribution by Thyroid Function and Recurrence Status Figure 6. : Thyroid dysfunction in recurrent cases skewed toward older Smoking vs. Recurrence Figure 6. : Smokers displayed a near-balanced distribution across recurrence outcomes, indicating smoking may be a risk amplifier. Thyroid Function vs. Recurrence Figure 6. Euthyroid patients were predominantly nonrecurrent, while recurrence cases exhibited a broader spread among those with thyroid dysfunction. International Journal of Advances in Artificial Intelligence and Machine Learning . Figure 6. Smoking vs Recurrence. Thyroid Function vs Recurrence. Age Distribution by Gender and Recurrence Status. Age Distribution by Smoking and Recurrence Status. Age Distribution by Thyroid Function and Recurrence Status The superior performance of XGBoost compared to other models in this study can be attributed to several First. XGBoost is well-suited to capturing complex, non-linear interactions among clinical features such as TNM staging. ATA risk, and treatment response, which may not be adequately modelled by linear approaches like logistic regression. Second, unlike Random Forest, which averages many deep trees and can be prone to overfitting on smaller datasets. XGBoost incorporates shrinkage . earning rat. and both L1/L2 regularization, which enhance its generalizability. Performance evaluation of the four models (Table . Logistic Regression. Random Forest. KNN, and XGBoost, was conducted on SMOTE-resampled data. XGBoost led in all performance metrics: 97% accuracy, 95% recall, 94% precision, 94% F1-score, and an AUCROC of 0. Logistic Regression and Random Forest also performed well, but fell short on recall and overall KNN lagged behind, particularly in recall . , which is problematic in a clinical setting. Figure 7 Ae 8. 5 show the AUC-ROC and Confusion Matrix of all models. International Journal of Advances in Artificial Intelligence and Machine Learning Third. XGBoost handles class imbalance effectively through scale-pos-weight adjustments and iterative reweighting, complementing our use of SMOTE. Finally, its boosting framework builds trees sequentially, focusing on difficult-to-predict cases such as rare recurrences, thereby improving recall, a clinically critical metric in recurrence prediction. Together, these properties explain why XGBoost demonstrated stronger recall and near-perfect AUCROC in this cohort, outperforming simpler linear models and bagging-based ensemble methods. Table 3. Performance Metrics Comparison Model XGBoost Logistic Regression Random Forest KNN Accuracy Precision Recall F1 Score AUC-ROC Figure 7. ROC Curve XGBoost. ROC Curve for Logistic Regression. ROC Curve for Random Forest. ROC Curv\me for KNN International Journal of Advances in Artificial Intelligence and Machine Learning . Figure 8. Confusion Matrix for XGBoost. Confusion Matrix for KNN. Confusion Matrix for Logistic Regression. Confusion Matrix for Random Forest XGBoost was ultimately selected for deployment due to its superior recall and AUC-ROC, robustness to overfitting through internal regularization, and its interpretability via feature importance ranking. Top contributing features included age, smoking, tumor staging (T and M). ATA risk level, and adenopathy, findings consistent with both the EDA and clinical These findings contribute a robust and interpretable decision-support tool for oncologists managing DTC The tool enables early risk stratification, optimized surveillance schedules, and targeted therapeutic planning. This is particularly valuable in low-resource settings where medical infrastructure and clinical personnel are limited. Importantly, the model augments clinical judgment rather than replacing it, fostering collaborative, data-driven care. Deployment of the final model was achieved via a lightweight, web-based Streamlit application. The webbased interface was successfully tested across standard desktop and mobile browsers. This application enables healthcare providers to input patient-specific variables and receive real-time recurrence risk estimates. Its intuitive design supports integration into outpatient workflows, and future integration into electronic health record (EHR) systems is planned. However, challenges persist. The training data, while comprehensive, does not originate from African populations (Borzooei et al. , 2. This raises generalizability concerns, especially given regional differences in genetics, disease presentation, and healthcare access. Data quality and infrastructure issues, such as fragmented records, paper-based systems, and a lack of standardized EMRs, further complicate the development and deployment of AI tools in African Predictions were generated in real time without requiring specialized hardware, and the interfacemaintained low-bandwidth These findings suggest that the model is not only accurate but also practical for integration into oncology workflows in resource-constrained settings, where traditional diagnostic infrastructure may be To move forward, a multipronged approach is needed. Health ministries must invest in digitized and standardized EMRs with oncology modules. Local institutions should curate and maintain high-quality datasets for training and validation. Ethical governance must guide the development and deployment of AI tools to prevent algorithmic bias. Models built in Western International Journal of Advances in Artificial Intelligence and Machine Learning settings should be re-trained using African data prior to implementation (Tang et al. , 2. Despite its strengths, the study has several limitations. First, the dataset used for model training originates from a Western population and may not generalize well to African or other underrepresented groups. In African populations, genetic polymorphisms such as differences in BRAF and RAS mutation prevalence (Mao et al. , environmental exposures such as variable iodine intake and aflatoxin exposure (Putatunda & Rama, 2. , and dietary patterns including cassava consumption and micronutrient variability (Adedinsewo et al. , 2. may influence both tumor biology and recurrence risks. Healthcare access factors such as delayed diagnosis, limited follow-up imaging, and heterogeneous treatment adherence could also modify predictive outcomes (Mahamadou et al. , 2. Without incorporating these variables, model generalizability to African cohorts remains limited. In conclusion, this study not only demonstrates the feasibility and utility of ML in recurrence prediction but also highlights the broader systemic requirements for equitable AI in oncology. XGBoost, when trained and deployed responsibly, can provide a powerful tool to enhance precision medicine in thyroid cancer care, especially in settings where personalized, scalable tools are most needed. 2 Discussion