Available online at https://journal. com/index. php/ijqrm/index International Journal of Quantitative Research and Modeling e-ISSN 2721-477X p-ISSN 2722-5046 Vol. No. 2, pp. 259-274, 2026 Applying Machine Learning Algorithms to Predict Employee Turnover Intention: A Comparative Model Analysis Siti Hadiaty1*. Fahmi Sidiq2. Yasir Salih3 Research Collaboration Community. Bandung. Indonesia Master of sains program. Faculty of Pharmacy, university Sultan Zainal Abidin. Kampung Gong Badak, 21300. Terengganu Department of Mathematics Education. Faculty of Education. Red Sea University. SUDAN *Corresponding author email: sitihadiatyy@gmail. Abstract Employee turnover represents a major challenge for organizations because it increases recruitment and training costs, disrupts operational continuity, and reduces organizational performance. Although machine learning has been widely applied to employee attrition prediction, most studies focus on comparing algorithms using the complete feature set, with limited attention to the predictive contribution of different employee information domains. This study aims to identify the most informative attribute domains for turnover prediction, compare the performance of ANN. RF, and SVM, and evaluate whether reduced-domain models can achieve performance comparable to full-feature models. The study utilized the IBM HR Analytics Employee Attrition dataset containing 1,470 employee records. Thirty predictive attributes were organized into six conceptual domains: Personal Information. Job Characteristics. Compensation. Work Environment. Career Development, and Relationship & Supervision. Twelve domain-based model configurations were developed and evaluated using ANN. RF, and SVM. Model development employed SMOTE to address class imbalance and repeated 10-fold cross-validation, while final evaluation was conducted on an independent holdout validation dataset. The results show that multi-domain models consistently outperform single-domain Compensation and Career Development emerged as the strongest standalone domains, while Work Environment was present in all top-performing models. The highest validation accuracy was achieved by M0-SVM . 01%), whereas M11SVM achieved comparable performance . 65%) using only 16 attributes. M11-ANN produced the highest ROC AUC . indicating superior discriminative capability. Feature importance analysis identified OverTime. MonthlyIncome. Age. TotalWorkingYears, and YearsAtCompany as the most influential predictors. These findings demonstrate that domain composition is as important as algorithm selection in employee turnover prediction and highlight the importance of work environment, compensation, and career development factors in supporting data-driven employee retention strategies. Keywords: Employee turnover prediction, machine learning, human resource analytics, feature domain analysis, employee Introduction Employee turnover intention, defined as an employeeAos conscious willingness to leave an organisation in the foreseeable future, remains one of the most persistent challenges in human resource management. High turnover rates impose substantial costs on organisations through recruitment, onboarding, training, and productivity losses, while also reducing organisational knowledge retention and workforce stability (Dhoopar et al. , 2. As labour markets become increasingly dynamic and competitive, organisations are under growing pressure to identify employees at risk of leaving before actual resignation occurs, enabling timely interventions and more effective retention strategies. Research on turnover intention has consistently demonstrated that employee departure decisions are influenced by a complex combination of personal, organisational, and job-related factors. Variables such as job satisfaction, workAelife balance, compensation, career advancement opportunities, supervisory relationships, and workload have repeatedly been identified as significant antecedents of turnover intention (Junaidi et al. , 2020. Saufi et al. , 2023. Jogi et al. These findings align with the Theory of Planned Behavior (Ajzen, 1. , which posits that behavioural intentions arise from attitudes, subjective norms, and perceived behavioural control. Consequently, turnover intention can be viewed as a measurable and predictable behavioural outcome shaped by multiple interacting determinants. The growing availability of employee data has accelerated the adoption of machine learning (ML) techniques in human resource analytics. Unlike traditional statistical approaches. ML algorithms can model complex nonlinear Hadiaty et al. / International Journal of Mathematics. Statistics, and Computing. Vol. No. 2, pp. 259-274, 2026 relationships, capture interactions among variables, and process large numbers of employee attributes simultaneously without strict distributional assumptions (Talebi et al. , 2. These capabilities make ML particularly suitable for predicting employee attrition, where behavioural decisions often emerge from multifaceted and interdependent factors. Recent studies have applied a wide range of machine learning algorithms to employee attrition prediction, including Random Forest (RF). Support Vector Machine (SVM). Artificial Neural Networks (ANN), ensemble learning approaches, and explainable artificial intelligence frameworks (Fallucchi et al. , 2020. Chung et al. , 2023. Park et al. Al-Ali et al. , 2. Although many studies report promising predictive performance, their primary focus has generally been algorithm comparison and accuracy optimisation. Consequently, relatively little attention has been devoted to understanding which groups of employee attributes contribute most strongly to predictive performance. This limitation is important because organisations often face practical constraints when collecting employee Building predictive systems using all available attributes may increase implementation costs, reduce interpretability, and complicate decision-making. From a managerial perspective, identifying a smaller subset of highly informative attribute domains may be more valuable than marginal improvements in predictive accuracy. Therefore, understanding the relative predictive contribution of different employee-related domains represents an important yet underexplored research question. Several studies have suggested that turnover intention is influenced by distinct dimensions of employee experience, including personal characteristics, job characteristics, compensation, work environment, career development, and supervisory relationships (Mobley, 1977. Jogi et al. , 2. However, empirical evidence remains limited regarding how these domains perform individually and in combination when used as inputs for machine learning models. Existing studies typically utilise the complete feature set and evaluate model performance as a whole, making it difficult to determine which domains provide the strongest predictive signal and whether certain domains can be omitted without substantial performance degradation. To address this gap, this study adopts a domain-based modelling strategy inspired by Hernandez and Sandoval . The six attribute domains were established based on the conceptual framework proposed in that study and further adapted to the characteristics of the IBM HR Analytics dataset. Using the IBM HR Analytics Employee Attrition dataset (IBM, 2. , the available employee attributes are organised into six conceptual domains: Personal Information. Job Characteristics. Compensation. Work Environment. Career Development, and Relationship & Supervision. These domains provide a structured framework for investigating the behavioural and organisational factors underlying employee turnover intention. Based on these domains, twelve classification models are constructed, ranging from single-domain configurations to multi-domain combinations. Three machine learning algorithms. ANN. RF, and SVM, are employed to evaluate each model configuration. This design enables a systematic assessment of how predictive performance changes as different domains are included or combined. Accordingly, this study pursues three objectives. First, it seeks to identify which employee attribute domains contain the strongest predictive information for turnover intention. Second, it evaluates the comparative performance of ANN. RF, and SVM in predicting employee attrition. Third, it investigates whether a reduced set of strategically selected domains can achieve predictive performance comparable to models that utilise the complete attribute set. addressing these objectives, the study contributes to both the machine learning and HR analytics literature by providing a clearer understanding of the relationship between domain composition, algorithm selection, and predictive The findings are expected to support the development of more efficient and interpretable employee attrition prediction systems while offering practical guidance for organisations seeking data-driven approaches to workforce retention. In addition, the study provides evidence regarding which aspects of employee experience should be prioritised for monitoring and intervention, thereby supporting more targeted and proactive human resource management strategies. Literature Review Employee Turnover Intention Employee turnover intention refers to an employeeAos conscious and deliberate willingness to leave an organisation within a foreseeable period (Mobley, 1. It is widely regarded as the strongest psychological precursor of actual turnover behaviour and therefore represents a critical target for organisational intervention. High turnover intention is associated with substantial organisational costs, including recruitment expenditure, onboarding activities, productivity disruption, and loss of organisational knowledge (Dhoopar et al. , 2. Consequently, understanding and predicting turnover intention has become a central objective in contemporary human resource management. The theoretical foundation most commonly used to explain turnover intention is the Theory of Planned Behavior (TPB) proposed by Ajzen . According to TPB, behavioural intention is shaped by attitudes toward the behaviour, subjective norms, and perceived behavioural control. In organisational settings, employees who develop negative attitudes toward their jobs, perceive favourable external employment opportunities, and believe they have the ability to change employment are more likely to exhibit turnover intention. This framework provides a conceptual basis for modelling turnover intention as a measurable and predictable outcome influenced by multiple employeerelated factors. Hadiaty et al. / International Journal of Mathematics. Statistics, and Computing. Vol. No. 2, pp. 259-274, 2026 Empirical studies consistently identify several key antecedents of turnover intention. Job satisfaction remains the most robust negative predictor, indicating that employees who are satisfied with their work are less likely to consider leaving (Jogi et al. , 2. WorkAelife balance, compensation, career advancement opportunities, supervisory relationships, and organisational support have also been shown to significantly influence turnover intention (Saufi et , 2. Furthermore, excessive workload and overtime work increase psychological strain and reduce employee well-being, thereby increasing the likelihood of attrition (Junaidi et al. , 2. These findings suggest that turnover intention is a multidimensional phenomenon arising from the interaction of personal, occupational, and organisational factors. To reflect this multidimensional nature, employee attributes can be grouped into several conceptual domains, including personal information, job characteristics, compensation, work environment, career development, and supervisory relationships. Examining these domains individually and collectively provides an opportunity to understand which aspects of employee experience contribute most strongly to turnover prediction. Machine Learning for Employee Turnover Prediction The increasing availability of workforce data has accelerated the adoption of ML techniques in HR analytics. Machine learning enables computers to learn predictive patterns directly from historical data and improve performance without explicit rule-based programming (Talebi et al. , 2. Among the various ML paradigms, supervised learning is the most widely used for employee turnover prediction because historical employee records contain known attrition outcomes that can be used as training labels (Fallucchi et al. , 2. Compared with conventional statistical approaches, machine learning algorithms offer several advantages in turnover prediction. They are capable of capturing nonlinear relationships, modelling complex interactions among employee attributes, and handling high-dimensional datasets without requiring strong distributional assumptions (AlAli et al. , 2. These characteristics are particularly valuable in HR analytics, where employee behaviour is often influenced by multiple interconnected factors rather than isolated variables. A wide range of machine learning algorithms has been applied to employee attrition prediction. Fallucchi et al. demonstrated that tree-based and neural-network approaches outperform traditional statistical models in identifying employees at risk of leaving. Chung et al. further improved predictive performance through stacking ensemble learning, while Park et al. incorporated additional behavioural and organisational variables to enhance classification accuracy. More recently. Al-Ali et al. integrated machine learning with explainable artificial intelligence (XAI), highlighting the importance of model interpretability in HR decision-making. Despite these advances, existing studies primarily focus on improving predictive accuracy through algorithm selection, hyperparameter optimisation, or ensemble strategies. Relatively limited attention has been given to understanding the predictive contribution of different employee attribute domains. Most studies utilise the entire feature set and evaluate model performance as a whole, making it difficult to determine which categories of employee information are most influential for turnover prediction. This limitation motivates the domain-based modelling approach adopted in the present study. Classification Algorithms Classification is a supervised learning task that assigns observations to predefined categories based on patterns learned from labelled data. In employee turnover prediction, the objective is to classify employees into two categories: Stay and Leave. Model performance is commonly evaluated using Accuracy. Precision. Recall, and ROC AUC, which together provide a comprehensive assessment of predictive capability (Davis & Goadrich, 2. Among classification algorithms. ANN. RF, and SVM have consistently demonstrated strong performance in employee attrition prediction. ANN models are capable of learning highly complex nonlinear relationships through interconnected hidden layers and have shown competitive predictive performance on HR datasets (Fallucchi et al. RF, introduced by Breiman . , combines multiple decision trees through ensemble learning and is valued for its robustness, resistance to overfitting, and ability to estimate feature importance. SVM, based on VapnikAos . statistical learning theory, construct optimal decision boundaries that maximise class separation and are particularly effective in high-dimensional classification problems. Although these algorithms have frequently been compared in the literature, findings remain inconsistent regarding which approach performs best under different feature configurations. This suggests that predictive performance may depend not only on algorithm selection but also on the composition of employee attribute domains used as model inputs. Class Imbalance. SMOTE, and Model Validation Employee turnover datasets commonly exhibit substantial class imbalance because the number of employees who leave is typically much smaller than the number who remain. Such imbalance can bias classification models toward the majority class, resulting in deceptively high accuracy but poor identification of attrition cases (Saito & Rehmsmeier, 2. Since the minority class represents the primary target of HR intervention, improving minorityclass detection is critical for practical deployment. The Synthetic Minority Over-sampling Technique (SMOTE) Hadiaty et al. / International Journal of Mathematics. Statistics, and Computing. Vol. No. 2, pp. 259-274, 2026 addresses this challenge by generating synthetic minority-class observations through interpolation between existing minority samples (Chawla et al. , 2. Previous studies have shown that SMOTE can substantially improve recall and minority-class discrimination in employee attrition prediction tasks (Xu et al. , 2023. Park et al. , 2. To avoid information leakage. SMOTE should be applied exclusively to the training dataset prior to model construction. Another important challenge in predictive modelling is overfitting, which occurs when a model learns patterns specific to the training data but fails to generalise to unseen observations. Cross-validation provides an effective mechanism for mitigating this problem by repeatedly partitioning the training data into multiple subsets and averaging performance across iterations (Powers & Atyabi, 2. Consequently, combining SMOTE with rigorous crossvalidation offers a robust framework for developing reliable employee turnover prediction models. Research Gap Although previous studies have successfully demonstrated the potential of ML for employee turnover prediction, three important gaps remain. First, most studies concentrate on algorithm comparison while treating employee attributes as a single undifferentiated feature set. Second, limited evidence exists regarding the relative predictive value of different employee attribute domains such as compensation, work environment, career development, and supervisory relationships. Third, it remains unclear whether reduced domain combinations can achieve predictive performance comparable to models that utilise the complete attribute set. To address these gaps, this study organises employee attributes into six conceptual domains and systematically evaluates twelve domain-based model configurations using ANN. RF, and SVM. This approach enables a deeper understanding of how domain composition influences predictive performance and identifies the most informative and practically deployable attribute combinations for employee turnover prediction. Materials and Methods Dataset Characteristics This study utilised the IBM HR Analytics Employee Attrition and Performance dataset comprising of 1,470 employee records and 35 original variables. Following preprocessing, four constant-value variables (EmployeeCount. EmployeeNumber. Over18, and StandardHour. were removed because they contained no predictive information, resulting in 31 usable variables consisting of 30 predictive attributes and one binary target variable (Attritio. The target variable was encoded as Stay . and Leave . , transforming the problem into a supervised binary classification task. A notable characteristic of the dataset is the substantial class imbalance between employees who remained in the organisation and those who left. Specifically, 1,233 employees . 9%) belonged to the Stay class, whereas only 237 employees . 1%) belonged to the Leave class, yielding an imbalance ratio of approximately 2:1. This distribution is consistent with real-world organisational settings where employee attrition typically represents a minority event. Consequently, the Synthetic Minority Over-sampling Technique (SMOTE) was applied exclusively to the training subset during model development to mitigate classification bias while preventing information leakage into the validation dataset. To facilitate systematic investigation of employee turnover determinants, the 30 predictive attributes were organised into six conceptual domains following the structured domain-based framework proposed by Hernandez and Sandoval . These domains represent complementary dimensions of employee characteristics, workplace conditions, and organisational experiences. As shown in Table 1, the domains include Personal Information (D. Job Characteristics (D. Compensation (D. Work Environment (D. Career Development (D. , and Relationship & Supervision (D. The Compensation and Career Development domains contain the largest number of attributes . ix eac. , whereas the Relationship & Supervision domain contains three attributes. The distributional characteristics of the dataset are presented in Figure 1. Figure 1A confirms the substantial imbalance between the Stay and Leave Table 1: Domain groupings of the IBM HR Analytics Employee Attrition dataset Domain Domain Name Personal Information Job Characteristics Compensation No. of Attributes Work Environment Career Development Relationship & Supervision Total Key Attributes Age. Gender. MaritalStatus. Education. EducationField JobRole. Department. JobLevel. JobInvolvement. BusinessTravel MonthlyIncome. HourlyRate. DailyRate. MonthlyRate. StockOptionLevel. PercentSalaryHike EnvironmentSatisfaction. WorkLifeBalance. OverTime. JobSatisfaction. DistanceFromHome YearsAtCompany. YearsSinceLastPromotion. TrainingTimesLastYear. NumCompaniesWorked. TotalWorkingYears. YearsInCurrentRole RelationshipSatisfaction. YearsWithCurrManager. PerformanceRating Hadiaty et al. / International Journal of Mathematics. Statistics, and Computing. Vol. No. 2, pp. 259-274, 2026 Figure 1B illustrates the age distribution by attrition status, showing that employees who left the organisation are predominantly concentrated within the 25Ae35 year age range. This finding suggests that early- and mid-career employees may exhibit greater workforce mobility and a higher propensity to seek alternative employment Figure 1C compares monthly income distributions between the two attrition groups. Employees who left the organisation exhibit a substantially lower median monthly income than those who remained (USD 4,787 versus USD 6,. , indicating that compensation may play an important role in turnover behaviour. This observation is consistent with previous studies identifying remuneration adequacy as a significant predictor of employee retention (Jogi et al. , 2. The influence of work environment factors is evident in Figure 1D and Figure 1E. Employees working overtime exhibit an attrition rate of approximately 30. 5%, nearly three times higher than the rate observed among employees who do not work overtime . 4%). Similarly, employees reporting lower job satisfaction levels . evels 1Ae. are disproportionately represented among attrition cases. These findings support previous evidence suggesting that excessive workload, reduced workAelife balance, and job dissatisfaction contribute significantly to turnover intention (Junaidi et al. , 2020. Saufi et al. , 2. Finally. Figure 1F presents attrition rates across tenure categories. The highest attrition rate occurs among employees with less than two years of organisational tenure, reaching approximately 34%. Attrition decreases steadily as tenure increases, suggesting that employee retention improves once individuals become more embedded within the organisation. This pattern is consistent with turnover theory, which proposes that employees are most likely to leave during the early stages of organisational membership when commitment and organisational attachment remain relatively weak (Mobley, 1. Figure 1: Dataset distributional profile. (A) Target class distribution . =1,. (B) Age distribution by attrition (C) Monthly income by attrition group. (D) Overtime frequency vs. (E) Job satisfaction levels by (F) Attrition rate by tenure bracket Experimental Design The research adopted a comparative experimental design consisting of four stages: data preprocessing, domainbased model design, model construction, and independent validation. As illustrated in Figure 2, the proposed workflow begins with dataset preparation and domain definition, followed by model configuration design, machine learning model development, and final performance evaluation on an independent validation dataset. This structured framework ensures a systematic assessment of both algorithm performance and the predictive contribution of individual employee attribute domains. Hadiaty et al. / International Journal of Mathematics. Statistics, and Computing. Vol. No. 2, pp. 259-274, 2026 Figure 2: Four-stage research workflow for employee turnover prediction Initially, the dataset was partitioned using stratified random sampling into a model-development subset . and an independent holdout validation subset . n = . Stratification ensured that the original class distribution was preserved in both subsets. As shown in Figure 2, the validation subset was isolated throughout the model construction phase and used exclusively during the final evaluation stage to provide an unbiased estimate of model generalisation performance. To examine the predictive value of employee attribute domains, twelve model configurations were developed. Based on the domain framework presented in Table 1. Model M0 utilised all six domains . , representing the complete feature set. Models M1AeM6 were constructed using individual domains only, enabling evaluation of the standalone predictive power of each domain. Models M7AeM11 combined two or three complementary domains to investigate whether integrating information from multiple employee dimensions could improve classification performance beyond that achieved by single-domain models. As summarised in Table 2, the twelve model configurations range from highly parsimonious models containing only three to six attributes (M1AeM. to more comprehensive multi-domain models containing up to sixteen attributes (M. This experimental structure enables direct comparison of domain-specific and cross-domain predictive capabilities while simultaneously evaluating the trade-off between model complexity and predictive performance. Table 2: Domain combinations used for the twelve classification models Model M10 M11 ue ue Ae Ae Ae Ae Ae Ae Ae Ae ue Ae ue Ae ue Ae Ae Ae Ae Ae Ae Ae Ae ue ue Ae Ae ue Ae Ae Ae Ae ue Ae Ae Ae ue Ae Ae Ae ue Ae Ae ue ue ue ue ue ue Ae Ae Ae Ae ue Ae ue Ae ue Ae ue ue Ae Ae Ae Ae Ae ue Ae Ae ue Ae Ae No. of Attributes Following the design presented in Table 2, each model configuration was evaluated using three machine learning algorithms. ANN. RF, and SVM, resulting in a total of 36 modelAealgorithm combinations. This factorial design Hadiaty et al. / International Journal of Mathematics. Statistics, and Computing. Vol. No. 2, pp. 259-274, 2026 allows simultaneous investigation of domain composition effects and algorithmic differences in employee turnover Model Construction To address the class imbalance problem, the SMOTE was applied exclusively to the training subset after data Applying SMOTE only to the training data prevents information leakage and ensures unbiased evaluation on the holdout validation dataset (Chawla et al. , 2. Following oversampling, the training dataset increased from 1,176 observations to approximately 1,972 balanced instances. Subsequently, all predictor variables were standardised using z-score normalisation. Scaling parameters were estimated exclusively from the training data and then applied to the validation dataset to maintain consistency. Model training employed repeated stratified 10-fold cross-validation. For each model configuration, the SMOTEbalanced training dataset was divided into ten folds, with nine folds used for training and one fold used for testing. This procedure was repeated ten times using different random seeds to minimise partitioning bias and provide robust estimates of model stability (Powers & Atyabi, 2. The mean and standard deviation of classification accuracy across replications were recorded for each modelAealgorithm combination. Three machine learning algorithms were investigated. The ANN classifier was implemented as a Multilayer Perceptron comprising two hidden layers with 100 and 50 neurons, respectively, using ReLU activation and early The RF classifier consisted of 100 decision trees employing Gini impurity as the splitting criterion and balanced class weights. The SVM classifier utilised a Radial Basis Function (RBF) kernel with a regularisation parameter of C = 1 and balanced class weights. Hyperparameter values were determined through grid-search optimization using repeated cross-validation, with the search ranges guided by recommendations from previous employee attrition studies (Talebi et al. , 2. The final parameter values were selected based on the highest validation performance. Validation and Model Interpretation Following cross-validation, each classifier was retrained using the complete SMOTE-balanced training dataset and subsequently evaluated on the independent holdout validation dataset . = . This stage assessed model generalisation performance on previously unseen employee records and provided an unbiased estimate of real-world predictive capability. Performance evaluation employed four complementary metrics: Accuracy. Precision. Recall, and Area Under the Receiver Operating Characteristic Curve (ROC AUC). Accuracy measures overall classification correctness. Precision evaluates the reliability of attrition predictions. Recall quantifies the proportion of actual attrition cases correctly identified, and ROC AUC measures discriminative capability across all classification Because employee attrition prediction represents an imbalanced classification problem. Receiver Operating Characteristic (ROC) curves and PrecisionAeRecall (PR) curves were additionally generated for the highest-performing PR analysis was included because it provides a more informative assessment of minority-class performance than ROC analysis when positive cases are relatively rare (Saito & Rehmsmeier, 2. To further investigate classification behaviour, confusion matrices were constructed for all model configurations. These matrices enabled examination of true positives, true negatives, false positives, and false negatives, providing insight into each modelAos ability to identify employees at risk of leaving. Finally, model interpretability was assessed through feature importance analysis using the Random Forest classifier trained on the complete feature set (M. This analysis enabled identification of the most influential predictors of employee turnover and facilitated interpretation of the relative contribution of the six conceptual domains. Evaluation Metrics The mathematical definitions of the evaluation metrics are given as follows: where TP denotes true positives. TN true negatives. FP false positives, and FN false negatives. In the context of employee turnover prediction. Recall is particularly important because false negatives correspond to employees who Hadiaty et al. / International Journal of Mathematics. Statistics, and Computing. Vol. No. 2, pp. 259-274, 2026 are at risk of leaving but remain undetected by the prediction system, thereby preventing timely organisational Results and Discussion Cross-Validation Results To evaluate model stability and predictive ability during the development phase, repeated stratified 10-fold crossvalidation was performed on the balanced SMOTE training dataset. After oversampling, the training subset increased from 1,176 to approximately 1,972 observations. Each model configuration was evaluated ten times using different random seeds, and the mean classification accuracy and standard deviation were recorded. This procedure provides a robust estimate of model performance while mitigating the influence of random partitioning effects (Powers & Atyabi, 2. Cross-validation results for all 12 model configurations across the three machine learning algorithms are presented in Table 3. Overall, substantial performance differences were observed among the feature configurations and classification algorithms, indicating that the domain composition of employee attributes plays a significant role in predicting employee turnover. These results further demonstrate that predictive performance generally improves with the inclusion of additional complementary domains in the model. Table 3: Cross-validation performance of ANN. RF, and SVM on the SMOTE balanced training dataset . = 1,. , where results are reported as mean accuracy (%) A standard deviation across 10 replicates of 10-fold cross-validation Model M10 M11 Attributes ANN Mean (%) ANN SD RF Mean (%) RF SD SVM Mean (%) SVM SD As shown in Table 3. Random Forest consistently achieved the highest cross-validation accuracy across all feature The strongest overall training performance was achieved by M0-RF, which achieved 90. 92% accuracy with a standard deviation of only 0. 25%, demonstrating excellent stability across cross-validation iterations. SVM followed with 89. 76% accuracy for M0, while ANN achieved 87. Standard deviation values were generally below 1%, indicating that all three algorithms exhibited stable behavior across different training-test partitions. The low variability observed in Table 2 indicates that the reported results are not significantly affected by random sampling effects and therefore provide reliable estimates of model performance. These findings support the effectiveness of the iterative cross-validation strategy in reducing partition bias. A clear performance hierarchy emerged among the single-domain models (M1AeM. Among these configurations, the Compensation domain (M. achieved the highest predictive performance, achieving 82. 35% accuracy with RF. The Career Development domain (M. closely followed at 82. Conversely, the Relationship & Supervision domain (M. consistently produced the weakest results across all algorithms, with accuracy ranging from 63. 31% to These findings indicate that compensation and career-related information carries significantly more predictive information about employee turnover than the supervisory relationship variable alone. The results further demonstrate the benefits of combining information from multiple employee domains. Models M7AeM11 consistently outperformed all single-domain configurations regardless of the classification algorithm used. For example. M11 achieved accuracies of 84. 10%, 89. 32%, and 86. 03% for ANN. RF, and SVM, respectively, substantially outperforming any individual domain. This pattern suggests that employee turnover behavior is influenced by multiple interacting factors rather than a single organizational dimension. The superiority of multi-domain models is particularly evident for configurations containing the Work Environment domain (D. Models M7. M8. M9, and M11Aiall of which include D4Airank among the strongestperforming configurations. This observation provides preliminary evidence that work environment factors may play a central role in employee turnover prediction. The importance of D4 becomes even more apparent during the independent validation stage discussed in the following section. The overall trends are visualised in Figure 3A, which compares mean accuracy and standard deviation for all modelAealgorithm combinations. The figure shows a clear separation between single-domain and multi-domain configurations, with the latter consistently achieving higher predictive performance. Meanwhile. Figure 3B illustrates the relationship between mean accuracy and model stability. Hadiaty et al. / International Journal of Mathematics. Statistics, and Computing. Vol. No. 2, pp. 259-274, 2026 revealing that the highest-performing models also maintain relatively low variability across replications. Taken together, the cross-validation results indicate that both domain composition and algorithm selection substantially influence predictive performance. While Random Forest achieves the strongest training accuracy, the superiority of multi-domain configurationsAiparticularly those incorporating Work Environment (D. Compensation (D. , and Career Development (D. Aisuggests that employee turnover is best understood as a multidimensional phenomenon requiring information from multiple organisational domains. These findings motivate further evaluation using the independent holdout validation dataset presented in the next section. Figure 3: Model performance on the training and test dataset. (A) Grouped bar chart showing mean accuracy with error bars (ASD) for all 12 models y 3 algorithms. The dashed line marks the 70% threshold, (B) Mean accuracy vs. standard deviation across replications illustrating the accuracyAestability trade-off Holdout Validation Results While cross-validation provides an estimate of model performance during development, the primary test of predictive ability is the ability to generalize to previously unseen data. Therefore, all 36 model-algorithm combinations were evaluated on an independent holdout validation dataset . = . , which was completely excluded from model training. SMOTE balancing, and cross-validation procedures. This evaluation provides an unbiased assessment of real-world predictive performance and allows for direct comparison of domain configurations under realistic application conditions. Validation results are presented in Table 4, which reports Accuracy. Precision. Recall, and ROC AUC for all model configurations. These metrics were chosen to provide a comprehensive assessment of classification performance, particularly under class imbalance conditions where accuracy alone may not adequately reflect the predictive ability of the minority class. Table 4: Validation performance of all model configurations on an independent holdout dataset . = . where Precision and Recall are reported for the minority class (Attrition = Ye. Model Attr. M10 M11 ANN Acc% ANN Prec ANN Rec ANN AUC Acc% Prec Rec AUC SVM Acc% SVM Prec SVM Rec SVM AUC The results demonstrate that predictive performance varies substantially across domain configurations, confirming that the composition of employee attribute domains has a direct influence on model effectiveness. In general, multidomain models consistently outperform single-domain models across all evaluation metrics, indicating that employee turnover intention is influenced by multiple interacting factors rather than isolated organisational dimensions. Among all evaluated configurations. M0-SVM achieved the highest validation accuracy of 84. 01%, representing the strongest overall classification performance. However, this model requires all 30 available predictor attributes. M11-SVM, which utilises only 16 attributes from Job Characteristics (D. Work Environment (D. , and Hadiaty et al. / International Journal of Mathematics. Statistics, and Computing. Vol. No. 2, pp. 259-274, 2026 Career Development (D. , achieved an accuracy of 82. 65%, only 1. 36 percentage points lower than M0-SVM. This finding is particularly important because M11 achieves comparable predictive performance while reducing the number of required attributes by approximately 46. 7%, making it substantially more practical for organisational Beyond classification accuracy, discriminative performance provides additional insight into model quality. The highest ROC AUC was achieved by M11-ANN . , followed by M7-ANN . M11-SVM . , and M7SVM . These results indicate that M11 possesses the strongest overall ability to distinguish between employees who remain and employees who leave across different classification thresholds. Consequently. M11 can be considered the most parsimonious high-performing configuration, balancing predictive capability, model simplicity, and practical applicability. The comparative validation performance across all models is illustrated in Figure 4, which presents Accuracy. Precision. Recall, and ROC AUC for every modelAealgorithm combination. The figure clearly demonstrates the superior performance of multi-domain configurations relative to single-domain models and highlights the strong generalisation capability of M0. M8. M9, and M11. Figure 4: Validation performance of all 12 model configurations across three machine learning algorithms on the independent holdout dataset . = . (A) Accuracy (%), (B) Precision, (C) Recall, and (D) ROC AUC The performance of single-domain models provides further insight into the predictive value of individual employee attribute categories. Among these configurations, the Compensation domain (M. achieves the strongest standalone performance, reaching 73. 81% accuracy using RF, followed closely by the Career Development domain (M. at In contrast, the Relationship & Supervision domain (M. consistently exhibits the weakest performance, with validation accuracies ranging from 51. 36% to 60. These findings suggest that compensation- and careerrelated variables contain substantially more predictive information regarding employee attrition than supervisory relationship variables alone. Precision and recall values reveal additional differences in minority-class detection capability. Single-domain models generally exhibit low precision values, ranging from 0. 167 to 0. 303, indicating a high proportion of falsepositive attrition predictions. Such performance would limit their usefulness in practical HR settings where intervention resources are often constrained. Conversely. M0. M8. M9, and M11 achieve substantially higher precision while maintaining moderate recall, resulting in a more balanced ability to identify employees genuinely at risk of leaving. Confusion Matrix Analysis While aggregate performance metrics such as accuracy, precision, recall, and ROC AUC provide a useful summary of model effectiveness, they do not reveal the specific classification errors made by each model. Therefore, confusion matrix analysis was conducted for the top-performing models identified in the validation stage. Specifically. M0SVM was selected as the model with the highest validation accuracy. M11-SVM as the best parsimonious model, and M11-ANN as the model with the highest ROC AUC. These three configurations collectively represent the strongest predictive models developed in this study. The corresponding confusion matrices are summarised in Table 5. In each matrix. TN and TP represent correctly classified employees, whereas FP and FN indicate classification errors. From Hadiaty et al. / International Journal of Mathematics. Statistics, and Computing. Vol. No. 2, pp. 259-274, 2026 an HR perspective, false negatives are particularly important because they correspond to employees who are at risk of leaving but are incorrectly predicted to stay, thereby preventing timely intervention. Table 5: Confusion matrices of the top-performing models on the validation dataset . = . Model M0-SVM M11-SVM M11-ANN As shown in Table 5, all three models demonstrate strong classification performance for the majority class (Sta. , correctly identifying more than 85% of stay employees. Among them. M0-SVM achieves the highest overall accuracy, correctly classifying 247 of the 294 validation instances. However, this performance is achieved using all 30 available predictor attributes. In contrast. M11-SVM correctly classifies 243 employees while utilising only 16 attributes drawn from the Job Characteristics (D. Work Environment (D. , and Career Development (D. Despite reducing the number of predictor variables by nearly half. M11-SVM produces performance that is highly comparable to the full-feature M0-SVM model. This finding reinforces the potential of domain-based feature selection to improve model efficiency without substantial loss of predictive performance. The strongest minority-class detection capability is observed for M11-ANN, which correctly identifies 27 employees who actually left the organisation. As indicated in Table 5, this model produces the lowest number of false negatives among the three configurations while simultaneously achieving the highest ROC AUC reported in the study . This result suggests that M11-ANN provides superior discrimination between stay and leave employees across different classification thresholds. ROC and PrecisionAeRecall Analysis Although accuracy, precision, and recall provide useful summaries of classification performance, they depend on a specific decision threshold and may therefore provide an incomplete picture of model discrimination capability. further evaluate the ability of the models to distinguish between employees who stay and those who leave. ROC and PrecisionAeRecall (PR) analyses were performed. Because employee attrition prediction involves a highly imbalanced dataset. PR curves are particularly informative and complement ROC analysis by focusing on minority-class detection performance (Saito & Rehmsmeier, 2. Based on the validation results presented in Table 4, the four strongestperforming configurations. M7. M8. M9, and M11, were selected for further analysis. These models consistently achieved superior performance across multiple evaluation metrics and represent the most competitive domain combinations identified in this study. The resulting ROC and PR curves are presented in Figure 6. Figure 6: Performance evaluation of the top-performing models on the validation dataset. (A) ROC curves for M7. M8. M9, and M11 across ANN. RF, and SVM classifiers. The dashed diagonal line represents random classification performance, (B) PR curves for the same models As shown in Figure 6A, all selected models perform substantially better than random classification, with ROC curves consistently positioned above the diagonal reference line. Among all evaluated configurations. M11-ANN achieves the highest discriminative performance with a ROC AUC of 0. 782, followed by M7-ANN . M11SVM . , and M7-SVM . These results indicate that ANN and SVM provide stronger separation between stay and leave employees than RF for the most informative domain combinations. The ROC analysis also reveals an interesting difference between training and validation performance. During cross-validation. RF consistently achieved the highest classification accuracy across nearly all model configurations. However, as shown in Figure 6A. RF is generally surpassed by ANN and SVM in terms of validation ROC AUC. Hadiaty et al. / International Journal of Mathematics. Statistics, and Computing. Vol. No. 2, pp. 259-274, 2026 This finding suggests that RF may exhibit mild overfitting when trained on complex multi-domain feature sets, whereas ANN and SVM demonstrate stronger generalisation to previously unseen employee records. The PrecisionAeRecall curves shown in Figure 6B provide additional evidence regarding minority-class prediction Because only 16. 1% of employees belong to the attrition class. PR analysis offers a more stringent assessment of model usefulness than ROC analysis alone. Consistent with the ROC results. M11-ANN and M7-ANN achieve the strongest PR performance, maintaining higher precision across a wide range of recall values. This behaviour indicates superior identification of employees at risk of leaving while simultaneously limiting the number of false-positive predictions. The strong PR performance of M11-ANN is particularly noteworthy because the model relies on only 16 attributes distributed across three domains: Job Characteristics (D. Work Environment (D. , and Career Development (D. The ability to achieve both the highest ROC AUC and the strongest PR performance using a reduced feature set highlights the effectiveness of these domains in capturing the underlying drivers of employee A further observation from Figure 6 is the recurring importance of the Work Environment domain (D. All topperforming models analysed in this section incorporate D4, either alone or in combination with other domains. This finding suggests that variables such as job satisfaction, workAelife balance, overtime status, environmental satisfaction, and commuting distance contribute substantially to model discrimination capability. The consistent presence of D4 among the strongest-performing configurations reinforces the notion that work environment factors represent some of the most immediate behavioural antecedents of employee turnover. Feature Importance and Algorithm Comparison To better understand the factors driving employee turnover prediction and to compare the behavior of the three machine learning algorithms, feature importance analysis and performance heatmaps were examined. Feature importance was obtained from a Random Forest model trained on the full feature set (M. , while validation performance patterns across all model configurations were visualized using ROC AUC and accuracy heatmaps. The results are presented in Figure 7. Figure 7: Feature importance and algorithm comparison. (A) The 20 most important features obtained from a Random Forest model trained on the full feature set (M. , grouped by six attribute domains. (B) ROC AUC heatmaps for all 12 model configurations across ANN. RF, and SVM. (C) Validation accuracy heatmaps for all 12 model configurations across ANN. RF, and SVM As shown in Figure 7A, the five most influential predictors are OverTime (D. MonthlyIncome (D. Age (D. TotalWorkingYears (D. , and YearsAtCompany (D. Collectively, these variables accounted for the majority of the total predictive significance. Among them. Overtime emerged as the single most influential predictor, indicating that employees who regularly work overtime are significantly more likely to exhibit turnover behavior than those with standard work schedules. The significance of Overtime is consistent with previous findings that excessive workload and prolonged working hours negatively impact employee well-being, work-life balance, and organizational commitment (Junaidi et al. , 2. Similarly, the high significance of MonthlyIncome supports previous evidence suggesting that compensation remains a significant determinant of employee retention decisions (Jogi et al. , 2. The significance of TotalWorkingYears and YearsAtCompany further suggests that career maturity and tenure in the organization play a significant role in shaping employee mobility patterns. Hadiaty et al. / International Journal of Mathematics. Statistics, and Computing. Vol. No. 2, pp. 259-274, 2026 A broader examination of Figure 7A reveals that the most influential variables largely come from the domains of Work Environment (D. Compensation (D. , and Career Development (D. This observation aligns with the validation results presented in Section 4. 2, where model configurations containing these domains consistently achieved the strongest predictive performance. The convergence of feature importance analysis and validation performance therefore strengthens the evidence that these domains contain the most informative predictors of employee turnover. A comparison of the performance of the three machine learning algorithms is illustrated through the heatmaps in Figure 7B and Figure 7C. The ROC AUC heatmap in Figure 7B shows that ANN and SVM generally achieve stronger discrimination performance than RF for the most informative multi-domain configurations. Specifically. M11-ANN achieves the highest ROC AUC value observed in this study . , followed by M11-SVM . These findings suggest that ANN and SVM are more effective in separating retained and departed employees when there are complex interactions between employee attributes. Conversely, the validation accuracy heatmap in Figure 7C shows that RF often achieves competitive or even superior accuracy for some configurations, particularly those involving the compensation and career domains. However, this superiority was not consistently reflected in the ROC/AUC performance, suggesting that RF may favor majority class classification and exhibit slightly weaker minority class discrimination compared to ANN and SVM. Comparing performance trends across all model configurations reveals distinct algorithm characteristics. achieved the highest cross-validation accuracy throughout model development and demonstrated strong robustness across feature subsets. ANN provided the strongest overall discriminatory ability, as reflected by the highest ROC/AUC values. SVM demonstrated the most consistent validation performance across increasing levels of model complexity and achieved the highest overall validation accuracy . 01%) when all attributes were used. To further investigate the relative contribution of individual attribute domains and the consistency of algorithm performance across different feature configurations, the results are summarised in Figure 8. As shown in Figure 8A, substantial differences exist among the six domain-specific models (M1AeM. , indicating that certain employee information categories contain considerably stronger predictive signals than others. Meanwhile. Figure 8B illustrates the progression of validation accuracy across all model configurations, providing a direct comparison of algorithm behaviour as model complexity increases from single-domain to multi-domain representations. Figure 8: Domain contribution and algorithm consistency analysis. (A) Validation accuracy of single-domain models (M1AeM. across ANN. RF, and SVM, (B) Validation accuracy across all model configurations, highlighting the bestperforming model for each algorithm As shown in Figure 8A, the Compensation domain (D. achieved the strongest standalone predictive performance, 81% validation accuracy using RF, followed closely by the Career Development domain (D. In contrast, the Relationship & Supervision domain (D. consistently produced the weakest results across all three These findings indicate that compensation and career-related variables contain the most informative independent signals associated with employee attrition, whereas supervisory relationship variables alone provide limited predictive value. Furthermore. Figure 8A demonstrates that none of the individual domains achieved predictive performance comparable to the strongest multi-domain configurations. Even the best single-domain models remained substantially below the performance levels observed for M8. M9. M11, and M0, reinforcing the multidimensional nature of employee turnover behaviour. The trend shown in Figure 8B further supports this observation. Validation accuracy generally increased as additional domains were incorporated into the models, indicating that complementary information from multiple employee dimensions improves prediction performance. Among the three algorithms. SVM exhibited the most consistent improvement with increasing model complexity, ultimately achieving the highest validation accuracy for both M0 . 01%) and M11 . 65%). RF achieved competitive performance for several intermediate configurations but tended to plateau as additional domains were added, whereas ANN demonstrated stronger discrimination capability despite slightly lower accuracy in some configurations. Taken together, the results presented in Figure 8 Hadiaty et al. / International Journal of Mathematics. Statistics, and Computing. Vol. No. 2, pp. 259-274, 2026 confirm that employee turnover intention cannot be adequately explained by any single domain in isolation. Instead, the integration of Job Characteristics (D. Work Environment (D. , and Career Development (D. provides the most effective balance between predictive performance and model simplicity, as demonstrated by the M11 configuration. These findings further support the domain-based modelling framework proposed in this study and highlight the importance of combining complementary employee information for accurate turnover prediction. Discussion The findings of this study provide several important contributions to the employee turnover prediction literature. First, the results demonstrate that the predictive performance of machine learning models depends not only on algorithm selection but also on the composition of employee attribute domains. While previous studies have primarily focused on comparing algorithms using the complete feature set (Fallucchi et al. , 2020. Chung et al. , 2023. Al-Ali et , 2. , the present study shows that different groups of employee attributes contribute unequally to turnover Specifically, the Compensation (D. Work Environment (D. , and Career Development (D. domains consistently emerged as the most informative predictors across both single-domain and multi-domain configurations. These findings directly address the research gap identified in Section 2. 5 and demonstrate the value of domain-based modelling for HR analytics. A notable finding is the strong predictive importance of work environment factors. Models incorporating the Work Environment domain (D. consistently ranked among the best-performing configurations, including M7. M8. M9, and M11. Furthermore, feature importance analysis identified OverTime as the single most influential predictor, followed by variables associated with job satisfaction and environmental satisfaction. This result is highly consistent with the work of Junaidi et al. , who reported that excessive workload and overtime significantly increase employeesAo intention to leave. Similarly. Saufi et al. found that workAelife balance serves as a critical mediating mechanism linking workplace conditions and turnover intention. The present findings therefore reinforce the argument that turnover behaviour is strongly influenced by employeesAo day-to-day work experiences rather than solely by demographic characteristics or supervisory relationships. The Compensation domain (D. also demonstrated substantial predictive power. Among all single-domain models. M3 achieved the highest standalone performance, reaching 73. 81% validation accuracy using RF. Feature importance analysis further identified MonthlyIncome as one of the strongest individual predictors. This finding supports the conclusions of Jogi et al. , who identified compensation adequacy as one of the most consistent determinants of employee retention across industries. Employees who perceive their compensation as insufficient relative to their effort or market opportunities are more likely to consider alternative employment options. The strong contribution of compensation-related variables observed in this study therefore aligns closely with established turnover theories and empirical evidence. Another important observation concerns the Career Development domain (D. Variables such as TotalWorkingYears. YearsAtCompany, and YearsSinceLastPromotion contributed substantially to model performance and feature importance rankings. This result suggests that employees evaluate not only their current working conditions but also their future prospects within the organisation. The finding is consistent with Park et al. , who demonstrated that career advancement opportunities significantly improve turnover prediction It also supports the broader literature indicating that employees are more likely to remain with organisations that provide clear pathways for professional growth and promotion. The domain-based analysis further reveals that employee turnover intention is inherently multidimensional. Although several individual domains achieved moderate predictive performance, none approached the effectiveness of the strongest multi-domain models. For example, the best single-domain model (M3-RF) achieved 73. validation accuracy, whereas M11-SVM achieved 82. 65% using a combination of Job Characteristics. Work Environment, and Career Development domains. This result supports the Theory of Planned Behavior (Ajzen, 1. , which argues that behavioural intentions arise from the interaction of multiple psychological and environmental influences rather than a single determinant. The findings also corroborate the observations of Mobley . , who conceptualised turnover intention as a multifaceted process involving job attitudes, organisational experiences, and perceived alternatives. From an algorithmic perspective, the results reveal an interesting distinction between training and validation Random Forest consistently achieved the highest cross-validation accuracy, reaching 90. 92% for M0, confirming previous studies that identified RF as one of the strongest algorithms for employee attrition prediction (Fallucchi et al. , 2020. Al-Ali et al. , 2. However, during holdout validation. SVM and ANN frequently achieved superior ROC AUC values and more balanced minority-class discrimination. This pattern suggests that RF may capture training patterns effectively but can exhibit mild overfitting when confronted with unseen employee records. Conversely. SVM demonstrated the strongest generalisation capability, achieving the highest overall validation accuracy . 01%), while ANN achieved the highest ROC AUC . These findings are broadly consistent with VapnikAos . theoretical argument that maximum-margin classifiers often generalise well in high-dimensional classification problems. One of the most practically significant findings concerns model efficiency. Although M0-SVM achieved the highest validation accuracy. M11-SVM achieved comparable performance while reducing the number of required Hadiaty et al. / International Journal of Mathematics. Statistics, and Computing. Vol. No. 2, pp. 259-274, 2026 predictor attributes from 30 to 16, representing a reduction of approximately 46. This result has important implications for organisational deployment. Collecting, maintaining, and monitoring fewer variables reduces implementation costs, simplifies data management, and improves model interpretability. Consequently. M11 may represent a more realistic solution for organisations seeking to implement predictive turnover monitoring systems without sacrificing substantial predictive performance. Finally, the results contribute to the growing literature on machine learning applications in human resource analytics by demonstrating that predictive effectiveness can be improved not only through algorithm optimisation but also through intelligent feature-domain selection. Whereas previous studies have primarily focused on developing increasingly sophisticated models (Chung et al. , 2023. Park et al. , 2024. Al-Ali et al. , 2. , the present study shows that understanding the structure and composition of employee information is equally important. The consistent importance of Work Environment. Compensation, and Career Development domains suggests that organisations should prioritise these areas when designing retention strategies and predictive monitoring systems. Collectively, these findings provide both theoretical support for multidimensional turnover frameworks and practical guidance for data-driven employee retention management. Conclussion This study developed and evaluated a domain-based machine learning framework for employee turnover prediction using the IBM HR Analytics Employee Attrition dataset. By organising employee attributes into six conceptual domains and systematically comparing twelve feature configurations across ANN. RF, and SVM classifiers, the study aimed to identify the most informative employee domains, determine the most effective classification algorithm, and assess whether reduced-domain models could achieve performance comparable to the full-feature configuration. The results demonstrate that employee turnover prediction is inherently multidimensional. Models incorporating information from multiple domains consistently outperformed single-domain configurations, indicating that employee attrition cannot be adequately explained by a single category of employee information. Among the individual domains. Compensation and Career Development exhibited the strongest standalone predictive capability, whereas Relationship & Supervision contributed the least predictive information. More importantly, the consistent presence of the Work Environment domain among the highest-performing models highlights the critical role of factors such as overtime, job satisfaction, workAelife balance, and environmental satisfaction in shaping employee turnover behaviour. From an algorithmic perspective, the findings reveal complementary strengths among the evaluated classifiers. Random Forest achieved the highest cross-validation accuracy during model development, demonstrating strong robustness and stability. However. Support Vector Machine exhibited the strongest generalisation capability on unseen data, achieving the highest validation accuracy, while Artificial Neural Network provided the strongest discriminative performance as reflected by the highest ROC AUC. These results suggest that model evaluation should consider both classification accuracy and minority-class detection capability rather than relying on a single performance metric. A particularly important finding is that the M11 configuration, which combined Job Characteristics. Work Environment, and Career Development domains, achieved performance comparable to the fullfeature model while using substantially fewer attributes. This result indicates that effective employee turnover prediction can be achieved without relying on the complete set of employee variables, thereby improving model interpretability, reducing data collection requirements, and enhancing practical applicability in organisational settings. The study also contributes to the employee turnover literature by extending previous machine learning research beyond algorithm comparison and demonstrating the importance of domain composition in predictive modelling. The findings support theoretical perspectives that view turnover intention as the outcome of multiple interacting organisational and individual factors rather than a single determinant. From a practical standpoint, organisations seeking to implement predictive retention systems should prioritise monitoring indicators related to work environment conditions, compensation adequacy, and career development opportunities, as these domains consistently exhibited the strongest relationship with employee attrition. Despite these contributions, several limitations should be acknowledged. First, the study utilised a single publicly available dataset, which may limit the generalisability of the findings across industries, organisational cultures, and geographical contexts. Second, the analysis was based on structured HR variables and did not incorporate unstructured data sources such as employee surveys, performance narratives, or communication records that may contain additional predictive information. Third, although SMOTE was employed to address class imbalance, minority-class prediction remains a challenging problem, as reflected by the moderate recall values observed across several models. Future research should therefore validate the proposed framework using datasets from multiple organisations and industries to assess its external validity. Further studies may also explore explainable artificial intelligence techniques to improve model transparency and support managerial decision-making. In addition, the integration of advanced ensemble methods, deep learning architectures, and unstructured employee data may provide deeper insights into the drivers of turnover intention and further improve predictive performance. Collectively, these directions offer promising opportunities for advancing data-driven employee retention strategies and strengthening the practical value of machine learning in human resource analytics. Hadiaty et al. / International Journal of Mathematics. Statistics, and Computing. Vol. No. 2, pp. 259-274, 2026 References