International Journal of Electrical and Computer Engineering (IJECE) Vol. No. June 2025, pp. ISSN: 2088-8708. DOI: 10. 11591/ijece. Deep learning for predicting drug-related problems in diabetes Fatima M. Smadi. Qasem A. Al-Radaideh Department of Information Systems. Yarmouk University. Irbid. Jordan Article Info ABSTRACT Article history: Machine learning and deep learning have made advances in the healthcare In this study, we aim to apply deep learning models to predict the drug-related problems (DRP. status for diabetes patients. Also, to determine the appropriate model to use for classification using deep learning algorithms or machine learning methods to investigate which one performed better results for tabular data by comparing the achieved deep learning results with the machine learning methods to figure out which one gives better results. To apply the deep learning models, the same criteria that were applied in the previous study have been implemented in this investigation, and the same dataset was used. The results show that the machine learning algorithms especially the random forests for predicting the DRPs status outperform the deep learning models. For classification tasks in healthcare for tabular data, the findings of this study show that machine learning methods are the appropriate model instead of using deep learning to apply Received Jun 28, 2024 Revised Jan 3, 2025 Accepted Jan 16, 2025 Keywords: Deep learning Deep neural networks Hyperparameter optimization Long-short term memory Machine learning Tabular data This is an open access article under the CC BY-SA license. Corresponding Author: Fatima M. Smadi Department of Information Systems. Yarmouk University Irbid 21163. Jordan Email: fatima. smadi13@gmail. INTRODUCTION Recently, with large amounts of data becoming available and the improvement in memory handling capabilities of the computers making the process of handling huge amounts of data easier, and the progress of machine learning and deep learning algorithms . , . However, tabular data is the most prevalent data type employed in many real-world applications. It is used in various fields, including medicine, fraud detection, computer vision, and disease diagnostic . In recent years, deep learning has gained a lot of traction in data mining during the last six years, the term Audeep learningAy was coined in 2006 . It is used to investigate healthcare topics in order to improve patient care . However, deep learning is a computational method that allows machines to analyze data like the human brain. Thus, it is a special type of machine learning that involves a deeper level. Also, deep learning differs from a neural network in that it has many layers and neurons in big numbers. On the other hand, the use of deep learning has evolved in many areas in the following applications such as speech recognition, medical research in cancer research to automatically detect cancer cells, medical imaging, drugs discovery, and natural language processing . Ae. Deep neural networks (DNN) trainee the network through feed-forward neural networks and use The structural building block of deep learning is the perceptron . Deep learning contains the input layer, hidden layers, and output layer . In addition, the DNN is characterized that it has more than one hidden layer, the direction of the information in the network is forward . here are no circle. Journal homepage: http://ijece. Int J Elec & Comp Eng ISSN: 2088-8708 from input to the output through hidden layers, also known as multi-layer perceptron (MLP) . Furthermore, the long-short term memory networks (LSTM) implement a more complex recurrent unit with gates to control what information is passed through. Also, it is considered a state-of-the-art model . However, the LSTM contains memory blocks that control the information flow, the key building block behind LSTM is the gates, the memory cells in LSTM layers interact with each other via gates, and the gates are responsible for enabling the LSTM to be able to add or remove information to its cell state. In addition. The LSTM follows the main four steps: forgetting irrelevant history, then performing computations to store relevant information, updating the internal state, and finally generating the output to be an input to the next LSTM layer. Training LSTM network applied the backpropagation method . , . As for using deep learning techniques in the Healthcare area. Kim et al. used the deep belief network (DBN) and different machine learning algorithms to predict cardiovascular risk prediction . ow risk and high ris. The authors evaluated the models using three performance metrics . ccuracy, specificity, and The nayve Bayes (NB) model obtained . %, 63%, 84%), logistic regression (LR) model obtained . %, 69%, 82%), support vector machine (SVM) model obtained . %, 100%, 71%), random forests (RF) model obtained . %, 61%, 82%), deep belief network (DBN) model obtained . %, 82%, 74%), and a statistical DBN model obtained . , 73%, 87%), the results showed that the statistical DBN model outperformed the other classifiers. Korotcov et al. used SVM. RF, and logistic linear regression to compare the performance of different machine learning algorithms to the deep neural networks (DNN) to predict drug discovery. The results of the performance metrics showed accuracy, precision, and recall, respectively. The DNN model obtained . %, 68%, 70%). SVM model obtained . %, 59%, 74%), logistic linear regression model obtained . %, 58%, 74%). RF model obtained . %, 59%, 70%), the results showed that the DNN model outperformed the machine learning algorithms. Other research used deep learning and machine learning to predict brain tumors in the brain . , the authors used deep learning and machine learning algorithms to predict brain tumors in the brain. They used the deep neural network and k-nearest neighbor algorithms . -NN) in their study. The results of the performance metrics showed accuracy, precision, and recall, respectively. The DNN model obtained . %, 97%, 97%), k-nearest neighbor model with k=1 obtained . %, 95%, 95%). The k-NN model with k=3 obtained . %, 89%, 86%). Ayon and Islam . used DNN to diagnose diabetes . bsence and presen. The results of the performance metrics for ten-folds showed accuracy, specificity, and sensitivity, respectively. The DNN model obtained . %, 96%, 97%). Massaro et al. used LSTM. LSTM using artificial records (LSTM-AR), and MLP to predict diabetes by classifying diabetes status and no-diabetes status. The reported result of the performance metrics showed accuracy. The LSTM model obtained . %). LSTM-AR model obtained . %), and MLP model obtained . %), by applying the LSTM-AR there was an improvement in results compared with MLP and LSTM. Tigga and Garg . used LR, k-NN. SVM. NB, decision tree (DT), and RF to predict type 2 diabetes The results of the performance metrics showed accuracy. The LR model obtained . %), k-NN model obtained . %). SVM model obtained . %). NB model obtained . %). DT model obtained . %), and the RF obtained . %). The results showed that the RF model outperformed the other algorithms. Shwartz-Ziv and Armor . demonstrated that when working with tabular data in classification and regression, deep learning models are not all you need. In their study, they compared tree ensemble models such as XGBoost with deep learning models to see which of them gives better results for tabular data. The study found that the XGBoost outperforms the deep models, it requires much less tuning. So, they recommend using ensemble models when using tabular data. Hairani and Priyanto . used SVM and RF combined with the synthetic minority over sampling technique (SMOTE), edited nearest neighbors (ENN), and hybrid SMOTE-ENN methods. The reported result of the performance metrics showed accuracy, sensitivity, and specificity, respectively. The SVM with SMOTE model obtained . %, 70%, 77%). SVM with ENN model obtained . %, 85%, 86%). SVM with SMOTE-ENN model obtained . %, 91%, 88%). RF with SMOTE model obtained . %, 86%, 78%). RF with ENN model obtained . %, 86%, 87%). RF with SMOTE-ENN model obtained . %, 98%, 92%). The findings showed that the RF with SMOTE-ENN model outperformed SMOTE and ENN individually in terms of average performance. Chu et al. used DNN to develop a cardiovascular disease (CVD) risk prediction. The results of the performance metrics showed accuracy, specificity, and sensitivity respectively. The DNN model obtained . 50%, 87. 23%, 88. 06%). The study contributes by applying deep learning techniques to new data that has not been applied before using deep learning models, the data was collected from six Majour hospitals in Jordan especially focusing on diabetic patients to predict the status of drug-related problems. By comparing the results of deep Deep learning for predicting drug-related problems in diabetes patients (Fatima M. Smad. A ISSN: 2088-8708 learning techniques with machine learning methods, the study achieves high accuracy in DRP status prediction and outperforms previous studies. This study highlights a foundation for future research in predicting drug-related problems. Many researchers have used deep learning models and machine learning methods in healthcare systems such as breast cancer, brain cancer, and drug discovery . , . When handling classification tasks in data mining with tabular data in healthcare systems, it becomes unclear to determine which model to use deep learning models or machine learning techniques. Deep learning models have shown the ability to handle large datasets, but they require computational resources and large amounts of data. However, machine learning methods, are known for their effective practical, and easy implementation. This study aims to investigate the best structure of the used deep learning techniques (DNN and LSTM) to predict the status of drug-related problems, and whether to apply deep learning models or machine learning methods when dealing with tabular data for classification. To perform this, the same dataset used in the . has been used. In addition, we compared the results obtained from deep learning models with the machine learning classifiers. RESEARCH METHOD The main purpose of this study is to investigate the best structure of the DNN and LSTM to predict the status of drug-related problems. Also, to find out the effect of applying deep learning compared with the machine learning methods. Additionally, whether to apply deep learning models or machine learning methods when dealing with tabular data for classification. Also, to find out which will give high performance for tabular data compared with the study in . To perform this study, the methodologyAos implementation is subdivided into several steps. Figure 1 summarizes the overall methodology steps in detail. The deep learning models that were experimented in this study are briefly described in the following subsection. Deep learning or neural networks books will provide a detailed description. Figure 1. The research methodology DNN are a kind of neural network inspired by the design of the human nervous system and the brainAos architecture . Additionally. DNN characterized that it has more than one hidden layer, the direction of the information in the network is forward . here are no circle. from input to the output through hidden layers, it also known as MLP . Ae. However, the layers of the DNNs are the input layer, hidden Int J Elec & Comp Eng. Vol. No. June 2025: 2998-3009 Int J Elec & Comp Eng ISSN: 2088-8708 ultiple hidden layer. , and output layer. The hidden layerAos location is located between the input layer and the output layer. Although the input layer is responsible for passing the inputs from the dataset to the next layer, the hidden layers are responsible for applying the non-linear transformation, and the output layer is responsible for producing the output. Furthermore, the essential elements of DNN are neurons, weights, non-linear activation function, and bias. In DNN there is a fully connected layer that is defined through the dense class while every neuron in the layer is considered an input to the neurons in the next layers, the model uses stochastic gradient descent in training the model. As well. The DNN training going through phases that starts with forward propagation followed by the backpropagation and the adjustment process . The forward propagation method in the neural network passes the data through the network starting with the input layer by passing multiple inputs at once. Each input value . cuycn ) needs to be multiplied by the corresponding weights . cycn ) and then added with all the other results for each neuron, then calculating the sum of the weighted inputs to each neuron, and the bias . is added in . , then applying a non-linear activation function. The result of the activation function is considered an output for the layer and an input for the neurons in the next hidden layer. This process will be applied in each layer until we reach the output layer to get the prediction . cU) . ycs = Ocycn ycuycn ycycn yca The backpropagation method is used to minimize the error in the network, when a neural network is trained using forward propagation by randomly initializing the weights of all neurons to apply the prediction, if the prediction for the model is incorrect . rror in the predictio. , the backpropagation must be used to minimize the error, which is calculated by the difference between the output from the forward propagation . rediction outpu. and the expected outputs . esired outpu. The goal is to let the error close to zero. When the prediction error is generated through the forward propagation, it will be in a high number, by applying the backward propagation from the output layer and updating the weights to achieve the first layer, then applying the forward propagation for the second time and generating the prediction error for the second time, the model keeps doing these steps for all samples in the training data until the specified epochs are reached and the error is minimized . The LSTM networks is considered a special type of recurrent neural networks (RNN). It has an input layer. LSTM hidden layers, and an output layer. According to Takeuchi et al. , the RNN has a difficulty in training because of the vanishing gradient problem, thus the LSTM overcomes this problem by composing an input gate, an output gate, and a forget gate. However, the LSTM architecture contains a set of memory blocks, each block containing recurrently connected memory cells that are connected via gates that allow the LSTM to add or remove information from the cell. The memory cellAos activation function allows storing a state for either a short moment or an extended amount of time. Also, each memory cell applies several steps to move the state to the next LSTM hidden layer and implement the same process to reach the output layer. The LSTM is trained through forward propagation and backpropagation methods . There are many applications of LSTM such as Automatic image caption generation and automatic translation of the Also, the LSTM has become increasingly used in health care in recent years. Experiments design In this study, we used the same method that was used in the previous study . , the cross-validation method with 10 folds has been applied to build the deep learning predictive models. The following steps were followed to perform this study: Using the cross-validation with 10 folds in deep learning models. Tuning the hyper-parameters to get the best structure for the DNN and LSTM through building and training the models by the training data. Using the testing data to apply the prediction. Using the confusion matrix metrics to evaluate the deep learning models. Comparing the evaluated model results with the previous study . Performance measures to evaluate the models To evaluate the deep learning models, the same criteria that were applied for the previous study . have been implemented in this investigation. Moreover, the overall performance metrics were generated via the confusion metrics to discover the modelAos performance including accuracy, sensitivity, and specificity. The confusion metrics were calculated using . Accuracy = ycNycE ycNycA ycNycE ycNycA yaycE yaycA Deep learning for predicting drug-related problems in diabetes patients (Fatima M. Smad. A ISSN: 2088-8708 Sensitivity = Specificity = ycNycE ycNycE yaycA ycNycA ycNycA yaycE RESULTS AND DISCUSSION We attempted to tune the best network structure to achieve the best performance. There were many different values from hyper-parameters to train the DNN and LSTM models to enhance it and get the best performance of the DNN and LSTM models based on many experiments to choose the best structure of the DNN and LSTM. Applying the process of training the model with the write hyper-parameters is based on experiments and results by trial and error. It depends on the experience with many experiments by tuning the classifier model, it takes time and effort to tune the hyper-parameters, when tuning hyper-parameters, we have to select some parameters with the correct configuration for each parameter one at a time and then continue to configure the next parameter and so on to get the best performance you are looking for . Deep neural networks experiments and results The DNN algorithm was used to create the classifier model, we applied the cross-validation method. Based on several experiments, we started by defining the baseline architecture to start building the model of the deep neural networks model with the hyper-parameters specified. The baseline architecture that was used to start building the DNN model: a fully connected neural network with three hidden layers with 31 neurons in each hidden layer, respectively, the activation function for each layer was rectified linear unit (ReLU), the dropout was 0. 5, for the optimizer was Adam, the epochs was 50, and the batch size was 32. The results for the baseline were accuracy: 88. 54, sensitivity: 87. 50, and specificity: 89. The following experiments were performed to get the best structure for the DNN model. Tune the number of hidden neurons in each hidden layer The following experiments were used to tune the number of hidden neurons in each hidden layer. Table 1 displays the experiment results on different numbers of hidden neurons. When using 78, 109, 124, 155, and 171 hidden neurons in the three hidden layers, we tested these neurons with epochs . , and batch size . , the best result occurred when we used the 124 hidden neurons, there was an improvement in the results until we reached the 124 hidden neurons used in the DNN model, after . we noted that there was no improvement in the results, we concluded that . hidden neurons used in the DNN model was the best. Table 1. The hidden neurons experiment for DNN Number of hidden neurons Accuracy % Sensitivity % Specificity % Tune the number of hidden layers The following experiments were used to tune the number of hidden layers. Table 2 displays the experiment results on different numbers of hidden layers. When using 3, 4, 5, 6, and 7 hidden layers, we tested these hidden layers with . hidden neurons in each hidden layer respectively, epochs . , and batch size . , the best result occurred when we used three hidden layers, we concluded that the number of hidden layers used in the DNN model was the best, as there was no improvement in the results when increasing the number of hidden layers. Table 2. Number of hidden layers experiment for DNN Number of hidden layers Accuracy % Sensitivity % Int J Elec & Comp Eng. Vol. No. June 2025: 2998-3009 Specificity % Int J Elec & Comp Eng ISSN: 2088-8708 Tune the Epochs for DNN The following experiments were used to tune the epochs. Table 3 displays the experiment results on different Epochs. When using 50, 200, 300, 400, and 500 epochs, we tested these epochs with . hidden neurons in each hidden layer respectively, with three hidden layers, and batch size . , the best result occurred when we used 400 epochs, there was an improvement in the results until we reached the 400 epochs in the DNN model, after 400 epochs, we noted that there was no improvement in the results, we concluded that . epochs used in the DNN model were the best. Table 3. Number of Epochs experiment for DNN Number of Epochs Accuracy % Sensitivity % Specificity % Tune the batch size for DNN The following experiments were used to tune the batch size. Table 4 displays the experiment results on different batch sizes. When using 16, 32, 64, 128, and 256 batch size, we tested these batch sizes with . hidden neurons in each hidden layer respectively, with three hidden layers, and epochs . , the best result occurred when we used 32 batch size, we concluded that the 32-batch size used in the DNN model was the best, as there was no improvement in the results when testing different epochs. Table 4. Number of batch size experiment for DNN Number of batch size Accuracy % Sensitivity % Specificity % Tune the dropout rate for DNN The following experiments were used to tune the dropout. Table 5 displays the experiment results on different dropout rates. When using 20%, 30%, 40%, and 50% dropout rates, we tested these dropouts with . hidden neurons in each hidden layer respectively, with three hidden layers, and epochs . , and batch size . , the best result occurred when we used 50% . dropout rate, we concluded that the 0. 5 dropout used in the DNN model was the best, as there was no improvement in the results when testing different Table 5. Number of dropout experiments for DNN Number of dropouts Accuracy % Sensitivity % Specificity % Tune the activation function for DNN The following experiments were used to tune the activation function. Table 6 displays the experiment results on different activation functions. When using ReLU. Tanh, and Sigmoid activation functions, we tested these activation functions with . hidden neurons in each hidden layer respectively, with three hidden layers, and epochs . , dropout . , and batch size . , the best result occurred when we used ReLU activation function, we concluded that the ReLU activation function used in the DNN model was the best, as there was no improvement in the results when testing different activation functions. Accordingly. Figure 2 summarizes the structure of DNN. Deep learning for predicting drug-related problems in diabetes patients (Fatima M. Smad. A ISSN: 2088-8708 Table 6. The activation functions experiment for DNN The activation functions ReLU Tanh Sigmoid Accuracy % Sensitivity % Specificity % Figure 2. The structure of deep neural networks LSTM networks experiments and results Based on several experiments, we started by defining the baseline architecture to start building the model of the LSTM with the hyper-parameters specified. The results for the baseline architecture were Accuracy: 86. Sensitivity: 84. 37, and Specificity: 88. The number of LSTM layers was . , the hidden neurons were . in each LSTM layer, the epochs were . , the batch size was . , the dropout was . and the optimizer was Adam. The following experiments were performed to get the best structure for the LSTM model. Tune the number of hidden neurons in each LSTM layers The following experiments were used to tune the number of hidden neurons in each hidden layer. Table 7, displays the experiment results on different numbers of hidden neurons. When using 31, 62, 78, 109, and 124 hidden neurons in the three hidden layers, we tested these neurons with epochs . , batch size . , and dropout . , the best result occurred when we used the 31 hidden neurons, there was an improvement in the results when tested 31 hidden neurons used in the LSTM model, after . we noted that there was no improvement in the results, we concluded that . hidden neurons used in the LSTM model was the best. Table 7. The hidden neurons experiment for LSTM Number of hidden neurons Accuracy % Sensitivity % Specificity % Tune the number of LSTM layers The following experiments were used to tune the number of LSTM hidden layers. Table 8 displays the experiment results on different numbers of LSTM hidden layers. When using 3, 4, 5, 6, and 7 LSTM hidden layers, we tested these hidden layers with . hidden neurons in each LSTM hidden layer Int J Elec & Comp Eng. Vol. No. June 2025: 2998-3009 Int J Elec & Comp Eng ISSN: 2088-8708 respectively, epochs . , and batch size . , and dropout . , the best result occurred when we used five LSTM hidden layers, there was an improvement in the results until we reached the five LSTM layers used in the LSTM model, after five LSTM layers we noted that there was no improvement in the results, we concluded that five LSTM layers used in the LSTM model were the best. Table 8. Number of hidden layers experiment for LSTM Number of LSTM layers Accuracy % Sensitivity % Specificity % Tune the epochs for LSTM The following experiments were used to tune the epochs. Table 9 displays the experiment results on different Epochs. When using 50, 100, 200, 300, and 400 epochs, we tested these epochs with . hidden neurons in each LSTM hidden layers respectively, with five LSTM hidden layers, and batch size . , the best result occurred when we used 50 epochs, we concluded that the 50 epochs used in the LSTM model were the best, as there was no improvement in the results when testing different epochs. Table 9. Epochs experiment for LSTM Number of Epochs Accuracy % Sensitivity % Specificity % Tune the batch size for LSTM The following experiments were used to tune the batch size. Table 10 displays the experiment results on different batch sizes. When using 16, 32, 64, 128, and 256 batch size, we tested these batch sizes with . hidden neurons in each LSTM hidden layers respectively, with five LSTM hidden layers, dropout . , and epochs . , the best result occurred when we used 32 batch size, there was an improvement in the results when tested 32 batch size used in the LSTM model, after . we noted that there was no improvement in the results, we concluded that . batch size used in the LSTM model was the best. Table 10. Batch size experiment for LSTM Number of batch size Accuracy % Sensitivity % Specificity % Tune the dropout rate for LSTM The following experiments were used to tune the dropout. Table 11 displays the experiment results on different dropout rates. When using 20%, 30%, 40%, and 50% dropout rates, we tested these dropouts with . hidden neurons in each LSTM hidden layers respectively, with five LSTM hidden layers, and epochs . , and batch size . , the best result occurred when we used 40% . dropout rate, there was an improvement in the results when tested 0. 4 dropouts used in the LSTM model, after . we noted that there was no improvement in the results, we concluded that . dropout used in the LSTM model was the best. Accordingly. Figure 3 summarizes the structure of LSTM. After trying many experiments to find out the best structure, the trials showed that the best structure for the DNN algorithm was when using 124 hidden neurons in three hidden layers, with a dropout rate of 0. , batch size . , and ReLU activation function. Also, the trials showed that the best structure for the LSTM algorithm was when using 31 hidden neurons in five LSTM hidden layers, with a dropout rate 0. , and batch size . A similar specificity achieved by the DNN and LSTM models is observed. Deep learning for predicting drug-related problems in diabetes patients (Fatima M. Smad. A ISSN: 2088-8708 Whereas the accuracy and the sensitivity for the DNN model were higher than the LSTM model. Table 12 demonstrates the results for the DNN and LSTM to get the best structure. Table 11. Number dropout experiment for LSTM Number of dropouts Accuracy % Sensitivity % Specificity % Figure 3. The structure of long-short term memory networks Table 12. Summary of the results obtained by DNN and LSTM best structure Classifiers DNN LSTM Hyperparameters No. of hidden neurons No. of hidden layers Epochs Batch size Dropout Activation function DNN best structure No. of hidden neurons No. of LSTM layers Epochs Batch size Dropout LSTM best structure Options . ,109,124,155,. ,4,5,6,. ,200,300,400,. ,32,64,128,. 2,0. 3,0. 4,0. [ReLU. Tanh. Sigmoi. ,3,400,32,0. , 62,78,109,. ,4,5,6,. ,100,200,300,. ,32,64,128,. 2,0. 3,0. 4,0. ,5,50,32,0. Accuracy % Sensitivity % Specificity % The findings of this study indicate that the obtained results showed that the best structure for the DNN algorithm was when using 124 hidden neurons in three hidden layers, with a dropout rate of 0. 5, epochs . , batch size . , and ReLU activation function. Also, the best structure for the LSTM algorithm was when using 31 hidden neurons in five LSTM hidden layers, with a dropout rate . , epochs . , and batch size . Additionally, the experiment results showed that the DNN obtained an accuracy of 95. sensitivity of 95. 31%, and specificity of 95. On the other hand, the LSTM obtained an accuracy of 50%, sensitivity of 79. 16%, and specificity of 95. After comparing the results for the deep learning models and the machine learning algorithms particularly the random forests for predicting the drug-related Int J Elec & Comp Eng. Vol. No. June 2025: 2998-3009 Int J Elec & Comp Eng ISSN: 2088-8708 problems (DRP. status applied in . , the random forests algorithm outperformed the deep learning models in terms of accuracy and sensitivity when working with tabular data in classification tasks. In the healthcare field, accuracy is essential. Choosing the appropriate model for the data can have a major effect on patient Our study recommends that pharmacists should use machine learning methods when working on identifying the DRPs status of diabetic patients for classification tasks in healthcare to increase the quality of healthcare services and identify the DRPs status for diabetic patients. The primary objective of this study was to apply deep learning models to predict the DRPs status of diabetes patients. We investigate the best structure of the DNN and LSTM to predict the status of drugrelated problems. Moreover, to find out the effect of applying deep learning compared with machine learning Additionally, whether to apply deep learning models or machine learning methods when dealing with tabular data for classification and to find out which will give high performance for tabular data compared with the study in . Choosing the right model for the tabular data depends on our understanding of the nature of the data, whether to use deep learning or machine learning. As demonstrated in the study . , the authors compared tree ensemble models such as XGBoost with deep learning models to determine which performs better results for tabular data. Their results showed that the XGBoost outperforms the deep learning models and requires much less tuning. Therefore, instead of applying deep learning models when working with tabular data, they recommend using ensemble models. Additionally, our results show that machine learning, particularly the random forests method applied in . , performed better than previous studies with high accuracy . 39%), 83%), and sensitivity . 95%) as shown in Table 13. Table 13, summarizes the performance comparison of deep learning and machine learning with previous studies. According to Al-Radaideh et al. , the random forests method achieved the following results accuracy of 97. 39%, sensitivity of 98. 95%, and specificity of 95. 83%). In this study, the experiment results showed that the DNN obtained an accuracy of 95. 57%, sensitivity of 95. 31%, and specificity of 95. Additionally, the LSTM obtained an accuracy of 87. 50%, sensitivity of 79. 16%, and specificity of 95. When comparing the results of the DNN and LSTM with the random forests results, we noted that the random forests algorithm applied in . has achieved the same specificity metric results compared with the DNN and LSTM models as shown in Table 13. Additionally, we noted that the random forests outperformed the results of the DNN in terms of accuracy with an increase of . 82%) and sensitivity with an increase of . Also, it outperformed the results of the LSTM in terms of accuracy with an increase of . 89%) and sensitivity with an increase of . 79%). As a result, when comparing the results, it was found that using machine learning algorithms particularly the random forests used in . to predict the DRPs status obtained the best results compared to the deep learning models in terms of accuracy and sensitivity when working with tabular data in classification tasks. Compared to deep learning models, the machine learning algorithms applied in . appear to be more effective and offer a significant improvement over previous deep learning models in terms of accuracy and sensitivity as well as improved outcomes. In addition, tuning the deep learning hyperparameters to achieve the best structure is a complex process that depends on trial and error and requires time and effort. It also requires a long run-time to train the model, unlike machine learning which requires less time and effort. The data used in this study was collected from six major hospitals in Jordan. Because of this regional restriction, the dataset may be skewed toward the particular regions where these hospitals are located and might not represent large populations, such as those in the US. Future research should expand the data to include data from other regions and populations. This will improve the generalizability of the machine learning method and offer a deeper understanding of drug-related problems across different geographic Table 13. A performance comparison between the applied deep learning models and previous studies Studies . Our study (DNN) Our study (LSTM) Models Statistical DBN DNNs DNNs DNNs LSTM using AR (LSTM-AR) RF with SMOTE-ENN DNNs DNNs LSTM networks Accuracy % Specificity % Sensitivity % Deep learning for predicting drug-related problems in diabetes patients (Fatima M. Smad. A ISSN: 2088-8708 CONCLUSION The results of this study indicate that the machine learning algorithms appear to be more effective and offer a significant improvement over deep learning models in terms of accuracy and sensitivity, as well as showing improved outcomes. The results indicated that the random forests algorithm performed better than the deep learning models when compared to machine learning techniques. In the research field, these findings assist pharmacists in accurately identifying the drug-related problems status of diabetes patients, allowing them to enhance patient care and improve the standards of healthcare generally. Finally, we recommend that instead of using deep learning models for tabular data, pharmacists employ the random forests classifier, as machine learning has been demonstrated to produce better results. Moreover, pharmacists at hospitals in Jordan can implement the random forests classifier. This will enhance their ability to recognize any drug related problems quickly and accurately, which will lead to better patient outcomes, and reduced unwanted complications arising from DRPs, and improved patient outcomes. This classifier ensures greater accuracy in healthcare services. FUNDING INFORMATION The authors declare that no funding was received for this study. AUTHOR CONTRIBUTIONS STATEMENT This journal uses the Contributor Roles Taxonomy (CRediT) to recognize individual author contributions, reduce authorship disputes, and facilitate collaboration. Name of Author Fatima M. Smadi Qasem A. Al-Radaideh C : Conceptualization M : Methodology So : Software Va : Validation Fo : Formal analysis ue ue ue ue ue ue ue ue ue ue ue ue I : Investigation R : Resources D : Data Curation O : Writing - Original Draft E : Writing - Review & Editing ue ue ue ue ue ue ue Vi : Visualization Su : Supervision P : Project administration Fu : Funding acquisition CONFLICT OF INTEREST STATEMENT The authors state no conflict of interest. DATA AVAILABILITY The data that support the findings of this study are available from the author. QAAR, upon reasonable request. REFERENCES