Jurnal Ilmu Komputer dan Informatika (JIKI) P-ISSN: 2807-6664 E-ISSN: 2807-6591 Vol. No. June 2025. Page. https://jiki. jurnal-id. DOI: https://doi. org/10. 54082/jiki. Comparative Analysis of Gaussian Nayve Bayes and Categorical Nayve Bayes Algorithms with Laplace Smoothing in COVID-19 Detection Dila Saputra1. Abdul Aziz Fahmi AoAlauddin2. Mochamad Azizan3 1,2,3 Informatics. Universitas Jenderal Soedirman. Indonesia Email: 1dila. saputra@mhs. id, 2abdul. alauddin@mhs. azizan@mhs. Received: Feb 12, 2025. Revised: Aug 14, 2025. Accepted: Aug 15, 2025. Published: Aug 29, 2025 Abstract In January 2020, it was confirmed that COVID-19 can be transmitted from human to human through the upper respiratory tract with a high infection rate. The number of COVID-19 cases worldwide continued to increase rapidly through close contact, droplets, and airborne transmission. In response, governments and the WHO implemented preventive measures, including COVID-19 treatment preparation, increased emergency healthcare capacity, and patient screening. Early detection of COVID-19 became crucial in taking action, providing treatment, and protecting In the Nayve Bayes algorithm, a potential issue arises with the possibility of zero probabilities for some features or attributes in the COVID-19 prediction training data. Therefore. Laplace Smoothing is used to address this This study aims to compare the average accuracy rates of Gaussian Nayve Bayes and Categorical Nayve Bayes algorithms using different proportions of training data but the same testing data for COVID-19 detection. The methods used in this research are Gaussian Nayve Bayes and Categorical Nayve Bayes with Laplace Smoothing implemented using the Python library called scikit-learn. The research results show that the Gaussian Nayve Bayes algorithm without Laplace Smoothing has an average accuracy of 0. 902165, while with Laplace Smoothing, it has an average accuracy of 0. For the Categorical Nayve Bayes algorithm, without Laplace Smoothing, it has an average accuracy of 0. 983864, while with Laplace Smoothing, it has an average accuracy of 0. In conclusion. Laplace Smoothing plays a significant role in improving the average accuracy of Nayve Bayes algorithms. Categorical Nayve Bayes achieves the highest average accuracy of 0. ith and without Laplace Smoothin. , while Gaussian Nayve Bayes achieves 0. ith and without Laplace Smoothin. Categorical Nayve Bayes has a higher average accuracy compared to Gaussian Nayve Bayes. Keywords: COVID-19. Laplace Smoothing. Nayve Bayes. Python This work is an open access article licensed under a Creative Commons Attribution 4. 0 International License. INTRODUCTION In the early phase of the COVID-19 outbreak, the association of newly identified patients with their visits to the Seafood Wholesale Market suggested a possible zoonotic origin of the disease. Although the original and intermediate hosts of SARS-CoV-2 have not been definitively determined, the phylogenetic closeness between SARS-CoV-2 and coronaviruses of bat origin indicates the possibility that this new virus is related to coronaviruses in bats . As of January 2020, there is strong clinical evidence confirming human-to-human transmission of SARS-CoV-2. The relatively high rate of infection, the mode of transmission through the upper respiratory tract . nd also possibly through contac. , the relatively long incubation period, and the long shedding period of the virus, together with the current global travel pattern, have all been key elements that have allowed this virus to evolve quickly became a pandemic . The total number of COVID-19 cases in the world is still increasing, namely 504,571,336 cases on June 24 2022. Based on scientific evidence that the spread of COVID-19 is very fast and can be transmitted through close contact or droplets as well as through the air, the government and Health Organizations The world (WHO) has taken several preventive steps to help reduce COVID-19 cases, such as preparing COVID-19 treatment for infected patients, increasing emergency treatment capacity in health facilities, and organizing patient screening . Jurnal Ilmu Komputer dan Informatika (JIKI) P-ISSN: 2807-6664 E-ISSN: 2807-6591 Vol. No. June 2025. Page. https://jiki. jurnal-id. DOI: https://doi. org/10. 54082/jiki. Preventive measures have an important role in suppressing COVID-19 cases if protocol therapy is implemented from an early stage. Early detection of COVID-19 is one way to help speed up action for patients, whether to confirm their health condition or to require further testing related to COVID-19. The COVID-19 early detection system is considered very important for patients and the people around them to be able to fight the COVID-19 pandemic, because if the patient gets appropriate and fast treatment, other people around them will also be protected . Nayve Bayes Classifier is a classification method that is rooted in Bayes' theorem. The classification method uses probability and statistical methods proposed by the British scientist Thomas Bayes, namely predicting future opportunities based on previous experience, so it is known as Bayes' Theorem. The main characteristic of the Nayve Bayes Classifier is a very strong . assumption of the independence of each condition/event . Gaussian Nayve Bayes is a classification method that is included in the Nayve Bayes algorithm This method is used to classify data based on the assumption that the features in the data follow a Gaussian distribution . ormal distributio. independent of each other. In Gaussian Nayve Bayes, it is assumed that each feature in the data has a Gaussian distribution with a different mean and variance for each class. This model calculates the posterior probability of the class using Bayes' theorem and then predicts the class with the highest probability . Categorical Nayve Bayes is an implementation of the Categorical Nayve Bayes algorithm for categorically distributed data. This algorithm assumes that each feature, described by index i, has its own categorical distribution. For each feature i in the training data set X. Categorical Nayve Bayes estimates the categorical distribution for each feature i of X conditioned on class y. The index set of samples is defined as . , 2, . , . , with n being the number of samples. In the Categorical Nayve Bayes algorithm, the estimated categorical distribution for each feature i is computed using the maximum likelihood estimation (MLE) method from training data. Next, the class posterior probability is computed using Bayes' theorem by taking into account the categorical distribution of each feature. The aim of this research is to compare the average accuracy level of the Gaussian Nayve Bayes algorithm, and Categorical Nayve Bayes with Laplace Smoothing based on the proportion of training data with the same testing data in detecting COVID-19 in people with certain symptoms. METHOD Data mining is a process, so carrying out the process must comply with the CRISP-DM (CrossIndustry Standard Process for Data minin. CRISP-DM is a data mining standardization prepared by three initiators of the data mining market, namely Daimler Chrysler. SPSS. NCR. CRISPDM does not determine certain standards or characteristics because each data to be analyzed will be reprocessed in the phases within it. In this study, researchers used data mining with the CRISP-DM procedure, but in this study only used five stages out of the six existing stages . The following stages are used as in Figure 1 in this research: Figure 1. Stages of the CRISP-DM procedure . Business Understanding Jurnal Ilmu Komputer dan Informatika (JIKI) Vol. No. June 2025. Page. https://jiki. jurnal-id. DOI: https://doi. org/10. 54082/jiki. P-ISSN: 2807-6664 E-ISSN: 2807-6591 This stage focuses on understanding the project's goals and requirements from a business perspective, then turning this knowledge into a data mining problem definition and an initial plan designed to achieve the goals. The main aim of this research is to compare the level of data separation accuracy of the Gaussian Nayve Bayes and Categorical Nayve Bayes algorithms with Laplace Smoothing based on the proportion of separation between the training data and the testing data. In order to obtain the most optimal algorithm in predicting COVID-19. Data Understanding At this stage, data is collected, identified and understood to be used in this research. The dataset used in this research is Symptoms and COVID Presence (May 2020 dat. Contains the types of disease symptoms present in people suspected of having COVID-19, during the COVID-19 pandemic. This dataset was obtained online from Kaggle. Data Preparation Next, data preparation will be carried out to produce optimal data during modeling. In the data preparation stage there is also preprocessing, which is the process of removing duplicate data, checking inconsistent data and correcting errors in writing words . There are several stages in data preparation, here are examples: Data Cleaning to delete rows with missing values or fill with appropriate values Data Transformation to change categorical data into data that can be understood by algorithms. Dimensional Reduction to select optimal features to be used for modeling. Modelling At the modeling stage, a classification process is carried out with the models proposed in this research, namely Gaussian Nayve Bayes and Categorical Nayve Bayes with Google Collaboratory for grouping types of disease symptoms that exist in humans, during the COVID-19 pandemic. Nayve Bayes itself is a simple probabilistic-based prediction technique based on the application of Bayes' theorem (Bayes' rul. with strong . assumptions of independence. In other words, in Nayve Bayes the model used is an independent feature model . Gaussian Nayve Bayes Gaussian Nayve Bayes is a classification method in data mining which is based on the assumption that the features used in classification follow a normal (Gaussia. This method is often used to classify data with continuous features. Gaussian Nayve Bayes is different from ordinary Nayve Bayes, namely that Gaussian Nayve Bayes has a Gaussian distribution . The formula for Gaussian Nayve Bayes is as follows: cuycn O y. = Oo2yuUyuayc2 exp (Oe . cuycn OeyuNyc )2 2yuayc2 In equation 1, ycE. cUycn . is the probability of variable ycUycn given class y, yuU is a mathematical constant with a value of approximately 3. 14159, yuayc is the standard deviation of variable ycUycn in class y, yuNyc is the mean of variable ycUycn in class y, and exp is the exponential function that calculates yce ycu , where e is Euler's number with a value of approximately 2. To combine the smoothing variance in the Gaussian Nayve Bayes formula, we can add a smoothing variable . sually referred to as ) to the variance in equation 1. Thus, the modified equation 1 will be: = . cu OeyuN )2 Oo2yuU. uayc2 y. ycn yc exp (Oe 2. 2 y. ) yc Adding to the variance in equation 2 ensures that the variance is always greater than zero, so that the probability does not become undefined. This helps maintain the stability and reliability of the Jurnal Ilmu Komputer dan Informatika (JIKI) Vol. No. June 2025. Page. https://jiki. jurnal-id. DOI: https://doi. org/10. 54082/jiki. P-ISSN: 2807-6664 E-ISSN: 2807-6591 model in carrying out probability estimates by considering smoothing variance. The commonly used value is = 1, which is the value for the Laplace Smoothing method. Categorical Nayve Bayes Categorical Nayve Bayes is applied to categorically distributed data. It can be assumed that each feature, described by index i, has its own categorical distribution. The formula for Categorical Nayve Bayes is as follows: cuycn = yc O yc = yca . = ycAycycnyca yu ycAyca yuycuycn In Equation 3, the conditional probability ycE. cuycn = yc | yc = yc. is calculated by adding to the count of occurrences of the categorical value t in class c, and then dividing this by the total number of samples in class c plus multiplied by the number of possible categorical values for the feature ycuycn . Laplace Smoothing is a method that is widely used, as well as smoothing which is called default smoothing and the oldest smoothing ever implemented in Nayve Bayes . The Laplace Smoothing method is used to avoid zero probability when there are category values that do not appear in a particular class in the training data. With the adjustment, the conditional probability P. i = t | y = . will always be greater than zero, so there is no missing probability. For the Laplace Smoothing Method, the value of = 1. Evaluation At the evaluation stage, a classification process was carried out using several algorithms, namely Nayve Bayes. Gaussian Nayve Bayes, and Categorical Nayve Bayes with Laplace Smoothing to see accuracy results using Python on Google Colaboratory. Evaluation is carried out in depth with the aim that the results at the modeling stage are in line with the targets to be achieved in the business understanding stage. RESULT Business Understanding In this stage, the focus is on understanding the objectives of the project and the requirements from a business perspective within the context of COVID-19 case prediction. This understanding is then translated into a clear data mining problem definition and an initial plan is designed to achieve the stated The primary goal of this research is to predict COVID-19 cases using the Gaussian Nayve Bayes and Categorical Nayve Bayes algorithms with Laplace Smoothing. In this context, the study aims to identify the most optimal algorithm for predicting COVID-19 cases based on different proportions of training and testing data separation. The evaluation process involves interpreting the results of the data mining models, as demonstrated during the modeling phase in the previous stage . At this stage, activities are carried out to prepare an initial strategy for achieving the research This includes collecting relevant data related to COVID-19 cases, selecting appropriate parameters for both algorithms, and designing suitable evaluation methods to measure their predictive Data Understanding At this stage, data collection, identification, and understanding are carried out for the research. The dataset used in this study is "Symptoms and COVID Presence (May 2020 dat. " It contains types of symptoms experienced by individuals suspected of having COVID-19 during the pandemic period. Figure 2 represent an understanding of the data used in this research. In Figure 2, the data used includes Breathing Problem. Fever. Dry Cough. Sore Throat. Running Nose. Asthma. Chronic Lung Disease. Headache. Heart Disease. Diabetes. Hypertension. Fatigue. Gastrointestinal issues. Abroad Travel. Contact with COVID Patient. Attended Large Gathering. Visited Public Exposed Places. Family Working in Public Exposed Places. Wearing Masks. Sanitization from Market, and COVID-19. The dataset consists of 5434 rows and 21 columns. This stage begins with data Jurnal Ilmu Komputer dan Informatika (JIKI) P-ISSN: 2807-6664 E-ISSN: 2807-6591 Vol. No. June 2025. Page. https://jiki. jurnal-id. DOI: https://doi. org/10. 54082/jiki. collection, followed by processes to gain a deep understanding of the data, identify data quality issues, or detect interesting aspects of the data that can be used to hypothesize hidden information. Figure 2. COVID-19 Symptoms Dataset To facilitate understanding of the columns in each dataset used, there is a detailed breakdown of the variables in each column of the dataset along with their values or data listed in Table 1. Table 1. Indicators in the Symptoms and COVID Presence Dataset (May 2020 dat. Variabel Nama Variabel Tipe Data Value Breathing Problem Binomial Yes/No Fever Binomial Yes/No Dry Cough Binomial Yes/No Sore throat Binomial Yes/No Running Nose Binomial Yes/No Asthma Binomial Yes/No Chronic Lung Disease Binomial Yes/No Headache Binomial Yes/No Heart Disease Binomial Yes/No X10 Diabetes Binomial Yes/No X11 Hyper Tension Binomial Yes/No X12 Fatigue Binomial Yes/No X13 Gastrointestinal Binomial Yes/No X14 Abroad travel Binomial Yes/No X15 Contact with COVID Patient Binomial Yes/No X16 Attended Large Gathering Binomial Yes/No X17 Visited Public Exposed Places Binomial Yes/No X18 Family working in Public Exposed Places Binomial Yes/No X19 Wearing Mask Binomial Yes/No X20 Sanitazion from Market Binomial Yes/No COVID-19 Binomial Yes/No Data Preparation This is an example of the use of sub-chapters in a paper. Sub-chapters are allowed to be included in all chapters, except in the conclusion. In the data preparation stage, a process of cleaning and transforming the data is carried out. Initially, the data consists of a mix of text and numbers, which is then converted into numerical boolean This is intended to make the data easier for algorithms to read and understand. The following is the data preparation for this research: Jurnal Ilmu Komputer dan Informatika (JIKI) P-ISSN: 2807-6664 E-ISSN: 2807-6591 Vol. No. June 2025. Page. https://jiki. jurnal-id. DOI: https://doi. org/10. 54082/jiki. Figure 3. Dataset after conversion to numerical boolean Next, feature selection will be performed using the Python library SelectKBest with the chisquare method. SelectKBest uses the chi-square statistical method to measure the relationship between each feature and the target variable. The chi-square statistic is used to test the independence between two categorical variables. In the context of feature selection. SelectKBest with chi-square calculates the chi-square score for each feature and selects the top K features with the highest scores. The chi-square score reflects the extent to which a feature influences the target variable. Therefore, a chi-square score analysis will be conducted on each feature. Figure 4. Chi-squareAos Score As shown in Figure 4, the highest chi-square scores with a threshold score of 80 are found in the features Abroad travel. Attended Large Gathering. Sore throat. Breathing Problem. Contact with COVID Patient. Dry Cough. Fever, and Family working in Public Exposed Places. Therefore, these eight features will be selected for the modeling process using the Gaussian Nayve Bayes and Categorical Nayve Bayes algorithms. Below is the information on the dataset that has been transformed and feature selection has been performed, as presented in Table 2. Table 2. Symptoms and COVID Presence Dataset (May 2020 dat. - Variable X after Data Transformation and Feature Selection No Variabel Variable Name Data Type Value X14 Abroad travel Binomial 1 or 0 X16 Attended Large Gathering Binomial 1 or 0 Sore throat Binomial 1 or 0 Breathing Problem Binomial 1 or 0 X15 Contact with COVID Patient Binomial 1 or 0 Jurnal Ilmu Komputer dan Informatika (JIKI) P-ISSN: 2807-6664 E-ISSN: 2807-6591 X18 Dry Cough Fever Family working in Public Exposed Places Vol. No. June 2025. Page. https://jiki. jurnal-id. DOI: https://doi. org/10. 54082/jiki. Binomial Binomial Binomial 1 or 0 1 or 0 1 or 0 Modelling The modeling stage involves preparing the dataset for testing the accuracy of Gaussian Nayve Bayes and Categorical Nayve Bayes algorithms. For Gaussian Nayve Bayes, two models are developed: one without Laplace Smoothing and one with Laplace Smoothing. Similarly, for Categorical Nayve Bayes, two models are created: one without Laplace Smoothing and one with Laplace Smoothing. The implementation of these models utilizes the Python library scikit-learn . , which includes implementations of Gaussian Nayve Bayes and Categorical Nayve Bayes algorithms. This setup can be seen in Figure 5 below. Figure 5. Application of Gaussian Nayve Bayes and Categorical Nayve Bayes Models Testing is conducted by training the models with different proportions of training data while keeping the testing data consistent at 10% of the entire dataset. The proportions of training data range from 10% to 90% of the entire dataset. Below are the analysis results for each model as shown in Figure 6 and the accuracy results for each iteration of training data in Table 3. Figure 6. Accuracy Analysis Across Algorithms/Models Jurnal Ilmu Komputer dan Informatika (JIKI) P-ISSN: 2807-6664 E-ISSN: 2807-6591 Vol. No. June 2025. Page. https://jiki. jurnal-id. DOI: https://doi. org/10. 54082/jiki. Table 3. Comparison of Accuracy Among Algorithms/Models with Different Training Data Proportions and Consistent Testing Data Categorical Nayve Bayes Data Gaussian Gaussian Nayve Bayes Categorical dengan Laplace Training Nayve Bayes dengan Laplace Smoothing Nayve Bayes Smoothing In Figure 6 and Table 3. Gaussian Nayve Bayes without Laplace Smoothing shows fluctuating accuracy across iterations of training data from 10% to 60%, achieving a highest accuracy of 0. and a lowest of 0. From 70% to 90% of training data, the accuracy remains consistent at Gaussian Nayve Bayes with Laplace Smoothing improves its accuracy compared to without Laplace Smoothing, with accuracy stabilizing across iterations of training data from 10% to 90%, ranging from a highest of 0. 974265 to a lowest of 0. Both Categorical Nayve Bayes without Laplace Smoothing and with Laplace Smoothing show identical accuracy values from iterations of training data from 20% to 90%, with a slight difference observed at 10% training data: 0. accuracy for Categorical Nayve Bayes without Laplace Smoothing and 0. 985294 accuracy for Categorical Nayve Bayes with Laplace Smoothing. Evaluation Classification conducted on the types of symptoms present in individuals suspected of having COVID-19 during the COVID-19 pandemic period utilized several algorithm models, including Gaussian Nayve Bayes and Categorical Nayve Bayes with and without Laplace Smoothing. The evaluation process yields accuracy values and average accuracy values for the algorithms in Table 3 as Figure 7. Comparison of Average Accuracy Between Gaussian Nayve Bayes and Categorical Nayve Bayes Algorithms Based on Figure 7, the average accuracy of Gaussian Nayve Bayes without Laplace Smoothing is 902165, while the average accuracy of Gaussian Nayve Bayes with Laplace Smoothing is 0. The difference in average accuracy is 0. For Categorical Nayve Bayes, the average accuracy without Laplace Smoothing is 0. 983864, and with Laplace Smoothing, it is 0. The difference in Jurnal Ilmu Komputer dan Informatika (JIKI) P-ISSN: 2807-6664 E-ISSN: 2807-6591 Vol. No. June 2025. Page. https://jiki. jurnal-id. DOI: https://doi. org/10. 54082/jiki. average accuracy is 0. These results demonstrate that Laplace Smoothing significantly improves the accuracy of the Gaussian Nayve Bayes algorithm, with an average accuracy increase of 7. On the other hand. Laplace Smoothing has a minimal impact on improving the average accuracy of the Categorical Nayve Bayes algorithm, with an increase of only 0. DISCUSSIONS In analyzing the Symptoms and COVID Presence (May 2020 dat. using Nayve Bayes classifiers, specifically Categorical Nayve Bayes (CNB) and Gaussian Nayve Bayes (GNB), several key observations and considerations emerge: Nature of the Data: The dataset comprises symptoms categorized as either present or absent, represented in a binary or categorical format . This categorical nature inherently aligns well with the assumptions of Categorical Nayve Bayes, which models features as discrete probabilities based on their frequencies. Distribution Assumptions: A Categorical Nayve Bayes: Assumes each feature is independent and calculates probabilities based on the frequency of category values. This model is suitable for the binary or categorical representation of COVID-19 symptoms in the dataset. A Gaussian Nayve Bayes: Assumes a normal distribution of numerical features. Given that COVID-19 symptoms are encoded as binary or categorical variables. Gaussian Nayve Bayes may not accurately capture the true distribution of the data. Model Performance: A Categorical Nayve Bayes: In this experiment. Categorical Nayve Bayes demonstrates higher accuracy compared to Gaussian Nayve Bayes. This is evident from the higher average accuracy values observed in CNB without Laplace Smoothing, compared to GNB with Laplace Smoothing. CNB leverages simplicity and is well-suited for the categorical structure of the data. A Gaussian Nayve Bayes: Although Gaussian Nayve Bayes can improve accuracy with Laplace Smoothing, its performance remains lower than Categorical Nayve Bayes. This suggests that the normal distribution assumption of numerical features in GNB may not fully match the actual distribution of COVID-19 symptom data that is categorical in nature. Effect of Laplace Smoothing: In Gaussian Nayve Bayes. Laplace Smoothing helps mitigate issues with zero or very low probabilities that can impact model accuracy. However, its impact is not as pronounced as seen in Categorical Nayve Bayes, indicating that the effectiveness of Laplace Smoothing depends on how well the data distribution aligns with the model assumptions. CONCLUSION Based on the results and discussions conducted, it can be concluded that the use of Laplace Smoothing plays a crucial role in improving the accuracy of Nayve Bayes algorithms such as Gaussian Nayve Bayes and Categorical Nayve Bayes. In terms of average accuracy. Categorical Nayve Bayes achieved the highest average accuracy of 0. oth without Laplace Smoothing and with Laplace Smoothin. On the other hand. Gaussian Nayve Bayes achieved an average accuracy of 0. oth without Laplace Smoothing and with Laplace Smoothin. The increase in training data also proved to be a factor in enhancing the accuracy of Nayve Bayes algorithms with Laplace Smoothing. Therefore, the issue of COVID-19 detection based on symptom data during the COVID-19 pandemic is considered successfully addressed due to achieving high accuracy values. CONFLICT OF INTEREST The authors declares that there is no conflict of interest between the authors or with research object in this paper. REFERENCES