Vol. 3 No. November 2025 . ISSN: 2987-6990 DOI: https://doi. org/10. 60005/coreid. Natural Language Processing and Random Forest for Mental Health Symptom Identification Using Social Media Data Sigit Sugara1. Popon Dauni2. Novianti Indah Putri3. Yogi Saputra4 Nana Suryana5 Department of Information System. Univ. Kebangsaan Republik Indonesia. Indonesia Department of Informatics Engineering. UIN Sunan Gunung Djati Bandung. Indonesia 1,2,3,5 Article Info Article history: Received May 13, 2025 Revised October 26, 2025 Accepted November 26, 2025 Keywords: Natural Language Processing. Random Forest. Mental health. Text analysis. Machine learning, symptom detection ABSTRACT This study examines the use of Natural Language Processing (NLP) and the Random Forest algorithm to identify mental health symptoms from social media text data. The rapid growth of social media has produced large amounts of subjective and context-dependent text, making it challenging to detect expressions related to anxiety, depression, and stress. The objective of this research is to develop a text classification model that can recognize these mental health symptoms using machine learning techniques. The proposed approach applies a structured preprocessing pipeline, including case folding, text cleansing, language normalization, negation handling, stop word removal, and tokenization, followed by feature extraction using Term FrequencyAeInverse Document Frequency (TF-IDF). The processed data are then classified using a Random Forest model. Experimental results show that the model achieves an overall accuracy of approximately 80%, indicating reliable performance in identifying mental health-related content, although prediction confidence remains moderate due to overlapping and ambiguous language patterns. This study contributes to the academic field by demonstrating the applicability of classical NLP-based machine learning methods for mental health text analysis and provides practical value as a foundation for automated early screening and monitoring systems using social media data. This is an open access article under the CC BY-SA license. Corresponding Author: Popon Dauni Information System Department. Faculty of Computer Science and Information Systems. Univ. Kebangsaan Republik Indonesia. Jln. Terusan Halimun No. 37 (Pelajar Pejuang . Bandung. Jawa Barat. Indonesia. Email: popon. dauni@ukri. INTRODUCTION Social media usage has become an inseparable part of daily life, with over 4. 8 billion users as of 2023. Platforms like Facebook. Instagram, and Twitter influence many aspects of life, including users' mental health. Several studies show that interactions on social media can impact anxiety disorders, stress, and depression, especially among adolescents and young adults . Previous research has employed systematic review methods and qualitative approaches. However, these studies have limitations in terms of result generalization, as they rely on subjective interpretation of data. response to these challenges, the current research uses a quantitative approach by leveraging Machine Learning technology to identify mental health symptoms based on sentiments found in social media content . Mental health, which forms the basis of this research, focuses on human information processing, including perception and pattern recognition in text. This research aims to identify mental health symptoms for social media users . In this research, we implement Natural Language Processing (NLP) and Random Forest. , two branches of Machine Learning. NLP identifies language patterns related to mental health, while the Random Forest model is used as a classification method to determine whether content shows symptoms of CoreID Journal | Vol. No. November 2025: 99-106 anxiety, depression, or stress. The combination of NLP and Random Forest provides strength in processing text data and producing accurate classifications . Natural Language Processing (NLP) is an artificial intelligence field focusing on the interaction between computers and humans using natural language. NLP has experienced rapid development due to the availability of large amounts of text data and the need for more human-like communication between computers and humans . Random Forest is an ensemble learning approach developed by Breiman to address classification and regression problems, using multiple models to improve accuracy in solving the same problem . Despite the growing body of research on mental health analysis using social media data, significant gaps persist at the methodological and practical levels . Existing studies largely rely on qualitative assessments or sentiment-based analyses, which are inherently subjective and difficult to generalize across large and diverse datasets. Conversely, recent approaches employing deep learning models often prioritize high predictive performance at the expense of computational efficiency, interpretability, and deployment feasibility, particularly in real-time or early screening contexts. Moreover, there is a lack of rigorous empirical studies that systematically evaluate classical machine learning models integrated with structured NLP preprocessing pipelines for multi-class mental health symptom identification. This gap is especially evident in research that balances accuracy, model transparency, and computational cost while addressing overlapping linguistic patterns associated with anxiety, depression, and stress. Therefore, a critical research gap exists in developing and validating an efficient, interpretable, and scalable machine learning framework that can reliably identify mental health symptoms from large-scale social media text data and support practical implementation in real-world mental health monitoring systems. The main difference between this research and previous studies lies in the approach used. This research not only focuses on qualitative analysis but also applies quantitative methods through NLP and Random Forest to produce a faster, more accurate, and real-time detection system for monitoring mental health through social METHOD The methodology in this research is divided into several stages that follow the Agile approach, focusing on analysis, design, implementation, and testing. The data processing workflow includes data collection, preprocessing, feature extraction using TF-IDF, and classification using Random Forest. Figure 1. Flowchart Methodology 1 Data Collection Data was collected through web scraping from Twitter using Python libraries Selenium and Beautiful Soup. Additionally, labeled datasets from Kaggle were used to train the model. The total dataset included 93,720 entries, with 21,961 relevant labeled data points across three categories: anxiety, depression, and stress. 2 Text Preprocessing Text preprocessing is a critical step in text classification. The following preprocessing steps were Natural Language Processing and Random Forest for Mental Health A (Popon Dauni, et a. - 100 CoreID Journal | Vol. No. November 2025: 99-110 Converting all text to lowercase to standardize the text format. Removing non-alphabetic characters, symbols, and punctuation to reduce noise. Standardizing text by converting slang or abbreviations to their standard forms. Handling negation words by replacing them with standardized forms to maintain the semantic Eliminating non-descriptive words such as articles and prepositions that don't contribute to the Breaking text into smaller units . for further processing. Table 1. Case Folding Before 0 is upset that he can'tupdate his Facebook by . 1 @Kenichan I dived manytimes for the ball. Man. 2 my whole body feelsitchy and like its on fire 3 @nationwideclass no,it's not behaving at all. 4 @Kwesidei not the wholecrew After 0 is upset that he can'tupdate his facebook by . 1 @kenichan i dived manytimes for the ball. 2 my whole body feelsitchy and like its on fire 3 @nationwideclass no,it's not behaving at all. 4 @kwesidei not the wholecrew Table 2. Cleansing Before 0 is upset that he can'tupdate his facebook by . 1 @kenichan i dived manytimes for the ball. 2 my whole body feelsitchy and like its on fire 3 @nationwideclass no,it's not behaving at all. 4 @kwesidei not the wholecrew After is upset that he canAot update his facebook by . kenichan i dived many times for the ball man. my whole body feels itchy and like itAos on fire nationwide class no itAos not behaving at all. Kwesi Dei not the whole crew Table 3. Data Normalization Before is upset that he canAot update hisFacebook by . kenichan i dived many times for the ball man. my whole body feels itchy and likeitAos on fire nationwide class noitAos not behaving at all. Non-standard word Dived, man, feels, i t c h y , fire, itAos, not Behaving, no, crew. After He is upset that he cannot update his Facebook. Kenichan. I have dived many times for the ball, friend. My whole body feels itchy and like it is burning. Nationwideclass, no, it is not functioning at all. Kwesidei, not the whole team Table 4. Convert Negation Before is upset that he canAo tupdate his Facebook by . kenichan i dived manytimes for the ball manag. my whole body feelsitchy and like its on fire nationwideclass no it snot behaving at all i . kwesidei not the whole crew After is upset that he canAo tupdate his Facebook by . kenichan i dived manytimes for the ball manag. my whole body feelsitchy and like its on fire nationwideclass not itAos not behaving at all i. kwesidei not the whole crew Table 5. Stopword Removal Before is upset that he can t updatehis facebook by . kenichan i dived many timesfor the ball manag. my whole body feels itchyand like its on fire nationwideclass not it s notbehaving at all i. kwesidei not the whole crew After upset update facebook textingmight cry result. kenichan dived many times ballmanaged save re. whole body feels itchy like fire nationwideclass behaving m mad see4 kwesidei whole crew Table 6. Tokenizing Before upset update facebooktexting might cryresult. kenichan dived many times ball managed save re. whole body feels itchy like fire nationwideclass behaving m mad see kwesidei whole crew After . pset, update,facebook, texting, might, cry,. enichan, dived, many,times, ball, managed, . hole, body,feels, itchy, like, fir. ationwideclass,behaving, m, mad, se. wesidei, whole, cre. 3 Feature Extraction with TF-IDF Term Frequency-Inverse Document Frequency (TF-IDF) weighting was applied to represent the importance of each word in the document collection. This technique helped identify the most relevant terms for each mental health category. The implementation resulted in 65,433 unique terms extracted from the processed text . Natural Language Processing and Random Forest for Mental Health A (Popon Dauni, et a. - 101 CoreID Journal | Vol. No. November 2025: 99-106 # Lakukan pembobotan TF-IDF # Impor modul yang diperlukan from sklearn. feature_extraction. text import TfidfVectorizer # Inisialisasi TfidfVectorizer tfidf_vectorizer = TfidfVectorizer() # Lakukan pembobotan TF-IDF pada data pelatihan X_train_tfidf = tfidf_vectorizer. fit_transform(X_trai. # Transformasi data pengujian menggunakan vectorizer yangsama X_test_tfidf = tfidf_vectorizer. transform(X_tes. print("Pembobotan TF-IDF selesai dilakukan. ") print. "Jumlah fitur: {X_train_tfidf. }") # Latih model menggunakan data yang telah dibobotkan fit(X_train_tfidf, y_trai. # Evaluasi model menggunakan data pengujian yang telahdibobotkan y_pred_tfidf = model. predict(X_test_tfid. accuracy_tfidf = accuracy_score. _test, y_pred_tfid. "\nAkurasi model dengan pembobotan TF-IDF: ccuracy_tfidf:. ") Figure 2. TF-IDF formula 4 Random Forest Classification The Random Forest algorithm was employed for classification due to its ability to handle highdimensional data and reduce overfitting through ensemble learning. The model was trained on the preprocessed and TF-IDF weighted data to classify text into three categories: anxiety, depression, and stress. 5 System Architecture A web-based dashboard was developed using Flask and Dash to provide a user-friendly interface for real-time text analysis. The system allows users to input text and receive predictions about potential mental health symptoms, along with visualization of the prediction probabilities. Figure 3 Workflow System Figure 4 user interface design Natural Language Processing and Random Forest for Mental Health A (Popon Dauni, et a. - 102 CoreID Journal | Vol. No. November 2025: 99-110 RESULTS AND DISCUSSION 1 Model Performance The Random Forest model achieved an overall accuracy of 80% in identifying mental health The model performed particularly well in classifying depression, with a precision of 0. 80 and recall of 0. For anxiety, the model achieved a precision of 0. 85 and recall of 0. 42, while for stress, it achieved a precision of 0. 76 and recall of 0. Laporan Klasifikasi: Anxiety Depression Stress macro avg weighted avg f1-score The confusion matrix revealed that the model was most successful at identifying depression, followed by stress and anxiety. Some misclassifications occurred between anxiety and depression, which can be attributed to the overlap in symptoms and language patterns between these conditions. Figure 5. Confusion Matrix 2 Confidence Levels While the model achieved high accuracy, the average confidence level was around 58%. This reflects the inherent challenges in working with subjective text data, where the same expressions might indicate different mental states depending on context. Despite this limitation, the model provides valuable insights for initial screening of mental health symptoms. Figure 6. Class Distribution Chart 3 System Testing The system was tested with random sentences containing keywords related to mental health Natural Language Processing and Random Forest for Mental Health A (Popon Dauni, et a. - 103 CoreID Journal | Vol. No. November 2025: 99-106 Results showed varying degrees of accuracy in correctly identifying conditions, with better performance on depression-related content compared to anxiety and stress. This aligns with the model's overall performance metrics, where depression had the highest precision and recall values. Table 7 Anxiety Test Sentence Confidance Score Anxiety is a mentalhealth conditioncharacterized bypersistent feelings of worry and fear, which can Interfere with daily life People with Anxiety oftenexperience physical symptoms such asincreased heartrate, sweating, and trembling I can't seem to stopthinking about allthe things that could go wrong. My mind races constantly, and Ican't focus on I'm afraid of what might happen, andit keeps me up at night. Prediction Results True False Table 8 Depression Test Sentence I feel like I'm trapped in a dark hole with no way out Every day feels like a heavy burden, and I can'tfind the strength to carry on. I don't have theenergy or motivation to doanything, even the things I used to It feels like aconstant cloud of sadness is hangingover me, and I can't escape it I'm afraid of what might happen, andit keeps me up at night. Confidance Score Prediction Results True False Anxiety Thot of as Table 9. Stress Test Sentence Confidance Score Prediction Results True He's a complete mtherf**ker I just want to scream at Everyone beingsuch a dck about everything Aom at my breaking point, and I feel like IAom being treated like a psy for complaining This endless cycle is making me feel like a mtherf**ker whoAos lost control. Everything feels like a ctstorm of problems, and I can't escape the False Depression Depression Thot of as Depression Natural Language Processing and Random Forest for Mental Health A (Popon Dauni, et a. - 104 CoreID Journal | Vol. No. November 2025: 99-110 Figure 7. Implemented System CONCLUSION With an overall accuracy of over 80%, this study shows that combining Natural Language Processing with the Random Forest algorithm may successfully identify mental health symptoms from text data. These results show that, especially in large-scale, non-clinical environments like social media monitoring platforms, the suggested method has useful potential as an early screening tool to complement initial mental health assessment. By facilitating simple text analysis and result visualization, the web-based dashboard further improves usability and makes the system usable by both technical and non-technical users. However, this study has a number of drawbacks, such as a reliance on a single language and data source that may restrict generalizability and modest prediction confidence levels brought on by overlapping linguistic patterns among mental health problems. Furthermore, the model's capacity to extract more complex contextual and semantic information is constrained by the application of traditional machine learning In order to increase robustness and applicability across a variety of populations and real-world mental health screening scenarios, future research should concentrate on integrating contextual language representations, such as transformer-based models, performing comparative evaluations with cutting-edge deep learning approaches, and expanding multilingual support. ACKNOWLEDGEMENTS Thank you to all parties who have supported this research so that it can be completed properly. Especially, for the parents who have always supported and guided me in this research. I'm very grateful. REFERENCES