Inform : Jurnal Ilmiah Bidang Teknologi Informasi dan Komunikasi Vol. 10 No. 1 January 2025. P-ISSN : 2502-3470. E-ISSN : 2581-0367 Optimizing Face Recognition and Emotion Detection in Student Identification Using FaceNet and YOLOv8 Models Ghina Fairuz Mumtaz1. Junta Zeniarja2. Ardytha Luthfiarta3. Almas Najiib Imam Muttaqin4 1,2,3,4 Informatics Department. Universitas Dian Nuswantoro. Indonesia 1ghinafairuz2321@gmail. com(*) 2,3 . unta, ardytha. @dsn. id, 4almasnajiib27@gmail. Received: 2024-12-01. Accepted: 2025-01-13. Published: 2025-01-21 AbstractAi The growing need for efficient student identification systems has driven advancements in face recognition and emotion detection This research presents a single-photo-based system that integrates YOLOv8 for face detection and FaceNet for generating unique facial embeddings, ensuring high-precision student identification and consistent emotion detection under diverse conditions. YOLOv8 localizes faces within images, while FaceNet processes them to generate embeddings for recognition. Emotion detection is performed using these embeddings or an auxiliary emotion classification model. The methodology includes pre-processing images into 64x64 grayscale format, employing image augmentation to enhance model generalization, and evaluating performance using accuracy, precision, recall, and F1 score The experimental dataset comprises 10 formal student photos, with testing conducted on 100 images. Results demonstrate 94% accuracy in face recognition with augmentation, surpassing 92% without it. Emotion detection achieves 95% accuracy in identifying seven emotions, including angry, happy, sad, neutral, fear, disgust, and surprise, despite variations in expression, lighting, and angle. This system provides a scalable and efficient solution for educational applications such as automated student identification, attendance monitoring, and emotion-based learning management. Its potential spans short-term automation in attendance and mental health monitoring, medium-term improvements in personalized learning and campus security, and long-term AI-driven educational advancements while addressing privacy and social acceptance KeywordsAi Face Recognition. Emotion Detection. YOLOv8. Face Net. Image Augmentation. Student Identification includes emotion detection functionality, a key feature for applications requiring real-time emotional analysis. This capability is facilitated through a pre-constructed model and pre-trained weights designed explicitly for recognizing facial As detailed in the provided resources, users can leverage the emotion detection feature within the Deep-Face library to analyze facial expressions directly, enabling the identification of emotions such as angry, disgust, fear, happy, neutral, sad, and surprised. This integration enhances the adaptability of learning environments. It allows for a more nuanced understanding of user engagement and emotional responses during interactions, ultimately fostering a more responsive and supportive educational framework . Building on these advantages, users can focus on applications and research without dealing with technical complexities, which is especially beneficial for academic environments with limited resources . The framework provides high accuracy in face recognition and face emotion recognition, which is important in education . In addition, the use of existing models and simple interfaces reduces development time and costs . This research applies YOLOv8 as a face detector. YOLOv8 supports object detection in a single stage with high efficiency . YOLOv8's capabilities in terms of speed and accuracy make it highly suitable for applications that require real-time response, which is important in the educational context to enhance studentAos learning experience by analyzing their behavior and emotions during the learning process . With the capability of detecting faces in various sizes and INTRODUCTION In the growing digital era, face and emotion recognition technology have become very important in various fields, including security systems, health, and education . education, face identification has great potential to simplify the attendance process and monitor the emotional state of students, which can be an important indicator in the learning process . This technology creates a more adaptive learning environment that responds to students . Through the utilization of the Deep-Face Framework, this research endeavors to assist in the development of a learning environment that is more adaptable. , which is designed to help overcome resource limitations in academic environments by supporting various superior models such as FaceNet . Arc-Face. VGG-Face, and Open-Face, giving users the flexibility to choose the model that best suits their needs without having to develop algorithms from scratch. With the advantages offered by the DeepFace Framework. FaceNet stands out as one of the most prominent models in its According to studies . FaceNet demonstrates high performance in efficiently recognizing and comparing faces, even when faced with variations in lighting, facial expressions, or angles, making it highly suitable for attendance systems based on facial recognition. Furthermore, as highlighted in a study . FaceNet detects emotions through facial features in real-time by processing data locally, thereby enhancing privacy and operational efficiency. In addition to its robust face recognition capabilities, the DeepFace Framework DOI : https://doi. org/10. 25139/inform. Inform : Jurnal Ilmiah Bidang Teknologi Informasi dan Komunikasi Vol. 10 No. 1 January 2025. P-ISSN : 2502-3470. E-ISSN : 2581-0367 orientations. YOLOv8 is an ideal choice to be combined with FaceNet in a deep learning-based student identity recognition system, which not only identifies students but also analyzes their emotions during the learning process. The use of deep learning models for face recognition has shown promising results, with various models developed to improve the accuracy and efficiency of the face recognition One widely used model is FaceNet. FaceNet is highly effective in face recognition systems, achieving high accuracy through an embedding method that allows comparison of distances between facial features in high dimensions . FaceNet generates a vector representation or embedding of facial images, where identical faces are close to each other in high-dimensional space, while different faces are farther apart . This approach allows the system to recognize faces with high precision, even when there are variations in lighting and viewing angle . With a high degree of accuracy and consistency. FaceNet is an ideal choice for applications that require fast and reliable identity verification, such as automated attendance systems in educational institutions . This research integrates FaceNet for face recognition . and YOLOv8 for face detection . With the approach of using one reference photo and ten validation photos, this research aims to measure the accuracy and speed of the system in recognizing faces in various conditions . Facial expression-based emotion recognition technology is also in the spotlight, especially in education, as it can detect studentAos emotional responses to customize learning methods . The combination of these technologies is expected to support the engagement, well-being, and effectiveness of automated attendance systems on campus . Previous research . utilizing YOLOv8 for object detection and FaceNet for face recognition has shown promising results in developing face recognition systems using a combination of YOLO for face detection and FaceNet for feature extraction, as well as improving the security of electronic systems through biometric identification. Although the research . opens up new exploration space in face recognition technology, it faces significant limitations regarding dataset diversity and real-world testing scenarios. These limitations become particularly apparent when examining the research methodology, as the study lacks detailed performance comparisons with relevant current face recognition models, such as Arc-Face, which could provide a more complete perspective on the effectiveness of the developed system . Based on the literature analysis, the use of YOLOv8 and FaceNet in education, especially in student identification and monitoring studentAos emotional states, is still relatively minimal . The implementation of these two technologies in the learning environment has not been widely explored, especially in the aspect of integration for attendance analysis and real-time monitoring of studentAos emotional state . This point indicates a significant research opportunity in developing face recognition and facial emotion recognition designed for modern educational needs . This research creates superior face recognition and emotion recognition specifically for educational situations to overcome the constraints present in previous studies regarding these areas. This research utilizes comprehensive Image Augmentation techniques to improve FaceNet's ability to recognize various student facial characteristics. This research utilizes Image Augmentation, including horizontal inversion, brightness adjustment. Gaussian noise addition, and random cropping to improve the model's robustness to variations in real conditions in learning environments . Image augmentation is an important technique to improve FaceNetAos ability to recognize individuals, especially when identifying students with various facial variations. This technique helps FaceNet to overcome differences such as facial expressions, different lighting conditions, and only sometimes good camera quality . By using this technique. FaceNet can better detect faces and emotions with high accuracy . , thus providing an effective technological solution for recognizing and analyzing Through this innovative strategy, research proves that image augmentation is not just an additional technique but a key element in developing a reliable face recognition system that is reliable and responsive to the complexities of educational environments . The result is a technological solution providing deep insights into studentAos identities and emotional states with improved accuracy. This research aims to develop and apply face recognition and emotion detection technologies, particularly in the educational context, to improve the efficiency and accuracy of student By utilizing FaceNet techniques supported by image augmentation, the proposed system can handle a wide variety of facial characteristics and different environmental conditions . The main objective of this research is to automate the process of verifying attendance and monitoring studentAos emotional state. The system is expected to improve school operational efficiency and support student well-being in the short term. In the medium term, the system can reinforce personalized learning, improve campus security, and provide a more comprehensive learning analysis . In the long term, integration with AI Education technology will enable more personalized learning, early detection of student mental health issues, soft skills development, and improved school Despite challenges such as infrastructure, regulatory, and social acceptance issues, the opportunities offered by these technologies are enormous . These systems not only have the potential to change the way schools manage and support learning but also bring education closer to studentAos needs. This research is expected to serve as a foundation for further studies on integrating face and emotion detection models for wider applications, such as attendance management, mental health analysis, and more responsive learning. II. RESEARCH METHODOLOGY This research methodology is designed to develop a student face recognition and emotion detection system by combining FaceNet and YOLOv8 models in Fig. The method includes a series of interrelated stages to ensure the system can work DOI : https://doi. org/10. 25139/inform. Inform : Jurnal Ilmiah Bidang Teknologi Informasi dan Komunikasi Vol. 10 No. 1 January 2025. P-ISSN : 2502-3470. E-ISSN : 2581-0367 accurately and optimally, from data collection to final evaluation . The research begins with preparing a dataset containing student face images and applying them for the initial detection process . Furthermore, the data obtained through the detection process is reprocessed in the preprocessing stage to improve the model's efficiency . Each step in this research aims to ensure that the system can overcome the challenges of real environments, such as uneven lighting and variations in facial expressions . This cropped facial region is then passed to FaceNet, which transforms the facial images into high-dimensional vector representations, facilitating the comparison of facial The identified face undergoes expression analysis via a trained emotion recognition model leveraging the DeepFace framework . Specifically. DeepFace extracts pertinent features from the image, feeding them into an emotion classifier that discerns one of seven potential emotions, including anger, disgust, fear, happiness, sadness, surprise, or neutrality . This integrated methodology guarantees swift and precise identification of facial expressions, thereby creating a robust pipeline for effectively recognizing and analyzing emotional states. Upon completing the training phase with the extended dataset, the system undergoes a comprehensive evaluation using accuracy, precision, recall, and F1-score metrics to thoroughly assess its capability in recognizing faces and detecting student emotions . This evaluation process highlights the flow of a structured and systematic approach, essential for designing a robust and education-ready system . Dataset The dataset used in this study consists of formal photographs of students from Universitas Dian Nuswantoro taken during their first-year enrolment. These photographs depict students in formal poses wearing the university uniform. Each individual in the dataset is represented by a single photo, resulting in 10 photos from 10 students. The formal photographs shown in Fig. 2 exhibit variations in lighting and size to reflect conditions encountered in real-world applications, thereby adding challenges for the model to achieve accurate identification. These formal photographs are used as the training dataset, totalling 10 photos. After training, testing is conducted using testing data to evaluate face recognition and emotion detection. In the testing dataset, each student is represented by 10 additional photos, with the example that person 1 has 10 photos, person 2 has 10 photos, and so on. Thus, the total testing dataset consists of 100 photos with various emotions, incorporating variations in lighting and size to reflect conditions encountered in real-world applications, thereby increasing the challenges for the model to achieve accurate identification. Fig. Research Framework The stages of this methodology are outlined comprehensively in the research flowchart, beginning with dataset collection. The data undergoes pre-processing, including converting images to grayscale format . Data augmentation techniques enhance dataset diversity and ensure robust model performance across various conditions . The FaceNet model is employed for face recognition through vector representation, while YOLOv8 is an advanced object detector that efficiently and accurately identifies facial regions . This potent combination establishes a foundational framework for subsequent analysis, where the procedure for recognizing facial expressions is initiated by YOLOv8, which detects and extracts the region of interest (ROI) containing the face . Fig. Academic Student Photos student faces in camera images . The YOLOv8 architecture uses the basic principle of a single-shot detector, which allows detection in a single-pass inference process . In this process. YOLOv8 can immediately generate a bounding box, class label, and confidence level for each detected object . One of the YOLOv8 YOLOv8 is an advanced object detection model designed to detect objects in images or videos with high speed and accuracy . YOLOv8 is an object detection model to recognize DOI : https://doi. org/10. 25139/inform. Inform : Jurnal Ilmiah Bidang Teknologi Informasi dan Komunikasi Vol. 10 No. 1 January 2025. P-ISSN : 2502-3470. E-ISSN : 2581-0367 outstanding features of YOLOv8 is its anchor-free approach, which distinguishes it from previous generations of YOLO models . This approach allows YOLOv8 to perform predictions without configuring anchor points, simplifying the detection process, reducing calculation complexity, and improving the model's efficiency and accuracy in various scenarios, including object detection in complex environments . effectively integrate features from large to small scales . With this Neck module. YOLOv8 becomes more sensitive to object details of different sizes and positions in the image, contributing to an overall improvement in detection accuracy . The final part of YOLOv8 is the head, which generates the final prediction through bounding boxes, class labels, and confidence scores for each object. The head combines information from the backbone and neck and performs classification and localization of objects in the image. With the anchor-free configuration, the YOLOv8 head can generate predictions more quickly and accurately without the need to adjust anchor points as in previous YOLO versions . In the final stage, the YOLOv8 detect module compiles and displays the detection results in a bounding box with each object's label and confidence level. This complete process flow, as illustrated in Fig. 4, shows how YOLOv8 provides structured information about the detected objects in the image through its Fig. Conceptual Diagram of Human Detection and Face Recognition various architectural components. With its efficiently designed The processing stage of YOLOv8 starts from the input stage, architecture and stages. YOLOv8 can handle object detection where an input image with a consistent size, such as the one in tasks with high speed and accuracy, making it ideal for various Fig. 3, of 64x64 pixels, is fed into the model. This fixed input applications . size is essential to ensure dimensional consistency during the detection process and makes it easier for the model to handle various input data with a consistent scale . With the same input size. YOLOv8 can optimize calculations and reduce computational load in the early processing stages . After the input is received, the image data is passed to the backbone part of the model. The YOLOv8 backbone is designed to extract key features from images, and in this architecture, the CSPNet (Cross Stage Partial Networ. element is used as the core structure. CSPNet allows YOLOv8 to reduce information redundancy and improve data flow between network layers . This allowed YOLOv8 to extract object features more efficiently without losing important information, thus improving the accuracy of the model in detecting objects . YOLOv8 also utilizes the C2f (Cross Stage Partial with Bottleneck - CSP) layer, which aims to strengthen the feature extraction process in more detail. C2f helps to improve the feature representation of smaller, detailed objects in the image . This layer works by breaking features into stages and processing them incrementally to retrieve richer information, especially for complex images or objects that are difficult to recognize . Furthermore. YOLOv8 uses the SPPF (Spatial Pyramid Fig. Architecture of YOLOv8 Detection Model Pooling-Fas. module to extend the spatial context of the extracted features. SPPF allows the model to simultaneously Pre-Processing identify features from different scales and areas in the image. In the pre-processing stage, after the face area is successfully which is crucial for detecting objects of different sizes . It adds more details about the location and size of objects in the detected, the image is cropped according to the identified image, thus improving detection accuracy on different objects region to ensure that the focus is only on the face, as shown in Fig. 5, where a sample image undergoes grayscale conversion. in a single image . At the Neck stage. YOLOv8 implements the Path The image is then resized to a standard resolution of 64x64 Aggregation Network (PANe. to combine feature information pixels. This step aims to create consistency in the input size fed from different scales. PANet facilitates the detection of small to the model during the training process. This size consistency objects and overlapping objects in a single image because it can is very important so that the model can recognize patterns more DOI : https://doi. org/10. 25139/inform. Inform : Jurnal Ilmiah Bidang Teknologi Informasi dan Komunikasi Vol. 10 No. 1 January 2025. P-ISSN : 2502-3470. E-ISSN : 2581-0367 accurately without being affected by variations in the original size of the image . After the image is cropped, the next step is to convert the image to grayscale format. Grayscale processing was chosen because it only requires one colour channel, unlike colour images that require three channels (RGB). This results in more efficient memory usage and processing. In the context of face detection, colour details are usually not very significant. What is more necessary are geometric features, such as the distance between the eyes, nose, and mouth, as these features are more effective in improving face recognition accuracy than colour information susceptible to lighting changes . Grayscale processing is highly efficient for image classification tasks, where the focus is on geometric features rather than colour information . the original dataset through various transformations, such as rotation, brightness change, and noise addition . Dataset Training The dataset Training stage is an important step in preparing the data to train the model in college studentAos face recognition and emotion detection. This stage involves processing the augmented data to produce an adequate training dataset in quantity and quality. This process is designed to improve the model's ability to generalize so that it not only memorizes patterns in the training data but can also recognize faces in various real conditions, such as differences in lighting, viewing angles, and facial expressions. The augmented image becomes the main component in creating the training dataset at this stage. The augmentation techniques applied include rotation, brightness adjustment, zoom, horizontal inversion, and noise addition. These techniques produce visual variations that aim to replicate various real-world conditions, making the model more resilient to challenges that may occur in real-world applications . The total number of augmented images is adjusted according to the number of original images and the amount of augmentation per image. For illustration, if there are ten original images and each image is augmented with ten augmentations, the training dataset includes 100 face images . Each augmented image is structured in a specific folder. The Fig. Face Image After Being Grayscale file naming is customized with the individual identity. The next stage involves several steps to prepare the image augmentation number, and relevant labels to make tracking data to be more varied and consistent during model training. easier during training. In addition, the augmented images are The first step is determining the augmentation performed on converted to a standard size, such as 64x64 pixels, to ensure each original image. This amount of augmentation can vary, for uniformity of input dimensions during training. This consistent example, by 20, 30, or even up to 40 new images for each size not only supports the efficiency of the computational original image. These new images are created with various process but also improves the model's accuracy . modifications, such as rotation, brightness adjustment, zoom. After processing, this training dataset can use the FaceNet or horizontal inversion. The addition of these variations aims to model in the facial feature extraction stage. This model enrich the dataset, enable the model to be more robust in generates a numerical representation . of each face recognizing faces under various lighting conditions, viewpoints, image, which is then used as the main input in model training. and expressions, and improve its ability to generalize to new This stage ensures that the training data has optimal diversity, data . quality, and structure to support the model's performance in In addition, to ensure consistency in the augmentation recognizing faces and accurately detecting emotions in faceprocess and other processing involving random numbers, an based student identification research . initial value for the random number generator . with a value of 42 was set. By setting the seed value, the random F. FaceNet (Face Recognition. Emotion Detectio. results generated will be the same in each experiment so that In this research. FaceNet is a deep learning model used for experiments can be repeated with consistent and reliable results face recognition and emotion detection of each individual's face . in the dataset. The dataset was obtained from formal photos of the Universitas Dian Nuswantoro students in 2021. Each Image Augmentation individual in the dataset is represented by one formal photo. In the Image Augmentation stage, this process is performed which shows a formal appearance with a university uniform. to increase the variation in the dataset and train the face There are 10 photos from 10 different students, as shown in detection and recognition model more effectively. Fig. Augmentation techniques aim to mimic various real-life After the 10 photos of 10 students shown in Fig. 2 were conditions that may affect the appearance of faces, such as processed in the image augmentation and training dataset, changes in lighting, viewing angle, contrast, and image quality. testing was conducted using a testing dataset to evaluate face These are important so the model can recognize faces in various recognition and emotion detection. In the testing dataset, each situations . Augmenting the dataset has been proven to student is represented by 10 additional photos. For example, improve the accuracy of face recognition models by enriching person 1 has 10 photos, person 2 has 10 photos, and so on. This DOI : https://doi. org/10. 25139/inform. Inform : Jurnal Ilmiah Bidang Teknologi Informasi dan Komunikasi Vol. 10 No. 1 January 2025. P-ISSN : 2502-3470. E-ISSN : 2581-0367 results in a total of 100 photos in the testing dataset. These additional photos include variations in lighting conditions, facial expressions, and image sizes to simulate real-world challenges that may arise in practical applications. Fig. Triplet loss Ai Learning process The embedding generated by FaceNet is not only used to distinguish individual identities but is also utilized in facial emotion analysis. By utilizing the embedding, the system can analyze specific facial patterns, such as lip expression, eyebrow position, or patterns around the eyes, which are associated with various emotions . FaceNet works with a Convolution Neural Networks (CNN) based architecture, which consists of several important layers . These layers extract key features from the face, filter out noise, and represent the face as a fixed-dimensional numerical This approach makes FaceNet excellent at handling various real-world facial conditions, such as lighting changes or dynamic expressions . Fig. FaceNet Architecture FaceNet, whose architecture is illustrated in Fig. 6, is used as the primary pre-trained model in this research to create face embeddings that serve as a unique numerical representation of each individual. The FaceNet model is loaded using the DeepFace Framework, which supports various pre-trained models such as VGG-Face. Open-Face. Arc-Face, and others . FaceNet was chosen due to its ability to generate a fixeddimensional embedding . , which contains specific information about the unique features of the face . FaceNet is designed using CNN (Convolutional Neural Network. to extract unique features from faces and generate face embedding that can be used for various identification and classification tasks . The embedding extraction process using FaceNet starts from the convolutional layers, which extract the face's main spatial features, such as contours and textures, and the position of important elements, such as eyes, nose, and lips . Input images with a standard size of 64x64 pixels provide uniform dimensions, ensuring the model can work efficiently without losing important details on the face. Once the spatial features are extracted, this data is passed to the fully connected layer, where the features are converted into a fixed-dimensional vector of 128 dimensions. This embedding vector represents the unique characteristics of each face and allows the model to distinguish individuals with a high degree of accuracy. Fig. CNN (Convolution Neural Networ. Architecture The CNN workflow process is depicted in Fig. 9, where the Convolutional Layers extract spatial features relevant to emotions from the input image. Examples of these layers include smiles, which convey cheerful feelings, and furrowed brows, which convey angry emotions . Furthermore, the Pooling layer reduces the data size while retaining important features, thus improving computational efficiency without sacrificing accuracy . The output of the Convolution Layers is passed to the Fully Connected layer, which converts the features into numerical representations that Fig. Training Facenet With Triplet Loss Function For Face Recognition can be used for emotion classification . In the final stage, the Output layer uses the Soft-max During training. FaceNet uses a loss function called Triplet Loss, as shown in Fig. This function compares three elements: function to generate probabilities for the six emotion categories anchor . ain face imag. , positive . nother face image from listed in Fig. 10: angry, disgust fear, happy, sad, surprised, and the same individua. , and negative . ace image from a different neutral . In Fig. Triplet Loss ensures that the embedding of anchor and positive has a small distance, while the embedding of anchor and negative has a large distance. This approach allows FaceNet to produce consistent and reliable embedding despite variations in lighting, viewing angle, or facial expression . Fig. Sample Image from JAFFE Dataset DOI : https://doi. org/10. 25139/inform. Inform : Jurnal Ilmiah Bidang Teknologi Informasi dan Komunikasi Vol. 10 No. 1 January 2025. P-ISSN : 2502-3470. E-ISSN : 2581-0367 i. RESULT AND DISCUSSION This research uses 64x64 pixel images for all feature detection and extraction processes, including emotion detection. This size was chosen to maintain computational efficiency without sacrificing accuracy. The Cosine Similarity metric is used for face recognition to measure the angle between two face representation vectors, as defined in Equation . Where X and Y are face representation vectors, this research uses a Threshold 6 to determine face recognition. If the Cosine Similarity value is less than 0. 6, the face is considered recognized, while a value greater than 0. 6 indicates that the face is not recognized . The selection of this threshold aims to minimize detection errors, such as false positives, so that the system can provide more accurate identification . Evaluation The evaluation stage in this research aims to measure the system's performance in face recognition and emotion detection using FaceNet and YOLOv8 models. The evaluation dataset used in this study consists of 100 images, which are evenly divided into ten classes, each representing one student. Each student has ten images, each equipped with a ground truth label that indicates the student's true identity. The dataset is organized in a directory format, where each folder represents the student's name as the ground truth label. The arrangement of the dataset in this structured format aims to facilitate the validation and identity-matching process . This research uses various evaluation metrics to measure the performance of the model's Face Recognition and Emotion Recognition. These evaluation methods include Accuracy. Precision. Recall. F1-Score, and a Confusion Matrix . addition, the Confusion Matrix is used to visualize the model's performance by comparing the predicted and actual labels. This matrix consists of four main components: True Positive (TP). False Positive (FP). True Negative (TN), and False Negative (FN), which are visualized as heat maps for in-depth analysis . Accuracy measures the percentage of correct predictions against the total number of test images using Equation . This metric provides an overview of how often the model makes correct predictions. In the context of accuracy, the Unknown category affects the calculation because the model assigns this label when the tested face cannot be recognized, either because the embedding distance exceeds a specified threshold, the individual does not exist in the database, or face detection fails. Therefore, the Unknown label is treated as an incorrect prediction in the accuracy calculation . Precision is important to evaluate the accuracy of predictions in each class, especially in situations with data imbalance. It measures the proportion of correct positive predictions to all positive predictions using Equation . Recall, as outlined in Equation . , focuses on the ability of the model to detect all positive Recall is a metric used to measure the model's ability to detect all data that belongs to a particular class. To provide a balance between Precision and Recall. F1-Score is used, which is an integrated average of the two. To provide a balance between Precision and Recall, the F1-Score is used, as shown in Equation . This metric ensures that the model focuses on prediction accuracy and effectively recognizes all relevant data. yaycaycaycycycaycayc = ycAycycoycayceyc ycuyce yaycuycycyceycayc ycEycyceyccycnycaycycnycuycuyc ycNycuycycayco yaycoycayciyceyc ycNycycyce ycEycuycycnycycnycyce . cNycE) ycEycyceycaycnycycnycuycu = ycNycycyce ycEycuycycnycycnycyce . cNycE) yaycaycoycyce ycEycuycycnycycnycyce . aycE) ycNycycyce ycEycuycycnycycnycyce . cNycE) ycIyceycaycaycoyco = ycNycycyce ycEycuycycnycycnycyce . cNycE) yaycaycoycyceycAyceyciycaycycnycyce . aycA) ya1 ycIycaycuycyce = 2 yycEycyceycaycnycycnycuycu yycIyceycaycaycoyco ycEycyceycaycnycycnycuycu ycIyceycaycaycoyco ycuOoyc yaycuycycnycuyce = AnycuAnAnycAn Emotion detection is performed after face location detection by YOLOv8, using the generated feature representation . With an image size of 64x64 pixels, this system maintains consistency in feature processing for identity and emotion recognition . Image Augmentation Table I presents the various image augmentation techniques applied to increase the variety of the dataset to strengthen the model's ability to recognize faces and detect emotions. Type TABLE I AUGMENTATION TECHNIQUE Value Filter Detail Applied Multiply And Add To Brightness Multiply . 5, 1. Add (-30, . Affine Log Contrast Scale: 8, 1. Gain: 8, 1. Multiply Multiply: 9, 1. Poisson Noise Lambda , . Enhance Image sharpness Simulates lighting variations by adjusting image brightness Changes in object size in the image scale the image. Adjust contrast Adjusts colour intensity Adds random noise to simulate noise effects in the image Sharpen Alpha: 2, 0. Lightness: 8, 1. Flip Horizontal Probability: With Brightness Channels Add: (-50, . Adjusts the brightness in color Sigma: , 1. Adds blur effect to the image. Effect Gaussian Blur Enhances image details with increased sharpness Flips the image horizontally with a 50% probability DOI : https://doi. org/10. 25139/inform. Inform : Jurnal Ilmiah Bidang Teknologi Informasi dan Komunikasi Vol. 10 No. 1 January 2025. P-ISSN : 2502-3470. E-ISSN : 2581-0367 The first technique. Detail Filter, highlights important details such as facial lines and skin texture by adding sharpness to the Multiply and Add to Brightness adjusts the brightness level by a certain factor, creating lighting variations that resemble real-life conditions . Furthermore. Affine Transformation changes the size or position of objects in the image, ensuring that the model remains effective despite face shifting or rotation . Other techniques, such as Log Contrast and Multiply, focus on adjusting the contrast and intensity of colours to improve the clarity of image features, allowing the model to recognize faces better even if the image has varying luminance or Poisson Noise adds noise that resembles real environmental conditions, helping the model handle lowquality data. Meanwhile. Sharpen accentuates important lines such as nose and lip contours, improving the model's ability to recognize facial features. Flip Horizontal creates a variety of face orientations by flipping the image horizontally, allowing the model to recognize faces from different viewpoints. The With Brightness Channel technique provides additional adjustments to the brightness level to produce richer lighting variations. Finally. Gaussian Blur is applied to reduce certain details in the image, testing the model's ability to recognize faces even when the image is blurry. This combination of augmentation strategies, as seen in Figure 11, results in a dataset that is more diverse and adaptable to obstacles encountered in the real world. These challenges include differences in image size, noise, and lighting. Fig. Augmentation Results Testing Result Based on Table II, tests were conducted on the face recognition and face emotion recognition models using a dataset consisting of 100 photos from 10 students, where each student has ten formal photos. This test aims to evaluate the performance of the model with and without the application of augmentation, as well as the effect of the amount of augmentation on accuracy, precision, recall, and F1 score. TABLE II TESTING RESULTS The Number of Augmented Photos generated per formal photo Application Augmentation Accuracy Precision Recall F1-Score Not Applied 92/95 92/95 100/100 96/97 Applied 39/95 54/95 58/100 56/97 Applied 77/95 81/95 94/100 87/97 Applied 80/95 81/95 99/100 89/97 Applied 89/95 90/95 99/100 94/97 Applied 90/95 90/95 100/100 95/97 Applied 89/95 89/95 100/100 94/97 Applied 91/95 91/95 100/100 95/97 Applied 93/95 93/95 100/100 96/97 Not Applied 92/95 92/95 100/100 96/97 Applied 94/95 94/95 100/100 97/97 In the condition without augmentation . irst ro. , the model showed high performance in face recognition with 92% accuracy, 92% precision, 100% recall, and 96% F1 score. As for emotion detection, the model is consistent with 95% accuracy, 95% precision, 100% recall, and 97% F1 score. This indicates that the model works well on formal data with no additional variations. When augmentation is applied to one image per student . econd ro. , the face recognition model's performance drops significantly, with an accuracy of 39%, precision of 54%, recall at 58%, and F1 score of 56%. However, emotion recognition performance remains stable, with an accuracy of 95%, precision of 95%, recall of 100%, and F1 score of 97%. This decrease occurs because the model faces new variations that require adaptation, while more augmentation data is needed to help the model recognize new patterns. When the number of augmentations increases to 10 images per student . hird ro. , the face recognition performance increases with 77% accuracy, 81% precision, 94% recall, and 87% F1 score, while emotion recognition remains stable at the same metrics . % for all metrics except recall remains 100%). Increasing the augmentation provides more sufficient variation, helping the model recognize more complex patterns on faces. DOI : https://doi. org/10. 25139/inform. Inform : Jurnal Ilmiah Bidang Teknologi Informasi dan Komunikasi Vol. 10 No. 1 January 2025. P-ISSN : 2502-3470. E-ISSN : 2581-0367 In the ninth row of Table II, where 80 images per formal photo are used without applying augmentation, the face recognition performance achieves an accuracy of 92%, precision of 92%, recall of 100%, and an F1 score of 96%, while emotion recognition consistently performs at 95% accuracy, precision, recall, and an F1 score of 97%. model's performance in the previous test results table. This shows that the applied image augmentation helps the model better recognize facial patterns over various data . Fig. Comparison of Accuracy Increasing the amount of augmentation to 80 images per student . ast ro. results in the best performance for face recognition, with an accuracy of 94%, precision of 94%, recall of 100%, and F1 score of 97%, which is even better than without augmentation. As illustrated in Fig. 12, the diagram compares face recognition performance with and without augmentation, highlighting how augmentation significantly improves the model's ability to recognize patterns. This indicates that sufficiently large augmentation helped the model recognize patterns better without losing performance on face Emotion recognition performance remained consistent with 95% accuracy, 95% precision, 100% recall, and 97% F1 score, indicating whether augmentation was applied did not affect emotion detection. Overall, image augmentation improved the generalization ability of the face recognition model, especially when the amount of augmentation is large enough. However, for emotion recognition, the model's performance is unaffected by augmentation or changes in the number of images, indicating that emotion detection is more stable against dataset variations. This means that augmentation is very important for face recognition but does not affect facial emotion detection. Fig. 13 visualizes the process and results of face and emotion recognition using the FaceNet model applied to student data. The Fig. 13 at the top shows the initial stages of the face recognition process, where the original image of the student is converted into a 64x64 pixel grayscale image before going through the feature extraction process using FaceNet. This process produces an embedding vector representing the unique characteristics of the student's face, which is the model. Then, it is used to predict the identity and emotion of the face. At the bottom, the facial recognition results and emotions are shown for various scenarios, such as happy, fear, neutral, surprise, disgust, sad, and angry expressions. Prediction of the student's identity (Actual: Almas, predicted: Alma. is achieved with high accuracy despite variations in facial expressions and lighting conditions, as evidenced by the Fig. Face and Emotion Recognition Results Tests on face recognition and face emotion recognition using FaceNet proved to be highly accurate . , as reflected in Fig. The model was able to identify faces with an accuracy of up to 94% and detect emotions with a stable accuracy of 95%, as seen in Table II. Emotion prediction remained accurate even when there were changes in expression, lighting, or viewing angle in the image. Comparison of FaceNet with Other Models Table i shows that FaceNet consistently outperforms ArcFace in face recognition. FaceNet achieves 92% accuracy without augmentation, much higher than Arc-Face's 37%. With DOI : https://doi. org/10. 25139/inform. Inform : Jurnal Ilmiah Bidang Teknologi Informasi dan Komunikasi Vol. 10 No. 1 January 2025. P-ISSN : 2502-3470. E-ISSN : 2581-0367 augmentation on 80 student images. FaceNet improved its accuracy to 94%, while Arc-Face remained low at 38%. This result confirms that FaceNet is more adaptive to variable data than Arc-Face. experimental results demonstrate that the accuracy remains stable without augmentation at 92%. In contrast, emotion detection showed consistent robustness to data changes, confirming FaceNet's stability for this research. Compared to other models such as Arc-Face, the quantitative analysis TABLE i presented in Table i confirms that FaceNet proved superior FACENET COMPARISON WITH ANOTHER MODEL regarding data utilization efficiency and ability to deal with Number of Photos Application Model Accuracy (%) complex pattern variations. Augmentation Augmentation The integration of FaceNet and YOLOv8 for optimizing face Arc-Face Not Applied and emotion detection has successfully achieved its FaceNet demonstrating high performance across Arc-Face and proving its real-world applicability. Applied FaceNet While the system shows significant promise, notable challenges were identified, particularly regarding YOLOv8's substantial Arc-Face Applied computational requirements and hardware demands, which FaceNet could limit deployment in resource-constrained environments. Additionally. FaceNet's reliance on data augmentation Overall, the integration of FaceNet . and techniques highlights areas for potential improvement in model YOLOv8. offers an efficient and accurate solution Despite these limitations, the research findings for face and emotion recognition. With further optimizations, confirm that the integration of FaceNet and YOLOv8 holds such as more effective use of augmentation and computational considerable potential for educational systems and similar efficiency, this combination has great potential to be applied in applications, with future work focused on enhancing various scenarios, including automatic attendance systems and computational efficiency and reducing system complexity emotion monitoring in education. The comparison with Arcwhile maintaining high-performance standards. Face confirms FaceNetAos superiority as a more robust and The integration of FaceNet and YOLOv8 for optimizing face adaptive model to data variations. recognition and emotion detection has successfully achieved its research objectives, demonstrating high performance across IV. CONCLUSION various test scenarios and proving its real-world applicability. Integrating FaceNet and YOLOv8 models provides an While the system shows significant promise, notable challenges efficient and accurate face recognition and emotion detection were identified, particularly regarding YOLOv8's substantial solution in student identification systems. FaceNet generates computational requirements and hardware demands, which embedding-based facial feature representations, achieving a could limit deployment in resource-constrained environments. face recognition accuracy of 94%. Table i highlights that Additionally. FaceNet's reliance on data augmentation FaceNet consistently outperforms Arc-Face in face recognition techniques highlights areas for potential improvement in model FaceNet achieves 92% accuracy without augmentation. Despite these limitations, the research findings significantly higher than Arc-Face's 37%. When augmentation confirm that the integration of FaceNet and YOLOv8 holds is applied to 80 Photos Augmentation. FaceNet improves its considerable potential for educational systems and similar accuracy to 94%, while Arc-Face remains at a much lower 38%. applications, with future work focused on enhancing This result confirms FaceNet's superior adaptability to variable computational efficiency and reducing system complexity data compared to Arc-Face. while maintaining high-performance standards. Additionally, the emotion detection capability of the system demonstrates stable performance, reaching 95% accuracy ACKNOWLEDGMENT under various conditions, including changes in lighting. The authors sincerely thank the Department of Informatics expression, and viewing angle. The system effectively detects Engineering for their invaluable support and guidance emotions such as angry, fear, neutral, happy, sad, and surprise. throughout the research process. Their unwavering YOLOv8, the face detector, supports FaceNet through fast and commitment to promoting innovation and academic excellence accurate face detection, making the system suitable for realhas been instrumental in completing this project. time scenarios like automatic attendance or emotion monitoring The authors also extend their heartfelt appreciation to their in educational environments. families for their unwavering love, encouragement, and This research also identified the importance of image understanding during the countless hours dedicated to this augmentation in improving the model's generalization ability. research endeavour. Their steadfast support has been a constant Tests show that a large amount of augmentation helps FaceNet source of motivation and inspiration, enabling the authors to better recognize patterns, significantly improving its overcome the challenges and achieve this milestone. performance compared to scenarios without augmentation. Table II indicates that FaceNet achieves 94% accuracy in face REFERENCES