ELKHA : Jurnal Teknik Elektro. Vol. 18 No. April 2026, pp. 65 - 73 ISSN: 1858-1463 . , 2580-6807 . Predicting Breakdown Voltage of Transformer Oil under Copper/Iron Contamination: A Comparative Study of Gradient vs Metaheuristic Training Giovanni Dimas Prenata1*) Faculty of Intelligent Electrical and Informatics Technology. Universitas 17 Agustus 1945 Surabaya. Indonesia Corresponding Email: *)gprenata@untag-sby. Abstract Ae Transformer oil functions as an insulating and cooling medium in high-voltage power systems, whose dielectric condition degrades over service life due to thermal aging, moisture ingress, and metallic contamination, leading to reduced Breakdown Voltage (BDV) and increased insulation failure risk that may necessitate oil regeneration, replacement, or indicate transformer end-of-life. Unlike Dissolved Gas Analysis (DGA), which evaluates transformer faults based on gas decomposition products. BDV directly reflects the dielectric strength of insulating oil and is more sensitive to particulate contamination such as Cu and Fe, making it more suitable for assessing material-level insulation degradation. This study investigates the influence of copper (C. and iron (F. particle contamination on BDV. It compares three Artificial Neural Network (ANN) training strategies for BDV prediction: gradient-based training (DFFNN-Pur. Genetic Algorithm optimization (DFFNNGA), and Grey Wolf Optimizer-based training (DFFNNGWO), using experimental data from 36 transformer oil samples obtained in accordance with IEC 60156:2018. The comparison represents a before-and-after modeling perspective on training strategy rather than on repeated physical testing. The results show that DFFNN-Pure achieved the highest prediction accuracy (RA = 0. RMSE = 0. MAE = 0. 238 kV), while DFFNN-GWO demonstrated stable convergence with competitive accuracy (RA = 0. RMSE = 0. 886 kV), whereas DFFNN-GA exhibited unstable convergence and poor generalization. Unlike previous studies that primarily focus on transformer remaining useful life estimation at the system level, this work emphasizes predicting BDV of transformer oil at the material level under metallic contamination. It provides a systematic comparison between gradient-based and metaheuristic training within the same DFFNN framework, supporting non-destructive condition monitoring and predictive maintenance. Keywords: Breakdown Voltage (BDV). Artificial Neural Network (ANN). Genetic Algorithm (GA). Grey Wolf Optimizer (GWO), metaheuristic optimization. INTRODUCTION Transformer oil serves a dual function as both an electrical insulating medium and a cooling agent in highvoltage power systems. The dielectric reliability of transformer oil is therefore critical to maintaining system stability, as insulation failure may lead to severe Manuscript received 2025-11-08. revised 2026-03-28. accepted 2026-03-30 operational disturbances, irreversible damage, or catastrophic transformer outages. One of the most widely used indicators to evaluate transformer oil quality is the Breakdown Voltage (BDV), which represents the maximum electric stress the oil can withstand before electrical discharge occurs . BDV testing is standardized under IEC 60156, in which an alternating electric field is gradually increased across a defined electrode gap until breakdown is observed, making BDV a fundamental criterion for assessing the suitability of both new and in-service transformer oils. During service operation, transformer oil undergoes gradual degradation due to thermal-oxidative aging, moisture absorption, and contamination by metallic particles originating from copper windings and iron core Thermal and oxidative aging processes generate polar by-products such as organic acids and water, which increase oil conductivity and reduce its dielectric strength . Moisture is particularly detrimental, as it enhances charge-carrier mobility and accelerates discharge-channel formation. In addition, suspended copper (C. and iron (F. particles distort the local electric field and promote particle-bridging mechanisms, significantly reducing BDV and accelerating insulation deterioration . These combined effects not only lower dielectric performance but also shorten the operational lifespan of both liquid and solid insulation To describe the probabilistic behavior of transformer oil breakdown, statistical models such as the Weibull distribution have been widely employed. The shape () and scale () parameters of the Weibull model characterize BDV variability and mean dielectric strength under different aging or contamination conditions . Although effective for reliability estimation and failure probability analysis. Weibull-based approaches remain empirical and limited in capturing complex nonlinear interactions among physicochemical properties, contamination characteristics, and dielectric performance. Artificial Neural Networks (ANN. have increasingly been adopted to model nonlinear relationships between transformer oil condition parameters such as moisture content, contamination level, and physicochemical properties and BDV values. ANN-based models can learn - 65 - This work is licensed under a Creative Commons Attribution 4. 0 License For more information, see https://creativecommons. org/licenses/by-nc-sa/4. Predicting Breakdown Voltage of Transformer Oil under Copper/Iron Contamination (G. D Prenata, et al. complex inputAeoutput mappings directly from experimental data and have shown promising results in insulation diagnostics and condition assessment . Unlike the study in . , which focuses on BDV prediction under general operating conditions using machine learning models, this research specifically investigates the impact of metallic particle contamination (Cu and F. on BDV. evaluates the effectiveness of different ANN training strategies within the same DFFNN architecture. However, the predictive accuracy and generalization capability of ANN models strongly depend on the training algorithm Conventional gradient-based methods, including backpropagation, perform well on smooth error surfaces but are prone to premature convergence and local minima, particularly when datasets are small, noisy, or highly nonlinear . To address these limitations, metaheuristic optimization algorithms have been introduced as alternative training strategies for neural networks. Algorithms such as Genetic Algorithm (GA). Particle Swarm Optimization (PSO), and Grey Wolf Optimizer (GWO) optimize network weights without relying on gradient information, thereby improving robustness to local minima . Among these. GWO is inspired by the hierarchical leadership structure and cooperative hunting behavior of grey wolves, where solution candidates are guided by three dominant agents (, , and ) to balance global exploration and local exploitation during optimization . Previous studies have demonstrated that GWO-based training can improve convergence stability and prediction accuracy in both regression and classification problems, including ANN and Multi-Layer Perceptron (MLP) models . Despite the growing use of ANN and metaheuristic methods in transformer diagnostics, comparative investigations focusing on gradient-based versus metaheuristic-based training strategies for BDV prediction under copper and iron contamination remain limited. particular, systematic comparisons between DFFNN-Pure. DFFNN-GA, and DFFNN-GWO using the same experimental BDV dataset are rarely reported . While previous studies . have applied metaheuristic algorithms such as GA and GWO to optimize neural networks for general prediction and classification tasks, they do not specifically address BDV prediction under metallic contamination, nor evaluate the performance differences between gradient-based and metaheuristic training methods within a unified experimental dataset and network structure. Therefore, this study presents a comparative evaluation of three DFFNN training strategies gradient descent (DFFNN-Pur. Genetic Algorithm optimization (DFFNN-GA), and Grey Wolf Optimizerbased training (DFFNN-GWO for predicting the breakdown voltage of transformer oil contaminated with Cu and Fe particles. The results are expected to provide insights into algorithm performance, convergence behavior, and predictive accuracy, contributing to AIbased condition monitoring and predictive maintenance of power transformers. The main novelty of this study lies in three key aspects. First, this research focuses specifically on material-level BDV prediction under controlled metallic contamination (Cu and F. , which is less explored than system-level diagnostics such as DGA. Second, this study provides a gradient-based metaheuristic training methods within the same DFFNN architecture and dataset, ensuring a fair and controlled Third, the integration of GWO and GA within a unified experimental framework enables detailed analysis of convergence behavior, stability, and predictive performance under nonlinear and small-sample conditions. These contributions distinguish this work from previous studies that primarily apply ANN or metaheuristic optimization independently without direct comparative analysis under identical experimental II. METHODOLOGY Weibull Distribution for Breakdown The Weibull distribution is widely used for reliability analysis and for predicting failure probability . in insulating materials such as transformer oil. This distribution is expressed through the probability density function (PDF) and the cumulative distribution function (CDF) as follows: = yu ( )yuOe1 yce Oe( )yu . Oe ( )yu ya . = 1 Oe yce : breakdown voltage . bserved valu. : shape parameter, which defines the form of the distribution curve : scale parameter, representing the characteristic breakdown voltage f . : probability density function F . : cumulative probability of failure The parameters and are typically estimated from experimental breakdown voltage data using the Maximum Likelihood Estimation (MLE) method or the linear regression approach based on the logAelog transformation: ln(Oe ln. Oe ya . cycn ))) = yu ln ycycn Oe yu ln . By plotting ln(Oe ln. Oe y. ) ycyc ln yc, the slope of the resulting line, which corresponds to while the intercept yields ln . The Weibull distribution helps quantify the probability that a sample of transformer oil will fail . reak dow. under a given voltage stress, and it is employed as a comparative feature or analytical tool in studies of electrical breakdown phenomena in transformer oil. Margins Feedforward Neural Network (DFFNN): Fundamental Formulation Figure 1 shows how to present a figure. The title should be located under the figure. The Figure must be easy to read and have a good resolution/quality. It should be noted that most readers may read the paper in black/white, although the online version is presented with the original color set from the final manuscript. - 66 - Predicting Breakdown Voltage of Transformer Oil under Copper/Iron Contamination (G. D Prenata, et al. A Deep Feedforward Neural Network (DFFNN) consists of an input layer, one or more hidden layers, and an output layer. Each neuron computes the following: ycyc = Ocycn ycuycn ycycnyc ycayc . ycayc = yua . cyc ) . xi : input signal . r activation from the previous wij : weight connecting neuron bj : bias of neuron j zj : linear combination of the input signals a j : activation of the neuron . : activation function : E. = Mutation: Small random perturbations are introduced into the genes to maintain population diversity and avoid premature convergence. Regeneration: The least fit chromosomes are replaced by the newly generated offspring from the crossover and mutation processes. This iterative process continues until a stopping criterion is met, such as the MSE falling below a predefined threshold or the maximum number of generations is reached. Grey Wolf Optimizer (GWO) for Network Training The Grey Wolf Optimizer (GWO) is a swarm-based algorithm that simulates the social hierarchy and cooperative hunting behavior of grey wolves. Its fundamental principles are as follows: Hierarchy: The three best-performing wolves are designated as . , . , and . , while the remaining wolves are classified as O . eOez ) If two hidden layers are employed, the forward propagation process is executed sequentially through these layers until the final output . For regression tasks such as BDV prediction, the Mean Squared Error (MSE) is commonly used as the loss function: ycAycIya = OcycA . c Oe yco )2 . ycA yco=1 yco The backpropagation process computes the gradient of the loss function with respect to the network weights as Output layer gradient: = . Oe y. yua A . ) . Hidden layer gradient: = . cO . )ycN yu . yua A . ) . Weight and bias update: O= yc . Oe yca. coOe. O= yca . Oe yu . : learning rate yua A . : derivative of the activation function . or the sigmoid function: yua . Oe yua . )) This method is efficient when the error surface is smooth, and the dataset is sufficiently large. however, it is prone to being trapped in local minima, particularly for small datasets or highly nonlinear error functions. Search phase: The control coefficient yca decreases linearly from 2 to 0 throughout iterations, shifting the search process from global exploration to local For each agent i, its position is updated according to : yayu = . ycUyu Oe ycUycn | , ycU1 = ycUyu Oe ya1 . yayu = . ycUyu Oe ycUycn | , ycU2 = ycUyu Oe ya2 . yayu = . ycUyu Oe ycUycn | , ycU3 = ycUyu Oe ya3 . ycU ycU ycU ycUycn O= 1 2 3. yayc = 2ycayuayc1 Oe yca , yayc = 2yc2 , yc yun. u , yu , y. and yc1 , yc2 are random numbers uniformly distributed within the range . Evaluation: Each agent . represents a vector of ANN weights. its fitness is computed based on the Mean Squared Error (MSE), and the wolves are ranked as , , . Iteration: The process is repeated until convergence or until the maximum number of iterations is reached. Genetic Algorithm (GA) for Weight Optimization The Genetic Algorithm (GA) is an evolutionary optimization technique inspired by natural selection, employing genetic operations such as selection, crossover, and mutation. For optimizing the weights of an Artificial Neural Network (ANN), the following steps are typically Chromosome representation: Each chromosome encodes a set of network weights and biases . ll weights are flattened into a single gene arra. Fitness function: The fitness of each chromosome is evaluated based on the Mean Squared Error (MSE) or another error function. chromosomes producing lower MSE values are assigned higher fitness scores. Selection: Parent chromosomes are selected based on their fitness probability using methods such as the roulette wheel or tournament selection. Crossover: Genetic material from two parents is combined to generate offspring using single-point, multi-point, or uniform crossover strategies. The GWO algorithm provides an effective balance between exploration and exploitation and does not rely on gradient information, making it suitable for optimizing highly nonlinear objective functions. In the context of ANN training, each wolf represents a set of network weights, and their positions are iteratively updated according to the positions of the , , and wolves until The control parameter a decreases linearly from 2 to 0 over iterations to maintain the balance between global exploration and local exploitation. Cinar . demonstrated that GWO effectively minimizes prediction error on nonlinear datasets such as wind speed and achieves faster convergence than GA. Furthermore. Liu . proposed an Improved Grey Wolf Optimizer (IGWO), which accelerates the search for optimal solutions through dynamic adaptation of control parameters, demonstrating significant performance improvements in ANN parameter optimization tasks. - 67 - Predicting Breakdown Voltage of Transformer Oil under Copper/Iron Contamination (G. D Prenata, et al. Research Design This study aims to compare the performance of the Gradient-based training method (Backpropagatio. with two metaheuristic training approaches, the Genetic Algorithm (GA) and the Grey Wolf Optimizer (GWO) in predicting the Breakdown Voltage (BDV) of transformer oil contaminated with copper (C. and iron (F. The research adopts a quantitative experimental approach, carried out in systematic stages that include the collection of laboratory test data, data preprocessing. DFFNN model training using three distinct optimization strategies, and performance analysis based on statistical metrics. crossover, and mutation, the population evolves toward optimal solutions. The best chromosome is selected as the final model and evaluated using RA. RMSE, and MAE. DFFNN Pure The DFFNN-Pure model is trained using a conventional gradient-based backpropagation algorithm. The process includes data normalization, forward propagation to generate BDV predictions, error computation using Mean Squared Error (MSE), and iterative weight updates using gradient descent. The training continues until convergence or the maximum number of epochs is reached. Model performance is evaluated using RA. RMSE, and MAE. Figure 2. Flowchart of DFFNN-GA Figure 2 presents the operational flow of the DFFNNGA model, in which network weight optimization is performed using a Genetic Algorithm. Instead of direct gradient updates, the model encodes neural weights as chromosomes and evaluates their fitness using the Mean Squared Error (MSE). Through iterative selection, crossover, and mutation processes, the population evolves toward improved solutions. This figure highlights the stochastic nature of GA, which enables global exploration but may introduce convergence instability when population diversity is insufficient. DFFNN-GWO The DFFNN-GWO model integrates the Grey Wolf Optimizer to optimize weights. Each wolf represents a candidate solution, and the best three wolves (, , ) guide the search process. The algorithm iteratively updates positions to balance exploration and exploitation until The final model is evaluated using RA. RMSE, and MAE. Figure 3 depicts the training mechanism of the DFFNNGWO model, which integrates the Grey Wolf Optimizer into neural network training. Each wolf represents a candidate weight configuration, and the hunting behavior, guided by the , , and wolves drives the optimization The linearly decreasing control parameter ensures a balance between global exploration and local This flowchart demonstrates how GWO effectively navigates complex error landscapes without relying on gradient information. Figure 1. Flowchart of DFFNN Pure Figure 1 illustrates the workflow of the DFFNN-Pure model based on gradient descent learning. The process begins with data loading and normalization, followed by forward propagation to estimate the breakdown voltage (BDV). The prediction error is then propagated backward to update network weights using a fixed learning rate. This iterative process continues until the convergence criterion or maximum epoch is reached. The flowchart emphasizes the dependence of gradient-based learning on local error gradients, which may limit convergence performance when the error surface is highly nonlinear. DFFNN-GA The DFFNN-GA model employs a Genetic Algorithm to optimize network weights. Each chromosome represents a set of weights and biases, and fitness is evaluated based on MSE. Through iterative selection, - 68 - Predicting Breakdown Voltage of Transformer Oil under Copper/Iron Contamination (G. D Prenata, et al. DFFNN Model Architecture The primary model employed in this study is a Deep Feedforward Neural Network (DFFNN) configured as Input layer: 6 neurons . orresponding to the number of input feature. Hidden layer: 1-2 layers with 8-12 neurons . mpirically tune. Output layer: 1 neuron . redicted BDV valu. Activation function: Sigmoid Loss function: Mean Squared Error (MSE) The initial weights were randomly initialized, and three distinct training methods were implemented: Table 2. The differences between each method Method Brief Description DFFNN-Pure (Gradient Descen. DFFNN-GA Figure 3. Flowchart of DFFNN-GWO DFFNN-GWO Research Data The dataset used in this study was obtained from breakdown voltage tests of transformer oil conducted in accordance with IEC 60156:2018. The samples were subjected to controlled contamination using copper (C. and iron (F. The dataset comprises 36 test sets compiled in the file MinyakTrafo. xlsx, with each sheet representing a distinct testing condition, such as before and after purification, and with varying levels of metallic The consolidated results from all sheets were merged into a single file for processing and analysis. Conventional backpropagation training using a fixed learning Weight optimization using a population-based Genetic Algorithm involving selection, crossover, and mutation. Weight training using the Grey Wolf Optimizer with three hierarchical roles (, , ) that iteratively update weight positions during optimization. Table 2 compares the fundamental characteristics of the three training strategies applied to the DFFNN model. Gradient-based training relies on deterministic error minimization, whereas GA and GWO utilize populationbased search mechanisms. The table highlights that metaheuristic approaches do not rely on gradient information, making them more suitable for highly nonlinear and multimodal optimization problems. Table 1. List of variables used in the research Variable Description Unit XCA Operating time of the oil XCC Testing temperature AC XCE Moisture content XCE Copper (C. concentration XCI Iron (F. concentration XCI Purification condition . = unpurified, 1 = purifie. Breakdown Voltage (BDV) Training and Validation Procedure The training process was carried out through the following steps: Data Preparation: All numerical data were cleaned of missing values and The dataset was divided into 80% training and 20% testing data. Model Training: Each model was trained for up to 5000 iterations or until the MSE converged. For DFFNN-GA and DFFNN-GWO, the metaheuristic parameters were defined as follows: population size: 10 individuals . olves or chromosome. , maximum generations: 5000 and convergence threshold: MSE < 0. Model Evaluation: After training, all models were tested using the test Predicted BDV values were compared with the actual experimental results. Performance Metrics: The models were evaluated based on three primary performance measures: Table 1 summarizes the input and output variables employed in this study. The selected input features represent operational, thermal, and contamination-related parameters that directly influence transformer oil dielectric The output variable. BDV, serves as a quantitative indicator of insulation performance. This structured variable selection ensures that both physical and chemical degradation factors are incorporated into the predictive model. The dataset consists of several input variables . and one output variable: the average breakdown voltage measured in kilovolts . V). Before model training, all data were normalized to the range . , . to ensure numerical stability and improve convergence during training. ycIycAycIya = Oo OcycA ycn=1. cycn Oe ycn ) . ycA ycAyaya = - 69 - ycA OcycA ycn=1 . cycn Oe ycn |. Predicting Breakdown Voltage of Transformer Oil under Copper/Iron Contamination (G. D Prenata, et al. Oc. c Oe ycn )2 ycI2 = 1 Oe Oc. cycn ycn Oe ycn ) i. RESULTS AND DISCUSSION This study compared three training strategies for the Deep Feedforward Neural Network (DFFNN)-Gradient Descent (DFFNN-Pur. Genetic Algorithm (DFFNNGA), and Grey Wolf Optimizer (DFFNN-GWO) to predict the Breakdown Voltage (BDV) of transformer oil contaminated with copper (C. and iron (F. All models were trained using BDV test data obtained under IEC 60156 standards, which were consolidated from 36 experimental spreadsheets into a single dataset (MinyakTrafo_Final36. The dataset included metal concentration parameters, average BDV test results, and two Weibull distribution parameters ( and ) that Increasing contamination led to a greater reduction in BDV compared to Cu, as Fe exhibits catalytic properties that accelerate oil This phenomenon alters the molecular structure and lowers the critical electric field strength, consistent with the dielectric degradation theory of hydrocarbonbased insulating oils. The MSE convergence curves in Figure 6 reveal distinct optimization characteristics among the three The DFFNN-Pure model stabilized at MSE OO 11 after 3000 iterations, showing a monotonic yet relatively slow convergence. This behavior is typical of gradient-based algorithms that depend on the derivative of the error function. When the error surface is non-convex, gradients tend to approach zero at local minima, causing the learning process to stagnate before reaching the global The DFFNN-GA model exhibited significant MSE oscillations due to the stochastic nature of crossover and mutation processes. Although GA initially explored a wide solution space, the small population size . caused premature convergence in suboptimal regions. contrast, the DFFNN-GWO model displayed a smooth exponential decline in MSE from an initial value of 14 to 00255 by the final iteration. This dynamic reflects an adaptive balance between exploration and exploitation through the linearly decreasing parameter a . Ie . The AeAe leadership mechanism facilitated cooperative learning: guided global exploration, maintained directional stability, and refined local exploitation steps. This behavior aligns with the findings of Mirjalili . and Cinar et al. , who reported that GWO maintains greater population diversity and exhibits superior convergence compared to GA. Experimental Scheme The experiments were conducted through the following steps: Training the DFFNN-Pure model: The model was trained using the gradient-based backpropagation algorithm with a constant learning rate of 0. Training the DFFNN-GA model: The network weights and biases were optimized through a population of chromosomes using selectionAe crossoverAemutation mechanisms. The fitness function was defined as the Mean Squared Error (MSE). Training the DFFNN-GWO model: The weights and biases were optimized through the collaborative behavior of the , , and wolves. The control parameter AuaAy decreased linearly from 2 to 0 over 5000 iterations. Performance Evaluation: The performance of the three models was compared based on MSE. RMSE. MAE, and RA. Table 3. List of variables used in the research Parameter Number of Iterations Population Size Learning Rate Mutation Rate MSE Threshold Activation Function DFFNNPure DFFNNGA DFFNNGWO Ae Ae Sigmoid Ae Sigmoid Ae Ae Sigmoid Table 3 outlines the hyperparameters used for each training method. To ensure a fair comparison, the maximum number of iterations and convergence thresholds were kept identical across all models. Differences in learning mechanisms, such as the learning rate in DFFNN-Pure and the mutation rate in DFFNN-GA, reflect the inherent characteristics of each optimization Training and Validation Procedure To improve the robustness and reliability of the experimental results, each model (DFFNN-Pure. DFFNNGA, and DFFNN-GWO) was executed multiple times . independent run. using different random initializations. The final performance metrics were reported as the mean and standard deviation values of RA. RMSE, and MAE. In addition, a k-fold cross-validation approach . = . was applied to evaluate the modelsAo generalization The dataset was randomly partitioned into five subsets, with each subset used once as test data while the remaining subsets were used for training. This process was repeated until all subsets had been used as testing data. This validation strategy ensures that the reported results are not dependent on a single data split and provides a more reliable assessment of model performance, particularly considering the limited dataset size. Figure 4. MSE convergence curve of three DFFNN training - 70 - Predicting Breakdown Voltage of Transformer Oil under Copper/Iron Contamination (G. D Prenata, et al. The DFFNN-Pure model produced predictions that closely followed the trend of the actual data, achieving RA = 0. RMSE = 0. 296, and MAE = 0. In contrast, the DFFNN-GA model exhibited poor performance with RA = Ae2. 26 and a high error value (RMSE = 9. indicating that the evolutionary process failed to identify an effective weight configuration due to excessive exploration and insufficient elitism. Meanwhile, the DFFNN-GWO model demonstrated the best trade-off between accuracy and stability, with RA = 0. RMSE = 886, and MAE = 0. Figure 4 illustrates the relationship between the predicted and actual BDV values. The blue points (DFFNN-Pur. lie almost entirely along the ideal diagonal line, indicating high prediction accuracy. The green points (DFFNN-GWO) show slight dispersion but remain close to the ideal line. In contrast, the orange points (DFFNNGA) deviate significantly, indicating poor agreement between the predicted and actual BDV values. Figure 5. Relationship between actual BDV values and predicted results of the three DFFNN models The reduction in BDV caused by metallic particles exhibits a highly nonlinear behavior. When voltage is applied, the metallic particles suspended in the oil act as local centers of electric-field intensification. Charge accumulation around these particles increases the field enhancement factor (-. , thereby reducing the effective breakdown voltage according to the following relation: Table 4. Comparison of predictive performance among the three DFFNN models Model DFFNN-Pure DFFNN-GA DFFNN-GWO Ae2. RMSE MAE Table 4 quantitatively compares the predictive performance of the three models using RA. RMSE, and MAE metrics. The results indicate that DFFNN-Pure achieves the highest accuracy, while DFFNN-GWO provides a strong balance between accuracy and stability. Conversely, the poor performance of DFFNN-GA suggests that genetic optimization is less effective for small experimental datasets. To ensure the statistical reliability of the results, each model was evaluated over 10 independent runs. The average performance and standard deviation are summarized as follows: DFFNN-Pure: RA = 0. 996 A 0. RMSE = 0. MAE = 0. 238 A 0. DFFNN-GA: RA = -2. 10 A 0. RMSE = 9. 15 A 1. MAE = 9. 05 A 1. DFFNN-GWO: RA = 0. 971 A 0. RMSE = 0. A 0. MAE = 0. 760 A 0. The results indicate that DFFNN-Pure consistently achieves the highest accuracy with low variance, while DFFNN-GWO demonstrates stable and reliable In contrast. DFFNN-GA exhibits high variability, indicating unstable optimization performance. Figure 5 shows the relationship between actual BDV values and predictions obtained from the three DFFNN The DFFNN-Pure predictions closely follow the ideal diagonal line, indicating high prediction accuracy. The DFFNN-GWO model exhibits slight dispersion but maintains a strong linear relationship with the actual In contrast, the DFFNN-GA predictions deviate significantly, confirming its limited generalization capability under the given dataset conditions. yayca = yaycC 1 yuyce . Such relationships are difficult to represent analytically, making machine learningAebased approaches highly relevant. Gradient-based models (DFFNN-Pur. can approximate smooth continuous functions but tend to lose sensitivity to complex cross-interactions among nonlinear parameters. In contrast. GWO optimizes network weights collectively across the entire solution space rather than locally, which explains why DFFNNGWO produces a more uniform approximation surface and exhibits greater resistance to overfitting. Furthermore, because GWO does not require derivative information, it is more robust to experimental noise and measurement uncertainty. These findings support the results reported by Liu . , indicating that population-based metaheuristics are more adaptive to nonconvex data structures and small sample sizes than gradient-based optimization methods. From the overall findings, three key points can be Numerical stability - GWO demonstrates stable and predictable convergence behavior, whereas GA tends to oscillate due to its random mutation mechanism. Resilience to small datasets - DFFNN-GWO achieves low error rates with only 36 training samples, indicating efficient global search performance. Practical application - DFFNN-GWO can be implemented as a soft-sensor system to predict the dielectric condition of transformer oil without requiring direct destructive testing, serving as a crucial component of modern Transformer Health Index assessment systems. Overall, these results indicate that DFFNN-GWO is the most efficient and reliable method for predicting the - 71 - Predicting Breakdown Voltage of Transformer Oil under Copper/Iron Contamination (G. D Prenata, et al. calculations,Ay Crystals, vol. 11, no. 2, 2021. breakdown voltage of transformer oil contaminated with metallic particles. The model not only excels in optimization stability but also demonstrates strong generalization capability for nonlinear physical phenomena and limited experimental datasets. From a practical perspective, the proposed model can be implemented as a soft-sensor system for transformer condition monitoring. By integrating measurable input parameters such as temperature, moisture content, and metallic contamination levels obtained from online sensors, the trained DFFNN model can estimate BDV in real time without requiring destructive laboratory testing. This approach complements existing diagnostic techniques such as Dissolved Gas Analysis (DGA), which focuses on fault detection at the system level. In contrast, the proposed BDV prediction model provides materiallevel insight into insulation degradation, enabling early detection of dielectric weakening and supporting predictive maintenance strategies in modern smart grid Cheng et al. Auand Breakdown Characteristics of Natural Ester and,Ay 2020. Negara. Fahmi. Asfani. Satriyadi Hernanda. Wahyudi, and M. Ibrahim. AuEffect of floating metallic particles in pre-breakdown and breakdown characteristics of oil transformer under dc voltage,Ay Energies, vol. 14, no. 12, 2021. Wang and M. Jin. AuConvolutional Domain Adaptation Network for Fault Diagnosis of Thermal System under Different Loading Conditions,Ay Chinese Control Conf. c, vol. 2020AeJuly, no. 3, pp. 4193Ae4197, 2020. Hussain. Siddique. Malik. Haider, and Hussain. AuMachine Learning-Based Prediction of Breakdown Voltage in High-Voltage Transmission Lines Under Ambient Conditions,Ay Eng, vol. 7, no. 1, p. 36, 2026. Kamyab. Daealhaq. Ghahfarokhi. Beheshtinejad, and E. Salajegheh. AuCombination of Genetic Algorithm and Neural Network to Select Facial Features in Face Recognition Technique,Ay Int. Robot. Control Syst. , vol. 3, no. 1, pp. 50Ae58, 2023. Taha. Saafan, and S. Ayyad. AuRevisiting natural selection: evolving dynamic neural networks using genetic algorithms for complex control tasks,Ay Artif. Intell. Rev. , vol. 58, no. 11, 2025. Bi et al. AuOptimizing a Multi-Layer Perceptron Based on an Improved Gray Wolf Algorithm to Identify Plant Diseases,Ay Mathematics, vol. 11, no. 15, pp. 1Ae36, 2023. IV. CONCLUSION This study presented a comparative evaluation of gradient-based and metaheuristic training strategies for predicting the Breakdown Voltage (BDV) of transformer oil contaminated with copper and iron particles, addressing the challenge of modeling nonlinear dielectric behavior with limited experimental data. Experimental results demonstrate that the DFFNN-Pure model achieves the highest numerical accuracy under smooth error conditions. In contrast, the DFFNN-GWO model provides the best balance between prediction accuracy, convergence stability, and robustness against nonlinear contamination In contrast, the DFFNN-GA model exhibits unstable convergence and poor generalization, making it unsuitable for BDV prediction using small datasets. Overall, the findings confirm that metaheuristic-based training particularly the Grey Wolf Optimizer effectively solves the BDV prediction problem by enhancing learning stability and reliability, thereby supporting non-destructive condition monitoring and predictive maintenance of power Despite the promising results, this study has several limitations. The dataset used consists of only 36 samples, which may limit the modelsAo generalization In addition, the evaluation is conducted under controlled laboratory conditions and may not fully represent real-world transformer operating environments. Future work should focus on expanding the dataset, incorporating additional diagnostic parameters such as dissolved gas analysis (DGA), and validating the model using real-time field data to enhance its practical . Azis, . Budy Santoso, and J. Jeffry. AuPenerapan Grey Wolf Optimizer Dalam Pelatihan Multi Layer Perceptron Untuk Menangani Masalah Klasifikasi Dan Regresi,Ay Adv. Comput. Syst. Innov. , vol. 2, no. 3, pp. 108Ae118, 2025. Tiwari. Agrawal. Belkhode. Ruatpuia, and S. Rokhum. AuHazardous effects of waste transformer oil and its prevention: A review,Ay Next Sustain. , vol. 3, no. January, p. 100026, 2024. Cinar and N. Natarajan. AuAn artificial neural network optimized by grey wolf optimizer for prediction of hourly wind speed in Tamil Nadu. India,Ay Intell. Syst. with Appl. 16, no. October, p. 200138, 2022. Hou. Gao. Wang, and C. Du. AuImproved Grey Wolf Optimization Algorithm and Application,Ay Sensors, 22, no. 10, pp. 1Ae19, 2022. Lau. Mohamad, and A. Suleiman. AuBreakdown characteristics of transformer oil with different contaminants,Ay Ie Transactions on Dielectrics and Electrical Insulation, vol. 22, no. 5, pp. 2766Ae2773, . Duval. AuDissolved gas analysis: It can save your transformer,Ay Ie Electrical Insulation Magazine, vol. 6, pp. 22Ae27, 1989. Gubanski et al. AuModern diagnostics of transformer insulation,Ay CIGRE Technical Brochure. Paris. France. REFERENCES