Augmentasi Moderat untuk Peningkatan Kemampuan Deteksi Anomali Serangan pada Dataset UNSW-NB15
Abstract
Class imbalance remains a persistent challenge in intrusion detection because infrequent attacks may carry substantial security impact while being poorly represented during model training. This study evaluates a moderate augmentation strategy for UNSW-NB15 by selectively limiting synthetic data growth instead of forcing every class to match the majority class. Four extremely underrepresented classesAnalysis, Backdoor, Shellcode, and Wormswere augmented using the target T_k = min(10,000, 15 x n_k). Network records were first mapped into a latent representation using a Stacked Denoising Autoencoder (SDAE), followed by five generative scenarios: Hybrid SDAE-GAN, SDAE Feature Extraction, SDAE Dimensionality Reduction, SDAE-WGAN-GP, and SDAE-CTGAN. No-augmentation and SMOTE baselines were included for comparison. S5-CTGAN achieved the highest macro F1-score of 0.4135, a 5.9% improvement over the no-augmentation baseline (0.3906), with an MCC of 0.5897. The most visible class-level gains were obtained for Analysis (F1 from 0.018 to 0.105) and DoS (0.206 to 0.375). S2-Feature Extraction produced the closest synthetic distribution with a mean Wasserstein Distance of 0.0172. A one-way ANOVA confirmed highly significant differences among experimental setups (p = 4.37 x 10^-41; eta-squared = 0.957), supported by Friedman and Tukey HSD tests. The results indicate that generating more synthetic samples is not inherently beneficial: aggressive full balancing can cause over-amplification, whereas controlled latent-space augmentation provides more stable gains while preserving the role of genuine observations.
Keywords
Full Text:
PDFReferences
N. Moustafa and J. Slay, UNSW-NB15: A comprehensive data set for network intrusion detection systems (UNSW-NB15 network data set), in 2015 Military Communications and Information Systems Conference (MilCIS), 2015, pp. 1-6, doi: 10.1109/MilCIS.2015.7348942.
N. Moustafa and J. Slay, The evaluation of Network Anomaly Detection Systems: Statistical analysis of the UNSW-NB15 data set and the comparison with the KDD99 data set, Information Security Journal: A Global Perspective, vol. 25, no. 1-3, pp. 18-31, 2016, doi: 10.1080/19393555.2015.1125974.
V. Shanmugam, R. Razavi-Far, and E. Hallaji, Addressing Class Imbalance in Intrusion Detection: A Comprehensive Evaluation of Machine Learning Approaches, Electronics, vol. 14, no. 1, art. 69, 2025, doi: 10.3390/electronics14010069.
C. Wheelus, E. Bou-Harb, and X. Zhu, Tackling Class Imbalance in Cyber Security Datasets, in 2018 IEEE International Conference on Information Reuse and Integration (IRI), 2018, pp. 229-232, doi: 10.1109/IRI.2018.00041.
N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, SMOTE: Synthetic Minority Over-sampling Technique, Journal of Artificial Intelligence Research, vol. 16, pp. 321-357, 2002, doi: 10.1613/jair.953.
P. Soltanzadeh and M. Hashemzadeh, RCSMOTE: Range-Controlled synthetic minority over-sampling technique for handling the class imbalance problem, Information Sciences, vol. 542, pp. 92-111, 2021, doi: 10.1016/j.ins.2020.07.014.
I. Goodfellow et al., Generative Adversarial Nets, in Advances in Neural Information Processing Systems 27, 2014, pp. 2672-2680.
P. Vincent, H. Larochelle, I. Lajoie, Y. Bengio, and P.-A. Manzagol, Stacked Denoising Autoencoders: Learning Useful Representations in a Deep Network with a Local Denoising Criterion, Journal of Machine Learning Research, vol. 11, pp. 3371-3408, 2010.
J. H. Lee and K. H. Park, AE-CGAN Model based High Performance Network Intrusion Detection System, Applied Sciences, vol. 9, no. 20, art. 4221, 2019, doi: 10.3390/app9204221.
I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin, and A. C. Courville, Improved Training of Wasserstein GANs, in Advances in Neural Information Processing Systems 30, 2017.
L. Xu, M. Skoularidou, A. Cuesta-Infante, and K. Veeramachaneni, Modeling Tabular Data using Conditional GAN, in Advances in Neural Information Processing Systems 32, 2019.
D. Li, D. Kotani, and Y. Okabe, Improving Attack Detection Performance in NIDS Using GAN, in 2020 IEEE 44th Annual Computers, Software, and Applications Conference (COMPSAC), 2020, pp. 817-825, doi: 10.1109/COMPSAC48688.2020.0-162.
C. Park, J. Lee, Y. Kim, J.-G. Park, H. Kim, and D. Hong, An Enhanced AI-Based Network Intrusion Detection System Using Generative Adversarial Networks, IEEE Internet of Things Journal, vol. 10, no. 3, pp. 2330-2345, 2023, doi: 10.1109/JIOT.2022.3211346.
H. Ding, Y. Sun, N. Huang, Z. Shen, and X. Cui, TMG-GAN: Generative Adversarial Networks-Based Imbalanced Learning for Network Intrusion Detection, IEEE Transactions on Information Forensics and Security, vol. 19, pp. 1156-1167, 2024, doi: 10.1109/TIFS.2023.3331240.
X. Zhao, K. W. Fok, and V. L. L. Thing, Enhancing network intrusion detection performance using generative adversarial networks, Computers & Security, art. 104005, 2024, doi: 10.1016/j.cose.2024.104005.
J. Cui et al., A novel multi-module integrated intrusion detection system for high-dimensional imbalanced data, Applied Intelligence, vol. 53, pp. 272-288, 2023, doi: 10.1007/s10489-022-03361-2.
S. Bourou, A. El Saer, T.-H. Velivassaki, A. Voulkidis, and T. Zahariadis, A Review of Tabular Data Synthesis Using GANs on an IDS Dataset, Information, vol. 12, no. 9, art. 375, 2021, doi: 10.3390/info12090375.
W. Tian, Y. Shen, N. Guo, J. Yuan, and Y. Yang, VAE-WACGAN: An Improved Data Augmentation Method Based on VAEGAN for Intrusion Detection, Sensors, vol. 24, art. 6035, 2024, doi: 10.3390/s24186035.
DOI: https://doi.org/10.5281/zenodo.22211310
Refbacks
- There are currently no refbacks.
Copyright (c) 2026 Sebastianus Hartantyo, Alva Hendi Muhammad

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
ISSN : 3025-6704





