Adaptive hybrid ensemble-based DDoS detection using reinforcement learning-guided optimization and deep learning

(1) Maha I. Raheem Mail (College of Engineering, University of Information Technology and Communication, Al-Mansour, 3071, Baghdad, Iraq)
(2) * Shouket A. Ahmed Mail (Department of Medical Instrumentation Techniques Engineering, Technical Engineering College, Al-Kitab University, Altun Kupri, Kirkuk, 36001, Iraq)
(3) Enas F. Aziz Mail (Department of Cybersecurity Engineering Technologies, Technical Engineering College for Computer and Artificial Intelligence/Kirkuk, Northern Technical University, 36001, Kirkuk, Iraq)
(4) Saad A. Assi Mail (Software Department, College of Computer Science and Information Technology, University of Kirkuk, Kirkuk, Iraq)
(5) Sinan Q. Salih Mail (Technical College of Engineering, Al-Bayan University, Baghdad 10011, Iraq)
(6) Ahmed Dheyaa Radhi Mail (College of Pharmacy, University of Al-Ameed, Karbala PO Box 198, Iraq)
(7) Hilal A. Fadhil Mail (Department of Electrical and Computer Engineering, Sohar University, Sohar, Oman)
(8) Taha Almulaisi Mail (Renewable Energy Research Unit, Polytechnic College Hawija, Northern Technical University, Hawija, 36007, Iraq)
*corresponding author

Abstract


Distributed Denial-of-Service (DDoS) attacks remain among the most disruptive network threats, and detectors that generalize across attack families with low false-alarm rates are still an open problem. Propose an adaptive hybrid ensemble that unifies two gradient-boosting learners (Random Forest and Gradient Boosting) with three deep neural base learners (DNN, CNN-1D, and LSTM) under a weighted soft-voting rule whose weights are produced by a Reinforcement Learning (RL) policy. The RL agent treats the ensemble-weight simplex as its action space, observes a state vector built from validation-set diagnostic statistics, and is trained by REINFORCE-with-baseline to maximize a reward equal to validation F1 minus a small calibration penalty. The framework is formalized as a Markov decision process with one stochastic step per training episode, which decouples ensemble-weight learning from the (non-differentiable) outer F1 objective. On a 10,000-sample, 25-feature, five-class benchmark with 7% label noise, the proposed system reaches weighted F1 = 0.846, accuracy = 84.7%, MCC = 0.781, AUC = 0.952, and ECE = 0.039. Friedman and Nemenyi post-hoc tests over 50 CV folds confirm the RL-guided ensemble is significantly better than every individual base learner and uniform voting at α = 0.05 (Cohen's d = 0.96). An ablation isolates the RL policy and gradient boosting as the main drivers; a label-noise robustness study shows graceful degradation up to 20%; a head-to-head comparison against the Bonobo Optimizer (BO), GA, PSO, GWO, and WOA shows the best F1/wallclock trade-off.

Keywords


DDoS detection; reinforcement learning; policy gradient; adaptive ensemble; deep learning; network intrusion detection

   

DOI

https://doi.org/10.26555/ijain.v12i3.2109
      

Article metrics

Abstract views : 17

   

Cite

   


Creative Commons License
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.

___________________________________________________________
International Journal of Advances in Intelligent Informatics
ISSN 2442-6571  (print) | 2548-3161 (online)
Organized by UAD and ASCEE Computer Society
Published by Universitas Ahmad Dahlan
W: http://ijain.org
E: info@ijain.org (paper handling issues)
 andri.pranolo.id@ieee.org (publication issues)

View IJAIN Stats

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0