ARTICLE

A Generalized Reinforcement Learning Controller for Metaheuristic Hyperparameter Control

Victor-Deian Balutoiu, Roxana Teodora Mafteiu-Scai, Liviu Octavian Mafteiu-Scai


© 2026 Liviu Octavian Mafteiu-Scai, published by UIKTEN. This work is licensed under the Creative Commons Attribution-NonCommercial 4.0 International. (CC BY-NC 4.0).

Citation Information: SAR Journal. Volume 9, Issue 2, Pages 89-99, ISSN 2619-9955, https://doi.org/10.18421/SAR92-01, June 2026.

Received: 29 April 2026.
Revised: 03 June 2026.
Accepted: 10 June 2026
Published: 27 June 2026.

Abstract:

It is well known that the performance of population-based metaheuristics depends strongly on how their hyperparameters are tuned. Among the various techniques available, one of the most prominent is Reinforcement Learning (RL). In general, existing RL approaches are specialized and tightly coupled with a single optimization algorithm. This paper proposes a generalized hyperparameter control method for metaheuristics, hereinafter referred to as PPO-based RL agent. The agent operates through an abstraction layer that maps generic control actions to algorithm-specific parameters. The model was trained exclusively on Differential Evolution (DE) using a set of six benchmark functions. It was evaluated on both the problems used for training and unseen problems, over 30 independent runs each. The experiments showed that the RL agent learned a generalized tuning policy that consistently outperformed static parameters on deceptive landscapes (such as the Schwefel and Rastrigin functions) and achieved comparable results to LSHADE, a specialized, state-of-the-art variant of DE. Another objective of this work was to evaluate the zero-shot transfer potential to other metaheuristics, such as Particle Swarm Optimization (PSO) and Genetic Algorithms (GA).


Keywords – reinforcing learning, metaheuristics, hyperparameters, tuning.

                   

                                                                      Full text PDF