Resolving Reward Misspecification in Deep Reinforcement Learning Agents Deployed Across Non-Stationary Industrial Control Environments
Keywords:
deep reinforcement learning, reward misspecification, non-stationary environments, meta-learning reward correction, distributional reinforcement learning, variational autoencoder latent embedding, industrial process control, covariate shift robustness, uncertainty-aware policy optimizationAbstract
Reward misspecification remains a critical failure mode for deep reinforcement learning (DRL) agents transitioning from simulated training pipelines to non-stationary industrial control environments, where dynamic process shifts invalidate handcrafted reward shaping assumptions. This paper proposes an Adaptive Reward Correction Network (ARCN), a meta-learning augmented framework that continuously re-calibrates reward signals via latent environment embeddings derived from variational autoencoders. ARCN integrates a distributional critic architecture with uncertainty-aware policy gradient updates, enabling robust agent behavior under covariate shift and partial observability. Empirical evaluations conducted on three benchmark industrial control testbeds—chemical reactor regulation, autonomous grid load balancing, and multi-axis robotic assembly—demonstrate a 31.4% reduction in cumulative constraint violations and a 22.7% improvement in long-horizon task completion relative to state-of-the-art baselines. Ablation studies confirm the necessity of each architectural component.
References
Semeniuk, V. V. (2025). OPTIMIZATION OF LOCAL DEVELOPMENT PROCESS USING DOCKER PHP IMAGE THAT COMES WITH A FULL SET OF TOOLS OUT OF THE BOX–PERFORMANCE AND OPTIMIZATION EXTENSIONS. ІНФОРМАЦІЙНЕ ЗАБЕЗПЕЧЕННЯ БАГАТОІНДЕКСНОЇ ТРАНСПОРТНОЇ ЗАДАЧІ З НЕЧІТКИМИ ІНТЕРВАЛАМИ.
Semeniuk, V. V. (2025). OPTIMIZATION OF LOCAL DEVELOPMENT PROCESS USING DOCKER PHP IMAGE THAT COMES WITH A FULL SET OF TOOLS OUT OF THE BOX: DATABASE AND INTERNATIONALIZATION EXTENSIONS. ВЧЕНІ ЗАПИСКИ, 12025226.