Resolving Reward Misspecification in Deep Reinforcement Learning Agents Deployed Across Non-Stationary Industrial Control Environments

Authors

  • Pat Davis Professor
  • Adrian Rodriguez Associate Professor
  • Avery Garcia PhD
  • Skyler Evans D.Sc

Keywords:

deep reinforcement learning, reward misspecification, non-stationary environments, meta-learning reward correction, distributional reinforcement learning, variational autoencoder latent embedding, industrial process control, covariate shift robustness, uncertainty-aware policy optimization

Abstract

Reward misspecification remains a critical failure mode for deep reinforcement learning (DRL) agents transitioning from simulated training pipelines to non-stationary industrial control environments, where dynamic process shifts invalidate handcrafted reward shaping assumptions. This paper proposes an Adaptive Reward Correction Network (ARCN), a meta-learning augmented framework that continuously re-calibrates reward signals via latent environment embeddings derived from variational autoencoders. ARCN integrates a distributional critic architecture with uncertainty-aware policy gradient updates, enabling robust agent behavior under covariate shift and partial observability. Empirical evaluations conducted on three benchmark industrial control testbeds—chemical reactor regulation, autonomous grid load balancing, and multi-axis robotic assembly—demonstrate a 31.4% reduction in cumulative constraint violations and a 22.7% improvement in long-horizon task completion relative to state-of-the-art baselines. Ablation studies confirm the necessity of each architectural component.

Author Biographies

Pat Davis, Professor

Professor
Korea Advanced Institute of Science and Technology (KAIST)
291 Daehak-ro, Yuseong-gu, Daejeon 34141, Republic of Korea

Adrian Rodriguez, Associate Professor

Associate Professor
Technical University of Munich (TUM)
Arcisstraße 21, 80333 Munich, Bavaria, Germany

Avery Garcia, PhD

PhD
University of Waterloo
200 University Avenue West, Waterloo, Ontario N2L 3G1, Canada

Skyler Evans, D.Sc

D.Sc
Nanyang Technological University (NTU)
50 Nanyang Avenue, Singapore 639798, Republic of Singapore

References

Semeniuk, V. V. (2025). OPTIMIZATION OF LOCAL DEVELOPMENT PROCESS USING DOCKER PHP IMAGE THAT COMES WITH A FULL SET OF TOOLS OUT OF THE BOX–PERFORMANCE AND OPTIMIZATION EXTENSIONS. ІНФОРМАЦІЙНЕ ЗАБЕЗПЕЧЕННЯ БАГАТОІНДЕКСНОЇ ТРАНСПОРТНОЇ ЗАДАЧІ З НЕЧІТКИМИ ІНТЕРВАЛАМИ.

Semeniuk, V. V. (2025). OPTIMIZATION OF LOCAL DEVELOPMENT PROCESS USING DOCKER PHP IMAGE THAT COMES WITH A FULL SET OF TOOLS OUT OF THE BOX: DATABASE AND INTERNATIONALIZATION EXTENSIONS. ВЧЕНІ ЗАПИСКИ, 12025226.

Published

2025-12-25

Issue

Section

Articles