Reinforcement Learning for Dynamic Scheduling in Intelligent Manufacturing Systems: Theory, Methods, and Engineering Practice

Authors

  • Yuxin Dong School of Materials Science and Engineering, Nanjing Institute of Technology, China

Keywords:

Reinforcement Learning, Dynamic Scheduling, Job Shop Scheduling, Cyber-Physical Systems, Smart Manufacturing, Deep Q-Network, Production Control, Markov Decision Process

Abstract

Dynamic scheduling-constantly changing which machines and resources do what work, unexpected arrival, constantly changing working hours and equipment failures-has always been too difficult to solve with perfect mathematical formulas. Therefore, the manufacturing system uses simple and fixed rules, but when the production environment changes greatly, these rules cannot work normally. Reinforcement learning (RL) regards scheduling as a series of decisions, in which agents learn strategies (state-to-action strategies) by interacting with the production environment, which has become the main AI method to solve this problem. This paper reviews the dynamic scheduling and RL in intelligent manufacturing system, the theory of network physical system (CPS), system engineering and information theory. Scheduling strategy is regarded as a compressed version of production status information, which is enough to make nearly perfect decisions. We follow the theoretical basis of RL-based scheduling, including markov decision processes, state-action representation and reward design. We also analyzed four transformation paths for RL to change the manufacturing industry: from standardized production to customized production, from reactive scheduling to active scheduling, from manual scheduling to autonomous scheduling, and from independent machines to integrated network physical scheduling ecosystem. The specific methods we see include deep Q-network agent, graph-neural-network strategy learning for job shop problems, compound reward shape and explainable RL substitution for production control. We also studied the application of fault detection in planning, optimization of quality-related processes and man-machine allocation, challenging issues, such as sample efficiency, generalization of different problem sizes, and security of deployment. We finally come to a big framework, which says that reward design is the main interface between manufacturing goals and policy learning behavior.

Downloads

Published

2025-10-01

How to Cite

Dong , Y. (2025). Reinforcement Learning for Dynamic Scheduling in Intelligent Manufacturing Systems: Theory, Methods, and Engineering Practice. CPS Digital Library - Series of Conferences, 4(2), 1–5. Retrieved from https://seriesofconference.com/index.php/SCJ/article/view/177