This thesis investigates the Flexible Flow Shop Scheduling Problem considering heterogeneous, parallel machines and sequence-dependent setup times to minimize the total tardiness objective in deterministic and stochastic problem settings. The problem is modeled as a Markov Decision Process and Deep Reinforcement Learning is applied to train a Neural Dispatching Rule. This thesis proposes Policy ShapedProximalPolicyOptimization (PS-PPO), whichextendsthewell-established Proximal Policy Optimization (PPO) algorithm to leverage the knowledge of exist ing guide policies. By applying policy shaping to change the sampling of actions during experience collection, existing policies can guide the exploration of PPO efficiently. By experiencing chains of good actions sooner during training, the training of PPO can be improved both, in terms of speed of convergence and final quality in most investigated instances. PS-PPO is not only able to leverage existing DRL policies as a guide policy but also a priority dispatching rule.
«
This thesis investigates the Flexible Flow Shop Scheduling Problem considering heterogeneous, parallel machines and sequence-dependent setup times to minimize the total tardiness objective in deterministic and stochastic problem settings. The problem is modeled as a Markov Decision Process and Deep Reinforcement Learning is applied to train a Neural Dispatching Rule. This thesis proposes Policy ShapedProximalPolicyOptimization (PS-PPO), whichextendsthewell-established Proximal Policy Optimizatio...
»