Deep Reinforcement Learning has shown great potential to solve complex sequential
decision-making problems and has been applied successfully to different intractable in-
ventory control problems. Therefore, this paper models a multi-product, multi-echelon
inventory control problem as a Markov Decision Process and applies Proximal Pol-
icy Optimization to solve it. We consider a realistic supply chain network of general
structure provided in the scope of the ASML inventory challenge. Different application
approaches of PPO are evaluated regarding their capability to handle the large com-
binatorial complexity inherent to the problem. We propose a continuous action space
to improve scalability and a Combinatorial Optimization Layer to transform continu-
ous actions into feasible, discrete ones. Further, Graph Neural Networks are leveraged
to extract structured feature embeddings and to predict continuous actions in a DRL
framework using a supervised pre-training approach. Our results show, that DRL and
GNN’s have the potential to match a benchmark heuristic, but have problems when
being faced with many inter-dependencies of components.
«
Deep Reinforcement Learning has shown great potential to solve complex sequential
decision-making problems and has been applied successfully to different intractable in-
ventory control problems. Therefore, this paper models a multi-product, multi-echelon
inventory control problem as a Markov Decision Process and applies Proximal Pol-
icy Optimization to solve it. We consider a realistic supply chain network of general
structure provided in the scope of the ASML inventory challenge. Differe...
»