Abstract
This paper addresses the challenge of achieving stable and constraint-satisfying control in high-order nonlinear systems. While terminal sliding mode control ensures the finite-time convergence, it often suffers from chattering and input saturation in practical implementations. In order to improve this, we propose a reinforcement learning framework based on constrained policy optimization to replace the discontinuous switching term with a learned policy, preserving sliding convergence while ensuring smooth and bounded control inputs. In addition, a theoretical input bound for high-order systems is derived and incorporated into the learning process as a safety constraint. Simulations validate the approach, showing reduced control effort and improved convergence within state limits across various initial conditions. This work provides a unified framework for safe and adaptive control, with potential applications in robotics and autonomous systems.
Keywords
Get full access to this article
View all access options for this article.
