The problem
Recovering a quadrotor from extreme states such as high velocities, high body rates, or a fully inverted attitude is where classical controllers struggle. Learned policies can handle these states, but they come with no guarantee that the drone stays out of unsafe regions.
Approach
- Learning: PPO (Stable-Baselines3) in an extended Flightmare simulator with Agilicious dynamics, using action-history observations.
- Reward and curriculum: a dual-scale Cauchy reward combined with goal-directed velocity objectives, plus curriculum learning, so the reward stays informative across the whole flight envelope.
- Robustness: domain randomization of mass, motor time constant and command delay, plus calibrated disturbances (wind, drag, sensor noise) for sim-to-real transfer.
- Safety: a minimum-intervention CBF quadratic program, built from Lie derivatives of the nonlinear quadrotor dynamics, that corrects the RL action at each step. It is used only at deployment, so training is unchanged, and it runs in under 1 ms per step.
- Baselines: a nonlinear MPC and a soft-constrained variant that share the same dynamics model and solver, for a fair comparison.
Results
- Recovers from up to ±8 m/s, ±10 rad/s, and upside-down starts, where the NMPC baseline fails.
- Stays reliable when the real mass or motor dynamics differ from the nominal model.
- With the CBF filter, the drone stays inside the safe region, where plain RL and NMPC cross the boundary.
- Deployed on a real Agilicious quadrotor, confirming real-time performance and sim-to-real transfer.
Publication
Safe Reinforcement Learning with Control Barrier Functions for Quadrotor Recovery.
O. Eşen, Y. Yuan, M. Ryll. Technical University of Munich. Submitted to an international controls conference.