1. Home
  2. Projects
  3. Safe RL for Quadrotor Recovery
Semester thesis · Autonomous Aerial Systems Lab, TUM · 2025–2026

Safe Reinforcement Learning for Quadrotor Recovery

I trained a reinforcement-learning policy that flies a quadrotor back to a goal from almost any starting state, including upside down, and added a Control Barrier Function safety filter that keeps it inside a safe region at deployment. The policy was validated on a real Agilicious quadrotor.

  • PPO
  • Control Barrier Functions
  • Sim-to-Real
  • NMPC
  • Stable-Baselines3

The problem

Recovering a quadrotor from extreme states such as high velocities, high body rates, or a fully inverted attitude is where classical controllers struggle. Learned policies can handle these states, but they come with no guarantee that the drone stays out of unsafe regions.

Approach

  • Learning: PPO (Stable-Baselines3) in an extended Flightmare simulator with Agilicious dynamics, using action-history observations.
  • Reward and curriculum: a dual-scale Cauchy reward combined with goal-directed velocity objectives, plus curriculum learning, so the reward stays informative across the whole flight envelope.
  • Robustness: domain randomization of mass, motor time constant and command delay, plus calibrated disturbances (wind, drag, sensor noise) for sim-to-real transfer.
  • Safety: a minimum-intervention CBF quadratic program, built from Lie derivatives of the nonlinear quadrotor dynamics, that corrects the RL action at each step. It is used only at deployment, so training is unchanged, and it runs in under 1 ms per step.
  • Baselines: a nonlinear MPC and a soft-constrained variant that share the same dynamics model and solver, for a fair comparison.

Results

  • Recovers from up to ±8 m/s, ±10 rad/s, and upside-down starts, where the NMPC baseline fails.
  • Stays reliable when the real mass or motor dynamics differ from the nominal model.
  • With the CBF filter, the drone stays inside the safe region, where plain RL and NMPC cross the boundary.
  • Deployed on a real Agilicious quadrotor, confirming real-time performance and sim-to-real transfer.
Quadrotor recovering from a fully inverted start, shown for the RL policy alone and with the CBF safety filter
Recovery from a fully inverted start: (a) RL policy, (b) RL policy with the CBF filter.
Altitude trajectories showing the CBF keeping the quadrotor between ceiling and ground limits
Vertical safety constraints with the goal placed outside the safe region. Plain RL and NMPC cross the boundary; RL + CBF stays inside it.

Publication

Safe Reinforcement Learning with Control Barrier Functions for Quadrotor Recovery.
O. Eşen, Y. Yuan, M. Ryll. Technical University of Munich. Submitted to an international controls conference.
More work

Other projects

UAV Cave Exploration

ROS 2 quadrotor that autonomously explores and maps an unknown cave with OctoMap, frontier-based exploration, sampling-based planning and a geometric SE(3) controller.

Read more