A small neural network was trained with PPO using Stable-Baselines3 to compare aircraft fault recovery against classical controllers. Simulated actuator faults tested how each controller responded, with recorded replays showing the network’s activity alongside the aircraft’s motion.
The policy has 15 inputs, two hidden layers of 128 neurons and four outputs. Line thickness represents the learned connection weights. To keep the connections readable, it shows the eight strongest incoming weights per neuron; line thickness is scaled within each layer.