Case file · RipWing Sim + architecture complete
FAULT-TOLERANT QUADCOPTER FLIGHT CONTROLLER  //  RUST · STM32 · RTIC

Project RipWing

A flight controller I built from scratch. It holds a hard real-time attitude loop while an edge ML monitor watches for sensor degradation and motor faults, but never gets to fly the craft itself. One constraint shaped everything: my hardware feedback loop is long, so I designed and proved almost all of it in simulation before a board existed.

Role
Controls · firmware · simulation
Collaborator
Remote · all electrical + PCB
Target
STM32F446RE dev · H743 flight
Stack
Rust · no_std · RTIC · embedded-hal
OBJECTIVES
What it has to do
Five aims. On a project like this, the reasoning is the deliverable.
01 Fly a stable quad on a loop that provably meets its deadlines
02 Design control laws simulation-first, validated before hardware
03 Detect faults in flight with an ML monitor that observes but never actuates
04 Enforce the host-testable / target-only boundary in the build system
05 Produce a flight recorder that survives the crash it recorded
Built with · Rust no_std RTIC embedded-hal cortex-m defmt flip-link TF Lite Micro · planned
00

The premise

A remote collaborator handles all of the electrical and PCB work, which makes the hardware feedback loop long. A change I can prototype in software in minutes takes days to see on a board.

So I pushed as much logic as I could into a half that runs on a laptop, and made the layering strict enough that the compiler keeps it honest instead of me. Nearly every choice below is the same trade in a new place: give up some peak performance for a guarantee I can prove.

The core bet
Build the aircraft in simulation before you build the controller.

A controller tuned against a model that lies to you is worse than no controller at all. So the simulator isn't a toy. It's where the control logic actually gets designed, and the same pure-math crates it runs are the ones that compile into the firmware. The line between "logic I can test on my laptop" and "code that needs the chip" is real, enforced, and wide.

01

The plant

Before any controller, I need the physics it acts on. The first model is the standard rigid-body quadcopter: twelve states, four rotor speeds as instant inputs, checked by a 24-test suite built bottom-up from analytical facts.

Four rotors, six degrees of freedom, so the craft is underactuated: every rotation is bought with a pattern of thrust. Differential thrust on a pair makes roll and pitch; the drag imbalance between the CW and CCW pairs makes yaw. This is what the tuner in the next station steers. Switch views to isolate each axis.

// BODY-FRAME FORCE & TORQUE MAP · X CONFIG ΣTᵢ = T_B
FRONTFL1.00×CCWFR1.00×CWRL1.00×CWRR1.00×CCW

All four rotors spin equal. Thrust sums straight up the body +z axis — this is hover / vertical throttle, the baseline every other axis perturbs from.

FIG.01 — Body-frame force & torque map (after Gibiansky). Thrust Tᵢ = k·ωᵢ² sums to T_B along +z. The model is honest about rigid-body dynamics, but it hides two things that matter enormously for tuning. That is the whole subject of the next station.

02

The fidelity ladder

Finding out where a model lies is worth more than trusting where it tells the truth. So the simulator grew in deliberate layers, and the gap between each one is itself the lesson.

The 6-DoF model treats rotor speed as an instant input. The second one promotes the four rotor speeds to real states with first-order lag, sixteen states total, and adds physics the base model structurally can't: motor lag, blade flapping, rotor gyroscopics, rate damping.

EQ.01 · Motor as a first-order lag
ω̇ = (ωcmd − ω) / τm // τ_m ≈ 50 ms · forces use the actual rotor speed, not the commanded one

Blade flapping is the subtle one. A rotor moving through the air makes more lift on its advancing blade than its retreating one, tilting the rotor away from the direction of motion. That gives a horizontal drag force (the dominant velocity damping on a real multirotor) plus a pitch-up moment in forward flight. It settles within one rotor revolution (~4 ms) against ~100 ms body dynamics, so I model it quasi-statically, the way the control literature does. The parameters are representative values from the literature, not measured off a specific airframe. I say so plainly, because passing borrowed numbers off as measured ones is the kind of overclaim that discredits a whole simulation.

▸ COMPARISON STUDY · what the harder model buys SAME EXPERIMENTS · BOTH PLANTS
Experiment 6-DoF 10-DOF What it means
Velocity decay from 4 m/s never decays t½ = 2.33 s Flapping is the real damping; the 6-DoF drone coasts forever
Roll step, cmd vs actual rpm identical 1659 rpm gap Motor lag is a real actuator, not a wire
Rate-loop gain sweep ×1→×48 ripple flat 5.36 °/s buzz The motor pole caps usable gain; 6-DoF rewards infinite gain
Steady 1.5 m/s cruise 0.76° pitch 1.61° pitch ~2× tilt needed to overcome flapping drag

The third row is the one that matters for tuning. Gains tuned on the 6-DoF model are wrong in exactly the direction tuning pushes them. Higher gain always looks free on the simpler model, because it has no actuator pole to ring against. A test confirms it: with flapping on, the hover mode is open-loop unstable (eigenvalues at +0.59 ± 1.34j), the classic unstable hover mode every real rotorcraft has and the 6-DoF model can't show. The two suites total 41 passing tests, and with the new coefficients zeroed the 16-state model reproduces the 6-DoF one to within 1e-9, so the upgrade is a strict superset.

The physics that owns the landing

Ground effect

Within about one rotor radius of the ground, downwash can't escape freely, so the same rotor speed makes more thrust. A quad settling onto a pad feels a rising thrust surplus right as it tries to land. If the controller doesn't know about it, the craft floats and won't touch down cleanly.

The trap

Cheeseman–Bennett has a singularity

The famous textbook relation blows up to infinity as height nears a quarter of the rotor radius, exactly the range a landing controller lives in. Feed a model that diverges into a simulator you mean to land, and you get a numerical explosion at the worst possible moment.

The choice

A bounded exponential instead

Stays finite all the way to the ground, a sensible ceiling at contact instead of divergence, while matching the same falloff higher up. Picking the well-behaved model over the famous one only pays off once you actually try to simulate the hard case.

A four-experiment study turns landing from a surprise into a modeled, testable scenario. The thrust-surplus float becomes a disturbance I can design the altitude loop against, ground effect makes the craft mildly self-leveling near the pad (a quirk to exploit), and the hover-trim shift on approach is exactly the operating-point change the gain-scheduling map is there to absorb.

03

Automated PID tuning

Drag the craft, fire a disturbance, command a step. This live roll-axis loop runs the same rigid-body dynamics, J·θ̈ = τ, at the 400 Hz firmware rate. It also shows off the exact problem I refused to solve by hand.

// ATTITUDE LOOP · ROLL AXIS 400 Hz · J 0.02 kg·m²
drag to perturb
THETA RESPONSE → SETPOINT
0.0°
peak error
0.00s
settle
0%
saturation
UNDERDAMPED
ζ ≈ 0.56

FIG.02 — Live closed-loop sim. Watch the saturation readout: when the motor differential hits its ±0.6 N·m limit, an unbounded integral term keeps piling up error it can't act on. That is integral windup, and it's why the integral here, and in the firmware, is clamped.

Sliding three gains by feel until the trace looks right is how most quads get tuned. It's also unreproducible, only one noise seed deep, and slow to redo on every airframe change. So instead of hand-tuning, I built a harness: staged Nelder–Mead in log-gain space, searching against the noise-injected 10-DOF plant. Stage 1 tunes the inner attitude and rate loops on a 15° roll doublet; stage 2 freezes those and tunes the outer position loop on a 2 m step. The gains come out reproducible, with robustness measured across noise seeds. That's a stronger claim than "it flew on my bench."

Key decision
Optimize the gains. Don't guess them.

The interesting part wasn't the optimizer. It was finding the four ways a naive setup produces gains that are optimal for the wrong problem. Each is a spot where an honest model quietly rewards a controller that would be dangerous on real hardware.

01
Sensor noise is mandatory

My first noise-free run came back with a derivative gain near 400. With perfect measurements, the optimizer had found that derivative gain is free damping. Adding realistic noise (0.02 m, 0.4°, 1°/s) and a motor-thrash penalty makes it cost what it costs on hardware, and the tuned Kd shrinks instead of exploding.

02
Give the integral a job

With no persistent disturbance, the integral gain is unidentifiable. A small shifted-battery torque bias makes Ki earn its keep, and it stays interpretable: with high rate stiffness the bias gets rejected through a tiny attitude offset, so Ki correctly falls toward zero.

03
Cap yaw inside the envelope

Yaw authority is only ±0.26 N·m across the whole motor range. My original limit let the controller demand 4× the yaw torque that exists; the allocator "solved" that by slamming one motor pair to zero and the other to max, a positive-feedback yaw ratchet that noise reliably tripped. The fix clamps yaw demand to half the physical authority before allocation.

04
Filter the D-term

Numerically differentiating a noisy error at 200 Hz amplifies measurement noise by roughly 280×. Every real flight controller filters the derivative term, so the PID class grew a configurable low-pass and the tuning uses it.

▸ EMBED · CODE · control/src/pid.rs / step() LST.01 · Rust · no_std
1 // the two fixes the tuner made unavoidable: anti-windup + filtered D
2 let err = setpoint - theta;
3 self.integral += err * dt;
4 self.integral = self.integral.clamp(-I_CLAMP, I_CLAMP); // no windup
5 // low-pass the derivative · raw diff at 200 Hz amplifies noise ~280x
6 self.d_lp += self.alpha * ((theta - self.prev) / dt - self.d_lp);
7 self.prev = theta;
8 let tau = self.kp*err + self.ki*self.integral - self.kd*self.d_lp;
9 tau.clamp(-TAU_MAX, TAU_MAX) // actuator saturation is real
The clamp on line 4 is the windup guard you can feel in the tuner above; the harness picked alpha and the gains against the noisy plant.
▸ RESULTS · default vs tuned, on the noisy plant 15° ROLL DOUBLET
Metric Default Tuned
Doublet overshoot 58% 1.1%
Rate ripple 20.5 °/s 1.1 °/s
Rise time 0.28 s 0.33 s

Cross-seed validation on five unseen noise realizations improves the ITAE metric 7–8× every time, so the gains aren't overfit to one lucky draw.

04

Firmware architecture

A strict layered stack, built as a Rust workspace. Each layer is a crate that can only depend on the layers below it, and Cargo's dependency graph makes breaking that rule a compile error instead of a convention I have to remember.

LAYER 5
Application RTIC app module · tasks, flight modes, anomaly response
target-only
LAYER 4
Control pure no_std math · PID, estimator, mixer, setpoint
host-testable
LAYER 3
Drivers generic over embedded-hal traits · IMU, baro, ESC, SD
trait-generic
LAYER 2
Board pin map, clock tree, peripheral init for this board
target-only
LAYER 1
HAL / Vendor stm32-hal, embedded-hal, cortex-m, rtic
target-only

The key property: the control crate and the ML crate depend on nothing embedded. They compile and test on my laptop with cargo test. The driver crate depends only on the embedded-hal traits, never a specific chip's HAL, so the same IMU driver runs on any board whose HAL implements them. Only the firmware binary is allowed to name the STM32.

The payoff of the long feedback cycle

The Cargo workspace is the call graph, and Rust's dependency rules enforce it: the pure-math crate literally can't import a HAL, and Cargo forbids circular dependencies. That's a stronger guarantee than a diagram a reviewer has to police by hand. I develop and debug the whole control and ML path on my laptop, and CI fails the moment anyone tries to sneak a hardware dependency into a host-testable crate.

▸ EMBED · CODE · control/Cargo.toml LST.02 · TOML
1 [package]
2 name = "control"
3  
4 [dependencies]
5 nalgebra = { version = "0.33", default-features = false } # libm, no std
6 # NO stm32-hal, NO cortex-m · won't compile if added.
7 # `cargo test` runs this crate on the host, every push.
Data flow

Sensor data comes in at up to 1 kHz from a dual IMU over SPI, feeds the attitude estimator and cascaded PID loops, and leaves as DShot motor commands. A telemetry path fans out from a lock-free ring buffer to three separate consumers (SD logger, ML detector, downlink), so a slow one can never back-pressure the 1 kHz producer.

Each transport is chosen for its job: an SPSC queue for RC events, a byte-range ring buffer for variable-rate telemetry, and a single atomic for the ML severity flag the control loop reads on its hot path.

RULE No heap allocation anywhere in the control path · every buffer sized at compile time
RULE The 1 kHz control chain runs at the highest priority · nothing preempts it
RULE Slow consumers drop data explicitly · an SD-card stall can't reach the control loop
RULE Failsafe checks live inside the control task · they depend on no other task being healthy
05

Real-time by design

Real-time does not mean fast. It means predictable. Every decision here gives up peak performance for a guarantee I can prove.

Key decision
Preemptive scheduling, not cooperative.

Under cooperative scheduling, if the SD-card logger hits a slow flash block-erase (which flash does, unpredictably, for tens of milliseconds), it holds the CPU and the attitude loop doesn't run. Now the craft is in free fall because of the logging subsystem. The control loop can't depend on the logger behaving, and preemption is what structurally guarantees that.

I looked at three ways the firmware could schedule work. RipWing's tasks are almost all periodic and run-to-completion (read IMU, fuse, run PID, write ESC), which is exactly the shape RTIC is built for.

▸ DECISION MATRIX · execution framework HARD REAL-TIME ATTITUDE CONTROL
Criterion Custom kernel Embassy RTIC
Preemption software cooperative hardware (NVIC)
Fit for periodic run-to-completion manual good native
Data-race safety by discipline async-checked compile-time proof
Preemption cost ~1 µs sw varies ~100 ns hw

The fixed-priority preemptive scheduling rate-monotonic analysis assumes is exactly what the NVIC does in hardware. Put simply: the scheduler is the interrupt controller. Embassy's cooperative executor brings back the hazard I just rejected; the custom kernel stays a side learning track, but it doesn't fly the aircraft.

Priority assignment · rate-monotonic

Shorter period means higher priority. I chose RMS over earliest-deadline-first on purpose, taking the well-known ~69% utilization ceiling in exchange for analyzability and graceful behavior under overload. Take the guarantee you can bound; refuse the one you can't.

HIGH Safety monitor › Attitude control › Sensor fusion › Anomaly detect › Logging › Telemetry LOW
Fault containment

Turn silent corruption into a loud, localized, debuggable fault

You can't reproduce an in-flight fault on the bench, so anything that turns a silent corruption into an immediately debuggable one is worth its overhead. Three, layered:

MPU guards

Guard regions below every task stack. Neither RTIC nor Embassy sets up the MPU by default, so this one is mine. It turns a stack overrun that would corrupt a neighbor into an immediate MemManage fault, tagged with the offending task's ID.

Hard-fault handler

Built early, before it is needed. Captures where it faulted, why (fault status registers), and who (thread ID), writes it to persistent storage, and cuts motor outputs before resetting. Power-cycle, read the record, five-minute fix.

flip-link

Already in the toolchain, a zero-cost linker-level version of the same idea: it flips the memory layout so a stack overflow hits unmapped memory and faults instead of smashing static data.

Threads use the process stack pointer; the kernel and interrupt handlers use the main one, so a blown thread stack faults cleanly instead of taking out the kernel and leaving a dead aircraft with no data. I verify timing with two or three GPIO pins reserved purely for instrumentation: one held high for the control loop shows its runtime and jitter on a scope, one toggled per IMU sample proves a true 1 kHz, one per thread makes priority inversion visible. Unlike a breakpoint, it works in flight. It's also how I'll get the worst-case execution times rate-monotonic analysis needs, since the theory can't give me those numbers. I have to measure them.

06

The flight recorder

A recorder that logs the anomaly detector's outputs and the raw sensor stream is only useful if it survives the incident it recorded.

A filesystem is really just three mappings (offset to block, name to location, used versus free), and the right one is set by the access pattern. RipWing's is write-once, sequential, reliability-critical, special enough to strip the design to its essentials rather than carry a general filesystem.

▸ COMPARISON · log format MUST SURVIVE A CRASH
Bespoke append-only FAT
Free-space mgmt single monotonic pointer free list, fragmentation
Crash safety trivial (pointer only advances) complex
Wear on raw flash near-ideal (sequential) needs management
Laptop-mountable no yes

I chose bespoke. Since I never delete in flight, free-space management collapses to one "next free block" pointer that only advances: no free list, no fragmentation, trivial to make crash-safe, and near-ideal wear on raw flash. The one thing FAT would buy, mounting the card straight onto a laptop, I get back with a small offline tool, and it isn't worth complicating the one data structure that has to survive a crash.

Surviving the crash

The format is append-only with per-record CRCs and a forward-recoverable structure, so a crash of the software or the aircraft leaves a readable trail up to the last instant. Paired with the hard-fault handler flushing its fault frame to this same log, the result is a genuine black box.

Transfer uses DMA with double buffering so the slow SD path never stalls the control loop, and the whole stack sits on the same trait system as everything else, so I unit-test the format against a RAM-backed fake block device with zero hardware.

▸ EMBED · CODE · log record LST.03 · Rust
1 #[repr(C)]
2 struct Record {
3   seq:  u32,        // monotonic
4   kind: u8,        // sensor | anomaly | fault
5   len:  u8,
6   data: [u8; N],
7   crc:  u16,        // per-record
8 }   // write ptr only advances
07

The sampling path

Half the battle is won or lost in the microseconds between the sensor's ADC and the first line of my code.

A real hazard, not a formality

Anti-aliasing before it aliases to DC

The airframe vibrates at the prop blade-pass frequency: motors at 6000 rpm with two-blade props put vibration near 200 Hz plus harmonics. Sample the IMU at 200 Hz with no anti-aliasing and that tone folds straight down to DC. The attitude estimate picks up a constant fake tilt that changes with throttle, the drone flies crooked, and I lose a week blaming the fusion filter for a bug that happened in the ADC before my code ran. The defenses, in order: mechanically isolate the IMU, set the sensor's on-chip anti-alias filter deliberately, and oversample well above the loop rate then decimate in software.

Hardware-timed

Samples come off a timer trigger or the IMU's data-ready interrupt, never a software delay loop. Jitter in the sample timing is its own noise source: it makes the rate non-uniform and smears the frequency content every downstream filter assumes is clean.

Fixed-point to start

Deterministic in timing, which helps the execution-time bound, and far faster on parts with no FPU. The dev target has one, but committing to fixed-point keeps the door open to the Cortex-M0+ port and keeps timing predictable where it matters most.

Filter · median then average

The median goes first. It kills impulsive faults: a corrupted SPI read, a single-bit error, a dropout on one IMU. Averaging first would let a spike smear across the whole window before the median could reject it. The running average then handles the steady vibration floor as an O(1) running sum.

I keep the window short on purpose: every sample of averaging is a sample of delay, and in a fast attitude loop that delay is phase margin spent on stability. The filter length directly sets the crossover frequency I can hit, so it's a value to tune against measured noise, not guess up front.

▸ EMBED · CODE · filter chain LST.04 · Rust
1 // median first: reject the spike before it smears
2 let m = median3(s0, s1, s2);
3 let filtered = avg.push(m);   // O(1) running sum
4 // but log the RAW sample · the detector must see the fault
5 recorder.push(Raw(s2));
Filter for control, log the raw for detection.

There's a deliberate synergy with the ML subsystem: the median rejects the spike so it can't destabilize the aircraft, while the logger still keeps the raw sample so the anomaly detector can learn the fault happened. Every sensor gets characterized before it enters the loop, too: a thousand still samples give the bias to subtract and the noise variance to feed the estimator. Gyro bias drifts with temperature, which is exactly why fusion exists: calibration sets the starting point, fusion tracks the drift from there.

08

Control architecture

A cascade: a fast, analyzable PID loop inside for attitude, a heavier optimizer planned for outside, and a learned model that's never allowed to touch the controls.

Owed before building · near-term

Before I commit to controllers I need to run controllability and observability on the linearized hover model. Controllability confirms the four motors give authority over the full rigid-body state; observability, per candidate sensor set, is the formal backing for which suite makes the state reconstructable. I'm flagging this as near-term, not claiming it's done. The difference between "I verified this" and "I assumed this" is the difference between engineering and hoping.

Inner loop · PID

What the vast majority of real quads fly. I have a reproducible tuning harness for it, and it meets the constraint that matters most on a flight-critical loop: I can analyze it, and I can't certify what I can't analyze. That's why heuristic controllers stay out of the flight-critical path.

Estimator · Kalman / EKF

Nonlinear dynamics make it an EKF. It comes with a warning I take seriously: Doyle's 1978 result shows an optimal estimator plus an optimal controller can have no guaranteed stability margins. So I design it to track stability margins, not just nominal performance. Optimal-in-simulation isn't the same as trustworthy-in-flight.

Outer loop · toward MPC

MPC handles actuator and safety limits natively (geofence, max tilt, min altitude all become constraints it respects) and can see a turn coming over its horizon. The catch is compute, so it belongs on the slower outer loop while the cheap deterministic PID keeps attitude fast. It's a later milestone, not a bring-up requirement.

Gain scheduling

Motor authority falls as the battery sags, so gains tuned at full charge are wrong at low charge. I handle it as a performance map: PID gains stored by operating condition (voltage, throttle, airspeed) and interpolated between. It's the firmware form of the Linear Parameter Varying idea.

Key decision
The model monitors. Classical control actuates.

The anomaly detector observes and classifies. Its output triggers a state transition whose actual response is a deterministic, analyzable fallback controller. The ML never commands the aircraft directly, which is far safer than putting a learned model in the flight-critical path. It has a neat dual use, too: the fault-injection framework that makes the detector's labeled training data is the same rig that tests the fault response. The thing that lets me test the fault safely is the thing that generates the data the detector learns from.

The dev board is an STM32F446RE on a Nucleo; the flight target is the STM32H743, picked for the headroom the ML inference and outer-loop MPC will need. I'm staging that commitment on purpose, confirming the target against measured execution-time budgets once TFLite inference and any MPC actually run, instead of guessing up front. Don't commit the airframe's compute until the utilization numbers are real.

Status board

What's done, what's next

A case study that only shows finished work hides the judgment that lives in sequencing. Nothing is on hardware yet, and that's by design, not omission.

Complete & tested
✓ Control theory foundation · state-space, LQR, Kalman, robustness
✓ 6-DoF simulator + 24-test verification suite
✓ 10-DOF simulator (motor lag + blade flapping) + 17-test suite
✓ Ground-effect model (bounded exponential) + landing / trim study
✓ Model-comparison study · what the added fidelity buys
✓ Automated PID tuning harness (staged Nelder–Mead, noise-injected)
✓ Firmware architecture · data flow, task graph, call graph, crate layout
✓ Architecture decision register (settled / open / collaborator-facing)
Next, in dependency order
01 Controllability / observability verification on the hover model
02 Hardware interface spec for the collaborator · hard deadline at board layout freeze
03 Firmware bring-up on STM32 (RTIC skeleton → control loop → drivers)
04 WCET measurement + rate-monotonic schedulability proof
05 Kalman / EKF state estimator
06 ML anomaly-detection subsystem + fault-injection training data
07 Physical flight testing

What I'd want a reviewer to take away: with a long hardware feedback cycle, I put the work into simulation and host-testability up front. I accepted a slower start in exchange for being able to validate control and kernel logic without hardware, and I enforced that boundary in Cargo's dependency graph instead of by discipline. Preemptive over cooperative, RMS over EDF, bespoke append-only over FAT, PID over anything I can't analyze: every one is the same trade in a different place. Give up peak performance for a guarantee I can prove.

Open channel

Want the deep dive on RipWing?

Happy to walk the simulator, the Cargo-enforced layering, the RTIC scheduling case, or the ML-monitors-classical-actuates split.