Project RipWing
A flight controller I built from scratch. It holds a hard real-time attitude loop while an edge ML monitor watches for sensor degradation and motor faults, but never gets to fly the craft itself. One constraint shaped everything: my hardware feedback loop is long, so I designed and proved almost all of it in simulation before a board existed.
The premise
A remote collaborator handles all of the electrical and PCB work, which makes the hardware feedback loop long. A change I can prototype in software in minutes takes days to see on a board.
So I pushed as much logic as I could into a half that runs on a laptop, and made the layering strict enough that the compiler keeps it honest instead of me. Nearly every choice below is the same trade in a new place: give up some peak performance for a guarantee I can prove.
A controller tuned against a model that lies to you is worse than no controller at all. So the simulator isn't a toy. It's where the control logic actually gets designed, and the same pure-math crates it runs are the ones that compile into the firmware. The line between "logic I can test on my laptop" and "code that needs the chip" is real, enforced, and wide.
The plant
Before any controller, I need the physics it acts on. The first model is the standard rigid-body quadcopter: twelve states, four rotor speeds as instant inputs, checked by a 24-test suite built bottom-up from analytical facts.
Four rotors, six degrees of freedom, so the craft is underactuated: every rotation is bought with a pattern of thrust. Differential thrust on a pair makes roll and pitch; the drag imbalance between the CW and CCW pairs makes yaw. This is what the tuner in the next station steers. Switch views to isolate each axis.
All four rotors spin equal. Thrust sums straight up the body +z axis — this is hover / vertical throttle, the baseline every other axis perturbs from.
FIG.01 — Body-frame force & torque map (after Gibiansky). Thrust Tᵢ = k·ωᵢ² sums to T_B along +z. The model is honest about rigid-body dynamics, but it hides two things that matter enormously for tuning. That is the whole subject of the next station.
The fidelity ladder
Finding out where a model lies is worth more than trusting where it tells the truth. So the simulator grew in deliberate layers, and the gap between each one is itself the lesson.
The 6-DoF model treats rotor speed as an instant input. The second one promotes the four rotor speeds to real states with first-order lag, sixteen states total, and adds physics the base model structurally can't: motor lag, blade flapping, rotor gyroscopics, rate damping.
Blade flapping is the subtle one. A rotor moving through the air makes more lift on its advancing blade than its retreating one, tilting the rotor away from the direction of motion. That gives a horizontal drag force (the dominant velocity damping on a real multirotor) plus a pitch-up moment in forward flight. It settles within one rotor revolution (~4 ms) against ~100 ms body dynamics, so I model it quasi-statically, the way the control literature does. The parameters are representative values from the literature, not measured off a specific airframe. I say so plainly, because passing borrowed numbers off as measured ones is the kind of overclaim that discredits a whole simulation.
| Experiment | 6-DoF | 10-DOF | What it means |
|---|---|---|---|
| Velocity decay from 4 m/s | never decays | t½ = 2.33 s | Flapping is the real damping; the 6-DoF drone coasts forever |
| Roll step, cmd vs actual rpm | identical | 1659 rpm gap | Motor lag is a real actuator, not a wire |
| Rate-loop gain sweep ×1→×48 | ripple flat | 5.36 °/s buzz | The motor pole caps usable gain; 6-DoF rewards infinite gain |
| Steady 1.5 m/s cruise | 0.76° pitch | 1.61° pitch | ~2× tilt needed to overcome flapping drag |
The third row is the one that matters for tuning. Gains tuned on the 6-DoF model are wrong in exactly the direction tuning pushes them. Higher gain always looks free on the simpler model, because it has no actuator pole to ring against. A test confirms it: with flapping on, the hover mode is open-loop unstable (eigenvalues at +0.59 ± 1.34j), the classic unstable hover mode every real rotorcraft has and the 6-DoF model can't show. The two suites total 41 passing tests, and with the new coefficients zeroed the 16-state model reproduces the 6-DoF one to within 1e-9, so the upgrade is a strict superset.
Ground effect
Within about one rotor radius of the ground, downwash can't escape freely, so the same rotor speed makes more thrust. A quad settling onto a pad feels a rising thrust surplus right as it tries to land. If the controller doesn't know about it, the craft floats and won't touch down cleanly.
Cheeseman–Bennett has a singularity
The famous textbook relation blows up to infinity as height nears a quarter of the rotor radius, exactly the range a landing controller lives in. Feed a model that diverges into a simulator you mean to land, and you get a numerical explosion at the worst possible moment.
A bounded exponential instead
Stays finite all the way to the ground, a sensible ceiling at contact instead of divergence, while matching the same falloff higher up. Picking the well-behaved model over the famous one only pays off once you actually try to simulate the hard case.
A four-experiment study turns landing from a surprise into a modeled, testable scenario. The thrust-surplus float becomes a disturbance I can design the altitude loop against, ground effect makes the craft mildly self-leveling near the pad (a quirk to exploit), and the hover-trim shift on approach is exactly the operating-point change the gain-scheduling map is there to absorb.
Automated PID tuning
Drag the craft, fire a disturbance, command a step. This live roll-axis loop runs the same rigid-body dynamics, J·θ̈ = τ, at the 400 Hz firmware rate. It also shows off the exact problem I refused to solve by hand.
FIG.02 — Live closed-loop sim. Watch the saturation readout: when the motor differential hits its ±0.6 N·m limit, an unbounded integral term keeps piling up error it can't act on. That is integral windup, and it's why the integral here, and in the firmware, is clamped.
Sliding three gains by feel until the trace looks right is how most quads get tuned. It's also unreproducible, only one noise seed deep, and slow to redo on every airframe change. So instead of hand-tuning, I built a harness: staged Nelder–Mead in log-gain space, searching against the noise-injected 10-DOF plant. Stage 1 tunes the inner attitude and rate loops on a 15° roll doublet; stage 2 freezes those and tunes the outer position loop on a 2 m step. The gains come out reproducible, with robustness measured across noise seeds. That's a stronger claim than "it flew on my bench."
The interesting part wasn't the optimizer. It was finding the four ways a naive setup produces gains that are optimal for the wrong problem. Each is a spot where an honest model quietly rewards a controller that would be dangerous on real hardware.
My first noise-free run came back with a derivative gain near 400. With perfect measurements, the optimizer had found that derivative gain is free damping. Adding realistic noise (0.02 m, 0.4°, 1°/s) and a motor-thrash penalty makes it cost what it costs on hardware, and the tuned Kd shrinks instead of exploding.
With no persistent disturbance, the integral gain is unidentifiable. A small shifted-battery torque bias makes Ki earn its keep, and it stays interpretable: with high rate stiffness the bias gets rejected through a tiny attitude offset, so Ki correctly falls toward zero.
Yaw authority is only ±0.26 N·m across the whole motor range. My original limit let the controller demand 4× the yaw torque that exists; the allocator "solved" that by slamming one motor pair to zero and the other to max, a positive-feedback yaw ratchet that noise reliably tripped. The fix clamps yaw demand to half the physical authority before allocation.
Numerically differentiating a noisy error at 200 Hz amplifies measurement noise by roughly 280×. Every real flight controller filters the derivative term, so the PID class grew a configurable low-pass and the tuning uses it.
| Metric | Default | Tuned |
|---|---|---|
| Doublet overshoot | 58% | 1.1% |
| Rate ripple | 20.5 °/s | 1.1 °/s |
| Rise time | 0.28 s | 0.33 s |
Cross-seed validation on five unseen noise realizations improves the ITAE metric 7–8× every time, so the gains aren't overfit to one lucky draw.
Firmware architecture
A strict layered stack, built as a Rust workspace. Each layer is a crate that can only depend on the layers below it, and Cargo's dependency graph makes breaking that rule a compile error instead of a convention I have to remember.
The key property: the control crate and the ML crate depend on
nothing embedded. They compile and test on my laptop with
cargo test. The driver crate depends only on the
embedded-hal traits, never a specific chip's HAL, so the same IMU
driver runs on any board whose HAL implements them. Only the firmware binary is
allowed to name the STM32.
The Cargo workspace is the call graph, and Rust's dependency rules enforce it: the pure-math crate literally can't import a HAL, and Cargo forbids circular dependencies. That's a stronger guarantee than a diagram a reviewer has to police by hand. I develop and debug the whole control and ML path on my laptop, and CI fails the moment anyone tries to sneak a hardware dependency into a host-testable crate.
Sensor data comes in at up to 1 kHz from a dual IMU over SPI, feeds the attitude estimator and cascaded PID loops, and leaves as DShot motor commands. A telemetry path fans out from a lock-free ring buffer to three separate consumers (SD logger, ML detector, downlink), so a slow one can never back-pressure the 1 kHz producer.
Each transport is chosen for its job: an SPSC queue for RC events, a byte-range ring buffer for variable-rate telemetry, and a single atomic for the ML severity flag the control loop reads on its hot path.
Real-time by design
Real-time does not mean fast. It means predictable. Every decision here gives up peak performance for a guarantee I can prove.
Under cooperative scheduling, if the SD-card logger hits a slow flash block-erase (which flash does, unpredictably, for tens of milliseconds), it holds the CPU and the attitude loop doesn't run. Now the craft is in free fall because of the logging subsystem. The control loop can't depend on the logger behaving, and preemption is what structurally guarantees that.
I looked at three ways the firmware could schedule work. RipWing's tasks are almost all periodic and run-to-completion (read IMU, fuse, run PID, write ESC), which is exactly the shape RTIC is built for.
| Criterion | Custom kernel | Embassy | RTIC |
|---|---|---|---|
| Preemption | software | cooperative | hardware (NVIC) |
| Fit for periodic run-to-completion | manual | good | native |
| Data-race safety | by discipline | async-checked | compile-time proof |
| Preemption cost | ~1 µs sw | varies | ~100 ns hw |
The fixed-priority preemptive scheduling rate-monotonic analysis assumes is exactly what the NVIC does in hardware. Put simply: the scheduler is the interrupt controller. Embassy's cooperative executor brings back the hazard I just rejected; the custom kernel stays a side learning track, but it doesn't fly the aircraft.
Shorter period means higher priority. I chose RMS over earliest-deadline-first on purpose, taking the well-known ~69% utilization ceiling in exchange for analyzability and graceful behavior under overload. Take the guarantee you can bound; refuse the one you can't.
Turn silent corruption into a loud, localized, debuggable fault
You can't reproduce an in-flight fault on the bench, so anything that turns a silent corruption into an immediately debuggable one is worth its overhead. Three, layered:
Guard regions below every task stack. Neither RTIC nor Embassy sets up the MPU by default, so this one is mine. It turns a stack overrun that would corrupt a neighbor into an immediate MemManage fault, tagged with the offending task's ID.
Built early, before it is needed. Captures where it faulted, why (fault status registers), and who (thread ID), writes it to persistent storage, and cuts motor outputs before resetting. Power-cycle, read the record, five-minute fix.
Already in the toolchain, a zero-cost linker-level version of the same idea: it flips the memory layout so a stack overflow hits unmapped memory and faults instead of smashing static data.
Threads use the process stack pointer; the kernel and interrupt handlers use the main one, so a blown thread stack faults cleanly instead of taking out the kernel and leaving a dead aircraft with no data. I verify timing with two or three GPIO pins reserved purely for instrumentation: one held high for the control loop shows its runtime and jitter on a scope, one toggled per IMU sample proves a true 1 kHz, one per thread makes priority inversion visible. Unlike a breakpoint, it works in flight. It's also how I'll get the worst-case execution times rate-monotonic analysis needs, since the theory can't give me those numbers. I have to measure them.
The flight recorder
A recorder that logs the anomaly detector's outputs and the raw sensor stream is only useful if it survives the incident it recorded.
A filesystem is really just three mappings (offset to block, name to location, used versus free), and the right one is set by the access pattern. RipWing's is write-once, sequential, reliability-critical, special enough to strip the design to its essentials rather than carry a general filesystem.
| Bespoke append-only | FAT | |
|---|---|---|
| Free-space mgmt | single monotonic pointer | free list, fragmentation |
| Crash safety | trivial (pointer only advances) | complex |
| Wear on raw flash | near-ideal (sequential) | needs management |
| Laptop-mountable | no | yes |
I chose bespoke. Since I never delete in flight, free-space management collapses to one "next free block" pointer that only advances: no free list, no fragmentation, trivial to make crash-safe, and near-ideal wear on raw flash. The one thing FAT would buy, mounting the card straight onto a laptop, I get back with a small offline tool, and it isn't worth complicating the one data structure that has to survive a crash.
The format is append-only with per-record CRCs and a forward-recoverable structure, so a crash of the software or the aircraft leaves a readable trail up to the last instant. Paired with the hard-fault handler flushing its fault frame to this same log, the result is a genuine black box.
Transfer uses DMA with double buffering so the slow SD path never stalls the control loop, and the whole stack sits on the same trait system as everything else, so I unit-test the format against a RAM-backed fake block device with zero hardware.
The sampling path
Half the battle is won or lost in the microseconds between the sensor's ADC and the first line of my code.
Anti-aliasing before it aliases to DC
The airframe vibrates at the prop blade-pass frequency: motors at 6000 rpm with two-blade props put vibration near 200 Hz plus harmonics. Sample the IMU at 200 Hz with no anti-aliasing and that tone folds straight down to DC. The attitude estimate picks up a constant fake tilt that changes with throttle, the drone flies crooked, and I lose a week blaming the fusion filter for a bug that happened in the ADC before my code ran. The defenses, in order: mechanically isolate the IMU, set the sensor's on-chip anti-alias filter deliberately, and oversample well above the loop rate then decimate in software.
Samples come off a timer trigger or the IMU's data-ready interrupt, never a software delay loop. Jitter in the sample timing is its own noise source: it makes the rate non-uniform and smears the frequency content every downstream filter assumes is clean.
Deterministic in timing, which helps the execution-time bound, and far faster on parts with no FPU. The dev target has one, but committing to fixed-point keeps the door open to the Cortex-M0+ port and keeps timing predictable where it matters most.
The median goes first. It kills impulsive faults: a corrupted SPI read, a single-bit error, a dropout on one IMU. Averaging first would let a spike smear across the whole window before the median could reject it. The running average then handles the steady vibration floor as an O(1) running sum.
I keep the window short on purpose: every sample of averaging is a sample of delay, and in a fast attitude loop that delay is phase margin spent on stability. The filter length directly sets the crossover frequency I can hit, so it's a value to tune against measured noise, not guess up front.
There's a deliberate synergy with the ML subsystem: the median rejects the spike so it can't destabilize the aircraft, while the logger still keeps the raw sample so the anomaly detector can learn the fault happened. Every sensor gets characterized before it enters the loop, too: a thousand still samples give the bias to subtract and the noise variance to feed the estimator. Gyro bias drifts with temperature, which is exactly why fusion exists: calibration sets the starting point, fusion tracks the drift from there.
Control architecture
A cascade: a fast, analyzable PID loop inside for attitude, a heavier optimizer planned for outside, and a learned model that's never allowed to touch the controls.
Before I commit to controllers I need to run controllability and observability on the linearized hover model. Controllability confirms the four motors give authority over the full rigid-body state; observability, per candidate sensor set, is the formal backing for which suite makes the state reconstructable. I'm flagging this as near-term, not claiming it's done. The difference between "I verified this" and "I assumed this" is the difference between engineering and hoping.
What the vast majority of real quads fly. I have a reproducible tuning harness for it, and it meets the constraint that matters most on a flight-critical loop: I can analyze it, and I can't certify what I can't analyze. That's why heuristic controllers stay out of the flight-critical path.
Nonlinear dynamics make it an EKF. It comes with a warning I take seriously: Doyle's 1978 result shows an optimal estimator plus an optimal controller can have no guaranteed stability margins. So I design it to track stability margins, not just nominal performance. Optimal-in-simulation isn't the same as trustworthy-in-flight.
MPC handles actuator and safety limits natively (geofence, max tilt, min altitude all become constraints it respects) and can see a turn coming over its horizon. The catch is compute, so it belongs on the slower outer loop while the cheap deterministic PID keeps attitude fast. It's a later milestone, not a bring-up requirement.
Motor authority falls as the battery sags, so gains tuned at full charge are wrong at low charge. I handle it as a performance map: PID gains stored by operating condition (voltage, throttle, airspeed) and interpolated between. It's the firmware form of the Linear Parameter Varying idea.
The anomaly detector observes and classifies. Its output triggers a state transition whose actual response is a deterministic, analyzable fallback controller. The ML never commands the aircraft directly, which is far safer than putting a learned model in the flight-critical path. It has a neat dual use, too: the fault-injection framework that makes the detector's labeled training data is the same rig that tests the fault response. The thing that lets me test the fault safely is the thing that generates the data the detector learns from.
The dev board is an STM32F446RE on a Nucleo; the flight target is the STM32H743, picked for the headroom the ML inference and outer-loop MPC will need. I'm staging that commitment on purpose, confirming the target against measured execution-time budgets once TFLite inference and any MPC actually run, instead of guessing up front. Don't commit the airframe's compute until the utilization numbers are real.
What's done, what's next
A case study that only shows finished work hides the judgment that lives in sequencing. Nothing is on hardware yet, and that's by design, not omission.
What I'd want a reviewer to take away: with a long hardware feedback cycle, I put the work into simulation and host-testability up front. I accepted a slower start in exchange for being able to validate control and kernel logic without hardware, and I enforced that boundary in Cargo's dependency graph instead of by discipline. Preemptive over cooperative, RMS over EDF, bespoke append-only over FAT, PID over anything I can't analyze: every one is the same trade in a different place. Give up peak performance for a guarantee I can prove.
Want the deep dive on RipWing?
Happy to walk the simulator, the Cargo-enforced layering, the RTIC scheduling case, or the ML-monitors-classical-actuates split.