Forensic Failure Files #6 — When the Margin for Luck Disappears
Case file: Boeing 737 MAX, MCAS, 2018–2019
Forensic Failure Files #5 made the case with a drag strip and a Las Vegas landing ramp: some people find the truth in the data on purpose, and some people get lucky. Drag strips and stunt jumps are where that gap shows up cheap. Aviation is where it shows up expensive — and where nobody gets a second flight to eyeball it and go with their gut.
The Case File
To fit a larger, more fuel-efficient engine on the 737 MAX, Boeing mounted it farther forward and higher on the wing than on earlier 737s.[^1] That relocation changed the airplane's aerodynamics: under certain flaps-retracted, low-speed, nose-up conditions, the new engine position gave the aircraft an unwanted tendency to pitch up.[^2] Boeing's fix was software — the Maneuvering Characteristics Augmentation System, MCAS — designed to trim the nose back down automatically so the MAX would handle like the 737s pilots were already certified on.[^3]
MCAS took its input from a single angle-of-attack sensor, with no cross-check against the second sensor mounted on the other side of the fuselage, even though the aircraft carried two.[^4] An AoA disagree alert was designed to be a standard feature warning pilots when the two sensors disagreed — but a software error linked it instead to an optional angle-of-attack indicator, so the alert only activated on the roughly 20% of MAX aircraft whose airlines had paid for that extra display. Neither crashed aircraft had it.[^5] The AoA sensor at the center of both accidents had itself been flagged in more than 200 incident reports to the FAA before the crashes, and Boeing did not flight-test the failure scenario that ultimately occurred.[^6]
On October 29, 2018, Lion Air Flight 610 crashed into the Java Sea shortly after takeoff, killing all 189 aboard.[^7] On March 10, 2019, Ethiopian Airlines Flight 302 crashed shortly after takeoff from Addis Ababa, killing all 157 aboard.[^8] In both cases, a single faulty AoA reading triggered MCAS, which repeatedly drove the nose down. The Lion Air crew had no prior knowledge that MCAS existed. The Ethiopian crew, flying five months after an emergency directive following Lion Air had briefed pilots on the runaway-stabilizer procedure, correctly diagnosed the problem and cut out electric trim as instructed — but under the aerodynamic loads at their speed, manual trim wouldn't move the stabilizer, and when they re-engaged electric power to try to correct it, MCAS reactivated and drove the nose down again before they could recover.[^9] The two crashes killed 346 people combined.[^10]
The evidence that matters most: Late in MCAS's development, its authority to move the horizontal stabilizer per activation was quietly increased from an original limit of 0.6 degrees to 2.5 degrees — nearly half the stabilizer's total range of motion, and repeatable every time the system reset. Boeing's safety assessment, however, was never updated to reflect it; the certification submitted to the FAA still classified an MCAS malfunction as a "Major" hazard rather than "Hazardous" or "Catastrophic," a classification that permitted a single-sensor design with no redundancy requirement.[^11] That classification is why a single point of sensor failure passed review at all. The engineers didn't skip a check. The process itself, as submitted, no longer matched what the system could actually do.
The forensic point: Boeing had data. They automated a response to it. Nobody re-ran the hazard classification once the system's authority changed underneath it. That is Armstrong's clutch problem exactly — an assumption tuned around instead of tested — except here the assumption was baked into flight-critical software, verified against an outdated version of itself, and the consequence was not a slower run. It was 346 deaths.
The Clock Gets Faster From Here
What separates MCAS from the clutch and the ramp isn't the underlying failure — trusting an unverified input — it's the amount of time available to catch it. MCAS could reactivate roughly five seconds after a pilot released the manual trim switch if the faulty AoA reading was still present; crews had a narrow, repeatedly-closing window to diagnose and override it, and in both accidents ran out of altitude first.[^12]
Self-driving cars compress that window into single-digit milliseconds, with no crew and no post-flight investigator to read the logs — the judgment has to be built in before the fact, because there's no time to build it in after. Rockets push the same problem further, generating more telemetry per launch than a person could review in a lifetime, sometimes on hardware that doesn't survive to be inspected. Same problem, faster clock, less room for a wrong guess.
Is AI Different, Or Just Higher-Stakes Software?
MCAS raises a fair question: is AI actually a new problem, or just software with worse consequences? Software has always had this failure mode, and MCAS is the proof case — a system trusted one input, acted automatically, and nobody built in the check.
That part isn't new, and AI inherits it. But there's a place the comparison breaks down. MCAS's logic was fully specified — an engineer wrote the rule: if AoA exceeds a given value, trim nose-down by a given amount. Once investigators pulled the software, they could point to the line. The failure was bad, but it was auditable.
A lot of what AI produces doesn't come from a rule anyone wrote. It comes from a model that learned a pattern from training data, and when it's wrong there frequently isn't a single line to point to — the "why" is distributed across weights nobody can read like a rulebook. AI doesn't just carry MCAS's failure mode forward. It can stack a second one on top: not just trusting bad data automatically, but being unable to fully explain, even after the fact, why the system did what it did.
There's a second distinction worth naming, and it's the one that actually changes the shape of the problem: every instrument in the drag-racing and aviation cases — Armstrong's data recorder, the flight data recorder, rocket telemetry — was built to answer one defined question first. The data had a job. A lot of what AI generates doesn't start from a question at all; it's output produced because the system can produce it. That's a failure mode upstream of "misread the data" — nobody ever decided what would count as signal in the first place.
Afterthought: The Standard for Critical Software Corrections
This is the practical rule the MCAS case file argues for, stated plainly: any software correction touching a safety-critical function has to be justified against a stated, sufficient body of data — not a plausible-sounding assumption, and not a single input trusted because nobody thought to ask what would happen if it lied.
That standard has three parts.
How much data is enough has to be answered, not assumed. "We had a sensor" is not a data standard. "We had one sensor, no cross-check, and no flight test of what happens if it fails" is what MCAS actually had, and it was not enough — not because more data is always better in the abstract, but because nobody stated in advance what quantity and diversity of input the correction needed to be trusted, so nobody could tell they'd fallen short until 346 people were dead.
If the data doesn't exist, building the means to collect it is part of the engineering job — not a step that gets skipped because the deadline is close. Armstrong didn't have clutch-lockup data until he built a way to get it. Boeing had the AoA sensor hardware already on the airframe and chose not to cross-check it, and chose not to flight-test the failure scenario that killed two planes. The difference between those two stories is the difference between a craftsman and a shortcut.
A single reading, or a handful of them, is not a statistical standard. One data point can tell you what happened once. It cannot tell you what the system does across its full operating envelope, under the range of conditions a critical correction will actually see in service. That requires continuous, constant data collection — not a snapshot from a test flight or a bench run — sustained long enough and broadly enough to establish a real statistical picture of normal versus failing behavior. MCAS treated a single sensor's momentary reading as ground truth. A statistically sound system treats any single reading as one data point in a distribution, and knows what "outside the distribution" looks like before it trusts the reading enough to move a control surface.
Put together: a critical software correction is only as sound as the data behind it, the data has to be built if it doesn't already exist, and one reading — or even a few — is never a substitute for a standard established across enough continuous data to know what normal actually looks like.
The Call to Action: Sloppy Software Has to Be Cleaned Up
MCAS wasn't inevitable. A single point of sensor trust with no cross-check is not an unavoidable cost of complexity — it's sloppy engineering, the kind that gets caught in review on any system where someone is personally accountable for the design. It got built and certified anyway.
A Professional Engineer's stamp on a bridge or a pressure vessel means a specific, licensed individual reviewed the design and is personally answerable if it fails. Safety-critical software already has its own version of rigor — standards like DO-178C and ARP4761 exist precisely to classify failure conditions and set the assurance level a system has to meet before it flies. MCAS went through that process. The standards didn't fail; the classification feeding them did, and it went unrevised after the system's authority was quietly expanded. That's the real gap: not the absence of a standard, but the absence of one accountable, licensed individual whose name and liability were on the line when that classification stopped matching the design. Nobody's license was on the line when MCAS's authority tripled without its hazard rating being revisited.
References
[^1]: Joint Authorities Technical Review (JATR), Boeing 737 MAX Flight Control System: Observations, Findings, and Recommendations, October 2019.
[^2]: Ibid.
[^3]: Ibid.
[^4]: U.S. House Committee on Transportation and Infrastructure, The Design, Development & Certification of the Boeing 737 MAX, Final Committee Report, September 2020.
[^5]: Boeing press statement, "Boeing Statement on AOA Disagree Alert," May 5, 2019 (software linked the standard AOA Disagree alert to the optional AOA indicator, so only aircraft with the paid indicator had the alert active); CNBC, "Boeing says disabled alert on 737 Max wasn't necessary for safe operation," May 5, 2019.
[^6]: CNN, "Boeing relied on single sensor for 737 Max that had been flagged 216 times to FAA," April 30, 2019.
[^7]: Komite Nasional Keselamatan Transportasi (KNKT), Final Aircraft Accident Investigation Report, Lion Air PK-LQP, KNKT.18.10.35.04, October 2019.
[^8]: Ethiopian Civil Aviation Authority, Aircraft Accident Investigation Bureau, Interim Investigation Report, Ethiopian Airlines ET-302, 2019/2020.
[^9]: Ethiopian AAIB Interim Report, op. cit.; U.S. House Committee Final Report, op. cit., on the Lion Air crew's lack of prior MCAS awareness versus the Ethiopian crew's briefing under the post–Lion Air emergency airworthiness directive.
[^10]: KNKT Final Report and Ethiopian AAIB Interim Report, op. cit. (189 + 157 = 346).
[^11]: U.S. House Committee Final Report, op. cit., and U.S. DOT Office of Inspector General report on FAA's 737 MAX certification, on the increase in MCAS stabilizer authority from 0.6 to 2.5 degrees and the "Major" hazard classification (Development Assurance Level C) that authority increase was never re-assessed against.
[^12]: Ethiopian AAIB Interim Report, op. cit., on MCAS reactivation timing relative to electric trim input.
Herbert Roberts, P.E. is a licensed professional engineer with 30+ years in aviation research and development across three companies, and has spent eight years analyzing accidents under his P.E. license.


