Affiliate disclosure
Book titles on this page link to Amazon. As an Amazon Associate, DataField.Dev earns from qualifying purchases — at no additional cost to you.
Further Reading: Reliability, Testing, and Why Rockets Fail
Reliability engineering sits at the intersection of probability, systems engineering, and hard-won operational experience. The sources below are Tier 1 (canonical works we are confident exist) or Tier 2 (a real, named report or resource whose exact edition/URL we do not pin down here). The failure case studies in particular reward going to the primary investigation reports.
Core textbook treatments
Wertz, Everett & Puschell, Space Mission Engineering: The New SMAD. Our systems-engineering anchor. Its chapters on reliability, redundancy, and the test program cover exactly this chapter's material at the level a real spacecraft program uses, including how reliability allocation feeds the mass budget. The reference for the whole progressive project. Tier 1.
NASA Systems Engineering Handbook (SP-2016-6105). NASA's own definitive guide to the review lifecycle (MDR/PDR/CDR), verification and validation, and the "test like you fly" philosophy. Free from NASA. Read it alongside Chapter 29 and this chapter to see how a real agency gates a design. Tier 1 — a real, freely available NASA document.
Sutton & Biblarz, Rocket Propulsion Elements (9th ed.). For the propulsion-reliability side — engine testing, hot-fire and static-fire campaigns, and the failure modes of turbopumps and combustion — Sutton is the standard. Pairs with the engine-out case study. Tier 1.
On reliability and FMEA specifically
MIL-STD-1629A, "Procedures for Performing a Failure Mode, Effects and Criticality Analysis." The classic military standard that codified FMECA. Superseded in places but still the clearest statement of the method's logic and the severity/criticality scales. Tier 2 — a real standard; find the current governing document for a live program.
"Reliability Engineering" — any of the standard texts (e.g., Ebeling, An Introduction to Reliability and Maintainability Engineering). For the mathematics behind $R(t) = e^{-\lambda t}$, the bathtub curve, series/parallel systems, and $k$-of-$n$ redundancy at more depth than we had room for. Tier 2 — a representative, widely used text.
The failure case studies (go to the primary reports)
J. L. Lions et al., Ariane 5 Flight 501 Failure — Report by the Inquiry Board (1996). Short, lucid, and devastating — the single best case study in software reliability ever written for engineers. Read the whole thing; it is only a few pages. Tier 2 — a real, widely circulated report.
NASA, Mars Climate Orbiter Mishap Investigation Board Phase I Report (1999). The units-mismatch loss, from the source. Note especially the process findings about unescalated navigation concerns — the reliability lesson is organizational as much as technical. Tier 2.
Rogers Commission, Report of the Presidential Commission on the Space Shuttle Challenger Accident (1986), and the Columbia Accident Investigation Board Report (2003). The two reports every space engineer should read. Feynman's Appendix F to the Rogers report — the epigraph of this chapter — is the finest short essay on engineering honesty in print. These set up Chapter 37. Tier 1 — landmark public documents.
Watch
Scott Manley, YouTube — videos on launch failures and "why rockets blow up." Clear, technically honest reconstructions of real failures, including the N-1 and several modern anomalies. A good visual companion to §32.5. Tier 2.
Suggested order
- Read this chapter's §32.2 again, then the Ariane 5 inquiry report — it makes every idea in §32.2 and §32.5 concrete in a few pages.
- Skim the NASA Systems Engineering Handbook chapters on the review lifecycle and on verification/test.
- Read Feynman's Appendix F. Then, if you have time, the CAIB report's chapters on organizational cause.
- Do the reliability allocation of Case Study 2 for your own mission, using a reliability text for the $k$-of-$n$ math if you want to go beyond the two-unit case.