Engineering for Failure: Why the Best Machines Are Not Designed to Never Fail

How Fail-Safe Design, Damage Tolerance, Redundancy, Fracture Mechanics and Predictive Maintenance Turn Failure from a Catastrophe into a Manageable Engineering Event

Engr. Faisal Abbas

BSc Electrical Engineering

Table of Contents

  1. The Dangerous Idea of the “Failure-Proof” Machine
  2. Failure Prevention vs. Failure Management
  3. Why Machines Eventually Fail
  4. From Strength to Damage Tolerance
  5. Fail-Safe Design and Redundancy
  6. Fatigue and Fracture: When Small Damage Becomes Critical
  7. Detecting Failure Before It Becomes Catastrophic
  8. Graceful Degradation and Safe States
  9. Predictive Maintenance: Managing the Remaining Life
  10. The Hidden Problem of Common-Cause Failure
  11. Designing the Failure Path, Not Just the Operating Path
  12. Rethinking What “Reliability” Really Means
  13. Conclusion: The Machine That Knows How to Fail

The Dangerous Idea of the “Failure-Proof” Machine

Engineering has traditionally been associated with strength, safety, reliability and durability. We calculate stresses, apply factors of safety, specify fatigue lives, select corrosion-resistant materials and test components under increasingly severe conditions.

All of this is necessary. But there is a dangerous assumption hiding behind the language of reliability: that the ultimate objective of engineering is to make failure disappear.

It cannot. Every physical machine operates in an uncertain world. Materials contain microscopic imperfections. Loads vary. Temperatures fluctuate. Lubricants deteriorate. Bearings wear. Welds contain discontinuities. Fasteners loosen. Electrical connections corrode. Operators make mistakes. Sensors drift. Manufacturing processes introduce variation. Maintenance is sometimes late, incomplete or incorrectly performed. Even when the original design is excellent, the machine does not remain in its original condition.

The more realistic engineering question is therefore not simply: “How do we prevent this machine from failing?”

It is: “If something begins to fail, what will the machine do next?”

That change in perspective is fundamental. A high-quality engineering system should prevent failures where reasonably possible. But it should also detect failures, contain them, tolerate them, communicate them, degrade predictably and, when necessary, shut down safely.

NASA’s systems-engineering guidance explicitly treats failure management as a system-level concern, defining failure tolerance in terms of retaining capability despite specified failures and describing fault management as the ability to detect, diagnose, isolate, respond to and recover from conditions that threaten system operation. This is the deeper meaning of engineering for failure.

Failure Prevention vs. Failure Management

The distinction can be summarized simply:

Failure preventionFailure management
Prevent the fault from occurringAssume some faults will occur
Increase strengthProvide residual strength
Reduce stressControl damage propagation
Improve material qualityDetect defects and deterioration
Prevent overloadManage overload consequences
Extend fatigue lifeDetect fatigue damage
Prevent corrosionMonitor and control corrosion
Increase reliabilityProvide fault tolerance
Design for normal operationDesign for abnormal operation
Ask “How long will it last?”Ask “What happens when it starts to deteriorate?”

The two approaches are not competitors. Failure prevention comes first. Failure management is what makes the system resilient when prevention is imperfect. The mistake is to treat the first as sufficient.

Consider a rotating shaft. An engineer may calculate the shaft’s nominal stress and demonstrate that it has an acceptable factor of safety. That addresses one question: Can the shaft carry the intended load? But several other questions remain.

What happens if a fatigue crack begins at a keyway? What happens if lubrication deteriorates? What happens if the shaft experiences an unexpected transient load? What happens if a bearing begins generating excessive vibration? Can the developing fault be detected? Can the machine reduce its speed or load? Can the damaged component continue operating temporarily? Can it fail without damaging neighboring components? Those questions belong to failure management.

Why Machines Eventually Fail

Failure is not a single phenomenon. It is the accumulated consequence of mechanical, environmental, manufacturing, operational and human factors.

Common mechanisms include:

  • fatigue from repeated or fluctuating stresses;
  • wear caused by repeated contact and material removal;
  • corrosion caused by chemical or electrochemical interaction with the environment;
  • creep under sustained stress at elevated temperature;
  • thermal fatigue caused by repeated temperature changes;
  • fracture caused by crack initiation and propagation;
  • buckling caused by instability rather than material strength;
  • vibration-induced damage caused by dynamic loading;
  • manufacturing defects such as inclusions, voids or dimensional deviations;
  • assembly errors such as incorrect torque or alignment;
  • overloads beyond design assumptions;
  • human and maintenance errors;
  • unexpected environmental conditions.

This creates an important engineering reality: A machine’s original design condition is temporary. Its condition evolves throughout its life.

The machine as a changing system

A useful conceptual model is:

New machine → operating loads → wear/damage → detectable degradation → intervention OR continued degradation → functional failure → safe shutdown/recovery OR catastrophic failure

The objective of modern engineering is not necessarily to stop this entire chain from existing. It is to interrupt the chain before the consequences become unacceptable.

From Strength to Damage Tolerance

Traditional strength calculations ask whether a structure can withstand a specified load. Damage-tolerant engineering asks a harder question: What if the structure is already damaged?

This is a profound shift. A component may contain a crack that is too small to affect immediate structural strength but large enough to grow under cyclic loading. Fracture mechanics provides a framework for understanding this process.

One important parameter is the stress-intensity factor:

K = Yσ√(πa)

where:

  • (K) = stress-intensity factor;
  • (Y) = geometry factor;
  • (σ) = applied stress;
  • (a) = characteristic crack size.

Fracture becomes critical when the relevant stress intensity approaches the material’s fracture toughness:

K → K_IC

The significance is enormous.

The question is no longer simply whether the material is “strong enough.” It becomes:

How large can a flaw become before the remaining structure can no longer safely carry the load?

That is the foundation of damage tolerance.

The aviation industry provides one of the clearest examples. The FAA’s fatigue and damage-tolerance discipline explicitly combines fatigue, fracture mechanics, material behavior, environmental effects, inspection, structural health monitoring and life-cycle management.

FAA guidance for transport aircraft also requires damage-tolerance evaluations that account for structural damage and establish inspection strategies. This is engineering built around an uncomfortable assumption: damage may exist. The design must therefore remain safe long enough for that damage to be detected and corrected.

Fail-Safe Design and Redundancy

Fail-safe design is perhaps the clearest expression of engineering for failure. A fail-safe system does not necessarily prevent the first failure. Instead, it prevents that failure from immediately becoming catastrophic. Imagine a structure carrying a load through two independent paths.

If one path fails:

Load Path A → Failure

the remaining path can temporarily carry the load:

Load Path B → Residual Load

This creates failure tolerance.

FAA guidance describes fail-safe aircraft structures in terms of retaining required residual strength after failure or partial failure of a principal structural element, commonly through redundant or multiple load paths.

However, redundancy is often misunderstood. Two components are not automatically independent simply because there are two of them.

Suppose two pumps share:

  • the same electrical supply;
  • the same cooling system;
  • the same control software;
  • the same pipe;
  • the same mounting structure.

A single common failure can disable both. This is called a common-cause or common-mode failure. NASA guidance makes this point explicitly: redundancy alone does not guarantee failure tolerance if the redundant elements share vulnerabilities.

The redundancy paradox

Adding redundancy can therefore produce:

More components → more potential individual failures

while simultaneously producing:

More independent paths → greater system resilience

The engineering objective is not simply to maximize the number of components. It is to maximize useful independence.

Fatigue and Fracture: When Small Damage Becomes Critical

One of the most dangerous forms of mechanical failure is fatigue because the component may look perfectly acceptable for much of its life. Fatigue is driven by repeated or fluctuating loading. A simplified fatigue relationship is often represented using an S–N curve, where stress amplitude is related to cycles to failure.

Conceptually:

Stress
  ↑
  │\
  │ \
  │  \
  │   \
  │    \________
  │             \____
  └────────────────────→ Number of cycles
       Low life      High life

But fatigue is not simply a countdown timer. Real components may experience variable-amplitude loading, corrosion, temperature changes, surface damage and manufacturing defects. Once a crack exists, fracture mechanics can be used to estimate crack-growth behavior. A classic representation is the Paris law:

da/dN = C(ΔK)^m

where:

  • (a) = crack size;
  • (N) = number of load cycles;
  • (C,m) = material/environment parameters;
  • (ΔK) = cyclic stress-intensity range.

The engineering implication is powerful: failure is often a process rather than an instantaneous event. A crack can begin microscopic, grow gradually and eventually reach a critical size. That creates an opportunity. If the system can detect the deterioration before critical failure, the failure can become a maintenance event rather than an accident. FAA’s damage-tolerance guidance specifically incorporates crack growth, residual strength and inspection opportunities into structural safety.

Detecting Failure Before It Becomes Catastrophic

Failure management depends on information. A machine cannot manage a failure it cannot detect. This is why modern engineering increasingly combines mechanical design with sensors, control systems and data analysis. Depending on the machine, useful signals can include:

  • vibration;
  • temperature;
  • pressure;
  • acoustic emissions;
  • electrical current;
  • rotational speed;
  • oil debris;
  • strain;
  • displacement;
  • leakage;
  • energy consumption;
  • operating cycles.

A developing bearing fault, for example, may produce abnormal vibration before complete seizure. A deteriorating electrical connection may produce abnormal temperature. A pump may show changing pressure and power characteristics before complete loss of function. This leads to a hierarchy:

Detect → Diagnose → Predict → Act

The objective is not merely to generate more data. It is to convert physical deterioration into an actionable engineering decision. NIST describes condition monitoring as the detection, diagnosis and prediction of faults or failures, and its recent systematic review highlights the growing role of AI and IoT technologies in condition-based maintenance. NASA applies the same broader principle at system level: critical faults should be detected, isolated and, where possible, recovered before they produce catastrophic consequences.

Graceful Degradation and Safe States

A particularly sophisticated machine does not move directly from:

FULL FUNCTION → TOTAL FAILURE

Instead, it can move through controlled states:

FULL PERFORMANCE
       │
       ▼
MINOR DEGRADATION
       │
       ▼
WARNING / DETECTION
       │
       ▼
REDUCED PERFORMANCE
       │
       ▼
SAFE OPERATING STATE
       │
       ▼
CONTROLLED SHUTDOWN
       │
       ▼
REPAIR / RECOVERY

This is graceful degradation. It matters because failure severity depends not only on whether a component fails, but also on how the system responds to that failure. A failed component that causes an immediate loss of control is fundamentally different from one that causes a warning followed by reduced performance and controlled shutdown.

NASA guidance specifically identifies predictable degradation as desirable because it can provide time for detection and recovery. This principle can be applied far beyond aerospace. Industrial machinery can reduce speed. Power systems can isolate faulty sections. Vehicles can enter restricted operating modes. Robotic systems can disable a malfunctioning actuator while maintaining control through remaining actuators. Computing systems can restart failed services while keeping other services operational.

The best system does not merely survive failure. It fails in an understandable way.

Predictive Maintenance: Managing the Remaining Life

Maintenance has historically been dominated by two simple strategies.

Reactive maintenance

Wait until it breaks. Cheap in some circumstances, disastrous in others.

Preventive maintenance

Replace it at a predetermined interval. Better, but imperfect. A component may fail before its scheduled replacement, or it may be replaced while it still has substantial useful life. Predictive maintenance introduces a different question: What condition is the component actually in?

The basic concept is:

Measured condition → Health assessment → Remaining useful life → Maintenance decision

NIST describes condition-based maintenance as an alternative to purely failure-based or schedule-based maintenance, using sensor information to identify abnormal behavior and estimate remaining life.

The U.S. Department of Energy similarly distinguishes reactive, preventive, predictive and reliability-centered maintenance, with reliability-centered maintenance selecting failure-management strategies according to the system’s reliability characteristics and operating context.

But predictive maintenance has a limitation

Prediction is not magic. A sensor can fail. A model can be wrong. Historical data may not represent future conditions. An anomaly may not correspond to an imminent failure. A false alarm can trigger unnecessary maintenance. A missed alarm can be much more dangerous. NASA explicitly recognizes that fault-detection systems themselves must tolerate false alarms and that not every fault can necessarily be detected or recovered from in time.

Therefore:

A predictive maintenance system must itself be engineered as a safety-critical system when its decisions affect safety.

The Hidden Problem of Common-Cause Failure

One of the biggest weaknesses in simplistic reliability thinking is the assumption that failures are independent. They often are not. Consider two redundant cooling pumps. If Pump A fails independently, Pump B may continue operating. But what if both pumps lose cooling because:

  • the same electrical bus fails;
  • the same control signal is corrupted;
  • the same contaminated fluid enters both;
  • the same pipe ruptures upstream;
  • the same software command disables both;
  • the same environmental condition overheats both?

Then the apparent redundancy was largely an illusion.

A simplified reliability architecture illustrates the difference:

ArchitectureFailure characteristic
Single componentOne failure can cause system failure
Parallel independent componentsOne failure may be tolerated
Parallel components with shared powerShared failure can defeat redundancy
Diverse redundant systemsLess vulnerable to identical failure mechanisms
Redundant + isolated + monitoredMuch stronger failure-management architecture

This is why good engineering asks not merely: “Do we have a backup?”

but: “What can destroy the primary and backup at the same time?”

NASA’s safety guidance emphasizes isolation, segregation and consideration of common-cause failures when establishing failure tolerance.

Designing the Failure Path, Not Just the Operating Path

This may be the most important idea in the entire subject. Traditional design often concentrates on the nominal operating path:

Input → Component → Output

Failure-oriented engineering adds another diagram:

Fault → Detection → Isolation → Degradation → Safe State → Repair

Both paths need engineering.

A system that performs brilliantly under normal conditions but behaves unpredictably after a fault is not necessarily a good system.

A failure-management design matrix

Failure eventImmediate consequenceDetectionContainmentRecovery
Bearing degradationIncreased vibrationVibration sensorLoad/speed reductionBearing replacement
Hydraulic leakPressure lossPressure/flow monitoringIsolation valveRepair/refill
Structural crackReduced residual strengthInspection/SHMLoad-path redundancyRepair/replacement
Motor overheatingThermal damageTemperature sensorAutomatic shutdownCooling/repair
Sensor failureIncorrect informationCross-check/diagnosticsFault isolationSensor replacement
Power-supply failureLoss of subsystemVoltage monitoringBackup supplyPower restoration

This is where engineering becomes genuinely systemic. The component designer, controls engineer, maintenance engineer, software engineer, operator and safety engineer are no longer solving separate problems. They are designing one failure-response architecture. NASA’s systems-engineering framework emphasizes this life-cycle, multidisciplinary view rather than treating design as an isolated calculation exercise.

Rethinking What “Reliability” Really Means

Reliability is often reduced to a single question: How long does the machine operate without failing?

That is useful, but incomplete. Suppose Machine A operates for 10,000 hours and then fails catastrophically. Machine B operates for 8,000 hours, develops a detectable fault, automatically reduces its load, remains stable long enough for scheduled maintenance and never produces a dangerous secondary failure.

Which machine is better engineered? If reliability is defined only as uninterrupted operating time, Machine A appears superior. If reliability is understood as successful delivery of the required function under realistic conditions, Machine B may be the better system.

This suggests a broader engineering scorecard:

PropertyFundamental question
StrengthCan it withstand the load?
DurabilityHow does it deteriorate with use?
ReliabilityHow consistently does it perform its function?
MaintainabilityHow easily can it be restored?
DetectabilityCan developing faults be identified?
Fault toleranceCan it continue after specified failures?
Damage toleranceCan it safely tolerate existing damage?
RecoverabilityCan it return to operation after failure?
Graceful degradationDoes performance decline predictably?
SafetyAre unacceptable consequences prevented?
ResilienceCan the whole system absorb and recover from disruption?

This changes the meaning of engineering quality. A machine should not be judged solely by the absence of failure. It should also be judged by the quality of its response to failure.

The Economics of Designing for Failure

There is, however, an important limit. Failure management is not free. Redundancy adds components. Sensors add cost. Monitoring systems add software and electronics. Inspection requires labor. Damage-tolerant structures may require more sophisticated materials and geometries. Additional containment can increase weight. Predictive-maintenance infrastructure requires data, connectivity and analytical capability.

Therefore, designing for failure is not equivalent to adding every possible safety mechanism. Engineering is an optimization problem. A simplified conceptual objective can be written as:

C_design + C_maintenance + C_failure + C_downtime + C_risk

A design that minimizes initial manufacturing cost may create enormous downstream costs.

Conversely, a highly redundant system may be technically excellent but economically irrational if the consequences of failure are negligible.

The appropriate level of failure tolerance depends on:

  • consequence of failure;
  • probability of failure;
  • detectability;
  • time available for intervention;
  • operating environment;
  • repairability;
  • redundancy;
  • regulatory requirements;
  • life-cycle cost;
  • human exposure.

That is why engineering cannot simply ask:

“Can we make it more reliable?”

The better question is:

“What level of reliability and failure tolerance is justified by the consequences and operating context?”

The Deeper Lesson: Failure Is a System Property

Perhaps the greatest mistake is to treat failure as belonging exclusively to a component. A bearing does not exist in isolation.

Its failure depends partly on:

  • shaft alignment;
  • lubrication;
  • load;
  • temperature;
  • installation;
  • contamination;
  • vibration;
  • maintenance;
  • operating speed.

Likewise, a structural member does not have a meaningful failure behavior independent of its surrounding load paths. A sensor failure does not have the same consequence in a system with independent cross-checking as it does in a system that trusts the sensor completely. Failure therefore belongs to the system architecture, not merely to the failed part. This is why failure-mode analysis, fault trees, reliability analysis, fracture mechanics, structural health monitoring, condition monitoring and maintenance strategy should not be treated as disconnected specialist activities.

They are different ways of asking the same fundamental question:

How does the system move from normal operation toward failure, and where can engineering intervene?

From “Failure-Proof” to “Failure-Aware”

The language we use in engineering matters because it influences what designers look for. “Failure-proof” thinking encourages the search for a perfect component. “Failure-aware” thinking encourages the search for vulnerabilities.

A failure-aware engineer asks:

  • What can fail?
  • How can it fail?
  • How quickly can it fail?
  • What happens immediately afterward?
  • Can the failure be detected?
  • Can it be isolated?
  • Can the remaining system continue operating?
  • Does the system degrade predictably?
  • Can the operator understand what happened?
  • Can maintenance intervene before catastrophic escalation?
  • What happens if the backup fails too?
  • What common cause could defeat the redundancy?
  • What happens when the original assumptions are wrong?

These questions are harder than simply calculating a factor of safety. They are also much closer to the reality of engineering practice.

Conclusion: The Machine That Knows How to Fail

The best machines are not necessarily those that never fail. In the real world, no material, component, sensor, bearing, structure, control system or manufacturing process is perfectly immune to deterioration. The difference between a mediocre system and an exceptional one is often what happens after the first abnormal condition appears. A poorly engineered machine can turn a small defect into a chain reaction.

A well-engineered machine can:

Detect → Isolate → Tolerate → Degrade → Warn → Shut Down Safely → Recover

That is not an admission of engineering failure. It is an acknowledgment of engineering reality. Modern damage-tolerance practice, fail-safe structures, redundancy, fracture mechanics, fault detection, condition monitoring and reliability-centered maintenance all point toward the same principle: designing for failure is part of designing for safety. The goal is not to pretend that failure can be eliminated. The goal is to ensure that when failure inevitably enters the system, it does not automatically become disaster. That distinction separates failure prevention from failure management.

And ultimately, it suggests a more demanding definition of engineering excellence: A great machine is not one that has been designed as though failure is impossible. A great machine is one that has been designed with enough intelligence, tolerance and foresight to remain safe when failure becomes possible.

The strongest design is therefore not necessarily the one with the greatest strength. It is the one with the best understanding of what happens when strength, assumptions or components eventually run out.

Key research references

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top