Suppose you inspect a machine one component at a time.

The sensors pass their tests. The software follows its specification. The valves open when commanded. The operators are trained. The backup system is available. The maintenance records look normal.

Then the machine fails.

The first instinct is to search for the broken part. Which component was defective? Who made the wrong decision? Which procedure was violated?

Sometimes that search works. A bearing cracks. A wire shorts. A person skips a required step.

But complex systems contain another possibility: every part can behave in a way that seems reasonable on its own while their interaction creates a state nobody intended.

The failure is not hidden inside a component.

It is hidden in the arrangement.

If every part can be working, where does the failure live?

The Hidden Assumption

We often assume that a safe system is a collection of safe parts.

That sounds almost unavoidable. If the components are reliable, the people competent, and the procedures correct, what else is there?

The answer is the relationship between them.

A traffic network is more than a collection of roads. A hospital is more than a collection of clinicians and devices. A power grid is more than a collection of generators and transmission lines. A spacecraft is more than a collection of tested subsystems.

At system level, new properties appear: congestion, feedback, timing, dependency, bottlenecks, coordination, escalation, and the possibility that one locally sensible action changes the environment for everyone else.

Systems engineer Nancy Leveson makes this distinction explicit in her work on system safety. In a component-failure accident, something breaks. In a component-interaction accident, the components may remain individually functional while their combined behavior becomes unsafe.

That changes the question investigators need to ask.

Instead of only asking, “What failed?”, they also have to ask, “What relationships allowed correct-looking behavior to become dangerous?”

Nancy G. Leveson, Engineering a Safer World: Systems Thinking Applied to Safety · MIT Press

A reliable component can participate in an unsafe system

Reliability and safety are related, but they are not the same property.

A component is reliable when it performs its intended function consistently. A system is safe when its behavior remains within acceptable constraints.

Those can diverge.

A sensor can accurately report a signal, software can process that signal according to its logic, and an actuator can faithfully carry out the resulting command. If the meaning of the signal changes with context, the whole chain can still produce the wrong outcome.

Leveson uses examples of this kind to argue that increasing component reliability alone cannot eliminate accidents caused by unsafe interactions.

This is counterintuitive because engineering culture often rewards decomposition. Break a large problem into smaller pieces. Assign requirements. Verify each piece. Reassemble them.

Decomposition is essential. But a system can acquire behaviors during reassembly that were not visible while the parts were isolated.

The test report for every component can be green while the system-level assumption connecting them is wrong.

Leveson, Engineering a Safer World · MIT Press

Interfaces are where facts change meaning

The loss of NASA's Mars Polar Lander in 1999 is a useful example because the suspected failure did not require an exotic component breakdown.

The lander's legs were designed to produce touchdown signals when they contacted the surface. But deploying the legs could also generate brief false signals.

NASA's later investigation concluded that the most probable failure was premature shutdown of the descent engines after software treated a spurious signal as if touchdown had already occurred. The engines may have shut down roughly forty metres above the surface.

The troubling part is not simply that a sensor produced noise. The system already had reason to expect transient signals during deployment.

The failure crossed interfaces: mechanical movement produced an electrical signal; the signal reached software; the software carried state forward; a landing-control decision was made from that state.

The review also found that a wiring error interfered with an earlier system test and that the full test was not repeated after the wiring was corrected.

No single sentence captures the whole failure. “The sensor failed” is too simple. “The software failed” is too simple. “Testing failed” is too simple.

The accident emerged from how those pieces met.

NASA Lessons Learned · Mars Polar Lander premature shutdown

NASA Science · Mars Polar Lander / Deep Space 2

Success can teach the wrong lesson

Complex systems do not only learn from failure. They learn from survival.

That can be dangerous.

Suppose a team notices an anomaly. The system survives. The anomaly appears again. The system survives again. Each successful outcome can make the anomaly feel less alarming, even if the underlying risk has not disappeared.

NASA's Columbia accident investigation found that foam had detached from the external tank on earlier Shuttle flights. Because previous strikes had not destroyed an orbiter, the organization increasingly treated foam shedding as an accepted maintenance problem rather than a potential loss-of-vehicle hazard.

The Columbia Accident Investigation Board did not describe the disaster as a piece of foam plus bad luck. It concluded that management practices were as much a cause of the accident as the physical strike itself.

This is a strange feature of risk: repeated success can reduce fear without reducing danger.

If a system operates outside its intended safety margin and survives, the absence of catastrophe can be mistaken for evidence that the margin was unnecessarily conservative.

The system has not necessarily become safer.

The organization may simply have become more comfortable with the risk.

Columbia Accident Investigation Board Report, Volume I · NASA

NASA · Columbia Accident Investigation Board Synopsis

A bad decision can be produced by reasonable local decisions

Organizations divide responsibility because no one person can understand every part of a complex operation.

That division creates expertise. It also creates boundaries.

An engineer sees a technical uncertainty. A manager sees a schedule commitment. A finance team sees rising cost. An operations group sees a queue of delayed work. A regulator sees compliance categories. Each perspective can be legitimate.

The danger appears when local goals are optimized independently and nobody owns the combined risk.

The Challenger investigation found that the launch decision on January 28, 1986 was flawed. Engineers at Morton Thiokol had raised concerns about O-ring performance in the unusually cold conditions and initially recommended against launch below 53 degrees Fahrenheit. The launch proceeded in much colder weather.

It would be easy to reduce that history to one person ignoring one warning. The Commission instead examined a decision process in which information, authority, and interpretation moved through multiple organizational layers.

Complex failures often look obvious in retrospect because the final sequence has already selected which facts matter.

Before the event, those facts compete with hundreds of other concerns.

That is why good system design cannot depend on the hope that one unusually perceptive person will always recognize the decisive signal in time.

Report of the Presidential Commission on the Space Shuttle Challenger Accident, Chapter V · NASA

Rogers Commission, Chapter VI · NASA

Efficiency can remove the space that catches mistakes

An efficient system tries to eliminate waste: idle capacity, waiting time, duplicated resources, unused inventory, redundant steps.

Often that is exactly what we want.

But some apparent waste is actually margin.

A hospital bed that is empty today may be capacity for tomorrow's surge. A second communication path may look unnecessary until the first fails. Extra time in a schedule may look inefficient until an unexpected inspection finds something important.

Research on large networked systems describes a related tension. NIST work on systemic risk notes that economic pressure can push systems toward fuller utilization while increasing interconnection so resources can be shared more efficiently. Those changes can improve ordinary performance while also creating conditions for abrupt system-wide instability.

This does not mean efficiency is bad. It means efficiency is usually measured under a model of normal operation.

Resilience asks a different question:

What happens when normal assumptions stop being true?

A system optimized for the most likely day may have less room to absorb the unlikely one.

NIST · Towards Systemic Risk Aware Engineering of Large-Scale Networks

Dependencies turn small failures into large ones

A component can fail quietly when nothing important depends on it.

The same component can become critical when it sits upstream of several other systems.

Electricity powers water pumps. Communications depend on electricity. Emergency response depends on communications. Hospitals depend on all three.

NIST's community resilience work describes cascading failure as a process in which one failure triggers failures in other components or systems. The 2003 Northeast blackout, for example, disrupted transportation and communications in addition to electric service.

The important property is not only the strength of each network. It is the pattern of dependency between networks.

That is why mapping components without mapping dependencies can produce a false sense of completeness.

Two systems can each meet their own reliability targets and still share a vulnerability neither target captures.

The larger the web, the harder it becomes to see where one local loss will reappear as someone else's critical input.

NIST Disaster Resilience Framework · Infrastructure dependencies and cascading failures

NIST · Design Methods for Resilient Community Systems

Redundancy only helps when it is genuinely independent

Engineers often respond to risk by adding backups.

Two pumps instead of one. Two servers. Two sensors. Two people checking the same decision.

Redundancy is powerful, but it contains an assumption: the backups must not fail for the same reason.

Two servers in the same flooded room are not independent. Two teams using the same incorrect data are not independent. Two algorithms trained on the same biased labels may agree because they inherited the same error.

The number of backups therefore matters less than the structure of their dependence.

This is another reason complex systems cannot be understood by counting components alone.

A diagram may show three separate boxes while the real world contains one shared power supply, one shared vendor, one shared mental model, or one shared deadline.

Visible redundancy can hide invisible common cause.

Human error is often the end of the explanation too early

After a failure, the sentence “the operator made a mistake” is emotionally satisfying.

It gives the event a location.

But people inside complex systems make decisions using the information, incentives, interfaces, time pressure, and procedures available to them at that moment.

A retrospective observer sees the outcome. The operator did not.

This does not eliminate personal responsibility. Negligence, recklessness, and deliberate rule-breaking are real.

It does mean that labeling an action “human error” may describe the last visible step without explaining why the system made that error likely, difficult to detect, or catastrophic.

If several competent people repeatedly make the same mistake, replacing the people may leave the mechanism untouched.

A stronger investigation asks what made the wrong action look reasonable at the time, what evidence was missing, what feedback arrived too late, and why the consequence was not contained.

Root cause is useful — and sometimes misleading

The phrase root cause suggests that an accident is a tree with one deepest root.

Sometimes it is.

A defective component can be traced to a manufacturing problem. A calculation error can be corrected. A missing procedure can be added.

But in tightly connected sociotechnical systems, causal structure can look more like a network than a chain.

Software changes operator behavior. Operator workarounds change what managers observe. Management incentives change maintenance choices. Maintenance choices change equipment condition. Equipment behavior changes software assumptions.

Leveson's systems-theoretic approach was developed partly because linear event-chain models can miss these feedback relationships.

The practical risk of a single-root story is that it invites a single-root fix.

Replace the component. Retrain the person. Add a warning.

Those actions may help. But if the accident emerged from the control structure around them, the system can produce a different failure through the same relationships.

Nancy G. Leveson · An Introduction to System Safety Engineering · MIT Press

Resilience begins by assuming something will go wrong

A fragile design asks how to prevent every failure.

A resilient design also asks what happens after prevention fails.

Can the problem be isolated?

Can the system degrade gradually instead of collapsing suddenly?

Can operators see what state the system is actually in?

Can a decision be reversed before its consequences spread?

Is there enough spare capacity to absorb a surprise?

NIST's resilience guidance points to several concrete strategies for reducing cascading effects, including redundancy, added capacity, and mechanisms that can isolate parts of a system rather than allowing every disturbance to propagate.

The deeper principle is containment.

You do not need to predict every future failure if the system is designed so that many failures remain local.

That is a different philosophy from perfection. It accepts that components, people, forecasts, and assumptions will sometimes be wrong.

The design goal becomes preventing one error from acquiring the power of the whole system.

NIST Disaster Resilience Framework · Dependencies and mitigation

Where the Fields Collide

Systems engineering tells us that safety can be an emergent property rather than a component property.

Human factors reminds us that people act inside interfaces, incentives, workloads, and incomplete information.

Organizational research shows how repeated success can normalize risk and how authority can separate technical knowledge from decision power.

Infrastructure research shows how dependencies allow failure to travel across systems that appear separate on an organizational chart.

Safety engineering adds one more uncomfortable idea: a system can become more reliable in ordinary operation while becoming more dangerous under unusual conditions.

Together, these fields point toward a different definition of failure.

Failure is not always the moment when a part breaks.

Sometimes it is the moment when relationships that were always present finally line up badly enough to become visible.

What We Know — and What We Don't

We know that complex systems can fail through interactions among components, not only through isolated component defects.

We know that tightly connected dependencies can produce cascading effects across infrastructure and organizations.

We know from major accident investigations that communication, schedule pressure, testing gaps, organizational structure, and normalized anomalies can matter as much as the immediate physical trigger.

We also know that redundancy, spare capacity, isolation, and better feedback can improve resilience.

What we do not know is how to eliminate complex-system failure entirely.

More monitoring can create alarm overload. More redundancy can create new interactions. More procedures can create new coordination burdens. More automation can remove one class of error while introducing another.

Complexity does not make safety impossible. It makes simple guarantees less credible.

The relevant question is not whether a system has risks.

It is whether the system can see, absorb, and contain risks before they become system-wide.

Back to the Green Check Marks

Return to the inspection.

The sensors pass. The software passes. The valves pass. The operators are qualified. The backup is online.

You could stop there and conclude that the system is ready.

Or you could ask a second set of questions.

What assumptions do these parts share?

Which signals change meaning as they cross an interface?

What happens if two normal events occur at the same time?

Which local goal can quietly increase global risk?

What failure would all the backups share?

Where is there no slack?

And if one part becomes wrong, how far can the wrongness travel?

Those questions are harder because they do not belong to any one component.

That is exactly why they belong to the system.

The Next Question

Complex systems blur responsibility because outcomes are produced by many interacting contributions.

Artificial intelligence creates a similar problem in a different form. A person provides an idea. A model generates possibilities. A platform supplies the model. Training data shaped what the model could produce. A human selects, edits, and publishes the result.

When creation becomes a system rather than a single act, ownership becomes harder to locate.

Who owns something created by a human and an AI together?

Sources & Further Reading

  1. Nancy G. Leveson, Engineering a Safer World: Systems Thinking Applied to Safety · MIT Press
  2. NASA Lessons Learned · Mars Polar Lander premature shutdown
  3. NASA Science · Mars Polar Lander / Deep Space 2
  4. Columbia Accident Investigation Board Report, Volume I · NASA
  5. NASA · Columbia Accident Investigation Board Synopsis
  6. Report of the Presidential Commission on the Space Shuttle Challenger Accident, Chapter V · NASA
  7. Rogers Commission, Chapter VI · NASA
  8. NIST · Towards Systemic Risk Aware Engineering of Large-Scale Networks
  9. NIST Disaster Resilience Framework · Infrastructure dependencies and cascading failures
  10. NIST · Design Methods for Resilient Community Systems
  11. Nancy G. Leveson · An Introduction to System Safety Engineering · MIT Press

Beyond the Question is an interdisciplinary series by Arin Vale.

Read the editorial approach

THE NEXT QUESTION

Who owns something created by a human and an AI together?

Continue to No. 015