
Common-cause failure occurs when one cause disables several components that appear to be independent. It reminds us that the number of backups is not the same as the independence of those backups.
Imagine a basement fitted with two drainage pumps. If one breaks, the other can take over, so the arrangement appears redundant. But if both pumps use the same electrical circuit, a flood that cuts power at the switchboard can stop them together. The issue is not two coincidental pump failures. A shared power supply has tied the supposedly separate protection paths back together.
This is related to, but not identical with, a single point of failure. A single point of failure means one component can fail and leave no alternative path. Common-cause failure shows that several alternatives may still be struck by the same environment, design defect or maintenance mistake. It is also more specific than ordinary correlation: we need to identify a pathway capable of producing the joint failure.
So checking redundancy means more than counting copies. We should ask whether they share a location, power source, software, supplier or operating procedure. Reliable redundancy is not merely adding another unit; it requires enough separation between critical failure paths.
Discover more from Geoffrey Chen
Subscribe to get the latest posts sent to your email.