philosophy

Explain it: What Is Survivorship Bias?

  • SHARE
Explain it

... like I'm 5 years old

Survivorship bias happens when we judge a situation by looking only at the people or things that made it through, while overlooking those that failed, disappeared, or were never recorded. Because survivors are easier to see, they can make success appear more common—and failure less important—than it really is.

Imagine reading interviews with five wealthy entrepreneurs. Each says that determination, early mornings, and bold risks produced their success. Their advice may be sincere, but thousands of equally determined entrepreneurs may have followed similar routines and lost their businesses. Their stories rarely become books, podcasts, or conference speeches.

The bias does not mean successful people have nothing useful to teach. It means their experiences represent only part of the evidence. Luck, timing, personal connections, starting resources, and simple randomness may separate visible winners from forgotten failures. This is one reason survivorship bias matters in behavioral economics, which examines how mental shortcuts affect decisions.

The same problem appears when someone praises an old product because “they don’t make things like they used to.” The durable objects from the past remain visible precisely because they lasted. The badly made ones may have broken and entered the trash years ago.

To resist the bias, pause whenever evidence seems unusually positive. Ask who began the journey, who vanished along the way, and why only certain outcomes remain visible.

Survivorship bias is like judging the quality of an entire basket of fruit by examining the few apples that remained fresh after every rotten one was quietly thrown away.

Explain it

... like I'm in College

During the Second World War, statistician Abraham Wald worked with the Statistical Research Group at Columbia University. Among its military problems was estimating aircraft vulnerability from damage observed on planes that returned from missions.

A simple reading of the evidence might treat the most frequently damaged areas as the greatest weaknesses. Wald’s analysis started from a more important observation: every aircraft being inspected had survived. Areas showing many hits were therefore places where damage was often compatible with returning. Areas showing fewer hits could be more vulnerable because aircraft struck there were less likely to appear in the dataset at all. His 1943 memoranda formally addressed estimating vulnerability from damage recorded on surviving planes.

This story reveals the mechanism behind survivorship bias. A selection process separates all original cases into visible survivors and missing non-survivors. If the likelihood of remaining visible is related to the outcome being studied, the surviving sample stops representing the original population.

The distortion commonly appears in:

  • Investment records that exclude funds that closed after poor performance
  • Workplace studies that survey only employees who stayed
  • Product testimonials collected from satisfied customers
  • Medical research that overlooks patients lost before observation
  • Career advice based only on prominent high achievers

Survivorship bias is therefore a form of selection bias. It often produces excessive optimism, but it can create other errors too. The direction depends on who disappears and why. The practical remedy is to define the original population, trace losses, examine unsuccessful cases, and test whether missing observations differ systematically from the survivors.

EXPLAIN IT with

Picture an adult building a Lego business district with 100 small companies. Every company begins as an identical gray platform. Over several rounds, builders add towers representing staff, investment, products, and revenue.

Most structures eventually collapse. Some receive weak foundations, some run out of bricks, and others are accidentally knocked over. Each failed model is immediately swept beneath the table. After an hour, only eight impressive towers remain in view.

A consultant enters after the cleanup. Seeing that every surviving tower contains a red window, the consultant announces, “Red windows create successful companies.” The conclusion sounds reasonable because the visible evidence is consistent. Yet dozens of collapsed buildings also had red windows. Those bricks could not prevent poor foundations, shortages, accidents, or bad designs.

The Lego table now represents a survivor-only dataset. The eight towers are observed cases, while the 92 removed structures are missing observations. The cleanup process is the selection mechanism. Because the consultant arrived after selection occurred, the visible sample gives a misleading account of which design choices mattered.

To investigate properly, the consultant would retrieve the discarded models, record when each collapsed, compare their designs, and reconstruct the original group of 100. If the failed models cannot be recovered, any conclusion should acknowledge that uncertainty.

Survivorship bias is not caused by the surviving towers lying. Their red windows are real. The mistake occurs when an observer treats the remaining towers as though they were the whole experiment. Good analysis begins by looking at what remains—but it becomes trustworthy only after asking what was removed, when it disappeared, and why.

Explain it

... like I'm an expert

Formally, survivorship bias arises when inference about a target population is based on observations conditioned on inclusion, persistence, detection, or survival. Let (S=1) denote observation. Analysts estimate a relationship within (P(X,Y\mid S=1)), although the quantity of interest belongs to (P(X,Y)) or another target distribution.

The resulting bias is not merely a small sample problem. Increasing the number of survivors can reduce sampling variance while leaving systematic selection error untouched. A vast, precisely measured survivor dataset may still support a confidently incorrect conclusion.

In a causal graph, conditioning on (S) becomes dangerous when selection is a collider, or a descendant of one, influenced by variables connected to both exposure and outcome. Restricting analysis to (S=1) can then open a noncausal path, induce associations absent from the target population, attenuate real effects, or reverse their apparent direction. Survivorship bias is consequently better understood as a data-generating and identification problem than as a purely psychological mistake.

Related manifestations include attrition bias, informative censoring, left truncation, publication bias, and the exclusion of delisted securities from historical financial databases. These mechanisms are not interchangeable, but all can make availability depend on variables relevant to inference.

Correction requires more than telling analysts to “remember the failures.” Depending on the design, useful approaches include reconstructing the sampling frame, retaining outcome data for departures, inverse-probability weighting, selection models, multiple imputation, bounds, and sensitivity analysis. Each method depends on assumptions about the selection mechanism and available covariates; missing non-survivors cannot always be recovered from survivor data alone.

The central expert question is therefore not simply, “Is this sample large?” It is: “What process determined entry into the sample, and does conditioning on that process alter the estimand?” This complements broader reasoning tools such as Hanlon’s Razor, which similarly encourages scrutiny before causal attribution.

  • SHARE