The plane that made it back isn't the whole story. Neither is the customer who left a five-star review, the CEO who dropped out of college, or the diet that "worked" for everyone who stuck with it. Every one of these stories has a hidden flaw, and that flaw usually gets labeled "survivorship bias" — even when it's actually something else entirely: selection bias.

These two terms get thrown around interchangeably in articles, podcasts, and business meetings. That's a problem, because they aren't the same thing. One is a specific type of the other. Mixing them up doesn't just muddy your vocabulary — it can make you draw the wrong conclusion from your own data.

This article breaks the two apart for good, with four examples that make the distinction stick.


What Is Selection Bias?

Selection bias is the umbrella term. It happens whenever the group of people, items, or data points you're studying isn't randomly chosen — meaning it doesn't fairly represent the larger population you're trying to understand.

Anytime your sample got filtered by some non-random process before you looked at it, selection bias is in play. That filter could be almost anything: who chose to respond to a survey, who a hospital happened to admit, who applied for a program, or who simply had access to the internet when you posted your poll.


What Is Survivorship Bias?

Survivorship bias is one specific flavor of selection bias. It happens when your sample only includes things that "survived" some process — while the things that didn't survive quietly disappeared from view.

The key word is disappeared. In survivorship bias, the missing data isn't just underrepresented — it's often gone entirely, destroyed, shut down, or otherwise removed from the record. You're not just looking at a skewed slice of the population; you're looking at the winners, with the losers erased.


The Core Difference, in One Line

Selection bias is about who or what got included in the first place. Survivorship bias is about who or what stuck around long enough to still be visible. Every survivorship bias is a selection bias — but not every selection bias involves survival at all.

Think of it this way: selection bias can happen at the starting line. Survivorship bias only shows up at the finish line, after some silent elimination has already occurred.


Example 1: The WWII Bomber Paradox (Survivorship Bias)


This is the example that made survivorship bias famous, and it's worth revisiting because it's so precise.

During World War II, engineers studied returning bombers to decide where to add armor. They mapped the bullet holes and found the wings and tail were hit most often — so the obvious move was to reinforce those areas.

Statistician Abraham Wald pointed out the flaw. The planes being studied were the ones that made it home. The bombers hit in the engine or cockpit never returned at all — they went down. The armor didn't belong where the surviving planes had holes; it belonged where the surviving planes had none, because a hit there meant the plane never survived to be counted.

This is survivorship bias in its purest form. The missing data — the destroyed planes — vanished completely from the sample. Nobody excluded them on purpose; they simply didn't survive to be observed.


Example 2: Mutual Fund Performance Reports (Survivorship Bias)


Financial firms love to publish glowing average returns for their mutual funds. What they often don't mention is that poorly performing funds get quietly shut down or merged into other funds every year.

So when a company reports "our funds averaged a 9% return over the last decade," that number typically reflects only the funds that survived the full ten years. The funds that lost money and got closed early are removed from the dataset entirely — they're not in the average because they no longer exist to be measured.

This inflates performance numbers industry-wide, a well-documented phenomenon researchers call "survivorship bias in mutual fund returns." An investor comparing today's fund lineup to historical averages is comparing apples to a cherry-picked basket of winners.


Example 3: Five-Star Product Reviews (Selection Bias)


Now shift to a case where nothing "died" or disappeared — the sample was just skewed from the start.

Most people who buy a product never leave a review at all. The ones who do tend to fall into two camps: people who loved it enough to say so, and people who were angry enough to vent. The large, quiet middle — people who thought the product was fine, forgettable, or mildly disappointing — rarely bothers typing a review.

That means the reviews you read online are not a random sample of buyers. They're a self-selected group with unusually strong opinions in either direction. This is classic selection bias: the population (all buyers) and the sample (people who chose to review) were never the same to begin with, and no elimination process was involved — people simply opted in or out.


Example 4: Gym-Based Health Surveys (Selection Bias)


Imagine a fitness researcher wants to know how often "the average person" exercises, so they hand out a survey at a gym.

The problem is obvious once you say it out loud: everyone in that gym already chose to be there. People who never exercise, who can't afford a membership, or who have mobility limitations are structurally absent from the sample. The resulting data will suggest people exercise far more often than they actually do — not because anything "survived," but because the sampling location itself excluded huge chunks of the real population from ever being asked.

This is selection bias operating purely at the collection stage. There's no attrition, no dropout, no disappearance over time — just a starting sample that was never representative.


Why This Distinction Actually Matters

It's tempting to treat this as a semantic nitpick, but the difference changes how you fix the problem.

If you're dealing with selection bias, the fix usually starts at data collection: randomize your sample, widen where you're recruiting from, or weight your results to match the true population.

If you're dealing with survivorship bias specifically, the fix requires reconstructing what's missing: tracking down the failed startups, the closed funds, the planes that didn't return. You have to actively go looking for the absence, because it won't show up in the data on its own.

Calling a survivorship bias problem a "selection bias" isn't technically wrong — it's just too vague to point you toward the right solution. And calling a plain selection bias problem "survivorship bias" can send you hunting for missing data that was never there to begin with; the sample was just narrow from day one.


Frequently Asked Questions


Is survivorship bias a type of selection bias?

Yes. Survivorship bias is a specific subcategory of selection bias, defined by the fact that non-surviving items are removed or excluded from the visible dataset over time.


Can selection bias happen without survivorship being involved?

Absolutely. Any time a sample is gathered in a way that isn't random — through self-selection, convenience sampling, or biased recruitment — selection bias exists, even if nothing was eliminated or "died off" along the way.


What's a simple way to remember the difference?

Selection bias asks, "Who got into the sample in the first place?" Survivorship bias asks, "Who's still around to be counted?" If the missing data disappeared over time, think survivorship. If it was never included to begin with, think selection.


Why do people confuse these two terms so often?

Because survivorship bias is the more famous, more quotable concept — largely thanks to the WWII bomber story — people default to it as shorthand for almost any biased sample, even when the actual mechanism is different.


How can businesses avoid both types of bias in their data?

Actively seek out missing groups: churned customers, failed projects, non-respondents, and shut-down initiatives. If a dataset only shows current winners, ask what happened to everyone who isn't in it anymore, and design data collection to include a genuinely representative slice of the population from the start.


Final Thoughts

Survivorship bias and selection bias aren't rival concepts — they're a family relationship. Selection bias is the broad category; survivorship bias is the specific case where the missing pieces vanished along the way rather than being excluded from the outset.

Once you can tell them apart, you start spotting them everywhere: in success-story books that only interview the winners, in app store ratings, in "top performer" case studies, in any claim built on a sample you didn't choose yourself. The habit worth building isn't just naming the bias correctly — it's asking, every time, who or what isn't in the room, and why.