Basic statistical concepts - bias
Definition
- Systematic deviation of a result from the truth
- Contrast random error, which is reduced by increasing sample size. Bias is not - a larger biased study is simply more precisely wrong
Three families
| Family | When it arises | Examples |
|---|---|---|
| Selection bias | Choosing who enters or stays in the study | Healthy worker, volunteer, referral, non-response, attrition/loss to follow-up, survivor bias, Berkson bias |
| Information (measurement) bias | Measuring exposure or outcome | Recall, interviewer, observer, misclassification, detection, Hawthorne effect |
| Publication / reporting bias | What gets published | Positive results, outcome switching, language bias, duplicate publication |
Where bias shows up
- Recall bias is the characteristic weakness of case-control studies; loss to follow-up is the characteristic weakness of cohort studies
- Publication bias: positive trials are published faster, more often, and more prominently - the basis for funnel plot asymmetry testing in meta-analysis
- Attrition of >20% seriously threatens the validity of a randomised trial
The screening biases - the most examined group
| Bias | Mechanism | Effect |
|---|---|---|
| Lead-time bias | Screening advances the date of diagnosis without changing the date of death | Apparent survival from diagnosis lengthens; lifespan unchanged |
| Length-time bias | Screening preferentially detects slow-growing disease with a long asymptomatic detectable phase | Screen-detected cases look more indolent and appear to have better prognosis, even with no true benefit of early detection or treatment |
| Overdiagnosis | The extreme of length-time bias: detection of disease that would never have caused symptoms or death | Inflates incidence and apparent survival; causes real treatment harm |
| Healthy volunteer bias | People who attend screening are healthier | Favours the screened group |
- -> the only unbiased endpoint for a screening programme is disease-specific (and ideally all-cause) MORTALITY in a randomised comparison - never 5-year survival
- The neuroblastoma and thyroid cancer screening experiences: incidence rose sharply, survival appeared excellent, mortality did not change at all
Misclassification
- Non-differential (random, unrelated to the other variable) -> biases toward the null - dilutes a true effect
- Differential (related to exposure or outcome status) -> biases in either direction, unpredictably. The dangerous kind
Other named biases worth knowing
- Immortal time bias - a period during which the outcome cannot occur is misallocated to the treated group (the classic artefact making many drugs look protective in database studies)
- Berkson bias - hospital-based controls are not representative of the population
- Neyman (prevalence-incidence) bias - fatal or rapidly resolving cases are missed by a prevalent-case study
- Confirmation and expectation bias - unblinded assessors
- Regression to the mean - extreme values naturally move toward the average on retesting; the reason uncontrolled before-after studies of any intervention started at a symptom peak look effective
Appraisal questions
- Selection: How were participants chosen? Who was excluded and why? What proportion was lost to follow-up, and did losses differ between arms?
- Information: Were exposure and outcome measured the same way in every group? Was the assessor blinded? Are the outcomes objective or subjective?
- Publication: Was the trial prospectively registered? Does the reported primary outcome match the registered one? Is there a funnel plot?
- Direction: In which direction would this bias push the result? A study biased against the finding it reports is more credible, not less
Prevention by design
Prevention - by design, because bias cannot be adjusted away afterwards
- Randomisation with concealed allocation - selection bias at entry
- Blinding of participants, clinicians and especially outcome assessors - information bias
- Objective, pre-specified, standardised outcome definitions and identical follow-up in every arm
- Minimise and account for loss to follow-up; intention-to-treat analysis with sensitivity analyses for missing data (best-case/worst-case imputation)
- Prospective trial registration and publication of protocols and negative results
- Population-based sampling and appropriate control selection in observational studies
- Use routinely collected or pre-existing records rather than recall where possible
In evidence synthesis
- Cochrane Risk of Bias 2 for trials, ROBINS-I for non-randomised studies, QUADAS-2 for diagnostic accuracy
- Funnel plot and Egger test for publication bias; trim-and-fill sensitivity analysis
Related concepts
- Confounding (a distinct concept - bias is a flaw in the study, confounding is a feature of the exposure-outcome relationship)
- Internal versus external validity
- Random error, precision, confidence intervals
- Screening programme evaluation, overdiagnosis, sojourn time
- CONSORT, STROBE, PRISMA
- Risk of bias tools: RoB 2, ROBINS-I, QUADAS-2
Pitfalls
- Bias cannot be fixed in the analysis. If it was not designed out, it is in the result
- Increasing the sample size increases precision around a biased estimate - it makes the wrong answer more confident
- Five-year survival is the most abused statistic in oncology screening debates - it is inflated by lead-time and length-time bias even when the programme saves no lives
- An uncontrolled before-after study is uninterpretable: regression to the mean, natural history and placebo effect all point the same way
- When a very large observational study contradicts a randomised trial, the trial is usually right
🔒
10 more sections, plus exam facts
Premium unlocks every note across every specialty, and the full exam fact library behind it.
Get premium access