Interpretation of the results of clinical trials involving medicines
Design hierarchy - and what each can and cannot answer
| Design | Answers | Cannot |
|---|---|---|
| RCT | Causation; randomisation balances known AND unknown confounders | Rare/late harms; external validity |
| Cohort | Incidence, prognosis, harm | Residual confounding |
| Case-control | Rare outcomes, efficient | Recall and selection bias; no incidence |
| Cross-sectional | Prevalence | Temporality |
| Meta-analysis | Precision, consistency | Cannot fix biased primary studies - "garbage in, garbage out" |
Trial framings
- Superiority - is A better than B?
- Non-inferiority - is A no worse than B by more than a pre-specified margin delta?
- The margin is a judgement, chosen by the investigators - interrogate it
- Per-protocol analysis is the conservative analysis here (ITT biases toward non-inferiority by diluting differences)
- Equivalence - within delta in both directions
- Pragmatic vs explanatory - real-world effectiveness vs idealised efficacy
- Adaptive / platform (e.g. RECOVERY, REMAP-CAP) - pre-specified rules allow arms to be added or dropped
Measures of effect - and which one is being hidden
| Measure | Meaning |
|---|---|
| ARR (absolute risk reduction) | Control risk - treatment risk. The clinically meaningful number |
| RRR (relative risk reduction) | ARR / control risk. Constant across baseline risks -> always looks impressive |
| NNT | 1 / ARR. Quote with the time horizon |
| NNH | 1 / absolute risk increase |
| HR | Instantaneous relative rate over time; assumes proportional hazards |
| OR | Odds ratio; approximates RR only when the outcome is rare (<10%) - otherwise exaggerates |
- *The commonest manipulation is reporting relative benefit and absolute harm*
- "38% reduction in events" with an ARR of 1.2% -> NNT 83
- Baseline risk drives absolute benefit - the same RRR gives a very different NNT in a low-risk population
Statistical significance vs importance
- p value = probability of data this extreme if the null were true
- Not the probability the null is true; not the size of the effect; not clinical importance
- Confidence interval is more informative than the p value
- Width -> precision; the boundary nearest the null answers "could the true effect be trivial?"; the far boundary answers "could it be large?"
- A non-significant result is not evidence of no effect - absence of evidence vs evidence of absence
- Fragility index - how many outcome events would need to switch to lose significance (often 1-3 in "positive" trials)
Bias - systematic, not fixed by a larger sample
How a trial produces a wrong answer.
- Selection bias - allocation not concealed; concealment (before randomisation) differs from blinding (after)
- Performance bias - differential co-intervention
- Detection bias - unblinded, subjective outcome assessment
- Attrition bias - differential loss to follow-up; check the CONSORT flow diagram
- Reporting bias - outcome switching (compare the paper with the registered protocol), publication bias (funnel plot asymmetry)
- Immortal time bias, lead-time and length-time bias (screening trials)
- Confounding by indication - the dominant problem in observational drug studies
Chance
- Type I (alpha) - false positive; inflated by multiple comparisons, subgroup analyses and repeated interim looks
- Type II (beta) - false negative; power = 1 - beta, conventionally 80-90%
- Underpowered trials both miss real effects and, when positive, exaggerate effect size (winner's curse)
Design choices that mislead
- Composite endpoints - driven by the softest, commonest component ("MACE" driven by revascularisation, not death)
- Surrogate endpoints - validated only if the treatment effect on the surrogate reliably predicts the clinical effect; CAST (antiarrhythmics suppressed ectopy and increased mortality) is the canonical failure
- Inappropriate comparator - placebo where an active standard exists; a suboptimal dose of the comparator
- Short follow-up, run-in periods that exclude non-adherers, highly selected populations
Appraisal sequence
1. PICO - population, intervention, comparator, outcome. Does it match my patient?
2. Was randomisation concealed? Was the allocation truly random?
3. Blinding - participants, clinicians, outcome assessors, analysts
4. Were the groups balanced at baseline? (Table 1; a p value in Table 1 is meaningless in a randomised trial)
5. Was the primary outcome pre-specified and registered? - check the trial registry
6. Intention-to-treat analysis? - preserves randomisation; the default for superiority
7. Was follow-up complete? - >20% loss threatens validity
8. Effect size with confidence interval, and the absolute numbers
9. Harms reported as thoroughly as benefits?
10. Funding and conflicts; who held the data and who wrote the paper?
11. Was the trial stopped early for benefit? - early stopping systematically overestimates effect
Subgroup analyses - the standard exam trap
- Believe a subgroup only if: pre-specified, few in number, a significant test for interaction (not merely significant within the subgroup), biologically plausible, and consistent across trials
- ISIS-2 famously showed no aspirin benefit in Gemini and Libra - to make the point
Meta-analysis
- Heterogeneity: I^2 (>50% substantial), Cochrane Q, prediction interval
- Fixed effect (one true effect) vs random effects (distribution of effects; wider CI)
- Risk-of-bias assessment (RoB 2, GRADE), sensitivity analysis, funnel plot
Applying a result to the patient in front of you
Applying a result to the patient in front of you.
- Is my patient like the trial population? Age, comorbidity, renal function, ethnicity, disease severity
- Australian practice: check whether the trial population resembles the Australian population and whether the comparator is the Australian standard of care
- What is my patient's baseline risk?
- Higher baseline risk -> lower NNT for the same RRR; low-risk patients may get almost no absolute benefit while carrying the same harm
- Are the outcomes ones the patient cares about? Mortality, function, symptoms - not a surrogate
- Weigh NNT against NNH over the same time horizon
- Feasibility and cost - is the drug PBS-listed, is the monitoring achievable?
- Shared decision-making: present absolute numbers, natural frequencies ("3 in 100 over 5 years"), and both benefit and harm
- GRADE the certainty of evidence: high / moderate / low / very low, downgraded for risk of bias, inconsistency, indirectness, imprecision and publication bias
- A guideline recommendation is a value judgement layered on evidence - separate the two
Reporting standards and registration
- CONSORT (RCT reporting), PRISMA (systematic reviews), STROBE (observational), SPIRIT (protocols)
- Trial registration - ANZCTR, ClinicalTrials.gov; the defence against outcome switching
- GRADE, RoB 2, Cochrane
- Human Research Ethics Committees, clinical equipoise, informed consent
- Data and Safety Monitoring Boards, pre-specified stopping rules (O'Brien-Fleming boundaries)
- Industry funding and sponsorship bias
- PBAC (Pharmaceutical Benefits Advisory Committee) - cost-effectiveness, ICER, QALY - the Australian translation step from evidence to access
- Real-world evidence, registries, target trial emulation
- Bradford Hill criteria - for causal inference from observational data
Why two trials of the same intervention disagree
Usually the trials were not asking the same question.
- Different populations - selection criteria, disease definition (e.g. what counted as "cryptogenic" stroke in the PFO closure trials)
- Different intervention - device generation, dose, timing, operator experience
- Different background therapy - the comparator improved between trials
- Different outcomes or follow-up duration - a benefit that takes 5 years to emerge is invisible in a 2-year trial
- Chance - especially with small event numbers
- Bias in one of them
- Discordance is far more often explained by these than by a genuinely different treatment effect
Evidence over time
- Early small positive trials regress toward the null as larger trials appear (the Proteus phenomenon)
- Effects seen on surrogates frequently do not survive hard-endpoint testing
- Meta-analyses are updated, and conclusions reverse - check the date
- The registrar's habit worth keeping: read the trial that changed practice, not the summary of it
13 more sections, plus exam facts
Premium unlocks every note across every specialty, and the full exam fact library behind it.
Get premium access