← Back to Blog

Error types

Type I and Type II Error

Statistics for clinical researchers and surgical trainees

In short

Type I error is a false positive: your analysis says a real difference exists when it doesn't. Type II error is a false negative: your analysis says nothing is going on when something real is. You already weigh this exact trade-off at the bedside every time you decide whether an equivocal finding is worth acting on. Alpha is your tolerance for a false positive, beta is your tolerance for a false negative, and tightening one loosens the other. Formalising that trade-off is all a hypothesis test does.

Residents learn Type I and Type II error as a 2x2 box on a slide, usually right before they forget it. That is a shame, because the underlying logic is not new to anyone who has ever stood at a bedside deciding whether an ambiguous finding is worth acting on. You have been reasoning about false positives and false negatives since your first year of training. Hypothesis testing just gives that reasoning a name and a number.

A decision you already make without a formula

Consider a patient two days post-laparotomy with a low-grade fever, mild tachycardia, and a slightly tender abdomen. Nothing is definitive. You have two ways to be wrong. You can act on the finding (image, culture, take the patient back to theatre) and discover nothing was wrong: a false alarm that cost the patient risk and cost the system resources. Or you can watch and wait, and the patient has an early anastomotic leak that declares itself two days later, sicker than it needed to be: a missed signal that cost the patient a delayed diagnosis.

Every clinician intuitively weighs which error is worse in that specific situation, and the weighing changes with context. A missed leak is usually worse than a negative re-look. A missed compartment syndrome is worse than an unnecessary fasciotomy. That asymmetry, deciding how much you are willing to tolerate one kind of wrong answer versus the other, is precisely what Type I and Type II error formalise in a statistical test. The bedside version has no name attached to it. The version in your results section does.

What alpha and beta actually mean in your output

A hypothesis test starts from a null hypothesis: no real difference between your groups. Type I error (denoted alpha) is rejecting that null when it is actually true, concluding a difference exists when it does not. It is the statistical equivalent of a false-positive diagnostic test. Type II error (denoted beta) is failing to reject the null when it is actually false, concluding no difference exists when one really does. It is the equivalent of a false-negative test that misses real disease. Power, which is 1 minus beta, is the test's sensitivity: its ability to detect a true effect when one is present.

That parallel is not loose. Sensitivity and specificity are the vocabulary you already use to judge a troponin assay or a screening test. A hypothesis test is doing the same job on a hypothesis instead of a patient: it is a diagnostic instrument for whether an observed difference reflects a real effect or noise, and it has the same two failure modes a diagnostic test has.

Convention sets alpha at .05 and targets beta at .20 (80% power), but these are conventions, not laws of nature, in the same way "positive if troponin exceeds 0.04 ng/mL" is a convention tuned to a clinical context, not a physical constant. A trial screening for a rare, catastrophic complication may reasonably tolerate a higher alpha to avoid missing it, just as a triage protocol for a lethal but rare diagnosis accepts more false positives to avoid a false negative.

Why Type II error is the one surgical literature keeps missing

Journals and reviewers scrutinise P values closely, which polices Type I error. Type II error gets far less attention, and the surgical literature shows it. A review of 117 randomised trials in orthopaedic trauma found a mean type II error rate of 90.5% for the primary outcome, with mean study power of only about 25% (24.65%, ranging from 2% to 99% across trials).1 Most of those trials that reported "no significant difference" were not statistically equipped to have found one even if it existed.

This is not a new observation. A 1978 survey of 71 "negative" trials in the New England Journal of Medicine found that most lacked adequate power to detect even a 25% therapeutic improvement, let alone a smaller but still clinically meaningful one.2 Absence of evidence gets read as evidence of absence in journal clubs constantly, and an underpowered "negative" trial is the statistical equivalent of ruling out a disease with a test that was never sensitive enough to catch it in the first place.

Reading a "negative" trial

Type I and Type II errors are properties of a single sample measured against an unknown population truth, so you can never know from one trial's result alone whether either has actually occurred, only the probability of each given the sample size the authors chose.4 Before treating a non-significant result as reassurance, check the power calculation the authors report, or the confidence interval around the effect estimate. A wide confidence interval that comfortably includes a clinically meaningful effect is a Type II error warning sign, regardless of what the P value says.

Setting the threshold before you look at the data

The formal two-error framework comes from Jerzy Neyman and Egon Pearson, who in 1933 argued that a test should be judged by both error rates simultaneously rather than by a single significance threshold considered in isolation.3 Their point translates directly to practice: decide, before you collect data, how much false-positive risk and how much false-negative risk you are willing to accept, given what each error would cost your specific patients. That is exactly how a surgeon sets a threshold for re-exploration before symptoms evolve, rather than moving the goalposts once the picture is ambiguous. Pre-specifying alpha and target power in a protocol is the research version of deciding your action threshold before you are standing at the bedside, uncertain, with the clock running.

References

  1. Lochner HV, Bhandari M, Tornetta P 3rd. Type-II error rates (beta errors) of randomized trials in orthopaedic trauma. J Bone Joint Surg Am. 2001;83(11):1650–1655. https://doi.org/10.2106/00004623-200111000-00005
  2. Freiman JA, Chalmers TC, Smith H Jr, Kuebler RR. The importance of beta, the type II error, and sample size in the design and interpretation of the randomized control trial: survey of 71 "negative" trials. N Engl J Med. 1978;299(13):690–694. https://doi.org/10.1056/NEJM197809282991304
  3. Neyman J, Pearson ES. On the problem of the most efficient tests of statistical hypotheses. Philos Trans R Soc Lond A. 1933;231:289–337. https://doi.org/10.1098/rsta.1933.0009
  4. Sedgwick P. Pitfalls of statistical hypothesis testing: type I and type II errors. BMJ. 2014;349:g4287. https://doi.org/10.1136/bmj.g4287

When StatsPlease runs your comparison, whichever preset fits your design, Group Comparison, Before vs After, or Correlation, it reports the exact P value and effect size behind it, computed directly from your data, so the alpha and power you state in your methods section reflect your actual analysis rather than a borrowed convention.

Try StatsPlease free