Test selection & reporting · ~7 min read
Kruskal-Wallis Test and Dunn's Post Hoc
The omnibus test confirms a difference exists somewhere. Localising it is a second, separate calculation.
Published
In short
The Kruskal-Wallis test is the non-parametric answer to one-way ANOVA for three or more independent groups, and like ANOVA it is an omnibus test: a significant result means at least one group's ranks differ from the others, not which one.1 It is also, strictly, a test of stochastic dominance across the ranked distributions rather than a test of medians specifically — it only collapses cleanly to "the medians differ" when the groups' distributions have the same shape and spread.3 Getting from a significant H statistic to a usable Results sentence requires a second step: Dunn's test, run with a multiple-comparison correction, on the same pooled ranks the omnibus test used.2 Papers that report the Kruskal-Wallis P value and stop there are reporting half a result, and it's specifically the half a peer reviewer is trained to notice is missing.
What the omnibus H-test actually claims
Kruskal-Wallis takes every observation across all groups, ranks them together from smallest to largest ignoring group membership, then asks whether the average rank in each group is close enough to what you'd expect if group assignment were irrelevant.1 That single H statistic and its P value answer one question only: is there evidence that at least one group's ranks are systematically higher or lower than the others? They say nothing about which group, or how many of the groups involved are actually different from each other. A three-group study with a significant Kruskal-Wallis result could have one group departing from two similar ones, all three pairwise-different, or any other pattern consistent with "not all the same" — the omnibus test can't distinguish between them.
It's also worth being precise about what "differ" means here. Kruskal-Wallis is a test of stochastic dominance across the ranked distributions, not, strictly, a test of medians.3 If the groups' distributions have noticeably different shapes or spreads — one skewed, one bimodal, one with far more variance than the others — a significant result can reflect that difference in shape rather than a shift in central tendency, and a rejected null hypothesis doesn't automatically license the sentence "the medians differed." The median-comparison interpretation is only safe once the groups look like the same distribution shifted, and that's worth a quick look at the data before the sentence gets written, not an assumption.
Why the omnibus result isn't the end of the analysis
The natural instinct after a significant Kruskal-Wallis result is to run pairwise Mann-Whitney U tests on each pair of groups. This is the wrong follow-up, for a specific reason: Mann-Whitney ranks only the two groups being compared, while Kruskal-Wallis ranked all the groups together. Running separate two-group tests throws away the pooled ranking the omnibus test was actually built on and reintroduces the multiple-comparisons problem with no correction attached, inflating the false-positive rate as more pairs get tested.
Dunn's test avoids both problems. It reuses the same rank sums the Kruskal-Wallis test already computed and compares pairs of groups using the pooled variance implied by the omnibus null hypothesis, then applies a correction — Bonferroni is the most common, though Holm and Benjamini-Hochberg variants exist — across however many pairwise comparisons the design calls for.2 That correction is not optional bookkeeping; it's the part of the analysis that keeps the overall false-positive rate at the level the study claims, and it's the specific thing a statistically literate reviewer checks for when three or more groups are being compared non-parametrically.
A significant Kruskal-Wallis test with no post-hoc comparisons reported is one of the more recognisable incomplete patterns in a Methods and Results section — it tells a reviewer the omnibus test was run and the follow-up question was never answered, not that there was nothing left to answer.
An illustrative example
Consider a hypothetical three-arm comparison of 24-hour postoperative opioid consumption (oral morphine milligram equivalents) across three multimodal analgesia protocols in 84 patients — illustrative numbers, constructed for demonstration, not drawn from a real dataset. Opioid consumption data of this kind is typically right-skewed, with a handful of patients needing rescue doses well above the rest, which is exactly the shape that makes rank-based testing the more defensible choice over a one-way ANOVA on the raw milligrams.
| Protocol | n | Median | IQR |
|---|---|---|---|
| A | 28 | 32.0 | 22.0–48.0 |
| B | 28 | 18.0 | 10.0–29.0 |
| C | 28 | 21.0 | 14.0–33.0 |
The Kruskal-Wallis test on this data returns H(2) = 14.6, P < .001 — evidence that consumption isn't distributed the same way across the three protocols. Dunn's test with Bonferroni correction across the three pairwise comparisons then localises that finding: Protocol A differs from Protocol B (Padj = .001) and from Protocol C (Padj = .014), while B and C do not differ from each other (Padj = 1.00). Without that second step, the honest sentence available from the omnibus result alone is "consumption differed across protocols somewhere" — which is not a sentence any Results section should have to settle for when the pairwise data is sitting right there.
Writing the AMA sentence
The complete version reports both steps, in order: "Twenty-four-hour opioid consumption differed significantly across the three analgesia protocols (Kruskal-Wallis H[2] = 14.6, P < .001). Dunn's post hoc pairwise comparisons with Bonferroni correction showed that Protocol A patients consumed significantly more opioid than Protocol B (Padj = .001) and Protocol C (Padj = .014) patients; Protocols B and C did not differ significantly (Padj = 1.00)." That single paragraph does what the omnibus test alone can't: it tells the reader which comparisons actually drove the result, with the correction method named so the reviewer doesn't have to ask.
How to run it — SPSS vs StatsPlease
In SPSS Statistics (matched here to the current 32.0.0 documentation), the omnibus test lives in two different menus, and they are not equivalent. The older route, Analyze ▸ Nonparametric Tests ▸ Legacy Dialogs ▸ K Independent Samples, asks for a Test Variable List and a Grouping Variable with a Define Range for the group codes, ticks the Kruskal-Wallis H box, and returns exactly that — an H statistic, degrees of freedom, and a P value, with no pairwise comparisons attached. The newer route, Analyze ▸ Nonparametric Tests ▸ Independent Samples, is the one that gets you Dunn's test: on the Fields tab, the outcome goes into Test Fields and the grouping variable into Groups; on the Settings tab, selecting Customize Tests and ticking Kruskal-Wallis 1-way ANOVA runs it. The output opens in the Model Viewer, and double-clicking the summary view opens a Pairwise Comparisons table with a Sig. column and, next to it, a Bonferroni-adjusted Adj. Sig. column. The number that supports a post-hoc claim is the Adj. Sig. column — quoting the unadjusted Sig. column instead is the most common misread of this output, and the two can differ enough to flip a comparison from significant to not.
Upload the dataset (or the relevant columns). StatsPlease's deterministic engine identifies the appropriate test from the variable type and study design, checks the relevant assumptions, computes the result using fixed, non-LLM algorithms, and drafts the Methods/Results sentence in AMA format — the same number a reader would get running the test by hand in SPSS.
You might also read
References
- Kruskal WH, Wallis WA. Use of ranks in one-criterion variance analysis. J Am Stat Assoc. 1952;47(260):583-621.
- Dunn OJ. Multiple comparisons using rank sums. Technometrics. 1964;6(3):241-252.
- Vargha A, Delaney HD. The Kruskal-Wallis test and stochastic homogeneity. J Educ Behav Stat. 1998;23(2):170-192.
A Kruskal-Wallis result and its pairwise follow-up are exactly the kind of two-step calculation where a transcription slip or a misread output column introduces an error — and the two belong together, not looked up separately.
StatsPlease's deterministic engine computes the Kruskal-Wallis H statistic from the ranks in your uploaded data and, when the omnibus result is significant, the Dunn's pairwise comparisons with the multiple-comparison correction applied, then drafts the Methods/Results sentence for both in AMA format — the same numbers you would get running each step by hand in SPSS.
Try StatsPlease free