How the U statistic works
Pool both groups and rank every value from smallest to largest (tied values share the average rank). U₁ = R₁ − n₁(n₁ + 1)/2 counts how many times a value from group 1 beats a value from group 2 across all n₁ × n₂ pairs. If the groups come from the same distribution, U₁ should be near n₁n₂/2; the further it is, the stronger the evidence of a difference.
Effect size
The rank-biserial correlation r = (U₁ − U₂)/(n₁n₂) runs from −1 to 1. The common-language effect size U₁/(n₁n₂) is the probability that a random member of group 1 scores higher than a random member of group 2, with ties counting half (McGraw & Wong, 1992).
Mann–Whitney or t-test?
For roughly normal data, the Welch t-test is slightly more powerful. For skewed data, ordinal scales, or outliers, Mann–Whitney is safer and loses little: its asymptotic relative efficiency versus the t-test is 95.5% even when the data are normal (Hodges & Lehmann, 1956).
Comparing three or more groups? Use the Kruskal–Wallis test, which extends Mann–Whitney and includes Dunn post-hoc comparisons.
Frequently asked questions
When should I use the Mann–Whitney U test?
To compare two independent groups when the outcome is ordinal (for example Likert ratings) or continuous but clearly non-normal with small samples. It is the nonparametric alternative to the independent-samples t-test.
Does Mann–Whitney compare medians?
Only if both groups have the same distribution shape. In general it tests whether a value from one group tends to be larger than a value from the other (stochastic dominance). The common-language effect size reports exactly that probability.
Exact or normal approximation?
With no ties and both groups of 30 or fewer, this calculator computes the exact p-value from the permutation distribution of U. Otherwise it uses the normal approximation with tie and continuity corrections, as R’s wilcox.test does.
How do I report it in APA style?
Give medians for each group, then U, z (for larger samples), p and an effect size such as the rank-biserial correlation: “U = 15, p = .254, r = .40.”
What is the difference between U and W?
Wilcoxon’s rank-sum W is the sum of ranks in one group; U = W − n(n + 1)/2. They are equivalent tests. R reports W, which equals U for the first group.