An employer audits pay across 5,000 people and finds no significant gender gap. The number is correct. The conclusion is not. Underneath it, two business units are pulling hard in opposite directions — and the aggregate reports their sum, not their content.
Absence of Disparity ≠ Equity
Opposite biases across units can cancel in aggregate, producing a misleading zero net effect.
This is where many audits stop. Pool every employee, compare mean base pay by gender, run the test. The gap is small, the p-value is nowhere near a conventional threshold, and the review concludes.
The confidence interval is the tell. It spans roughly $14,000. An aggregate this imprecise is consistent with unit-level effects several times larger than the point estimate — in either direction, and in both at once. A null result on a wide interval is not evidence of parity; it is an absence of resolution.
Same employees. Same variable. Same test. The only thing that changes is the unit of analysis. Reveal the units one at a time and watch what the organization-wide average was averaging over.
Two units carry significant gaps in opposite directions, a third is significant and positive, and a fourth cannot be distinguished from zero. All four are inside the organization-wide null. Tests are Welch's unequal-variance t-tests on the illustrative parameters in Section 06.
The two largest units are equal in size and close in magnitude, pointing opposite ways. Each unit's contribution to the organization-wide gap is its own gap weighted by its share of the workforce. Sum the contributions and the two large effects offset each other almost exactly.
| Business unit | N | Women's share | Mean base pay | Gap (M−W) | p | Weight | Contribution |
|---|
Two units of identical size (1,500 each) carry gaps of +$20,000 and −$22,000. Weighted at 30% each they contribute +$6,000 and −$6,600, offsetting all but $600 of each other. Product Design and Client Support add back +$3,600, leaving an employment-weighted total of +$3,000.
The actual pooled gap is smaller still, +$1,143, because women's share is not constant across units: Client Support is only 30% women, so men are over-weighted in a mid-paying unit while women are over-weighted in low-paying Field Service. That composition difference accounts for the remaining $1,857. Either way, the organization-wide figure describes no unit in the company.
Leave every unit's internal pay structure untouched — the same means, the same gender shares, the same gaps — and change only how many people each unit employs. The organization-wide gap moves, and it will change sign. No individual's pay changed.
Pooled means are recomputed live: each unit contributes its own men's and women's means, weighted by its scaled headcount and its fixed gender share. Within-unit pay structure is held constant throughout. This is the classroom point — an organization can move its headline number without changing anyone's pay, simply by growing or shrinking units.
A gap is a symptom. The object of inquiry is the process that generated it. Once the units disagree, the analytical task is to specify competing data-generating processes and identify what evidence would distinguish among them. These four are the candidates worth putting in front of students — deliberately posed as questions, not answers.
Are starting salaries set by competing offers, or by any external reference that itself varies by group?
What would settle it: longitudinal salary histories back to hire date. Cross-sectional pay cannot separate an entry-point effect from subsequent differences in raises. If the data start at today's salary, the honest answer is that the hypothesis is untestable and the data request is the finding.
Does the gap sit in base pay, or in the discretionary layer — bonus, equity, spot awards — where judgment has more room?
What would settle it: decompose total compensation and test each component separately against a measure of contribution. A unit can pay a base premium to one group while favoring another in bonus; testing only total pay conceals both.
Are the groups differently distributed across job levels, grades, or credential tiers within the unit?
What would settle it: compare within level and grade, then ask why the level distribution looks the way it does. A gap that attenuates on controlling for level has been relocated, not explained — the level assignment is itself an outcome to interrogate.
Are the groups concentrated in different roles, so that a pay gap is really a job-allocation gap?
What would settle it: role-by-group composition alongside a within-role pay test. Note that these are two separate claims: a unit can be visibly segregated by role and still show no significant pay gap. Segregation and disparity each have to be demonstrated on their own evidence.
Explained ≠ Justified. A decomposition tells you which variable accounts for the gap. It does not tell you that the process generating that variable was fair.
The Xs Are Not Innocent. Level, tenure, credentials, and performance ratings are themselves produced by processes that may carry bias. Whether to adjust for them is a normative and strategic choice, not a purely statistical one.
If aggregation reverses or erases the pattern, Simpson's Paradox is present.
| Unit | n men | n women | mean men | mean women | SD men | SD women |
|---|
Welch's t-test depends only on these means, standard deviations, and sample sizes, so the statistics on this page are fully determined by the table — no random draw is involved. Any dataset matching these moments reproduces every figure shown.
The illustration above is deliberately thin — one variable, four units, no institutional context. If you want to teach the paradox as a decision problem, with a protagonist, a board deadline, and real analytical ambiguity, there is a published case built around exactly this structure, with a teaching note and a supplemental dataset.
Rider, C. I., Choi, E., & Kim, Y. (2023). The Quest for Gender Pay Equity at Elemental Systems. WDI Publishing, case 5-154-986. 8 pages.
Listed learning objectives include conducting an organizational pay equity analysis, applying the equity analytics framework to distinguish differential treatment from disparate impact, and recognizing how organization-level data can obscure disparities within individual units.
This interactive is a companion to a graduate course in equity analytics. The course hub sets out the 2×2 that organizes the candidate mechanisms in Section 05 — Process (allocations × valuations) crossed with Behavior (differential treatment × disparate impact) — along with the seven-step analytical workflow and the full set of maxims.
Using this page in your own course is welcome — link to it, project it, or rebuild it from the parameters in Section 06. Adaptations and corrections are welcome by email.