Methodology
How we measure the gap — and how we make it defensible
Every number Paritir produces comes from a documented model you can inspect. This page sets out the job-evaluation scoring, the adjusted-gap regression, the automated defensibility checks — and, just as important, what the method does not claim.
The pipeline
From workforce data to a sealed report
Five stages. Each is deterministic and reproducible: the same inputs always produce the same result.
Import
HRIS employee data — pay, hours, contract, location, start date — normalised to FTE-annualised total remuneration in a single reporting currency.
Job evaluation
A gender-neutral job-evaluation survey scores each role's work value on four weighted factors, giving an objective measure of work of equal value.
Measure
A log-linear regression estimates the adjusted gap — the pay difference between women and men that remains after legitimate factors are accounted for.
Check
A battery of automated defensibility checks runs on every segment, producing a pass / warn / fail verdict per category of workers.
Report & seal
The report, its inputs and the check verdicts are sealed into a tamper-evident evidence pack for audit.
Step 1 — work of equal value
Job evaluation: scoring the value of the work
- Skills
- 40% — knowledge, qualifications, problem-solving and interpersonal demands
- Responsibility
- 30% — for people, resources, information and decisions
- Effort
- 20% — physical, mental and emotional demand
- Working conditions
- 10% — environment, hazards and unsocial hours
Weights are the platform default and can be overridden per organisation where a justified, gender-neutral scheme requires it. Where a role's own cohort is too small to score privately, its work-value score is rolled up from a wider cohort (entity → country → organisation) — and every roll-up is disclosed as a defensibility check, never hidden.
Step 2 — the adjusted gap
The adjusted (unexplained) gap
We report two numbers and never conflate them: the raw gap (a simple average difference) and the adjusted gap (what remains after legitimate factors).
An ordinary-least-squares regression of log pay. δ, the coefficient on the female indicator, is the adjusted gap: the average pay difference between women and men of otherwise-equal profile. Positive means women are paid less.
- Dependent variable
- Natural log of FTE-annualised total remuneration — base pay plus all cash and in-kind components.
- Default controls
- Work-value (job-evaluation) score, tenure and tenure², location (country) and contract type — the legitimate ‘equal value’ factors.
- Reference categories
- Deterministic: the most frequent level in each category, so results are reproducible.
- Standard errors
- HC3 heteroskedasticity-consistent (robust) — reliable on the unequal group sizes typical of workforce data.
- Precision reported
- A 95% confidence interval on every gap, the probability the true gap exceeds the 5% threshold, and the minimum detectable effect for the sample.
- Off by default
- Job family, seniority, part-time status and site are not controlled by default — each can absorb, and so hide, part of a real gap (e.g. occupational segregation). They are available as clearly-cautioned exploratory toggles.
Where useful, the raw gap can be decomposed into the share attributable to each factor (sequential, Shapley and Gelbach methods), separating the part explained by legitimate factors from the unexplained remainder.
The 5% threshold
When a gap becomes a joint pay assessment
Step 3 — defensibility
Automated defensibility checks
Every segment is tested before its number is trusted. Each check returns pass, warn, fail or info; the worst status across a report's categories is its verdict.
- Robust standard errors
- Confirms HC3 robust estimation was used. (info)
- Confidence interval
- Reports the 95% CI, the probability the gap exceeds 5%, and the minimum detectable effect. (info)
- Multicollinearity
- Maximum VIF and the design's condition number. Warns above VIF 5 or κ 30; fails above VIF 10, κ 100, or a rank-deficient design.
- Reference categories
- Records the reference level chosen for each control. (info)
- Single-gender
- Fails a segment that contains only one gender — no gap is estimable.
- Cell size (k-anonymity)
- Fails below 20 employees in the segment, or fewer than 10 of either gender — protecting both privacy and statistical validity.
- Field coverage
- Measures how completely a control is populated (e.g. tenure from start dates). Warns below 80%; fails on a flagged low-coverage segment. Coverage is measured — never silently dropped from the model.
- Job-evaluation completeness
- Analysis-level: the share of the workforce carrying a job-evaluation score. (info / warn)
- Work-value score scope
- Analysis-level: how much of the workforce carries a score rolled up from a wider cohort. Warns above 25%. (info / warn)
- Outlier surfacing (IQR)
- Flags raw-pay and residual outliers at 1.5× and 3× the interquartile range. Always surfaced for review, never removed from the analysis. (info / warn)
The publish gate
Advisory everywhere, blocking where it counts
Every published report — its inputs, its coefficients and its check verdicts — is embedded in a tamper-evident ‘defensibility seal’ (retained under GDPR Article 17(3)). If a works council, auditor or authority asks how a number was produced, the answer is reproducible from the sealed pack.
Honesty
Assumptions and limits
What this method can and cannot tell you.
Association, not proof
The adjusted gap is the pay difference not explained by the factors in the model. It is strong evidence to investigate — not, by itself, proof of discrimination.
Controls can hide gaps
Adjusting for factors that are themselves gender-correlated (job family, seniority, part-time, site) can absorb part of a real gap. That is why they are off by default and cautioned when on.
Small groups can't be measured
k-anonymity floors mean very small categories are not assessed. This protects individuals but leaves some pay unmeasured.
Only what's in the data
The model uses the controls present in your HRIS. Missing fields lower coverage (and are flagged); genuinely unobserved factors cannot be accounted for.
A model, not a lawyer
Results are a statistical and compliance tool, not legal advice. A flagged category calls for a joint pay assessment and, usually, counsel.
Reproducible by design
Standardised inputs, deterministic references and seeded computation mean the same data always yields the same result — and the same seal.
See it on your own data
Explore the full pipeline with sample data — no sign-up call — or read how we secure it.