Methodology

How we measure the gap — and how we make it defensible

Every number Paritir produces comes from a documented model you can inspect. This page sets out the job-evaluation scoring, the adjusted-gap regression, the automated defensibility checks — and, just as important, what the method does not claim.

The pipeline

From workforce data to a sealed report

Five stages. Each is deterministic and reproducible: the same inputs always produce the same result.

01

Import

HRIS employee data — pay, hours, contract, location, start date — normalised to FTE-annualised total remuneration in a single reporting currency.

02

Job evaluation

A gender-neutral job-evaluation survey scores each role's work value on four weighted factors, giving an objective measure of work of equal value.

03

Measure

A log-linear regression estimates the adjusted gap — the pay difference between women and men that remains after legitimate factors are accounted for.

04

Check

A battery of automated defensibility checks runs on every segment, producing a pass / warn / fail verdict per category of workers.

05

Report & seal

The report, its inputs and the check verdicts are sealed into a tamper-evident evidence pack for audit.

Step 1 — work of equal value

Job evaluation: scoring the value of the work

The Directive requires pay to be compared for equal work and work of equal value, assessed on gender-neutral criteria. Paritir scores every role with a structured questionnaire built on the four factor groups named in Article 4(4) — skills, effort, responsibility and working conditions. Each answer maps to a raw score; factor scores are normalised to a 0–100 scale and combined with the platform-default weighting below.
Skills
40% — knowledge, qualifications, problem-solving and interpersonal demands
Responsibility
30% — for people, resources, information and decisions
Effort
20% — physical, mental and emotional demand
Working conditions
10% — environment, hazards and unsocial hours

Weights are the platform default and can be overridden per organisation where a justified, gender-neutral scheme requires it. Where a role's own cohort is too small to score privately, its work-value score is rolled up from a wider cohort (entity → country → organisation) — and every roll-up is disclosed as a defensibility check, never hidden.

Step 2 — the adjusted gap

The adjusted (unexplained) gap

We report two numbers and never conflate them: the raw gap (a simple average difference) and the adjusted gap (what remains after legitimate factors).

An ordinary-least-squares regression of log pay. δ, the coefficient on the female indicator, is the adjusted gap: the average pay difference between women and men of otherwise-equal profile. Positive means women are paid less.

Dependent variable
Natural log of FTE-annualised total remuneration — base pay plus all cash and in-kind components.
Default controls
Work-value (job-evaluation) score, tenure and tenure², location (country) and contract type — the legitimate ‘equal value’ factors.
Reference categories
Deterministic: the most frequent level in each category, so results are reproducible.
Standard errors
HC3 heteroskedasticity-consistent (robust) — reliable on the unequal group sizes typical of workforce data.
Precision reported
A 95% confidence interval on every gap, the probability the true gap exceeds the 5% threshold, and the minimum detectable effect for the sample.
Off by default
Job family, seniority, part-time status and site are not controlled by default — each can absorb, and so hide, part of a real gap (e.g. occupational segregation). They are available as clearly-cautioned exploratory toggles.

Where useful, the raw gap can be decomposed into the share attributable to each factor (sequential, Shapley and Gelbach methods), separating the part explained by legitimate factors from the unexplained remainder.

The 5% threshold

When a gap becomes a joint pay assessment

Under Article 10 of the Directive, an unjustified gap of 5% or more in any category of workers — not closed within six months — triggers a joint pay assessment with workers' representatives. Paritir evaluates the threshold per segment and direction-agnostically (a gap favouring either sex counts), and flags every category that crosses it, with its confidence interval so you can see how firmly the estimate sits above or below the line.

Step 3 — defensibility

Automated defensibility checks

Every segment is tested before its number is trusted. Each check returns pass, warn, fail or info; the worst status across a report's categories is its verdict.

Robust standard errors
Confirms HC3 robust estimation was used. (info)
Confidence interval
Reports the 95% CI, the probability the gap exceeds 5%, and the minimum detectable effect. (info)
Multicollinearity
Maximum VIF and the design's condition number. Warns above VIF 5 or κ 30; fails above VIF 10, κ 100, or a rank-deficient design.
Reference categories
Records the reference level chosen for each control. (info)
Single-gender
Fails a segment that contains only one gender — no gap is estimable.
Cell size (k-anonymity)
Fails below 20 employees in the segment, or fewer than 10 of either gender — protecting both privacy and statistical validity.
Field coverage
Measures how completely a control is populated (e.g. tenure from start dates). Warns below 80%; fails on a flagged low-coverage segment. Coverage is measured — never silently dropped from the model.
Job-evaluation completeness
Analysis-level: the share of the workforce carrying a job-evaluation score. (info / warn)
Work-value score scope
Analysis-level: how much of the workforce carries a score rolled up from a wider cohort. Warns above 25%. (info / warn)
Outlier surfacing (IQR)
Flags raw-pay and residual outliers at 1.5× and 3× the interquartile range. Always surfaced for review, never removed from the analysis. (info / warn)

The publish gate

Advisory everywhere, blocking where it counts

The checks are advisory throughout the workflow — they never stop you exploring your data. But a report with a fail verdict cannot be published until either the underlying issue is resolved or an authorised administrator records a written override reason, which is stored with the report. A verdict that has gone stale — computed before the latest analysis — blocks publication the same way, so a published number always reflects the current data.

Every published report — its inputs, its coefficients and its check verdicts — is embedded in a tamper-evident ‘defensibility seal’ (retained under GDPR Article 17(3)). If a works council, auditor or authority asks how a number was produced, the answer is reproducible from the sealed pack.

Honesty

Assumptions and limits

What this method can and cannot tell you.

Association, not proof

The adjusted gap is the pay difference not explained by the factors in the model. It is strong evidence to investigate — not, by itself, proof of discrimination.

Controls can hide gaps

Adjusting for factors that are themselves gender-correlated (job family, seniority, part-time, site) can absorb part of a real gap. That is why they are off by default and cautioned when on.

Small groups can't be measured

k-anonymity floors mean very small categories are not assessed. This protects individuals but leaves some pay unmeasured.

Only what's in the data

The model uses the controls present in your HRIS. Missing fields lower coverage (and are flagged); genuinely unobserved factors cannot be accounted for.

A model, not a lawyer

Results are a statistical and compliance tool, not legal advice. A flagged category calls for a joint pay assessment and, usually, counsel.

Reproducible by design

Standardised inputs, deterministic references and seeded computation mean the same data always yields the same result — and the same seal.

See it on your own data

Explore the full pipeline with sample data — no sign-up call — or read how we secure it.