- The paper conducts a five-year, seven-stage audit of nearly 497,000 hiring pipeline records, revealing that aggregate gender parity—49.8% of hires were women—masked significant disparities in shortlisting and representation.
- The analysis identifies statistically significant adverse impacts for women in mid-salary roles (DIR 0.786), full-time roles (DIR 0.758), and among those aged 46–55 (DIR 0.77), alongside a €1,147 within-sector salary gap and the near absence of adults aged 55 and over.
- The paper shows that vendor opacity, missing demographic and qualification data, undocumented analyst decisions, and changing shortlisting rates limit causal interpretation and demonstrate why public-sector hiring systems require continuous, end-to-end fairness monitoring.
This paper presents an independent, end-to-end fairness audit of a semi-automated hiring system operated by Barcelona Activa, the labor intermediation service of the Barcelona City Council, which uses the proprietary TalentClue platform for candidate search and shortlisting. Drawing on approximately 497,000 candidate-vacancy pipeline entries spanning September 2017 to September 2022, the audit examines seven stages of the recruitment process rather than the automated component alone. Its central empirical claim is that aggregate gender parity in hiring outcomes coexists with substantial disparities at specific pipeline stages, salary bands, sector strata, and demographic intersections—disparities that a conventional model-output or compliance-style audit would not detect.
System and data
Barcelona Activa operates a human-in-the-loop intermediation process: employers submit vacancies; an analyst translates requirements into structured filters and free-text keywords; TalentClue executes searches and returns ranked profiles; the analyst manually constructs shortlists; and shortlists are handed to employers. The paper maps this as a seven-stage pipeline and documents pervasive visibility gaps: no record of keyword selection (Stage 3), no access to TalentClue's matching and ranking logic (Stages 4–5), no audit trail for shortlist inclusion or exclusion decisions (Stage 6), and no post-shortlist outcome data returned by employers (Stage 7). This vendor–deployer information asymmetry is itself treated as an audit finding: Barcelona Activa cannot independently assess whether returned profiles are representative of the eligible pool or how matching weights are assigned.
The dataset includes declared sex, age, origin, education, sector, salary band, contract type, and pipeline status. Two missingness issues are flagged as material: 24.0% of records lack gender information and 14.5% lack country of origin. Candidate-level qualification, skills, and experience data are unavailable, which constrains conditional parity analysis throughout.
Methodology
The audit applies Eticas' post-deployment impact assessment methodology, restricted by funding conditions to the Bias and Fairness category of its AI risk taxonomy. Three analytical layers are used. Pre-processing benchmarks the candidate pool against Barcelona's active labor force using the Spanish Labor Force Survey (EPA) for 2019–2021. In-processing computes Disparate Impact Ratios (DIR) across pipeline transitions, using the EEOC four-fifths threshold of 0.80 explicitly as a practitioner benchmark rather than a legal determination, complemented by chi-square or Fisher's exact tests at α=0.05. Stratified conditional analysis substitutes for multivariate regression on grounds of interpretability, operational actionability, and the limited covariate set. Post-processing examines temporal stability of shortlisting rates by gender from 2017 to 2022. Findings are tiered by robustness: large-sample, stratified, statistically significant results are presented as robust; small-subgroup results are flagged as indicative.
Representativeness findings
Binary gender composition tracks the labor market closely (51.5% women registered vs. 47.6% in the EPA benchmark). The most consequential pre-processing finding is structural: adults aged 55 and over constitute 15.6% of Barcelona's active labor force but are effectively absent from the platform at every stage—a zero representation that indicates either registration coverage failure, systematic exclusion through search practices or employer requirements, or dataset truncation rules that themselves constitute an auditability problem. Non-binary candidates appear only 285 times over five years, plausibly reflecting form design, non-response, or disclosure reluctance. Occupational composition shows strong gendered clustering (e.g., women underrepresented by 26.0 percentage points in Real Estate and Architecture relative to the labor market), though the audit cannot determine whether the platform reproduces or merely mirrors occupational segregation because matching logic is unobservable.
Disparities masked by aggregate parity
Women constitute 51.5% of registered candidates and 49.8% of hired candidates, a difference that is not statistically significant. Under an aggregate model-output framework—the scope of NYC Local Law 144-style audits—this system would be reported as fair. The pipeline-level analysis contradicts this conclusion along five dimensions:
| Finding |
Metric |
Significance |
| Mid-salary (€15k–24k) shortlisting |
DIR = 0.786 |
p < 0.001 |
| Full-time role shortlisting |
DIR = 0.758 |
p < 0.001 |
| Within-sector salary gap |
€1,147 aggregate; persistent in 15/20 sectors |
p < 0.001 |
| Women aged 46–55 shortlisting |
DIR = 0.77 |
below 0.80 threshold |
| Adults 55+ representation |
Zero vs. 15.6% labor force share |
structural |
Two methodological points deserve emphasis. First, the within-sector salary gap (women matched to vacancies averaging €16,492 vs. €17,639 for men) persists after sector stratification, demonstrating that sectoral composition alone does not explain the aggregate disparity. Second, because matching is analyst-driven rather than candidate-selected, the full-time/part-time asymmetry (adverse impact concentrated in full-time roles, none in part-time) cannot be attributed to differential candidate preferences. Shortlisting also reduces women's representation within male-typed sectors where they were already matched—for example, a 12.4 percentage-point drop in Real Estate and Architecture—indicating a stage-internal filtering effect rather than self-selection.
Three further findings are flagged as indicative due to sample limitations. Non-binary candidates were shortlisted at 3.51% versus 11.89% for men (DIR = 0.295, Fisher's exact test p < 0.001), with zero forwarded to employers across the entire period; the authors caution that N = 285 limits precision even though chance is unlikely to explain the pattern. Origin-based disparities show Spanish candidates shortlisted at 10.27% versus 5.80% for EU-origin (DIR = 0.565) and 6.48% for non-EU candidates (DIR = 0.631), widening sharply at forwarding (DIRs of 0.386 and 0.149); these estimates rest on 14.5% missing origin data and a very small non-EU subgroup (N = 242), without controls for language proficiency or credential recognition. An "education paradox" emerges in which higher-educated candidates are shortlisted at rates 40–42% lower than high school graduates (DIRs 0.576–0.600); since higher-education categories skew female (58–62%), this mechanism disproportionately affects women, though causal attribution requires qualification data the audit lacks.
Temporal non-stationarity
The gender shortlisting gap narrowed from approximately 6.5 percentage points in 2017 to 1.3 percentage points in 2022. Taken alone this suggests improving fairness, but it coincides with overall shortlisting rates halving for both groups (roughly 18% to 9%), consistent with rising application volume relative to vacancies. Whether convergence reflects substantive improvement or selection effects in a tightening pipeline cannot be determined from aggregated data. The implication drawn is structural: point-in-time audits produce snapshots of a moving target—an audit in 2017 and one in 2022 would report materially different conclusions about the same system, neither capturing the trajectory. This provides direct empirical support for continuous evaluation in production, which the authors position as complementary to pre-deployment assessment, periodic independent auditing, and incident-based investigation, and grounded in Article 72 post-market monitoring obligations under the EU AI Act.
Limitations and open questions
The paper is explicit that its findings evidence differential outcomes, not causal effects attributable to any single component. Three limitations bear directly on interpretation. First, missing variables—qualifications, skills, experience, analyst keywords, and post-shortlist employer decisions—prevent conditional parity analyses and reviewer-level analysis, and leave the three candidate mechanisms (pool composition, platform filtering, analyst discretion) empirically entangled. Second, substantial missingness in gender (24%) and origin (14.5%) fields may itself bias group-specific estimates. Third, the audit's funding-mandated scope excluded privacy, governance, reliability, and other risk categories identified during risk mapping, so the fairness findings should not be read as a comprehensive system assessment. Open questions left by the paper include whether the 55+ absence originates in registration coverage or search practices, whether the temporal gender-gap convergence will persist or reverse, and what magnitude of non-binary disadvantage the current sample size conceals.
Conclusion
By tracing five years of operational decisions across seven pipeline stages, this audit quantifies the gap between aggregate and pipeline-level fairness assessment in a deployed public-sector hiring system. It demonstrates concretely what component-level evaluation misses, documents vendor opacity as a governance constraint on deployers, and supplies longitudinal evidence that fairness properties are non-stationary. The practical upshot is twofold: deployers need auditable human decision processes and procurement-level access to third-party system documentation, and regulators should anchor accountability in deployment-context outcomes and continuous monitoring rather than fixed periodic procedures alone.