← Back to research
J. Financial Services Research · Under review · 2026

Who Bears the Burden?

Heterogeneous Racial Approval Differentials in U.S. Mortgage Lending — causal forest double ML on 42.3 million applications.

Rajveer Singh Pall · 2026

ProblemBlack mortgage applicants are approved far less often than White applicants — the open question is how much of that survives once creditworthiness is accounted for.
ApproachCausal machine learning estimates an individual-level penalty for every applicant, then checks who specifically carries it.
ResultA 9.39-point penalty survives controls for 90.7% of Black applicants — and it's far larger under manual review than automated scoring.
SignificancePoints the cause at human discretion in underwriting, not the algorithms — the opposite of where bias concerns usually focus.

The Discovery in One Figure

Manual underwriting−14.8 pp
Automated underwriting−6.2 pp

The gap is driven by human discretion, not automated scoring

The Paper in Five Minutes

Knowing the average approval gap is not enough — averages hide who actually pays. Using methods from modern causal inference (the same family behind clinical-trial analysis), this paper estimates the approval penalty for each applicant profile across 42 million mortgage applications. The distribution is wide: nine in ten Black applicants face some estimated penalty, and the decisive factor is not the applicant but the process — applications handled by human underwriters carry more than double the penalty of those decided by automated systems. That points the fairness question at something a regulator can act on: how applications are routed.

The Research Question

Across 42.3 million HMDA mortgage applications (2020–2024), the raw Black–White approval gap is 14.95 percentage points. The interesting question is causal, not descriptive: how much of that gap survives once genuine creditworthiness is accounted for — and who, specifically, bears it?

How It Works

I estimate individual-level conditional average treatment effects with double machine learning (LightGBM nuisance models, cross-fitting) and causal forests, adjusting for 33 creditworthiness controls — income, LTV, DTI, loan purpose and type, automated-vs-manual underwriting, and geography. Cross-fitting removes regularization bias so the effect estimate is not contaminated by the prediction models.

The headline estimate is triangulated across five independent identification strategies — DFL decomposition, within-lender fixed effects, RDD at the 80% LTV PMI threshold, difference-in-differences across the 2022–24 tightening, and Manski partial-identification bounds. A race-shuffle placebo yields a 17.9× signal-to-noise ratio; the DR-Learner replicates the estimate within 0.15 pp.

01
Engineer at scale42M HMDA applications → 2M/1.5M estimation samples
02
Isolate the differentialdouble machine learning, LightGBM nuisances, 5-fold cross-fitting
03
Map who bears itcausal forest estimates the penalty per applicant profile
04
Find the mechanismmanual vs automated underwriting · same lender, same year
05
Attack the resultplacebo shuffles · Oster bounds · Cinelli–Hazlett sensitivity

The Results

42.3MHMDA applications
−9.39 ppConditional penalty
90.7%Applicants penalised
17.9×Placebo signal-to-noise

A −9.39 pp mean conditional penalty remains after controls — 62.8% of the raw gap is left unexplained by observable creditworthiness. 90.7% of Black applicants face a negative conditional effect, and the penalty is largest under manual underwriting (−14.8 pp) versus automated systems (−6.2 pp), implicating human discretion over algorithmic scoring as the dominant channel.

Conditional approval penalty by underwriting channel

Manual underwriting−14.79 pp
Pooled (all channels)−9.39 pp
Same lender, same year−7.13 pp
Automated systems−6.17 pp

Bar length = size of the estimated Black–White differential (all negative). The channel, not the applicant, is decisive.

Double-machine-learning estimates across specifications
Double-machine-learning estimates of the conditional approval differential across specifications.
Distribution of individual-level estimated effects
The distribution of individual-level estimated effects: wide heterogeneity, overwhelmingly negative.
SHAP attribution over the causal-forest estimates
SHAP attribution over the causal-forest estimates: what drives who bears the penalty.

Level treated as an upper bound (no credit scores in HMDA); the channel contrast is the robust object.

What This Changes

It points the fairness question at an actionable mechanism — lender-controlled routing and handling of applications — rather than at applicant characteristics. A regulator can act on how applications are routed; an applicant cannot.

↗ View on GitHub ↗ SSRN preprint ↗ Full research page

Citation

Rajveer Singh Pall. "Who Bears the Burden? Heterogeneous Racial Approval Differentials in U.S. Mortgage Lending: Causal Forest DML on 42 Million HMDA Applications." Under journal review, 2026. SSRN preprint 6984959.

Related Research