Do School Bureaucrats Discriminate against Foreigners? Experimental Evidence from Italy

This page is an experiment, inspired by a tweet from MIT Sloan professor Anna Stansbury: what can a paper be when it isn't a PDF? Check out below my (admittedly experimental) answer to that question!

One · what varies

The same email, three signatures

Every school received an identical inquiry about enrollment. Only the sender's name changed.

Email text translated from Italian. The municipality named in the body was personalized to each school. I report only three names here as an example; the experiment used four Italian names as controls, and two French and two Arabic names for treatment.

Two · how it is assigned

Randomization happened at the province stratum

Italian signature · 50% French signature · 25% Arabic signature · 25%

7,120state schools

103provinces

50 / 25 / 25drawn within each

7,120 schools · 103 provinces · 18 regions · sending randomized again across 9 working days

Three · the headline

The French name fares worse than the Arabic one

80.9% of schools answered the Italian-name control. A French signature loses six points of that; an Arabic one loses two and a half. Both gaps are significant, and so is the distance between them.

Data table

Four · the same answer, any way you ask

Build the comparison yourself

Effects on the response rate, in percentage points relative to the Italian-name control. Whiskers are 95% intervals. Pick any two specifications; the axis never rescales, so distances stay comparable.

Specification A · filled

Specification B · hollow

Every specification includes province and day-of-sending fixed effects; standard errors are heteroskedasticity-robust. Logit coefficients are presented as average marginal effects. Post-double LASSO follows Belloni, Chernozhukov and Hansen (2014). The one-day outcome was only estimated with the LPM. N = 7,120 throughout.

Five · the replies themselves

When they do answer, what do they say?

Answering is only half of the behavior. Every reply was classified as useful or not, and polite or not.

500 replies hand-labeled as useful / polite, to teach the model
BERT, Italian version fine-tuned on those labels
5,107 replies labeled by the model: the estimation sample

Estimated on the 5,107 schools that replied, once the 500 emails used to train the NLP model are set aside. All coefficients are conditional on getting an answer at all. Controls are always on in these specifications. Every specification includes province and day-of-sending fixed effects; standard errors are heteroskedasticity-robust. The two classifiers are not equally good: politeness was learned at 0.88 accuracy, usefulness at 0.68, so the usefulness labels carry more noise.

What counted as useful, what counted as polite
UsefulContains any information beyond the administrative office's opening hours; a phone number to call was enough. Not useful if even the opening hours are missing.
PoliteIncludes openings and greetings, uses the formal Lei or Voi form, no capslock.
ClassifierUsefulness: accuracy 0.680 · F1 0.675 · precision 0.684 · recall 0.680.
Politeness: accuracy 0.88 · F1 0.874 · precision 0.883 · recall 0.88.

Six · where it bites

The same experiment, region by region

Conditional average treatment effects from a linear AIPW model. Purple means the region treats foreign names worse than Italian ones; green means better. Hover, tap or tab onto a region for its estimate.

Outcome Arm

Estimator: linear AIPW (Chernozhukov et al. 2018), nuisance parameters from a random forest with the honesty approach (Athey and Imbens 2016), outcome and treatment modeled conditional on province and day-of-sending, ten cross-fitting folds. Usefulness and politeness are conditional on the school replying at all. Trentino-Alto Adige and Valle d'Aosta are not in the experiment: absent from the Ministry source data, and their large German- and French-speaking minorities would have confounded the alias manipulation. Region boundaries: ISTAT via openpolis, CC-BY.

Data table for the current selection