Published
-
Predicting Well-Being with Mobile Phone Data: Evidence from Four Countries
M. Merritt Smith, Emily Aiken, Joshua E. Blumenstock, Sveta Milusheva
AEA Papers and Proceedings, 2026
Using data from Afghanistan, Côte d’Ivoire, Malawi, and Togo, we run parallel, standardized machine learning experiments to assess which measures of welfare can be predicted from mobile phone data, which types of phone data are most useful, and how much training data is required. Long-term poverty measures like wealth indices and multidimensional poverty are predicted more accurately than consumption, while transient vulnerability measures like food security and mental health are very difficult to predict.
Working Papers
-
Generated Data, Real Consequences: Why LLMs Are Inappropriate for Development Aid Targeting
M. Merritt Smith
2026
We evaluate proposals to use large language models to generate synthetic household survey data for development aid targeting, across 17 nationally representative household surveys (131,906 households). When asked to answer survey questions in the persona of a specific household, LLMs stereotype by country: country identity explains about half the variance in their consumption predictions, versus 12% in true consumption. Fine-tuning removes the stereotyping but not the targeting errors, and we argue that even without stereotyping, LLM-generated data would be inappropriate for allocative decisions about specific households.
Research Assistance
-
Infrastructure deficits and informal settlements in sub-Saharan Africa
Luís M. A. Bettencourt and Nicholas Marchio
Nature, 2025
I contributed to the building- and block-level population estimates underlying this paper as an RA at the Mansueto Institute.
-
Street access, Informality and Development: A block level analysis across all of sub‑Saharan Africa
Luís Bettencourt and Nicholas Marchio
arXiv, 2023
I contributed to the data that underlies this paper as an RA while at the Mansueto Institute. We used a combination of Open Street Map, LandScan, and WorldPop data to estimate the population of every block in sub-Saharan Africa, and from there estimate a measure called the “K-complexity” of the block, which is the number of other building parcels a person would have to walk through to get to a road. This tells us the degree to which a given block is a slum. We have since shared these data results with multiple international humanitarian organizations. The data is visualized here for public consumption.
-
Criminal charges, risk assessment, and violent recidivism in cases of domestic abuse
Dan A. Black, et al.
NBER Working Paper Series, 2023
I contributed to this paper as an RA while at Harris. It studies domestic abuse recidivism using data from the Greater Manchester Police. In this project, we studied two treatments: charging the perpetrator with a crime, or using a formal risk assessment created by the police to determine if the perpetrator should be labeled as high-risk. I worked on the propensity score weighting and analyzed the treatments for heterogeneity.