Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
23 changes: 23 additions & 0 deletions .github/workflows/draft-pdf.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,23 @@
on:
push:
paths:
- paper.md
- paper.bib
pull_request:
paths:
- paper.md
- paper.bib

jobs:
paper:
runs-on: ubuntu-latest
name: Draft PDF
steps:
- uses: actions/checkout@v4
- uses: openjournals/openjournals-draft-action@master
with:
journal: joss
- uses: actions/upload-artifact@v4
with:
name: paper
path: paper.pdf
31 changes: 31 additions & 0 deletions CITATION.cff
Original file line number Diff line number Diff line change
@@ -0,0 +1,31 @@
cff-version: 1.2.0
message: "If you use this software, please cite it as below."
type: software
title: "microimpute: A Model-Agnostic Tool for Cross-Survey Imputation"
url: "https://github.com/PolicyEngine/microimpute"
repository-code: "https://github.com/PolicyEngine/microimpute"
abstract: "A Python package for imputing variables from one survey onto another, implementing five imputation methods behind a common interface with cross-validated automatic method selection."
authors:
- family-names: Ahmadi
given-names: Vahid
orcid: "https://orcid.org/0009-0004-1093-6272"
affiliation: "PolicyEngine, Washington, DC, United States"
- family-names: Ghenis
given-names: Max
orcid: "https://orcid.org/0000-0002-1335-8277"
affiliation: "PolicyEngine, Washington, DC, United States"
- family-names: Juaristi
given-names: María
orcid: "https://orcid.org/0009-0007-4946-2248"
affiliation: "PolicyEngine, Washington, DC, United States"
- family-names: Woodruff
given-names: Nikhil
orcid: "https://orcid.org/0009-0009-5004-4910"
affiliation: "PolicyEngine, Washington, DC, United States"
keywords:
- imputation
- statistical matching
- survey microdata
- quantile regression
- microsimulation
- Python
64 changes: 64 additions & 0 deletions CODE_OF_CONDUCT.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,64 @@
# Contributor covenant code of conduct

## Our pledge

We as members, contributors, and leaders pledge to make participation in our
community a harassment-free experience for everyone, regardless of age, body
size, visible or invisible disability, ethnicity, sex characteristics, gender
identity and expression, level of experience, education, socio-economic status,
nationality, personal appearance, race, religion, or sexual identity
and orientation.

We pledge to act and interact in ways that contribute to an open, welcoming,
diverse, inclusive, and healthy community.

## Our standards

Examples of behavior that contributes to a positive environment for our
community include:

* Demonstrating empathy and kindness toward other people
* Being respectful of differing opinions, viewpoints, and experiences
* Giving and gracefully accepting constructive feedback
* Accepting responsibility and apologizing to those affected by our mistakes,
and learning from the experience
* Focusing on what is best not just for us as individuals, but for the
overall community

Examples of unacceptable behavior include:

* The use of sexualized language or imagery, and sexual attention or
advances of any kind
* Trolling, insulting or derogatory comments, and personal or political attacks
* Public or private harassment
* Publishing others' private information, such as a physical or email
address, without their explicit permission
* Other conduct which could reasonably be considered inappropriate in a
professional setting

## Enforcement responsibilities

Community leaders are responsible for clarifying and enforcing our standards of
acceptable behavior and will take appropriate and fair corrective action in
response to any behavior that they deem inappropriate, threatening, offensive,
or harmful.

## Scope

This Code of Conduct applies within all community spaces, and also applies when
an individual is officially representing the community in public spaces.

## Enforcement

Instances of abusive, harassing, or otherwise unacceptable behavior may be
reported to the community leaders responsible for enforcement at
hello@policyengine.org. All complaints will be reviewed and investigated
promptly and fairly.

## Attribution

This Code of Conduct is adapted from the [Contributor Covenant][homepage],
version 2.0, available at
https://www.contributor-covenant.org/version/2/0/code_of_conduct.html.

[homepage]: https://www.contributor-covenant.org
1 change: 1 addition & 0 deletions changelog.d/joss-paper.added.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- Added a JOSS paper, a citation file, a code of conduct, and a workflow that builds a draft PDF of the paper.
124 changes: 124 additions & 0 deletions paper.bib
Original file line number Diff line number Diff line change
@@ -0,0 +1,124 @@
@misc{arnold_ventures,
title={Public Finance Program},
author={{Arnold Ventures}},
year={2023},
note={Grant to PolicyEngine for congressional district-level policy analysis},
url={https://www.arnoldventures.org/work/public-finance}
}

@misc{nsf_pose,
title={{POSE}: Phase {I}: {PolicyEngine} -- Advancing Public Policy Analysis},
author={{National Science Foundation}},
year={2025},
note={Award 2518372. PI: Max Ghenis, PSL Foundation. \$299,974},
url={https://www.nsf.gov/awardsearch/showAward.jsp?AWD_ID=2518372}
}

@online{neo_philanthropy,
title={{NEO Philanthropy} Awards \$200,000 Grant to {PolicyEngine}},
author={Ghenis, Max},
year={2024},
url={https://policyengine.org/us/research/neo-philanthropy}
}

@software{claude2026,
title={{Claude}},
author={{Anthropic}},
year={2026},
note={Opus 4 model used for code refactoring assistance},
url={https://www.anthropic.com/claude}
}

@misc{nuffield2024grant,
title={Enhancing, localising and democratising tax-benefit policy analysis},
author={{Nuffield Foundation}},
year={2024},
note={General Election Analysis and Briefing Fund grant to PolicyEngine},
url={https://www.nuffieldfoundation.org/project/enhancing-localising-and-democratising-tax-benefit-policy-analysis}
}
@article{koenker1978regression,
author = {Koenker, Roger and Bassett, Gilbert},
title = {Regression Quantiles},
journal = {Econometrica},
volume = {46},
number = {1},
pages = {33--50},
year = {1978},
doi = {10.2307/1913643}
}

@article{meinshausen2006qrf,
author = {Meinshausen, Nicolai},
title = {Quantile Regression Forests},
journal = {Journal of Machine Learning Research},
volume = {7},
pages = {983--999},
year = {2006},
url = {https://www.jmlr.org/papers/v7/meinshausen06a.html}
}

@techreport{bishop1994mdn,
author = {Bishop, Christopher M.},
title = {Mixture Density Networks},
institution = {Aston University},
year = {1994},
url = {https://publications.aston.ac.uk/id/eprint/373/}
}

@article{pedregosa2011scikit,
author = {Pedregosa, Fabian and Varoquaux, Ga{\"e}l and Gramfort, Alexandre and Michel, Vincent and Thirion, Bertrand and Grisel, Olivier and Blondel, Mathieu and Prettenhofer, Peter and Weiss, Ron and Dubourg, Vincent and Vanderplas, Jake and Passos, Alexandre and Cournapeau, David and Brucher, Matthieu and Perrot, Matthieu and Duchesnay, {\'E}douard},
title = {Scikit-learn: Machine Learning in Python},
journal = {Journal of Machine Learning Research},
volume = {12},
pages = {2825--2830},
year = {2011},
url = {https://www.jmlr.org/papers/v12/pedregosa11a.html}
}

@inproceedings{seabold2010statsmodels,
author = {Seabold, Skipper and Perktold, Josef},
title = {statsmodels: Econometric and statistical modeling with Python},
booktitle = {Proceedings of the 9th Python in Science Conference},
year = {2010},
doi = {10.25080/Majora-92bf1922-011}
}

@article{vanbuuren2011mice,
author = {van Buuren, Stef and Groothuis-Oudshoorn, Karin},
title = {mice: Multivariate Imputation by Chained Equations in R},
journal = {Journal of Statistical Software},
volume = {45},
number = {3},
pages = {1--67},
year = {2011},
doi = {10.18637/jss.v045.i03}
}

@manual{dorazio2022statmatch,
author = {D'Orazio, Marcello},
title = {StatMatch: Statistical Matching or Data Fusion},
year = {2022},
note = {R package},
url = {https://CRAN.R-project.org/package=StatMatch}
}

@unpublished{juaristi2026microimpute,
author = {Juaristi, Mar\'ia},
title = {microimpute: Benchmarking Cross-Survey Imputation Methods for U.S. Household Wealth},
year = {2026},
note = {Working paper, PolicyEngine},
url = {https://github.com/PolicyEngine/microimpute/blob/main/paper/main.pdf}
}

@article{policyengine_py,
author = {Ahmadi, Vahid and Ghenis, Max and Woodruff, Nikhil and Makarchuk, Pavel and Volk, Anthony},
title = {policyengine: A Microsimulation Tool for Tax-Benefit Policy Analysis},
journal = {Journal of Open Source Software},
publisher = {The Open Journal},
volume = {11},
number = {125},
pages = {11115},
year = {2026},
doi = {10.21105/joss.11115},
url = {https://doi.org/10.21105/joss.11115}
}
97 changes: 97 additions & 0 deletions paper.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,97 @@
---
title: "microimpute: A Model-Agnostic Tool for Cross-Survey Imputation"
tags:
- Python
- imputation
- statistical matching
- survey microdata
- quantile regression
- microsimulation
authors:
- name: Vahid Ahmadi
orcid: 0009-0004-1093-6272
affiliation: '1'
corresponding: true
- name: Max Ghenis
orcid: 0000-0002-1335-8277
affiliation: '1'
- name: María Juaristi
orcid: 0009-0007-4946-2248
affiliation: '1'
- name: Nikhil Woodruff
orcid: 0009-0009-5004-4910
affiliation: '1'
affiliations:
- name: PolicyEngine, Washington, DC, United States
index: '1'
date: 14 September 2026
bibliography: paper.bib
---

# Summary

`microimpute` imputes variables from one survey onto another and, more importantly, makes the choice of imputation method an empirical question rather than a convention. Policy microdata routinely lacks variables an analysis needs: a labour force survey records earnings but not wealth, a household survey records spending but not assets. The standard remedy is to borrow the variable from a richer donor survey conditional on characteristics both surveys observe. Many methods do this, they disagree, and the disagreement matters for the resulting estimates.

The package implements five approaches behind one `fit`/`predict` interface — statistical matching, ordinary least squares, quantile regression [@koenker1978regression], quantile regression forests [@meinshausen2006qrf], and mixture density networks [@bishop1994mdn] — and adds `autoimpute`, which cross-validates the available methods on the user's own data under five-fold cross-validation, optionally tuning hyperparameters, and selects by quantile loss for numerical targets or log loss for categorical ones. Statistical matching wraps R's `StatMatch` through `rpy2` and mixture density networks require PyTorch; both are optional extras, and `autoimpute` compares whichever methods are installed. Ordinary least squares and statistical matching accept survey weights directly, fitting by weighted least squares and by weighted donor selection respectively, so survey design need not be discarded at the imputation step; quantile regression and mixture density networks raise an explicit error rather than silently returning an unweighted fit.

The design follows from an empirical finding rather than a preference. Benchmarking across six further datasets, alongside the wealth application, shows no method dominating across all of them: quantile regression forests win where relationships are nonlinear, and matching better preserves marginal distributions because it draws from the donor pool directly, with ordinary least squares and quantile regression occupying middle ranks [@juaristi2026microimpute]. With six benchmark datasets, the rank differences are not robust to the inclusion or exclusion of any single dataset. If method performance is dataset-specific, the useful tool is one that measures it.

# Statement of Need

Imputation choices are usually invisible in published analysis. A study reports a distributional result; the imputation method that produced the underlying variable is a sentence in an appendix, if it appears at all. Yet the choice can move headline numbers. In PolicyEngine's US model, imputing household wealth from the Survey of Consumer Finances onto the Current Population Survey is what makes asset tests bind at all: the model otherwise defaults countable resources to zero, so the baseline count of Supplemental Security Income recipients is overestimated by 167%, and a reform to the SSI asset limit cannot be simulated [@juaristi2026microimpute].

Analysts nonetheless tend to pick one method and keep it, because comparing methods is laborious. Each has a different API, different hyperparameters, and different output — a conditional mean from a regression, a donor record from matching, a predictive distribution from a forest. Building a like-for-like comparison means writing adapters and a cross-validation harness before any comparison happens, which is enough friction that the comparison usually is not done.

`microimpute` removes that friction. Because every method returns quantiles of the conditional distribution rather than a point prediction, they can be scored on the same footing with quantile loss, and the comparison is a function call rather than a project. The package also makes the imputation reproducible: hyperparameter tuning, cross-validation, and selection run from a single entry point that records what was chosen.

# State of the Field

| Tool | Multiple methods | Automated selection | Quantile-based evaluation | Survey weights | Language |
|---|---|---|---|---|---|
| `microimpute` | 5 (3 without optional extras) | Yes | Yes | Partly | Python |
| `scikit-learn` `IterativeImputer` [@pedregosa2011scikit] | 1 family | No | No | No | Python |
| `statsmodels` MICE [@seabold2010statsmodels] | 1 family | No | No | No | Python |
| R `mice` [@vanbuuren2011mice] | Several | No | No | No | R |
| R `StatMatch` [@dorazio2022statmatch] | Matching | No | No | Yes | R |

`scikit-learn` and `statsmodels` treat imputation as filling missing values within a dataset, which is a different problem from borrowing a variable across two surveys with no overlapping records. R's `mice` is the reference implementation for multiple imputation by chained equations, and `StatMatch` for statistical matching, but neither compares across method families or selects between them, and using both means working in two idioms.

The gap `microimpute` fills is comparison. Its contribution is not a new estimator but a harness that makes existing estimators commensurable on a user's data, with an evaluation metric appropriate to distributional imputation.

# Software Design

Every model implements `fit(X_train, predictors, imputed_variables, weight_col=None)` and `predict(X_test, quantiles)`, returning quantiles of the conditional distribution. That uniformity is what makes the comparison possible: a regression and a donor-matching procedure are not obviously comparable until both are expressed as predictive distributions. Imputation is framed throughout as a donor-to-receiver problem: the donor survey observes both the predictors and the target variables, the receiver survey observes only the predictors, and the two share no records. Categorical predictors are encoded and numeric predictors standardised consistently across the two frames, so a model fitted on the donor can be applied to the receiver without the analyst reconciling schemas by hand.


```python
from microimpute.comparisons import autoimpute

result = autoimpute(
donor_data=scf,
receiver_data=cps,
predictors=["age", "income", "education"],
imputed_variables=["net_worth"],
)
```

`autoimpute` runs each available method under five-fold cross-validation on the donor data, scores it by average quantile loss across a grid of quantiles, refits the winner on the full donor sample, and applies it to the receiver, returning the imputed values together with the comparison that justified them. Categorical and boolean targets are handled with log loss, and the target type is inferred rather than declared. Because the result carries the full per-method cross-validation table, the selection is auditable after the fact rather than buried in the run.

Alongside the imputers, the package provides diagnostics for the step that usually determines imputation quality more than the estimator does: the choice of predictors. `compute_predictor_correlations` reports Pearson and Spearman correlations among candidate predictors and mutual information between each predictor and each target, while `leave_one_out_analysis` and `progressive_predictor_inclusion` measure contribution by loss: the first by the degradation when a predictor is dropped, the second by building the predictor set up one addition at a time to find an ordering and a subset. A predictor set can then be defended rather than assumed. The zero-inflated wrapper composes a model for the probability of a zero with a model for the positive part, which matters for variables such as asset holdings where a large share of the population is at zero.

Results are inspectable rather than final: the package reports per-method losses so an analyst can see how close the decision was, and a companion web dashboard, distributed separately, renders the comparison for exploration.

# Research Impact Statement

`microimpute` builds the imputed variables in the microdata underlying `policyengine` [@policyengine_py], the microsimulation model behind the analyses published at [policyengine.org](https://policyengine.org). Its SCF-to-CPS wealth imputation supplies the countable-resource inputs on which US asset-tested programme modelling depends, and its quantile regression forests impute variables into the UK microdata. It is also used in standalone studies, including a UK trade shock study and an analysis of a National Insurance contributions exemption.

The accompanying research paper documents the benchmarking exercise and the SSI application in full [@juaristi2026microimpute]; this paper describes the software.

# Acknowledgements

We thank Ben Ogorek for contributions to the package. Arnold Ventures [@arnold_ventures], NEO Philanthropy [@neo_philanthropy], the Gerald Huff Fund for Humanity, and the National Science Foundation (NSF POSE Phase I, Award 2518372) [@nsf_pose] funded this work in the US; the Nuffield Foundation has funded the UK work since September 2024 [@nuffield2024grant]. These funders had no involvement in the design, development, or content of this software or paper. All authors are employed by PolicyEngine and may benefit reputationally from the software's adoption; this relationship is disclosed as a potential conflict of interest.

# AI Usage Disclosure

The authors used generative AI tools, specifically Claude by Anthropic [@claude2026], to assist with code refactoring, test authoring, and drafting of this paper. Human authors reviewed, edited, and validated all AI-assisted outputs, and made all decisions regarding method implementations, evaluation design, and software architecture. The authors remain fully responsible for the accuracy, originality, and correctness of all submitted materials.

# References
Loading