diff --git a/.github/workflows/draft-pdf.yml b/.github/workflows/draft-pdf.yml new file mode 100644 index 00000000..d057c09a --- /dev/null +++ b/.github/workflows/draft-pdf.yml @@ -0,0 +1,23 @@ +on: + push: + paths: + - paper.md + - paper.bib + pull_request: + paths: + - paper.md + - paper.bib + +jobs: + paper: + runs-on: ubuntu-latest + name: Draft PDF + steps: + - uses: actions/checkout@v4 + - uses: openjournals/openjournals-draft-action@master + with: + journal: joss + - uses: actions/upload-artifact@v4 + with: + name: paper + path: paper.pdf diff --git a/CITATION.cff b/CITATION.cff new file mode 100644 index 00000000..f7783bdc --- /dev/null +++ b/CITATION.cff @@ -0,0 +1,31 @@ +cff-version: 1.2.0 +message: "If you use this software, please cite it as below." +type: software +title: "microimpute: A Model-Agnostic Tool for Cross-Survey Imputation" +url: "https://github.com/PolicyEngine/microimpute" +repository-code: "https://github.com/PolicyEngine/microimpute" +abstract: "A Python package for imputing variables from one survey onto another, implementing five imputation methods behind a common interface with cross-validated automatic method selection." +authors: + - family-names: Ahmadi + given-names: Vahid + orcid: "https://orcid.org/0009-0004-1093-6272" + affiliation: "PolicyEngine, Washington, DC, United States" + - family-names: Ghenis + given-names: Max + orcid: "https://orcid.org/0000-0002-1335-8277" + affiliation: "PolicyEngine, Washington, DC, United States" + - family-names: Juaristi + given-names: María + orcid: "https://orcid.org/0009-0007-4946-2248" + affiliation: "PolicyEngine, Washington, DC, United States" + - family-names: Woodruff + given-names: Nikhil + orcid: "https://orcid.org/0009-0009-5004-4910" + affiliation: "PolicyEngine, Washington, DC, United States" +keywords: + - imputation + - statistical matching + - survey microdata + - quantile regression + - microsimulation + - Python diff --git a/CODE_OF_CONDUCT.md b/CODE_OF_CONDUCT.md new file mode 100644 index 00000000..56cdab0f --- /dev/null +++ b/CODE_OF_CONDUCT.md @@ -0,0 +1,64 @@ +# Contributor covenant code of conduct + +## Our pledge + +We as members, contributors, and leaders pledge to make participation in our +community a harassment-free experience for everyone, regardless of age, body +size, visible or invisible disability, ethnicity, sex characteristics, gender +identity and expression, level of experience, education, socio-economic status, +nationality, personal appearance, race, religion, or sexual identity +and orientation. + +We pledge to act and interact in ways that contribute to an open, welcoming, +diverse, inclusive, and healthy community. + +## Our standards + +Examples of behavior that contributes to a positive environment for our +community include: + +* Demonstrating empathy and kindness toward other people +* Being respectful of differing opinions, viewpoints, and experiences +* Giving and gracefully accepting constructive feedback +* Accepting responsibility and apologizing to those affected by our mistakes, + and learning from the experience +* Focusing on what is best not just for us as individuals, but for the + overall community + +Examples of unacceptable behavior include: + +* The use of sexualized language or imagery, and sexual attention or + advances of any kind +* Trolling, insulting or derogatory comments, and personal or political attacks +* Public or private harassment +* Publishing others' private information, such as a physical or email + address, without their explicit permission +* Other conduct which could reasonably be considered inappropriate in a + professional setting + +## Enforcement responsibilities + +Community leaders are responsible for clarifying and enforcing our standards of +acceptable behavior and will take appropriate and fair corrective action in +response to any behavior that they deem inappropriate, threatening, offensive, +or harmful. + +## Scope + +This Code of Conduct applies within all community spaces, and also applies when +an individual is officially representing the community in public spaces. + +## Enforcement + +Instances of abusive, harassing, or otherwise unacceptable behavior may be +reported to the community leaders responsible for enforcement at +hello@policyengine.org. All complaints will be reviewed and investigated +promptly and fairly. + +## Attribution + +This Code of Conduct is adapted from the [Contributor Covenant][homepage], +version 2.0, available at +https://www.contributor-covenant.org/version/2/0/code_of_conduct.html. + +[homepage]: https://www.contributor-covenant.org diff --git a/architecture.png b/architecture.png new file mode 100644 index 00000000..155cd215 Binary files /dev/null and b/architecture.png differ diff --git a/architecture.svg b/architecture.svg new file mode 100644 index 00000000..4c84c962 --- /dev/null +++ b/architecture.svg @@ -0,0 +1,50 @@ + + + + + + + + + + + + + + + Donor + Observes predictors + and target variables + (e.g. Survey of Consumer Finances) + + + + Receiver + Observes only the + shared predictors + (e.g. Current Population Survey) + + + + + + + + Methods + Matching, ordinary least squares, quantile + regression, quantile regression forests, MDN + (one fit/predict interface, scored by quantile loss) + + + Cross-validation + (five folds on the donor) + + + + + + + Result + Imputed variables in the receiver + (with per-method losses and the fitted model) + diff --git a/changelog.d/joss-paper.added.md b/changelog.d/joss-paper.added.md new file mode 100644 index 00000000..abbe4e56 --- /dev/null +++ b/changelog.d/joss-paper.added.md @@ -0,0 +1 @@ +- Added a JOSS paper, a citation file, a code of conduct, and a workflow that builds a draft PDF of the paper. diff --git a/paper.bib b/paper.bib new file mode 100644 index 00000000..efd8b5e1 --- /dev/null +++ b/paper.bib @@ -0,0 +1,124 @@ +@misc{arnold_ventures, + title={Public Finance Program}, + author={{Arnold Ventures}}, + year={2023}, + note={Grant to PolicyEngine for congressional district-level policy analysis}, + url={https://www.arnoldventures.org/work/public-finance} +} + +@misc{nsf_pose, + title={{POSE}: Phase {I}: {PolicyEngine} -- Advancing Public Policy Analysis}, + author={{National Science Foundation}}, + year={2025}, + note={Award 2518372. PI: Max Ghenis, PSL Foundation. \$299,974}, + url={https://www.nsf.gov/awardsearch/showAward.jsp?AWD_ID=2518372} +} + +@online{neo_philanthropy, + title={{NEO Philanthropy} Awards \$200,000 Grant to {PolicyEngine}}, + author={Ghenis, Max}, + year={2024}, + url={https://policyengine.org/us/research/neo-philanthropy} +} + +@software{claude2026, + title={{Claude}}, + author={{Anthropic}}, + year={2026}, + note={Opus 4 model used for code refactoring assistance}, + url={https://www.anthropic.com/claude} +} + +@misc{nuffield2024grant, + title={Enhancing, localising and democratising tax-benefit policy analysis}, + author={{Nuffield Foundation}}, + year={2024}, + note={General Election Analysis and Briefing Fund grant to PolicyEngine}, + url={https://www.nuffieldfoundation.org/project/enhancing-localising-and-democratising-tax-benefit-policy-analysis} +} +@article{koenker1978regression, + author = {Koenker, Roger and Bassett, Gilbert}, + title = {Regression Quantiles}, + journal = {Econometrica}, + volume = {46}, + number = {1}, + pages = {33--50}, + year = {1978}, + doi = {10.2307/1913643} +} + +@article{meinshausen2006qrf, + author = {Meinshausen, Nicolai}, + title = {Quantile Regression Forests}, + journal = {Journal of Machine Learning Research}, + volume = {7}, + pages = {983--999}, + year = {2006}, + url = {https://www.jmlr.org/papers/v7/meinshausen06a.html} +} + +@techreport{bishop1994mdn, + author = {Bishop, Christopher M.}, + title = {Mixture Density Networks}, + institution = {Aston University}, + year = {1994}, + url = {https://publications.aston.ac.uk/id/eprint/373/} +} + +@article{pedregosa2011scikit, + author = {Pedregosa, Fabian and Varoquaux, Ga{\"e}l and Gramfort, Alexandre and Michel, Vincent and Thirion, Bertrand and Grisel, Olivier and Blondel, Mathieu and Prettenhofer, Peter and Weiss, Ron and Dubourg, Vincent and Vanderplas, Jake and Passos, Alexandre and Cournapeau, David and Brucher, Matthieu and Perrot, Matthieu and Duchesnay, {\'E}douard}, + title = {Scikit-learn: Machine Learning in Python}, + journal = {Journal of Machine Learning Research}, + volume = {12}, + pages = {2825--2830}, + year = {2011}, + url = {https://www.jmlr.org/papers/v12/pedregosa11a.html} +} + +@inproceedings{seabold2010statsmodels, + author = {Seabold, Skipper and Perktold, Josef}, + title = {statsmodels: Econometric and statistical modeling with Python}, + booktitle = {Proceedings of the 9th Python in Science Conference}, + year = {2010}, + doi = {10.25080/Majora-92bf1922-011} +} + +@article{vanbuuren2011mice, + author = {van Buuren, Stef and Groothuis-Oudshoorn, Karin}, + title = {mice: Multivariate Imputation by Chained Equations in R}, + journal = {Journal of Statistical Software}, + volume = {45}, + number = {3}, + pages = {1--67}, + year = {2011}, + doi = {10.18637/jss.v045.i03} +} + +@manual{dorazio2022statmatch, + author = {D'Orazio, Marcello}, + title = {StatMatch: Statistical Matching or Data Fusion}, + year = {2022}, + note = {R package}, + url = {https://CRAN.R-project.org/package=StatMatch} +} + +@unpublished{juaristi2026microimpute, + author = {Juaristi, Mar\'ia}, + title = {microimpute: Benchmarking Cross-Survey Imputation Methods for U.S. Household Wealth}, + year = {2026}, + note = {Working paper, PolicyEngine}, + url = {https://github.com/PolicyEngine/microimpute/blob/main/paper/main.pdf} +} + +@article{policyengine_py, + author = {Ahmadi, Vahid and Ghenis, Max and Woodruff, Nikhil and Makarchuk, Pavel and Volk, Anthony}, + title = {policyengine: A Microsimulation Tool for Tax-Benefit Policy Analysis}, + journal = {Journal of Open Source Software}, + publisher = {The Open Journal}, + volume = {11}, + number = {125}, + pages = {11115}, + year = {2026}, + doi = {10.21105/joss.11115}, + url = {https://doi.org/10.21105/joss.11115} +} diff --git a/paper.md b/paper.md new file mode 100644 index 00000000..02cbaefd --- /dev/null +++ b/paper.md @@ -0,0 +1,103 @@ +--- +title: "microimpute: A Model-Agnostic Tool for Cross-Survey Imputation" +tags: + - Python + - imputation + - statistical matching + - survey microdata + - quantile regression + - microsimulation +authors: + - name: Vahid Ahmadi + orcid: 0009-0004-1093-6272 + affiliation: '1' + corresponding: true + - name: Max Ghenis + orcid: 0000-0002-1335-8277 + affiliation: '1' + - name: María Juaristi + orcid: 0009-0007-4946-2248 + affiliation: '1' + - name: Nikhil Woodruff + orcid: 0009-0009-5004-4910 + affiliation: '1' +affiliations: + - name: PolicyEngine, Washington, DC, United States + index: '1' +date: 14 September 2026 +bibliography: paper.bib +--- + +# Summary + +`microimpute` imputes variables from one survey onto another and, more importantly, makes the choice of imputation method an empirical question rather than a convention. Policy microdata routinely lacks variables an analysis needs: a labour force survey records earnings but not wealth, a household survey records spending but not assets. The standard remedy is to borrow the variable from a richer donor survey conditional on characteristics both surveys observe. Many methods do this, they disagree, and the disagreement matters for the resulting estimates. + +The package implements five approaches behind one `fit`/`predict` interface — statistical matching, ordinary least squares, quantile regression [@koenker1978regression], quantile regression forests [@meinshausen2006qrf], and mixture density networks [@bishop1994mdn] — and adds `autoimpute`, which cross-validates the available methods on the user's own data under five-fold cross-validation, optionally tuning hyperparameters, and selects by quantile loss for numerical targets or a categorical loss for categorical ones. Statistical matching wraps R's `StatMatch` through `rpy2` and mixture density networks require PyTorch; both are optional extras, and `autoimpute` compares whichever methods are installed. Ordinary least squares accepts survey weights directly and fits by weighted least squares; quantile regression forests accept a weight column, which is passed to the underlying forest. Quantile regression and mixture density networks raise an explicit error rather than silently returning an unweighted fit, so survey design is never discarded without the analyst knowing. + +The design follows from an empirical finding rather than a preference. Benchmarking across six further datasets, alongside the wealth application, shows no method dominating across all of them: quantile regression forests win where relationships are nonlinear, while matching achieves the lowest mean rank overall because it draws from the donor pool directly and better preserves marginal distributions, with ordinary least squares and quantile regression occupying middle ranks [@juaristi2026microimpute]. With six benchmark datasets, the rank differences are not robust to the inclusion or exclusion of any single dataset. If method performance is dataset-specific, the useful tool is one that measures it. + +# Statement of Need + +Imputation choices are usually invisible in published analysis. A study reports a distributional result; the imputation method that produced the underlying variable is a sentence in an appendix, if it appears at all. Yet the choice can move headline numbers. In PolicyEngine's US model, imputing household wealth from the Survey of Consumer Finances onto the Current Population Survey is what makes asset tests bind at all: the model otherwise defaults countable resources to zero, so the baseline count of Supplemental Security Income recipients is overestimated by 167%, and a reform to the SSI asset limit cannot be simulated [@juaristi2026microimpute]. + +Analysts nonetheless tend to pick one method and keep it, because comparing methods is laborious. Each has a different API, different hyperparameters, and different output — a conditional mean from a regression, a donor record from matching, a predictive distribution from a forest. Building a like-for-like comparison means writing adapters and a cross-validation harness before any comparison happens, which is enough friction that the comparison usually is not done. + +`microimpute` removes that friction. Because the methods are expressed as predictive distributions rather than point predictions, they are scored on the same footing with quantile loss across a common grid, and the comparison is a function call rather than a project. The package also makes the imputation reproducible: hyperparameter tuning, cross-validation, and selection run from a single entry point that records what was chosen. + +# State of the Field + +\renewcommand{\arraystretch}{1.5} + +| | `microimpute` | `scikit-learn` | `statsmodels` | R `mice` | R `StatMatch` | +|---|---|---|---|---|---| +| Multiple methods | 5 | 1 family | 1 family | Several | Matching | +| Automated selection | Yes | No | No | No | No | +| Quantile-based evaluation | Yes | No | No | No | No | +| Survey weights | Partly | No | No | No | Partly | +| Language | Python | Python | Python | R | R | + +\renewcommand{\arraystretch}{1.0} + +Three of `microimpute`'s five methods install with the package; statistical matching and mixture density networks are optional extras. `scikit-learn`'s `IterativeImputer` [@pedregosa2011scikit] and `statsmodels`' MICE [@seabold2010statsmodels] treat imputation as filling missing values within a dataset, which is a different problem from borrowing a variable across two surveys with no overlapping records. R's `mice` [@vanbuuren2011mice] is the reference implementation for multiple imputation by chained equations, and `StatMatch` [@dorazio2022statmatch] for statistical matching — the latter supporting donor weights in its random and rank hot-deck routines, though not in its distance-based nearest-neighbour hot deck — but neither compares across method families or selects between them, and using both means working in two idioms. + +The gap `microimpute` fills is comparison. Its contribution is not a new estimator but a harness that makes existing estimators commensurable on a user's data, with an evaluation metric appropriate to distributional imputation. + +# Software Design + +Every model implements `fit(X_train, predictors, imputed_variables, weight_col=None)` and `predict(X_test, quantiles)`, returning quantiles of the conditional distribution. That uniformity is what makes the comparison possible: a regression and a donor-matching procedure are not obviously comparable until both are expressed as predictive distributions. Imputation is framed throughout as a donor-to-receiver problem: the donor survey observes both the predictors and the target variables, the receiver survey observes only the predictors, and the two share no records. Categorical predictors are encoded consistently across the two frames, so a model fitted on the donor can be applied to the receiver without the analyst reconciling schemas by hand. Optional numeric transformations — log, inverse hyperbolic sine and standardisation — are available for both frames. + +![How `microimpute` works. A donor survey observing both the predictors and the targets, a receiver observing only the predictors, and a set of candidate methods feed a cross-validated comparison, which returns the imputed variables alongside the losses that chose the method.](architecture.png){width="62%"} + + +```python +from microimpute.comparisons import autoimpute + +result = autoimpute( + donor_data=scf, + receiver_data=cps, + predictors=["age", "income", "education"], + imputed_variables=["net_worth"], +) +``` + +`autoimpute` runs each available method under five-fold cross-validation on the donor data, scores it by average quantile loss across a grid of quantiles, refits the winner on the full donor sample, and applies it to the receiver, returning the imputed values together with the comparison that justified them. Categorical and boolean targets are handled with log loss, and the target type is inferred rather than declared. Because the result carries the full per-method cross-validation table, the selection is auditable after the fact rather than buried in the run. + +Alongside the imputers, the package provides diagnostics for the step that usually determines imputation quality more than the estimator does: the choice of predictors. `compute_predictor_correlations` reports Pearson and Spearman correlations among candidate predictors and mutual information between each predictor and each target, while `leave_one_out_analysis` and `progressive_predictor_inclusion` measure contribution by loss: the first by the degradation when a predictor is dropped, the second by building the predictor set up one addition at a time to find an ordering and a subset. A predictor set can then be defended rather than assumed. The zero-inflated wrapper composes a model for the probability of a zero with a model for the positive part, which matters for variables such as asset holdings where a large share of the population is at zero. + +Results are inspectable rather than final: the package reports per-method losses so an analyst can see how close the decision was, and a companion web dashboard renders the comparison for exploration. + +# Research Impact Statement + +`microimpute` builds the imputed variables in the microdata underlying `policyengine` [@policyengine_py], the microsimulation model behind the analyses published at [policyengine.org](https://policyengine.org). Its SCF-to-CPS wealth imputation supplies the countable-resource inputs on which US asset-tested programme modelling depends, and its quantile regression forests impute variables into the UK microdata. It is also used in standalone studies, including a UK trade shock study and an analysis of a National Insurance contributions exemption. + +The accompanying research paper documents the benchmarking exercise and the SSI application in full [@juaristi2026microimpute]; this paper describes the software. + +# Acknowledgements + +We thank Ben Ogorek for contributions to the package. Arnold Ventures [@arnold_ventures], NEO Philanthropy [@neo_philanthropy], the Gerald Huff Fund for Humanity, and the National Science Foundation (NSF POSE Phase I, Award 2518372) [@nsf_pose] funded this work in the US; the Nuffield Foundation has funded the UK work since September 2024 [@nuffield2024grant]. These funders had no involvement in the design, development, or content of this software or paper. All authors are employed by PolicyEngine and may benefit reputationally from the software's adoption; this relationship is disclosed as a potential conflict of interest. + +# AI Usage Disclosure + +The authors used generative AI tools, specifically Claude by Anthropic [@claude2026], to assist with code refactoring, test authoring, and drafting of this paper. Human authors reviewed, edited, and validated all AI-assisted outputs, and made all decisions regarding method implementations, evaluation design, and software architecture. The authors remain fully responsible for the accuracy, originality, and correctness of all submitted materials. + +# References