Motivation
Max's question (Slack, 2026-07-23): the current approach doesn't generate synthetic firm microdata — how do RAND/Urban do it for the health-insurance models? Researched against the primary methodology documents (HIPSM 2011/2013/2020 docs, RAND WR-650/OP313/IJM/PMC, CBO WP 2021-15 series, Census CES WP-14-46), with adversarial verification of 25 extracted claims (22 confirmed, 3 refuted). Full report posted on #192. This issue scopes what we adopt.
What the field actually does (verified)
No incumbent model has firm microdata or worker-firm linkage either. All three major models synthesize the firm side from survey inputs and certify it against aggregates:
- Urban HIPSM: firms built entirely from the person file — workers partitioned by ESI-offer status × region × industry × firm size; each worker seeds a firm ("core employee" / "nucleus") populated with statistically similar coworkers; firm weights calibrated to SUSB counts; offer/eligibility imputed by regression (2005 CPS supplement, the last with offer data) calibrated to published MEPS-IC tables; premiums from covered lives' expenditures + size/industry admin load, calibrated to MEPS-IC and Kaiser/HRET.
- RAND COMPARE: real Kaiser/HRET employer-survey records statistically matched to SIPP workers on region / firm size / industry / offer status; sequential reweighting (SIPP workforce = truth → Kaiser/HRET weights → firm counts to SUSB by size); premiums in region×size pools (P = E(me)×AV×(1+δ)); offer decision is firm-level utility maximization validated against %-of-payroll (~11%) and offer-elasticity (≈ −0.5) benchmarks.
- CBO HISIM2 (WP 2021-15; extracted but not fully verified — read before citing): worker-anchored synthetic firms where coworker-draw probabilities are estimated from linked administrative tax data — the strongest precedent for a defensibly non-random assignment.
None of them audits worker-to-firm assignment directly. Validation everywhere is status-quo replication + aggregate targets: SUSB/BDS firm counts by size class, MEPS-IC offer/take-up/premium tables by firm size, KFF/HRET premiums and self-funding shares, literature offer elasticities by size, %-of-payroll benchmarks. All public.
Proposed scope (not for the current lock)
Path A — HIPSM-style partition-and-populate, zero new microdata. Person-side partition axes from our existing readers (imputed offer status × region × industry section × IC2 firm-size band); seed firms from core workers; populate with similar coworkers; calibrate firm weights to SUSB counts by size/industry/region (targets already in the B-side stack). Division: partitioning + offer imputation = Workstream A; SUSB weighting + validation harness = Workstream B.
Path B (COMPARE-style match to KFF EHBS microdata) stays contingent on open question 2 below.
The E12 question — proposed as a §13 referee item
E12 (the gate that catches random worker-to-firm assignment) can't lock as a linkage gate, and the current end-state reading is "phase 2 explicitly uncertified." The research adds a third option: since no incumbent certifies assignment by linkage — the field-standard bar is aggregate replication — E12 could be re-scoped to an aggregate assignment audit: the synthetic assignment must reproduce QWI/J2J earnings-by-firm-size-by-industry joint distributions (targets already committed), which would exceed HIPSM/COMPARE practice. Options for the referee round:
- Keep E12 as a linkage gate → unlockable → phase 2 stays uncertified (status quo, stricter than field practice — a legitimate choice, but it should be chosen knowingly).
- Re-scope E12 to the aggregate assignment audit → phase 2 certifiable at (above) the field-standard bar.
- Both: aggregate-E12 locks now as the certification bar; the linkage formulation is retained as a registered aspiration pending a data source that could ever power it.
Proposing this as an explicit §13 item so round 2 rules on it rather than hardening option 1 by default. No change to the current IC3 block is proposed — first-lock stays firm-size-keyed as drafted.
Open questions (close before design freeze)
- Read CBO WP 2021-15 in full — does the admin-tax-estimated coworker draw imply an estimation our SIPP job-level data could support an analogue of?
- Is KFF EHBS microdata obtainable on terms compatible with an open-source project? (Determines whether Path B exists.)
- What does the current HIPSM vintage use in place of the 2005 CPS supplement for offer imputation — is there a newer public offer-data source?
- Exact form of the aggregate-E12 statistic and its noise floor (same machinery as the existing floor batteries).
Sequencing
After the IC3 lock and the five _v1 floor promotions. Nothing here blocks or modifies round 2 except the proposed §13 item, which is a question, not a design.
🤖 Generated with Claude Code
Motivation
Max's question (Slack, 2026-07-23): the current approach doesn't generate synthetic firm microdata — how do RAND/Urban do it for the health-insurance models? Researched against the primary methodology documents (HIPSM 2011/2013/2020 docs, RAND WR-650/OP313/IJM/PMC, CBO WP 2021-15 series, Census CES WP-14-46), with adversarial verification of 25 extracted claims (22 confirmed, 3 refuted). Full report posted on #192. This issue scopes what we adopt.
What the field actually does (verified)
No incumbent model has firm microdata or worker-firm linkage either. All three major models synthesize the firm side from survey inputs and certify it against aggregates:
None of them audits worker-to-firm assignment directly. Validation everywhere is status-quo replication + aggregate targets: SUSB/BDS firm counts by size class, MEPS-IC offer/take-up/premium tables by firm size, KFF/HRET premiums and self-funding shares, literature offer elasticities by size, %-of-payroll benchmarks. All public.
Proposed scope (not for the current lock)
Path A — HIPSM-style partition-and-populate, zero new microdata. Person-side partition axes from our existing readers (imputed offer status × region × industry section × IC2 firm-size band); seed firms from core workers; populate with similar coworkers; calibrate firm weights to SUSB counts by size/industry/region (targets already in the B-side stack). Division: partitioning + offer imputation = Workstream A; SUSB weighting + validation harness = Workstream B.
Path B (COMPARE-style match to KFF EHBS microdata) stays contingent on open question 2 below.
The E12 question — proposed as a §13 referee item
E12 (the gate that catches random worker-to-firm assignment) can't lock as a linkage gate, and the current end-state reading is "phase 2 explicitly uncertified." The research adds a third option: since no incumbent certifies assignment by linkage — the field-standard bar is aggregate replication — E12 could be re-scoped to an aggregate assignment audit: the synthetic assignment must reproduce QWI/J2J earnings-by-firm-size-by-industry joint distributions (targets already committed), which would exceed HIPSM/COMPARE practice. Options for the referee round:
Proposing this as an explicit §13 item so round 2 rules on it rather than hardening option 1 by default. No change to the current IC3 block is proposed — first-lock stays firm-size-keyed as drafted.
Open questions (close before design freeze)
Sequencing
After the IC3 lock and the five
_v1floor promotions. Nothing here blocks or modifies round 2 except the proposed §13 item, which is a question, not a design.🤖 Generated with Claude Code