Skip to content

Replace legacy SPM calculations with canonical rolling CE and ACS forecasts - #36

Merged
MaxGhenis merged 62 commits into
mainfrom
max/spm-release-rebuild-20260908
Sep 11, 2026
Merged

MaxGhenis merged 62 commits into
mainfrom
max/spm-release-rebuild-20260908

Conversation

@MaxGhenis

@MaxGhenis MaxGhenis commented Sep 8, 2026 •

Copy link
Copy Markdown
Contributor

SPM projections need to advance both expenditure and rent windows. This replaces the legacy threshold, county and congressional-district paths with one portable rolling CE/ACS artifact shared by standalone calculations and the PolicyEngine, Microcosm Frame and native Axiom adapters.

The candidate uses published BLS national bases and housing shares through 2025, then conditional CE-trend and zero-real-growth scenarios through 2035. Counties select year-specific SPM estimation areas—MSAs and Census residual groups. National bases, housing shares, geography, assumptions and source vintages carry separate status and hashes. Missing geography and unsupported inputs fail explicitly.

The standalone Next app uses the same artifact, with area search, one shared header, forecast/scenario controls, annual comparisons, support warnings and reproducible Python examples. Its initial data payload is about 110 KB gzip; the complete 26.6 MB audit artifact is a separate download. Application provenance stays in the results footnote.

Validation includes:

  • All six Python3.9–3.14 suite jobs pass at exact-head run 34502068721, alongside lint and web checks. The current3.14 job uses Python3.14.7 and reports507passed/81existing optional skips. Earlier recorded505-pass runs on specific patch versions remain historical evidence. The publisher runs the same six-version matrix. Runtime/scientific bytes are unchanged by these CI additions.
  • 115 web tests and both standalone/subpath Next production builds pass. Hosted Chromium run 34502068729 passes10/10 cases per mount at b5c821f, with no skipped, unexpected or flaky cases. Coverage includes search/header/mobile/forecast/provenance, real failed and stalled request recovery, unavailable-area transitions, persistent Massachusetts/Sumter warnings and thin-support/topcoding diagnostics. Both mounts downloaded the exact canonical artifact. All eight standard-CI jobs at this head pass; final public-URL acceptance remains a post-publication check.
  • The scheduled BLS drift alert no longer depends on an unprovisioned GitHub label. Three workflow tests execute the actual issue-creation/comment shell against a controlled gh command, including absence of the label. Independent reviews verify these changes and the browser assertions.
  • Earlier recorded qualification includes 2,261,520 exact canonical comparisons, 140 real Frame/Axiom checks, 28 documentation Python examples and 14 CLI checks. Native Axiom executes the actual rules engine; its documented decimal/binary-float boundary differences and adapter scope remain explicit.

On the fixed BuildP population with native SPM role enrichment, the isolated 2025 SPM change raises the updated country model's poverty rate from 12.7543% to 13.2354% (+0.4811 percentage points). Separate country-model changes offset part of this, giving a combined change from the preserved production baseline of 13.0697% to 13.2354%. For 2026 CE trend, the combined change is 13.1732% to 13.6345%; the isolated SPM effect is +1.2189 percentage points and has a separate decomposition. These are authenticated historical development runs on Core3.30.1, not deployed results or qualification of the final Core3.32.5 runtime. That final runtime requires fresh qualification.

Six historical country-candidate and four unpublished-wrapper development runs preserve poverty masks, rates and decompositions. All five SPM amounts and the raw housing cap match the canonical storage conversion. Three 2026 person-resource values in one Montana SPM unit differ by one float32 ULP (1.56 cents), with no poverty-classification effect. The narrow exception was proposed after observing that result; it is not a predeclared or previously approved tolerance. The preserved original outputs and explicit disposition remain subject to final Fable review.

Proposed package version: 1.0.0. Forecast content SHA256: 3d86d5c4c0423480e6b69b75d222ffa4a7a2639e4094df5ba2504af01be17173. Final Fable agreement, protective consumer releases, actual registry publication, certification using published bytes, and coordinated downstream promotion remain required. The app honestly labels this build local_preview. Historical forecast commitments and retrospective experiments remain immutable.

Related: PolicyEngine/spm-threshold-paper#2 and #3; PolicyEngine/policyengine-us#9428; PolicyEngine/microcosm#894; protective pins PolicyEngine/policyengine-api#3824 and PolicyEngine/policyengine-sim-api#675.

MaxGhenis and others added 30 commits July 17, 2026 20:44
BLS reissued corrected 2019-2024 SPM thresholds on 2026-07-17 after
finding errors in its threshold-generation code, re-anchoring the
revised-methodology series at 82% (previously 83%). Auditing this
package against the correction revealed that the hand-entered
HISTORICAL_THRESHOLDS dict matched no BLS or Census publication for
2019-2020 and 2022-2023, with errors up to 8% — several times larger
than the BLS correction itself (at most 1.6%).

- Package the official corrected workbook (SHA-256 recorded) and
  generate the threshold series from it via
  scripts/build_threshold_series.py: full precision, 2005-2024,
  standard errors and tenure population shares included.
- Bundle three series: bls-corrected-2026-07-17 (default),
  census-published-pre-correction (cross-verified against P60-275/277/
  280/283/287), and package-legacy-0.3 (verbatim, for reproducibility).
- Add a weekly drift-watch workflow that re-downloads the BLS workbook
  and opens an issue on divergence.
- Fix four bugs in the CE-PUMD replication that benchmarking surfaced:
  recall-window annualization (x2 -> x4), telephone double-count (UTIL
  already contains TELEPH), phantom MRTPRIN/INFOTECH columns (real
  principal outlays are EMRTPNO*/MRTPRNO*; FMLI has no internet
  summary), and the threshold formula itself (BLS computes
  0.82 * (1.2 * FCSUti_E - SU_E + SU_Eh) over a pooled 47-53rd
  percentile estimation sample, not per-tenure percentiles).
- Rebuild the CE downloader on the current PUMD year-bundle layout
  (per-quarter zips no longer exist) with a persistent cache and a
  curl-cffi fallback for bls.gov's TLS bot detection ([ce] extra).
- Benchmark the replication against both reference series
  (scripts/benchmark_bls_replication.py): 1-4.5% mean absolute
  deviation per year with no in-kind imputation. Analysis in
  docs/bls-2026-correction.md.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
BLS re-estimates thresholds from the rolling CE window each year, so
the published series moves with consumption as well as prices; pure
CPI aging under-projected by 2.2%/yr on average over 2020-2024.
Backtesting three projection rules against the corrected series
(scripts/backtest_threshold_projection.py) selects a 50/50 blend of
realized FCSUti-composite CPI aging and the CE replication growth
ratio: 1.35%/yr mean absolute error vs 2.23% for All-Items CPI-U.

- Package a 2025 nowcast built with the blend (nowcast_thresholds(),
  data/nowcast/nowcast_2025.json) with method, per-tenure components,
  and caveats; superseded when BLS publishes actual 2025 thresholds
  ~September 2026.
- Fix a silent BLS API truncation: unregistered requests cap at
  10-year spans and return only the FIRST years of longer requests;
  fetch_bls_cpi_series now chunks long spans.
- Handle the April 2023 CE food redesign per row: 2024Q2+ FMLI files
  replace FOOD/FDHOME with GROCER (food and nonfood groceries); food
  at home is 80% of GROCER per the BLS errata. A frame-wide column
  check zeroed food for redesign-era quarters in pooled windows,
  making replicated 2025 thresholds fall 4-5% nominal before the fix.
- 2025 CPI annual averages are 11-month means (October 2025 release
  canceled during the federal shutdown); documented in the nowcast.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The 2.2%/yr figure is mean absolute error; CPI-U aging understated
threshold growth in four of five backtest years (signed mean -1.9%/yr)
and overstated it in 2021. Caught by the paper's red-team review.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Year selector labels 2025 "(nowcast)" and later years "(forecast)";
  the results header and base-threshold card show a warning badge and
  an artifact-derived disclaimer (label text ships in spm_config.json
  straight from the packaged nowcast document) with a link to the
  working paper at spm-threshold-paper.vercel.app.
- generate_geoadj_data.py now imports thresholds, projections, and the
  nowcast from the package instead of carrying a fourth hand-copied
  threshold dict — the failure mode this whole change removes. Metro
  adjustment ratios keep the pre-correction workbook vintage in both
  numerator and denominator (same-vintage rule documented inline);
  composed metro thresholds equal the workbook rescaled onto the
  corrected base, unchanged from the tested convention.
- spm_config.json regenerated from 0.4.0: corrected full-precision
  bases 2005-2024, nowcast 2025, price forecasts 2026-2030.
- Methodology explainer states the corrected series and the BLS CE
  window (T-5)Q2-(T)Q1, derives the 2024 renter base from data instead
  of a hardcoded $39,430, and links the paper.
- Census metro workbook download falls back to the committed asset for
  offline regeneration.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
PRECOMPUTED_FCSUTI_FACTORS was a table of two-significant-digit
guesses ("~4.0% annual") serving as the offline inflation path, and
the last resort was a flat 4%/yr estimate — the same hand-entered-data
genre as the threshold errors this branch corrects. Both are gone:

- scripts/build_cpi_store.py generates data/bls/cpi_annual.json
  (2005-2025 annual averages for all nine CPI series, including the
  internet series the composite previously lacked offline) from the
  BLS API with per-series FRED fallback and cross-source validation
  at 0.1% tolerance; provenance records each series' source. 2025
  values are BLS's official eleven-month annual averages (October
  2025 release not published during the shutdown).
- get_fcsuti_inflation_factor resolves live API -> packaged store ->
  ValueError. No guesses remain; a year outside every source raises
  instead of fabricating.
- CPI_PROJECTIONS is now the package's only hand-entered table and is
  labeled as a stated assumption pending a reproducibly fetchable CBO
  vintage (cbo.gov blocks non-browser clients).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… review)

A cross-model adversarial review (gpt-5.6-sol via codex) of the working
paper found that the FCSUti composite summed raw CPI index levels across
series with different reference bases, weighting components by level as
well as share: shelter's effective 2024 weight was 55 percent against a
stated 47, telephone's 1.2 against 4. The same construction appeared in
get_fcsuti_cpi and in the backtest/nowcast scripts. Every component is
now rebased to the composite's base year before weighting.

Regenerated downstream: replication benchmark (deflators changed),
projection backtest (composite MAE 1.40 -> 1.57 percent, matching the
reviewer's predicted repair; the replication growth ratio improves to
0.41 percent and now ranks first), and the packaged 2025 nowcast
(blend values move ~0.1-0.3 percent: renter 40,755.98, owner with
mortgage 41,036.34, owner without 34,135.99). The blend remains the
committed primary estimate — re-selecting the rule after two looks at
the backtest would be selection on noise — and the September
evaluation reports all rules, as already committed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Observed 2026-08-06: spm_thresholds_2024.htm again displays the
pre-correction 2024 values (39,430/39,068/32,586) with no correction
notice, while the 2023 page is stubbed and the corrected workbook
carries different numbers. The workbook comparison cannot see that
inconsistency, so the weekly check now also fetches the latest
per-year page and warns (non-fatal - the inconsistency is BLS-side)
when the superseded vintage appears. Verified live: workbook still
matches all 20 years within $0.005; the new warning fires on the
current 2024 page.

Documented in the paper as a dated footnote (spm-threshold-paper
b77fe52).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…view)

David Trimmer's review of the working paper caught the pre-repair
blend MAE (1.35%) surviving in nowcast_2025.json's own method field,
compute_nowcast_2025.py, and the nowcast/forecast docstrings, plus a
caveat still carrying the falsified "all rules biased low in
2022-2024" claim and the pre-repair -0.6/-2.4 range. A stale value
surviving in an artifact's provenance field is this project's thesis
happening to the project.

All five sites now state the repaired numbers (blend 0.76%/yr,
replication 0.41% ranked first, composite 1.57%, CPI-U 2.23%) with
the committed-primary rationale. nowcast_2025.json regenerated;
values and components verified byte-identical - only method/caveats
changed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@MaxGhenis MaxGhenis changed the title Rebuild SPM releases with rolling expenditure and rent forecasts Replace legacy SPM calculations with canonical rolling CE and ACS forecasts Sep 9, 2026
@MaxGhenis
MaxGhenis marked this pull request as ready for review September 11, 2026 15:31
@MaxGhenis
MaxGhenis merged commit 22bab36 into main Sep 11, 2026
13 checks passed

This branch was successfully deployed

1 active deployment
Preview — b5c821f6 Deployed Sep 10, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant