Replace legacy SPM calculations with canonical rolling CE and ACS forecasts - #36
Merged
Merged
Conversation
BLS reissued corrected 2019-2024 SPM thresholds on 2026-07-17 after finding errors in its threshold-generation code, re-anchoring the revised-methodology series at 82% (previously 83%). Auditing this package against the correction revealed that the hand-entered HISTORICAL_THRESHOLDS dict matched no BLS or Census publication for 2019-2020 and 2022-2023, with errors up to 8% — several times larger than the BLS correction itself (at most 1.6%). - Package the official corrected workbook (SHA-256 recorded) and generate the threshold series from it via scripts/build_threshold_series.py: full precision, 2005-2024, standard errors and tenure population shares included. - Bundle three series: bls-corrected-2026-07-17 (default), census-published-pre-correction (cross-verified against P60-275/277/ 280/283/287), and package-legacy-0.3 (verbatim, for reproducibility). - Add a weekly drift-watch workflow that re-downloads the BLS workbook and opens an issue on divergence. - Fix four bugs in the CE-PUMD replication that benchmarking surfaced: recall-window annualization (x2 -> x4), telephone double-count (UTIL already contains TELEPH), phantom MRTPRIN/INFOTECH columns (real principal outlays are EMRTPNO*/MRTPRNO*; FMLI has no internet summary), and the threshold formula itself (BLS computes 0.82 * (1.2 * FCSUti_E - SU_E + SU_Eh) over a pooled 47-53rd percentile estimation sample, not per-tenure percentiles). - Rebuild the CE downloader on the current PUMD year-bundle layout (per-quarter zips no longer exist) with a persistent cache and a curl-cffi fallback for bls.gov's TLS bot detection ([ce] extra). - Benchmark the replication against both reference series (scripts/benchmark_bls_replication.py): 1-4.5% mean absolute deviation per year with no in-kind imputation. Analysis in docs/bls-2026-correction.md. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
BLS re-estimates thresholds from the rolling CE window each year, so the published series moves with consumption as well as prices; pure CPI aging under-projected by 2.2%/yr on average over 2020-2024. Backtesting three projection rules against the corrected series (scripts/backtest_threshold_projection.py) selects a 50/50 blend of realized FCSUti-composite CPI aging and the CE replication growth ratio: 1.35%/yr mean absolute error vs 2.23% for All-Items CPI-U. - Package a 2025 nowcast built with the blend (nowcast_thresholds(), data/nowcast/nowcast_2025.json) with method, per-tenure components, and caveats; superseded when BLS publishes actual 2025 thresholds ~September 2026. - Fix a silent BLS API truncation: unregistered requests cap at 10-year spans and return only the FIRST years of longer requests; fetch_bls_cpi_series now chunks long spans. - Handle the April 2023 CE food redesign per row: 2024Q2+ FMLI files replace FOOD/FDHOME with GROCER (food and nonfood groceries); food at home is 80% of GROCER per the BLS errata. A frame-wide column check zeroed food for redesign-era quarters in pooled windows, making replicated 2025 thresholds fall 4-5% nominal before the fix. - 2025 CPI annual averages are 11-month means (October 2025 release canceled during the federal shutdown); documented in the nowcast. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The 2.2%/yr figure is mean absolute error; CPI-U aging understated threshold growth in four of five backtest years (signed mean -1.9%/yr) and overstated it in 2021. Caught by the paper's red-team review. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Year selector labels 2025 "(nowcast)" and later years "(forecast)"; the results header and base-threshold card show a warning badge and an artifact-derived disclaimer (label text ships in spm_config.json straight from the packaged nowcast document) with a link to the working paper at spm-threshold-paper.vercel.app. - generate_geoadj_data.py now imports thresholds, projections, and the nowcast from the package instead of carrying a fourth hand-copied threshold dict — the failure mode this whole change removes. Metro adjustment ratios keep the pre-correction workbook vintage in both numerator and denominator (same-vintage rule documented inline); composed metro thresholds equal the workbook rescaled onto the corrected base, unchanged from the tested convention. - spm_config.json regenerated from 0.4.0: corrected full-precision bases 2005-2024, nowcast 2025, price forecasts 2026-2030. - Methodology explainer states the corrected series and the BLS CE window (T-5)Q2-(T)Q1, derives the 2024 renter base from data instead of a hardcoded $39,430, and links the paper. - Census metro workbook download falls back to the committed asset for offline regeneration. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
PRECOMPUTED_FCSUTI_FACTORS was a table of two-significant-digit
guesses ("~4.0% annual") serving as the offline inflation path, and
the last resort was a flat 4%/yr estimate — the same hand-entered-data
genre as the threshold errors this branch corrects. Both are gone:
- scripts/build_cpi_store.py generates data/bls/cpi_annual.json
(2005-2025 annual averages for all nine CPI series, including the
internet series the composite previously lacked offline) from the
BLS API with per-series FRED fallback and cross-source validation
at 0.1% tolerance; provenance records each series' source. 2025
values are BLS's official eleven-month annual averages (October
2025 release not published during the shutdown).
- get_fcsuti_inflation_factor resolves live API -> packaged store ->
ValueError. No guesses remain; a year outside every source raises
instead of fabricating.
- CPI_PROJECTIONS is now the package's only hand-entered table and is
labeled as a stated assumption pending a reproducibly fetchable CBO
vintage (cbo.gov blocks non-browser clients).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… review) A cross-model adversarial review (gpt-5.6-sol via codex) of the working paper found that the FCSUti composite summed raw CPI index levels across series with different reference bases, weighting components by level as well as share: shelter's effective 2024 weight was 55 percent against a stated 47, telephone's 1.2 against 4. The same construction appeared in get_fcsuti_cpi and in the backtest/nowcast scripts. Every component is now rebased to the composite's base year before weighting. Regenerated downstream: replication benchmark (deflators changed), projection backtest (composite MAE 1.40 -> 1.57 percent, matching the reviewer's predicted repair; the replication growth ratio improves to 0.41 percent and now ranks first), and the packaged 2025 nowcast (blend values move ~0.1-0.3 percent: renter 40,755.98, owner with mortgage 41,036.34, owner without 34,135.99). The blend remains the committed primary estimate — re-selecting the rule after two looks at the backtest would be selection on noise — and the September evaluation reports all rules, as already committed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Observed 2026-08-06: spm_thresholds_2024.htm again displays the pre-correction 2024 values (39,430/39,068/32,586) with no correction notice, while the 2023 page is stubbed and the corrected workbook carries different numbers. The workbook comparison cannot see that inconsistency, so the weekly check now also fetches the latest per-year page and warns (non-fatal - the inconsistency is BLS-side) when the superseded vintage appears. Verified live: workbook still matches all 20 years within $0.005; the new warning fires on the current 2024 page. Documented in the paper as a dated footnote (spm-threshold-paper b77fe52). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…view) David Trimmer's review of the working paper caught the pre-repair blend MAE (1.35%) surviving in nowcast_2025.json's own method field, compute_nowcast_2025.py, and the nowcast/forecast docstrings, plus a caveat still carrying the falsified "all rules biased low in 2022-2024" claim and the pre-repair -0.6/-2.4 range. A stale value surviving in an artifact's provenance field is this project's thesis happening to the project. All five sites now state the repaired numbers (blend 0.76%/yr, replication 0.41% ranked first, composite 1.57%, CPI-U 2.23%) with the committed-primary rationale. nowcast_2025.json regenerated; values and components verified byte-identical - only method/caveats changed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
SPM projections need to advance both expenditure and rent windows. This replaces the legacy threshold, county and congressional-district paths with one portable rolling CE/ACS artifact shared by standalone calculations and the PolicyEngine, Microcosm Frame and native Axiom adapters.
The candidate uses published BLS national bases and housing shares through 2025, then conditional CE-trend and zero-real-growth scenarios through 2035. Counties select year-specific SPM estimation areas—MSAs and Census residual groups. National bases, housing shares, geography, assumptions and source vintages carry separate status and hashes. Missing geography and unsupported inputs fail explicitly.
The standalone Next app uses the same artifact, with area search, one shared header, forecast/scenario controls, annual comparisons, support warnings and reproducible Python examples. Its initial data payload is about 110 KB gzip; the complete 26.6 MB audit artifact is a separate download. Application provenance stays in the results footnote.
Validation includes:
b5c821f, with no skipped, unexpected or flaky cases. Coverage includes search/header/mobile/forecast/provenance, real failed and stalled request recovery, unavailable-area transitions, persistent Massachusetts/Sumter warnings and thin-support/topcoding diagnostics. Both mounts downloaded the exact canonical artifact. All eight standard-CI jobs at this head pass; final public-URL acceptance remains a post-publication check.On the fixed BuildP population with native SPM role enrichment, the isolated 2025 SPM change raises the updated country model's poverty rate from 12.7543% to 13.2354% (+0.4811 percentage points). Separate country-model changes offset part of this, giving a combined change from the preserved production baseline of 13.0697% to 13.2354%. For 2026 CE trend, the combined change is 13.1732% to 13.6345%; the isolated SPM effect is +1.2189 percentage points and has a separate decomposition. These are authenticated historical development runs on Core3.30.1, not deployed results or qualification of the final Core3.32.5 runtime. That final runtime requires fresh qualification.
Six historical country-candidate and four unpublished-wrapper development runs preserve poverty masks, rates and decompositions. All five SPM amounts and the raw housing cap match the canonical storage conversion. Three 2026 person-resource values in one Montana SPM unit differ by one float32 ULP (1.56 cents), with no poverty-classification effect. The narrow exception was proposed after observing that result; it is not a predeclared or previously approved tolerance. The preserved original outputs and explicit disposition remain subject to final Fable review.
Proposed package version: 1.0.0. Forecast content SHA256:
3d86d5c4c0423480e6b69b75d222ffa4a7a2639e4094df5ba2504af01be17173. Final Fable agreement, protective consumer releases, actual registry publication, certification using published bytes, and coordinated downstream promotion remain required. The app honestly labels this buildlocal_preview. Historical forecast commitments and retrospective experiments remain immutable.Related: PolicyEngine/spm-threshold-paper#2 and #3; PolicyEngine/policyengine-us#9428; PolicyEngine/microcosm#894; protective pins PolicyEngine/policyengine-api#3824 and PolicyEngine/policyengine-sim-api#675.