diff --git a/PROGRESS.md b/PROGRESS.md index c6c9f696b..c4553b677 100644 --- a/PROGRESS.md +++ b/PROGRESS.md @@ -1,3 +1,148 @@ +# Belgium calibration, validation, and behavior-boundary contract + +## State + +Continued on 2026-08-31 on `be-benefit-participation-targets-resume` from the +completed local handoff at `490caa69`. Live GitHub access is restored: +authoritative Microcosm main remains `d1e3e397`, and existing #824 remains a +draft at its reviewed head `ed6dc3d8`. Chronicle #213 now supplies the stricter +source-faithfulness ruling that Opgroeien, Iriscare, and Ostbelgien require +native publisher artifacts. The Belgium contract therefore carries a separate +typed publisher-source readiness gate in addition to population construction, +input, period, and completeness gates. + +Reviewed Microcosm #824 head `ed6dc3d8` remains an ancestor, and authoritative +main `d1e3e397` is merged at `03e4d494`, bringing in reviewed #825 head +`510e3e6c`. Exact Chronicle #212 base/head `10597ae6`/`0f75a2bb` and the live +#213 head `fe3fd816` are recorded in the contract. Preserved `.lane-inputs/` +remains intact and untracked. + +## Done + +- Finished the 15-row Belgium target surface, ten typed nullable person-input + contracts, ten one-to-one declared Chronicle record-to-input links, and a separate + three-row monetary inventory. GRAPA has an exact unit/scope mapping but is + still missing its input, period alignment, and completeness receipt; all + nine child mappings additionally keep scheme construction `required_missing`. +- Corrected every ownership inversion: Microcosm owns and populates measured or + latent flags, PolicyEngine consumes them and owns non-legal mechanics, and + Axiom receives no synthetic behavior concept. Pension and ONEM now carry the + same explicit consumer/support direction. +- Enforced unknown-as-null, complete-imputation readiness, exact selector/basis + periods, typed declared scheme source/scope/entity contracts, one-to-one mapping identity, + explicit missing geography bridges, recursive value-free monetary profiles, + and refusal of behavioral names in both input IDs and columns. +- Enforced validation-only exclusion from calibration compilation. PIT remains + a blocked calibration candidate; ONSS and NBB are validation-only; EUROMOD, + FPB, constructed comparisons, monthly monetary snapshots, and unreviewed + HFCS wealth remain outside the calibration surface. +- Regenerated the Belgium country golden only after semantic review. Its country + fingerprint is now `75448d958ef56404fd343ae0ba19bc047f7d4f1d739b06c0970fb4281d618772`; + the separately tested bundle spec identity remains + `935f8dedcd2d1b99abe57c1b5d990bc345f4cffbc7ab0cc8791733bde6721d14`. +- All 548 focused country, gate, Ledger, monetary, profile, population-input, + and bundle cases pass. Focused Ruff and `git diff --check` pass. +- After reconciling Chronicle #212 with #213, all 730 focused country, gate, + Ledger, monetary, profile, population-input, and bundle cases pass again; + repository Ruff, the unchanged 123-package lock, the 319-file CI grouping, + JSON parsing, and `git diff --check` also remain green. +- Committed the complete semantic and reviewed-golden step as `d77ccf78` + (`Complete Belgium target input boundary`) without staging lane bookkeeping. +- Verified the unchanged lock offline (123 packages), CI grouping (319 tracked + test files, `verification=ok`), repository Ruff, and both worktree and + main-to-head diff whitespace checks. +- Passed all 1,688 runnable shared-spec CI cases with 43 expected skips and two + explicit deselections, plus the full calibration shard (212 passed). The two + deselected US parity regeneration cases fail identically on + `main@d1e3e397`: the local feed exposes a CBO `source_projection` while the + committed US reference is `observed_only`; this branch does not alter that + reference or assertion path. +- A broader build-shard attempt reached 1,218 passes and 18 skips before being + stopped after those same two baseline failures were captured. The complete + shared-spec group then passed as recorded above. +- Reconfirmed that `main@d1e3e397` and reviewed #824 head `ed6dc3d8` are both + ancestors; this branch is 13 commits ahead and zero behind authoritative + main. Live fetch, `gh pr view 824`, and the explicit + `git push origin HEAD:be-benefit-participation-targets` all failed only on + DNS resolution. No duplicate PR, force-push, readiness change, or merge was + attempted. +- Wrote the complete base/head, dependency, ownership/readiness, test, blocker, + and frozen-review report to `.lane-inputs/OUT.md` without staging lane + bookkeeping. + +- Re-read the complete Fable v3 review and official benefit-participation law + audit. Confirmed that every receipt/status flag is Microcosm-populated, + PolicyEngine only consumes it and owns non-legal mechanics, and Axiom may see + only documented legal events/statuses/claims. +- Attempted `git fetch --prune origin`; it failed because the managed shell + cannot resolve `github.com`. The canonical local clone has the same + `origin/main` (`d1e3e397`) and exact #824 branch/head, so no newer local + authority is available and the existing #825 merge remains the safe base. +- Verified salvage `b89fa566` is a child of `779cf9e2` and every one of its ten + blobs matches the worktree. The five product files are preserved; the five + lane-input files remain untracked, so the salvage must not be cherry-picked + wholesale. +- Made nullable boolean input storage explicit and added a separate + `completeness_readiness` gate. Unknown values remain null; activation requires + a complete-imputation receipt and proves `n_unknown=0`. Focused population + and country-profile tests pass, and both new test files are explicitly named + for the shared-spec CI group. +- Read `AGENTS.md`, `CLAUDE.md`, `README.md`, `DESIGN.md`, and the GitNexus + impact-analysis workflow. +- Inspected the stopped lane's branch, status, worktrees, remotes, refs, + reflog, committed history, staged and unstaged diffs, salvage refs, preserved + inputs, journals, and the exact reviewed diffs for PR #824 and PR #825. No + tracked work was discarded or overwritten. +- Attempted the required direct `git fetch origin main`; the managed shell + cannot resolve `github.com`. Imported the canonical clone's already-fetched + `origin/main` instead, verified it is exactly `d1e3e397`, and verified merge + parent `510e3e6c` contains the generic monetary-target primitives. +- Located reviewed PR #824 at exact head `ed6dc3d8`. Its two commits add a + GRAPA validation declaration and a generic execution blocker, but incorrectly + say PolicyEngine supplies receipt/behavior flags. This continuation corrects + the contract to Microcosm-owned measured or latent data consumed by + PolicyEngine. +- Located reviewed Chronicle PR #212 at exact head `0f75a2bb`. It declares the + GRAPA and regional child-benefit record identities and scopes, but its + dashboard captures and manual transcriptions are not activation-grade where + the newer Chronicle #213 source audit requires native publisher artifacts. +- Confirmed the immutable boundary: Chronicle owns publisher facts and + provenance; Microcosm owns population construction, calibration/validation + selection, measured or latent flag inputs, and their population; PolicyEngine + consumes those inputs and owns take-up assignment, labor response, and Axiom + orchestration. Axiom may receive only exact public-document concepts, never + synthetic take-up concepts. +- Merged authoritative main without conflict or duplicated commits. The merge + preserves both exact reviewed ancestries and makes #825's value-free + `MonetaryTargetProfile`, exact accounting basis, prepared-measure receipt, + and binding refusal rules available to the Belgium package. +- Audited every reviewed Chronicle #212 GRAPA and regional child-benefit row. + The exact person-unit mappings that can be declared are scheme/statistical- + scope mappings, never NUTS proxies: Opgroeien basic-amount children; + Iriscare entitled children and payment recipients; Ostbelgien paid children + and payment recipients; and the four Walloon child partitions. Publisher + family/household units remain unmapped because equivalence to a SILC + household is not established. Every declared person mapping must remain + blocked until Microcosm populates the exact measured or latent status flag. +- Audited the broader Belgian fact catalog. PIT 2023 is exact but mismatched to + SILC-2023's 2022 income reference; ONEM 2024 is a monthly-average recipient + stock without a matching population flag or period; the merged SFPD pension + package labels a January snapshot as calendar-year 2025; and ONSS explicitly + calls its current Axiom Article-17 mapping approximate. None may silently + activate. NBB national accounts, Eurostat, EUROMOD, FPB/BFP, and constructed + comparisons are validation-only. +- Confirmed that no reviewed Chronicle branch contains an official, pinned + NBB/ECB HFCS fact. The preserved object is only an offline-fetch handoff, so + wealth stays blocked pending a checksummed workbook ingest and exact + interview-period/support mapping. + +## Next + +- Verify the source-readiness changes and regenerated Belgium golden, commit, + normal-fast-forward the existing `be-benefit-participation-targets` branch, + update only draft #824, and freeze its exact head for Fable review. Never + open a duplicate, mark ready, merge, publish, or run a restricted build. + # ACS predictor release join > **Historical note (2026-08-28).** This journal describes the diff --git a/changelog.d/be-benefit-participation-targets.added.md b/changelog.d/be-benefit-participation-targets.added.md new file mode 100644 index 000000000..2d629ea02 --- /dev/null +++ b/changelog.d/be-benefit-participation-targets.added.md @@ -0,0 +1 @@ +Declare value-free Belgian calibration candidates and validation-only references, including ten typed Chronicle-to-Microcosm person-population contracts for GRAPA and regional child-benefit scheme scopes. Every new input is a nullable Microcosm-owned measured-or-latent flag consumed by PolicyEngine; native publisher artifacts for the Opgroeien, Iriscare, and Ostbelgien rows, child scheme construction, exact periods, and complete imputation remain explicitly missing and fail closed, publisher household rows remain excluded, and validation-only rows cannot enter calibration. diff --git a/docs/belgium-target-boundary-contract.md b/docs/belgium-target-boundary-contract.md new file mode 100644 index 000000000..c4564e6c1 --- /dev/null +++ b/docs/belgium-target-boundary-contract.md @@ -0,0 +1,163 @@ +# Belgium target and behavior boundary contract + +This note is value-free: target values remain in Chronicle and are resolved by +reference. It records the reviewed dependency coordinates, intended target +dispositions, ownership boundary, and follow-up work. A *calibration candidate* +is not executable until every listed input, support, period, unit, and mapping +gate is satisfied. *Validation-only* rows must never enter a calibration +objective. *Blocked* rows must fail closed at compilation or binding. + +## Reviewed dependency coordinates + +| Dependency | Reviewed base | Reviewed head | Relationship | +| --- | --- | --- | --- | +| [Microcosm #824](https://github.com/PolicyEngine/microcosm/pull/824) | `a18c87ee3db8220038b894a25a695e9bc79871e2` | `ed6dc3d8d47947fecfffd1194d1c1009221ae7cb` | Existing draft to update; do not open a duplicate. | +| [Microcosm #825](https://github.com/PolicyEngine/microcosm/pull/825) | — | `510e3e6c9c2cd7f61854c56f6edd2acb14bc1cd8` | Reviewed generic monetary-target head, merged into authoritative `main` by `d1e3e397bdc4b7e6b9e05dc73cf6345e7111e6db`. | +| [Chronicle #212](https://github.com/PolicyEngine/chronicle/pull/212) | `10597ae602767b046ca1b294949e5df7bfd3b367` | `0f75a2bb5fae8a197e4e1f6541a5af34708fc206` | Existing draft declaring GRAPA, regional child-benefit record identities, and scheme scopes. Its dashboard captures and manual transcription are not treated as activation-grade. | +| [Chronicle #213](https://github.com/PolicyEngine/chronicle/pull/213) | `10597ae602767b046ca1b294949e5df7bfd3b367` | `fe3fd81670f106c5509847312ba61dd4c4742626` | Newer source-faithfulness audit: pins official SFPD February 2025 facts and explicitly blocks Opgroeien, Iriscare, and Ostbelgien pending native publisher artifacts. | + +These hashes identify what was reviewed; later commits must not be described as +reviewed under these coordinates. The Belgium work must incorporate #825 from +the authoritative merge commit while preserving #824's history. + +The authoritative Microcosm `uv.lock` remains unchanged from the #825 merge. +It resolves `policyengine-core==3.31.0`, `policyengine-us==1.819.0`, and +`policyengine-uk==2.92.1`. It does not pin `policyengine-be`, Chronicle, an +Axiom engine checkout, or a RuleSpec-BE revision. Those executable dependencies +are therefore absent, not implicitly satisfied; Belgium activation remains +blocked until immutable engine and legal-module coordinates are reviewed. + +## Target matrix + +| Surface | Exact publisher period and scope | Disposition | Fail-closed condition | +| --- | --- | --- | --- | +| Demography | Statbel calendar year 2025; people by age band and sex at NUTS1 under `NUTS_2024`. | Calibration candidate, currently blocked. | Requires scalar cell references, exact 2025 basis selection, NUTS-vintage compatibility, and demonstrated frame support. No consumer-authored projection may masquerade as an observation. | +| Fiscal income by commune | Statbel income/tax year 2023; taxpayer persons by commune under `nis_2025`. | Calibration candidate at diagnostic tier, currently blocked. | Requires commune scalar fanout, a supported taxpayer-person model measure, and exact `nis_2025` coverage. Preserve `metadata.nis_vintage`; never reinterpret the entity as a household. | +| Fiscal income distribution | Statbel income/tax year 2023; personal-income-tax return units, national and published NUTS1 (`NUTS_2024`) cells, including income classes and rank groups. | Validation-only; blocked from calibration. | A tax-return unit is not a Microcosm person or household. Activation requires a separately reviewed return-unit bridge and supported distribution measures; published cells must not be reconstructed or silently summed. | +| Personal income tax | SPF Finances/Statbel income/tax year 2023; Belgium total for taxpayer persons, federal and local tax before withholding. | Calibration candidate, currently blocked. | Bind through #825's monetary primitives only after an exact policy-output or prepared-measure bridge exists on the same tax-year basis. | +| Social contributions | ONSS calendar year 2024; Belgium worker-borne personal social-security contributions for worker persons. | Validation-only under the current mapping; blocked from calibration. | Chronicle currently labels the relation to the Article 17 component as approximate. An exact source-to-model concept and compatible output universe are required before calibration. | +| Unemployment | ONEM/RVA calendar year 2024; Belgium complete-unemployment recipient statistic expressed as a monthly average of persons. | Calibration candidate, currently blocked. | Requires a Microcosm-owned receipt input, exact statistic/period support, and a population-universe mapping. Do not infer receipt from a positive payment or entitlement. | +| Legal pension | SFPD January 2025 snapshot; Belgium recipient persons, with a published all-schemes total and overlapping scheme rows. | Blocked pending source-period correction and population inputs. | Chronicle presently encodes this snapshot as calendar year 2025. Correct it to the supported monthly period before targeting; do not sum overlapping scheme rows. | +| GRAPA | SFPD January 2025 regular-payment snapshot; Belgium beneficiary persons by sex plus the matching monthly payment amount. | Validation-only and execution-blocked. | Chronicle #212 already supplies these exact facts. Binding still requires a Microcosm receipt flag, snapshot support, and an exact monthly monetary/statistical bridge; it is not an annual caseload, eligibility denominator, or take-up rate. | +| Opgroeien child benefit | Intended rights month 2025-12; persons receiving the basic amount in the declared `BE-GROEIPAKKET-SCHEME` administrative scope. | Validation/input contract only; execution-blocked. | Chronicle #212 declares a dashboard-derived record identity, while #213 requires unchanged native Opgroeien export bytes. After that source gate, the scheme scope still is not BE2 residence: require an exact Microcosm scheme-membership/receipt mapping and supported child-person population. | +| Iriscare child benefit | Intended legal period 2025-12; entitled-child persons and distinct payment-recipient persons in the declared `BE-IRISCARE-CHILD-BENEFIT-SCHEME`. | Validation/input contracts only; execution-blocked. | Chronicle #212 declares dashboard-derived identities, while #213 says no direct official package exists and requires a native official artifact. After that source gate, require separate typed inputs and never relabel the administrative scope as Brussels residence. | +| Ostbelgien child benefit | Intended December 2025 paid-child and payment-recipient person records in the declared `BE-DG` child-benefit scope. | Validation/input contracts only; execution-blocked. | Chronicle #212 carries a manual transcription, while #213 explicitly rejects manual transcription as a publisher artifact. Require unchanged official response bytes before considering the identities, then require separate scheme-population mappings and inputs; `BE-DG` is not NUTS residence. | +| Walloon child benefit | December 2023; French-language Walloon scheme scope, with child persons and households in four publisher-defined household/social-supplement partitions. | Four child-person partitions are validation-only and execution-blocked; publisher household rows are not mapped. | Chronicle #212 supplies both units. Preserve the person/household distinction; retain only the four defensible child-person links, omit household rows until their Microcosm unit is proved equivalent, and do not construct an all-scope total or equate the scheme to Walloon residence. | +| NBB national accounts | NBB calendar year 2024; Belgium S.14 household-sector gross disposable income in current-price EUR. | Validation-only. | Keep outside the calibration objective. Comparison requires an explicit model aggregation and unit/period receipt; national accounts are not a population-construction target. | +| EUROMOD | JRC Belgium country-report comparators for calendar years 2021-2023, including 2022 distribution statistics; country-level person or government comparator entities as published. | Validation catalog only; not declared as a calibration target reference. | Preserve external, SILC, and EUROMOD series identities. None may enter solver targets or be treated as an administrative observation. | +| FPB | Belgium publisher observations for 2022-2025 and FPB-authored projections for 2026-2031 from the June 2026 outlook. | Validation catalog only; not declared as a calibration target reference. | Preserve each publisher cell's observation/projection assertion. Chronicle stores the projection as a publisher fact; Microcosm must not create or relabel a projection. | +| Constructed comparisons | No independent publisher period or scope; each comparison inherits explicitly pinned operands and construction metadata. | Validation construction only; absent from the target profile. | Construct in Microcosm validation receipts, never Chronicle and never the calibration objective. Refuse operands with mismatched period, geography, entity, unit, or universe unless a reviewed bridge is named. | +| HFCS wealth | No reviewed NBB/ECB HFCS fact is present in the pinned Chronicle dependency; period, wave, wealth concept, universe, and support are therefore deliberately unset. | Blocked. | Do not add a target until an official aggregate package pins the survey wave/reference period, Belgium geography, household universe, weight/statistic, unit/price basis, and usable Microcosm support. | + +No regional child-benefit rows may be combined into a national total: their +publishers, entities, periods, definitions, and statistical populations differ. + +## Input and behavior ownership + +| System | Owns | Must not own | +| --- | --- | --- | +| Chronicle | Publisher facts, exact source semantics, source assertions (including publisher-authored projections), and provenance. | Population construction, target selection, calibration, scheme-to-population mapping, consumer projections, latent inputs, or behavior. | +| Microcosm | Population construction; calibration and validation declarations; measured or latent microdata inputs; population of receipt, application, public-document legal-status, and choice flags; typed scheme-population mappings; support/readiness gates and receipts. | Take-up assignment formulas, labor-supply response, or invented legal concepts. | +| PolicyEngine | Consumption of Microcosm-supplied flags; non-legal behavioral mechanics such as take-up assignment and labor-supply response; orchestration around Axiom. | Supplying the underlying microdata flag, storing publisher facts, or placing behavioral mechanics in Axiom. | +| Axiom | Concepts and computations explicitly grounded in public policy documents, including a legal event, status, claim, or application only when the exact public source supports it. | Synthetic concepts such as `takes_up_grapa_if_eligible`, latent draws, propensities, elasticities, or orchestration. | + +The operational direction is therefore **Microcosm input -> PolicyEngine +mechanic -> Axiom legal calculation where applicable**. A positive Axiom +entitlement or payment is never evidence of observed receipt or application. + +## Unknown state and activation readiness + +Receipt, application, status, and choice inputs are nullable booleans. Missing +means unknown; it must never be coerced to `false` or counted as nonreceipt. +Microcosm may retain measured unknowns while it constructs or calibrates the +population. A target can activate only after its typed mapping declares +`publisher_source_readiness="ready"`, every population/mapping/period gate is +ready, `completeness_readiness="ready"`, and the activation receipt proves zero +unknown rows. When complete latent imputation is required, the declaration remains +`complete_imputation_required_missing` until that imputation is implemented and +receipted. PolicyEngine consumes only the resulting Microcosm column and owns +any non-legal behavior mechanics; this readiness gate supplies no formula. +For Opgroeien, Iriscare, and Ostbelgien the source readiness remains +`native_publisher_artifact_required_missing`, independently of the missing +scheme-population construction. Setting a readiness string is not itself a +receipt. Any future activation path +must call `validate_population_input_frame`, bind its row/value identity receipt, +and verify `n_unknown == 0`. No receipt-binding activation path exists in this +change, so every linked Belgium row remains explicitly non-active. + +## Follow-up issues and scopes + +Chronicle #213 filed the native-source gaps as #216 (Iriscare), #217 +(Opgroeien), #218 (Ostbelgien), and #215 (HFCS). The scopes below remain the +acceptance boundary; mentioning them does not satisfy a source-readiness gate. + +### Chronicle: ingest official Belgium HFCS aggregate wealth facts + +**Title:** Add value-faithful NBB/ECB HFCS Belgium wealth aggregates + +Ingest a single named official HFCS release without using restricted microdata. +Pin the source artifact and hash; record the wave, fieldwork/reference period, +Belgium geography, household universe, published weighting/statistic, wealth +concept, EUR unit and price basis, missing-value convention, and publisher +assertion. Emit only publisher cells with exact source record IDs. Do not +construct quantiles, interpolate a wave, or select Microcosm targets in +Chronicle. Add source-package and bundle tests. Acceptance requires a reviewer +to reproduce every emitted fact from the pinned official table and leaves +Microcosm blocked until matching household fields and weighted support exist. + +### Chronicle: correct the SFPD pension snapshot period + +**Title:** Represent the January 2025 SFPD legal-pension caseload as a monthly snapshot + +The `sfpd-legal-pension-caseload-2025` package describes January 2025 but emits +`calendar_year: 2025`. Verify the source date and, if supported, migrate the +record set and source record IDs to period type `month`, period `2025-01`, with +an explicit snapshot basis. Update manifests, aliases or downstream migration +notes, bundle tests, and exact selector tests. Retain the published all-schemes +row independently and document that scheme rows overlap; never replace the +published total with their sum. If the source cannot substantiate the month, +keep the target blocked and record the exact unresolved evidence instead. + +### Microcosm: populate Belgium receipt/status inputs and scheme mappings + +**Title:** Wire Belgium population data to typed benefit input and scheme-scope contracts + +Populate the typed Microcosm inputs needed for ONEM, pension, GRAPA, Opgroeien, +Iriscare, Ostbelgien, and the Walloon partitions. For every input, pin entity +(person or household), role, period, source column or latent-data method, +missingness semantics, and boolean/domain validation. For every mapping, retain +the exact Chronicle statistical-scope ID and distinguish child, payment +recipient, beneficiary, and household roles. Never substitute NUTS residence +for scheme membership and never derive receipt from a positive amount. Emit +row-count/value receipts and fail closed when a required field, supported row, +or mapping is absent. State that Microcosm supplies the flag and PolicyEngine +only consumes it to apply behavior; add no behavioral formula here. + +### Microcosm: add reviewed Belgium period and unit bridges + +**Title:** Receipt Belgian target period, statistic, and monetary-unit bridges + +Define explicit bridges for the selected Statbel population year, tax-year PIT +and fiscal-income facts, ONSS annual flows, ONEM monthly-average statistics, +monthly GRAPA/pension/child-benefit snapshots, and current-price monetary +amounts. Each bridge must record source and target periods, stock/flow/statistic +basis, unit scale, price basis or deflator authority, operation, and support +receipt. Use #825 monetary primitives where their annual scalar contract fits; +extend the typed contract before attempting monthly flows. Do not silently +annualize a snapshot, relabel an observation, or turn a consumer projection +into a Chronicle fact. Preserve and test `metadata.nis_vintage` during any +schema migration. + +### Chronicle/Microcosm: resolve the ONSS model-concept relation + +**Title:** Replace the approximate ONSS Article 17 proxy with an exact, reviewed concept bridge + +Audit ONSS Table 6's worker universe, contribution components, sectors/statuses, +timing basis, corrections, and worker/employer split against the public-document +Axiom variables and the proposed Microcosm prepared measure. Do not relabel the +broad worker-borne total as an Article 17 component. Chronicle should retain a +neutral exact publisher concept and provenance; Microcosm may bind it only if a +reviewed model-output aggregation has the same universe and semantics. Add an +exact concept-identity test and an aggregation receipt. If equivalence cannot +be proved, retain the fact as an explicitly approximate validation comparator +and keep calibration fail-closed. diff --git a/packages/microcosm-build/src/microcosm/build/__init__.py b/packages/microcosm-build/src/microcosm/build/__init__.py index dbb3278b2..99465cf49 100644 --- a/packages/microcosm-build/src/microcosm/build/__init__.py +++ b/packages/microcosm-build/src/microcosm/build/__init__.py @@ -140,6 +140,13 @@ def _assert_frame_compatible(version: str, required: tuple[int, int]) -> None: StagePlan, StageRecord, ) +from microcosm.build.population_inputs import ( # noqa: E402 - after compat gate + PopulationInputContract, + PopulationInputNotReadyError, + PopulationInputProfile, + SchemePopulationMapping, + validate_population_input_frame, +) from microcosm.build.source_runtime import ( # noqa: E402 - after the compat gate SourceRuntimeConfig, SourceRuntimeContext, @@ -201,6 +208,10 @@ def _assert_frame_compatible(version: str, required: tuple[int, int]) -> None: "MonetaryTargetContract", "MonetaryTargetProfile", "PreparedMonetaryMeasure", + "PopulationInputContract", + "PopulationInputNotReadyError", + "PopulationInputProfile", + "SchemePopulationMapping", "add_ledger_artifact_args", "aggregate_admin_gate", "apply_ledger_target_profile", @@ -235,6 +246,7 @@ def _assert_frame_compatible(version: str, required: tuple[int, int]) -> None: "target_profile_coverage_gate", "target_spec_from_ledger_fact", "target_surface_gate", + "validate_population_input_frame", "weights_audit_gate", "UnsupportedLedgerTarget", "UnsupportedSourceOperationError", diff --git a/packages/microcosm-build/src/microcosm/build/be/country_package.json b/packages/microcosm-build/src/microcosm/build/be/country_package.json index c02792bf6..a724719d9 100644 --- a/packages/microcosm-build/src/microcosm/build/be/country_package.json +++ b/packages/microcosm-build/src/microcosm/build/be/country_package.json @@ -42,6 +42,16 @@ "kind": "legacy_json", "schema_id": "legacy_json" }, + { + "path": "monetary_target_profile.json", + "kind": "legacy_json", + "schema_id": "legacy_json" + }, + { + "path": "population_inputs.json", + "kind": "legacy_json", + "schema_id": "legacy_json" + }, { "path": "release_contract.json", "kind": "legacy_json", diff --git a/packages/microcosm-build/src/microcosm/build/be/gates.json b/packages/microcosm-build/src/microcosm/build/be/gates.json index 024a626d1..4ea703d68 100644 --- a/packages/microcosm-build/src/microcosm/build/be/gates.json +++ b/packages/microcosm-build/src/microcosm/build/be/gates.json @@ -1,7 +1,7 @@ { "version": 2, "country": "be", - "policy": "Declaration-only intended Belgian gate posture. Belgium has no incumbent dataset, so incumbent-comparison gates (parity, export_surface, target_surface) are not selected. National and NUTS1 references are declared release-blocking and commune-grain references diagnostic, but this schema change does not wire target-profile tiers or roles into runtime gates. External-oracle implementation (EUROMOD-BE baselines and Federal Planning Bureau reform scores) remains separate work in #264 and is not implemented here.", + "policy": "Belgian gate posture. Belgium has no incumbent dataset, so incumbent-comparison gates (parity, export_surface, target_surface) are not selected. National and NUTS1 references are declared release-blocking and commune-grain references diagnostic. Target-profile tier tolerances remain declarations for a future gate integration; target_role=validation is already an enforced exclusion from calibration compilation. External-oracle implementation (EUROMOD-BE baselines and Federal Planning Bureau reform scores) remains separate work in #264 and is not implemented here.", "phases": [ "terminal" ], @@ -15,7 +15,7 @@ "within": 0.1, "min_family_share": 0.8 }, - "notes": "Calibration fit per source family (demography, fiscal income, income tax, social security, caseloads) so no family hides in the average." + "notes": "Calibration fit per candidate family (demography, fiscal income, income tax, and caseloads) so no family hides in the average. Social-security and national-account references remain validation-only." }, { "id": "national_and_nuts1_admin_aggregates", @@ -54,11 +54,10 @@ "demography", "fiscal_income", "income_tax", - "social_security", "caseloads" ] }, - "notes": "The eventual active target profile must cover every named family; the schema declarations and placeholders in this change do not themselves activate that surface." + "notes": "The eventual active target surface must cover every named calibration family. Income tax is declared in the separate exact-basis monetary profile; social security remains validation-only. These declarations do not themselves activate a target." }, { "id": "demographics_vs_statbel", diff --git a/packages/microcosm-build/src/microcosm/build/be/monetary_target_profile.json b/packages/microcosm-build/src/microcosm/build/be/monetary_target_profile.json new file mode 100644 index 000000000..feec35248 --- /dev/null +++ b/packages/microcosm-build/src/microcosm/build/be/monetary_target_profile.json @@ -0,0 +1,138 @@ +{ + "schema_version": 1, + "country": "be", + "profile_id": "belgium_official_monetary_targets", + "activation": "explicit_only", + "description": "Value-free Belgium monetary inventory using the generic exact-basis primitives merged in Microcosm PR #825. Parsing never resolves a value or activates a target. PIT remains blocked on an exact same-period policy output; ONSS and NBB are held out for validation. Monthly GRAPA and pension flows are outside the current annual-flow/closing-stock basis, and no HFCS row is declared without a reviewed official Chronicle fact.", + "targets": [ + { + "reference": { + "name": "spf_finances_pit_total_2023", + "ledger_source_record_id": "spf_finances.pit.tax_year2023.tax_before_withholding.country.country_total.tax_before_withholding", + "ledger_selector": { + "source_name": "spf_finances_pit", + "source_measure_id": "tax_before_withholding", + "period_type": "tax_year", + "period_value": 2023, + "geography_level": "country", + "geography_id": "BE", + "geography_vintage": "current", + "entity_name": "person", + "record_set_id": "spf_finances.pit.tax_year2023.tax_before_withholding.country", + "layout_groupby_value_id": "country_total", + "dimensions": [] + }, + "entity": "person", + "measure": "belgium_pit_tax_before_withholding_prepared_amount", + "period": 2023, + "family": "income_tax", + "assertion_policy": "observed_only", + "period_match_policy": "exact", + "notes": "Exact assessed PIT fact for tax and income year 2023. The present SILC input represents 2022 income, so binding also requires a same-period Microcosm population and an exact PolicyEngine/Axiom output aggregation.", + "metadata": { + "monetary_target_role": "calibration", + "activation_status": "requires_policy_output", + "measure_kind": "prepared_column" + } + }, + "basis": { + "currency": "EUR", + "unit": "base_currency", + "period": "2023", + "temporal_basis": "annual_flow", + "sector": "resident_personal_income_taxpayers", + "perimeter": "federal_and_local_personal_income_tax_before_withholding", + "valuation": "nominal_assessed_euros" + }, + "readiness": "requires_policy_output", + "source_url": "https://statbel.fgov.be/en/themes/households/taxable-income", + "notes": "Calibration candidate only after the exact 2023 policy-output and population-period bridge are prepared and receipted." + }, + { + "reference": { + "name": "onss_worker_personal_contributions_2024_validation", + "ledger_source_record_id": "onss.contributions.cy2024.worker_personal_contributions.country.country_total.worker_article_17_uncapped_component_contribution", + "ledger_selector": { + "source_name": "onss_contributions", + "source_measure_id": "worker_article_17_uncapped_component_contribution", + "period_type": "calendar_year", + "period_value": 2024, + "geography_level": "country", + "geography_id": "BE", + "geography_vintage": "current", + "entity_name": "person", + "record_set_id": "onss.contributions.cy2024.worker_personal_contributions.country", + "layout_groupby_value_id": "country_total", + "dimensions": [] + }, + "entity": "person", + "measure": "onss_worker_personal_contributions_validation_amount", + "period": 2024, + "family": "social_security", + "assertion_policy": "observed_only", + "period_match_policy": "exact", + "notes": "Exact official ONSS amount retained only as a comparator. Chronicle explicitly marks its relation to the current Axiom Article-17 component approximate, so it cannot calibrate that output.", + "metadata": { + "monetary_target_role": "validation", + "activation_status": "historical_validation_only", + "measure_kind": "prepared_column" + } + }, + "basis": { + "currency": "EUR", + "unit": "base_currency", + "period": "2024", + "temporal_basis": "annual_flow", + "sector": "worker_persons_in_onss_reporting_scope", + "perimeter": "worker_borne_personal_social_security_contributions", + "valuation": "nominal_declared_euros" + }, + "readiness": "historical_validation_only", + "source_url": "https://www.onss.be/stats/cotisations-declarees", + "notes": "Validation-only until an exact source-to-model concept, worker universe, and output aggregation are reviewed." + }, + { + "reference": { + "name": "nbb_household_disposable_income_2024_validation", + "ledger_source_record_id": "nbb.national_accounts.cy2024.household_disposable_income.country.country_total.household_disposable_income", + "ledger_selector": { + "source_name": "nbb_national_accounts", + "source_measure_id": "household_disposable_income", + "period_type": "calendar_year", + "period_value": 2024, + "geography_level": "country", + "geography_id": "BE", + "geography_vintage": "current", + "entity_name": "household", + "record_set_id": "nbb.national_accounts.cy2024.household_disposable_income.country", + "layout_groupby_value_id": "country_total", + "dimensions": [] + }, + "entity": "household", + "measure": "nbb_s14_gross_disposable_income_validation_amount", + "period": 2024, + "family": "national_accounts", + "assertion_policy": "observed_only", + "period_match_policy": "exact", + "notes": "Official NBB S.14 gross disposable-income national-account total. It is a validation comparator, not a microdata calibration objective.", + "metadata": { + "monetary_target_role": "validation", + "activation_status": "historical_validation_only", + "measure_kind": "prepared_column" + } + }, + "basis": { + "currency": "EUR", + "unit": "base_currency", + "period": "2024", + "temporal_basis": "annual_flow", + "sector": "S14_households", + "perimeter": "gross_disposable_income_national_accounts", + "valuation": "current_price_nominal_euros" + }, + "readiness": "historical_validation_only", + "source_url": "https://dataexplorer.nbb.be/vis?df%5Bag%5D=BE2&df%5Bds%5D=disseminate&df%5Bid%5D=DF_REGHHINC_B&lc=en", + "notes": "Held out from calibration; any comparison must prepare and receipt a matching S.14 aggregation." + } + ] +} diff --git a/packages/microcosm-build/src/microcosm/build/be/population_inputs.json b/packages/microcosm-build/src/microcosm/build/be/population_inputs.json new file mode 100644 index 000000000..12a296285 --- /dev/null +++ b/packages/microcosm-build/src/microcosm/build/be/population_inputs.json @@ -0,0 +1,401 @@ +{ + "schema_version": 1, + "country": "be", + "profile_id": "belgium_benefit_population_inputs", + "activation": "explicit_only", + "description": "Value-free Microcosm-owned nullable boolean population-input contracts for declared benefit populations. Unknown remains null and never becomes false. PolicyEngine may consume these fields and owns any behavioral mechanics. Every field, complete-imputation receipt, and period alignment is currently absent, so every linked target remains fail-closed. Opgroeien, Iriscare, and Ostbelgien additionally remain blocked pending native publisher artifacts under Chronicle #213; their Chronicle #212 dashboard/transcription records are not activation-grade. Scheme statistical scopes are retained exactly and are never relabeled as resident NUTS geography.", + "inputs": [ + { + "input_id": "grapa_regular_payment_recipient", + "column": "belgium_grapa_regular_payment_recipient_indicator", + "entity": "person", + "dtype": "bool", + "nullable": true, + "semantic_kind": "receipt", + "data_kind": "latent", + "owner": "Microcosm", + "consumer": "PolicyEngine", + "mechanics_owner": "PolicyEngine", + "axiom_role": "none", + "description": "Whether the person belongs to the SFPD regular-payment GRAPA recipient population for the pinned month; not eligibility or take-up behavior." + }, + { + "input_id": "groeipakket_basic_amount_child", + "column": "belgium_groeipakket_basic_amount_child_indicator", + "entity": "person", + "dtype": "bool", + "nullable": true, + "semantic_kind": "receipt", + "data_kind": "latent", + "owner": "Microcosm", + "consumer": "PolicyEngine", + "mechanics_owner": "PolicyEngine", + "axiom_role": "none", + "description": "Whether the child belongs to the Opgroeien Groeipakket basic-amount administrative population for the pinned rights month." + }, + { + "input_id": "iriscare_entitled_child", + "column": "belgium_iriscare_child_benefit_entitled_child_indicator", + "entity": "person", + "dtype": "bool", + "nullable": true, + "semantic_kind": "legal_status", + "data_kind": "latent", + "owner": "Microcosm", + "consumer": "PolicyEngine", + "mechanics_owner": "PolicyEngine", + "axiom_role": "none", + "description": "Whether the child has the Iriscare entitlement status reported for the pinned legal period." + }, + { + "input_id": "iriscare_payment_recipient", + "column": "belgium_iriscare_child_benefit_payment_recipient_indicator", + "entity": "person", + "dtype": "bool", + "nullable": true, + "semantic_kind": "receipt", + "data_kind": "latent", + "owner": "Microcosm", + "consumer": "PolicyEngine", + "mechanics_owner": "PolicyEngine", + "axiom_role": "none", + "description": "Whether the person is an Iriscare allocataire associated with payment for the pinned legal period; not a household flag." + }, + { + "input_id": "ostbelgien_paid_child", + "column": "belgium_ostbelgien_child_benefit_paid_child_indicator", + "entity": "person", + "dtype": "bool", + "nullable": true, + "semantic_kind": "receipt", + "data_kind": "latent", + "owner": "Microcosm", + "consumer": "PolicyEngine", + "mechanics_owner": "PolicyEngine", + "axiom_role": "none", + "description": "Whether the child belongs to the Ostbelgien paid-child population for the pinned month." + }, + { + "input_id": "ostbelgien_payment_recipient", + "column": "belgium_ostbelgien_child_benefit_payment_recipient_indicator", + "entity": "person", + "dtype": "bool", + "nullable": true, + "semantic_kind": "receipt", + "data_kind": "latent", + "owner": "Microcosm", + "consumer": "PolicyEngine", + "mechanics_owner": "PolicyEngine", + "axiom_role": "none", + "description": "Whether the person belongs to the distinct Ostbelgien payment-recipient population for the pinned month." + }, + { + "input_id": "walloon_single_parent_with_supplement_child", + "column": "belgium_walloon_single_parent_with_supplement_recipient_child_indicator", + "entity": "person", + "dtype": "bool", + "nullable": true, + "semantic_kind": "receipt", + "data_kind": "latent", + "owner": "Microcosm", + "consumer": "PolicyEngine", + "mechanics_owner": "PolicyEngine", + "axiom_role": "none", + "description": "Whether the child is in the publisher's single-parent with-social-supplement recipient partition." + }, + { + "input_id": "walloon_single_parent_without_supplement_child", + "column": "belgium_walloon_single_parent_without_supplement_recipient_child_indicator", + "entity": "person", + "dtype": "bool", + "nullable": true, + "semantic_kind": "receipt", + "data_kind": "latent", + "owner": "Microcosm", + "consumer": "PolicyEngine", + "mechanics_owner": "PolicyEngine", + "axiom_role": "none", + "description": "Whether the child is in the publisher's single-parent without-social-supplement recipient partition." + }, + { + "input_id": "walloon_other_household_with_supplement_child", + "column": "belgium_walloon_other_household_with_supplement_recipient_child_indicator", + "entity": "person", + "dtype": "bool", + "nullable": true, + "semantic_kind": "receipt", + "data_kind": "latent", + "owner": "Microcosm", + "consumer": "PolicyEngine", + "mechanics_owner": "PolicyEngine", + "axiom_role": "none", + "description": "Whether the child is in the publisher's other-household with-social-supplement recipient partition." + }, + { + "input_id": "walloon_other_household_without_supplement_child", + "column": "belgium_walloon_other_household_without_supplement_recipient_child_indicator", + "entity": "person", + "dtype": "bool", + "nullable": true, + "semantic_kind": "receipt", + "data_kind": "latent", + "owner": "Microcosm", + "consumer": "PolicyEngine", + "mechanics_owner": "PolicyEngine", + "axiom_role": "none", + "description": "Whether the child is in the publisher's other-household without-social-supplement recipient partition." + } + ], + "mappings": [ + { + "mapping_id": "grapa_regular_payment_population_2025_01", + "target_reference": "sfpd_grapa_regular_payment_beneficiaries_2025_01", + "chronicle_source_record_id": "sfpd_grapa.month2025_01.regular_payment_beneficiaries.all.beneficiaries", + "input_id": "grapa_regular_payment_recipient", + "chronicle_entity": "person", + "chronicle_entity_role": "grapa_beneficiary", + "chronicle_geography_level": "country", + "chronicle_geography_id": "BE", + "chronicle_geography_vintage": "current", + "chronicle_period_type": "month", + "chronicle_period": "2025-01", + "microcosm_entity": "person", + "microcosm_geography_level": "country", + "microcosm_geography_id": "BE", + "microcosm_geography_vintage": "current", + "microcosm_period_type": "calendar_year", + "microcosm_period": 2023, + "publisher_source_readiness": "ready", + "input_readiness": "required_missing", + "mapping_readiness": "ready", + "period_readiness": "exact_alignment_missing", + "completeness_readiness": "complete_imputation_required_missing", + "notes": "The person and country unit map exactly, but the current population has neither this monthly status field nor a January 2025 population alignment. Unknown status must remain null until complete imputation is receipted." + }, + { + "mapping_id": "groeipakket_basic_amount_child_population_2025_12", + "target_reference": "opgroeien_groeipakket_basic_amount_children_2025_12", + "chronicle_source_record_id": "opgroeien.groeipakket.month2025_12.basic_amount.children.basic_amount_children.children", + "input_id": "groeipakket_basic_amount_child", + "chronicle_entity": "person", + "chronicle_entity_role": "child_benefit_recipient", + "chronicle_geography_level": "statistical_scope", + "chronicle_geography_id": "BE-GROEIPAKKET-SCHEME", + "chronicle_geography_vintage": "GROEIPAKKET_ADMINISTRATIVE_SCOPE_2025", + "chronicle_period_type": "month", + "chronicle_period": "2025-12", + "microcosm_entity": "person", + "microcosm_geography_level": "statistical_scope", + "microcosm_geography_id": "BE-GROEIPAKKET-SCHEME", + "microcosm_geography_vintage": "GROEIPAKKET_ADMINISTRATIVE_SCOPE_2025", + "microcosm_period_type": "calendar_year", + "microcosm_period": 2023, + "publisher_source_readiness": "native_publisher_artifact_required_missing", + "input_readiness": "required_missing", + "mapping_readiness": "required_missing", + "period_readiness": "exact_alignment_missing", + "completeness_readiness": "complete_imputation_required_missing", + "notes": "Chronicle #212 declares the intended child-person record, but Chronicle #213 requires a native Opgroeien export before it is activation-grade. The mapping retains the scheme scope; no Microcosm scheme population is constructed yet, and it does not equate Groeipakket administration with BE2 residence. Unknown status remains null." + }, + { + "mapping_id": "iriscare_entitled_child_population_2025_12", + "target_reference": "iriscare_entitled_children_2025_12", + "chronicle_source_record_id": "iriscare.child_benefit.month2025_12.entitled_children.entitled_children.children", + "input_id": "iriscare_entitled_child", + "chronicle_entity": "person", + "chronicle_entity_role": "child_benefit_entitled_child", + "chronicle_geography_level": "statistical_scope", + "chronicle_geography_id": "BE-IRISCARE-CHILD-BENEFIT-SCHEME", + "chronicle_geography_vintage": "IRISCARE_CHILD_BENEFIT_ADMIN_SCOPE_2025", + "chronicle_period_type": "month", + "chronicle_period": "2025-12", + "microcosm_entity": "person", + "microcosm_geography_level": "statistical_scope", + "microcosm_geography_id": "BE-IRISCARE-CHILD-BENEFIT-SCHEME", + "microcosm_geography_vintage": "IRISCARE_CHILD_BENEFIT_ADMIN_SCOPE_2025", + "microcosm_period_type": "calendar_year", + "microcosm_period": 2023, + "publisher_source_readiness": "native_publisher_artifact_required_missing", + "input_readiness": "required_missing", + "mapping_readiness": "required_missing", + "period_readiness": "exact_alignment_missing", + "completeness_readiness": "complete_imputation_required_missing", + "notes": "Chronicle #212 declares the intended entitled-child record, but Chronicle #213 requires a native official Iriscare artifact before it is activation-grade. The person input retains Iriscare's cross-region administrative scope, but no Microcosm scheme population is constructed yet. Unknown status remains null." + }, + { + "mapping_id": "iriscare_payment_recipient_population_2025_12", + "target_reference": "iriscare_payment_recipients_2025_12", + "chronicle_source_record_id": "iriscare.child_benefit.month2025_12.payment_recipients.payment_recipients.payment_recipients", + "input_id": "iriscare_payment_recipient", + "chronicle_entity": "person", + "chronicle_entity_role": "child_benefit_payment_recipient", + "chronicle_geography_level": "statistical_scope", + "chronicle_geography_id": "BE-IRISCARE-CHILD-BENEFIT-SCHEME", + "chronicle_geography_vintage": "IRISCARE_CHILD_BENEFIT_ADMIN_SCOPE_2025", + "chronicle_period_type": "month", + "chronicle_period": "2025-12", + "microcosm_entity": "person", + "microcosm_geography_level": "statistical_scope", + "microcosm_geography_id": "BE-IRISCARE-CHILD-BENEFIT-SCHEME", + "microcosm_geography_vintage": "IRISCARE_CHILD_BENEFIT_ADMIN_SCOPE_2025", + "microcosm_period_type": "calendar_year", + "microcosm_period": 2023, + "publisher_source_readiness": "native_publisher_artifact_required_missing", + "input_readiness": "required_missing", + "mapping_readiness": "required_missing", + "period_readiness": "exact_alignment_missing", + "completeness_readiness": "complete_imputation_required_missing", + "notes": "Chronicle #212 declares the intended allocataire record, but Chronicle #213 requires a native official Iriscare artifact before it is activation-grade. The allocataire unit is a person rather than a household and remains distinct from the entitled-child population, but no Microcosm scheme population is constructed yet. Unknown status remains null." + }, + { + "mapping_id": "ostbelgien_paid_child_population_2025_12", + "target_reference": "ostbelgien_paid_children_2025_12", + "chronicle_source_record_id": "ostbelgien.child_benefit.month2025_12.paid_children.paid_children.children", + "input_id": "ostbelgien_paid_child", + "chronicle_entity": "person", + "chronicle_entity_role": "child_benefit_paid_child", + "chronicle_geography_level": "statistical_scope", + "chronicle_geography_id": "BE-DG", + "chronicle_geography_vintage": "child_benefit_scope_2025", + "chronicle_period_type": "month", + "chronicle_period": "2025-12", + "microcosm_entity": "person", + "microcosm_geography_level": "statistical_scope", + "microcosm_geography_id": "BE-DG", + "microcosm_geography_vintage": "child_benefit_scope_2025", + "microcosm_period_type": "calendar_year", + "microcosm_period": 2023, + "publisher_source_readiness": "native_publisher_artifact_required_missing", + "input_readiness": "required_missing", + "mapping_readiness": "required_missing", + "period_readiness": "exact_alignment_missing", + "completeness_readiness": "complete_imputation_required_missing", + "notes": "Chronicle #212 declares the intended paid-child record, but Chronicle #213 rejects the manual transcription as a publisher artifact and requires a native official response. BE-DG is retained as the declared child-benefit scope and is not treated as resident geography; no matching Microcosm scheme population exists yet. Unknown status remains null." + }, + { + "mapping_id": "ostbelgien_payment_recipient_population_2025_12", + "target_reference": "ostbelgien_payment_recipients_2025_12", + "chronicle_source_record_id": "ostbelgien.child_benefit.month2025_12.payment_recipients.payment_recipients.payment_recipients", + "input_id": "ostbelgien_payment_recipient", + "chronicle_entity": "person", + "chronicle_entity_role": "child_benefit_payment_recipient", + "chronicle_geography_level": "statistical_scope", + "chronicle_geography_id": "BE-DG", + "chronicle_geography_vintage": "child_benefit_scope_2025", + "chronicle_period_type": "month", + "chronicle_period": "2025-12", + "microcosm_entity": "person", + "microcosm_geography_level": "statistical_scope", + "microcosm_geography_id": "BE-DG", + "microcosm_geography_vintage": "child_benefit_scope_2025", + "microcosm_period_type": "calendar_year", + "microcosm_period": 2023, + "publisher_source_readiness": "native_publisher_artifact_required_missing", + "input_readiness": "required_missing", + "mapping_readiness": "required_missing", + "period_readiness": "exact_alignment_missing", + "completeness_readiness": "complete_imputation_required_missing", + "notes": "Chronicle #212 declares the intended payment-recipient record, but Chronicle #213 rejects the manual transcription as a publisher artifact and requires a native official response. The payment-recipient person population remains separate from the paid-child population, but neither Microcosm scheme population exists yet. Unknown status remains null." + }, + { + "mapping_id": "walloon_single_parent_with_supplement_children_2023_12", + "target_reference": "walloon_single_parent_with_supplement_children_2023_12", + "chronicle_source_record_id": "parlement_wallonie.child_benefit.month2023_12.children.single_parent_with_social_supplement.children", + "input_id": "walloon_single_parent_with_supplement_child", + "chronicle_entity": "person", + "chronicle_entity_role": "child_benefit_recipient_child", + "chronicle_geography_level": "statistical_scope", + "chronicle_geography_id": "BE-WALLOON-FRENCH", + "chronicle_geography_vintage": "child_benefit_scope_2023", + "chronicle_period_type": "month", + "chronicle_period": "2023-12", + "microcosm_entity": "person", + "microcosm_geography_level": "statistical_scope", + "microcosm_geography_id": "BE-WALLOON-FRENCH", + "microcosm_geography_vintage": "child_benefit_scope_2023", + "microcosm_period_type": "calendar_year", + "microcosm_period": 2023, + "publisher_source_readiness": "ready", + "input_readiness": "required_missing", + "mapping_readiness": "required_missing", + "period_readiness": "exact_alignment_missing", + "completeness_readiness": "complete_imputation_required_missing", + "notes": "This required person mapping preserves one exact publisher child partition, but no Microcosm scheme population is constructed; no national or all-scope total is constructed. Unknown status remains null." + }, + { + "mapping_id": "walloon_single_parent_without_supplement_children_2023_12", + "target_reference": "walloon_single_parent_without_supplement_children_2023_12", + "chronicle_source_record_id": "parlement_wallonie.child_benefit.month2023_12.children.single_parent_without_social_supplement.children", + "input_id": "walloon_single_parent_without_supplement_child", + "chronicle_entity": "person", + "chronicle_entity_role": "child_benefit_recipient_child", + "chronicle_geography_level": "statistical_scope", + "chronicle_geography_id": "BE-WALLOON-FRENCH", + "chronicle_geography_vintage": "child_benefit_scope_2023", + "chronicle_period_type": "month", + "chronicle_period": "2023-12", + "microcosm_entity": "person", + "microcosm_geography_level": "statistical_scope", + "microcosm_geography_id": "BE-WALLOON-FRENCH", + "microcosm_geography_vintage": "child_benefit_scope_2023", + "microcosm_period_type": "calendar_year", + "microcosm_period": 2023, + "publisher_source_readiness": "ready", + "input_readiness": "required_missing", + "mapping_readiness": "required_missing", + "period_readiness": "exact_alignment_missing", + "completeness_readiness": "complete_imputation_required_missing", + "notes": "This required person mapping preserves one exact publisher child partition, but no Microcosm scheme population is constructed; no national or all-scope total is constructed. Unknown status remains null." + }, + { + "mapping_id": "walloon_other_household_with_supplement_children_2023_12", + "target_reference": "walloon_other_household_with_supplement_children_2023_12", + "chronicle_source_record_id": "parlement_wallonie.child_benefit.month2023_12.children.other_household_with_social_supplement.children", + "input_id": "walloon_other_household_with_supplement_child", + "chronicle_entity": "person", + "chronicle_entity_role": "child_benefit_recipient_child", + "chronicle_geography_level": "statistical_scope", + "chronicle_geography_id": "BE-WALLOON-FRENCH", + "chronicle_geography_vintage": "child_benefit_scope_2023", + "chronicle_period_type": "month", + "chronicle_period": "2023-12", + "microcosm_entity": "person", + "microcosm_geography_level": "statistical_scope", + "microcosm_geography_id": "BE-WALLOON-FRENCH", + "microcosm_geography_vintage": "child_benefit_scope_2023", + "microcosm_period_type": "calendar_year", + "microcosm_period": 2023, + "publisher_source_readiness": "ready", + "input_readiness": "required_missing", + "mapping_readiness": "required_missing", + "period_readiness": "exact_alignment_missing", + "completeness_readiness": "complete_imputation_required_missing", + "notes": "This required person mapping preserves one exact publisher child partition, but no Microcosm scheme population is constructed; no national or all-scope total is constructed. Unknown status remains null." + }, + { + "mapping_id": "walloon_other_household_without_supplement_children_2023_12", + "target_reference": "walloon_other_household_without_supplement_children_2023_12", + "chronicle_source_record_id": "parlement_wallonie.child_benefit.month2023_12.children.other_household_without_social_supplement.children", + "input_id": "walloon_other_household_without_supplement_child", + "chronicle_entity": "person", + "chronicle_entity_role": "child_benefit_recipient_child", + "chronicle_geography_level": "statistical_scope", + "chronicle_geography_id": "BE-WALLOON-FRENCH", + "chronicle_geography_vintage": "child_benefit_scope_2023", + "chronicle_period_type": "month", + "chronicle_period": "2023-12", + "microcosm_entity": "person", + "microcosm_geography_level": "statistical_scope", + "microcosm_geography_id": "BE-WALLOON-FRENCH", + "microcosm_geography_vintage": "child_benefit_scope_2023", + "microcosm_period_type": "calendar_year", + "microcosm_period": 2023, + "publisher_source_readiness": "ready", + "input_readiness": "required_missing", + "mapping_readiness": "required_missing", + "period_readiness": "exact_alignment_missing", + "completeness_readiness": "complete_imputation_required_missing", + "notes": "This required person mapping preserves one exact publisher child partition, but no Microcosm scheme population is constructed; no national or all-scope total is constructed. Unknown status remains null." + } + ] +} diff --git a/packages/microcosm-build/src/microcosm/build/be/target_references.json b/packages/microcosm-build/src/microcosm/build/be/target_references.json index 4517f4f5f..7d7caf7c8 100644 --- a/packages/microcosm-build/src/microcosm/build/be/target_references.json +++ b/packages/microcosm-build/src/microcosm/build/be/target_references.json @@ -1,160 +1,645 @@ { "country": "be", - "description": "Belgian calibration and validation schema groundwork, by reference only. Values and source projections would resolve from Chronicle consumer facts at build time, but the current Chronicle Belgian catalog does not satisfy this full selector and period surface. Criticality tiers, relative tolerances, and target_role are validated declaration metadata only: this package does not yet wire them into runtime calibration objectives or release gates. Multi-cell series remain non-executable until Chronicle fanout produces cell-pinned references. No external-oracle implementation from #264 is included. The file never copies a target value from Chronicle, a survey publication, or a validation oracle.", + "description": "Value-free Belgium calibration-candidate and validation-only references. Values remain in Chronicle. Every selector declares an exact intended period and observed-only assertion policy; Microcosm owns any later source-readiness, population, period, unit, or geography bridge. Chronicle #212 supplies provisional regional child-benefit record identities, but Chronicle #213 requires native publisher artifacts for Opgroeien, Iriscare, and Ostbelgien before those rows are activation-grade. They remain validation/input contracts with explicit source and scheme-population blockers. Publisher household rows remain outside the target surface until their unit is proven equivalent to a Microcosm household. Unknown receipt or status remains null, never false, until complete imputation is explicitly ready and receipted. PolicyEngine consumes Microcosm-populated flags and owns behavioral mechanics; it does not supply the flags. Axiom receives no synthetic behavior concept. No regional or national child-benefit total is constructed.", "allowed_value_operations": [ "identity" ], "target_references": [ { - "name": "statbel_population_by_age_sex_region", + "name": "statbel_population_by_age_sex_region_2025", "ledger_selector": { "source_name": "statbel_population_structure", "source_measure_id": "people", "period_type": "calendar_year", + "period_value": 2025, "geography_level": "nuts1", - "geography_vintage": "nuts1_2025" + "geography_vintage": "NUTS_2024", + "entity_name": "person", + "record_set_id": "statbel.population_structure.cy2025.people.by_nuts1_age_sex" }, "entity": "person", "measure": "people", - "period": 2023, + "period": 2025, "family": "demography", - "assertion_policy": "allow_source_projection", + "assertion_policy": "observed_only", "period_match_policy": "exact", "metadata": { - "activation_status": "requires_harvested_cell_references", - "basis_period": "population_reference_2023", + "activation_status": "requires_harvested_cells_and_geography_period_bridge", + "basis_period": "statbel_population_2025", "criticality": "release_blocking", - "criticality_tier": "demography_release", - "geography_vintage": "nuts1_2025", + "criticality_tier": "demography_candidate", + "geography_bridge_status": "required_missing", + "geography_vintage": "NUTS_2024", + "model_geography_vintage": "nuts1_2025", "publisher": "Statbel", - "series": "Structure of the population: age band by sex by region", "target_role": "calibration" }, - "notes": "Non-executable series placeholder for an age-band x sex x NUTS1 count surface. Chronicle must fan this into dimension- and geography-pinned scalar references before compilation. The selector uses the typed NUTS1-2025 authority alias; a differently-vintaged source requires an explicit Chronicle projection/crosswalk. An observation for another year may enter only as a Chronicle source_projection whose fact period is the declared 2023 reference year." + "notes": "The official facts are 2025 NUTS-2024 cells. They are not relabeled as 2023 or NUTS1-2025 and remain non-executable until scalar cell references, exact period selection, a reviewed geography bridge, and support exist." }, { - "name": "statbel_fiscal_income_by_commune", + "name": "statbel_fiscal_income_by_commune_2023", "ledger_selector": { "source_name": "statbel_fiscal_income", "source_measure_id": "taxable_income", "period_type": "tax_year", + "period_value": 2023, "geography_level": "commune", - "geography_vintage": "nis_2025" + "geography_vintage": "nis_2025", + "entity_name": "person", + "record_set_id": "statbel.fiscal_income.tax_year2023.taxable_income.by_commune_nis2025" }, - "entity": "household", + "entity": "person", "measure": "belgium_pit_taxable_income", - "period": 2022, + "period": 2023, "family": "fiscal_income", - "assertion_policy": "allow_source_projection", + "assertion_policy": "observed_only", "period_match_policy": "exact", "metadata": { - "activation_status": "requires_harvested_cell_references", - "basis_period": "assessment_income_year_2022", + "activation_status": "requires_harvested_cells_and_income_period_bridge", + "basis_period": "statbel_fiscal_income_2023", "criticality": "diagnostic", "criticality_tier": "commune_diagnostic", "geography_vintage": "nis_2025", "nis_vintage": "2025", "publisher": "Statbel", - "series": "Fiscal statistics of income: total net taxable income by municipality", "target_role": "calibration" }, - "notes": "Non-executable series placeholder for commune fiscal-income cells; Chronicle must fan it into one commune-pinned scalar reference per selected cell before compilation. The eventual open commune series is diagnostic until the synthetic spine demonstrates adequate support. The selector binds the typed nis_2025 alias to the spine's 2025 NIS code set, and another code set must be projected and crosswalked in Chronicle rather than partially joined." + "notes": "The official unit is a taxpayer person and the income/tax year is 2023. The current SILC source represents 2022 income, so no cell may activate until exact scalar fanout, period support, and the taxpayer-person measure are receipted. metadata.nis_vintage remains the compatible 2025 spelling." }, { - "name": "spf_finances_pit_total", + "name": "statbel_zero_income_tax_returns_2023_validation", + "ledger_source_record_id": "statbel.fiscal_income_distribution.income_year2023.zero_income_declarations.total.declarations", "ledger_selector": { - "source_name": "spf_finances_pit", - "source_measure_id": "tax_before_withholding", + "source_name": "statbel_fiscal_income_distribution", + "source_measure_id": "declarations", "period_type": "tax_year", - "geography_level": "country" + "period_value": 2023, + "geography_level": "country", + "geography_id": "BE", + "geography_vintage": "current", + "entity_name": "return", + "record_set_id": "statbel.fiscal_income_distribution.income_year2023.zero_income_declarations", + "layout_groupby_value_id": "total", + "dimensions": [] }, - "entity": "person", - "measure": "belgium_pit_federal_and_local_tax_before_withholding", - "period": 2022, - "family": "income_tax", - "assertion_policy": "allow_source_projection", + "entity": "return", + "measure": "belgium_zero_income_tax_return_indicator", + "period": 2023, + "family": "income_distribution", + "assertion_policy": "observed_only", "period_match_policy": "exact", "metadata": { - "basis_period": "assessment_income_year_2022", - "criticality": "release_blocking", - "criticality_tier": "core_fiscal_release", - "publisher": "SPF Finances / Statbel", - "series": "Personal income tax: assessed total", - "target_role": "calibration" + "activation_status": "requires_tax_return_unit_and_period_bridge", + "basis_period": "statbel_fiscal_distribution_2023", + "criticality": "diagnostic", + "criticality_tier": "validation_only", + "publisher": "Statbel", + "target_role": "validation" }, - "notes": "National assessed PIT is compared with the Axiom-computed federal and local liability on the 2022 income-year basis. Assessment-year facts must identify the represented income year in Chronicle; a period mismatch is a projection, never a silent relabeling." + "notes": "A personal-income-tax return is neither a Microcosm person nor household. This exact scalar identifies the distribution surface but remains validation-only and non-executable until a separately reviewed return unit exists; no distribution cells are reconstructed." }, { - "name": "onss_employee_contribution_total", + "name": "onem_complete_unemployment_monthly_average_2024", + "ledger_source_record_id": "onem_rva.unemployment.cy2024.complete_unemployment.country.country_total.receives_unemployment_benefit", "ledger_selector": { - "source_name": "onss_contributions", - "source_measure_id": "worker_article_17_uncapped_component_contribution", + "source_name": "onem_rva_unemployment", + "source_measure_id": "receives_unemployment_benefit", "period_type": "calendar_year", - "geography_level": "country" + "period_value": 2024, + "geography_level": "country", + "geography_id": "BE", + "geography_vintage": "current", + "entity_name": "person", + "record_set_id": "onem_rva.unemployment.cy2024.complete_unemployment.country", + "layout_groupby_value_id": "country_total", + "dimensions": [] }, "entity": "person", - "measure": "belgium_worker_article_17_uncapped_component_contribution", - "period": 2022, - "family": "social_security", - "assertion_policy": "allow_source_projection", + "measure": "belgium_complete_unemployment_monthly_average_recipient_indicator", + "period": 2024, + "family": "caseloads", + "assertion_policy": "observed_only", "period_match_policy": "exact", "metadata": { - "basis_period": "calendar_year_2022", + "activation_status": "requires_person_period_population_mapping_and_period_alignment", + "anti_proxy_rule": "do_not_derive_receipt_from_positive_amount_or_entitlement", + "axiom_behavior_ownership": "none", + "basis_period": "onem_monthly_average_2024", "criticality": "release_blocking", - "criticality_tier": "standard_admin_release", - "publisher": "ONSS/RSZ", - "series": "Employee social-security contributions", + "criticality_tier": "caseload_candidate", + "input_consumer": "PolicyEngine", + "input_owner": "Microcosm", + "mechanics_owner": "PolicyEngine", + "microcosm_population_input_status": "absent", + "policyengine_input_support_status": "absent", + "publisher": "ONEM/RVA", "target_role": "calibration" }, - "notes": "Worker-side contributions on the encoded slice. Employer-side and capped-regime totals can enter as distinct references when both the engine concepts and Ledger selectors exist." + "notes": "This is a monthly-average recipient statistic, not annual-any recipients or a take-up rate. It remains blocked because the Microcosm-populated unemployment receipt flag, the PolicyEngine input variable that consumes it, a typed person-period population, and an exact period mapping are absent." }, { - "name": "onem_unemployment_caseload", + "name": "sfpd_legal_pension_all_schemes_january_2025_blocked", + "ledger_source_record_id": "sfpd.legal_pension.cy2025.recipients.all_schemes.recipients", "ledger_selector": { - "source_name": "onem_rva_unemployment", - "source_measure_id": "receives_unemployment_benefit", + "source_name": "sfpd_pensions", + "source_measure_id": "recipients", "period_type": "calendar_year", - "geography_level": "country" + "period_value": 2025, + "geography_level": "country", + "geography_id": "BE", + "geography_vintage": "current", + "entity_name": "person", + "record_set_id": "sfpd.legal_pension.cy2025.recipients.by_scheme", + "layout_groupby_value_id": "all_schemes", + "dimensions": [] }, "entity": "person", - "measure": "receives_unemployment_benefit", - "period": 2022, + "measure": "belgium_legal_pension_recipient_indicator", + "period": 2025, "family": "caseloads", - "assertion_policy": "allow_source_projection", + "assertion_policy": "observed_only", "period_match_policy": "exact", "metadata": { - "basis_period": "calendar_year_2022", - "criticality": "release_blocking", - "criticality_tier": "caseload_release", - "publisher": "ONEM/RVA", - "series": "Unemployment benefit recipients", - "target_role": "calibration" + "activation_status": "requires_chronicle_month_period_correction_and_population_input", + "anti_proxy_rule": "do_not_derive_receipt_from_positive_amount_or_entitlement", + "axiom_behavior_ownership": "none", + "basis_period": "sfpd_pension_catalog_2025_mislabeled", + "criticality": "diagnostic", + "criticality_tier": "validation_only", + "input_consumer": "PolicyEngine", + "input_owner": "Microcosm", + "mechanics_owner": "PolicyEngine", + "microcosm_population_input_status": "absent", + "policyengine_input_support_status": "absent", + "publisher": "SFPD", + "target_role": "validation" }, - "notes": "The initial caseload family uses the calendar-year recipient count. Pension and regional child-benefit caseload references can join this family when their Chronicle packages and model mappings land." + "notes": "The merged Chronicle package types this row as calendar year 2025, while its source basis is January 2025. It is recorded only as a blocked dependency: correct the source period before targeting, and never sum the overlapping scheme rows in place of the published all-schemes total." }, { - "name": "nbb_household_disposable_income", + "name": "sfpd_grapa_regular_payment_beneficiaries_2025_01", + "ledger_source_record_id": "sfpd_grapa.month2025_01.regular_payment_beneficiaries.all.beneficiaries", "ledger_selector": { - "source_name": "nbb_national_accounts", - "source_measure_id": "household_disposable_income", - "period_type": "calendar_year", - "geography_level": "country" + "source_name": "sfpd_grapa", + "source_measure_id": "beneficiaries", + "period_type": "month", + "period_value": "2025-01", + "geography_level": "country", + "geography_id": "BE", + "geography_vintage": "current", + "entity_name": "person", + "record_set_id": "sfpd_grapa.month2025_01.regular_payment_beneficiaries.by_sex", + "layout_groupby_value_id": "all", + "dimensions": [] + }, + "entity": "person", + "measure": "belgium_grapa_regular_payment_recipient_indicator", + "period": "2025-01", + "family": "caseloads", + "assertion_policy": "observed_only", + "period_match_policy": "exact", + "metadata": { + "activation_status": "requires_microcosm_input_and_period_alignment", + "anti_proxy_rule": "do_not_derive_receipt_from_positive_amount_or_entitlement", + "axiom_behavior_ownership": "none", + "basis_period": "grapa_regular_payment_snapshot_2025_01", + "criticality": "diagnostic", + "criticality_tier": "validation_only", + "input_consumer": "PolicyEngine", + "input_owner": "Microcosm", + "mechanics_owner": "PolicyEngine", + "publisher_source_readiness": "ready", + "population_input_id": "grapa_regular_payment_recipient", + "population_input_readiness": "required_missing", + "population_mapping_readiness": "ready", + "population_period_readiness": "exact_alignment_missing", + "population_completeness_readiness": "complete_imputation_required_missing", + "publisher": "SFPD", + "scheme_population_mapping_id": "grapa_regular_payment_population_2025_01", + "target_role": "validation" + }, + "notes": "Exact January 2025 regular-payment beneficiary snapshot from Chronicle PR #212. It is not eligibility, annual-any receipt, a transaction count, or behavior. PolicyEngine must consume Microcosm's distinct measured or latent receipt flag through a PolicyEngine input variable and owns any residual non-legal behavior mechanics before Microcosm can estimate this row." + }, + { + "name": "opgroeien_groeipakket_basic_amount_children_2025_12", + "ledger_source_record_id": "opgroeien.groeipakket.month2025_12.basic_amount.children.basic_amount_children.children", + "ledger_selector": { + "source_name": "opgroeien_groeipakket_dashboard", + "source_measure_id": "children", + "period_type": "month", + "period_value": "2025-12", + "geography_level": "statistical_scope", + "geography_id": "BE-GROEIPAKKET-SCHEME", + "geography_vintage": "GROEIPAKKET_ADMINISTRATIVE_SCOPE_2025", + "entity_name": "person", + "record_set_id": "opgroeien.groeipakket.month2025_12.basic_amount.children", + "layout_groupby_value_id": "basic_amount_children", + "dimensions": [ + "groeipakket.component", + "publication_status" + ] + }, + "entity": "person", + "measure": "belgium_groeipakket_basic_amount_child_indicator", + "period": "2025-12", + "family": "caseloads", + "assertion_policy": "observed_only", + "period_match_policy": "exact", + "metadata": { + "activation_status": "requires_native_publisher_artifact_microcosm_input_and_period_alignment", + "anti_proxy_rule": "do_not_substitute_nuts_residence_or_positive_amount", + "axiom_behavior_ownership": "none", + "basis_period": "child_benefit_rights_month_2025_12", + "criticality": "diagnostic", + "criticality_tier": "validation_only", + "input_consumer": "PolicyEngine", + "input_owner": "Microcosm", + "mechanics_owner": "PolicyEngine", + "publisher_source_readiness": "native_publisher_artifact_required_missing", + "population_input_id": "groeipakket_basic_amount_child", + "population_input_readiness": "required_missing", + "population_mapping_readiness": "required_missing", + "population_period_readiness": "exact_alignment_missing", + "population_completeness_readiness": "complete_imputation_required_missing", + "publisher": "Opgroeien", + "scheme_population_mapping_id": "groeipakket_basic_amount_child_population_2025_12", + "target_role": "validation" + }, + "notes": "Chronicle #212 declares this provisional dashboard-derived record identity, but Chronicle #213 requires a pinned native Opgroeien export before it is activation-grade. The scheme-wide population includes cross-border categories and is not BE2 residence. The typed mapping retains the declared scope and blocks independently on publisher bytes, scheme-population construction, the exact child status, period alignment, and completeness." + }, + { + "name": "iriscare_entitled_children_2025_12", + "ledger_source_record_id": "iriscare.child_benefit.month2025_12.entitled_children.entitled_children.children", + "ledger_selector": { + "source_name": "iriscare_child_benefit_dashboard", + "source_measure_id": "children", + "period_type": "month", + "period_value": "2025-12", + "geography_level": "statistical_scope", + "geography_id": "BE-IRISCARE-CHILD-BENEFIT-SCHEME", + "geography_vintage": "IRISCARE_CHILD_BENEFIT_ADMIN_SCOPE_2025", + "entity_name": "person", + "record_set_id": "iriscare.child_benefit.month2025_12.entitled_children", + "layout_groupby_value_id": "entitled_children", + "dimensions": [ + "iriscare.statistic", + "publication_status" + ] }, - "entity": "household", - "measure": "household_disposable_income", - "period": 2022, - "family": "national_accounts", - "assertion_policy": "allow_source_projection", + "entity": "person", + "measure": "belgium_iriscare_child_benefit_entitled_child_indicator", + "period": "2025-12", + "family": "caseloads", + "assertion_policy": "observed_only", "period_match_policy": "exact", "metadata": { - "basis_period": "calendar_year_2022", + "activation_status": "requires_native_publisher_artifact_microcosm_input_and_period_alignment", + "anti_proxy_rule": "do_not_substitute_nuts_residence_or_positive_amount", + "axiom_behavior_ownership": "none", + "basis_period": "child_benefit_rights_month_2025_12", "criticality": "diagnostic", "criticality_tier": "validation_only", - "publisher": "NBB", - "series": "Household-sector disposable income (national accounts)", + "input_consumer": "PolicyEngine", + "input_owner": "Microcosm", + "mechanics_owner": "PolicyEngine", + "publisher_source_readiness": "native_publisher_artifact_required_missing", + "population_input_id": "iriscare_entitled_child", + "population_input_readiness": "required_missing", + "population_mapping_readiness": "required_missing", + "population_period_readiness": "exact_alignment_missing", + "population_completeness_readiness": "complete_imputation_required_missing", + "publisher": "Iriscare", + "scheme_population_mapping_id": "iriscare_entitled_child_population_2025_12", "target_role": "validation" }, - "notes": "Validation-tier level anchor only. The macro-realism gate may use it as a broad band, but neither the exact NBB level nor any survey- or EUROMOD-derived quantity enters the calibration objective." + "notes": "Chronicle #212 declares this provisional dashboard-derived record identity, but Chronicle #213 requires a pinned native official Iriscare artifact before it is activation-grade. The intended entitled-child person status remains separate from payment recipients and from BE1 residence; source, scheme-population, input, period, and completeness gates all remain fail-closed." + }, + { + "name": "iriscare_payment_recipients_2025_12", + "ledger_source_record_id": "iriscare.child_benefit.month2025_12.payment_recipients.payment_recipients.payment_recipients", + "ledger_selector": { + "source_name": "iriscare_child_benefit_dashboard", + "source_measure_id": "payment_recipients", + "period_type": "month", + "period_value": "2025-12", + "geography_level": "statistical_scope", + "geography_id": "BE-IRISCARE-CHILD-BENEFIT-SCHEME", + "geography_vintage": "IRISCARE_CHILD_BENEFIT_ADMIN_SCOPE_2025", + "entity_name": "person", + "record_set_id": "iriscare.child_benefit.month2025_12.payment_recipients", + "layout_groupby_value_id": "payment_recipients", + "dimensions": [ + "iriscare.statistic", + "publication_status" + ] + }, + "entity": "person", + "measure": "belgium_iriscare_child_benefit_payment_recipient_indicator", + "period": "2025-12", + "family": "caseloads", + "assertion_policy": "observed_only", + "period_match_policy": "exact", + "metadata": { + "activation_status": "requires_native_publisher_artifact_microcosm_input_and_period_alignment", + "anti_proxy_rule": "do_not_substitute_nuts_residence_or_positive_amount", + "axiom_behavior_ownership": "none", + "basis_period": "child_benefit_rights_month_2025_12", + "criticality": "diagnostic", + "criticality_tier": "validation_only", + "input_consumer": "PolicyEngine", + "input_owner": "Microcosm", + "mechanics_owner": "PolicyEngine", + "publisher_source_readiness": "native_publisher_artifact_required_missing", + "population_input_id": "iriscare_payment_recipient", + "population_input_readiness": "required_missing", + "population_mapping_readiness": "required_missing", + "population_period_readiness": "exact_alignment_missing", + "population_completeness_readiness": "complete_imputation_required_missing", + "publisher": "Iriscare", + "scheme_population_mapping_id": "iriscare_payment_recipient_population_2025_12", + "target_role": "validation" + }, + "notes": "Chronicle #212 declares this provisional dashboard-derived record identity, but Chronicle #213 requires a pinned native official Iriscare artifact before it is activation-grade. The intended allocataire unit is a payment-recipient person, not a household, and remains distinct from entitled children; source, scheme-population, input, period, and completeness gates all remain fail-closed." + }, + { + "name": "ostbelgien_paid_children_2025_12", + "ledger_source_record_id": "ostbelgien.child_benefit.month2025_12.paid_children.paid_children.children", + "ledger_selector": { + "source_name": "ostbelgien_child_benefit_statistics", + "source_measure_id": "children", + "period_type": "month", + "period_value": "2025-12", + "geography_level": "statistical_scope", + "geography_id": "BE-DG", + "geography_vintage": "child_benefit_scope_2025", + "entity_name": "person", + "record_set_id": "ostbelgien.child_benefit.month2025_12.paid_children", + "layout_groupby_value_id": "paid_children", + "dimensions": [ + "ostbelgien.statistic" + ] + }, + "entity": "person", + "measure": "belgium_ostbelgien_child_benefit_paid_child_indicator", + "period": "2025-12", + "family": "caseloads", + "assertion_policy": "observed_only", + "period_match_policy": "exact", + "metadata": { + "activation_status": "requires_native_publisher_artifact_microcosm_input_and_period_alignment", + "anti_proxy_rule": "do_not_substitute_nuts_residence_or_positive_amount", + "axiom_behavior_ownership": "none", + "basis_period": "child_benefit_rights_month_2025_12", + "criticality": "diagnostic", + "criticality_tier": "validation_only", + "input_consumer": "PolicyEngine", + "input_owner": "Microcosm", + "mechanics_owner": "PolicyEngine", + "publisher_source_readiness": "native_publisher_artifact_required_missing", + "population_input_id": "ostbelgien_paid_child", + "population_input_readiness": "required_missing", + "population_mapping_readiness": "required_missing", + "population_period_readiness": "exact_alignment_missing", + "population_completeness_readiness": "complete_imputation_required_missing", + "publisher": "Ostbelgien Statistik", + "scheme_population_mapping_id": "ostbelgien_paid_child_population_2025_12", + "target_role": "validation" + }, + "notes": "Chronicle #212 declares this manually transcribed record identity, but Chronicle #213 rejects manual transcription as a publisher artifact and requires a pinned native official response before activation. BE-DG remains the declared scheme scope, not resident geography. Paid children stay distinct from payment recipients, and every source, population, input, period, and completeness gate remains fail-closed." + }, + { + "name": "ostbelgien_payment_recipients_2025_12", + "ledger_source_record_id": "ostbelgien.child_benefit.month2025_12.payment_recipients.payment_recipients.payment_recipients", + "ledger_selector": { + "source_name": "ostbelgien_child_benefit_statistics", + "source_measure_id": "payment_recipients", + "period_type": "month", + "period_value": "2025-12", + "geography_level": "statistical_scope", + "geography_id": "BE-DG", + "geography_vintage": "child_benefit_scope_2025", + "entity_name": "person", + "record_set_id": "ostbelgien.child_benefit.month2025_12.payment_recipients", + "layout_groupby_value_id": "payment_recipients", + "dimensions": [ + "ostbelgien.statistic" + ] + }, + "entity": "person", + "measure": "belgium_ostbelgien_child_benefit_payment_recipient_indicator", + "period": "2025-12", + "family": "caseloads", + "assertion_policy": "observed_only", + "period_match_policy": "exact", + "metadata": { + "activation_status": "requires_native_publisher_artifact_microcosm_input_and_period_alignment", + "anti_proxy_rule": "do_not_substitute_nuts_residence_or_positive_amount", + "axiom_behavior_ownership": "none", + "basis_period": "child_benefit_rights_month_2025_12", + "criticality": "diagnostic", + "criticality_tier": "validation_only", + "input_consumer": "PolicyEngine", + "input_owner": "Microcosm", + "mechanics_owner": "PolicyEngine", + "publisher_source_readiness": "native_publisher_artifact_required_missing", + "population_input_id": "ostbelgien_payment_recipient", + "population_input_readiness": "required_missing", + "population_mapping_readiness": "required_missing", + "population_period_readiness": "exact_alignment_missing", + "population_completeness_readiness": "complete_imputation_required_missing", + "publisher": "Ostbelgien Statistik", + "scheme_population_mapping_id": "ostbelgien_payment_recipient_population_2025_12", + "target_role": "validation" + }, + "notes": "Chronicle #212 declares this manually transcribed record identity, but Chronicle #213 rejects manual transcription as a publisher artifact and requires a pinned native official response before activation. The intended unit is a distinct payment-recipient person in the declared scheme scope, not a child, family, household, or take-up count; every source, population, input, period, and completeness gate remains fail-closed." + }, + { + "name": "walloon_single_parent_with_supplement_children_2023_12", + "ledger_source_record_id": "parlement_wallonie.child_benefit.month2023_12.children.single_parent_with_social_supplement.children", + "ledger_selector": { + "source_name": "parlement_wallonie_child_benefit_response", + "source_measure_id": "children", + "period_type": "month", + "period_value": "2023-12", + "geography_level": "statistical_scope", + "geography_id": "BE-WALLOON-FRENCH", + "geography_vintage": "child_benefit_scope_2023", + "entity_name": "person", + "record_set_id": "parlement_wallonie.child_benefit.month2023_12.children_by_household_group", + "layout_groupby_value_id": "single_parent_with_social_supplement", + "dimensions": [ + "parlement_wallonie.household_type", + "parlement_wallonie.social_supplement_status" + ] + }, + "entity": "person", + "measure": "belgium_walloon_single_parent_with_supplement_recipient_child_indicator", + "period": "2023-12", + "family": "caseloads", + "assertion_policy": "observed_only", + "period_match_policy": "exact", + "metadata": { + "activation_status": "requires_microcosm_input_and_period_alignment", + "anti_proxy_rule": "preserve_publisher_partition_and_do_not_construct_total", + "axiom_behavior_ownership": "none", + "basis_period": "walloon_child_benefit_snapshot_2023_12", + "criticality": "diagnostic", + "criticality_tier": "validation_only", + "input_consumer": "PolicyEngine", + "input_owner": "Microcosm", + "mechanics_owner": "PolicyEngine", + "publisher_source_readiness": "ready", + "population_input_id": "walloon_single_parent_with_supplement_child", + "population_input_readiness": "required_missing", + "population_mapping_readiness": "required_missing", + "population_period_readiness": "exact_alignment_missing", + "population_completeness_readiness": "complete_imputation_required_missing", + "publisher": "Parlement de Wallonie", + "scheme_population_mapping_id": "walloon_single_parent_with_supplement_children_2023_12", + "target_role": "validation" + }, + "notes": "One of four exact publisher child partitions. It is not combined with the other partitions and is not mapped to the publisher household count." + }, + { + "name": "walloon_single_parent_without_supplement_children_2023_12", + "ledger_source_record_id": "parlement_wallonie.child_benefit.month2023_12.children.single_parent_without_social_supplement.children", + "ledger_selector": { + "source_name": "parlement_wallonie_child_benefit_response", + "source_measure_id": "children", + "period_type": "month", + "period_value": "2023-12", + "geography_level": "statistical_scope", + "geography_id": "BE-WALLOON-FRENCH", + "geography_vintage": "child_benefit_scope_2023", + "entity_name": "person", + "record_set_id": "parlement_wallonie.child_benefit.month2023_12.children_by_household_group", + "layout_groupby_value_id": "single_parent_without_social_supplement", + "dimensions": [ + "parlement_wallonie.household_type", + "parlement_wallonie.social_supplement_status" + ] + }, + "entity": "person", + "measure": "belgium_walloon_single_parent_without_supplement_recipient_child_indicator", + "period": "2023-12", + "family": "caseloads", + "assertion_policy": "observed_only", + "period_match_policy": "exact", + "metadata": { + "activation_status": "requires_microcosm_input_and_period_alignment", + "anti_proxy_rule": "preserve_publisher_partition_and_do_not_construct_total", + "axiom_behavior_ownership": "none", + "basis_period": "walloon_child_benefit_snapshot_2023_12", + "criticality": "diagnostic", + "criticality_tier": "validation_only", + "input_consumer": "PolicyEngine", + "input_owner": "Microcosm", + "mechanics_owner": "PolicyEngine", + "publisher_source_readiness": "ready", + "population_input_id": "walloon_single_parent_without_supplement_child", + "population_input_readiness": "required_missing", + "population_mapping_readiness": "required_missing", + "population_period_readiness": "exact_alignment_missing", + "population_completeness_readiness": "complete_imputation_required_missing", + "publisher": "Parlement de Wallonie", + "scheme_population_mapping_id": "walloon_single_parent_without_supplement_children_2023_12", + "target_role": "validation" + }, + "notes": "One of four exact publisher child partitions. It is not combined with the other partitions and is not mapped to the publisher household count." + }, + { + "name": "walloon_other_household_with_supplement_children_2023_12", + "ledger_source_record_id": "parlement_wallonie.child_benefit.month2023_12.children.other_household_with_social_supplement.children", + "ledger_selector": { + "source_name": "parlement_wallonie_child_benefit_response", + "source_measure_id": "children", + "period_type": "month", + "period_value": "2023-12", + "geography_level": "statistical_scope", + "geography_id": "BE-WALLOON-FRENCH", + "geography_vintage": "child_benefit_scope_2023", + "entity_name": "person", + "record_set_id": "parlement_wallonie.child_benefit.month2023_12.children_by_household_group", + "layout_groupby_value_id": "other_household_with_social_supplement", + "dimensions": [ + "parlement_wallonie.household_type", + "parlement_wallonie.social_supplement_status" + ] + }, + "entity": "person", + "measure": "belgium_walloon_other_household_with_supplement_recipient_child_indicator", + "period": "2023-12", + "family": "caseloads", + "assertion_policy": "observed_only", + "period_match_policy": "exact", + "metadata": { + "activation_status": "requires_microcosm_input_and_period_alignment", + "anti_proxy_rule": "preserve_publisher_partition_and_do_not_construct_total", + "axiom_behavior_ownership": "none", + "basis_period": "walloon_child_benefit_snapshot_2023_12", + "criticality": "diagnostic", + "criticality_tier": "validation_only", + "input_consumer": "PolicyEngine", + "input_owner": "Microcosm", + "mechanics_owner": "PolicyEngine", + "publisher_source_readiness": "ready", + "population_input_id": "walloon_other_household_with_supplement_child", + "population_input_readiness": "required_missing", + "population_mapping_readiness": "required_missing", + "population_period_readiness": "exact_alignment_missing", + "population_completeness_readiness": "complete_imputation_required_missing", + "publisher": "Parlement de Wallonie", + "scheme_population_mapping_id": "walloon_other_household_with_supplement_children_2023_12", + "target_role": "validation" + }, + "notes": "One of four exact publisher child partitions. It is not combined with the other partitions and is not mapped to the publisher household count." + }, + { + "name": "walloon_other_household_without_supplement_children_2023_12", + "ledger_source_record_id": "parlement_wallonie.child_benefit.month2023_12.children.other_household_without_social_supplement.children", + "ledger_selector": { + "source_name": "parlement_wallonie_child_benefit_response", + "source_measure_id": "children", + "period_type": "month", + "period_value": "2023-12", + "geography_level": "statistical_scope", + "geography_id": "BE-WALLOON-FRENCH", + "geography_vintage": "child_benefit_scope_2023", + "entity_name": "person", + "record_set_id": "parlement_wallonie.child_benefit.month2023_12.children_by_household_group", + "layout_groupby_value_id": "other_household_without_social_supplement", + "dimensions": [ + "parlement_wallonie.household_type", + "parlement_wallonie.social_supplement_status" + ] + }, + "entity": "person", + "measure": "belgium_walloon_other_household_without_supplement_recipient_child_indicator", + "period": "2023-12", + "family": "caseloads", + "assertion_policy": "observed_only", + "period_match_policy": "exact", + "metadata": { + "activation_status": "requires_microcosm_input_and_period_alignment", + "anti_proxy_rule": "preserve_publisher_partition_and_do_not_construct_total", + "axiom_behavior_ownership": "none", + "basis_period": "walloon_child_benefit_snapshot_2023_12", + "criticality": "diagnostic", + "criticality_tier": "validation_only", + "input_consumer": "PolicyEngine", + "input_owner": "Microcosm", + "mechanics_owner": "PolicyEngine", + "publisher_source_readiness": "ready", + "population_input_id": "walloon_other_household_without_supplement_child", + "population_input_readiness": "required_missing", + "population_mapping_readiness": "required_missing", + "population_period_readiness": "exact_alignment_missing", + "population_completeness_readiness": "complete_imputation_required_missing", + "publisher": "Parlement de Wallonie", + "scheme_population_mapping_id": "walloon_other_household_without_supplement_children_2023_12", + "target_role": "validation" + }, + "notes": "One of four exact publisher child partitions. It is not combined with the other partitions and is not mapped to the publisher household count." } ], "target_profile": { @@ -162,67 +647,86 @@ "required_families": [ "demography", "fiscal_income", - "income_tax", - "social_security", "caseloads" ], "criticality_tiers": { - "demography_release": { + "demography_candidate": { "criticality": "release_blocking", "relative_tolerance": 0.02, - "description": "Population-structure cells: plus or minus 2 percent." - }, - "core_fiscal_release": { - "criticality": "release_blocking", - "relative_tolerance": 0.05, - "description": "Core fiscal totals, including assessed PIT: plus or minus 5 percent." + "description": "Future exact-period population cells; currently blocked on cell, period, geography, and support contracts." }, - "standard_admin_release": { - "criticality": "release_blocking", + "commune_diagnostic": { + "criticality": "diagnostic", "relative_tolerance": 0.1, - "description": "Administrative contribution totals: plus or minus 10 percent." + "description": "Future commune fiscal-income cells, reported diagnostically after exact period and unit support." }, - "caseload_release": { + "caseload_candidate": { "criticality": "release_blocking", "relative_tolerance": 0.15, - "description": "Administrative caseload totals: plus or minus 15 percent." - }, - "commune_diagnostic": { - "criticality": "diagnostic", - "relative_tolerance": 0.1, - "description": "Commune fiscal-income cells: plus or minus 10 percent, reported but not release-blocking." + "description": "Future administrative caseload level after exact population-input, statistic, and period support." }, "validation_only": { "criticality": "diagnostic", "relative_tolerance": null, - "description": "Validation oracle only; no calibration tolerance." + "description": "Held out from calibration; no solver tolerance." } }, "basis_periods": { - "population_reference_2023": { - "period": 2023, - "basis": "reference_date", + "statbel_population_2025": { + "period": 2025, + "basis": "official_population_reference_year", "fact_period_type": "calendar_year", - "mismatch_policy": "requires_source_projection", - "description": "Population stock in the BE-SILC 2023 survey year." + "mismatch_policy": "exact_observation_only", + "description": "Statbel 2025 population cells at their published NUTS-2024 geography." }, - "assessment_income_year_2022": { - "period": 2022, - "basis": "income_year", + "statbel_fiscal_income_2023": { + "period": 2023, + "basis": "income_and_tax_year", "fact_period_type": "tax_year", - "survey_year": 2023, - "income_reference_offset_years": -1, - "mismatch_policy": "requires_source_projection", - "description": "BE-SILC survey year 2023 carries 2022 incomes; assessed fiscal facts bind by represented income year." - }, - "calendar_year_2022": { - "period": 2022, - "basis": "calendar_year", + "mismatch_policy": "exact_observation_only", + "description": "Statbel fiscal income for income/tax year 2023 under nis_2025." + }, + "statbel_fiscal_distribution_2023": { + "period": 2023, + "basis": "income_and_tax_year", + "fact_period_type": "tax_year", + "mismatch_policy": "exact_observation_only", + "description": "Statbel tax-return distribution fact for income year 2023." + }, + "onem_monthly_average_2024": { + "period": 2024, + "basis": "monthly_average_within_calendar_year", "fact_period_type": "calendar_year", - "survey_year": 2023, - "income_reference_offset_years": -1, - "mismatch_policy": "requires_source_projection", - "description": "Calendar-year administrative or national-accounts fact aligned to the 2022 income reference year." + "mismatch_policy": "exact_observation_only", + "description": "ONEM complete-unemployment monthly-average recipient statistic for 2024." + }, + "sfpd_pension_catalog_2025_mislabeled": { + "period": 2025, + "basis": "blocked_source_period_mismatch", + "fact_period_type": "calendar_year", + "mismatch_policy": "exact_observation_only", + "description": "Chronicle currently exposes calendar-year 2025, while the publisher basis is January 2025; the reference cannot activate." + }, + "grapa_regular_payment_snapshot_2025_01": { + "period": "2025-01", + "basis": "regular_payment_snapshot", + "fact_period_type": "month", + "mismatch_policy": "exact_observation_only", + "description": "SFPD January 2025 GRAPA regular-payment population." + }, + "child_benefit_rights_month_2025_12": { + "period": "2025-12", + "basis": "publisher_rights_or_reference_month", + "fact_period_type": "month", + "mismatch_policy": "exact_observation_only", + "description": "Opgroeien, Iriscare, and Ostbelgien December 2025 publisher populations, retained separately." + }, + "walloon_child_benefit_snapshot_2023_12": { + "period": "2023-12", + "basis": "publisher_reference_month", + "fact_period_type": "month", + "mismatch_policy": "exact_observation_only", + "description": "Walloon Parliament December 2023 publisher child partitions, retained separately." } }, "hierarchy_reconciliations": [] diff --git a/packages/microcosm-build/src/microcosm/build/country_spec.py b/packages/microcosm-build/src/microcosm/build/country_spec.py index 1eb45e39c..51168612c 100644 --- a/packages/microcosm-build/src/microcosm/build/country_spec.py +++ b/packages/microcosm-build/src/microcosm/build/country_spec.py @@ -22,6 +22,8 @@ - ``geography_spine.json`` — :class:`GeographySpineManifest` (this module) - ``target_references.json`` — Ledger references, validated as :class:`~microcosm.build.ledger_targets.LedgerTargetReference` rows +- ``population_inputs.json`` — :class:`~microcosm.build.population_inputs.PopulationInputProfile` +- ``monetary_target_profile.json`` — :class:`~microcosm.build.monetary_profile.MonetaryTargetProfile` - ``gates.json`` — :class:`GatesManifest` (this module) - ``release_contract.json`` — :class:`ReleaseContractManifest` (this module) @@ -47,7 +49,9 @@ period_type_hint, period_values_semantically_equal, ) +from microcosm.build.monetary_profile import MonetaryTargetProfile from microcosm.build.plan import DonorSpec, Stage, StagePlan +from microcosm.build.population_inputs import PopulationInputProfile from microcosm.build.source_manifest import ( SourceManifest, SupportSpineManifest, @@ -920,8 +924,9 @@ def _validate_target_profile( declaration fields needed to make a future target surface auditable without carrying target values: named criticality/tolerance tiers, named basis periods, an exact-period posture, and explicit geography vintages on - every subnational selector. This validates schema policy only; it does not - apply tier tolerances or ``target_role`` to runtime calibration or gates. + every subnational selector. Tier tolerances remain declaration metadata for + a separate gate integration. The shared target compiler does enforce + ``target_role='validation'`` as an exclusion from calibration. """ profile = raw.get("target_profile", {}) @@ -1033,7 +1038,7 @@ def _validate_target_profile( "target_references.json: target_profile.basis_periods must be a " "non-empty object." ) - normalized_periods: dict[str, tuple[object, str]] = {} + normalized_periods: dict[str, tuple[object, str, str]] = {} for raw_basis_id, raw_basis in basis_periods.items(): basis_id = _require_non_empty_string( raw_basis_id, @@ -1086,10 +1091,15 @@ def _validate_target_profile( field_name="fact_period_type", context=f"basis period {basis_id!r}", ) - if raw_basis.get("mismatch_policy") != "requires_source_projection": + mismatch_policy = raw_basis.get("mismatch_policy") + if mismatch_policy not in { + "requires_source_projection", + "exact_observation_only", + }: raise ValueError( f"target_references.json: basis period {basis_id!r} must set " - "mismatch_policy='requires_source_projection'." + "mismatch_policy to 'requires_source_projection' or " + "'exact_observation_only'." ) survey_year = raw_basis.get("survey_year") income_offset = raw_basis.get("income_reference_offset_years") @@ -1118,7 +1128,11 @@ def _validate_target_profile( f"{period!r} does not equal survey_year {survey_year!r} plus " f"income_reference_offset_years {income_offset!r}." ) - normalized_periods[basis_id] = (period, fact_period_type) + normalized_periods[basis_id] = ( + period, + fact_period_type, + str(mismatch_policy), + ) geography_vintage_aliases = _typed_geography_vintage_aliases( resolved_spec, @@ -1171,7 +1185,7 @@ def _validate_target_profile( f"target_references.json: {context} names unknown basis_period " f"{basis_id!r}." ) - basis_period, fact_period_type = normalized_periods[basis_id] + basis_period, fact_period_type, mismatch_policy = normalized_periods[basis_id] if not period_values_semantically_equal( reference.period, basis_period, declared_type=fact_period_type ): @@ -1200,21 +1214,45 @@ def _validate_target_profile( f"{selector_period_type!r} does not match basis period " f"{basis_id!r} fact_period_type {fact_period_type!r}." ) + selector_period = reference.ledger_selector.get("period_value") + if not period_values_semantically_equal( + selector_period, + reference.period, + declared_type=fact_period_type, + ): + raise ValueError( + f"target_references.json: {context} selector period " + f"{selector_period!r} does not match its declared period " + f"{reference.period!r}." + ) if reference.period_match_policy != "exact": raise ValueError( f"target_references.json: {context} must set " "period_match_policy='exact'; stale observations require an " "explicit Chronicle source_projection." ) - if reference.assertion_policy != "allow_source_projection": + expected_assertion_policy = ( + "allow_source_projection" + if mismatch_policy == "requires_source_projection" + else "observed_only" + ) + if reference.assertion_policy != expected_assertion_policy: raise ValueError( f"target_references.json: {context} must set " - "assertion_policy='allow_source_projection' so an explicitly " - "projected fact can satisfy the declared basis period." + f"assertion_policy={expected_assertion_policy!r} to agree with " + f"basis period {basis_id!r} mismatch_policy={mismatch_policy!r}." ) geography_level = str(reference.ledger_selector.get("geography_level", "")) - if geography_level and geography_level not in {"country", "national"}: + if geography_level == "statistical_scope": + scope_id = str(reference.ledger_selector.get("geography_id", "")) + scope_vintage = str(reference.ledger_selector.get("geography_vintage", "")) + if not scope_id or not scope_vintage: + raise ValueError( + f"target_references.json: statistical-scope {context} must " + "pin geography_id and geography_vintage." + ) + elif geography_level and geography_level not in {"country", "national"}: selector_vintage = str( reference.ledger_selector.get("geography_vintage", "") ) @@ -1232,12 +1270,22 @@ def _validate_target_profile( "geography layer-vintage registry does not declare it." ) if selector_vintage not in accepted_typed_aliases: - raise ValueError( - f"target_references.json: {context} geography vintage " - f"{selector_vintage!r} is not an exact typed authority alias " - f"for layer {geography_level!r}; expected one of " - f"{sorted(accepted_typed_aliases)!r}." + model_vintage = reference.metadata.get("model_geography_vintage", "") + explicit_missing_bridge = ( + model_vintage in accepted_typed_aliases + and reference.metadata.get("geography_bridge_status") + == "required_missing" + and bool(reference.metadata.get("activation_status")) + and reference.metadata["activation_status"] != "active" ) + if not explicit_missing_bridge: + raise ValueError( + f"target_references.json: {context} geography vintage " + f"{selector_vintage!r} is not an exact typed authority " + f"alias for layer {geography_level!r}; expected one of " + f"{sorted(accepted_typed_aliases)!r}, or an explicit " + "non-active required_missing bridge to one of them." + ) if geography_spine is not None and ( geography_level == geography_spine.geography_spine.geography_level ): @@ -1331,6 +1379,181 @@ def _validate_local_target_references( return references +def _validate_population_target_links( + references: tuple[LedgerTargetReference, ...], + profile: PopulationInputProfile | None, + *, + country: str, +) -> None: + """Bind target references to exact, typed population-input mappings. + + The population profile is value-free, but its identities are executable + authority: a target cannot silently change source record, statistical + scope, entity, period, or input column while retaining a mapping id. The + Belgium package requires a complete bijection between its declared scheme + mappings and mapped target rows; other countries may adopt the profile + before declaring Ledger targets. + """ + + mapped_reference_rows = tuple( + reference + for reference in references + if reference.metadata.get("scheme_population_mapping_id") + ) + mapped_references = { + reference.name: reference for reference in mapped_reference_rows + } + if profile is None: + if mapped_references: + raise ValueError( + "target_references.json: scheme-population mapping ids require " + "a declared population_inputs.json profile." + ) + return + + reference_by_name = {reference.name: reference for reference in references} + mapping_ids = {mapping.mapping_id for mapping in profile.mappings} + linked_names_by_mapping_id: dict[str, list[str]] = {} + for reference in mapped_reference_rows: + mapping_id = reference.metadata["scheme_population_mapping_id"] + if mapping_id not in mapping_ids: + raise ValueError( + f"target_references.json: target {reference.name!r} names " + f"unknown scheme-population mapping {mapping_id!r}." + ) + linked_names_by_mapping_id.setdefault(mapping_id, []).append(reference.name) + + duplicate_links = { + mapping_id: names + for mapping_id, names in linked_names_by_mapping_id.items() + if len(names) != 1 + } + if duplicate_links: + raise ValueError( + "target_references.json: scheme-population mappings require one " + f"exact target link each; duplicate links {duplicate_links!r}." + ) + + for mapping in profile.mappings: + reference = reference_by_name.get(mapping.target_reference) + if reference is None: + if country == "be": + raise ValueError( + "population_inputs.json: Belgium scheme-population mapping " + f"{mapping.mapping_id!r} names unknown target reference " + f"{mapping.target_reference!r}." + ) + continue + context = f"scheme-population mapping {mapping.mapping_id!r}" + linked_names = linked_names_by_mapping_id.get(mapping.mapping_id, []) + if linked_names != [mapping.target_reference]: + raise ValueError( + f"target_references.json: {context} must link only its declared " + f"target {mapping.target_reference!r}, got {linked_names!r}." + ) + input_contract = profile.input(mapping.input_id) + expected_metadata = { + "scheme_population_mapping_id": mapping.mapping_id, + "population_input_id": mapping.input_id, + "input_owner": "Microcosm", + "input_consumer": "PolicyEngine", + "mechanics_owner": "PolicyEngine", + "axiom_behavior_ownership": "none", + "publisher_source_readiness": mapping.publisher_source_readiness, + "population_input_readiness": mapping.input_readiness, + "population_mapping_readiness": mapping.mapping_readiness, + "population_period_readiness": mapping.period_readiness, + "population_completeness_readiness": mapping.completeness_readiness, + } + metadata_mismatches = { + key: (expected, reference.metadata.get(key)) + for key, expected in expected_metadata.items() + if reference.metadata.get(key) != expected + } + if metadata_mismatches: + raise ValueError( + f"target_references.json: {context} metadata mismatch " + f"{metadata_mismatches!r}." + ) + if reference.ledger_source_record_id != mapping.chronicle_source_record_id: + raise ValueError( + f"target_references.json: {context} source record id does not " + "match its target reference." + ) + if ( + reference.entity != mapping.microcosm_entity + or reference.measure != input_contract.column + or reference.filter is not None + ): + raise ValueError( + f"target_references.json: {context} must use its exact " + "Microcosm entity and boolean input column without a proxy filter." + ) + + selector = reference.ledger_selector + expected_selector = { + "entity_name": mapping.chronicle_entity, + "geography_level": mapping.chronicle_geography_level, + "geography_id": mapping.chronicle_geography_id, + "geography_vintage": mapping.chronicle_geography_vintage, + "period_type": mapping.chronicle_period_type, + } + selector_mismatches = { + key: (expected, selector.get(key)) + for key, expected in expected_selector.items() + if selector.get(key) != expected + } + if selector_mismatches: + raise ValueError( + f"target_references.json: {context} Chronicle selector mismatch " + f"{selector_mismatches!r}." + ) + if not period_values_semantically_equal( + selector.get("period_value"), + mapping.chronicle_period, + declared_type=mapping.chronicle_period_type, + ): + raise ValueError( + f"target_references.json: {context} selector period does not " + "match its Chronicle mapping period." + ) + if not period_values_semantically_equal( + reference.period, + mapping.chronicle_period, + declared_type=mapping.chronicle_period_type, + ): + raise ValueError( + f"target_references.json: {context} Chronicle period does not " + "match its target reference." + ) + if ( + reference.value_operation != "identity" + or reference.assertion_policy != "observed_only" + or reference.period_match_policy != "exact" + ): + raise ValueError( + f"target_references.json: {context} requires identity, " + "observed_only, exact-period resolution." + ) + if mapping.blockers and reference.metadata.get("activation_status") == "active": + raise ValueError( + f"target_references.json: {context} cannot be active while its " + f"population mapping is blocked by {mapping.blockers!r}." + ) + + if country == "be": + linked_mapping_ids = { + reference.metadata["scheme_population_mapping_id"] + for reference in mapped_references.values() + } + unlinked = sorted(mapping_ids - linked_mapping_ids) + if unlinked: + raise ValueError( + "population_inputs.json: Belgium mapping(s) have no exact target " + f"reference link: {unlinked!r}." + ) + + def _local_area_rosters( crosswalk: Mapping[str, Any], *, @@ -1472,8 +1695,15 @@ class ResolvedCountrySpec: geography_spine: The geography-spine manifest, when declared. target_references: Ledger target references, when declared. target_profile: The validated value-free target declaration carried by - ``target_references.json``. Tier tolerances and target roles remain - metadata until a separate runtime integration consumes them. + ``target_references.json``. Tier tolerances remain metadata until a + separate gate integration consumes them; validation roles are + already excluded from calibration compilation. + population_input_profile: The typed, value-free inventory of + Microcosm-owned population inputs and scheme mappings, when + declared. + monetary_target_profile: The typed, value-free inventory of monetary + target references, exact bases, and activation prerequisites, when + declared. gates: The gate selection, when declared. release_contract: The release contract, when declared. take_up_contract: The constants-era take-up compatibility view. For a @@ -1495,6 +1725,8 @@ class ResolvedCountrySpec: geography_spine: GeographySpineManifest | None target_references: tuple[LedgerTargetReference, ...] target_profile: Mapping[str, Any] + population_input_profile: PopulationInputProfile | None + monetary_target_profile: MonetaryTargetProfile | None local_target_references: tuple[LedgerTargetReference, ...] gates: GatesManifest | None release_contract: ReleaseContractManifest | None @@ -2125,6 +2357,27 @@ def load_country_spec(country: str | Path) -> ResolvedCountrySpec: if "target_references.json" in payloads else MappingProxyType({}) ) + population_input_profile = ( + PopulationInputProfile.from_mapping( + payloads["population_inputs.json"], + country=declared_country, + ) + if "population_inputs.json" in payloads + else None + ) + monetary_target_profile = ( + MonetaryTargetProfile.from_mapping( + payloads["monetary_target_profile.json"], + country=declared_country, + ) + if "monetary_target_profile.json" in payloads + else None + ) + _validate_population_target_links( + target_references, + population_input_profile, + country=declared_country, + ) local_target_references = ( _validate_local_target_references( payloads["local_target_references.json"], @@ -2161,6 +2414,8 @@ def load_country_spec(country: str | Path) -> ResolvedCountrySpec: geography_spine=geography_spine, target_references=target_references, target_profile=target_profile, + population_input_profile=population_input_profile, + monetary_target_profile=monetary_target_profile, local_target_references=local_target_references, gates=gates, release_contract=release_contract, diff --git a/packages/microcosm-build/src/microcosm/build/ledger_targets.py b/packages/microcosm-build/src/microcosm/build/ledger_targets.py index 83bca9306..169d9a88f 100644 --- a/packages/microcosm-build/src/microcosm/build/ledger_targets.py +++ b/packages/microcosm-build/src/microcosm/build/ledger_targets.py @@ -30,6 +30,26 @@ ) EXACT_PERIOD_VALUE_OPERATIONS = frozenset(("identity", "sum", "count_x_mean")) DEFAULT_HIERARCHY_MATCH_SPEC_FIELDS = ("entity", "period", "family", "filter") +EXECUTION_GUARDED_STATUS_KEYS = frozenset( + ( + "axiom_legal_output_status", + "behavioral_takeup_flag_status", + "geography_alignment_status", + "geography_bridge_status", + "microcosm_population_input_status", + "population_completeness_readiness", + "population_input_readiness", + "population_mapping_readiness", + "population_period_readiness", + "policyengine_behavior_input_status", + "policyengine_input_support_status", + "publisher_source_readiness", + "support_status", + ) +) +EXECUTION_RESOLVED_STATUS_VALUES = frozenset( + ("aligned", "available", "not_applicable", "ready", "resolved") +) @dataclass(frozen=True) @@ -423,6 +443,12 @@ def _require_executable_reference(reference: LedgerTargetReference) -> None: singleton input or become ambiguous as soon as the next cell arrived. """ + if reference.metadata.get("target_role") == "validation": + raise ValueError( + f"Ledger target reference {reference.name!r} is validation-only and " + "cannot enter a calibration TargetRegistry." + ) + activation_status = reference.metadata.get("activation_status", "") if activation_status and activation_status != "active": raise ValueError( @@ -432,6 +458,24 @@ def _require_executable_reference(reference: LedgerTargetReference) -> None: "fact reference) before compilation." ) + blockers = { + key: value + for key, value in reference.metadata.items() + if key in EXECUTION_GUARDED_STATUS_KEYS + and ( + not isinstance(value, str) or value not in EXECUTION_RESOLVED_STATUS_VALUES + ) + } + if blockers: + rendered = ", ".join( + f"{key}={value!r}" for key, value in sorted(blockers.items()) + ) + raise ValueError( + f"Ledger target reference {reference.name!r} has unresolved " + f"execution blockers: {rendered}. Activation is fail-closed until " + "every declared support and alignment status is resolved." + ) + def apply_ledger_target_profile( registry: TargetRegistry, diff --git a/packages/microcosm-build/src/microcosm/build/monetary_profile.py b/packages/microcosm-build/src/microcosm/build/monetary_profile.py index c9607f88b..83aa2d1be 100644 --- a/packages/microcosm-build/src/microcosm/build/monetary_profile.py +++ b/packages/microcosm-build/src/microcosm/build/monetary_profile.py @@ -11,7 +11,10 @@ from typing import Any from urllib.parse import urlparse -from microcosm.build.ledger_targets import LedgerTargetReference +from microcosm.build.ledger_targets import ( + LedgerTargetReference, + period_values_semantically_equal, +) from microcosm.build.monetary_targets import MonetaryBasis READINESS_STATES = frozenset( @@ -22,6 +25,9 @@ "historical_validation_only", } ) +_FORBIDDEN_CARRIED_VALUE_KEYS = frozenset( + {"value", "values", "observed", "observed_value"} +) @dataclass(frozen=True) @@ -92,6 +98,16 @@ def _closed_keys(raw: Any, keys: set[str], context: str) -> None: raise ValueError(f"{context}: expected exactly {sorted(keys)}") +def _carried_value_keys(value: Any) -> set[str]: + if isinstance(value, Mapping): + return (set(value) & _FORBIDDEN_CARRIED_VALUE_KEYS) | { + key for child in value.values() for key in _carried_value_keys(child) + } + if isinstance(value, (list, tuple)): + return {key for child in value for key in _carried_value_keys(child)} + return set() + + def _contract(raw: Any, context: str) -> MonetaryTargetContract: _closed_keys( raw, {"reference", "basis", "readiness", "source_url", "notes"}, context @@ -108,6 +124,12 @@ def _contract(raw: Any, context: str) -> MonetaryTargetContract: raw["basis"], Mapping ): raise ValueError(f"{context}: reference and basis must be mappings") + carried_values = _carried_value_keys(raw["reference"]) + if carried_values: + raise ValueError( + f"{context}: monetary references must be value-free; forbidden " + f"carried value keys {sorted(carried_values)!r}" + ) try: reference = LedgerTargetReference(**raw["reference"]) basis = MonetaryBasis(**raw["basis"]) @@ -133,6 +155,19 @@ def _contract(raw: Any, context: str) -> MonetaryTargetContract: raise ValueError(f"{context}: no implicit monetary period uprating") if isinstance(reference.period, bool) or str(reference.period) != basis.period[:4]: raise ValueError(f"{context}: model year must match the accounting basis") + selector = reference.ledger_selector + if selector: + selector_type = selector.get("period_type") + if not isinstance(selector_type, str) or not selector_type: + raise ValueError(f"{context}: selector must declare period_type") + if not period_values_semantically_equal( + selector.get("period_value"), + reference.period, + declared_type=selector_type, + ): + raise ValueError( + f"{context}: selector period must match the model year and basis" + ) role = "validation" if readiness == "historical_validation_only" else "calibration" metadata = reference.metadata if ( diff --git a/packages/microcosm-build/src/microcosm/build/population_inputs.py b/packages/microcosm-build/src/microcosm/build/population_inputs.py new file mode 100644 index 000000000..5e2767ef8 --- /dev/null +++ b/packages/microcosm-build/src/microcosm/build/population_inputs.py @@ -0,0 +1,631 @@ +"""Typed, fail-closed population inputs for behavioral adapters. + +Microcosm owns the measured or latent population columns declared here. +PolicyEngine may consume those columns and owns any behavioral mechanics that +use them. This module deliberately has no formula, probability, eligibility, +take-up-assignment, or Axiom-concept surface. +""" + +from __future__ import annotations + +import hashlib +import json +import re +from collections.abc import Mapping +from dataclasses import asdict, dataclass +from typing import Any, Literal + +import numpy as np + +from microcosm.frame import Frame + +__all__ = [ + "PopulationInputContract", + "PopulationInputNotReadyError", + "PopulationInputProfile", + "SchemePopulationMapping", + "validate_population_input_frame", +] + +PopulationSemanticKind = Literal["receipt", "application", "legal_status", "choice"] +PopulationDataKind = Literal["measured", "latent"] +InputReadiness = Literal["ready", "required_missing"] +MappingReadiness = Literal["ready", "required_missing"] +PeriodReadiness = Literal["ready", "exact_alignment_missing"] +CompletenessReadiness = Literal["ready", "complete_imputation_required_missing"] +PublisherSourceReadiness = Literal[ + "ready", + "native_publisher_artifact_required_missing", +] + +_INPUT_KEYS = frozenset( + { + "input_id", + "column", + "entity", + "dtype", + "nullable", + "semantic_kind", + "data_kind", + "owner", + "consumer", + "mechanics_owner", + "axiom_role", + "description", + } +) +_MAPPING_KEYS = frozenset( + { + "mapping_id", + "target_reference", + "input_id", + "chronicle_source_record_id", + "chronicle_entity", + "chronicle_entity_role", + "chronicle_geography_level", + "chronicle_geography_id", + "chronicle_geography_vintage", + "chronicle_period_type", + "chronicle_period", + "microcosm_entity", + "microcosm_geography_level", + "microcosm_geography_id", + "microcosm_geography_vintage", + "microcosm_period_type", + "microcosm_period", + "publisher_source_readiness", + "input_readiness", + "mapping_readiness", + "period_readiness", + "completeness_readiness", + "notes", + } +) +_PROFILE_KEYS = frozenset( + { + "schema_version", + "country", + "profile_id", + "activation", + "description", + "inputs", + "mappings", + } +) +_SEMANTIC_KINDS = frozenset({"receipt", "application", "legal_status", "choice"}) +_DATA_KINDS = frozenset({"measured", "latent"}) +_INPUT_READINESS = frozenset({"ready", "required_missing"}) +_MAPPING_READINESS = frozenset({"ready", "required_missing"}) +_PERIOD_READINESS = frozenset({"ready", "exact_alignment_missing"}) +_COMPLETENESS_READINESS = frozenset({"ready", "complete_imputation_required_missing"}) +_PUBLISHER_SOURCE_READINESS = frozenset( + {"ready", "native_publisher_artifact_required_missing"} +) +_BEHAVIORAL_NAME_TOKENS = ( + "take_up", + "takeup", + "takes_up", + "if_eligible", + "propensity", + "elasticity", +) + + +class PopulationInputNotReadyError(RuntimeError): + """A required population input cannot safely be used yet.""" + + +def _closed_mapping( + raw: object, + keys: frozenset[str], + *, + context: str, +) -> Mapping[str, Any]: + if not isinstance(raw, Mapping): + raise TypeError(f"{context} must be an object.") + actual = set(raw) + if actual != keys: + missing = sorted(keys - actual) + unknown = sorted(actual - keys) + raise ValueError( + f"{context} keys differ; missing={missing}, unknown={unknown}." + ) + return raw + + +def _text(value: object, *, field: str) -> str: + if not isinstance(value, str) or not value.strip(): + raise ValueError(f"{field} must be a non-empty string.") + return value + + +def _identifier(value: object, *, field: str) -> str: + result = _text(value, field=field) + if re.fullmatch(r"[a-z][a-z0-9_]*", result) is None: + raise ValueError(f"{field} must match [a-z][a-z0-9_]*, got {result!r}.") + return result + + +def _period(value: object, *, field: str) -> int | str: + if isinstance(value, bool) or not isinstance(value, (int, str)): + raise ValueError(f"{field} must be an integer or non-empty string.") + if isinstance(value, str) and not value.strip(): + raise ValueError(f"{field} must be an integer or non-empty string.") + return value + + +def _digest(value: object) -> str: + encoded = json.dumps( + value, + sort_keys=True, + separators=(",", ":"), + allow_nan=False, + ) + return hashlib.sha256(encoded.encode("utf-8")).hexdigest() + + +@dataclass(frozen=True) +class PopulationInputContract: + """One null-preserving Microcosm-owned boolean population input column.""" + + input_id: str + column: str + entity: str + dtype: str + nullable: bool + semantic_kind: PopulationSemanticKind + data_kind: PopulationDataKind + owner: str + consumer: str + mechanics_owner: str + axiom_role: str + description: str + + def __post_init__(self) -> None: + _identifier(self.input_id, field="population input input_id") + _identifier(self.column, field=f"population input {self.input_id!r} column") + _identifier(self.entity, field=f"population input {self.input_id!r} entity") + if self.dtype != "bool": + raise ValueError( + f"population input {self.input_id!r} dtype must be 'bool'." + ) + if self.nullable is not True: + raise ValueError( + f"population input {self.input_id!r} nullable must be true so " + "unknown status remains null until an explicit completeness gate." + ) + if self.semantic_kind not in _SEMANTIC_KINDS: + raise ValueError( + f"population input {self.input_id!r} has unsupported semantic_kind " + f"{self.semantic_kind!r}." + ) + if self.data_kind not in _DATA_KINDS: + raise ValueError( + f"population input {self.input_id!r} has unsupported data_kind " + f"{self.data_kind!r}." + ) + if self.owner != "Microcosm": + raise ValueError( + f"population input {self.input_id!r} owner must be 'Microcosm'." + ) + if self.consumer != "PolicyEngine" or self.mechanics_owner != "PolicyEngine": + raise ValueError( + f"population input {self.input_id!r} consumer and mechanics_owner " + "must both be 'PolicyEngine'." + ) + if self.axiom_role != "none": + raise ValueError( + f"population input {self.input_id!r} axiom_role must be 'none'; " + "this contract cannot invent an Axiom input or formula." + ) + _text(self.description, field=f"population input {self.input_id!r} description") + for field, value in (("input_id", self.input_id), ("column", self.column)): + normalized = value.casefold() + forbidden = [ + token for token in _BEHAVIORAL_NAME_TOKENS if token in normalized + ] + if forbidden: + raise ValueError( + f"population input {self.input_id!r} {field} {value!r} looks " + f"behavioral ({forbidden}); declare only receipt, application, " + "legal-status, or choice data." + ) + + @classmethod + def from_mapping(cls, raw: object) -> PopulationInputContract: + """Parse one closed-world population-input declaration.""" + + values = _closed_mapping(raw, _INPUT_KEYS, context="population input") + return cls(**values) + + +@dataclass(frozen=True) +class SchemePopulationMapping: + """Declared Chronicle scheme population to one Microcosm input column.""" + + mapping_id: str + target_reference: str + input_id: str + chronicle_source_record_id: str + chronicle_entity: str + chronicle_entity_role: str + chronicle_geography_level: str + chronicle_geography_id: str + chronicle_geography_vintage: str + chronicle_period_type: str + chronicle_period: int | str + microcosm_entity: str + microcosm_geography_level: str + microcosm_geography_id: str + microcosm_geography_vintage: str + microcosm_period_type: str + microcosm_period: int | str + publisher_source_readiness: PublisherSourceReadiness + input_readiness: InputReadiness + mapping_readiness: MappingReadiness + period_readiness: PeriodReadiness + completeness_readiness: CompletenessReadiness + notes: str + + def __post_init__(self) -> None: + _identifier(self.mapping_id, field="scheme-population mapping_id") + _text( + self.target_reference, + field=f"scheme-population mapping {self.mapping_id!r} target_reference", + ) + _identifier( + self.input_id, + field=f"scheme-population mapping {self.mapping_id!r} input_id", + ) + _text( + self.chronicle_source_record_id, + field=( + f"scheme-population mapping {self.mapping_id!r} " + "chronicle_source_record_id" + ), + ) + for field_name in ( + "chronicle_entity", + "chronicle_entity_role", + "chronicle_geography_level", + "chronicle_period_type", + "microcosm_entity", + "microcosm_geography_level", + "microcosm_period_type", + ): + _identifier( + getattr(self, field_name), + field=f"scheme-population mapping {self.mapping_id!r} {field_name}", + ) + for field_name in ( + "chronicle_geography_id", + "chronicle_geography_vintage", + "microcosm_geography_id", + "microcosm_geography_vintage", + ): + _text( + getattr(self, field_name), + field=f"scheme-population mapping {self.mapping_id!r} {field_name}", + ) + _period( + self.chronicle_period, + field=f"scheme-population mapping {self.mapping_id!r} chronicle_period", + ) + _period( + self.microcosm_period, + field=f"scheme-population mapping {self.mapping_id!r} microcosm_period", + ) + if self.publisher_source_readiness not in _PUBLISHER_SOURCE_READINESS: + raise ValueError( + f"scheme-population mapping {self.mapping_id!r} has unknown " + "publisher_source_readiness " + f"{self.publisher_source_readiness!r}." + ) + if self.input_readiness not in _INPUT_READINESS: + raise ValueError( + f"scheme-population mapping {self.mapping_id!r} has unknown " + f"input_readiness {self.input_readiness!r}." + ) + if self.mapping_readiness not in _MAPPING_READINESS: + raise ValueError( + f"scheme-population mapping {self.mapping_id!r} has unknown " + f"mapping_readiness {self.mapping_readiness!r}." + ) + if self.period_readiness not in _PERIOD_READINESS: + raise ValueError( + f"scheme-population mapping {self.mapping_id!r} has unknown " + f"period_readiness {self.period_readiness!r}." + ) + if self.completeness_readiness not in _COMPLETENESS_READINESS: + raise ValueError( + f"scheme-population mapping {self.mapping_id!r} has unknown " + "completeness_readiness " + f"{self.completeness_readiness!r}." + ) + _text(self.notes, field=f"scheme-population mapping {self.mapping_id!r} notes") + + if ( + self.chronicle_geography_level == "statistical_scope" + and self.microcosm_geography_level != "statistical_scope" + ): + raise ValueError( + f"scheme-population mapping {self.mapping_id!r} cannot reinterpret " + "a Chronicle statistical_scope as a resident NUTS or other " + "geography." + ) + if self.mapping_readiness == "ready": + source_identity = ( + self.chronicle_entity, + self.chronicle_geography_level, + self.chronicle_geography_id, + self.chronicle_geography_vintage, + ) + population_identity = ( + self.microcosm_entity, + self.microcosm_geography_level, + self.microcosm_geography_id, + self.microcosm_geography_vintage, + ) + if source_identity != population_identity: + raise ValueError( + f"scheme-population mapping {self.mapping_id!r} marked ready " + "without an exact entity/geography identity." + ) + if self.period_readiness == "ready" and ( + self.chronicle_period_type, + self.chronicle_period, + ) != (self.microcosm_period_type, self.microcosm_period): + raise ValueError( + f"scheme-population mapping {self.mapping_id!r} marked period " + "ready without an exact period identity." + ) + + @classmethod + def from_mapping(cls, raw: object) -> SchemePopulationMapping: + """Parse one closed-world scheme-population mapping.""" + + values = _closed_mapping( + raw, + _MAPPING_KEYS, + context="scheme-population mapping", + ) + return cls(**values) + + @property + def blockers(self) -> tuple[str, ...]: + """Readiness fields that still block execution.""" + + statuses = { + "publisher_source_readiness": self.publisher_source_readiness, + "input_readiness": self.input_readiness, + "mapping_readiness": self.mapping_readiness, + "period_readiness": self.period_readiness, + "completeness_readiness": self.completeness_readiness, + } + return tuple( + f"{key}={value!r}" for key, value in statuses.items() if value != "ready" + ) + + def require_ready(self) -> None: + """Refuse execution until source, input, mapping, and period are ready.""" + + if self.blockers: + raise PopulationInputNotReadyError( + f"scheme-population mapping {self.mapping_id!r} is not ready: " + + ", ".join(self.blockers) + + "." + ) + + +@dataclass(frozen=True) +class PopulationInputProfile: + """A value-free inventory of inputs and exact scheme mappings.""" + + country: str + profile_id: str + description: str + inputs: tuple[PopulationInputContract, ...] + mappings: tuple[SchemePopulationMapping, ...] + + @classmethod + def from_mapping( + cls, + raw: object, + *, + country: str, + ) -> PopulationInputProfile: + """Parse and cross-check a closed, explicitly activated profile.""" + + values = _closed_mapping(raw, _PROFILE_KEYS, context="population input profile") + if type(values["schema_version"]) is not int or values["schema_version"] != 1: + raise ValueError("population input profile schema_version must be 1.") + if values["country"] != country: + raise ValueError( + f"population input profile country must match {country!r}." + ) + if values["activation"] != "explicit_only": + raise ValueError( + "population input profile activation must be 'explicit_only'." + ) + _identifier(country, field="population input profile country") + profile_id = _identifier( + values["profile_id"], field="population input profile profile_id" + ) + description = _text( + values["description"], field="population input profile description" + ) + raw_inputs = values["inputs"] + raw_mappings = values["mappings"] + if not isinstance(raw_inputs, list) or not raw_inputs: + raise ValueError( + "population input profile inputs must be a non-empty list." + ) + if not isinstance(raw_mappings, list) or not raw_mappings: + raise ValueError( + "population input profile mappings must be a non-empty list." + ) + inputs = tuple(PopulationInputContract.from_mapping(row) for row in raw_inputs) + mappings = tuple( + SchemePopulationMapping.from_mapping(row) for row in raw_mappings + ) + cls._validate_links(inputs, mappings) + return cls(country, profile_id, description, inputs, mappings) + + @staticmethod + def _validate_links( + inputs: tuple[PopulationInputContract, ...], + mappings: tuple[SchemePopulationMapping, ...], + ) -> None: + input_ids = [row.input_id for row in inputs] + columns = [(row.entity, row.column) for row in inputs] + mapping_ids = [row.mapping_id for row in mappings] + if len(input_ids) != len(set(input_ids)): + raise ValueError("population input profile has duplicate input ids.") + if len(columns) != len(set(columns)): + raise ValueError( + "population input profile has duplicate entity/column inputs." + ) + if len(mapping_ids) != len(set(mapping_ids)): + raise ValueError("population input profile has duplicate mapping ids.") + + input_by_id = {row.input_id: row for row in inputs} + used_inputs: set[str] = set() + for mapping in mappings: + input_contract = input_by_id.get(mapping.input_id) + if input_contract is None: + raise ValueError( + f"scheme-population mapping {mapping.mapping_id!r} references " + f"unknown input {mapping.input_id!r}." + ) + if input_contract.entity != mapping.microcosm_entity: + raise ValueError( + f"scheme-population mapping {mapping.mapping_id!r} entity " + f"{mapping.microcosm_entity!r} does not match input entity " + f"{input_contract.entity!r}." + ) + used_inputs.add(mapping.input_id) + orphaned = sorted(set(input_ids) - used_inputs) + if orphaned: + raise ValueError( + "population input profile has inputs with no scheme-population " + f"mapping: {orphaned}." + ) + + def input(self, input_id: str) -> PopulationInputContract: + """Return a declared input by id, refusing omissions.""" + + for contract in self.inputs: + if contract.input_id == input_id: + return contract + raise KeyError(f"Unknown population input {input_id!r}.") + + def mapping(self, mapping_id: str) -> SchemePopulationMapping: + """Return a declared mapping by id, refusing bypass by omission.""" + + for mapping in self.mappings: + if mapping.mapping_id == mapping_id: + return mapping + raise KeyError(f"Unknown scheme-population mapping {mapping_id!r}.") + + +def _canonical_row_id(value: object, *, column: str) -> int | str: + if isinstance(value, np.generic): + value = value.item() + if isinstance(value, bool) or not isinstance(value, (int, str)): + raise ValueError( + f"population input row identity column {column!r} must contain only " + "integer or string ids." + ) + return value + + +def validate_population_input_frame( + frame: Frame, + profile: PopulationInputProfile, + *, + mapping_id: str, +) -> Mapping[str, object]: + """Validate one ready Frame column and return a row/value identity receipt. + + Readiness is checked before the Frame is inspected. The receipt contains + only counts and hashes, never raw microdata ids or values. + """ + + if not isinstance(profile, PopulationInputProfile): + raise TypeError("profile must be a PopulationInputProfile.") + mapping = profile.mapping(mapping_id) + mapping.require_ready() + if not isinstance(frame, Frame): + raise TypeError("frame must be a microcosm Frame.") + + contract = profile.input(mapping.input_id) + table = frame.table(contract.entity) + if contract.column not in table.columns: + raise ValueError( + f"ready population input {contract.input_id!r} is missing Frame column " + f"{contract.column!r} on entity {contract.entity!r}." + ) + values = table[contract.column] + if values.empty: + raise ValueError( + f"ready population input {contract.input_id!r} has no entity rows." + ) + if bool(values.isna().any()): + raise ValueError( + f"ready population input {contract.input_id!r} contains missing values." + ) + raw_values = values.to_numpy(dtype=object, copy=True) + if not all(isinstance(value, (bool, np.bool_)) for value in raw_values): + raise ValueError( + f"ready population input {contract.input_id!r} must contain only " + "boolean values; integer 0/1 and other proxies are refused." + ) + boolean_values = [bool(value) for value in raw_values] + + id_column = frame.schema.entity_id_column(contract.entity) + row_ids = [ + _canonical_row_id(value, column=id_column) + for value in table[id_column].to_numpy(dtype=object, copy=True) + ] + contract_payload = { + "input": asdict(contract), + "mapping": asdict(mapping), + } + unsigned_receipt: dict[str, object] = { + "schema_version": 1, + "country": profile.country, + "profile_id": profile.profile_id, + "mapping_id": mapping.mapping_id, + "target_reference": mapping.target_reference, + "input_id": contract.input_id, + "entity": contract.entity, + "column": contract.column, + "semantic_kind": contract.semantic_kind, + "data_kind": contract.data_kind, + "chronicle_entity_role": mapping.chronicle_entity_role, + "chronicle_source_record_id": mapping.chronicle_source_record_id, + "chronicle_geography_level": mapping.chronicle_geography_level, + "chronicle_geography_id": mapping.chronicle_geography_id, + "chronicle_geography_vintage": mapping.chronicle_geography_vintage, + "chronicle_period_type": mapping.chronicle_period_type, + "chronicle_period": mapping.chronicle_period, + "publisher_source_readiness": mapping.publisher_source_readiness, + "n_rows": len(boolean_values), + "n_true": sum(boolean_values), + "n_false": len(boolean_values) - sum(boolean_values), + "n_unknown": 0, + "contract_sha256": _digest(contract_payload), + "row_ids_sha256": _digest({"entity": contract.entity, "row_ids": row_ids}), + "values_sha256": _digest({"column": contract.column, "values": boolean_values}), + "row_values_sha256": _digest( + { + "entity": contract.entity, + "column": contract.column, + "row_values": list(zip(row_ids, boolean_values, strict=True)), + } + ), + } + return { + **unsigned_receipt, + "receipt_sha256": _digest(unsigned_receipt), + } diff --git a/packages/microcosm-build/tests/golden/be_country_spec.json b/packages/microcosm-build/tests/golden/be_country_spec.json index 5cfc5cea4..b650cd323 100644 --- a/packages/microcosm-build/tests/golden/be_country_spec.json +++ b/packages/microcosm-build/tests/golden/be_country_spec.json @@ -1,6 +1,6 @@ { "country": "be", - "fingerprint": "6f1dab96ccedb21fc261acf9ebbccfc036b2e354a78cf9314e394707a1888cf6", + "fingerprint": "75448d958ef56404fd343ae0ba19bc047f7d4f1d739b06c0970fb4281d618772", "gate_ids": [ "calibration_per_family_fit", "national_and_nuts1_admin_aggregates", @@ -28,9 +28,11 @@ "staging_repo": "policyengine/populace-be-staging" }, "resource_hashes": { - "country_package.json": "1dff9052dbdcdbceeb077d1212008226cf783aa5d33cbda52d07437aef69d30b", - "gates.json": "85c07f713168058a081b40cf94da6a33035108ba7807e04818486032fb2ad580", + "country_package.json": "73f1411a3e0541204bb1f6d060061dfbe8cd39477a2e0819c12d1bcb552c0036", + "gates.json": "2a607384834002a919f4068e297321fa68f925f13afefa4f8fe145f1be3c3945", "geography_spine.json": "0b39217386dbc9b053a036766004f10a3a34434a293c1e956e0f4f053a08e1bf", + "monetary_target_profile.json": "91d8114a1fb5896800cf15f2f524d5cb1090b0ee7e895382311174e0eab17f60", + "population_inputs.json": "bb67b50d099d3206f3d8db6c9ff7291cff06779617b081111f0e1f4780e7e7ca", "release_contract.json": "af956e05782cd2aa5df108b216f0ff2a835f2a4e87d68419239b537550082c25", "source_stages.json": "64e45f48b5c3bb00a710e56f805c7c8c441ddc1050cfb61d846d45e9cbbb16b2", "spec/bundle.yaml": "182e2a0f0b12cef6db6a0192438721b2e041b2c72ef16fbafde8f13fd7b93021", @@ -39,7 +41,7 @@ "spec/sources.yaml": "1ca516ad27ee815cb60f451901dcdc6b8989c4f1731d0f6f3f1e50714e29e35e", "spec/spine.yaml": "908d99488c3905806ec6878d96b0e4a4cbb67c6ec69d53fe5f1d160246fdb513", "spec/vintages.yaml": "c842490287404cf306907800b0bcadec5d954e7e9c510f4b293786c5bedd0910", - "target_references.json": "a8c9f875e3a04d082fbacfa8511c856b3879077d8a4276df50ea38c345e702ea" + "target_references.json": "31717d4d48cb94e9a1552ea80aa4e1e2b4c3c5f6b00f8aa7fa0faab146fffdb5" }, "resources": [ "spec/bundle.yaml", @@ -50,6 +52,8 @@ "spec/vintages.yaml", "gates.json", "geography_spine.json", + "monetary_target_profile.json", + "population_inputs.json", "release_contract.json", "source_stages.json", "target_references.json" @@ -58,11 +62,20 @@ "silc_load" ], "target_reference_names": [ - "statbel_population_by_age_sex_region", - "statbel_fiscal_income_by_commune", - "spf_finances_pit_total", - "onss_employee_contribution_total", - "onem_unemployment_caseload", - "nbb_household_disposable_income" + "statbel_population_by_age_sex_region_2025", + "statbel_fiscal_income_by_commune_2023", + "statbel_zero_income_tax_returns_2023_validation", + "onem_complete_unemployment_monthly_average_2024", + "sfpd_legal_pension_all_schemes_january_2025_blocked", + "sfpd_grapa_regular_payment_beneficiaries_2025_01", + "opgroeien_groeipakket_basic_amount_children_2025_12", + "iriscare_entitled_children_2025_12", + "iriscare_payment_recipients_2025_12", + "ostbelgien_paid_children_2025_12", + "ostbelgien_payment_recipients_2025_12", + "walloon_single_parent_with_supplement_children_2023_12", + "walloon_single_parent_without_supplement_children_2023_12", + "walloon_other_household_with_supplement_children_2023_12", + "walloon_other_household_without_supplement_children_2023_12" ] } diff --git a/packages/microcosm-build/tests/test_country_spec.py b/packages/microcosm-build/tests/test_country_spec.py index df94c31f6..292d00681 100644 --- a/packages/microcosm-build/tests/test_country_spec.py +++ b/packages/microcosm-build/tests/test_country_spec.py @@ -183,6 +183,7 @@ def _schema2_target_resource() -> dict[str, object]: "source_name": "official_population", "source_measure_id": "people", "period_type": "calendar_year", + "period_value": 2023, "geography_level": "country", }, "entity": "person", @@ -491,6 +492,8 @@ def test_loads_with_every_declared_resource(self, spec) -> None: "source_stages.json", "geography_spine.json", "target_references.json", + "population_inputs.json", + "monetary_target_profile.json", "gates.json", "release_contract.json", } @@ -524,16 +527,25 @@ def test_geography_spine_is_vintage_aware(self, spec) -> None: def test_targets_arrive_by_reference_with_no_values(self, spec) -> None: names = {reference.name for reference in spec.target_references} - assert { - "statbel_population_by_age_sex_region", - "statbel_fiscal_income_by_commune", - "spf_finances_pit_total", - "onss_employee_contribution_total", - "onem_unemployment_caseload", - "nbb_household_disposable_income", - } <= names + assert names == { + "statbel_population_by_age_sex_region_2025", + "statbel_fiscal_income_by_commune_2023", + "statbel_zero_income_tax_returns_2023_validation", + "onem_complete_unemployment_monthly_average_2024", + "sfpd_legal_pension_all_schemes_january_2025_blocked", + "sfpd_grapa_regular_payment_beneficiaries_2025_01", + "opgroeien_groeipakket_basic_amount_children_2025_12", + "iriscare_entitled_children_2025_12", + "iriscare_payment_recipients_2025_12", + "ostbelgien_paid_children_2025_12", + "ostbelgien_payment_recipients_2025_12", + "walloon_single_parent_with_supplement_children_2023_12", + "walloon_single_parent_without_supplement_children_2023_12", + "walloon_other_household_with_supplement_children_2023_12", + "walloon_other_household_without_supplement_children_2023_12", + } by_name = {reference.name: reference for reference in spec.target_references} - commune = by_name["statbel_fiscal_income_by_commune"] + commune = by_name["statbel_fiscal_income_by_commune_2023"] assert commune.metadata["nis_vintage"] == "2025" assert commune.metadata["geography_vintage"] == "nis_2025" assert commune.ledger_selector["geography_vintage"] == "nis_2025" @@ -541,72 +553,71 @@ def test_targets_arrive_by_reference_with_no_values(self, spec) -> None: assert { by_name[name].metadata["activation_status"] for name in { - "statbel_population_by_age_sex_region", - "statbel_fiscal_income_by_commune", + "statbel_population_by_age_sex_region_2025", + "statbel_fiscal_income_by_commune_2023", } - } == {"requires_harvested_cell_references"} + } == { + "requires_harvested_cells_and_geography_period_bridge", + "requires_harvested_cells_and_income_period_bridge", + } - payload = json.loads( - (COUNTRY_PACKAGE_ROOT / "be/target_references.json").read_text( - encoding="utf-8" + for resource_name, collection_name in ( + ("target_references.json", "target_references"), + ("monetary_target_profile.json", "targets"), + ): + payload = json.loads( + (COUNTRY_PACKAGE_ROOT / f"be/{resource_name}").read_text( + encoding="utf-8" + ) + ) + assert FORBIDDEN_TARGET_VALUE_KEYS.isdisjoint( + _nested_mapping_keys(payload[collection_name]) ) - ) - assert FORBIDDEN_TARGET_VALUE_KEYS.isdisjoint( - _nested_mapping_keys(payload["target_references"]) - ) def test_target_selectors_declare_the_intended_chronicle_vocabulary( self, spec ) -> None: references = {reference.name: reference for reference in spec.target_references} expected = { - "statbel_population_by_age_sex_region": ( + "statbel_population_by_age_sex_region_2025": ( "statbel_population_structure", "people", "calendar_year", - 2023, + 2025, "nuts1", - "nuts1_2025", + "NUTS_2024", ), - "statbel_fiscal_income_by_commune": ( + "statbel_fiscal_income_by_commune_2023": ( "statbel_fiscal_income", "taxable_income", "tax_year", - 2022, + 2023, "commune", "nis_2025", ), - "spf_finances_pit_total": ( - "spf_finances_pit", - "tax_before_withholding", + "statbel_zero_income_tax_returns_2023_validation": ( + "statbel_fiscal_income_distribution", + "declarations", "tax_year", - 2022, - "country", - None, - ), - "onss_employee_contribution_total": ( - "onss_contributions", - "worker_article_17_uncapped_component_contribution", - "calendar_year", - 2022, + 2023, "country", - None, + "current", ), - "onem_unemployment_caseload": ( + "onem_complete_unemployment_monthly_average_2024": ( "onem_rva_unemployment", "receives_unemployment_benefit", "calendar_year", - 2022, + 2024, "country", - None, + "current", ), - "nbb_household_disposable_income": ( - "nbb_national_accounts", - "household_disposable_income", - "calendar_year", - 2022, + "sfpd_grapa_regular_payment_beneficiaries_2025_01": ( + "sfpd_grapa", + "beneficiaries", + "month", + "2025-01", "country", - None, + "current", ), } @@ -624,68 +635,387 @@ def test_target_selectors_declare_the_intended_chronicle_vocabulary( assert reference.ledger_selector["period_type"] == period_type assert reference.period == period assert reference.period_match_policy == "exact" - assert reference.assertion_policy == "allow_source_projection" + assert reference.assertion_policy == "observed_only" assert reference.ledger_selector["geography_level"] == geography_level assert ( reference.ledger_selector.get("geography_vintage") == geography_vintage ) + def test_benefit_participation_selectors_pin_scalar_chronicle_facts( + self, spec + ) -> None: + references = {reference.name: reference for reference in spec.target_references} + expected_selector_fields = { + "sfpd_grapa_regular_payment_beneficiaries_2025_01": { + "record_set_id": ( + "sfpd_grapa.month2025_01.regular_payment_beneficiaries.by_sex" + ), + "period_value": "2025-01", + "geography_id": "BE", + "entity_name": "person", + "layout_groupby_value_id": "all", + "dimensions": [], + }, + } + + for name, expected in expected_selector_fields.items(): + selector = references[name].ledger_selector + assert {key: selector[key] for key in expected} == expected + + def test_grapa_participation_reference_is_a_support_aware_validation( + self, spec + ) -> None: + references = {reference.name: reference for reference in spec.target_references} + grapa = references["sfpd_grapa_regular_payment_beneficiaries_2025_01"] + assert grapa.measure == "belgium_grapa_regular_payment_recipient_indicator" + assert grapa.metadata["target_role"] == "validation" + assert grapa.metadata["criticality_tier"] == "validation_only" + assert grapa.metadata["input_owner"] == "Microcosm" + assert grapa.metadata["input_consumer"] == "PolicyEngine" + assert grapa.metadata["mechanics_owner"] == "PolicyEngine" + assert grapa.metadata["axiom_behavior_ownership"] == "none" + assert ( + grapa.metadata["anti_proxy_rule"] + == "do_not_derive_receipt_from_positive_amount_or_entitlement" + ) + assert grapa.metadata["activation_status"] != "active" + assert grapa.metadata["publisher_source_readiness"] == "ready" + assert grapa.entity == "person" + assert grapa.metadata["population_input_readiness"] == "required_missing" + assert grapa.metadata["population_mapping_readiness"] == "ready" + assert ( + grapa.metadata["population_completeness_readiness"] + == "complete_imputation_required_missing" + ) + assert "consume Microcosm's distinct measured or latent receipt flag" in ( + grapa.notes + ) + + unemployment = references["onem_complete_unemployment_monthly_average_2024"] + assert unemployment.metadata["activation_status"] != "active" + assert unemployment.metadata["input_owner"] == "Microcosm" + assert unemployment.metadata["input_consumer"] == "PolicyEngine" + assert unemployment.metadata["mechanics_owner"] == "PolicyEngine" + assert unemployment.metadata["axiom_behavior_ownership"] == "none" + assert unemployment.metadata["microcosm_population_input_status"] == "absent" + assert unemployment.metadata["policyengine_input_support_status"] == "absent" + assert ( + unemployment.metadata["anti_proxy_rule"] + == "do_not_derive_receipt_from_positive_amount_or_entitlement" + ) + assert "Microcosm-populated unemployment receipt flag" in unemployment.notes + assert "PolicyEngine input variable that consumes it" in unemployment.notes + assert all( + reference.metadata.get("activation_status") != "active" + for reference in spec.target_references + if reference.family == "caseloads" + ) + + def test_scheme_population_references_are_exact_typed_and_fail_closed( + self, spec + ) -> None: + profile = spec.population_input_profile + assert profile is not None + assert len(profile.inputs) == len(profile.mappings) == 10 + inputs = {row.input_id: row for row in profile.inputs} + references = {row.name: row for row in spec.target_references} + + expected_links = { + ( + "grapa_regular_payment_population_2025_01", + "sfpd_grapa_regular_payment_beneficiaries_2025_01", + "sfpd_grapa.month2025_01.regular_payment_beneficiaries.all.beneficiaries", + "grapa_beneficiary", + "BE", + "current", + "belgium_grapa_regular_payment_recipient_indicator", + ), + ( + "groeipakket_basic_amount_child_population_2025_12", + "opgroeien_groeipakket_basic_amount_children_2025_12", + "opgroeien.groeipakket.month2025_12.basic_amount.children.basic_amount_children.children", + "child_benefit_recipient", + "BE-GROEIPAKKET-SCHEME", + "GROEIPAKKET_ADMINISTRATIVE_SCOPE_2025", + "belgium_groeipakket_basic_amount_child_indicator", + ), + ( + "iriscare_entitled_child_population_2025_12", + "iriscare_entitled_children_2025_12", + "iriscare.child_benefit.month2025_12.entitled_children.entitled_children.children", + "child_benefit_entitled_child", + "BE-IRISCARE-CHILD-BENEFIT-SCHEME", + "IRISCARE_CHILD_BENEFIT_ADMIN_SCOPE_2025", + "belgium_iriscare_child_benefit_entitled_child_indicator", + ), + ( + "iriscare_payment_recipient_population_2025_12", + "iriscare_payment_recipients_2025_12", + "iriscare.child_benefit.month2025_12.payment_recipients.payment_recipients.payment_recipients", + "child_benefit_payment_recipient", + "BE-IRISCARE-CHILD-BENEFIT-SCHEME", + "IRISCARE_CHILD_BENEFIT_ADMIN_SCOPE_2025", + "belgium_iriscare_child_benefit_payment_recipient_indicator", + ), + ( + "ostbelgien_paid_child_population_2025_12", + "ostbelgien_paid_children_2025_12", + "ostbelgien.child_benefit.month2025_12.paid_children.paid_children.children", + "child_benefit_paid_child", + "BE-DG", + "child_benefit_scope_2025", + "belgium_ostbelgien_child_benefit_paid_child_indicator", + ), + ( + "ostbelgien_payment_recipient_population_2025_12", + "ostbelgien_payment_recipients_2025_12", + "ostbelgien.child_benefit.month2025_12.payment_recipients.payment_recipients.payment_recipients", + "child_benefit_payment_recipient", + "BE-DG", + "child_benefit_scope_2025", + "belgium_ostbelgien_child_benefit_payment_recipient_indicator", + ), + ( + "walloon_single_parent_with_supplement_children_2023_12", + "walloon_single_parent_with_supplement_children_2023_12", + "parlement_wallonie.child_benefit.month2023_12.children.single_parent_with_social_supplement.children", + "child_benefit_recipient_child", + "BE-WALLOON-FRENCH", + "child_benefit_scope_2023", + "belgium_walloon_single_parent_with_supplement_recipient_child_indicator", + ), + ( + "walloon_single_parent_without_supplement_children_2023_12", + "walloon_single_parent_without_supplement_children_2023_12", + "parlement_wallonie.child_benefit.month2023_12.children.single_parent_without_social_supplement.children", + "child_benefit_recipient_child", + "BE-WALLOON-FRENCH", + "child_benefit_scope_2023", + "belgium_walloon_single_parent_without_supplement_recipient_child_indicator", + ), + ( + "walloon_other_household_with_supplement_children_2023_12", + "walloon_other_household_with_supplement_children_2023_12", + "parlement_wallonie.child_benefit.month2023_12.children.other_household_with_social_supplement.children", + "child_benefit_recipient_child", + "BE-WALLOON-FRENCH", + "child_benefit_scope_2023", + "belgium_walloon_other_household_with_supplement_recipient_child_indicator", + ), + ( + "walloon_other_household_without_supplement_children_2023_12", + "walloon_other_household_without_supplement_children_2023_12", + "parlement_wallonie.child_benefit.month2023_12.children.other_household_without_social_supplement.children", + "child_benefit_recipient_child", + "BE-WALLOON-FRENCH", + "child_benefit_scope_2023", + "belgium_walloon_other_household_without_supplement_recipient_child_indicator", + ), + } + actual_links = { + ( + row.mapping_id, + row.target_reference, + row.chronicle_source_record_id, + row.chronicle_entity_role, + row.chronicle_geography_id, + row.chronicle_geography_vintage, + inputs[row.input_id].column, + ) + for row in profile.mappings + } + assert actual_links == expected_links + + assert all(row.entity == "person" and row.nullable for row in profile.inputs) + assert all(row.owner == "Microcosm" for row in profile.inputs) + assert all(row.consumer == "PolicyEngine" for row in profile.inputs) + assert all(row.mechanics_owner == "PolicyEngine" for row in profile.inputs) + assert all(row.axiom_role == "none" for row in profile.inputs) + assert all( + row.input_readiness == "required_missing" for row in profile.mappings + ) + assert all( + row.period_readiness == "exact_alignment_missing" + for row in profile.mappings + ) + assert all( + row.completeness_readiness == "complete_imputation_required_missing" + for row in profile.mappings + ) + assert {row.mapping_readiness for row in profile.mappings} == { + "ready", + "required_missing", + } + assert sum(row.mapping_readiness == "ready" for row in profile.mappings) == 1 + assert ( + sum( + row.publisher_source_readiness + == "native_publisher_artifact_required_missing" + for row in profile.mappings + ) + == 5 + ) + + for mapping in profile.mappings: + reference = references[mapping.target_reference] + input_contract = inputs[mapping.input_id] + assert reference.entity == mapping.microcosm_entity == "person" + assert reference.measure == input_contract.column + assert reference.filter is None + assert reference.ledger_source_record_id == ( + mapping.chronicle_source_record_id + ) + assert reference.period == mapping.chronicle_period + assert reference.ledger_selector["period_value"] == ( + mapping.chronicle_period + ) + assert reference.ledger_selector["geography_id"] == ( + mapping.chronicle_geography_id + ) + assert reference.metadata["publisher_source_readiness"] == ( + mapping.publisher_source_readiness + ) + assert reference.metadata["activation_status"] != "active" + + child_mappings = profile.mappings[1:] + assert all( + row.mapping_readiness == "required_missing" for row in child_mappings + ) + pending_native_sources = { + row.mapping_id + for row in child_mappings + if row.publisher_source_readiness + == "native_publisher_artifact_required_missing" + } + assert pending_native_sources == { + "groeipakket_basic_amount_child_population_2025_12", + "iriscare_entitled_child_population_2025_12", + "iriscare_payment_recipient_population_2025_12", + "ostbelgien_paid_child_population_2025_12", + "ostbelgien_payment_recipient_population_2025_12", + } + assert all( + "requires_native_publisher_artifact" + in references[row.target_reference].metadata["activation_status"] + for row in child_mappings + if row.mapping_id in pending_native_sources + ) + assert all( + row.chronicle_geography_level == "statistical_scope" + for row in child_mappings + ) + assert all( + references[row.target_reference].ledger_selector["entity_name"] == "person" + for row in child_mappings + ) + assert all( + references[row.target_reference].ledger_selector["source_measure_id"] + not in {"households", "families"} + for row in child_mappings + ) + walloon_ids = { + row.chronicle_source_record_id + for row in child_mappings + if row.chronicle_geography_id == "BE-WALLOON-FRENCH" + } + assert len(walloon_ids) == 4 + assert all(".month2023_12.children." in value for value in walloon_ids) + assert all("children_by_household_group" not in value for value in walloon_ids) + def test_target_profile_declares_tiers_and_income_basis(self, spec) -> None: profile = spec.target_profile assert profile["schema_version"] == 2 assert tuple(profile["required_families"]) == ( "demography", "fiscal_income", - "income_tax", - "social_security", "caseloads", ) tiers = profile["criticality_tiers"] - assert tiers["core_fiscal_release"]["relative_tolerance"] == 0.05 - assert tiers["caseload_release"]["relative_tolerance"] == 0.15 + assert tiers["demography_candidate"]["relative_tolerance"] == 0.02 + assert tiers["commune_diagnostic"]["relative_tolerance"] == 0.1 + assert tiers["caseload_candidate"]["relative_tolerance"] == 0.15 assert tiers["validation_only"]["relative_tolerance"] is None - income_basis = profile["basis_periods"]["assessment_income_year_2022"] - assert income_basis["period"] == 2022 + income_basis = profile["basis_periods"]["statbel_fiscal_income_2023"] + assert income_basis["period"] == 2023 assert income_basis["fact_period_type"] == "tax_year" - assert income_basis["survey_year"] == 2023 - assert income_basis["income_reference_offset_years"] == -1 - assert income_basis["mismatch_policy"] == "requires_source_projection" + assert income_basis["mismatch_policy"] == "exact_observation_only" + + grapa_basis = profile["basis_periods"]["grapa_regular_payment_snapshot_2025_01"] + assert grapa_basis["period"] == "2025-01" + assert grapa_basis["basis"] == "regular_payment_snapshot" + assert grapa_basis["fact_period_type"] == "month" references = {reference.name: reference for reference in spec.target_references} - assert ( - references["nbb_household_disposable_income"].metadata["target_role"] - == "validation" - ) assert { reference.family for reference in references.values() if reference.metadata["target_role"] == "calibration" - } >= set(profile["required_families"]) + } == set(profile["required_families"]) + + monetary = spec.monetary_target_profile + assert monetary is not None + monetary_by_name = {row.reference.name: row for row in monetary.targets} + assert set(monetary_by_name) == { + "spf_finances_pit_total_2023", + "onss_worker_personal_contributions_2024_validation", + "nbb_household_disposable_income_2024_validation", + } + pit = monetary_by_name["spf_finances_pit_total_2023"] + assert pit.reference.family == "income_tax" + assert pit.reference.metadata["monetary_target_role"] == "calibration" + assert pit.readiness == "requires_policy_output" + for name in ( + "onss_worker_personal_contributions_2024_validation", + "nbb_household_disposable_income_2024_validation", + ): + assert monetary_by_name[name].reference.metadata[ + "monetary_target_role" + ] == ("validation") + assert monetary_by_name[name].readiness == "historical_validation_only" + + coverage = next( + gate for gate in spec.gates.gates if gate.id == "target_profile_coverage" + ) + assert set(coverage.parameters["required_families"]) == { + *profile["required_families"], + "income_tax", + } + assert "social_security" not in coverage.parameters["required_families"] - def test_target_profile_tiers_and_roles_are_declaration_only(self, spec) -> None: + def test_target_profile_tiers_are_declarations_and_validation_is_excluded( + self, spec + ) -> None: assert all(reference.tolerance is None for reference in spec.target_references) description = json.loads( (COUNTRY_PACKAGE_ROOT / "be/target_references.json").read_text( encoding="utf-8" ) )["description"] - assert "validated declaration metadata only" in description - assert "does not yet wire them into runtime calibration" in description - assert "current Chronicle Belgian catalog does not satisfy" in description - assert "#264" in description - assert "Declaration-only intended Belgian gate posture" in spec.gates.policy + assert "Values remain in Chronicle" in description + assert "Unknown receipt or status remains null, never false" in description + assert "PolicyEngine consumes Microcosm-populated flags" in description + assert ( + "PolicyEngine" in description and "does not supply the flags" in description + ) + assert "Axiom receives no synthetic behavior concept" in description + assert "Target-profile tier tolerances remain declarations" in ( + spec.gates.policy + ) + assert "target_role=validation is already an enforced exclusion" in ( + spec.gates.policy + ) assert "not implemented here" in spec.gates.policy @pytest.mark.parametrize( ("reference_name", "aliases"), [ ( - "statbel_population_by_age_sex_region", + "statbel_population_by_age_sex_region_2025", ("be_nuts1_2025", "nuts1_2025", "2025_nuts1"), ), ( - "statbel_fiscal_income_by_commune", + "statbel_fiscal_income_by_commune_2023", ("be_nis_2025", "nis_2025", "2025_nis"), ), ], @@ -716,8 +1046,7 @@ def test_subnational_targets_accept_only_declared_typed_vintage_aliases( @pytest.mark.parametrize( ("reference_name", "invalid_vintage"), [ - ("statbel_population_by_age_sex_region", "NUTS_2024"), - ("statbel_fiscal_income_by_commune", "nis_2024"), + ("statbel_fiscal_income_by_commune_2023", "nis_2024"), ], ) def test_subnational_targets_refuse_vintages_outside_typed_registry( @@ -737,6 +1066,29 @@ def test_subnational_targets_refuse_vintages_outside_typed_registry( with pytest.raises(ValueError, match="not an exact typed authority alias"): load_country_spec(package_dir) + @pytest.mark.parametrize( + "missing_field", ["geography_bridge_status", "activation_status"] + ) + def test_population_source_vintage_requires_an_explicit_missing_bridge( + self, tmp_path, missing_field + ) -> None: + package_dir = tmp_path / "be" + shutil.copytree(COUNTRY_PACKAGE_ROOT / "be", package_dir) + target_path = package_dir / "target_references.json" + payload = json.loads(target_path.read_text(encoding="utf-8")) + reference = next( + row + for row in payload["target_references"] + if row["name"] == "statbel_population_by_age_sex_region_2025" + ) + reference["metadata"].pop(missing_field) + target_path.write_text(json.dumps(payload), encoding="utf-8") + + with pytest.raises( + ValueError, match="explicit non-active required_missing bridge" + ): + load_country_spec(package_dir) + def test_subnational_target_requires_a_typed_geography_layer( self, tmp_path ) -> None: @@ -752,22 +1104,215 @@ def test_subnational_target_requires_a_typed_geography_layer( load_country_spec(package_dir) @pytest.mark.parametrize( - "reference_name", + ("reference_name", "message"), [ - "statbel_population_by_age_sex_region", - "statbel_fiscal_income_by_commune", + ( + "statbel_population_by_age_sex_region_2025", + "non-executable placeholder", + ), + ( + "statbel_fiscal_income_by_commune_2023", + "non-executable placeholder", + ), + ( + "onem_complete_unemployment_monthly_average_2024", + "non-executable placeholder", + ), + ( + "sfpd_grapa_regular_payment_beneficiaries_2025_01", + "validation-only", + ), ], ) - def test_multicell_be_placeholders_cannot_compile_before_fanout( - self, spec, reference_name + def test_nonactive_be_target_references_cannot_compile( + self, spec, reference_name, message ) -> None: reference = next( row for row in spec.target_references if row.name == reference_name ) - with pytest.raises(ValueError, match="non-executable placeholder"): + with pytest.raises(ValueError, match=message): compile_ledger_target_references([], [reference], country="be") + @pytest.mark.parametrize( + ("reference_name", "blocker"), + [ + ( + "onem_complete_unemployment_monthly_average_2024", + "microcosm_population_input_status='absent'", + ), + ( + "sfpd_grapa_regular_payment_beneficiaries_2025_01", + "population_completeness_readiness=" + "'complete_imputation_required_missing'", + ), + ( + "opgroeien_groeipakket_basic_amount_children_2025_12", + "publisher_source_readiness=" + "'native_publisher_artifact_required_missing'", + ), + ], + ) + def test_be_target_activation_is_fail_closed_while_support_is_unresolved( + self, spec, reference_name, blocker + ) -> None: + reference = next( + row for row in spec.target_references if row.name == reference_name + ) + active_reference = replace( + reference, + metadata={ + **reference.metadata, + "activation_status": "active", + "target_role": "calibration", + }, + ) + + with pytest.raises(ValueError, match=blocker): + compile_ledger_target_references([], [active_reference], country="be") + + def test_be_target_activation_refuses_unknown_guarded_status(self, spec) -> None: + reference = next( + row + for row in spec.target_references + if row.name == "sfpd_grapa_regular_payment_beneficiaries_2025_01" + ) + active_reference = replace( + reference, + metadata={ + "activation_status": "active", + "publication_status": "provisional", + "support_status": "pending", + }, + ) + + with pytest.raises(ValueError, match="support_status='pending'"): + compile_ledger_target_references([], [active_reference], country="be") + + def test_be_target_activation_ignores_unrelated_publication_status( + self, spec + ) -> None: + reference = next( + row + for row in spec.target_references + if row.name == "sfpd_grapa_regular_payment_beneficiaries_2025_01" + ) + active_reference = replace( + reference, + metadata={ + "activation_status": "active", + "publication_status": "provisional", + "support_status": "ready", + "target_role": "calibration", + }, + ) + + with pytest.raises(ValueError, match="did not match a Ledger fact selector"): + compile_ledger_target_references([], [active_reference], country="be") + + def test_validation_only_reference_never_enters_calibration(self, spec) -> None: + reference = next( + row + for row in spec.target_references + if row.name == "statbel_zero_income_tax_returns_2023_validation" + ) + active_reference = replace( + reference, + metadata={**reference.metadata, "activation_status": "active"}, + ) + + with pytest.raises(ValueError, match="validation-only"): + compile_ledger_target_references([], [active_reference], country="be") + + @pytest.mark.parametrize( + ("mutation", "message"), + [ + ("metadata_owner", "metadata mismatch"), + ("publisher_source", "metadata mismatch"), + ("source_record", "source record id"), + ("selector_period", "selector period"), + ("selector_scope", "Chronicle selector mismatch"), + ("active_blocked", "cannot be active while"), + ], + ) + def test_scheme_population_cross_links_fail_closed( + self, tmp_path, mutation, message + ) -> None: + package_dir = tmp_path / "be" + shutil.copytree(COUNTRY_PACKAGE_ROOT / "be", package_dir) + target_path = package_dir / "target_references.json" + payload = json.loads(target_path.read_text(encoding="utf-8")) + reference = next( + row + for row in payload["target_references"] + if row["name"] == "sfpd_grapa_regular_payment_beneficiaries_2025_01" + ) + if mutation == "metadata_owner": + reference["metadata"]["input_owner"] = "PolicyEngine" + elif mutation == "publisher_source": + reference["metadata"]["publisher_source_readiness"] = ( + "native_publisher_artifact_required_missing" + ) + elif mutation == "source_record": + reference["ledger_source_record_id"] += ".wrong" + elif mutation == "selector_period": + reference["ledger_selector"]["period_value"] = "2025-02" + elif mutation == "selector_scope": + reference["ledger_selector"]["geography_id"] = "BE2" + else: + reference["metadata"]["activation_status"] = "active" + target_path.write_text(json.dumps(payload), encoding="utf-8") + + with pytest.raises(ValueError, match=message): + load_country_spec(package_dir) + + def test_unmapped_target_selector_period_cannot_drift_from_basis( + self, tmp_path + ) -> None: + package_dir = tmp_path / "be" + shutil.copytree(COUNTRY_PACKAGE_ROOT / "be", package_dir) + target_path = package_dir / "target_references.json" + payload = json.loads(target_path.read_text(encoding="utf-8")) + reference = next( + row + for row in payload["target_references"] + if row["name"] == "onem_complete_unemployment_monthly_average_2024" + ) + reference["ledger_selector"]["period_value"] = 2023 + target_path.write_text(json.dumps(payload), encoding="utf-8") + + with pytest.raises(ValueError, match="selector period 2023"): + load_country_spec(package_dir) + + def test_scheme_population_mapping_cannot_link_two_targets(self, tmp_path) -> None: + package_dir = tmp_path / "be" + shutil.copytree(COUNTRY_PACKAGE_ROOT / "be", package_dir) + target_path = package_dir / "target_references.json" + payload = json.loads(target_path.read_text(encoding="utf-8")) + duplicate = next( + row + for row in payload["target_references"] + if row["name"] == "sfpd_grapa_regular_payment_beneficiaries_2025_01" + ).copy() + duplicate = json.loads(json.dumps(duplicate)) + duplicate["name"] = "duplicate_grapa_validation" + payload["target_references"].append(duplicate) + target_path.write_text(json.dumps(payload), encoding="utf-8") + + with pytest.raises(ValueError, match="duplicate links"): + load_country_spec(package_dir) + + def test_belgium_population_mapping_cannot_be_orphaned(self, tmp_path) -> None: + package_dir = tmp_path / "be" + shutil.copytree(COUNTRY_PACKAGE_ROOT / "be", package_dir) + profile_path = package_dir / "population_inputs.json" + payload = json.loads(profile_path.read_text(encoding="utf-8")) + payload["mappings"][0]["target_reference"] = "missing_grapa_target" + profile_path.write_text(json.dumps(payload), encoding="utf-8") + + with pytest.raises(ValueError, match="names unknown target reference"): + load_country_spec(package_dir) + def test_gates_select_no_incumbent_comparison(self, spec) -> None: selected = {gate.gate for gate in spec.gates.gates} assert "parity" not in selected # no incumbent; #264 remains separate @@ -1699,6 +2244,7 @@ def test_schema2_target_profile_uses_shared_academic_period_semantics( reference = files["target_references.json"]["target_references"][0] reference["period"] = reference_period reference["ledger_selector"]["period_type"] = "academic_year" + reference["ledger_selector"]["period_value"] = reference_period basis = files["target_references.json"]["target_profile"]["basis_periods"][ "population_2023" ] @@ -2122,6 +2668,7 @@ def test_schema2_period_contract_preserves_valid_aliases_and_opaque_labels( reference = files["target_references.json"]["target_references"][0] reference["period"] = reference_period reference["ledger_selector"]["period_type"] = period_type + reference["ledger_selector"]["period_value"] = reference_period basis = files["target_references.json"]["target_profile"]["basis_periods"][ "population_2023" ] @@ -2148,7 +2695,7 @@ def test_schema2_be_geography_vintage_contract_survives_identifier_resolution( raw_reference = next( row for row in payload["target_references"] - if row["name"] == "statbel_fiscal_income_by_commune" + if row["name"] == "statbel_fiscal_income_by_commune_2023" ) raw_reference["metadata"]["activation_status"] = "active" raw_reference["ledger_selector"]["geography_id"] = "21004" @@ -2162,15 +2709,15 @@ def test_schema2_be_geography_vintage_contract_survives_identifier_resolution( reference = next( row for row in loaded.target_references - if row.name == "statbel_fiscal_income_by_commune" + if row.name == "statbel_fiscal_income_by_commune_2023" ) fact = { "aggregate_fact_key": fact_key, "lineage": {"source_record_id": record_id}, "value": 10.0, # Synthetic resolver probe, not a Belgian source value. - "period": {"type": "tax_year", "value": 2022}, + "period": {"type": "tax_year", "value": 2023}, "geography": {"level": "commune", "id": "21004"}, - "entity": {"name": "household"}, + "entity": {"name": "person"}, "observed_measure": { "source_name": "statbel_fiscal_income", "source_measure_id": "taxable_income", diff --git a/packages/microcosm-build/tests/test_spec_country_contract_profiles.py b/packages/microcosm-build/tests/test_spec_country_contract_profiles.py new file mode 100644 index 000000000..8c2908b3d --- /dev/null +++ b/packages/microcosm-build/tests/test_spec_country_contract_profiles.py @@ -0,0 +1,229 @@ +"""Specification tests for value-free population and monetary profiles.""" + +from __future__ import annotations + +import copy +import json +from pathlib import Path + +import pytest + +from microcosm.build.country_spec import load_country_spec +from microcosm.build.trace import sha256_file + + +def _population_profile() -> dict[str, object]: + return { + "schema_version": 1, + "country": "xx", + "profile_id": "benefit_population_inputs", + "activation": "explicit_only", + "description": "Value-free population input fixture.", + "inputs": [ + { + "input_id": "benefit_receipt", + "column": "receives_benefit", + "entity": "person", + "dtype": "bool", + "nullable": True, + "semantic_kind": "receipt", + "data_kind": "latent", + "owner": "Microcosm", + "consumer": "PolicyEngine", + "mechanics_owner": "PolicyEngine", + "axiom_role": "none", + "description": "Measured or latent receipt status, not behavior.", + } + ], + "mappings": [ + { + "mapping_id": "benefit_receipt_mapping", + "target_reference": "benefit_recipients", + "input_id": "benefit_receipt", + "chronicle_source_record_id": "official.benefit.recipients", + "chronicle_entity": "person", + "chronicle_entity_role": "recipient", + "chronicle_geography_level": "statistical_scope", + "chronicle_geography_id": "XX-SCHEME", + "chronicle_geography_vintage": "scheme_scope_2024", + "chronicle_period_type": "calendar_year", + "chronicle_period": 2024, + "microcosm_entity": "person", + "microcosm_geography_level": "statistical_scope", + "microcosm_geography_id": "XX-SCHEME", + "microcosm_geography_vintage": "scheme_scope_2024", + "microcosm_period_type": "calendar_year", + "microcosm_period": 2024, + "publisher_source_readiness": "ready", + "input_readiness": "required_missing", + "mapping_readiness": "required_missing", + "period_readiness": "ready", + "completeness_readiness": "complete_imputation_required_missing", + "notes": "Execution remains blocked until the input exists.", + } + ], + } + + +def _monetary_profile() -> dict[str, object]: + return { + "schema_version": 1, + "country": "xx", + "profile_id": "monetary_targets_2024", + "activation": "explicit_only", + "description": "Value-free monetary target fixture.", + "targets": [ + { + "reference": { + "name": "income/payroll", + "ledger_source_record_id": "official.payroll.2024", + "entity": "person", + "measure": "employment_income", + "period": 2024, + "family": "income", + "metadata": { + "monetary_target_role": "calibration", + "activation_status": "requires_prepared_measure", + "measure_kind": "prepared_column", + }, + }, + "basis": { + "currency": "XXX", + "unit": "base_currency", + "period": "2024", + "temporal_basis": "annual_flow", + "sector": "S14", + "perimeter": "resident employee payroll", + "valuation": "nominal", + }, + "readiness": "requires_prepared_measure", + "source_url": "https://example.test/payroll", + "notes": "Requires a prepared entity-aligned amount column.", + } + ], + } + + +def _write_package(root: Path, resources: dict[str, dict[str, object]]) -> Path: + package = root / "xx" + package.mkdir() + manifest = { + "schema_version": 1, + "country": "xx", + "policy": "spec-only test package", + "resources": list(resources), + } + (package / "country_package.json").write_text( + json.dumps(manifest, sort_keys=True), encoding="utf-8" + ) + for name, payload in resources.items(): + (package / name).write_text( + json.dumps(payload, sort_keys=True), encoding="utf-8" + ) + return package + + +def _profile_resources() -> dict[str, dict[str, object]]: + return { + "population_inputs.json": _population_profile(), + "monetary_target_profile.json": _monetary_profile(), + } + + +def test_profiles_are_optional_when_not_declared(tmp_path: Path) -> None: + package = _write_package(tmp_path, {"evidence.json": {"status": "present"}}) + + spec = load_country_spec(package) + + assert spec.population_input_profile is None + assert spec.monetary_target_profile is None + + +def test_declared_profiles_are_typed_and_content_hashed(tmp_path: Path) -> None: + resources = _profile_resources() + package = _write_package(tmp_path, resources) + + before = load_country_spec(package) + + assert before.population_input_profile is not None + assert before.population_input_profile.profile_id == "benefit_population_inputs" + assert before.population_input_profile.inputs[0].owner == "Microcosm" + assert before.monetary_target_profile is not None + assert before.monetary_target_profile.profile_id == "monetary_targets_2024" + assert before.monetary_target_profile.targets[0].basis.period == "2024" + for name in resources: + assert before.resource_hashes[name] == sha256_file(package / name) + + population_path = package / "population_inputs.json" + resources["population_inputs.json"]["description"] = ( + "Changed value-free population input fixture." + ) + population_path.write_text( + json.dumps(resources["population_inputs.json"], sort_keys=True), + encoding="utf-8", + ) + after = load_country_spec(package) + + assert ( + after.resource_hashes["population_inputs.json"] + != before.resource_hashes["population_inputs.json"] + ) + assert ( + after.resource_hashes["monetary_target_profile.json"] + == before.resource_hashes["monetary_target_profile.json"] + ) + assert after.fingerprint != before.fingerprint + + +@pytest.mark.parametrize( + ("profile_name", "failure"), + [ + ("population_inputs.json", "country"), + ("monetary_target_profile.json", "country"), + ("population_inputs.json", "malformed"), + ("monetary_target_profile.json", "malformed"), + ], +) +def test_declared_profiles_fail_on_country_mismatch_or_malformed_content( + tmp_path: Path, + profile_name: str, + failure: str, +) -> None: + resources = copy.deepcopy(_profile_resources()) + profile = resources[profile_name] + if failure == "country": + profile["country"] = "yy" + elif profile_name == "population_inputs.json": + del profile["activation"] + else: + profile["targets"][0]["readiness"] = "ready" + package = _write_package(tmp_path, resources) + + with pytest.raises(ValueError): + load_country_spec(package) + + +@pytest.mark.parametrize("container", ["ledger_selector", "metadata"]) +def test_monetary_profile_refuses_nested_carried_values( + tmp_path: Path, container: str +) -> None: + resources = copy.deepcopy(_profile_resources()) + reference = resources["monetary_target_profile.json"]["targets"][0]["reference"] + reference.setdefault(container, {})["observed_value"] = 42 + package = _write_package(tmp_path, resources) + + with pytest.raises(ValueError, match="must be value-free"): + load_country_spec(package) + + +def test_monetary_profile_refuses_selector_period_drift(tmp_path: Path) -> None: + resources = copy.deepcopy(_profile_resources()) + reference = resources["monetary_target_profile.json"]["targets"][0]["reference"] + reference["ledger_selector"] = { + "period_type": "calendar_year", + "period_value": 2023, + } + package = _write_package(tmp_path, resources) + + with pytest.raises(ValueError, match="selector period must match"): + load_country_spec(package) diff --git a/packages/microcosm-build/tests/test_spec_engine_country_bundles.py b/packages/microcosm-build/tests/test_spec_engine_country_bundles.py index 736925be2..1d457b2a1 100644 --- a/packages/microcosm-build/tests/test_spec_engine_country_bundles.py +++ b/packages/microcosm-build/tests/test_spec_engine_country_bundles.py @@ -45,7 +45,7 @@ ), ( "be", - "7062e38f4d623553fb0604380a8dac0edacb6261c155b6e31fc38ef7c0f1c57c", + "935f8dedcd2d1b99abe57c1b5d990bc345f4cffbc7ab0cc8791733bde6721d14", { "household.household_id", "person.person_id", @@ -192,6 +192,8 @@ def test_be_generation_zero_views_still_come_from_the_country_spec_seam() -> Non assert {row.path for row in spec.resource_rows if row.kind == "legacy_json"} == { "gates.json", "geography_spine.json", + "monetary_target_profile.json", + "population_inputs.json", "release_contract.json", "source_stages.json", "target_references.json", diff --git a/packages/microcosm-build/tests/test_spec_population_inputs.py b/packages/microcosm-build/tests/test_spec_population_inputs.py new file mode 100644 index 000000000..af6cd472c --- /dev/null +++ b/packages/microcosm-build/tests/test_spec_population_inputs.py @@ -0,0 +1,444 @@ +"""Specification tests for fail-closed population-input and mapping contracts.""" + +from __future__ import annotations + +import copy +import json + +import numpy as np +import pandas as pd +import pytest + +from microcosm.build import ( + PopulationInputContract, + PopulationInputNotReadyError, + PopulationInputProfile, + SchemePopulationMapping, + validate_population_input_frame, +) +from microcosm.frame import EntitySchema, Frame, WeightKind, Weights + + +def _payload() -> dict[str, object]: + return { + "schema_version": 1, + "country": "xx", + "profile_id": "benefit_population_inputs", + "activation": "explicit_only", + "description": "Synthetic value-free population input contract.", + "inputs": [ + { + "input_id": "regular_payment_recipient", + "column": "regular_payment_recipient_indicator", + "entity": "person", + "dtype": "bool", + "nullable": True, + "semantic_kind": "receipt", + "data_kind": "latent", + "owner": "Microcosm", + "consumer": "PolicyEngine", + "mechanics_owner": "PolicyEngine", + "axiom_role": "none", + "description": "Microcosm population status supplied to an adapter.", + } + ], + "mappings": [ + { + "mapping_id": "regular_payment_scheme_population", + "target_reference": "publisher_regular_payment_recipients", + "input_id": "regular_payment_recipient", + "chronicle_source_record_id": ( + "publisher.regular_payment.month2025_01.all.beneficiaries" + ), + "chronicle_entity": "person", + "chronicle_entity_role": "regular_payment_beneficiary", + "chronicle_geography_level": "statistical_scope", + "chronicle_geography_id": "XX-SCHEME", + "chronicle_geography_vintage": "SCHEME_ADMIN_SCOPE_2025", + "chronicle_period_type": "month", + "chronicle_period": "2025-01", + "microcosm_entity": "person", + "microcosm_geography_level": "statistical_scope", + "microcosm_geography_id": "XX-SCHEME", + "microcosm_geography_vintage": "SCHEME_ADMIN_SCOPE_2025", + "microcosm_period_type": "month", + "microcosm_period": "2025-01", + "publisher_source_readiness": "ready", + "input_readiness": "ready", + "mapping_readiness": "ready", + "period_readiness": "ready", + "completeness_readiness": "ready", + "notes": "Exact publisher scheme population and snapshot.", + } + ], + } + + +def _profile(payload: dict[str, object] | None = None) -> PopulationInputProfile: + return PopulationInputProfile.from_mapping(payload or _payload(), country="xx") + + +def _frame(values=(True, False, True), ids=(11, 12, 13)) -> Frame: + person = pd.DataFrame( + { + "person_id": ids, + "person_household_id": [1, 1, 2], + "regular_payment_recipient_indicator": values, + } + ) + household = pd.DataFrame({"household_id": [1, 2]}) + return Frame( + {"person": person, "household": household}, + EntitySchema(group_entities=("household",)), + {"household": Weights(np.array([2.0, 1.0]), WeightKind.DESIGN)}, + ) + + +def test_profile_parses_typed_microcosm_to_policyengine_boundary(): + profile = _profile() + + assert isinstance(profile.inputs[0], PopulationInputContract) + assert isinstance(profile.mappings[0], SchemePopulationMapping) + assert profile.inputs[0].owner == "Microcosm" + assert profile.inputs[0].consumer == "PolicyEngine" + assert profile.inputs[0].mechanics_owner == "PolicyEngine" + assert profile.inputs[0].axiom_role == "none" + assert profile.mappings[0].chronicle_entity_role == ("regular_payment_beneficiary") + assert profile.mappings[0].blockers == () + + +@pytest.mark.parametrize( + ("path", "key", "value"), + [ + ((), "automatic_activation", True), + (("inputs", 0), "formula", "eligible and random_draw"), + (("inputs", 0), "axiom_concept", "takes_up_if_eligible"), + (("mappings", 0), "geography_crosswalk", "statistical_scope_to_nuts1"), + ], +) +def test_every_schema_level_rejects_unknown_bypass_fields(path, key, value): + payload = copy.deepcopy(_payload()) + location = payload + for part in path: + location = location[part] + location[key] = value + + with pytest.raises(ValueError, match="unknown"): + PopulationInputProfile.from_mapping(payload, country="xx") + + +@pytest.mark.parametrize( + ("path", "key"), + [ + ((), "activation"), + (("inputs", 0), "owner"), + (("inputs", 0), "semantic_kind"), + (("mappings", 0), "chronicle_source_record_id"), + (("mappings", 0), "chronicle_entity_role"), + (("mappings", 0), "publisher_source_readiness"), + (("mappings", 0), "mapping_readiness"), + (("mappings", 0), "period_readiness"), + (("mappings", 0), "completeness_readiness"), + ], +) +def test_required_contract_fields_cannot_be_omitted(path, key): + payload = copy.deepcopy(_payload()) + location = payload + for part in path: + location = location[part] + del location[key] + + with pytest.raises(ValueError, match="missing"): + PopulationInputProfile.from_mapping(payload, country="xx") + + +@pytest.mark.parametrize( + ("field", "value", "message"), + [ + ("dtype", "int64", "dtype"), + ("nullable", False, "nullable"), + ("semantic_kind", "take_up", "semantic_kind"), + ("data_kind", "modeled_behavior", "data_kind"), + ("owner", "PolicyEngine", "owner"), + ("consumer", "Microcosm", "consumer"), + ("mechanics_owner", "Axiom", "mechanics_owner"), + ("axiom_role", "behavior_input", "axiom_role"), + ], +) +def test_input_contract_refuses_wrong_types_or_ownership(field, value, message): + payload = copy.deepcopy(_payload()) + payload["inputs"][0][field] = value + + with pytest.raises(ValueError, match=message): + PopulationInputProfile.from_mapping(payload, country="xx") + + +@pytest.mark.parametrize( + "column", + [ + "takes_up_grant_if_eligible", + "grant_take_up", + "grant_takeup_propensity", + "labor_supply_elasticity", + ], +) +def test_behavioral_or_eligibility_variable_synthesis_is_refused(column): + payload = copy.deepcopy(_payload()) + payload["inputs"][0]["column"] = column + + with pytest.raises(ValueError, match="looks behavioral"): + PopulationInputProfile.from_mapping(payload, country="xx") + + +def test_behavioral_input_id_is_refused_before_it_can_be_linked(): + payload = copy.deepcopy(_payload()) + payload["inputs"][0]["input_id"] = "takes_up_grant_if_eligible" + payload["mappings"][0]["input_id"] = "takes_up_grant_if_eligible" + + with pytest.raises(ValueError, match="input_id.*looks behavioral"): + PopulationInputProfile.from_mapping(payload, country="xx") + + +def test_statistical_scope_cannot_masquerade_as_nuts_geography(): + payload = copy.deepcopy(_payload()) + mapping = payload["mappings"][0] + mapping["microcosm_geography_level"] = "nuts1" + mapping["microcosm_geography_id"] = "XX1" + mapping["microcosm_geography_vintage"] = "NUTS_2024" + mapping["mapping_readiness"] = "required_missing" + + with pytest.raises(ValueError, match="statistical_scope"): + PopulationInputProfile.from_mapping(payload, country="xx") + + +@pytest.mark.parametrize( + ("field", "value", "message"), + [ + ("microcosm_entity", "household", "entity/geography identity"), + ("microcosm_geography_id", "XX-RESIDENTS", "entity/geography identity"), + ("microcosm_period", "2025-02", "exact period identity"), + ], +) +def test_ready_mapping_requires_exact_entity_geography_and_period( + field, value, message +): + payload = copy.deepcopy(_payload()) + payload["mappings"][0][field] = value + + with pytest.raises(ValueError, match=message): + PopulationInputProfile.from_mapping(payload, country="xx") + + +@pytest.mark.parametrize( + ("field", "value"), + [ + ("publisher_source_readiness", "dashboard_capture"), + ("input_readiness", "pending"), + ("mapping_readiness", "inferred"), + ("period_readiness", "projected"), + ("completeness_readiness", "assume_false"), + ], +) +def test_unknown_readiness_values_are_refused(field, value): + payload = copy.deepcopy(_payload()) + payload["mappings"][0][field] = value + + with pytest.raises(ValueError, match=field): + PopulationInputProfile.from_mapping(payload, country="xx") + + +def test_native_publisher_artifact_gap_is_an_execution_blocker(): + payload = copy.deepcopy(_payload()) + payload["mappings"][0]["publisher_source_readiness"] = ( + "native_publisher_artifact_required_missing" + ) + profile = PopulationInputProfile.from_mapping(payload, country="xx") + + assert profile.mappings[0].blockers == ( + "publisher_source_readiness='native_publisher_artifact_required_missing'", + ) + with pytest.raises( + PopulationInputNotReadyError, + match="publisher_source_readiness", + ): + profile.mappings[0].require_ready() + + +def test_profile_rejects_orphan_inputs_unknown_links_and_duplicate_columns(): + orphaned = copy.deepcopy(_payload()) + extra_input = copy.deepcopy(orphaned["inputs"][0]) + extra_input["input_id"] = "orphaned_application" + extra_input["column"] = "orphaned_application_indicator" + extra_input["semantic_kind"] = "application" + orphaned["inputs"].append(extra_input) + with pytest.raises(ValueError, match="no scheme-population mapping"): + PopulationInputProfile.from_mapping(orphaned, country="xx") + + unknown = copy.deepcopy(_payload()) + unknown["mappings"][0]["input_id"] = "omitted_input" + with pytest.raises(ValueError, match="unknown input"): + PopulationInputProfile.from_mapping(unknown, country="xx") + + duplicate = copy.deepcopy(_payload()) + second = copy.deepcopy(duplicate["inputs"][0]) + second["input_id"] = "same_column_choice" + second["semantic_kind"] = "choice" + duplicate["inputs"].append(second) + duplicate_mapping = copy.deepcopy(duplicate["mappings"][0]) + duplicate_mapping["mapping_id"] = "same_column_choice_mapping" + duplicate_mapping["input_id"] = "same_column_choice" + duplicate["mappings"].append(duplicate_mapping) + with pytest.raises(ValueError, match="duplicate entity/column"): + PopulationInputProfile.from_mapping(duplicate, country="xx") + + +def test_nonready_contract_fails_before_touching_a_frame(): + payload = copy.deepcopy(_payload()) + payload["mappings"][0]["input_readiness"] = "required_missing" + profile = _profile(payload) + + class FrameAccessTrap: + def table(self, entity): # pragma: no cover - must never run + raise AssertionError(f"Frame was accessed for {entity}") + + with pytest.raises(PopulationInputNotReadyError, match="input_readiness"): + validate_population_input_frame( + FrameAccessTrap(), + profile, + mapping_id="regular_payment_scheme_population", + ) + + +def test_unknown_values_remain_null_until_complete_imputation_is_receipted(): + payload = copy.deepcopy(_payload()) + payload["mappings"][0]["completeness_readiness"] = ( + "complete_imputation_required_missing" + ) + profile = _profile(payload) + frame = _frame(values=pd.array([True, pd.NA, False], dtype="boolean")) + + with pytest.raises( + PopulationInputNotReadyError, + match="completeness_readiness='complete_imputation_required_missing'", + ): + validate_population_input_frame( + frame, + profile, + mapping_id="regular_payment_scheme_population", + ) + + assert frame.table("person")[ + "regular_payment_recipient_indicator" + ].isna().tolist() == [ + False, + True, + False, + ] + + +def test_ready_boolean_column_emits_deterministic_row_value_identity_receipt(): + profile = _profile() + receipt = validate_population_input_frame( + _frame(), + profile, + mapping_id="regular_payment_scheme_population", + ) + repeated = validate_population_input_frame( + _frame(), + profile, + mapping_id="regular_payment_scheme_population", + ) + + assert receipt == repeated + assert receipt["n_rows"] == 3 + assert receipt["n_true"] == 2 + assert receipt["n_false"] == 1 + assert receipt["n_unknown"] == 0 + assert receipt["chronicle_source_record_id"] == ( + "publisher.regular_payment.month2025_01.all.beneficiaries" + ) + assert receipt["chronicle_geography_level"] == "statistical_scope" + for key in ( + "contract_sha256", + "row_ids_sha256", + "values_sha256", + "row_values_sha256", + "receipt_sha256", + ): + assert len(receipt[key]) == 64 + assert "row_ids" not in receipt + assert "values" not in receipt + + changed_values = validate_population_input_frame( + _frame(values=(False, False, True)), + profile, + mapping_id="regular_payment_scheme_population", + ) + assert changed_values["row_ids_sha256"] == receipt["row_ids_sha256"] + assert changed_values["values_sha256"] != receipt["values_sha256"] + assert changed_values["row_values_sha256"] != receipt["row_values_sha256"] + assert changed_values["receipt_sha256"] != receipt["receipt_sha256"] + + changed_ids = validate_population_input_frame( + _frame(ids=(21, 22, 23)), + profile, + mapping_id="regular_payment_scheme_population", + ) + assert changed_ids["row_ids_sha256"] != receipt["row_ids_sha256"] + assert changed_ids["values_sha256"] == receipt["values_sha256"] + assert changed_ids["row_values_sha256"] != receipt["row_values_sha256"] + + +@pytest.mark.parametrize( + ("values", "message"), + [ + ((True, None, False), "missing values"), + ((1, 0, 1), "boolean values"), + (("yes", "no", "yes"), "boolean values"), + ], +) +def test_ready_input_refuses_incomplete_or_proxy_vectors(values, message): + with pytest.raises(ValueError, match=message): + validate_population_input_frame( + _frame(values=values), + _profile(), + mapping_id="regular_payment_scheme_population", + ) + + +def test_ready_input_refuses_missing_column_and_mapping_omission(): + frame = _frame() + missing = Frame( + { + "person": frame.table("person").drop( + columns=["regular_payment_recipient_indicator"] + ), + "household": frame.table("household"), + }, + frame.schema, + {"household": frame.weights_for("household")}, + ) + with pytest.raises(ValueError, match="missing Frame column"): + validate_population_input_frame( + missing, + _profile(), + mapping_id="regular_payment_scheme_population", + ) + with pytest.raises(KeyError, match="Unknown scheme-population mapping"): + validate_population_input_frame( + frame, + _profile(), + mapping_id="undeclared_bypass", + ) + + +def test_receipt_is_json_safe_without_exposing_microdata_rows(): + receipt = validate_population_input_frame( + _frame(ids=(910001, 910002, 910003)), + _profile(), + mapping_id="regular_payment_scheme_population", + ) + rendered = json.dumps(receipt, sort_keys=True, allow_nan=False) + + assert "910001" not in rendered + assert "regular_payment_recipient_indicator" in rendered