diff --git a/COVERAGE_PROGRESS.md b/COVERAGE_PROGRESS.md new file mode 100644 index 00000000..aee92b3b --- /dev/null +++ b/COVERAGE_PROGRESS.md @@ -0,0 +1,51 @@ +# eCPS Coverage Progress + +Continuation checkpoints belong on the native `buildm-ecps-coverage` branch; +the temporary shadow repository used by an earlier linked-worktree sandbox has +been removed. + +- overtime/FLSA | complete (targeted pytest + Ruff green) | b6e58363457497fbdb5f3123539c09ee930cd122 +- SIPP tips | complete (targeted salvage pytest green) | 0c7bbfb691aed03cbfd5c985c40da237d5e2505b +- SCF auto loans | complete (targeted salvage pytest green) | 0c7bbfb691aed03cbfd5c985c40da237d5e2505b +- overtime/FLSA | verified green (104-test mandated set; 70-test family rerun + Ruff) | 68739314c6cc205bf28a3314348755a083eb3d82 +- education inputs | complete (archived published-path semantics verified; post-disaggregation tuition invariant fixed; donor-to-support smoke + mandated/expanded targeted pytest + Ruff green; AOTC abolition probe = $12.28B on pinned eCPS with policyengine-us 1.764.6) | this native checkpoint +- retirement contributions | complete (measured ASEC RETCB_VAL five-leaf split + archived CPS-to-PUF QRF/constraints restored; five post-reference hard requirements + Saver's Credit abolition probe; 127 targeted tests, live policyengine-us 1.764.6 graph check, real-QRF smoke, and Ruff green) | this native checkpoint +- casualty losses | complete (archived IRS PUF E20500 direct carry + post-disaggregation reconciliation + sparse PUF-support QRF restored; promoted to required with OBBBA reactivation probe; 284 targeted tests, 211,677-row real-PUF donor smoke, live policyengine-us 1.764.6 graph/household bind checks, and Ruff green) | this native checkpoint +- miscellaneous itemized | complete (archived IRS PUF E20400 unreimbursed-employee-expense proxy + post-disaggregation reconciliation + PUF-support QRF restored; promoted to required with OBBBA reactivation probe; 303 targeted tests, 211,677-row real-PUF donor smoke, live policyengine-us 1.764.6 graph/household bind checks, and Ruff green) | this native checkpoint +- childcare | complete (archived ASEC SPM_CHILDCAREXPNS direct carry + exact eight-predictor/5,000-row CPS-to-PUF QRF + first-person SPM-unit reduction restored; promoted with OBBBA CDCC reversion probe; 274 targeted tests, 1,000-household real-ASEC/QRF smoke, $16.18M live policyengine-us 1.764.6 probe on the staged artifact, independent review, and Ruff green) | this native checkpoint +- alimony income/expense | complete (archived ASEC OI_OFF=20/OI_VAL recipient split + code-20/code-12 miscellaneous-income exclusion and exact PUF E00800/E03500 carries restored with post-disaggregation reconciliation and sparse PUF support; promoted with ALD-abolition probe; 307 targeted tests, 5,000-household real-ASEC/all-PUF-donor QRF smoke, $55.70M live policyengine-us 1.764.6 probe on the staged artifact, independent review, and Ruff green) | this native checkpoint +- QBI / Section 199A inputs | complete (the pinned processed PUF 1.8.0 `qbi_simulation_version=1` leaves are carried without redrawing; count-preserving source-aligned boolean placement plus qualified-source SSTB, W-2/UBIA, and exposure reconciliation restored; promoted with REIT/PTP and wage/property guardrail probes; 211,677-return donor load, 1,000-household/10,000-donor real-QRF smoke, $5.85M and $185.99M live policyengine-us 1.764.6 probe effects, mandated/expanded targeted tests, independent review, and Ruff green) | this native checkpoint +- domestic production ALD | complete (archived IRS PUF E03240 tax-unit carry, exported field, retired weighted-QRF target, and versioned PUF 1.8.0 artifact verified; post-disaggregation reconciliation plus sparse PUF-only support restored; promoted with former Section 199 reactivation probe; 211,677-return donor audit, 1,000-household/10,000-donor real-QRF smoke with zero ASEC signal, $84.87M live policyengine-us 1.764.6 probe effect, mandated/expanded targeted tests, independent review, and Ruff green) | this native checkpoint +- child support received/expense | complete (archived ASEC CSP_VAL/CHSP_VAL positive annual person-level carries + exact eight-predictor/5,000-row joint weighted QRF on the PUF support half restored; promoted with separate SNAP receipt-source and paid-support-deduction probes; exact ASEC carry and nondefault ASEC/PUF channels verified on a 1,000-household real-data smoke, +$7.01M/+2.16M live policyengine-us 1.764.6 probe effects, 160 integrated targeted tests, independent review, and Ruff green) | this native checkpoint +- disability benefits | complete (archived ASEC two-slot DIS_VAL1/DIS_VAL2 derivation restored with DIS_SC source code 1 excluded as workers' compensation, plus the exact eight-predictor/5,000-row weighted QRF on the PUF support half; promoted with a SNAP unearned-income exclusion probe; full pooled ASEC audit found 0.776% weighted positive signal and $39.87B, the 1,000-household real-data smoke had exact ASEC carry and 0.696% PUF-channel signal, the live policyengine-us 1.764.6 probe scored +$21.18M, and 172 integrated targeted tests, independent review, and Ruff are green) | this native checkpoint +- educator expense | complete (archived IRS PUF E03220 direct carry, filer/spouse earnings-share allocation, exported field, and retired PUF-QRF treatment verified; the processed PUF 1.8.0 donor is carried through the shared sparse weighted QRF on Populace's dedicated PUF support channel and allocated by employment-income share, with ASEC explicitly source-zero; promoted with educator-ALD abolition probe; 211,677-tax-unit donor audit found 2.483% weighted positive signal and $1.379B, the 1,000-household real-data smoke had 2.130% PUF-channel signal, the live policyengine-us 1.764.6 probe scored +$1.056M, and 231 integrated targeted tests, independent review, and Ruff are green) | this native checkpoint +- health insurance premiums | complete (archived ASEC PHIP_VAL residual after CHIP, Marketplace, and Medicaid premiums plus the exact eight-predictor/5,000-row joint PUF-support QRF restored for `other_health_insurance_premiums`; promoted without a synthetic reform probe because no parameterized formula uniquely reads the leaf, with a live policyengine-us 1.764.6 graph regression instead; `employer_sponsored_insurance_premiums` remains a reviewed SOURCE UNAVAILABILITY WITH EVIDENCE exclusion because the retired derivation requires NOW_OWNGRP, NOW_HIPAID, and NOW_GRPFTYP, which all three SHA-recorded 2022-2024 hermetic HDF inputs omit; the 1,000-household real-data smoke has 39.262%/40.232% ASEC/PUF positive shares, batched and unbatched outputs match exactly and persist green, and the 176-test targeted set, independent review, and Ruff are green) | this native checkpoint +- farm operations/rent income | complete (corrected archived ASEC `FRSE_VAL` carry to signed `farm_operations_income`, kept the distinct PUF `T27800` Schedule J `farm_income`, and restored exact signed PUF `E02100` operations plus `E27200` rent carries after aggregate disaggregation and through weighted QRF support; promoted both with independent 2026 Section 199A probes scoring -$4.16M and +$9.14M baseline-minus-reform QBID; the persisted 1,000-household smoke retains positive/loss observations on every source-backed channel, keeps ASEC rent exactly zero, and passes the production gate after reload; 331 integrated targeted tests, independent review, and Ruff are green) | this native checkpoint +- financial assistance | reviewed exclusion — SOURCE UNAVAILABILITY WITH EVIDENCE (archived `cps.py` lines 1493-1496 requires person `FIN_VAL`, the archived loader makes it required, and the CPS-only PUF-clone QRF has no independent PUF target; all three SHA-locked 2022-2024 hermetic H5s omit `FIN_VAL`, `FIN_YN`, and `I_FINVAL`, while their `HFINVAL`/`FFINVAL` fields are only household/family aggregates; artifact-backed regressions rehash every file, prove 374/393/440 positive multi-person families, and prove household amounts equal summed family amounts, so choosing a recipient would synthesize data; the generated manifest retains the exclusion and 110 targeted tests, independent review, and Ruff are green) | this native checkpoint +- SIPP household vehicles | complete (archived December 2023 SIPP household transform, target-specific allocation masks, weighted count classifier, and count-conditioned value QRF restored from the immutable 3.73 GB donor revision `21280dca5995e978d706740a8a4b9b7860cfd7b6` / SHA-256 `5c30439e365fc26483318ef61d1d8f4bb2f0e9d6bb47c22c06756a7698733ee2`; the full donor yields 16,841 household support rows and the persisted 30,000-household Populace smoke carries 84.1%/84.4% count/value signal with no invalid values; promoted both inputs with a Texas SNAP additional-vehicle-exemption abolition probe scoring +$3.58M in 2026, while deliberately avoiding the retired pipeline's later full-balance-sheet `net_worth` fold; 249 integrated targeted tests, independent review, and Ruff are green) | this native checkpoint +- household weight | complete (archived commit `42ed5d45c56df80d754fbe24cce21cfeb8d05cbe` derives ASEC `HSUP_WGT / 100` at `cps.py` lines 1052-1075 and PUF `S006 / 100` at `puf.py` lines 636-638 and 789-799, after which `enhanced_cps.py` lines 796-864 persists optimized weights; Populace already carries the same source/calibrated values as the authoritative typed household `Weights` vector and its PolicyEngine adapter materializes that vector as `household_weight`, so the release gate now recognizes the export value while overriding any stale table copy; the real 337,704-household base has all-positive weights, 126,532 unique values, range 19.09-2,313.09, total 134.69M, no duplicate table column after load, and passes the promoted required-input gate; 134 targeted coverage/validation/plan/adapter tests, independent review, and Ruff are green) | this native checkpoint +- Form 4952 elected investment income | complete (archived commit `42ed5d45c56df80d754fbe24cce21cfeb8d05cbe` maps IRS PUF `E58990` directly to `investment_income_elected_form_4952` at `puf.py` line 708, exports it at lines 804-850, and includes it in both retired weighted-QRF target/override lists at `calibration/puf_impute.py` lines 90-198; the processed person array embodies the retired randomized `EARNSPLIT` allocation at `puf.py` lines 477-546 and 1513-1601, but the hermetic ASEC target lacks `EARNSPLIT`, so Populace aggregates the source-backed donor leaf to the tax-unit quantity PolicyEngine-US actually reads and uses explicit first-person within-unit placement rather than inventing an earnings split; the leaf is re-derived after aggregate-record disaggregation, carried only on the PUF support channel, and sparsified to the donor positive rate; the SHA-pinned PUF 1.8.0 artifact `7669f5b5281f20080e77204f9bd4aabfad0aa101fa283e22caf9ba8d61d4d6df` has 6,313 positive people / 3,688 positive tax units, a 0.139498% weighted positive rate, and $5.766B weighted source mass; a persisted 10,000-household real-QRF smoke retained 20 positive PUF-channel records at 0.074782%, $106.72M weighted mass, and a uniquely binding Form 4952 neutralization probe scored +$10.90M baseline-minus-reform 2024 income tax; promoted to required with 353 expanded targeted tests, independent review, and Ruff green) | this native checkpoint +- capital-gain detail inputs | complete (archived IRS PUF direct carries `long_term_capital_gains_on_collectibles = E24518` and `unrecaptured_section_1250_gain = E24515` at commit `42ed5d45c56df80d754fbe24cce21cfeb8d05cbe`, `puf.py` lines 654-702, plus export and retired weighted-QRF treatment at `puf.py` lines 804-850 and `calibration/puf_impute.py` lines 90-198 restored; processed PUF 1.8.0 SHA `7669f5b5281f20080e77204f9bd4aabfad0aa101fa283e22caf9ba8d61d4d6df` has 1,695 collectible-gain tax units / 0.044970% weighted support / $8.990B and 11,717 unrecaptured-1250 tax units / 0.790556% / $39.931B; a persisted 9,000-household real-QRF smoke retained seven and 83 PUF-channel records with zero ASEC leakage, and live neutralization probes scored +$2.68M and +$3.25M baseline-minus-reform income tax; `investment_interest_expense` is retained only as a reviewed SOURCE UNAVAILABILITY WITH EVIDENCE exclusion because archived `puf.py` line 654 and lines 1548-1550 make total and mortgage-deductible interest source-identical, `mortgage_interest.py` lines 268-320 derives only their residual, the retired positive signal is stochastic QRF divergence, and every SHA-locked hermetic source omits the two intermediates/raw E19200 while PUF 1.8.0 carries 499,045 all-zero leaf values; promoted both source-backed leaves with 350 targeted integration tests, independent review, and Ruff green) | this native checkpoint +- ASEC relationship inputs | complete (archived `cps.py` exact carries `is_household_head = P_SEQ == 1`, `is_separated = A_MARITL == 6`, and `is_surviving_spouse = A_MARITL == 4` restored through a strict manifest stage; the 865,046-person Build-J base has 41.3221%/1.4010%/4.7531% weighted signal and exactly one head in every 337,704 frame household, and the live household-head neutralization probe binds at -$265.50M baseline-minus-reform capped childcare expenses; `is_unmarried_partner_of_household_head` is retained only as a reviewed SOURCE UNAVAILABILITY WITH EVIDENCE exclusion because archived `cps.py` lines 1214-1221 requires `PERRP`, all three SHA-locked inputs omit it, 2022-2023 also omit `PECOHAB`/`A_EXPRRP`/`A_FAMREL`, and assigning the 2024-only partner pointer or combined partner/roommate code across the equal-third pooled years would synthesize statuses; 177 targeted tests, independent review, and Ruff green) | this native checkpoint +- WIC nutritional risk | reviewed exclusion — SOURCE UNAVAILABILITY WITH EVIDENCE (archived `cps.py` lines 684-701 generates the leaf by OR-ing reported receipt with a category-specific Bernoulli draw from fixed early-1980s priors rather than carrying an assessment; archived `census_cps.py` lines 293-365 and `cps.py` line 1558 expose only WIC value/receipt and map `WICYN == 1` to receipt; all three SHA-locked 2022-2024 inputs rehash and contain only person `SPM_WICVAL`/`WICYN` plus household `HRNUMWIC`/`HRWICYN`, with no nutritional-risk field; the Census questionnaire lets an adult woman report receipt for herself or on behalf of a child and most rows are code-0 not-in-universe, so WICYN does not identify the assessed person; USDA certification requires a competent professional assessment in a distinct administrative-record system, and FNS applied an all-category 1.0 risk adjustment to its revised CY2016-2021 estimates; assigning risk to the WICYN row, setting nonrecipients false, or replaying the retired draw would synthesize unavailable person assessments, and all-true is the 1.764.6 default; the generated manifest preserves the exclusion, 105 targeted tests, independent patch review, two source audits, and Ruff are green) | this native checkpoint +- retirement distributions | complete (archived `cps.py` lines 1448-1481 exact four-slot `DST_SC*`/`DST_VAL*` mappings restored for taxable 401(k), taxable 403(b), tax-exempt Roth IRA, taxable regular IRA, Keogh, and taxable SEP; the retired `extended_cps.py` lines 140-148, 639-745, and 1014-1073 PUF-support QRF is limited to the four populated non-IRA leaves, preserving PUF E01400 taxable-IRA ownership and copied Roth-IRA source values; the real 865,046-person Build-J stage passes exact ASEC reconciliation and all signal bands with $148.97M authoritative weighted Keogh source, and its uniquely binding neutralization probe scores +$24.47M 2024 federal income tax against a $1M floor; generated coverage is 140 required / 25 reviewed exclusions / 25 probes, and 204 expanded targeted tests, independent review, real-QRF smoke, Ruff, and diff-check are green) | this native checkpoint +- retirement distributions support-base follow-up | complete (independent audit found the reusable PUF-support builder omitted the already-restored family even though the final fiscal builder carried it; distributions now run before cloning and after PUF imputation, the post-PUF signal gate fails closed, and summary telemetry records the result; the four-slot mapping regression now exercises both material `DST_*_YNG` pairs; 18 support-builder tests, 24 family tests, seven changed-node regressions, Ruff, and diff-check are green) | this native checkpoint +- SCF household net worth | complete (retained the salvaged direct-anchor approach after verifying archived commit `42ed5d45`: `utils/asset_imputation.py` lines 15-19 maps summary-extract `networth` to `scf_net_worth`, `calibration/source_impute.py` lines 1146-1164 includes it in the exact eight-predictor weighted QRF, and lines 1324-1338 force construction-only components to sum back to that signed anchor before export; the SHA-pinned 22,975-row donor preserves positive and indebted households, and a 1,000-household real-ASEC/20-tree QRF smoke produced 816 unique values from -$142,900 to $46.557M, 99.952% weighted nonzero / 2.987% negative signal, and passed the combined SCF wealth gate; promoted `net_worth` to required with household L0 nonconstant enforcement and a PolicyEngine-US 1.764.6 input-contract regression, but no invented reform probe because no 1.764.6 formula consumes this standalone input leaf; 263 expanded targeted tests, Ruff, format-check, and diff-check are green) | this native checkpoint +- housing and tenure inputs | complete (archived ASEC exact carries `H_TENURE`, `SPM_CAPHOUSESUB`, and `SPM_TENMORTSTATUS` plus the ACS 2022 ten-predictor rent QRF and PUF-half housing-assistance QRF restored; the SHA-pinned 1,505,108-head ACS donor replays the retired joint rent/real-estate-tax sampler exactly at 10,000 shared rows / 9,905 rent rows / 9,469 tax rows, retaining 1,346 zero-weight rent heads while Populace deliberately applies design weights; the persisted 6,000-household/12,000-support-channel real-data smoke passes before and after H5 reload with 10.635% positive rent, 4.772% ASEC and 3.416% PUF housing-assistance shares, all tenure categories, and a live PolicyEngine-US 1.764.6 rent-neutralization SNAP effect of +$189.58M; promoted all four leaves with a PUF-channel-specific fail-closed gate and stale-base cache rejection; 274 expanded targeted tests, independent adversarial review, Ruff, Bash syntax, format, and diff checks are green) | this native checkpoint +- prior-year income | complete (archived adjacent-ASEC `PERIDNUM` join with allocation flags, pairwise sentinels, oldest-cohort defaults, current-value fallback, signed losses, and the exact eight-predictor/5,000-row joint PUF-support QRF restored; promoted `self_employment_income_last_year` and `previous_year_income_available` with a unique structural neutralization probe; the full SHA-locked 432,523-person audit found 16.9147% weighted availability, 2.4896% self-employment signal, 196 loss records, and $357.24B weighted net source amount; a 1,000-household-per-year real-QRF smoke retained both support channels with zero ASEC reconciliation mismatches and a +$2.703B live PolicyEngine-US 1.764.6 probe effect; selective stale-cache validation and source-aware rederivation close independent-review findings; 162 integrated targeted tests, Ruff, format, Bash syntax, and diff checks are green) | this native checkpoint +- SALT refund income | complete (archived IRS PUF `E00700` direct carry, exported leaf, retired target/override QRF treatment, post-disaggregation reconciliation, and PUF-only sparse support restored without recreating the unavailable randomized `EARNSPLIT`; promoted `salt_refund_income` with a uniquely binding Idaho/West Virginia/South Carolina state-tax neutralization probe; the SHA-pinned 211,677-return PUF has 56,584 positive tax units, 13.5799% weighted support, and $45.469B weighted source mass; a persisted 3,024-household target-state ASEC/10,000-donor weighted-QRF smoke retained 530 positive people / 213 unique values after H5 reload, kept ASEC exactly zero, passed the signal gate at 3.6807% overall / 7.3614% PUF-channel support, and scored -$44.282M baseline-minus-neutralized 2024 state income tax against a $1M floor; 275 integrated targeted tests, independent adversarial review, Ruff, format, and diff checks are green) | this native checkpoint +- SPM unit energy subsidy | complete (archived commit `42ed5d45c56df80d754fbe24cce21cfeb8d05cbe` ingests `ENGVAL` as replicated `SPM_ENGVAL` at `census_cps.py` lines 39-58, 77-181, and 240-305, maps it exactly to `spm_unit_energy_subsidy` at `cps.py` lines 1599-1622, and applies the retired eight-predictor/5,000-row CPS-only QRF plus first-person SPM-unit reduction to the PUF clone half at `extended_cps.py` lines 135-194, 234-248, 639-739, and 1014-1073; all three SHA-locked 2022-2024 ASEC artifacts carry finite, nonnegative, internally consistent replicas, and the full production-pooled 176,039-unit source has 6,033 positive units, 3.2951% weighted support, and $2.902B weighted mass; a persisted real-data/100-tree smoke retained 950 positive values / 154 unique values across 31,356 SPM units after H5 reload, passed channel gates at 3.2138% ASEC / 2.5046% PUF support with $221.125M weighted mass, and its uniquely binding 2024 `spm_unit_benefits` neutralization scored +$221.125M against a $100M floor; promoted to required with 215 integrated targeted tests, independent adversarial review, deterministic manifest regeneration, Ruff, format, JSON, and diff checks green) | this native checkpoint +- survivor benefits | reviewed exclusion — SOURCE UNAVAILABILITY WITH EVIDENCE (archived `cps.py` line 1495 requires person `SRVS_VAL`, the archived loader makes it required, and the CPS-only PUF-clone QRF has no independent donor target; all three SHA-locked 2022-2024 hermetic H5 person tables omit `SRVS_VAL`; their `FSURVAL`/`HSURVAL` values are family/household aggregates that reconcile exactly, while 657/647/610 positive families contain multiple people, so assigning a recipient would synthesize data; artifact-backed regressions rehash every file and prove the aggregate identity, the generated manifest retains the exclusion, and 108 targeted tests, independent archive review, Ruff, format, JSON, and diff checks are green) | this native checkpoint +- DC property-tax-credit take-up | reviewed exclusion — SOURCE UNAVAILABILITY WITH EVIDENCE (archived `cps.py` lines 565-599 generates `takes_up_dc_ptc` by a seeded Bernoulli draw; archived `dc_ptc.yaml` lines 1-11 divides 37,133 claims by a 131,791,388 PolicyEngine value estimate, which is dimensionally invalid, arithmetically not 0.32, and not an external eligible-population denominator; every SHA-locked ASEC/PUF input lacks claimant identity, the ASEC state-tax amounts are generic liabilities, and the PUF maps federal `P08000` to `other_credits` while carrying no state geography; replaying or count-allocating the draw would synthesize tax-unit claimants; six artifact-backed family tests, 109 mandated targeted tests, independent adversarial review, Ruff, format, JSON, and diff checks are green) | this native checkpoint +- Medicare take-up input | complete (archived `cps.py` lines 1579-1585 maps measured ASEC `MCARE == 1` to `medicare_enrolled`, `puf_impute.py` lines 608-629 copies non-imputed CPS values onto the PUF support half, and `extended_cps.py` lines 1747-1754 exports the value as `takes_up_medicare_if_eligible`; all three SHA-locked 2022-2024 sources contain exact nondefault signal at 18.6172%/18.7910%/19.0830% weighted enrollment; the persisted 337,704-household Build-J artifact re-derives 158,818 positive people at 19.5342% on both ASEC and PUF-support channels with zero source mismatches, and a deterministic 10,000-household live PolicyEngine-US 1.764.6 neutralization smoke removes $4.910B of 2024 Medicare cost against a $1B floor; promoted to required with stale degenerate/cache exemptions removed; 123 mandated tests, 96 fiscal-builder tests, 25 support/cache tests, Ruff, format, Bash syntax, and diff checks are green) | this native checkpoint +- Head Start / Early Head Start take-up | OPEN — not restored and not accepted as a terminal exclusion (archived commit `42ed5d45c56df80d754fbe24cce21cfeb8d05cbe` `datasets/cps/cps.py` lines 582-583 and 640-648 loads scalar rates and makes unanchored person-level Bernoulli draws; `parameters/take_up/head_start.yaml` lines 1-10 supplies 0.40 for 2020 and 0.30 for 2021+, while `parameters/take_up/early_head_start.yaml` lines 1-9 supplies 0.09 for 2020+, all citing NIEER; PolicyEngine-US 1.764.6 consumes the flags in `head_start.py` and `early_head_start.py` lines 10 and 13-22 over distinct model-specific age/pregnancy plus AGI-or-categorical-benefit eligibility universes; ACF FY2022 Program Facts publishes exact funded slots but only one combined 801,859 cumulative-enrollment count with rounded age shares, full PIR microdata require request/authenticated HSES access, and HHS/ASPE's published 63-per-100 measure covers only income-poor 3-4-year-olds and explicitly excludes Early Head Start; OPEN QUESTION: locate an ACF/HHS participation or slot-coverage rate, or an exact ACF numerator and independently published Census denominator, matching each 1.764.6 eligibility universe; reusing NIEER, applying the ASPE rate, dividing slots by Populace-modeled eligibles, or allocating slots to records would invent or scope-mismatch the rate) | research checkpoint +- housing-assistance take-up | complete (archived commit `42ed5d45c56df80d754fbe24cce21cfeb8d05cbe` `datasets/cps/cps.py` lines 585 and 664-682 preserves reported recipients and then fills eligible nonreporters with a scalar draw, but `parameters/take_up/housing_assistance.yaml` lines 1-15 explicitly says its 0.50 value is an ad hoc eCPS benchmark adjustment pending fuller calibration; archived `db/etl_housing_assistance.py` lines 35-169 instead reads a live HUD workbook into household-grain state counts, which cannot identify SPM-unit recipients and is not pinned in the hermetic build; Populace therefore restores the non-synthetic source surface by keeping `takes_up_housing_assistance_if_eligible` exactly equal to the existing `SPM_CAPHOUSESUB` receipt anchor before cloning and to the receipt QRF result on the PUF-support half, never removing a reporter or allocating an unobserved recipient; a production-ingredient smoke using 6,000 households from the three SHA-locked ASEC files, the hash-checked ACS donor, and pinned congressional-district/block ladder passed all five housing gates at 4.7724% weighted receipt/take-up with zero missing, invalid, or mismatched SPM units, and the live PolicyEngine-US 1.764.6 neutralization removed $202.795M of 2024 HUD assistance against a $100M floor; promoted to required with stale-cache reuse rejected, yielding 151 required / 14 exclusions / 31 probes; 138 mandated/inventory tests, 119 builder tests, 43 L0 export tests, Ruff, format, Bash syntax, JSON, and diff checks are green) | this native checkpoint +- SSI take-up | OPEN dependency — not restored and not accepted as a terminal exclusion (archived commit `42ed5d45c56df80d754fbe24cce21cfeb8d05cbe` `datasets/cps/cps.py` lines 1497-1499 anchors reporters on `SSI_VAL`, `datasets/cps/takeup.py` lines 10-35 preserves those reporters, and `cps.py` lines 650-657 fills additional people; `parameters/take_up/ssi.yaml` lines 1-9 applies 0.50 to all people even though its Urban citation covers adults 65+ only; the SHA-locked ASEC inputs contain real SSI_VAL signal, including 2,252 positive 2024 rows representing about 5.394 million weighted reporters, and SSA Statistical Supplement table 7.A1 publishes 2024 federal-payment recipient counts of 1,001,922 under 18, 3,905,779 ages 18-64, and 2,382,142 ages 65+; however PolicyEngine-US 1.764.6 eligibility requires the separate default-false input `meets_ssi_disability_criteria`, which Populace does not export, while archived `cps.py` lines 2853-2886 restores that prerequisite from a SIPP model; current modeled federal SSI eligibility is zero below age 65 before that restoration and remains below the 7,289,843 SSA target in the latest cached post-eligibility stage; OPEN DEPENDENCY: faithfully restore the archived SIPP disability criterion, then preserve SSI_VAL reporters and perform source-identity, age-band-specific weighted count calibration among people with `uncapped_ssi > 0`; using broad `is_disabled`, reusing the aged-only 0.50 rate, or calibrating only the aggregate would scope-mismatch the source and conceal missing child/disabled eligibility) | research checkpoint +- weeks unemployed | OPEN artifact gap — not restored and not accepted as an irreducible exclusion (archived commit `42ed5d45c56df80d754fbe24cce21cfeb8d05cbe` `datasets/cps/cps.py` lines 1438-1442 maps `weeks_unemployed` exclusively from person `LKWEEKS`, `datasets/cps/census_cps.py` lines 39-58 and 306-365 requires that field, and `calibration/puf_impute.py` lines 703-785 only QRF-imputes the PUF clone from CPS observations; the SHA-locked 2022 H5 `7ccca976...474e` omits LKWEEKS and every audited alias, while locked 2023/2024 contain 4,411/4,556 positive rows at 3.2283%/3.4003% weighted support; neither `UC_VAL` nor `52 - WKSWORK` reproduces LKWEEKS, and filling or QRF-imputing the missing pooled third would synthesize values; Census's official 2023 ASEC dictionary nevertheless documents LKWEEKS and archived `census_cps.py` line 245 maps income year 2022 to `asecpub23csv.zip`, whose 146,133-person count matches the incomplete locked H5, so the upstream source exists and this gap is reducible; OPEN DEPENDENCY: regenerate/re-pin the 2022 H5 or build a keyed SHA-pinned LKWEEKS sidecar from the official archive, then restore the direct carry and PUF-only QRF; no honest release probe exists yet because PolicyEngine-US 1.764.6 Pennsylvania UC also requires four default-zero wage/credit-week inputs absent from the reference) | research checkpoint +- workers' compensation | complete (archived commit `42ed5d45c56df80d754fbe24cce21cfeb8d05cbe` `datasets/cps/cps.py` lines 1559-1571 carries annual person `WC_VAL` directly while keeping code-1 `DIS_VAL*` slots out of the separate disability-benefits leaf, and `datasets/cps/extended_cps.py` lines 135-194, 234-248, 639-745, and 1014-1073 replaces the CPS-only output on the PUF support half with the retired eight-predictor/at-most-5,000-person QRF; all three SHA-locked 2022-2024 ASEC artifacts have finite nonnegative source signal at 0.2174%/0.2614%/0.2817% weighted support and $7.786B/$10.851B/$9.146B weighted mass; a production-ingredient 30,000-household/76,694-person real-QRF smoke passed exact ASEC reconciliation and both channel gates at 0.2048% ASEC / 0.0425% PUF weighted support with $459.024M combined mass, and its uniquely binding SNAP unearned-income exclusion probe scored +$28.261M reform-minus-baseline against a $10M floor; promoted to required, yielding 152 required / 13 reviewed exclusions / 32 probes; 122 mandated coverage/plan tests, 164 builder/export regressions, Ruff, format, JSON, and diff checks are green) | this native checkpoint +- WIC claim input | complete (archived commit `42ed5d45c56df80d754fbe24cce21cfeb8d05cbe` `datasets/cps/cps.py` lines 684-691, `parameters/take_up/wic_takeup.yaml` lines 1-33, and `utils/randomness.py` lines 5-28 derives `would_claim_wic` with stable category-rate draws; the archived rates exactly match USDA FNS CY2022 average-month coverage of 45.6% pregnant, 68.9% all postpartum, 66.3% breastfeeding, 78.4% infant, and 46.0% child; Populace applies the stage after eligibility inputs and pregnancy, correcting the retired file's pre-pregnancy ordering, copies source-identity outcomes across support clones, and uses the all-postpartum rate without inventing the unavailable breastfeeding status or treating adult-reporter `WICYN` as a participant anchor; a production 6,000-household/14,951-person smoke produced 514 positive people at 3.3005% weighted support, with 46.12%/74.19%/81.99%/43.87% realized pregnant/postpartum/infant/child claim shares, and its uniquely binding neutralization removed $57.192M of 2024 WIC benefits against a $25M floor; promoted only `would_claim_wic` while retaining the evidenced `is_wic_at_nutritional_risk` exclusion, yielding 153 required / 12 reviewed exclusions / 33 probes; 130 mandated coverage/plan tests, 192 builder/export/family regressions, independent integration audit, Ruff, format, JSON, and diff checks are green) | this native checkpoint +- voluntary tax filing | complete (archived commit `42ed5d45c56df80d754fbe24cce21cfeb8d05cbe` `datasets/cps/cps.py` lines 726-747 and `parameters/take_up/voluntary_filing.yaml` lines 1-43 derives `would_file_taxes_voluntarily` from an uncited child-by-wage-by-age probability table; Populace replaces that synthetic draw with measured December filing/expected-filing responses from the immutable full 2023 SIPP donor revision `21280dca5995e978d706740a8a4b9b7860cfd7b6`, SHA-256 `5c30439e365fc26483318ef61d1d8f4bb2f0e9d6bb47c22c06756a7698733ee2`, excluding 526 respondents claimed as dependents, collapsing reciprocal spouses, using the complete canonical unit's minimum-PNUM reference person/weight before response filtering, and carrying household under-18 count only as measured context rather than inventing dependent attachments; the source audit has 39,513 December people, 30,510 observed responses / 24,473 filing yes, 22,313 pre-weight canonical units / 16,820 yes, and 22,296 positive-weight units / 16,817 yes at a 76.0308456% weighted rate with zero spouse disagreements; one seeded 100-tree QRF prediction per `tax_unit_source_id` fans out identically to support clones; a production 6,000-source-household / 12,000-support-household Build-J smoke passed both filing channels at 81.9282% with zero clone mismatches, passed the ACA source gate, and its uniquely binding neutralization removed $197.082M of 2024 ACA PTC against a $100M floor; promoted to required, yielding 154 required / 11 reviewed exclusions / 34 probes; 122 mandated family/coverage/validation/plan tests, 140 builder/L0 regressions, full pinned-donor audit, live PolicyEngine-US 1.764.6 fixture, independent review, Ruff, format, JSON, and diff checks are green) | this native checkpoint +- SSI disability criterion | complete (archived commit `42ed5d45c56df80d754fbe24cce21cfeb8d05cbe` `datasets/sipp/sipp.py` lines 63-105, 404-428, 541-562, 565-630, 633-703, 724-795, and 892-949 defines the exact 77-column December SIPP contract, observed-label/allocation filters, 2024 SSI financial screen, nineteen predictors, postprediction disability-signal requirement, fixed WPFINWGT bootstrap, and unweighted 100-tree boolean QRF; `datasets/cps/cps.py` lines 2853-2886, `calibration/source_impute.py` lines 869-990, and `datasets/cps/extended_cps.py` lines 140-170, 392-424, and 981-992 separately predict ASEC and PUF-support people from fresh copies of the fitted model while preserving only direct under-65 ASEC SSI reporters; the immutable shared 2023 SIPP donor revision `21280dca5995e978d706740a8a4b9b7860cfd7b6`, SHA-256 `5c30439e365fc26483318ef61d1d8f4bb2f0e9d6bb47c22c06756a7698733ee2`, yields exactly 39,513 December rows and 9,346 candidates (577 positive / 8,769 negative), weighted share 5.56674699%, then the archived seed `8386123572872638692` resamples 9,346 rows / 5,314 unique / 524 positive before fixed forest seed 42; the production 6,000-source-household / 12,000-support-household smoke passes at 7.05424% weighted true overall, 5.68679% ASEC, and 8.42168% PUF with zero reporter-anchor loss and 785 source-person clone divergences, and the shipped neutralization gate scores +$586.393M 2024 SSI against a $100M floor; promoted the final post-reference export to required, bumped the target-frame checkpoint materializer, and generated 155 required / 11 reviewed exclusions / 35 probes; 262 mandated family/coverage/validation/plan/builder/L0 tests, pinned-donor audit, live PolicyEngine-US 1.764.6 fixture, independent adversarial review, Ruff, format, JSON, derived-manifest, and diff checks are green) | this native checkpoint +- SSI take-up | complete (archived commit `42ed5d45c56df80d754fbe24cce21cfeb8d05cbe` `datasets/cps/cps.py` lines 584, 650-657, and 1497-1499 plus `datasets/cps/takeup.py` lines 10-35 preserve direct ASEC `SSI_VAL` reporters and fill/export `takes_up_ssi_if_eligible`; the retired `parameters/take_up/ssi.yaml` 0.50 all-age rate is rejected because its cited 0.49 estimate covers adults 65+ only; archived `utils/ssi_targets.py` lines 41-74 instead identifies SSA SSI Monthly Statistics December 2024 Table 1, whose `Total with—Federal payment` row supplies exact targets of 1,001,922 under 18, 3,905,779 ages 18-64, and 2,382,142 ages 65+; Populace makes one deterministic decision per `person_source_id` inside `uncapped_ssi > 0`, preserves all 7,087 pre-L0 direct-reporter lineages even when only their PUF clone survives, and records rather than conceals an unreachable band; on the production sparse weights the child band saturates at 131,822.264 with an 870,099.736 evidenced modeled-capacity shortfall, while working-age and aged selected counts land at 3,910,990.716 and 2,391,621.337, each within one source-identity weight of target, with 82.4144%/83.4278% nonconstant ASEC/PUF flag shares and zero anchor/clone violations; dependency-aware bounded reconciliation now fixes SSI before replaying ACA, Medicaid, private-premium inputs, target materialization, and ordinary refit, then diagnoses final weights without post-optimization mutation; the persisted sparse artifact SHA-256 `9269360d3409fdc15c90c43dda394ada6c91eff5bb64c12ccd9def7d670dd077` scores +$57.115B 2024 SSI under the shipped neutralization probe against a $10B floor; promoted to required and default-publication diagnostics, yielding 156 required / 10 reviewed exclusions / 36 probes; 372 expanded family/coverage/validation/plan/contract/parity/Medicaid/builder/L0/publication tests, live PolicyEngine-US 1.764.6 fixture, production-ingredient smoke, independent adversarial review, Ruff, format, strict JSON, deterministic manifest regeneration, and diff checks are green) | this native checkpoint +- Head Start / Early Head Start take-up | complete restoration plus irreducible reviewed exclusion (archived commit `42ed5d45c56df80d754fbe24cce21cfeb8d05cbe` `datasets/cps/cps.py` lines 582-583 and 640-648, `parameters/take_up/head_start.yaml` lines 1-10, `parameters/take_up/early_head_start.yaml` lines 1-9, and `utils/randomness.py` lines 5-28 prove the retired exports were scalar-rate random draws; standard `takes_up_head_start_if_eligible` is restored from the immutable full 2023 SIPP donor revision `21280dca5995e978d706740a8a4b9b7860cfd7b6`, SHA-256 `5c30439e365fc26483318ef61d1d8f4bb2f0e9d6bb47c22c06756a7698733ee2`, using December age-3--5 direct `EEDHEADST` responses and strict upstream-reported structural negatives while excluding hot-decked target answers and imputed screeners; the instrument's broader federally sponsored-preschool wording and Head Start/Even Start/Fair Start examples are recorded explicitly, so the 785-row donor with 45 positive / 740 negative labels and 6.16620218% weighted signal is treated as a measured proxy rather than an exact program-only indicator; one seeded weighted-QRF decision per `person_source_id` fans identically to both support clones and leaves other ages false; the persisted production sparse artifact SHA-256 `67ad74b9ad9222ed342a0279dfc8175e872966fa59f86aeecb7fad52021ba500` has 326 positive rows, 3.833731% weighted age-domain take-up, and zero clone mismatches, and its live PolicyEngine-US 1.764.6 neutralization removes $4.575B of 2024 Head Start against a $100M floor; Early Head Start remains a SOURCE UNAVAILABILITY WITH EVIDENCE exclusion because all three SHA-locked ASEC files and the processed PUF omit person enrollment, SIPP has zero direct positives under age 3, its ambiguous child-care fields begin at age 2 and have no EHS distinction or pregnancy branch, and ACF cumulative enrollment/funded slots cannot identify point-in-time people; generated coverage is 157 required / 9 reviewed exclusions / 37 probes; 213 targeted family/exclusion/coverage/validation/plan/L0/parity/contract/builder/checkpoint tests pass with one intentional duplicate full-SIPP exclusion audit left opt-in, and pinned-source audits, Ruff, format, strict JSON, deterministic manifest regeneration, and diff checks are green) | this native checkpoint +- weeks unemployed | complete (archived commit `42ed5d45c56df80d754fbe24cce21cfeb8d05cbe` `datasets/cps/cps.py` lines 1438-1442 directly carries person `LKWEEKS` with only historical -1 mapped to zero, `datasets/cps/census_cps.py` lines 39-58, 89-133, 245, and 306-351 identifies and loads the 2023 ASEC public-use archive for income year 2022, and `calibration/puf_impute.py` lines 703-785 defines the special six-role/demographic plus optional-unemployment-compensation, at-most-5,000-row QRF and PUF-only zero rule; official archive SHA-256 `d2e000250782adfbdd7f29c82b66d866591a30f0d330496698ec19f9c784ce11` and member SHA-256 `19b56537e50e7663f954361ef2bb5ce9cef8d9d45f156fe1a69a99b654198ffe` yield 146,133 unique people, 4,200 positives, 3.14560943% weighted support, and 172,563,622.04 weighted person-weeks, and exact `PERIDNUM` plus redundant Census identity joins fill all 54,471 clone-expanded 2022 rows without statistical invention or clone disagreement; the persisted 166,324-person production sparse artifact SHA-256 `0bf7451e2d8bf1da07dd9524b178015494b30d5850baecc18e87e0a7c2f9e999` passes integer/range/source/UC reconciliation with zero mismatches, 3,717 positive rows, and 5.218208% ASEC / 0.559432% PUF weighted signal, while source-anchored channel-share/weighted-week bands reject a one-row PUF collapse and run again after frozen-support selection; no false monetary probe is added because PolicyEngine-US 1.764.6 reads the leaf only in Pennsylvania UC alongside four other missing direct inputs, so an engine-graph regression and the hard signal gate replace an inevitably structural-zero probe; generated coverage is 158 required / 8 reviewed exclusions / 37 probes, and 360 family/coverage/validation/plan/L0/parity/runtime/builder/checkpoint tests, the full pinned-source audit, Ruff, format, strict JSON, deterministic manifest regeneration, and diff checks are green) | this native checkpoint diff --git a/experiments/build_j_recert/buildj_base.sh b/experiments/build_j_recert/buildj_base.sh index cc2ae425..16e5153e 100755 --- a/experiments/build_j_recert/buildj_base.sh +++ b/experiments/build_j_recert/buildj_base.sh @@ -1,11 +1,9 @@ #!/bin/bash -# Build J STEP 1 — base pool rebuild on main's stages (populace#368 re-certification). -# Replicates Build F's base command EXACTLY on origin/main's base builder -# (byte-identical to Build F's tools/build_us_puf_support_base.py) to (a) satisfy the -# #368 "base pool rebuild" step and (b) EMPIRICALLY confirm determinism. The SNAP -# (#350/#352/#353) and SCF-wealth (#373) enrichments are RELEASE-TIME source stages -# (us_runtime/*), NOT base-builder stages, so the base sha is expected == Build F's -# 18833fb6 (documented finding vs the task's "new base sha expected" premise). +# Build J STEP 1 — rebuild the current source-staged base pool. +# The branch now restores CPS/ACS housing inputs in the reusable base builder, +# so Build F's pre-housing base and any pre-prior-year-income base are stale. +# Cache reuse is allowed only when every newly restored persisted leaf carries +# signal. # # Compute discipline: ~12 min (Build F was 12m14s), one chunk (<30 min). Detached via # setsid+caffeinate; pidfile + real rc + append log + memory-pressure sampler. @@ -27,6 +25,45 @@ TS=$(date -u +%Y%m%dT%H%M%SZ) PLOG="$LOGDIR/pressure_base_$TS.log" mkdir -p "$LOGDIR" "$OUTDIR" say() { echo "[$(date -u +%FT%TZ)] $1" | tee -a "$LOG"; } +base_has_restored_signal() { + .venv/bin/python -c ' +import sys +import pandas as pd + +required = { + "person": ( + "pre_subsidy_rent", + "self_employment_income_last_year", + "previous_year_income_available", + "salt_refund_income", + "takes_up_medicare_if_eligible", + "workers_compensation", + "would_claim_wic", + ), + "spm_unit": ( + "receives_housing_assistance", + "takes_up_housing_assistance_if_eligible", + "spm_unit_tenure_type", + "spm_unit_energy_subsidy", + ), + "household": ("tenure_type",), +} +try: + with pd.HDFStore(sys.argv[1], mode="r") as store: + tables = { + entity: store.select(entity, columns=list(names)) + for entity, names in required.items() + } +except (KeyError, TypeError, ValueError): + raise SystemExit(1) +ok = all( + name in tables[entity] and tables[entity][name].nunique(dropna=False) > 1 + for entity, names in required.items() + for name in names +) +raise SystemExit(0 if ok else 1) +' "$1" +} rm -f "$LOGDIR/base.rc" cd "$WT" || { say "FATAL cannot cd $WT"; echo 2 > "$LOGDIR/base.rc"; exit 2; } @@ -34,21 +71,25 @@ source .venv/bin/activate 2>/dev/null # ---- integrity preflight on inputs (reproduce Build F's exact base) ---- chk() { local got exp; got=$(shasum -a 256 "$1" | cut -c1-16); [ "$got" = "$2" ] || { say "FATAL sha mismatch $1: $got != $2"; echo 2 > "$LOGDIR/base.rc"; exit 2; }; } -for f in "$USD/census_cps_2024.h5" "$USD/census_cps_2023.h5" "$USD/census_cps_2022.h5" "$USD/puf_2024.h5" "$AGING_FACTS" "$CDX" "$BLADDER"; do +for f in "$USD/census_cps_2024.h5" "$USD/census_cps_2023.h5" "$USD/census_cps_2022.h5" "$USD/puf_2024.h5" "$USD/acs_2022.h5" "$AGING_FACTS" "$CDX" "$BLADDER"; do [ -f "$f" ] || { say "FATAL missing input $f"; echo 2 > "$LOGDIR/base.rc"; exit 2; } done chk "$AGING_FACTS" a5d34d4aad325d8c chk "$CDX" 383a666631aafd4f chk "$BLADDER" 7ba39b959068181b +chk "$USD/acs_2022.h5" 0b319b496f19a691 SHORT=$(git rev-parse --short HEAD) PEUS=$(.venv/bin/python -c "from importlib.metadata import version; print(version('policyengine-us'))" 2>/dev/null) say "BUILD J BASE START commit=$SHORT pe-us=$PEUS pid=$$ pressure_log=$PLOG" -say " inputs: asec 2024/2023/2022 + puf_2024; aging-facts a5d34d4a; cdx 383a6666; bladder 7ba39b95; seed 0 n-est 32" +say " inputs: asec 2024/2023/2022 + puf_2024 + acs_2022; aging-facts a5d34d4a; cdx 383a6666; bladder 7ba39b95; seed 0 n-est 32" -if [ -f "$BASE" ]; then - say "BASE: already present at $BASE — verifying sha only (delete to force rebuild)" +if [ -f "$BASE" ] && base_has_restored_signal "$BASE"; then + say "BASE: already present with nondefault restored input surface at $BASE — reusing" else + if [ -f "$BASE" ]; then + say "BASE: cached artifact is stale or restored-input-default-only — rebuilding in place" + fi # ---- memory-pressure sampler (self-terminates on python exit; SAMPLE ONLY, no kill) ---- ( echo "ts_utc,free_pct,swap_used_mb,py_rss_mb,py_pid" while :; do @@ -62,12 +103,13 @@ else done ) >> "$PLOG" 2>&1 & SAMPLER=$! - say "BASE: rebuilding (3-year ASEC pool 2024+2023+2022, 2x PUF clone, spec-less equal thirds)" + say "BASE: rebuilding (3-year ASEC pool 2024+2023+2022, pinned ACS 2022 rent donor, 2x PUF clone, spec-less equal thirds)" .venv/bin/python tools/build_us_puf_support_base.py \ --asec-h5 2024="$USD/census_cps_2024.h5" \ --asec-h5 2023="$USD/census_cps_2023.h5" \ --asec-h5 2022="$USD/census_cps_2022.h5" \ --puf-h5 "$USD/puf_2024.h5" \ + --acs-h5 "$USD/acs_2022.h5" \ --target-year 2024 \ --seed 0 --n-estimators 32 \ --ledger-facts "$AGING_FACTS" \ @@ -83,13 +125,21 @@ else if [ $rc -ne 0 ]; then echo "$rc" > "$LOGDIR/base.rc"; say "BASE FAILED rc=$rc — diagnose in base_j.log"; exit "$rc"; fi fi +if ! base_has_restored_signal "$BASE"; then + echo 2 > "$LOGDIR/base.rc" + say "BASE FAILED: output does not persist all restored nondefault inputs" + exit 2 +fi + BASE_SHA=$(shasum -a 256 "$BASE" | cut -d' ' -f1) echo "$BASE_SHA" > "$LOGDIR/base.sha" say "BASE sha: $BASE_SHA" if [ "$BASE_SHA" = "18833fb68e60ee74461608d81a5c5ab7d52435e17026d9e3b062d9de18d6871f" ]; then - say "BASE sha CONFIRMS Build F 18833fb6 — base pool unchanged (assets/SNAP are release-time source stages, not base-builder stages)" + echo 2 > "$LOGDIR/base.rc" + say "BASE FAILED: Build F's pre-housing SHA was reused" + exit 2 else - say "BASE sha DIFFERS from 18833fb6 — investigate before proceeding (base builder or deps changed the pool)" + say "BASE sha differs from pre-housing Build F as required" fi echo "0" > "$LOGDIR/base.rc" say "BUILD J BASE DONE rc=0 sha=$BASE_SHA" diff --git a/packages/populace-build/src/populace/build/source_manifest.py b/packages/populace-build/src/populace/build/source_manifest.py index 48a05921..63a229d1 100644 --- a/packages/populace-build/src/populace/build/source_manifest.py +++ b/packages/populace-build/src/populace/build/source_manifest.py @@ -50,26 +50,54 @@ "compute_ratio", "declare_income_reference_offset", "derive", + "derive_childcare_inputs", + "derive_child_support_inputs", + "derive_disability_benefits", + "derive_energy_subsidy", + "derive_education_inputs", "derive_eligibility_inputs", "derive_hours_worked", + "derive_housing_tenure_inputs", "derive_immigration_status", + "derive_medicare_take_up", + "derive_other_health_insurance_premiums", + "derive_prior_year_income", "derive_snap_abawd_discretionary_exemption", "derive_snap_take_up", "derive_puf_policyengine_variables", "derive_mortgage_balance_hints", "derive_pregnancy", + "derive_relationship_inputs", + "derive_retirement_distributions", + "derive_retirement_contributions", + "derive_workers_compensation", + "derive_weeks_unemployed", + "derive_wic_claim", "disaggregate_aggregate_records", "fit_labor_market_models", "fit_tip_income_model", + "fit_weighted_acs_rent_qrf", "fit_vehicle_model", "fit_weighted_imputer", "fit_weighted_qrf", "fold_into", "head_carry", "join", + "impute_retirement_contributions_to_puf_support", + "impute_childcare_to_puf_support", + "impute_child_support_to_puf_support", + "impute_disability_benefits_to_puf_support", + "impute_energy_subsidy_to_puf_support", + "impute_housing_assistance_to_puf_support", + "impute_other_health_insurance_premiums_to_puf_support", + "impute_prior_year_income_to_puf_support", + "impute_retirement_distributions_to_puf_support", + "impute_workers_compensation_to_puf_support", + "impute_weeks_unemployed_to_puf_support", "map_columns", "read_table", "read_tables", + "read_acs_rent_donor", "replace_sentinels", "split_component_by_share", "support_clip", diff --git a/packages/populace-build/src/populace/build/us/ecps_parity_known_gaps.json b/packages/populace-build/src/populace/build/us/ecps_parity_known_gaps.json index 98587b7e..5ae778c9 100644 --- a/packages/populace-build/src/populace/build/us/ecps_parity_known_gaps.json +++ b/packages/populace-build/src/populace/build/us/ecps_parity_known_gaps.json @@ -2,357 +2,596 @@ "schema_version": 1, "description": "Reason'd, issue-linked parity exemption register (the debt ledger). Each entry names a layer the incumbent eCPS populates that the current Populace candidate does not, with the reason it is absent and the tracking issue that owns closing it. parity_gate is run with known_gaps = the keys of this file, and fails on any entry whose layer the candidate now populates (stale exemptions cannot rot silently); the launch condition is that this register shrinks to deliberate scope decisions, not TODOs. Seeded from the first parity run of populace_us_2024.h5 against the pinned reference (see ecps_parity_reference.json). Every issue ref is an existing open PolicyEngine/populace issue.", "known_gaps": { - "alimony_expense": { - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "issue": "PolicyEngine/populace#38" - }, - "alimony_income": { - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "issue": "PolicyEngine/populace#38" - }, - "attends_eligible_educational_institution_for_american_opportunity_credit": { - "reason": "American Opportunity / tuition education-credit input essentially unimputed on the candidate frame (education credits understated).", - "issue": "PolicyEngine/populace#253" - }, - "auto_loan_balance": { - "reason": "Vehicle-ownership / auto-loan input (populace#49 cites us-data #267/#281; see also #252 for the OBBBA auto-loan deduction); not yet on the frame.", - "issue": "PolicyEngine/populace#49" - }, - "auto_loan_interest": { - "reason": "Vehicle-ownership / auto-loan input (populace#49 cites us-data #267/#281; see also #252 for the OBBBA auto-loan deduction); not yet on the frame.", - "issue": "PolicyEngine/populace#49" - }, - "business_is_sstb": { - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "issue": "PolicyEngine/populace#298" - }, - "casualty_loss": { - "reason": "SPM MOOP / work-expense / deduction input (childcare-MOOP-work-expense family of #32); not yet on the candidate frame.", - "issue": "PolicyEngine/populace#32" - }, - "child_support_expense": { - "reason": "HHS OCSE child-support received/paid SPM input; not yet on the frame.", - "issue": "PolicyEngine/populace#32" - }, - "child_support_received": { - "reason": "HHS OCSE child-support received/paid SPM input; not yet on the frame.", - "issue": "PolicyEngine/populace#32" - }, - "cps_race": { - "reason": "Demographic / occupation / labor-status descriptor input carried by the incumbent but not yet on the candidate frame; remaining input-layer work (#38).", - "issue": "PolicyEngine/populace#38" - }, - "detailed_occupation_recode": { - "reason": "Demographic / occupation / labor-status descriptor input carried by the incumbent but not yet on the candidate frame; remaining input-layer work (#38).", - "issue": "PolicyEngine/populace#38" - }, - "disability_benefits": { - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "issue": "PolicyEngine/populace#38" - }, - "domestic_production_ald": { - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "issue": "PolicyEngine/populace#298" - }, - "educational_assistance": { - "reason": "Educational-assistance income input tied to the education-credit surface; not yet on the candidate frame.", - "issue": "PolicyEngine/populace#253" - }, - "educator_expense": { - "reason": "SPM MOOP / work-expense / deduction input (childcare-MOOP-work-expense family of #32); not yet on the candidate frame.", - "issue": "PolicyEngine/populace#32" - }, "employer_sponsored_insurance_premiums": { - "reason": "SPM MOOP / work-expense / deduction input (childcare-MOOP-work-expense family of #32); not yet on the candidate frame.", - "issue": "PolicyEngine/populace#32" - }, - "estate_income_would_be_qualified": { - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "issue": "PolicyEngine/populace#298" - }, - "farm_operations_income": { - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "issue": "PolicyEngine/populace#38" - }, - "farm_operations_income_would_be_qualified": { - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "issue": "PolicyEngine/populace#298" - }, - "farm_rent_income": { - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "issue": "PolicyEngine/populace#38" - }, - "farm_rent_income_would_be_qualified": { - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "issue": "PolicyEngine/populace#298" + "reason": "SOURCE UNAVAILABILITY WITH EVIDENCE: archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe datasets/cps/cps.py lines 197-271 and 1575-1581 requires NOW_OWNGRP, NOW_HIPAID, NOW_GRPFTYP, and PHIP_VAL; datasets/cps/census_cps.py lines 13-55 declares the three NOW_* fields optional. All three SHA-recorded 2022-2024 HDF inputs consumed by the hermetic --asec-h5 build contain PHIP_VAL but omit all three required NOW_* fields, so the retired owner/payment/plan-type derivation cannot execute. has_esi and NOW_GRP are not semantically valid substitutes.", + "issue": "PolicyEngine/populace#32", + "evidence": { + "classification": "source_unavailability", + "retired_derivation": { + "repository_owner": "PolicyEngine", + "repository_name_parts": ["policyengine-", "us-data"], + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "datasets", "cps", "cps.py"], + "lines": "197-271,1575-1581" + }, + "optional_source_fields": { + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "datasets", "cps", "census_cps.py"], + "lines": "13-55" + }, + "required_columns": ["NOW_OWNGRP", "NOW_HIPAID", "NOW_GRPFTYP", "PHIP_VAL"], + "hermetic_inputs": [ + { + "filename": "census_cps_2022.h5", + "sha256": "7ccca976284bb47815d84460cc4f75a0a65d26d7754ab0a0f417de351b3d474e", + "present_columns": ["PHIP_VAL"], + "missing_columns": ["NOW_OWNGRP", "NOW_HIPAID", "NOW_GRPFTYP"] + }, + { + "filename": "census_cps_2023.h5", + "sha256": "cb57817327799f42b741caed5f9be94d04021c2e6809c1ad7bd0686da5428d88", + "present_columns": ["PHIP_VAL"], + "missing_columns": ["NOW_OWNGRP", "NOW_HIPAID", "NOW_GRPFTYP"] + }, + { + "filename": "census_cps_2024.h5", + "sha256": "ec36604cb735a660b51b0b2f90be27d803b5878f3464fb30d0eacead59c1260d", + "present_columns": ["PHIP_VAL"], + "missing_columns": ["NOW_OWNGRP", "NOW_HIPAID", "NOW_GRPFTYP"] + } + ], + "hermetic_build_contract": "experiments/build_j_recert/buildj_base.sh lines 65-69 passes the three prebuilt 2022-2024 artifacts via --asec-h5; experiments/build_j_recert/base_j.summary.json lines 55-75 records their SHA-256 digests; tools/build_us_puf_support_base.py lines 85-88, 647-687, and 748-755 consumes those mappings and does not refresh omitted raw Census columns." + } }, "financial_assistance": { - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "issue": "PolicyEngine/populace#38" - }, - "has_american_opportunity_credit_1098_t_or_exception": { - "reason": "American Opportunity / tuition education-credit input essentially unimputed on the candidate frame (education credits understated).", - "issue": "PolicyEngine/populace#253" - }, - "has_american_opportunity_credit_institution_ein": { - "reason": "American Opportunity / tuition education-credit input essentially unimputed on the candidate frame (education credits understated).", - "issue": "PolicyEngine/populace#253" - }, - "has_never_worked": { - "reason": "Demographic / occupation / labor-status descriptor input carried by the incumbent but not yet on the candidate frame; remaining input-layer work (#38).", - "issue": "PolicyEngine/populace#38" - }, - "hourly_wage": { - "reason": "Hours-worked / hourly-wage / labor input not yet carried through Populace US outputs (#242).", - "issue": "PolicyEngine/populace#242" - }, - "household_vehicles_owned": { - "reason": "Vehicle-ownership / auto-loan input (populace#49 cites us-data #267/#281; see also #252 for the OBBBA auto-loan deduction); not yet on the frame.", - "issue": "PolicyEngine/populace#49" - }, - "household_vehicles_value": { - "reason": "Vehicle-ownership / auto-loan input (populace#49 cites us-data #267/#281; see also #252 for the OBBBA auto-loan deduction); not yet on the frame.", - "issue": "PolicyEngine/populace#49" - }, - "household_weight": { - "reason": "Incumbent stored a household_weight input column; Populace carries weights as a typed Frame weight vector, not a stored input variable. Classify as reported-observation vs reviewed-exclusion under #38.", - "issue": "PolicyEngine/populace#38" - }, - "investment_income_elected_form_4952": { - "reason": "Capital-gains component detail (collectibles / unrecaptured 1250 / 4952 election / investment-interest) not yet split onto the candidate frame.", - "issue": "PolicyEngine/populace#274" + "reason": "SOURCE UNAVAILABILITY WITH EVIDENCE: archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe datasets/cps/cps.py lines 1493-1496 derives person-level financial_assistance directly from FIN_VAL; datasets/cps/census_cps.py lines 39-58 and 306-359 make FIN_VAL a required person source. All three SHA-recorded 2022-2024 ASEC H5 inputs consumed by the hermetic --asec-h5 build omit FIN_VAL, FIN_YN, and I_FINVAL from their person tables. HFINVAL and FFINVAL are household/family aggregates and cannot recover the person recipient without synthetic allocation. The retired PUF-clone QRF at datasets/cps/extended_cps.py lines 140-194 and 639-745 trains financial_assistance from CPS person observations, so PUF provides no independent replacement source.", + "issue": "PolicyEngine/populace#38", + "evidence": { + "classification": "source_unavailability", + "retired_derivation": { + "repository_owner": "PolicyEngine", + "repository_name_parts": ["policyengine-", "us-data"], + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "datasets", "cps", "cps.py"], + "lines": "1493-1496" + }, + "required_person_source_declaration": { + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "datasets", "cps", "census_cps.py"], + "lines": "39-58,306-359" + }, + "puf_clone_qrf": { + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "datasets", "cps", "extended_cps.py"], + "lines": "140-194,639-745", + "target_line": 166 + }, + "no_independent_puf_source": { + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "calibration", "puf_impute.py"], + "tax_detail_target_lines": "90-149,158-198", + "financial_assistance_occurrences": 0 + }, + "required_person_columns": ["FIN_VAL"], + "missing_person_context_columns": ["FIN_YN", "I_FINVAL"], + "official_variable_dictionary": { + "url": "https://api.census.gov/data/2024/cps/asec/mar/variables.html", + "FIN_VAL": {"entity": "person", "label": "Financial assistance income amount"}, + "HFINVAL": {"entity": "household", "label": "Financial assistance amount - Household"}, + "FFINVAL": {"entity": "family", "label": "Financial assistance amount - Family"} + }, + "policyengine_variable": { + "version": "1.764.6", + "entity": "person", + "definition_period": "year", + "documentation": "Cash financial assistance from outside the household." + }, + "semantic_non_substitutes": { + "HFINVAL": "Household aggregate; it does not identify the recipient person or preserve allocation when household members belong to different program units.", + "FFINVAL": "Family aggregate; it does not identify the recipient person or preserve allocation when family members belong to different program units.", + "rejection": "Assigning an aggregate to a head or first member, or splitting it among members, would synthesize a person-level allocation that the locked source does not report." + }, + "hermetic_inputs": [ + { + "filename": "census_cps_2022.h5", + "sha256": "7ccca976284bb47815d84460cc4f75a0a65d26d7754ab0a0f417de351b3d474e", + "missing_person_columns": ["FIN_VAL", "FIN_YN", "I_FINVAL"], + "present_household_columns": ["HFINVAL", "HFIN_YN"], + "present_family_columns": ["FFINVAL", "FINC_FIN"], + "positive_multi_person_families": 374 + }, + { + "filename": "census_cps_2023.h5", + "sha256": "cb57817327799f42b741caed5f9be94d04021c2e6809c1ad7bd0686da5428d88", + "missing_person_columns": ["FIN_VAL", "FIN_YN", "I_FINVAL"], + "present_household_columns": ["HFINVAL", "HFIN_YN"], + "present_family_columns": ["FFINVAL", "FINC_FIN"], + "positive_multi_person_families": 393 + }, + { + "filename": "census_cps_2024.h5", + "sha256": "ec36604cb735a660b51b0b2f90be27d803b5878f3464fb30d0eacead59c1260d", + "missing_person_columns": ["FIN_VAL", "FIN_YN", "I_FINVAL"], + "present_household_columns": ["HFINVAL", "HFIN_YN"], + "present_family_columns": ["FFINVAL", "FINC_FIN"], + "positive_multi_person_families": 440 + } + ], + "aggregate_identity": "For every locked artifact, HFINVAL equals the household sum of FFINVAL; neither aggregate identifies a recipient person.", + "hermetic_build_contract": "experiments/build_j_recert/buildj_base.sh lines 65-69 passes only the three prebuilt 2022-2024 artifacts via --asec-h5; experiments/build_j_recert/base_j.summary.json lines 55-75 records their SHA-256 digests; tools/build_us_puf_support_base.py lines 85-88 and 659-699 consumes only those mappings; packages/populace-build/src/populace/build/us_runtime/asec_pool.py lines 61-75 reads the prebuilt person/household tables without refreshing omitted Census columns." + } }, "investment_interest_expense": { - "reason": "Capital-gains component detail (collectibles / unrecaptured 1250 / 4952 election / investment-interest) not yet split onto the candidate frame.", - "issue": "PolicyEngine/populace#274" - }, - "is_computer_scientist": { - "reason": "Demographic / occupation / labor-status descriptor input carried by the incumbent but not yet on the candidate frame; remaining input-layer work (#38).", - "issue": "PolicyEngine/populace#38" - }, - "is_enrolled_at_least_half_time_for_american_opportunity_credit": { - "reason": "American Opportunity / tuition education-credit input essentially unimputed on the candidate frame (education credits understated).", - "issue": "PolicyEngine/populace#253" - }, - "is_executive_administrative_professional": { - "reason": "Demographic / occupation / labor-status descriptor input carried by the incumbent but not yet on the candidate frame; remaining input-layer work (#38).", - "issue": "PolicyEngine/populace#38" - }, - "is_farmer_fisher": { - "reason": "Demographic / occupation / labor-status descriptor input carried by the incumbent but not yet on the candidate frame; remaining input-layer work (#38).", - "issue": "PolicyEngine/populace#38" - }, - "is_hispanic": { - "reason": "Demographic / occupation / labor-status descriptor input carried by the incumbent but not yet on the candidate frame; remaining input-layer work (#38).", - "issue": "PolicyEngine/populace#38" - }, - "is_household_head": { - "reason": "Demographic / occupation / labor-status descriptor input carried by the incumbent but not yet on the candidate frame; remaining input-layer work (#38).", - "issue": "PolicyEngine/populace#38" - }, - "is_military": { - "reason": "Demographic / occupation / labor-status descriptor input carried by the incumbent but not yet on the candidate frame; remaining input-layer work (#38).", - "issue": "PolicyEngine/populace#38" - }, - "is_paid_hourly": { - "reason": "Hours-worked / hourly-wage / labor input not yet carried through Populace US outputs (#242).", - "issue": "PolicyEngine/populace#242" - }, - "is_pursuing_credential_for_american_opportunity_credit": { - "reason": "American Opportunity / tuition education-credit input essentially unimputed on the candidate frame (education credits understated).", - "issue": "PolicyEngine/populace#253" - }, - "is_separated": { - "reason": "Demographic / occupation / labor-status descriptor input carried by the incumbent but not yet on the candidate frame; remaining input-layer work (#38).", - "issue": "PolicyEngine/populace#38" - }, - "is_surviving_spouse": { - "reason": "Demographic / occupation / labor-status descriptor input carried by the incumbent but not yet on the candidate frame; remaining input-layer work (#38).", - "issue": "PolicyEngine/populace#38" - }, - "is_union_member_or_covered": { - "reason": "Hours-worked / hourly-wage / labor input not yet carried through Populace US outputs (#242).", - "issue": "PolicyEngine/populace#242" + "reason": "SOURCE UNAVAILABILITY WITH EVIDENCE: archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe maps the sole interest field used by the retired derivation, E19200, to interest_deduction at datasets/puf/puf.py line 654 and explicitly assigns that identical amount as deductible mortgage interest at lines 1548-1550. utils/mortgage_interest.py lines 141-155 and 268-320 derives investment_interest_expense only as max(total_interest_deduction - tax_unit_deductible, 0); the retired QRF separately predicted the two source-identical aliases before conversion, so any positive residual is stochastic prediction divergence rather than observed source detail. The SHA-pinned PUF 1.8.0 artifact consumed by the hermetic build has 499,045 all-zero investment_interest_expense values and omits interest_deduction, deductible_mortgage_interest, and raw E19200; all three locked ASEC inputs omit them too. A nonzero restoration would invent an unobserved mortgage-versus-investment split.", + "issue": "PolicyEngine/populace#274", + "evidence": { + "classification": "source_unavailability", + "retired_derivation": { + "repository_owner": "PolicyEngine", + "repository_name_parts": ["policyengine-", "us-data"], + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "utils", "mortgage_interest.py"], + "lines": "141-155,268-320,327-349" + }, + "retired_puf_source_identity": { + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "datasets", "puf", "puf.py"], + "interest_mapping_lines": "636-656", + "identical_mortgage_assignment_lines": "1548-1550", + "raw_source": "E19200" + }, + "retired_qrf_and_conversion_order": { + "qrf_path_parts": ["policyengine_", "us_data", "calibration", "puf_impute.py"], + "qrf_lines": "90-198,940-1075", + "conversion_path_parts": ["policyengine_", "us_data", "datasets", "cps", "extended_cps.py"], + "conversion_lines": "1178-1190", + "finding": "The QRF predicts interest_deduction and deductible_mortgage_interest separately even though the PUF construction makes their tax-unit source amounts identical; a positive post-QRF residual is not observed microdata detail." + }, + "required_intermediates": [ + "interest_deduction", + "deductible_mortgage_interest" + ], + "processed_puf": { + "filename": "puf_2024.h5", + "release": "policyengine/irs-soi-puf/1.8.0", + "sha256": "7669f5b5281f20080e77204f9bd4aabfad0aa101fa283e22caf9ba8d61d4d6df", + "rows": 499045, + "positive_values": 0, + "weighted_total": 0.0, + "missing_columns": [ + "interest_deduction", + "deductible_mortgage_interest", + "E19200" + ] + }, + "hermetic_asec_inputs": [ + { + "filename": "census_cps_2022.h5", + "sha256": "7ccca976284bb47815d84460cc4f75a0a65d26d7754ab0a0f417de351b3d474e", + "missing_columns": ["interest_deduction", "deductible_mortgage_interest", "E19200"] + }, + { + "filename": "census_cps_2023.h5", + "sha256": "cb57817327799f42b741caed5f9be94d04021c2e6809c1ad7bd0686da5428d88", + "missing_columns": ["interest_deduction", "deductible_mortgage_interest", "E19200"] + }, + { + "filename": "census_cps_2024.h5", + "sha256": "ec36604cb735a660b51b0b2f90be27d803b5878f3464fb30d0eacead59c1260d", + "missing_columns": ["interest_deduction", "deductible_mortgage_interest", "E19200"] + } + ], + "semantic_rejection": "SCF mortgage hints and aggregate SOI totals do not identify a return-level mortgage-versus-investment interest split. Re-running two stochastic predictions solely to manufacture a residual would synthesize values rather than restore a source observation.", + "hermetic_build_contract": "experiments/build_j_recert/buildj_base.sh lines 65-72 passes only the three locked ASEC H5 files and processed PUF H5; tools/build_us_puf_support_base.py lines 240-247 reads those arrays and cannot refresh omitted restricted raw-PUF columns." + } }, "is_unmarried_partner_of_household_head": { - "reason": "Demographic / occupation / labor-status descriptor input carried by the incumbent but not yet on the candidate frame; remaining input-layer work (#38).", - "issue": "PolicyEngine/populace#38" + "reason": "SOURCE UNAVAILABILITY WITH EVIDENCE: archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe datasets/cps/cps.py lines 1214-1221 derives is_unmarried_partner_of_household_head exclusively from person PERRP; datasets/cps/census_cps.py lines 306-318 declares that person source. All three SHA-recorded 2022-2024 ASEC HDF inputs consumed by the hermetic build omit PERRP. The 2022 and 2023 inputs also omit PECOHAB, A_EXPRRP, and A_FAMREL; only 2024 carries those alternatives, and its A_EXPRRP code 13 conflates partner with roommate. Populace's older-year relationship fallback uses only line, spouse, sex, and parent pointers and never identifies a partner. Using the 2024 pointer for one-third of the pool while treating the other two-thirds as false, or assigning partner status from the combined partner/roommate code, would synthesize statuses the locked sources do not report.", + "issue": "PolicyEngine/populace#38", + "evidence": { + "classification": "source_unavailability", + "retired_derivation": { + "repository_owner": "PolicyEngine", + "repository_name_parts": ["policyengine-", "us-data"], + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "datasets", "cps", "cps.py"], + "lines": "1214-1221" + }, + "required_person_source_declaration": { + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "datasets", "cps", "census_cps.py"], + "lines": "306-318" + }, + "required_person_columns": ["PERRP"], + "official_variable_dictionary": { + "url": "https://api.census.gov/data/2024/cps/asec/mar/variables.html", + "PERRP": "Expanded relationship categories", + "PECOHAB": "Demographics Line Number of Cohabiting Partner", + "A_EXPRRP": "Demographics, Relationship expanded code" + }, + "hermetic_inputs": [ + { + "filename": "census_cps_2022.h5", + "sha256": "7ccca976284bb47815d84460cc4f75a0a65d26d7754ab0a0f417de351b3d474e", + "missing_person_columns": ["PERRP", "PECOHAB", "A_EXPRRP", "A_FAMREL"] + }, + { + "filename": "census_cps_2023.h5", + "sha256": "cb57817327799f42b741caed5f9be94d04021c2e6809c1ad7bd0686da5428d88", + "missing_person_columns": ["PERRP", "PECOHAB", "A_EXPRRP", "A_FAMREL"] + }, + { + "filename": "census_cps_2024.h5", + "sha256": "ec36604cb735a660b51b0b2f90be27d803b5878f3464fb30d0eacead59c1260d", + "missing_person_columns": ["PERRP"], + "present_person_columns": ["PECOHAB", "A_EXPRRP", "A_FAMREL"], + "positive_pecohab_rows": 8114, + "a_exprrp_partner_or_roommate_rows": 4257 + } + ], + "pooled_source_year_shares": { + "2022": 0.3333333333333333, + "2023": 0.3333333333333333, + "2024": 0.3333333333333333 + }, + "semantic_non_substitutes": { + "A_EXPRRP": "Only the 2024 locked input carries it, and code 13 is the combined CPSRelationshipCode.PARTNER_OR_ROOMMATE category rather than the retired PERRP unmarried-partner category.", + "PECOHAB": "Only the 2024 locked input carries the cohabiting-partner line pointer; it cannot recover partner status for the equally weighted 2022 and 2023 source years.", + "derived_A_EXPRRP": "packages/populace-build/src/populace/build/us_runtime/asec_pool.py lines 397-422 derives older-year relationship codes from line, spouse, sex, and parent pointers and never emits code 13.", + "rejection": "Using only the 2024 alternatives and filling the equally weighted 2022-2023 source years as false, or splitting the combined partner/roommate category, would synthesize person statuses absent from the locked inputs." + }, + "hermetic_build_contract": "experiments/build_j_recert/buildj_base.sh lines 65-69 passes only the three prebuilt 2022-2024 artifacts via --asec-h5; experiments/build_j_recert/base_j.summary.json lines 55-75 records their SHA-256 digests and equal source-year shares; packages/populace-build/src/populace/build/us_runtime/asec_pool.py lines 61-75 reads those tables without refreshing omitted Census columns." + } }, "is_wic_at_nutritional_risk": { - "reason": "WIC take-up / eligibility-screen input the incumbent seeds; not yet produced by Populace (take-up family).", - "issue": "PolicyEngine/populace#312" - }, - "keogh_distributions": { - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "issue": "PolicyEngine/populace#38" - }, - "long_term_capital_gains_on_collectibles": { - "reason": "Capital-gains component detail (collectibles / unrecaptured 1250 / 4952 election / investment-interest) not yet split onto the candidate frame.", - "issue": "PolicyEngine/populace#274" - }, - "net_worth": { - "reason": "SCF-derived household asset/net-worth stock; not yet imputed onto the candidate frame (wealth imputation backlog).", - "issue": "PolicyEngine/populace#49" - }, - "other_health_insurance_premiums": { - "reason": "SPM MOOP / work-expense / deduction input (childcare-MOOP-work-expense family of #32); not yet on the candidate frame.", - "issue": "PolicyEngine/populace#32" - }, - "partnership_s_corp_income_would_be_qualified": { - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "issue": "PolicyEngine/populace#298" - }, - "pre_subsidy_rent": { - "reason": "HUD housing-assistance / tenure SPM input feeding capped housing subsidy; not yet on the candidate frame.", - "issue": "PolicyEngine/populace#32" - }, - "previous_year_income_available": { - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "issue": "PolicyEngine/populace#38" - }, - "qualified_bdc_income": { - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "issue": "PolicyEngine/populace#298" - }, - "qualified_reit_and_ptp_income": { - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "issue": "PolicyEngine/populace#298" - }, - "qualified_tuition_expenses": { - "reason": "American Opportunity / tuition education-credit input essentially unimputed on the candidate frame (#253).", - "issue": "PolicyEngine/populace#253" - }, - "receives_housing_assistance": { - "reason": "HUD housing-assistance / tenure SPM input feeding capped housing subsidy; not yet on the candidate frame.", - "issue": "PolicyEngine/populace#32" - }, - "rental_income_would_be_qualified": { - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "issue": "PolicyEngine/populace#298" - }, - "salt_refund_income": { - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "issue": "PolicyEngine/populace#38" - }, - "self_employment_income_last_year": { - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "issue": "PolicyEngine/populace#38" - }, - "self_employment_income_would_be_qualified": { - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "issue": "PolicyEngine/populace#298" - }, - "spm_unit_energy_subsidy": { - "reason": "SPM energy-subsidy (LIHEAP) resource input; not yet on the candidate frame.", - "issue": "PolicyEngine/populace#32" - }, - "spm_unit_tenure_type": { - "reason": "HUD housing-assistance / tenure SPM input feeding capped housing subsidy; not yet on the candidate frame.", - "issue": "PolicyEngine/populace#32" - }, - "sstb_self_employment_income_before_lsr": { - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "issue": "PolicyEngine/populace#298" - }, - "sstb_self_employment_income_would_be_qualified": { - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "issue": "PolicyEngine/populace#298" - }, - "sstb_unadjusted_basis_qualified_property": { - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "issue": "PolicyEngine/populace#298" - }, - "sstb_w2_wages_from_qualified_business": { - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "issue": "PolicyEngine/populace#298" + "reason": "SOURCE UNAVAILABILITY WITH EVIDENCE: archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe datasets/cps/cps.py lines 684-701 does not map a measured nutritional-risk assessment; it loads category priors, makes a seeded Bernoulli draw, and ORs that synthetic draw with reported WIC receipt. parameters/take_up/wic_nutritional_risk.yaml lines 1-13 gives every fixed category prior a 1980 effective key (the infant factor was historically adopted in 1991), while parameters/__init__.py lines 38-50 merely carries the latest value forward. datasets/cps/census_cps.py lines 293-365 loads only WIC receipt/value fields, and cps.py line 1558 maps WICYN solely to receives_wic. All three SHA-recorded 2022-2024 ASEC H5 inputs consumed by the hermetic build contain only SPM_WICVAL/WICYN on person and HRNUMWIC/HRWICYN on household; none contains an individual nutritional-risk assessment. The Census questionnaire gives WICYN an adult-female universe and allows receipt reporting for the respondent or on behalf of a child, so a yes does not identify the assessed person and code 0 is not-in-universe rather than a negative assessment. USDA requires a competent professional to assess and document risk in separate WIC certification records; FNS adopted an all-category 1.0 risk adjustment while producing its CY2020 estimates and applied it to the revised CY2016-2021 series. Assigning risk to a WICYN row, assigning False to nonrecipients, or recreating the retired random False outcomes would synthesize person-level assessments absent from the locked sources, while assigning True to all unknowns is the PolicyEngine-US default and carries no non-default signal.", + "issue": "PolicyEngine/populace#312", + "evidence": { + "classification": "source_unavailability", + "retired_derivation": { + "repository_owner": "PolicyEngine", + "repository_name_parts": ["policyengine-", "us-data"], + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "datasets", "cps", "cps.py"], + "lines": "684-701", + "method": "receives_wic OR category-specific seeded Bernoulli draw" + }, + "retired_risk_rates": { + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "parameters", "take_up", "wic_nutritional_risk.yaml"], + "lines": "1-13", + "vintage": "YAML effective key 1980 for all values; historical infant factor adopted by USDA in 1991", + "values": { + "PREGNANT": 0.913, + "POSTPARTUM": 0.933, + "BREASTFEEDING": 0.889, + "INFANT": 0.95, + "CHILD": 0.752, + "NONE": 0 + }, + "historical_source": "https://www.ncbi.nlm.nih.gov/books/NBK221951/" + }, + "category_rate_source": { + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "parameters", "__init__.py"], + "lines": "38-50", + "method": "latest category value with year less than or equal to build year" + }, + "retired_receipt_mapping": { + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "datasets", "cps", "cps.py"], + "lines": "1553-1559", + "mapping": "receives_wic = person.WICYN == 1" + }, + "retired_source_declaration": { + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "datasets", "cps", "census_cps.py"], + "lines": "293-365", + "available_wic_fields": ["SPM_WICVAL", "WICYN"] + }, + "official_census_dictionary": { + "url": "https://www2.census.gov/programs-surveys/cps/techdocs/cpsmar24.pdf", + "field": "WICYN", + "definition": "Who received WIC?", + "universe": "Adult female", + "questionnaire_context": "Receipt may be reported for the respondent herself or on behalf of a child.", + "codes": {"0": "Not in universe", "1": "Received WIC", "2": "Did not receive WIC"} + }, + "official_certification_source": { + "url": "https://www.fns.usda.gov/research/wic/participant-program-characteristics-2022", + "finding": "A competent professional authority assesses and documents at least one nutritional-risk criterion; WIC PC is a census of active certified participants built from state administrative records.", + "record_linkage_inference": "The FNS page describes a distinct administrative-record system; the three locked ASEC schemas contain neither an assessment field nor a link key. Therefore the hermetic build cannot attach those certification records to ASEC people." + }, + "current_fns_estimation_method": { + "url": "https://www.fns.usda.gov/sites/default/files/resource-files/wic-eligibility-report-vol1-2021.pdf", + "lines": "PDF page 84 (printed report page 70)", + "adopted_when": "while producing the CY2020 estimates", + "applied_period": "revised CY2016-CY2021 estimates in the cited report", + "risk_adjustment": 1.0, + "scope": "all participant categories" + }, + "policyengine_variable": { + "version": "1.764.6", + "entity": "person", + "definition_period": "month", + "value_type": "bool", + "default": true, + "consumer": "is_wic_eligible" + }, + "hermetic_inputs": [ + { + "filename": "census_cps_2022.h5", + "sha256": "7ccca976284bb47815d84460cc4f75a0a65d26d7754ab0a0f417de351b3d474e", + "person_wic_columns": ["SPM_WICVAL", "WICYN"], + "household_wic_columns": ["HRNUMWIC", "HRWICYN"], + "wicyn_counts": {"0": 109063, "1": 1349, "2": 35721} + }, + { + "filename": "census_cps_2023.h5", + "sha256": "cb57817327799f42b741caed5f9be94d04021c2e6809c1ad7bd0686da5428d88", + "person_wic_columns": ["SPM_WICVAL", "WICYN"], + "household_wic_columns": ["HRNUMWIC", "HRWICYN"], + "wicyn_counts": {"0": 108017, "1": 1342, "2": 34906} + }, + { + "filename": "census_cps_2024.h5", + "sha256": "ec36604cb735a660b51b0b2f90be27d803b5878f3464fb30d0eacead59c1260d", + "person_wic_columns": ["SPM_WICVAL", "WICYN"], + "household_wic_columns": ["HRNUMWIC", "HRWICYN"], + "wicyn_counts": {"0": 106402, "1": 1348, "2": 34375} + } + ], + "semantic_non_substitutes": { + "WICYN": "Its adult-female respondent may report receipt for herself or on behalf of a child, so the row does not identify the assessed person; nonreceipt also does not prove absence of nutritional risk.", + "SPM_WICVAL": "A replicated SPM-unit benefit amount identifies neither an assessed person nor a negative nutritional-risk determination.", + "retired_bernoulli_draw": "The early-1980s category priors are aggregate eligibility-estimation factors, not person-level assessment observations.", + "rejection": "Assigning False to nonrecipients, allocating an SPM amount to a person, or replaying the retired Bernoulli draw would synthesize individual nutritional-risk assessments absent from the hermetic sources." + }, + "hermetic_build_contract": "experiments/build_j_recert/buildj_base.sh lines 65-69 passes only the three prebuilt 2022-2024 artifacts via --asec-h5; experiments/build_j_recert/base_j.summary.json lines 55-75 records their SHA-256 digests; packages/populace-build/src/populace/build/us_runtime/asec_pool.py lines 61-75 reads only their person and household tables without refreshing or linking separate WIC certification records." + } }, "survivor_benefits": { - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "issue": "PolicyEngine/populace#38" + "reason": "SOURCE UNAVAILABILITY WITH EVIDENCE: archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe datasets/cps/cps.py line 1495 derives person-level survivor_benefits exclusively from SRVS_VAL; datasets/cps/census_cps.py lines 39-58 and 306-359 make SRVS_VAL a required person source. All three SHA-recorded 2022-2024 ASEC H5 inputs consumed by the hermetic --asec-h5 build omit SRVS_VAL from their person tables. Their FSURVAL/HSURVAL fields are only family/household aggregates: each household value exactly equals summed family values, and 657/647/610 positive families have multiple people, so the recipient cannot be recovered. The retired PUF-clone QRF at datasets/cps/extended_cps.py lines 140-194 and 639-745 trains survivor_benefits from CPS person observations, so PUF provides no independent replacement source. Restoring a nonzero person-year amount without SRVS_VAL would synthesize data.", + "issue": "PolicyEngine/populace#38", + "evidence": { + "classification": "source_unavailability", + "retired_derivation": { + "repository_owner": "PolicyEngine", + "repository_name_parts": ["policyengine-", "us-data"], + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "datasets", "cps", "cps.py"], + "lines": "1495" + }, + "required_person_source_declaration": { + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "datasets", "cps", "census_cps.py"], + "lines": "39-58,306-359" + }, + "puf_clone_qrf": { + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "datasets", "cps", "extended_cps.py"], + "lines": "140-194,639-745", + "target_line": 167 + }, + "no_independent_puf_source": { + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "calibration", "puf_impute.py"], + "tax_detail_target_lines": "90-149,158-198", + "survivor_benefits_occurrences": 0 + }, + "required_person_columns": ["SRVS_VAL"], + "aggregate_only_columns": { + "family": ["FSURVAL", "FINC_SUR"], + "household": ["HSURVAL", "HSUR_YN"] + }, + "policyengine_variable": { + "version": "1.764.6", + "entity": "person", + "definition_period": "year", + "documentation": "Survivor benefits other than Social Security survivor benefits." + }, + "hermetic_inputs": [ + { + "filename": "census_cps_2022.h5", + "sha256": "7ccca976284bb47815d84460cc4f75a0a65d26d7754ab0a0f417de351b3d474e", + "missing_person_columns": ["SRVS_VAL"], + "present_family_columns": ["FSURVAL", "FINC_SUR"], + "present_household_columns": ["HSURVAL", "HSUR_YN"], + "positive_multi_person_families": 657 + }, + { + "filename": "census_cps_2023.h5", + "sha256": "cb57817327799f42b741caed5f9be94d04021c2e6809c1ad7bd0686da5428d88", + "missing_person_columns": ["SRVS_VAL"], + "present_family_columns": ["FSURVAL", "FINC_SUR"], + "present_household_columns": ["HSURVAL", "HSUR_YN"], + "positive_multi_person_families": 647 + }, + { + "filename": "census_cps_2024.h5", + "sha256": "ec36604cb735a660b51b0b2f90be27d803b5878f3464fb30d0eacead59c1260d", + "missing_person_columns": ["SRVS_VAL"], + "present_family_columns": ["FSURVAL", "FINC_SUR"], + "present_household_columns": ["HSURVAL", "HSUR_YN"], + "positive_multi_person_families": 610 + } + ], + "aggregate_identity": "For every locked artifact, HSURVAL equals the household sum of FSURVAL; neither aggregate identifies the recipient person.", + "semantic_rejection": "The retired derivation requires a reported person amount. Its CPS-only QRF cannot train without that leaf, and assigning FSURVAL or HSURVAL to a head, first member, or arbitrary split would synthesize a survivor-benefit observation absent from every locked person table.", + "hermetic_build_contract": "experiments/build_j_recert/buildj_base.sh lines 65-69 passes only the three prebuilt 2022-2024 artifacts via --asec-h5; experiments/build_j_recert/base_j.summary.json lines 55-75 records their SHA-256 digests; packages/populace-build/src/populace/build/us_runtime/asec_pool.py lines 61-75 reads the prebuilt person tables without refreshing omitted Census columns." + } }, "takes_up_dc_ptc": { - "reason": "Stochastic take-up flag the incumbent seeds so take-up-gated programs do not ship at 100% participation; Populace does not yet produce it.", - "issue": "PolicyEngine/populace#312" + "reason": "SOURCE UNAVAILABILITY WITH EVIDENCE: archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe datasets/cps/cps.py lines 565-599 does not carry an observed DC property-tax-credit claim; it loads a scalar rate and assigns takes_up_dc_ptc by an independent seeded Bernoulli draw. parameters/take_up/dc_ptc.yaml lines 1-11 divides an aggregate 37,133-claim count by a 131,791,388 PolicyEngine value estimate, which is neither dimensionally a participation denominator nor mathematically 0.32; utils/randomness.py lines 5-28 confirms the output is generated from a variable-name-seeded random stream. Archived census_cps.py lines 251-262 exposes only generic state-tax liability amounts, not a Schedule H claim; archived puf.py lines 704-719 maps other_credits from federal P08000, and the locked PUF has no state geography. All three SHA-recorded 2022-2024 ASEC H5 inputs and the SHA-recorded processed PUF 1.8.0 input consumed by the hermetic build omit takes_up_dc_ptc and any DC Schedule H claim indicator. The aggregate claim count cannot identify claiming tax units, so replaying the retired draw or selecting modeled-eligible units would synthesize claimant identities absent from every locked source.", + "issue": "PolicyEngine/populace#312", + "evidence": { + "classification": "source_unavailability", + "retired_derivation": { + "repository_owner": "PolicyEngine", + "repository_name_parts": ["policyengine-", "us-data"], + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "datasets", "cps", "cps.py"], + "lines": "565-599", + "method": "seeded Bernoulli draw below a scalar dc_ptc_rate" + }, + "retired_rate_parameter": { + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "parameters", "take_up", "dc_ptc.yaml"], + "lines": "1-11", + "value": 0.32, + "administrative_claim_count": 37133, + "comment_model_estimate": 131791388, + "denominator_description": "PolicyEngine DC PTC value estimate", + "external_eligible_population_denominator": false, + "dimensionally_consistent_participation_rate": false, + "arithmetically_consistent_with_value": false + }, + "retired_randomness": { + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "utils", "randomness.py"], + "lines": "5-28", + "key": "takes_up_dc_ptc" + }, + "required_tax_unit_observation": "takes_up_dc_ptc or an observed DC Schedule H claim indicator", + "missing_claim_column_patterns": ["takes_up_dc_ptc", "dc_ptc", "schedule_h", "property_tax_credit"], + "archived_asec_tax_columns": { + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "datasets", "cps", "census_cps.py"], + "lines": "251-262", + "generic_state_tax_columns": ["STATETAX_A", "STATETAX_B"] + }, + "archived_puf_credit_mapping": { + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "datasets", "puf", "puf.py"], + "lines": "704-719", + "other_credits_source": "P08000", + "scope": "federal PUF credit amount, not a DC Schedule H claim indicator" + }, + "official_census_variable_dictionary": { + "url": "https://api.census.gov/data/2025/cps/asec/mar/variables.html", + "STATETAX_A": "State income tax liability, after credits", + "STATETAX_B": "State income tax liability, before credits", + "SPM_STTAX": "SPM unit's state tax" + }, + "policyengine_variable": { + "version": "1.764.6", + "entity": "tax_unit", + "definition_period": "year", + "value_type": "bool", + "default": true + }, + "hermetic_asec_inputs": [ + { + "filename": "census_cps_2022.h5", + "sha256": "7ccca976284bb47815d84460cc4f75a0a65d26d7754ab0a0f417de351b3d474e", + "missing_columns": ["takes_up_dc_ptc"], + "present_generic_tax_amount_columns": ["STATETAX_A", "STATETAX_B", "SPM_STTAX"] + }, + { + "filename": "census_cps_2023.h5", + "sha256": "cb57817327799f42b741caed5f9be94d04021c2e6809c1ad7bd0686da5428d88", + "missing_columns": ["takes_up_dc_ptc"], + "present_generic_tax_amount_columns": ["STATETAX_A", "STATETAX_B", "SPM_STTAX"] + }, + { + "filename": "census_cps_2024.h5", + "sha256": "ec36604cb735a660b51b0b2f90be27d803b5878f3464fb30d0eacead59c1260d", + "missing_columns": ["takes_up_dc_ptc"], + "present_generic_tax_amount_columns": ["STATETAX_A", "STATETAX_B", "SPM_STTAX"] + } + ], + "processed_puf": { + "filename": "puf_2024.h5", + "release": "policyengine/irs-soi-puf/1.8.0", + "sha256": "7669f5b5281f20080e77204f9bd4aabfad0aa101fa283e22caf9ba8d61d4d6df", + "missing_arrays": ["takes_up_dc_ptc", "state_fips", "state_code", "state_code_str"] + }, + "current_take_up_contract": { + "path": "packages/populace-build/src/populace/build/us/take_up_contract.json", + "treatment": "rate_unsourced", + "rate_status": "model_relative" + }, + "semantic_non_substitutes": { + "STATETAX_A_STATETAX_B": "Generic state income-tax liabilities before/after credits do not identify which credit changed the liability or whether Schedule H was claimed.", + "SPM_STTAX": "An SPM-unit aggregate state-tax amount is neither tax-unit claimant identity nor a DC Schedule H claim indicator.", + "PUF_other_credits": "Archived puf.py maps federal P08000 to other_credits; the locked PUF lacks state geography and cannot identify a DC credit claimant.", + "rejection": "A published aggregate number of claims supplies no tax-unit claimant identity. Replaying a seeded draw, sorting eligible DC units, or otherwise allocating the count would synthesize the boolean values the locked microdata do not report." + }, + "hermetic_build_contract": "experiments/build_j_recert/buildj_base.sh lines 70-81 and 103-108 passes only the three prebuilt ASEC artifacts and processed PUF via --asec-h5/--puf-h5; experiments/build_j_recert/base_j.summary.json lines 55-75 and 317-318 records their SHA-256 digests." + } }, "takes_up_early_head_start_if_eligible": { - "reason": "Stochastic take-up flag the incumbent seeds so take-up-gated programs do not ship at 100% participation; Populace does not yet produce it.", - "issue": "PolicyEngine/populace#312" - }, - "takes_up_head_start_if_eligible": { - "reason": "Stochastic take-up flag the incumbent seeds so take-up-gated programs do not ship at 100% participation; Populace does not yet produce it.", - "issue": "PolicyEngine/populace#312" - }, - "takes_up_housing_assistance_if_eligible": { - "reason": "Stochastic take-up flag the incumbent seeds so take-up-gated programs do not ship at 100% participation; Populace does not yet produce it.", - "issue": "PolicyEngine/populace#312" - }, - "takes_up_medicare_if_eligible": { - "reason": "Stochastic take-up flag the incumbent seeds so take-up-gated programs do not ship at 100% participation; Populace does not yet produce it.", - "issue": "PolicyEngine/populace#312" - }, - "takes_up_ssi_if_eligible": { - "reason": "Stochastic take-up flag the incumbent seeds so take-up-gated programs do not ship at 100% participation; Populace does not yet produce it.", - "issue": "PolicyEngine/populace#312" - }, - "tax_exempt_ira_distributions": { - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "issue": "PolicyEngine/populace#38" - }, - "taxable_401k_distributions": { - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "issue": "PolicyEngine/populace#38" - }, - "taxable_403b_distributions": { - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "issue": "PolicyEngine/populace#38" - }, - "taxable_sep_distributions": { - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "issue": "PolicyEngine/populace#38" - }, - "tenure_type": { - "reason": "HUD housing-assistance / tenure SPM input feeding capped housing subsidy; not yet on the candidate frame.", - "issue": "PolicyEngine/populace#32" - }, - "tip_income": { - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "issue": "PolicyEngine/populace#38" - }, - "treasury_tipped_occupation_code": { - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "issue": "PolicyEngine/populace#38" - }, - "unadjusted_basis_qualified_property": { - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "issue": "PolicyEngine/populace#298" - }, - "unrecaptured_section_1250_gain": { - "reason": "Capital-gains component detail (collectibles / unrecaptured 1250 / 4952 election / investment-interest) not yet split onto the candidate frame.", - "issue": "PolicyEngine/populace#274" - }, - "unreimbursed_business_employee_expenses": { - "reason": "SPM MOOP / work-expense / deduction input (childcare-MOOP-work-expense family of #32); not yet on the candidate frame.", - "issue": "PolicyEngine/populace#32" - }, - "w2_wages_from_qualified_business": { - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "issue": "PolicyEngine/populace#298" - }, - "weeks_unemployed": { - "reason": "Demographic / occupation / labor-status descriptor input carried by the incumbent but not yet on the candidate frame; remaining input-layer work (#38).", - "issue": "PolicyEngine/populace#38" - }, - "workers_compensation": { - "reason": "Workers' compensation benefit SPM input; not yet on the candidate frame.", - "issue": "PolicyEngine/populace#32" - }, - "would_claim_wic": { - "reason": "WIC take-up / eligibility-screen input the incumbent seeds; not yet produced by Populace (take-up family).", - "issue": "PolicyEngine/populace#312" - }, - "would_file_taxes_voluntarily": { - "reason": "Voluntary-filing take-up flag the incumbent seeds; not yet produced by Populace (take-up family).", - "issue": "PolicyEngine/populace#312" + "reason": "SOURCE UNAVAILABILITY WITH EVIDENCE: archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe datasets/cps/cps.py lines 582-583 and 640-648 does not carry an observed Early Head Start enrollment indicator; it loads a scalar rate and assigns takes_up_early_head_start_if_eligible by an unanchored person-level draw. parameters/take_up/early_head_start.yaml lines 1-9 supplies 0.09 from a NIEER research chart, not an ACF administrative participation fact, and utils/randomness.py lines 5-28 confirms variable-name-seeded randomness. PolicyEngine-US 1.764.6 applies the input to people under age 3 or pregnant. All three SHA-recorded 2022-2024 ASEC H5 inputs and the processed PUF consumed by the hermetic build omit both Head Start take-up flags and any individual Head Start/Early Head Start enrollment indicator. The independently SHA-pinned full 2023 SIPP observes standard Head Start only through a nursery/preschool education item with no reported positives under age 3; its child-care fields cover only ages 2-7, contain no Early Head Start distinction or pregnancy branch, and have a Census-documented universe inconsistency. ACF's exact EHS cumulative-enrollment total includes every participant served during the year and turnover replacements, while its funded enrollment is capacity; neither identifies point-in-time person recipients. Assigning either aggregate to modeled-eligible infants, toddlers, or pregnant people, mapping ambiguous age-2 SIPP Head Start responses to EHS, or replaying the retired 0.09 draw would synthesize the missing person values.", + "issue": "PolicyEngine/populace#312", + "evidence": { + "classification": "source_unavailability", + "retired_derivation": { + "repository_owner": "PolicyEngine", + "repository_name_parts": ["policyengine-", "us-data"], + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "datasets", "cps", "cps.py"], + "lines": "582-583,640-648" + }, + "retired_rate": { + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "parameters", "take_up", "early_head_start.yaml"], + "lines": "1-9", + "value": 0.09, + "citation_scope": "NIEER research chart; no ACF administrative participation fact" + }, + "retired_randomness": { + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "utils", "randomness.py"], + "lines": "5-28", + "key": "takes_up_early_head_start_if_eligible" + }, + "policyengine_variable": { + "version": "1.764.6", + "entity": "person", + "definition_period": "year", + "value_type": "bool", + "default": true, + "eligibility_domain": "age < 3 or is_pregnant, plus Head Start income or categorical eligibility" + }, + "required_person_observation": "Individual Early Head Start enrollment covering infants, toddlers, and pregnant participants", + "hermetic_asec_inputs": [ + { + "filename": "census_cps_2022.h5", + "sha256": "7ccca976284bb47815d84460cc4f75a0a65d26d7754ab0a0f417de351b3d474e", + "missing_columns": ["takes_up_early_head_start_if_eligible", "takes_up_head_start_if_eligible", "is_enrolled_in_head_start"] + }, + { + "filename": "census_cps_2023.h5", + "sha256": "cb57817327799f42b741caed5f9be94d04021c2e6809c1ad7bd0686da5428d88", + "missing_columns": ["takes_up_early_head_start_if_eligible", "takes_up_head_start_if_eligible", "is_enrolled_in_head_start"] + }, + { + "filename": "census_cps_2024.h5", + "sha256": "ec36604cb735a660b51b0b2f90be27d803b5878f3464fb30d0eacead59c1260d", + "missing_columns": ["takes_up_early_head_start_if_eligible", "takes_up_head_start_if_eligible", "is_enrolled_in_head_start"] + } + ], + "processed_puf": { + "filename": "puf_2024.h5", + "release": "policyengine/irs-soi-puf/1.8.0", + "sha256": "7669f5b5281f20080e77204f9bd4aabfad0aa101fa283e22caf9ba8d61d4d6df", + "missing_columns": ["takes_up_early_head_start_if_eligible", "takes_up_head_start_if_eligible", "is_enrolled_in_head_start"] + }, + "pinned_sipp": { + "revision": "21280dca5995e978d706740a8a4b9b7860cfd7b6", + "sha256": "5c30439e365fc26483318ef61d1d8f4bb2f0e9d6bb47c22c06756a7698733ee2", + "direct_education_item": "EEDHEADST is nursery/preschool only and has zero reported positives under age 3", + "child_care_scope": "RDAYHS, RHEADST, and RNURHS cover ages 2-7 and do not distinguish Early Head Start", + "pregnancy_signal": "none", + "census_user_note": "https://www.census.gov/programs-surveys/sipp/tech-documentation/user-notes/2023-usernotes/2023-small-inconsist-child-care.html" + }, + "administrative_non_substitutes": { + "cumulative_enrollment": "Counts every participant served at any point in the program year, including turnover replacements; it is not point-in-time take-up and supplies no person identity.", + "funded_enrollment": "Capacity or federally supported slots, not observed occupancy or person identity.", + "source": "https://headstart.gov/sites/default/files/pdf/service-snapshot-ehs-2023-2024.pdf" + }, + "hermetic_build_contract": "experiments/build_j_recert/buildj_base.sh lines 65-81 passes only the three prebuilt ASEC artifacts and processed PUF via --asec-h5/--puf-h5; experiments/build_j_recert/base_j.summary.json lines 55-75 and 317-318 records their SHA-256 digests. The separately pinned SIPP donor does not observe the required Early Head Start domain." + } } } } diff --git a/packages/populace-build/src/populace/build/us/release_input_coverage_manifest.json b/packages/populace-build/src/populace/build/us/release_input_coverage_manifest.json index 84d9b50f..da61eb9e 100644 --- a/packages/populace-build/src/populace/build/us/release_input_coverage_manifest.json +++ b/packages/populace-build/src/populace/build/us/release_input_coverage_manifest.json @@ -4,29 +4,19 @@ "status": "required" }, "alimony_expense": { - "issue": "PolicyEngine/populace#38", - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "status": "reviewed_exclusion" + "status": "required" }, "alimony_income": { - "issue": "PolicyEngine/populace#38", - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "status": "reviewed_exclusion" + "status": "required" }, "attends_eligible_educational_institution_for_american_opportunity_credit": { - "issue": "PolicyEngine/populace#253", - "reason": "American Opportunity / tuition education-credit input essentially unimputed on the candidate frame (education credits understated).", - "status": "reviewed_exclusion" + "status": "required" }, "auto_loan_balance": { - "issue": "PolicyEngine/populace#49", - "reason": "Vehicle-ownership / auto-loan input (populace#49 cites us-data #267/#281; see also #252 for the OBBBA auto-loan deduction); not yet on the frame.", - "status": "reviewed_exclusion" + "status": "required" }, "auto_loan_interest": { - "issue": "PolicyEngine/populace#49", - "reason": "Vehicle-ownership / auto-loan input (populace#49 cites us-data #267/#281; see also #252 for the OBBBA auto-loan deduction); not yet on the frame.", - "status": "reviewed_exclusion" + "status": "required" }, "bank_account_assets": { "note": "SSI countable-resource asset input; required with NO reviewed exclusion per PolicyEngine/populace#368 so the gate fails until the asset stage is restored (Deliverable 2). Currently absent — this is the intended red gate.", @@ -40,14 +30,10 @@ "status": "required" }, "business_is_sstb": { - "issue": "PolicyEngine/populace#298", - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "status": "reviewed_exclusion" + "status": "required" }, "casualty_loss": { - "issue": "PolicyEngine/populace#32", - "reason": "SPM MOOP / work-expense / deduction input (childcare-MOOP-work-expense family of #32); not yet on the candidate frame.", - "status": "reviewed_exclusion" + "status": "required" }, "charitable_cash_donations": { "status": "required" @@ -56,14 +42,10 @@ "status": "required" }, "child_support_expense": { - "issue": "PolicyEngine/populace#32", - "reason": "HHS OCSE child-support received/paid SPM input; not yet on the frame.", - "status": "reviewed_exclusion" + "status": "required" }, "child_support_received": { - "issue": "PolicyEngine/populace#32", - "reason": "HHS OCSE child-support received/paid SPM input; not yet on the frame.", - "status": "reviewed_exclusion" + "status": "required" }, "congressional_district_geoid": { "status": "required" @@ -72,38 +54,26 @@ "status": "required" }, "cps_race": { - "issue": "PolicyEngine/populace#38", - "reason": "Demographic / occupation / labor-status descriptor input carried by the incumbent but not yet on the candidate frame; remaining input-layer work (#38).", - "status": "reviewed_exclusion" + "status": "required" }, "detailed_occupation_recode": { - "issue": "PolicyEngine/populace#38", - "reason": "Demographic / occupation / labor-status descriptor input carried by the incumbent but not yet on the candidate frame; remaining input-layer work (#38).", - "status": "reviewed_exclusion" + "status": "required" }, "disability_benefits": { - "issue": "PolicyEngine/populace#38", - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "status": "reviewed_exclusion" + "status": "required" }, "domestic_production_ald": { - "issue": "PolicyEngine/populace#298", - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "status": "reviewed_exclusion" + "status": "required" }, "educational_assistance": { - "issue": "PolicyEngine/populace#253", - "reason": "Educational-assistance income input tied to the education-credit surface; not yet on the candidate frame.", - "status": "reviewed_exclusion" + "status": "required" }, "educator_expense": { - "issue": "PolicyEngine/populace#32", - "reason": "SPM MOOP / work-expense / deduction input (childcare-MOOP-work-expense family of #32); not yet on the candidate frame.", - "status": "reviewed_exclusion" + "status": "required" }, "employer_sponsored_insurance_premiums": { "issue": "PolicyEngine/populace#32", - "reason": "SPM MOOP / work-expense / deduction input (childcare-MOOP-work-expense family of #32); not yet on the candidate frame.", + "reason": "SOURCE UNAVAILABILITY WITH EVIDENCE: archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe datasets/cps/cps.py lines 197-271 and 1575-1581 requires NOW_OWNGRP, NOW_HIPAID, NOW_GRPFTYP, and PHIP_VAL; datasets/cps/census_cps.py lines 13-55 declares the three NOW_* fields optional. All three SHA-recorded 2022-2024 HDF inputs consumed by the hermetic --asec-h5 build contain PHIP_VAL but omit all three required NOW_* fields, so the retired owner/payment/plan-type derivation cannot execute. has_esi and NOW_GRP are not semantically valid substitutes.", "status": "reviewed_exclusion" }, "employment_income_before_lsr": { @@ -113,36 +83,26 @@ "status": "required" }, "estate_income_would_be_qualified": { - "issue": "PolicyEngine/populace#298", - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "status": "reviewed_exclusion" + "status": "required" }, "farm_income": { "status": "required" }, "farm_operations_income": { - "issue": "PolicyEngine/populace#38", - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "status": "reviewed_exclusion" + "status": "required" }, "farm_operations_income_would_be_qualified": { - "issue": "PolicyEngine/populace#298", - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "status": "reviewed_exclusion" + "status": "required" }, "farm_rent_income": { - "issue": "PolicyEngine/populace#38", - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "status": "reviewed_exclusion" + "status": "required" }, "farm_rent_income_would_be_qualified": { - "issue": "PolicyEngine/populace#298", - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "status": "reviewed_exclusion" + "status": "required" }, "financial_assistance": { "issue": "PolicyEngine/populace#38", - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", + "reason": "SOURCE UNAVAILABILITY WITH EVIDENCE: archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe datasets/cps/cps.py lines 1493-1496 derives person-level financial_assistance directly from FIN_VAL; datasets/cps/census_cps.py lines 39-58 and 306-359 make FIN_VAL a required person source. All three SHA-recorded 2022-2024 ASEC H5 inputs consumed by the hermetic --asec-h5 build omit FIN_VAL, FIN_YN, and I_FINVAL from their person tables. HFINVAL and FFINVAL are household/family aggregates and cannot recover the person recipient without synthetic allocation. The retired PUF-clone QRF at datasets/cps/extended_cps.py lines 140-194 and 639-745 trains financial_assistance from CPS person observations, so PUF provides no independent replacement source.", "status": "reviewed_exclusion" }, "first_home_mortgage_balance": { @@ -154,15 +114,14 @@ "first_home_mortgage_origination_year": { "status": "required" }, + "fsla_overtime_premium": { + "status": "required" + }, "has_american_opportunity_credit_1098_t_or_exception": { - "issue": "PolicyEngine/populace#253", - "reason": "American Opportunity / tuition education-credit input essentially unimputed on the candidate frame (education credits understated).", - "status": "reviewed_exclusion" + "status": "required" }, "has_american_opportunity_credit_institution_ein": { - "issue": "PolicyEngine/populace#253", - "reason": "American Opportunity / tuition education-credit input essentially unimputed on the candidate frame (education credits understated).", - "status": "reviewed_exclusion" + "status": "required" }, "has_champva_health_coverage_at_interview": { "status": "required" @@ -183,9 +142,7 @@ "status": "required" }, "has_never_worked": { - "issue": "PolicyEngine/populace#38", - "reason": "Demographic / occupation / labor-status descriptor input carried by the incumbent but not yet on the candidate frame; remaining input-layer work (#38).", - "status": "reviewed_exclusion" + "status": "required" }, "has_non_marketplace_direct_purchase_health_coverage_at_interview": { "status": "required" @@ -209,66 +166,48 @@ "status": "required" }, "hourly_wage": { - "issue": "PolicyEngine/populace#242", - "reason": "Hours-worked / hourly-wage / labor input not yet carried through Populace US outputs (#242).", - "status": "reviewed_exclusion" + "status": "required" }, "hours_worked_last_week": { "status": "required" }, "household_vehicles_owned": { - "issue": "PolicyEngine/populace#49", - "reason": "Vehicle-ownership / auto-loan input (populace#49 cites us-data #267/#281; see also #252 for the OBBBA auto-loan deduction); not yet on the frame.", - "status": "reviewed_exclusion" + "status": "required" }, "household_vehicles_value": { - "issue": "PolicyEngine/populace#49", - "reason": "Vehicle-ownership / auto-loan input (populace#49 cites us-data #267/#281; see also #252 for the OBBBA auto-loan deduction); not yet on the frame.", - "status": "reviewed_exclusion" + "status": "required" }, "household_weight": { - "issue": "PolicyEngine/populace#38", - "reason": "Incumbent stored a household_weight input column; Populace carries weights as a typed Frame weight vector, not a stored input variable. Classify as reported-observation vs reviewed-exclusion under #38.", - "status": "reviewed_exclusion" + "status": "required" }, "immigration_status_str": { "status": "required" }, "investment_income_elected_form_4952": { - "issue": "PolicyEngine/populace#274", - "reason": "Capital-gains component detail (collectibles / unrecaptured 1250 / 4952 election / investment-interest) not yet split onto the candidate frame.", - "status": "reviewed_exclusion" + "status": "required" }, "investment_interest_expense": { "issue": "PolicyEngine/populace#274", - "reason": "Capital-gains component detail (collectibles / unrecaptured 1250 / 4952 election / investment-interest) not yet split onto the candidate frame.", + "reason": "SOURCE UNAVAILABILITY WITH EVIDENCE: archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe maps the sole interest field used by the retired derivation, E19200, to interest_deduction at datasets/puf/puf.py line 654 and explicitly assigns that identical amount as deductible mortgage interest at lines 1548-1550. utils/mortgage_interest.py lines 141-155 and 268-320 derives investment_interest_expense only as max(total_interest_deduction - tax_unit_deductible, 0); the retired QRF separately predicted the two source-identical aliases before conversion, so any positive residual is stochastic prediction divergence rather than observed source detail. The SHA-pinned PUF 1.8.0 artifact consumed by the hermetic build has 499,045 all-zero investment_interest_expense values and omits interest_deduction, deductible_mortgage_interest, and raw E19200; all three locked ASEC inputs omit them too. A nonzero restoration would invent an unobserved mortgage-versus-investment split.", "status": "reviewed_exclusion" }, "is_blind": { "status": "required" }, "is_computer_scientist": { - "issue": "PolicyEngine/populace#38", - "reason": "Demographic / occupation / labor-status descriptor input carried by the incumbent but not yet on the candidate frame; remaining input-layer work (#38).", - "status": "reviewed_exclusion" + "status": "required" }, "is_disabled": { "status": "required" }, "is_enrolled_at_least_half_time_for_american_opportunity_credit": { - "issue": "PolicyEngine/populace#253", - "reason": "American Opportunity / tuition education-credit input essentially unimputed on the candidate frame (education credits understated).", - "status": "reviewed_exclusion" + "status": "required" }, "is_executive_administrative_professional": { - "issue": "PolicyEngine/populace#38", - "reason": "Demographic / occupation / labor-status descriptor input carried by the incumbent but not yet on the candidate frame; remaining input-layer work (#38).", - "status": "reviewed_exclusion" + "status": "required" }, "is_farmer_fisher": { - "issue": "PolicyEngine/populace#38", - "reason": "Demographic / occupation / labor-status descriptor input carried by the incumbent but not yet on the candidate frame; remaining input-layer work (#38).", - "status": "reviewed_exclusion" + "status": "required" }, "is_female": { "status": "required" @@ -277,78 +216,59 @@ "status": "required" }, "is_hispanic": { - "issue": "PolicyEngine/populace#38", - "reason": "Demographic / occupation / labor-status descriptor input carried by the incumbent but not yet on the candidate frame; remaining input-layer work (#38).", - "status": "reviewed_exclusion" + "status": "required" }, "is_household_head": { - "issue": "PolicyEngine/populace#38", - "reason": "Demographic / occupation / labor-status descriptor input carried by the incumbent but not yet on the candidate frame; remaining input-layer work (#38).", - "status": "reviewed_exclusion" + "status": "required" }, "is_military": { - "issue": "PolicyEngine/populace#38", - "reason": "Demographic / occupation / labor-status descriptor input carried by the incumbent but not yet on the candidate frame; remaining input-layer work (#38).", - "status": "reviewed_exclusion" + "status": "required" }, "is_paid_hourly": { - "issue": "PolicyEngine/populace#242", - "reason": "Hours-worked / hourly-wage / labor input not yet carried through Populace US outputs (#242).", - "status": "reviewed_exclusion" + "status": "required" }, "is_pregnant": { "status": "required" }, "is_pursuing_credential_for_american_opportunity_credit": { - "issue": "PolicyEngine/populace#253", - "reason": "American Opportunity / tuition education-credit input essentially unimputed on the candidate frame (education credits understated).", - "status": "reviewed_exclusion" + "status": "required" }, "is_separated": { - "issue": "PolicyEngine/populace#38", - "reason": "Demographic / occupation / labor-status descriptor input carried by the incumbent but not yet on the candidate frame; remaining input-layer work (#38).", - "status": "reviewed_exclusion" + "status": "required" }, "is_surviving_spouse": { - "issue": "PolicyEngine/populace#38", - "reason": "Demographic / occupation / labor-status descriptor input carried by the incumbent but not yet on the candidate frame; remaining input-layer work (#38).", - "status": "reviewed_exclusion" + "status": "required" }, "is_union_member_or_covered": { - "issue": "PolicyEngine/populace#242", - "reason": "Hours-worked / hourly-wage / labor input not yet carried through Populace US outputs (#242).", - "status": "reviewed_exclusion" + "status": "required" }, "is_unmarried_partner_of_household_head": { "issue": "PolicyEngine/populace#38", - "reason": "Demographic / occupation / labor-status descriptor input carried by the incumbent but not yet on the candidate frame; remaining input-layer work (#38).", + "reason": "SOURCE UNAVAILABILITY WITH EVIDENCE: archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe datasets/cps/cps.py lines 1214-1221 derives is_unmarried_partner_of_household_head exclusively from person PERRP; datasets/cps/census_cps.py lines 306-318 declares that person source. All three SHA-recorded 2022-2024 ASEC HDF inputs consumed by the hermetic build omit PERRP. The 2022 and 2023 inputs also omit PECOHAB, A_EXPRRP, and A_FAMREL; only 2024 carries those alternatives, and its A_EXPRRP code 13 conflates partner with roommate. Populace's older-year relationship fallback uses only line, spouse, sex, and parent pointers and never identifies a partner. Using the 2024 pointer for one-third of the pool while treating the other two-thirds as false, or assigning partner status from the combined partner/roommate code, would synthesize statuses the locked sources do not report.", "status": "reviewed_exclusion" }, "is_wic_at_nutritional_risk": { "issue": "PolicyEngine/populace#312", - "reason": "WIC take-up / eligibility-screen input the incumbent seeds; not yet produced by Populace (take-up family).", + "reason": "SOURCE UNAVAILABILITY WITH EVIDENCE: archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe datasets/cps/cps.py lines 684-701 does not map a measured nutritional-risk assessment; it loads category priors, makes a seeded Bernoulli draw, and ORs that synthetic draw with reported WIC receipt. parameters/take_up/wic_nutritional_risk.yaml lines 1-13 gives every fixed category prior a 1980 effective key (the infant factor was historically adopted in 1991), while parameters/__init__.py lines 38-50 merely carries the latest value forward. datasets/cps/census_cps.py lines 293-365 loads only WIC receipt/value fields, and cps.py line 1558 maps WICYN solely to receives_wic. All three SHA-recorded 2022-2024 ASEC H5 inputs consumed by the hermetic build contain only SPM_WICVAL/WICYN on person and HRNUMWIC/HRWICYN on household; none contains an individual nutritional-risk assessment. The Census questionnaire gives WICYN an adult-female universe and allows receipt reporting for the respondent or on behalf of a child, so a yes does not identify the assessed person and code 0 is not-in-universe rather than a negative assessment. USDA requires a competent professional to assess and document risk in separate WIC certification records; FNS adopted an all-category 1.0 risk adjustment while producing its CY2020 estimates and applied it to the revised CY2016-2021 series. Assigning risk to a WICYN row, assigning False to nonrecipients, or recreating the retired random False outcomes would synthesize person-level assessments absent from the locked sources, while assigning True to all unknowns is the PolicyEngine-US default and carries no non-default signal.", "status": "reviewed_exclusion" }, "keogh_distributions": { - "issue": "PolicyEngine/populace#38", - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "status": "reviewed_exclusion" + "status": "required" }, "long_term_capital_gains_before_response": { "status": "required" }, "long_term_capital_gains_on_collectibles": { - "issue": "PolicyEngine/populace#274", - "reason": "Capital-gains component detail (collectibles / unrecaptured 1250 / 4952 election / investment-interest) not yet split onto the candidate frame.", - "status": "reviewed_exclusion" + "status": "required" + }, + "meets_ssi_disability_criteria": { + "status": "required" }, "miscellaneous_income": { "status": "required" }, "net_worth": { - "issue": "PolicyEngine/populace#49", - "reason": "SCF-derived household asset/net-worth stock; not yet imputed onto the candidate frame (wealth imputation backlog).", - "status": "reviewed_exclusion" + "status": "required" }, "non_qualified_dividend_income": { "status": "required" @@ -357,9 +277,7 @@ "status": "required" }, "other_health_insurance_premiums": { - "issue": "PolicyEngine/populace#32", - "reason": "SPM MOOP / work-expense / deduction input (childcare-MOOP-work-expense family of #32); not yet on the candidate frame.", - "status": "reviewed_exclusion" + "status": "required" }, "other_medical_expenses": { "status": "required" @@ -371,74 +289,64 @@ "status": "required" }, "partnership_s_corp_income_would_be_qualified": { - "issue": "PolicyEngine/populace#298", - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "status": "reviewed_exclusion" + "status": "required" }, "pre_subsidy_rent": { - "issue": "PolicyEngine/populace#32", - "reason": "HUD housing-assistance / tenure SPM input feeding capped housing subsidy; not yet on the candidate frame.", - "status": "reviewed_exclusion" + "status": "required" }, "previous_year_income_available": { - "issue": "PolicyEngine/populace#38", - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "status": "reviewed_exclusion" + "status": "required" }, "qualified_bdc_income": { - "issue": "PolicyEngine/populace#298", - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "status": "reviewed_exclusion" + "status": "required" }, "qualified_dividend_income": { "status": "required" }, + "qualified_passenger_vehicle_loan_interest": { + "status": "required" + }, "qualified_reit_and_ptp_income": { - "issue": "PolicyEngine/populace#298", - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "status": "reviewed_exclusion" + "status": "required" }, "qualified_tuition_expenses": { - "issue": "PolicyEngine/populace#253", - "reason": "American Opportunity / tuition education-credit input essentially unimputed on the candidate frame (#253).", - "status": "reviewed_exclusion" + "status": "required" }, "real_estate_taxes": { "status": "required" }, "receives_housing_assistance": { - "issue": "PolicyEngine/populace#32", - "reason": "HUD housing-assistance / tenure SPM input feeding capped housing subsidy; not yet on the candidate frame.", - "status": "reviewed_exclusion" + "status": "required" }, "rental_income": { "status": "required" }, "rental_income_would_be_qualified": { - "issue": "PolicyEngine/populace#298", - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "status": "reviewed_exclusion" + "status": "required" + }, + "roth_401k_contributions_desired": { + "status": "required" + }, + "roth_ira_contributions_desired": { + "status": "required" }, "salt_refund_income": { - "issue": "PolicyEngine/populace#38", - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "status": "reviewed_exclusion" + "status": "required" }, "selected_marketplace_plan_benchmark_ratio": { "status": "required" }, + "self_employed_pension_contributions_desired": { + "status": "required" + }, "self_employment_income_before_lsr": { "status": "required" }, "self_employment_income_last_year": { - "issue": "PolicyEngine/populace#38", - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "status": "reviewed_exclusion" + "status": "required" }, "self_employment_income_would_be_qualified": { - "issue": "PolicyEngine/populace#298", - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "status": "reviewed_exclusion" + "status": "required" }, "short_term_capital_gains": { "status": "required" @@ -456,40 +364,28 @@ "status": "required" }, "spm_unit_energy_subsidy": { - "issue": "PolicyEngine/populace#32", - "reason": "SPM energy-subsidy (LIHEAP) resource input; not yet on the candidate frame.", - "status": "reviewed_exclusion" + "status": "required" }, "spm_unit_pre_subsidy_childcare_expenses": { "status": "required" }, "spm_unit_tenure_type": { - "issue": "PolicyEngine/populace#32", - "reason": "HUD housing-assistance / tenure SPM input feeding capped housing subsidy; not yet on the candidate frame.", - "status": "reviewed_exclusion" + "status": "required" }, "ssn_card_type": { "status": "required" }, "sstb_self_employment_income_before_lsr": { - "issue": "PolicyEngine/populace#298", - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "status": "reviewed_exclusion" + "status": "required" }, "sstb_self_employment_income_would_be_qualified": { - "issue": "PolicyEngine/populace#298", - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "status": "reviewed_exclusion" + "status": "required" }, "sstb_unadjusted_basis_qualified_property": { - "issue": "PolicyEngine/populace#298", - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "status": "reviewed_exclusion" + "status": "required" }, "sstb_w2_wages_from_qualified_business": { - "issue": "PolicyEngine/populace#298", - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "status": "reviewed_exclusion" + "status": "required" }, "state_fips": { "status": "required" @@ -503,7 +399,7 @@ }, "survivor_benefits": { "issue": "PolicyEngine/populace#38", - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", + "reason": "SOURCE UNAVAILABILITY WITH EVIDENCE: archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe datasets/cps/cps.py line 1495 derives person-level survivor_benefits exclusively from SRVS_VAL; datasets/cps/census_cps.py lines 39-58 and 306-359 make SRVS_VAL a required person source. All three SHA-recorded 2022-2024 ASEC H5 inputs consumed by the hermetic --asec-h5 build omit SRVS_VAL from their person tables. Their FSURVAL/HSURVAL fields are only family/household aggregates: each household value exactly equals summed family values, and 657/647/610 positive families have multiple people, so the recipient cannot be recovered. The retired PUF-clone QRF at datasets/cps/extended_cps.py lines 140-194 and 639-745 trains survivor_benefits from CPS person observations, so PUF provides no independent replacement source. Restoring a nonzero person-year amount without SRVS_VAL would synthesize data.", "status": "reviewed_exclusion" }, "takes_up_aca_if_eligible": { @@ -511,42 +407,34 @@ }, "takes_up_dc_ptc": { "issue": "PolicyEngine/populace#312", - "reason": "Stochastic take-up flag the incumbent seeds so take-up-gated programs do not ship at 100% participation; Populace does not yet produce it.", + "reason": "SOURCE UNAVAILABILITY WITH EVIDENCE: archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe datasets/cps/cps.py lines 565-599 does not carry an observed DC property-tax-credit claim; it loads a scalar rate and assigns takes_up_dc_ptc by an independent seeded Bernoulli draw. parameters/take_up/dc_ptc.yaml lines 1-11 divides an aggregate 37,133-claim count by a 131,791,388 PolicyEngine value estimate, which is neither dimensionally a participation denominator nor mathematically 0.32; utils/randomness.py lines 5-28 confirms the output is generated from a variable-name-seeded random stream. Archived census_cps.py lines 251-262 exposes only generic state-tax liability amounts, not a Schedule H claim; archived puf.py lines 704-719 maps other_credits from federal P08000, and the locked PUF has no state geography. All three SHA-recorded 2022-2024 ASEC H5 inputs and the SHA-recorded processed PUF 1.8.0 input consumed by the hermetic build omit takes_up_dc_ptc and any DC Schedule H claim indicator. The aggregate claim count cannot identify claiming tax units, so replaying the retired draw or selecting modeled-eligible units would synthesize claimant identities absent from every locked source.", "status": "reviewed_exclusion" }, "takes_up_early_head_start_if_eligible": { "issue": "PolicyEngine/populace#312", - "reason": "Stochastic take-up flag the incumbent seeds so take-up-gated programs do not ship at 100% participation; Populace does not yet produce it.", + "reason": "SOURCE UNAVAILABILITY WITH EVIDENCE: archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe datasets/cps/cps.py lines 582-583 and 640-648 does not carry an observed Early Head Start enrollment indicator; it loads a scalar rate and assigns takes_up_early_head_start_if_eligible by an unanchored person-level draw. parameters/take_up/early_head_start.yaml lines 1-9 supplies 0.09 from a NIEER research chart, not an ACF administrative participation fact, and utils/randomness.py lines 5-28 confirms variable-name-seeded randomness. PolicyEngine-US 1.764.6 applies the input to people under age 3 or pregnant. All three SHA-recorded 2022-2024 ASEC H5 inputs and the processed PUF consumed by the hermetic build omit both Head Start take-up flags and any individual Head Start/Early Head Start enrollment indicator. The independently SHA-pinned full 2023 SIPP observes standard Head Start only through a nursery/preschool education item with no reported positives under age 3; its child-care fields cover only ages 2-7, contain no Early Head Start distinction or pregnancy branch, and have a Census-documented universe inconsistency. ACF's exact EHS cumulative-enrollment total includes every participant served during the year and turnover replacements, while its funded enrollment is capacity; neither identifies point-in-time person recipients. Assigning either aggregate to modeled-eligible infants, toddlers, or pregnant people, mapping ambiguous age-2 SIPP Head Start responses to EHS, or replaying the retired 0.09 draw would synthesize the missing person values.", "status": "reviewed_exclusion" }, "takes_up_eitc": { "status": "required" }, "takes_up_head_start_if_eligible": { - "issue": "PolicyEngine/populace#312", - "reason": "Stochastic take-up flag the incumbent seeds so take-up-gated programs do not ship at 100% participation; Populace does not yet produce it.", - "status": "reviewed_exclusion" + "status": "required" }, "takes_up_housing_assistance_if_eligible": { - "issue": "PolicyEngine/populace#312", - "reason": "Stochastic take-up flag the incumbent seeds so take-up-gated programs do not ship at 100% participation; Populace does not yet produce it.", - "status": "reviewed_exclusion" + "status": "required" }, "takes_up_medicaid_if_eligible": { "status": "required" }, "takes_up_medicare_if_eligible": { - "issue": "PolicyEngine/populace#312", - "reason": "Stochastic take-up flag the incumbent seeds so take-up-gated programs do not ship at 100% participation; Populace does not yet produce it.", - "status": "reviewed_exclusion" + "status": "required" }, "takes_up_snap_if_eligible": { "status": "required" }, "takes_up_ssi_if_eligible": { - "issue": "PolicyEngine/populace#312", - "reason": "Stochastic take-up flag the incumbent seeds so take-up-gated programs do not ship at 100% participation; Populace does not yet produce it.", - "status": "reviewed_exclusion" + "status": "required" }, "takes_up_tanf_if_eligible": { "status": "required" @@ -555,22 +443,16 @@ "status": "required" }, "tax_exempt_ira_distributions": { - "issue": "PolicyEngine/populace#38", - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "status": "reviewed_exclusion" + "status": "required" }, "tax_exempt_private_pension_income": { "status": "required" }, "taxable_401k_distributions": { - "issue": "PolicyEngine/populace#38", - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "status": "reviewed_exclusion" + "status": "required" }, "taxable_403b_distributions": { - "issue": "PolicyEngine/populace#38", - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "status": "reviewed_exclusion" + "status": "required" }, "taxable_interest_income": { "status": "required" @@ -582,84 +464,66 @@ "status": "required" }, "taxable_sep_distributions": { - "issue": "PolicyEngine/populace#38", - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "status": "reviewed_exclusion" + "status": "required" }, "tenure_type": { - "issue": "PolicyEngine/populace#32", - "reason": "HUD housing-assistance / tenure SPM input feeding capped housing subsidy; not yet on the candidate frame.", - "status": "reviewed_exclusion" + "status": "required" }, "tip_income": { - "issue": "PolicyEngine/populace#38", - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "status": "reviewed_exclusion" + "status": "required" }, "tract_geoid": { "status": "required" }, + "traditional_401k_contributions_desired": { + "status": "required" + }, + "traditional_ira_contributions_desired": { + "status": "required" + }, "treasury_tipped_occupation_code": { - "issue": "PolicyEngine/populace#38", - "reason": "Residual US tax-input / income-source layer not yet sourced onto the candidate frame; tracked as remaining input-layer work in #38.", - "status": "reviewed_exclusion" + "status": "required" }, "unadjusted_basis_qualified_property": { - "issue": "PolicyEngine/populace#298", - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "status": "reviewed_exclusion" + "status": "required" }, "unemployment_compensation": { "status": "required" }, "unrecaptured_section_1250_gain": { - "issue": "PolicyEngine/populace#274", - "reason": "Capital-gains component detail (collectibles / unrecaptured 1250 / 4952 election / investment-interest) not yet split onto the candidate frame.", - "status": "reviewed_exclusion" + "status": "required" }, "unreimbursed_business_employee_expenses": { - "issue": "PolicyEngine/populace#32", - "reason": "SPM MOOP / work-expense / deduction input (childcare-MOOP-work-expense family of #32); not yet on the candidate frame.", - "status": "reviewed_exclusion" + "status": "required" }, "veterans_benefits": { "status": "required" }, "w2_wages_from_qualified_business": { - "issue": "PolicyEngine/populace#298", - "reason": "Section 199A QBI / passthrough-qualification input; not yet on the candidate frame (QBI base already off vs targets, #298).", - "status": "reviewed_exclusion" + "status": "required" }, "weekly_hours_worked_before_lsr": { "status": "required" }, "weeks_unemployed": { - "issue": "PolicyEngine/populace#38", - "reason": "Demographic / occupation / labor-status descriptor input carried by the incumbent but not yet on the candidate frame; remaining input-layer work (#38).", - "status": "reviewed_exclusion" + "status": "required" }, "workers_compensation": { - "issue": "PolicyEngine/populace#32", - "reason": "Workers' compensation benefit SPM input; not yet on the candidate frame.", - "status": "reviewed_exclusion" + "status": "required" }, "would_claim_wic": { - "issue": "PolicyEngine/populace#312", - "reason": "WIC take-up / eligibility-screen input the incumbent seeds; not yet produced by Populace (take-up family).", - "status": "reviewed_exclusion" + "status": "required" }, "would_file_taxes_voluntarily": { - "issue": "PolicyEngine/populace#312", - "reason": "Voluntary-filing take-up flag the incumbent seeds; not yet produced by Populace (take-up family).", - "status": "reviewed_exclusion" + "status": "required" } }, "counts": { - "required": 70, - "reviewed_exclusion": 88, - "total": 158 + "required": 158, + "reviewed_exclusion": 8, + "total": 166 }, - "derivation": "Required surface = ecps_parity_reference.json populated layers (input columns the pinned, sha-verified reference eCPS populates). status='reviewed_exclusion' for ecps_parity_known_gaps.json entries (reason+issue from that register); EXCEPT the SSI countable-resource asset inputs (bank_account_assets, stock_assets, bond_assets), which are status='required' with NO exclusion per PolicyEngine/populace#368 so the gate fails on today's artifacts and asset restoration (Deliverable 2) turns it green. All other populated layers are 'required'. Regenerate with tools/build_us_release_input_coverage_manifest.py.", + "derivation": "Required surface = input columns in the pinned, sha-verified ecps_parity_reference.json populated layers, plus the documented post-reference fsla_overtime_premium, qualified_passenger_vehicle_loan_interest, five desired retirement-contribution inputs, and meets_ssi_disability_criteria required by shipped validation probes. status='reviewed_exclusion' for ecps_parity_known_gaps.json entries (reason+issue from that register); EXCEPT every primary-source restoration pinned by RESTORED_REFERENCE_ECPS_REQUIRED_INPUTS (including the Section 199A QBI family), and the SSI countable-resource asset inputs (bank_account_assets, stock_assets, bond_assets), which are status='required' with NO exclusion per PolicyEngine/populace#368 so the gate fails on today's artifacts and asset restoration (Deliverable 2) turns it green. All other populated layers are 'required'. Regenerate with tools/build_us_release_input_coverage_manifest.py.", "description": "Declared full-coverage contract for a US release: every input column the reference eCPS exports must be persisted as a key with non-default signal, or carry a reviewed exclusion. Enforced as a hard release gate (populace.build.us_runtime.release_input_coverage) that generalizes assert_required_us_release_source_columns from 5 columns to the full eCPS input surface.", "issue": "PolicyEngine/populace#368", "reference": { @@ -672,6 +536,86 @@ "vintage": "2024" }, "reform_coverage_probes": [ + { + "binding_inputs": [ + "self_employment_income_last_year" + ], + "budget_measure": "tax_unit_earned_income_last_year", + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "id": "prior_year_self_employment_neutralization", + "issue": "PolicyEngine/populace#38", + "min_abs_effect": 1000000000.0, + "name": "Prior-year self-employment support neutralization", + "neutralized_variable": "self_employment_income_last_year", + "parameter_changes": {}, + "period": 2024, + "reason": "PolicyEngine-US adds self_employment_income_last_year to earned_income_last_year and then aggregates it over tax-unit nondependents. The Wyden-Smith ACTC lookback is the downstream policy consumer, but its parameter also reads formula-owned prior-year wages, so neutralizing this one leaf is the unique coverage probe. The SHA-locked strict equal-share 2022-2024 ASEC pool carries $357.24 billion of weighted net source amount; without the adjacent-year carry the neutralization is a structural zero. The distinct previous_year_income_available flag has no formula consumer in PolicyEngine-US 1.764.6 and remains protected by the hard non-default column gate." + }, + { + "binding_inputs": [ + "keogh_distributions" + ], + "budget_measure": "income_tax", + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "id": "keogh_distribution_neutralization", + "issue": "PolicyEngine/populace#38", + "min_abs_effect": 1000000.0, + "name": "Keogh distribution neutralization", + "neutralized_variable": "keogh_distributions", + "parameter_changes": {}, + "period": 2024, + "reason": "PolicyEngine-US includes keogh_distributions directly in taxable retirement distributions and federal gross income. Neutralizing only the measured code-5 ASEC leaf lowers federal income tax. On the 865,046-person staged Build-J support, the locked ASEC source carries $148.97 million of weighted Keogh distributions and the baseline-minus-neutralized income-tax effect is +$24.47 million. Without the restored DST_SC*/DST_VAL* mapping the effect is a structural zero." + }, + { + "binding_inputs": [ + "is_household_head" + ], + "budget_measure": "spm_unit_capped_work_childcare_expenses", + "effect_direction": "baseline_minus_reform", + "expected_sign": "negative", + "id": "household_head_childcare_cap_neutralization", + "issue": "PolicyEngine/populace#38", + "min_abs_effect": 1000000.0, + "name": "Household-head childcare earned-cap neutralization", + "neutralized_variable": "is_household_head", + "parameter_changes": {}, + "period": 2024, + "reason": "PolicyEngine-US uses is_household_head to identify whose earnings cap SPM work-childcare expenses when an SPM unit contains multiple tax units. Neutralizing only the measured head flag falls back to tax-unit roles and changes the cap. On the 865,046-person staged Build-J artifact, baseline-minus-neutralized capped expenses are -$265.50 million. Without the restored P_SEQ input the neutralization is a structural zero." + }, + { + "binding_inputs": [ + "spm_unit_energy_subsidy" + ], + "budget_measure": "spm_unit_benefits", + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "id": "spm_unit_energy_subsidy_neutralization", + "issue": "PolicyEngine/populace#32", + "min_abs_effect": 100000000.0, + "name": "SPM energy-subsidy neutralization", + "neutralized_variable": "spm_unit_energy_subsidy", + "parameter_changes": {}, + "period": 2024, + "reason": "PolicyEngine-US adds this measured LIHEAP resource dollar-for-dollar to spm_unit_benefits. Neutralizing the leaf must therefore lower benefits by its weighted source mass; without the restored SPM_ENGVAL carry, the effect is a structural zero. No OBBBA provision consumes this SPM resource, so the direct neutralization is the uniquely isolating policy-engine probe." + }, + { + "binding_inputs": [ + "takes_up_medicare_if_eligible" + ], + "budget_measure": "medicare_cost", + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "id": "medicare_take_up_neutralization", + "issue": "PolicyEngine/populace#312", + "min_abs_effect": 1000000000.0, + "name": "Measured Medicare enrollment neutralization", + "neutralized_variable": "takes_up_medicare_if_eligible", + "parameter_changes": {}, + "period": 2024, + "reason": "PolicyEngine-US computes medicare_enrolled from the measured takes_up_medicare_if_eligible leaf and modeled eligibility, then gates Medicare costs on enrollment. Neutralizing only the restored MCARE == 1 leaf must reduce aggregate Medicare cost; without the measured carry the probe is a structural zero." + }, { "binding_inputs": [ "bank_account_assets", @@ -680,6 +624,7 @@ ], "budget_measure": "ssi", "effect_direction": "reform_minus_baseline", + "expected_sign": "positive", "id": "ssi_asset_limit_10k_20k", "issue": "PolicyEngine/populace#356", "min_abs_effect": 1000000000.0, @@ -692,7 +637,702 @@ "2024-01-01.2100-12-31": 10000 } }, + "period": 2024, "reason": "Raising the SSI resource limit is a pure relaxation that only changes who passes meets_ssi_resource_test, which compares ssi_countable_resources = bank_account_assets + stock_assets + bond_assets against the limit. With those asset inputs absent, countable resources are 0 for every record, everyone already passes the resource test, and raising the limit scores exactly $0. Dense-native reference: +$1.6B at $10k/$20k, ~+$16.1B with no limit (PolicyEngine/populace#356)." + }, + { + "binding_inputs": [ + "meets_ssi_disability_criteria" + ], + "budget_measure": "ssi", + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "id": "ssi_disability_criteria_neutralization", + "issue": "PolicyEngine/populace#312", + "min_abs_effect": 100000000.0, + "name": "SSI disability-criteria neutralization", + "neutralized_variable": "meets_ssi_disability_criteria", + "parameter_changes": {}, + "period": 2024, + "reason": "PolicyEngine-US requires the person-level disability criterion for non-aged SSI eligibility. Neutralizing only the restored SIPP-imputed criterion must therefore remove SSI from otherwise eligible disabled or blind people, so baseline-minus-neutralized SSI is positive. If the post-reference exported input is absent or degenerate, this isolated eligibility channel scores exactly $0. A deterministic 6,000-source-household Build-J smoke with the pinned SIPP and SCF donors scored +$586.393 million baseline-minus-neutralized SSI; the $100 million floor retains ample sampling margin while rejecting a materially weakened criterion channel." + }, + { + "binding_inputs": [ + "takes_up_ssi_if_eligible" + ], + "budget_measure": "ssi", + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "id": "ssi_take_up_neutralization", + "issue": "PolicyEngine/populace#312", + "min_abs_effect": 10000000000.0, + "name": "SSI take-up neutralization", + "neutralized_variable": "takes_up_ssi_if_eligible", + "parameter_changes": {}, + "period": 2024, + "reason": "PolicyEngine-US gates SSI benefits on the restored person-level take-up leaf after eligibility. Neutralizing only that leaf must therefore remove SSI from source-reported and SSA-count-calibrated recipients, so baseline-minus-neutralized SSI is positive. If the restored exported input is absent, all false, or not persisted, this isolated channel scores exactly $0. A production-ingredient sparse smoke (staged artifact sha256 c5939dad81153da51b2cc57081ddb3e729700366144868df742b3ad86eafcd7c; restored artifact sha256 9269360d3409fdc15c90c43dda394ada6c91eff5bb64c12ccd9def7d670dd077) measured +$57,114,569,526.38 of baseline-minus-neutralized 2024 SSI. The $10 billion floor retains over 5.7x observed margin while rejecting a materially degenerate persisted flag." + }, + { + "binding_inputs": [ + "takes_up_head_start_if_eligible" + ], + "budget_measure": "head_start", + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "id": "head_start_take_up_neutralization", + "issue": "PolicyEngine/populace#312", + "min_abs_effect": 100000000.0, + "name": "Measured Head Start take-up neutralization", + "neutralized_variable": "takes_up_head_start_if_eligible", + "parameter_changes": {}, + "period": 2024, + "reason": "PolicyEngine-US gates Head Start benefits on the person-level take-up leaf after modeled age, income, and categorical eligibility. Neutralizing only the restored SIPP response model must therefore remove Head Start from measured-proxy recipients, so baseline-minus-neutralized Head Start is positive. If the restored export is absent, all false, or not persisted, this isolated channel scores exactly $0. A production-ingredient sparse smoke (staged artifact sha256 67ad74b9ad9222ed342a0279dfc8175e872966fa59f86aeecb7fad52021ba500) measured +$4,575,181,976.69 of baseline-minus-neutralized 2024 Head Start. The $100 million floor retains over 45x observed margin while remaining far above numerical noise." + }, + { + "binding_inputs": [ + "qualified_tuition_expenses", + "is_pursuing_credential_for_american_opportunity_credit", + "attends_eligible_educational_institution_for_american_opportunity_credit", + "is_enrolled_at_least_half_time_for_american_opportunity_credit", + "has_american_opportunity_credit_1098_t_or_exception", + "has_american_opportunity_credit_institution_ein" + ], + "budget_measure": "american_opportunity_credit", + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "id": "aotc_abolition", + "issue": "PolicyEngine/populace#253", + "min_abs_effect": 100000000.0, + "name": "American Opportunity Tax Credit abolition", + "parameter_changes": { + "gov.irs.credits.education.american_opportunity_credit.abolition": { + "2024-01-01.2100-12-31": true + } + }, + "period": 2024, + "reason": "Abolishing the American Opportunity Tax Credit sets the credit to zero, so baseline-minus-reform AOTC must be positive. With qualified tuition or any of the five affirmative AOTC factual inputs absent or degenerate, the baseline credit is a structural zero and the abolition scores exactly $0." + }, + { + "binding_inputs": [ + "traditional_401k_contributions_desired", + "roth_401k_contributions_desired", + "traditional_ira_contributions_desired", + "roth_ira_contributions_desired", + "self_employed_pension_contributions_desired" + ], + "budget_measure": "savers_credit", + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "id": "savers_credit_abolition", + "issue": "PolicyEngine/populace#278", + "min_abs_effect": 100000000.0, + "name": "Retirement Saver's Credit abolition", + "parameter_changes": { + "gov.irs.credits.retirement_saving.contributions_cap": { + "2024-01-01.2100-12-31": 0 + } + }, + "period": 2024, + "reason": "Setting the Saver's Credit contribution cap to zero abolishes the credit, so baseline-minus-reform Saver's Credit must be positive. PolicyEngine-US builds qualified contributions from the realized forms of all five desired retirement-contribution inputs. If the desired-input family is absent or degenerate, the baseline credit is a structural zero and abolition scores $0." + }, + { + "binding_inputs": [ + "qualified_reit_and_ptp_income" + ], + "budget_measure": "qualified_business_income_deduction", + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "id": "qbi_reit_ptp_rate_abolition", + "issue": "PolicyEngine/populace#298", + "min_abs_effect": 1000000.0, + "name": "Section 199A qualified REIT/PTP component abolition", + "parameter_changes": { + "gov.irs.deductions.qbi.max.reit_ptp_rate": { + "2024-01-01.2100-12-31": 0 + } + }, + "period": 2024, + "reason": "Setting only the qualified REIT/PTP component rate to zero removes that component from the Section 199A deduction, so baseline-minus-reform QBID must be positive. Without populated qualified_reit_and_ptp_income the change is a structural zero." + }, + { + "binding_inputs": [ + "w2_wages_from_qualified_business", + "unadjusted_basis_qualified_property" + ], + "budget_measure": "qualified_business_income_deduction", + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "id": "qbi_wage_property_guardrails_zeroed", + "issue": "PolicyEngine/populace#298", + "min_abs_effect": 1000000.0, + "name": "Section 199A W-2 wage and UBIA guardrails zeroed", + "parameter_changes": { + "gov.irs.deductions.qbi.max.business_property.rate": { + "2024-01-01.2100-12-31": 0 + }, + "gov.irs.deductions.qbi.max.w2_wages.alt_rate": { + "2024-01-01.2100-12-31": 0 + }, + "gov.irs.deductions.qbi.max.w2_wages.rate": { + "2024-01-01.2100-12-31": 0 + } + }, + "period": 2024, + "reason": "Zeroing all W-2 wage and UBIA cap rates tightens the Section 199A deduction for high-income qualified businesses, so baseline-minus-reform QBID must be positive. If the total W-2 and UBIA inputs are absent, both baseline and reform guardrails are zero and the change is a structural zero. The archived all-or-nothing SSTB routing leaves its SSTB-allocable copies to the hard signal gate rather than overclaiming reform coverage." + }, + { + "binding_inputs": [ + "farm_operations_income" + ], + "budget_measure": "qualified_business_income_deduction", + "effect_direction": "baseline_minus_reform", + "expected_sign": "negative", + "id": "qbi_farm_operations_income_exclusion", + "issue": "PolicyEngine/populace#298", + "min_abs_effect": 1000000.0, + "name": "Exclude farm-operations income from Section 199A QBI", + "parameter_changes": { + "gov.irs.deductions.qbi.income_definition": { + "2026-01-01.2026-12-31": [ + "self_employment_income", + "partnership_s_corp_income", + "farm_rent_income", + "rental_income", + "estate_income" + ] + } + }, + "period": 2026, + "reason": "Removing only farm_operations_income from the 2026 Section 199A income definition isolates the restored signed Schedule F leaf. The real staged candidate is loss-heavy, so excluding it raises QBID and baseline-minus-reform is negative (-$4.16M). Without farm_operations_income the reform is a structural zero." + }, + { + "binding_inputs": [ + "farm_rent_income" + ], + "budget_measure": "qualified_business_income_deduction", + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "id": "qbi_farm_rent_income_exclusion", + "issue": "PolicyEngine/populace#298", + "min_abs_effect": 1000000.0, + "name": "Exclude farm-rent income from Section 199A QBI", + "parameter_changes": { + "gov.irs.deductions.qbi.income_definition": { + "2026-01-01.2026-12-31": [ + "self_employment_income", + "partnership_s_corp_income", + "farm_operations_income", + "rental_income", + "estate_income" + ] + } + }, + "period": 2026, + "reason": "Removing only farm_rent_income from the 2026 Section 199A income definition isolates the restored signed E27200 leaf. The real staged candidate produces +$9.14M baseline-minus-reform QBID. Without farm_rent_income the reform is a structural zero." + }, + { + "binding_inputs": [ + "domestic_production_ald" + ], + "budget_measure": "income_tax", + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "id": "domestic_production_ald_reactivation", + "issue": "PolicyEngine/populace#298", + "min_abs_effect": 1000000.0, + "name": "Former Section 199 domestic-production deduction reactivation", + "parameter_changes": { + "gov.irs.ald.deductions": { + "2024-01-01.2024-12-31": [ + "loss_ald", + "self_employment_tax_ald", + "student_loan_interest_ald", + "early_withdrawal_penalty", + "alimony_expense_ald", + "educator_expense", + "health_savings_account_ald", + "self_employed_health_insurance_ald", + "self_employed_pension_contribution_ald", + "traditional_ira_contributions", + "qualified_adoption_assistance_expense", + "us_bonds_for_higher_ed", + "specified_possession_income", + "puerto_rico_income", + "domestic_production_ald" + ] + } + }, + "period": 2024, + "reason": "PolicyEngine-US 1.764.6 excludes the former Section 199 deduction from current-law above-the-line deductions. This probe preserves the exact 2024 list and adds only domestic_production_ald, so baseline-minus-reform income tax must be positive. Without the restored E03240 input, reactivation is a structural zero." + }, + { + "binding_inputs": [ + "investment_income_elected_form_4952" + ], + "budget_measure": "income_tax", + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "id": "form_4952_election_neutralization", + "issue": "PolicyEngine/populace#274", + "min_abs_effect": 1000000.0, + "name": "Form 4952 elected investment income neutralization", + "neutralized_variable": "investment_income_elected_form_4952", + "parameter_changes": {}, + "period": 2024, + "reason": "PolicyEngine-US subtracts the tax-unit sum of investment_income_elected_form_4952 from net capital gain. Neutralizing only that leaf increases preferential net capital gain and lowers income tax, so baseline-minus-reform income tax must be positive. Without the restored E58990 input the neutralization is a structural zero." + }, + { + "binding_inputs": [ + "salt_refund_income" + ], + "budget_measure": "state_income_tax", + "effect_direction": "baseline_minus_reform", + "expected_sign": "negative", + "id": "salt_refund_income_neutralization", + "issue": "PolicyEngine/populace#38", + "min_abs_effect": 1000000.0, + "name": "State and local tax refund income neutralization", + "neutralized_variable": "salt_refund_income", + "parameter_changes": {}, + "period": 2024, + "reason": "PolicyEngine-US 1.764.6 includes salt_refund_income in the South Carolina, Idaho, and West Virginia subtraction lists. Neutralizing only that leaf removes the state subtraction and raises state income tax, so baseline-minus-reform state income tax must be negative. Without the restored E00700 input the neutralization is a structural zero." + }, + { + "binding_inputs": [ + "long_term_capital_gains_on_collectibles" + ], + "budget_measure": "income_tax", + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "id": "collectibles_gain_neutralization", + "issue": "PolicyEngine/populace#274", + "min_abs_effect": 1000000.0, + "name": "Long-term collectibles gain neutralization", + "neutralized_variable": "long_term_capital_gains_on_collectibles", + "parameter_changes": {}, + "period": 2024, + "reason": "PolicyEngine-US includes collectibles in capital_gains_28_percent_rate_gain. Neutralizing only the E24518 memo leaf reclassifies those gains from the special 28-percent bucket to the ordinary preferential capital-gain schedule, so baseline-minus-reform income tax must be positive. Without the restored leaf the neutralization is a structural zero." + }, + { + "binding_inputs": [ + "unrecaptured_section_1250_gain" + ], + "budget_measure": "income_tax", + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "id": "unrecaptured_section_1250_gain_neutralization", + "issue": "PolicyEngine/populace#274", + "min_abs_effect": 1000000.0, + "name": "Unrecaptured section 1250 gain neutralization", + "neutralized_variable": "unrecaptured_section_1250_gain", + "parameter_changes": {}, + "period": 2024, + "reason": "PolicyEngine-US taxes the E24515 memo leaf at the special unrecaptured-section-1250 rate. Neutralizing only that leaf reclassifies the same net gain onto the ordinary preferential capital-gain schedule, so baseline-minus-reform income tax must be positive. Without the restored leaf the neutralization is a structural zero." + }, + { + "binding_inputs": [ + "child_support_received" + ], + "budget_measure": "snap", + "effect_direction": "reform_minus_baseline", + "expected_sign": "positive", + "id": "child_support_received_snap_exclusion", + "issue": "PolicyEngine/populace#32", + "min_abs_effect": 1000000.0, + "name": "Exclude child-support receipts from SNAP unearned income", + "parameter_changes": { + "gov.usda.snap.income.sources.unearned": { + "2024-01-01.2024-12-31": [ + "ssi", + "tanf", + "general_assistance", + "pension_income", + "veterans_benefits", + "unemployment_compensation", + "disability_benefits", + "workers_compensation", + "social_security", + "retirement_distributions", + "rental_income", + "alimony_income", + "financial_assistance", + "survivor_benefits", + "dividend_income", + "interest_income", + "miscellaneous_income" + ] + } + }, + "period": 2024, + "reason": "Removing only child_support_received from SNAP unearned-income sources lowers countable income and must increase SNAP for some recipients. Without the measured/QRF child-support receipt leaf, the source-list reform is a structural zero." + }, + { + "binding_inputs": [ + "child_support_expense" + ], + "budget_measure": "snap", + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "id": "child_support_expense_snap_deduction_abolition", + "issue": "PolicyEngine/populace#32", + "min_abs_effect": 1000000.0, + "name": "Abolish the SNAP child-support expense deduction", + "parameter_changes": { + "gov.usda.snap.income.deductions.allowed": { + "2024-01-01.2024-12-31": [ + "snap_standard_deduction", + "snap_earned_income_deduction", + "snap_dependent_care_deduction", + "snap_excess_medical_expense_deduction", + "snap_excess_shelter_expense_deduction" + ] + } + }, + "period": 2024, + "reason": "Removing only snap_child_support_deduction raises countable net income and must reduce SNAP in states that take the expense as a net-income deduction. Without the measured/QRF positive expense leaf, abolition is a structural zero." + }, + { + "binding_inputs": [ + "disability_benefits" + ], + "budget_measure": "snap", + "effect_direction": "reform_minus_baseline", + "expected_sign": "positive", + "id": "disability_benefits_snap_exclusion", + "issue": "PolicyEngine/populace#38", + "min_abs_effect": 1000000.0, + "name": "Exclude disability benefits from SNAP unearned income", + "parameter_changes": { + "gov.usda.snap.income.sources.unearned": { + "2024-01-01.2024-12-31": [ + "ssi", + "tanf", + "general_assistance", + "pension_income", + "veterans_benefits", + "unemployment_compensation", + "workers_compensation", + "social_security", + "retirement_distributions", + "rental_income", + "child_support_received", + "alimony_income", + "financial_assistance", + "survivor_benefits", + "dividend_income", + "interest_income", + "miscellaneous_income" + ] + } + }, + "period": 2024, + "reason": "Removing only disability_benefits from SNAP unearned-income sources lowers countable income and must increase SNAP for some recipients. Without the measured/QRF non-workers-compensation benefit leaf, the source-list reform is a structural zero." + }, + { + "binding_inputs": [ + "workers_compensation" + ], + "budget_measure": "snap", + "effect_direction": "reform_minus_baseline", + "expected_sign": "positive", + "id": "workers_compensation_snap_exclusion", + "issue": "PolicyEngine/populace#32", + "min_abs_effect": 10000000.0, + "name": "Exclude workers' compensation from SNAP unearned income", + "parameter_changes": { + "gov.usda.snap.income.sources.unearned": { + "2024-01-01.2024-12-31": [ + "ssi", + "tanf", + "general_assistance", + "pension_income", + "veterans_benefits", + "unemployment_compensation", + "disability_benefits", + "social_security", + "retirement_distributions", + "rental_income", + "child_support_received", + "alimony_income", + "financial_assistance", + "survivor_benefits", + "dividend_income", + "interest_income", + "miscellaneous_income" + ] + } + }, + "period": 2024, + "reason": "Removing only workers_compensation from SNAP unearned-income sources lowers countable income and must increase SNAP for some recipients. A production-ingredient 30,000-household smoke scored +$28.26M reform-minus-baseline; without the measured WC_VAL carry and PUF-half QRF, the source-list reform is a structural zero." + }, + { + "binding_inputs": [ + "would_claim_wic" + ], + "budget_measure": "wic", + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "id": "wic_claim_neutralization", + "issue": "PolicyEngine/populace#312", + "min_abs_effect": 25000000.0, + "name": "WIC claim neutralization", + "neutralized_variable": "would_claim_wic", + "parameter_changes": {}, + "period": 2024, + "reason": "PolicyEngine-US multiplies each eligible person's monthly WIC food package by would_claim_wic. A 6,000-household production-ingredient smoke with the FNS category-rate stage scored +$57.19M baseline-minus-neutralized; without the restored claim surface the probe is a structural zero." + }, + { + "binding_inputs": [ + "educator_expense" + ], + "budget_measure": "income_tax", + "effect_direction": "reform_minus_baseline", + "expected_sign": "positive", + "id": "educator_expense_ald_abolition", + "issue": "PolicyEngine/populace#32", + "min_abs_effect": 1000000.0, + "name": "Abolish the educator-expense above-the-line deduction", + "parameter_changes": { + "gov.irs.ald.deductions": { + "2024-01-01.2024-12-31": [ + "loss_ald", + "self_employment_tax_ald", + "student_loan_interest_ald", + "early_withdrawal_penalty", + "alimony_expense_ald", + "health_savings_account_ald", + "self_employed_health_insurance_ald", + "self_employed_pension_contribution_ald", + "traditional_ira_contributions", + "qualified_adoption_assistance_expense", + "us_bonds_for_higher_ed", + "specified_possession_income", + "puerto_rico_income" + ] + } + }, + "period": 2024, + "reason": "Removing only educator_expense from the above-the-line deduction source list raises taxable income and must increase income tax for some filers. Without the restored PUF E03220 leaf, abolition is a structural zero." + }, + { + "binding_inputs": [ + "alimony_expense" + ], + "budget_measure": "income_tax", + "effect_direction": "baseline_minus_reform", + "expected_sign": "negative", + "id": "alimony_expense_ald_abolition", + "issue": "PolicyEngine/populace#38", + "min_abs_effect": 1000000.0, + "name": "Alimony expense above-the-line deduction abolition", + "parameter_changes": { + "gov.irs.ald.alimony_expense.divorce_year_threshold[0].amount": { + "2024-01-01.2100-12-31": false + } + }, + "period": 2024, + "reason": "The retired export has no nondefault divorce_year input, so PolicyEngine-US applies its default year 0 through the first eligibility bracket. Setting that bracket's amount to false abolishes the alimony-expense above-the-line deduction on the release, so baseline-minus-reform income tax must be negative. With alimony_expense absent or degenerate, the abolition scores exactly $0." + }, + { + "binding_inputs": [ + "casualty_loss" + ], + "budget_measure": "income_tax", + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "id": "obbba_casualty_loss_limit", + "issue": "PolicyEngine/populace#32", + "min_abs_effect": 1000000.0, + "name": "OBBBA casualty-loss deduction reactivation", + "parameter_changes": { + "gov.irs.deductions.itemized.casualty.active": { + "2026-01-01.2026-12-31": true + } + }, + "period": 2026, + "reason": "Reactivating the casualty-loss deduction lowers income tax only for tax units with casualty_loss above the statutory AGI floor, so baseline-minus-reform income tax must be positive. With the casualty-loss input absent or degenerate, the reactivation scores exactly $0." + }, + { + "binding_inputs": [ + "unreimbursed_business_employee_expenses" + ], + "budget_measure": "income_tax", + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "id": "obbba_misc_itemized_deductions", + "issue": "PolicyEngine/populace#32", + "min_abs_effect": 100000000.0, + "name": "OBBBA miscellaneous-itemized deduction reactivation", + "parameter_changes": { + "gov.irs.deductions.itemized.misc.applies": { + "2026-01-01.2026-12-31": true + } + }, + "period": 2026, + "reason": "Reactivating the miscellaneous itemized deduction lowers income tax only for tax units with qualifying expenses above its AGI floor, so baseline-minus-reform income tax must be positive. Without unreimbursed_business_employee_expenses, the retired pipeline's only populated miscellaneous-expense input, the reactivation is a structural zero on the export." + }, + { + "binding_inputs": [ + "spm_unit_pre_subsidy_childcare_expenses" + ], + "budget_measure": "income_tax", + "effect_direction": "baseline_minus_reform", + "expected_sign": "negative", + "id": "obbba_cdcc", + "issue": "PolicyEngine/populace#278", + "min_abs_effect": 1000000.0, + "name": "OBBBA Child and Dependent Care Credit reversion", + "parameter_changes": { + "gov.irs.credits.cdcc.phase_out.amended_structure.applies": { + "2026-01-01.2026-12-31": false + }, + "gov.irs.credits.cdcc.phase_out.max": { + "2026-01-01.2026-12-31": 0.35 + }, + "gov.irs.credits.cdcc.phase_out.min": { + "2026-01-01.2026-12-31": 0.2 + } + }, + "period": 2026, + "reason": "Reverting the OBBBA CDCC enhancement raises income tax, so baseline-minus-reform income tax must be negative. With measured pre-subsidy childcare expenses absent or degenerate, no filer has qualifying care expenses and the reversion scores exactly $0." + }, + { + "binding_inputs": [ + "tip_income", + "treasury_tipped_occupation_code" + ], + "budget_measure": "income_tax", + "effect_direction": "baseline_minus_reform", + "expected_sign": "negative", + "id": "obbba_no_tax_on_tips", + "issue": "PolicyEngine/populace#38", + "min_abs_effect": 100000000.0, + "name": "OBBBA no-tax-on-tips deduction", + "parameter_changes": { + "gov.irs.deductions.tip_income.cap": { + "2026-01-01.2026-12-31": 0 + } + }, + "period": 2026, + "reason": "Setting the OBBBA tip-deduction cap to zero removes the deduction, so baseline-minus-reform income tax must be negative in 2026. With tip_income or the Treasury tipped-occupation code absent, qualified tip income is zero and the repeal scores exactly $0." + }, + { + "binding_inputs": [ + "fsla_overtime_premium" + ], + "budget_measure": "income_tax", + "effect_direction": "baseline_minus_reform", + "expected_sign": "negative", + "id": "obbba_no_tax_on_overtime", + "issue": "PolicyEngine/populace#242", + "min_abs_effect": 100000000.0, + "name": "OBBBA no-tax-on-overtime deduction", + "parameter_changes": { + "gov.irs.deductions.overtime_income.cap.HEAD_OF_HOUSEHOLD": { + "2026-01-01.2026-12-31": 0 + }, + "gov.irs.deductions.overtime_income.cap.JOINT": { + "2026-01-01.2026-12-31": 0 + }, + "gov.irs.deductions.overtime_income.cap.SEPARATE": { + "2026-01-01.2026-12-31": 0 + }, + "gov.irs.deductions.overtime_income.cap.SINGLE": { + "2026-01-01.2026-12-31": 0 + }, + "gov.irs.deductions.overtime_income.cap.SURVIVING_SPOUSE": { + "2026-01-01.2026-12-31": 0 + } + }, + "period": 2026, + "reason": "Setting every OBBBA overtime-deduction cap to zero removes the deduction, so reform income tax rises and baseline-minus-reform must be negative in 2026. With fsla_overtime_premium absent or degenerate, qualified overtime is zero and the repeal scores $0." + }, + { + "binding_inputs": [ + "qualified_passenger_vehicle_loan_interest" + ], + "budget_measure": "income_tax", + "effect_direction": "baseline_minus_reform", + "expected_sign": "negative", + "id": "obbba_auto_loan_interest", + "issue": "PolicyEngine/populace#252", + "min_abs_effect": 100000000.0, + "name": "OBBBA no-tax-on-auto-loan-interest deduction", + "parameter_changes": { + "gov.irs.deductions.auto_loan_interest.cap": { + "2026-01-01.2026-12-31": 0 + } + }, + "period": 2026, + "reason": "Setting the OBBBA auto-loan-interest deduction cap to zero removes the deduction, so reform income tax rises and baseline-minus-reform must be negative in 2026. With qualified passenger-vehicle loan interest absent or degenerate, the repeal scores exactly $0." + }, + { + "binding_inputs": [ + "household_vehicles_owned", + "household_vehicles_value" + ], + "budget_measure": "snap", + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "id": "tx_snap_additional_vehicle_exemption_abolition", + "issue": "PolicyEngine/populace#49", + "min_abs_effect": 1000000.0, + "name": "Texas SNAP additional-vehicle exemption abolition", + "parameter_changes": { + "gov.hhs.tanf.non_cash.tx_additional_vehicle_exemption": { + "2026-01-01.2100-12-31": 0 + } + }, + "period": 2026, + "reason": "Setting the Texas TANF non-cash additional-vehicle exemption to zero tightens the asset test used by Texas SNAP categorical eligibility, so baseline-minus-reform SNAP must be positive. PolicyEngine-US computes the exemption from both household vehicle count and value; if either restored SIPP vehicle input is absent or degenerate, this vehicle-specific reform loses its intended binding channel. A persisted 30,000-household Populace smoke scored +$3.58 million in 2026 SNAP." + }, + { + "binding_inputs": [ + "would_file_taxes_voluntarily" + ], + "budget_measure": "aca_ptc", + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "id": "voluntary_filing_aca_ptc_neutralization", + "issue": "PolicyEngine/populace#312", + "min_abs_effect": 100000000.0, + "name": "Voluntary tax filing ACA PTC neutralization", + "neutralized_variable": "would_file_taxes_voluntarily", + "parameter_changes": {}, + "period": 2024, + "reason": "PolicyEngine-US includes would_file_taxes_voluntarily in tax_unit_is_filer alongside required and credit filers. Neutralizing only the restored SIPP filing-response leaf therefore removes ACA premium tax credits from otherwise eligible voluntary filers, so baseline-minus-neutralized aca_ptc must be positive. With the filing leaf absent or degenerate, this isolated response channel scores exactly $0." + }, + { + "binding_inputs": [ + "pre_subsidy_rent" + ], + "budget_measure": "snap", + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "id": "pre_subsidy_rent_neutralization", + "issue": "PolicyEngine/populace#32", + "min_abs_effect": 1000000.0, + "name": "Pre-subsidy rent neutralization", + "neutralized_variable": "pre_subsidy_rent", + "parameter_changes": {}, + "period": 2024, + "reason": "PolicyEngine-US includes pre_subsidy_rent in the SNAP shelter deduction. Neutralizing only the restored ACS rent leaf therefore reduces SNAP; on the pinned retired small eCPS artifact under PolicyEngine-US 1.764.6, baseline-minus-neutralized SNAP is +$11.731 billion. Without nondefault ACS rent, that effect is a structural zero. The other source-mapped housing leaves are enforced by their exact ASEC mappings and signal gate; household tenure_type has no standalone PolicyEngine-US 1.764.6 formula consumer." + }, + { + "binding_inputs": [ + "takes_up_housing_assistance_if_eligible" + ], + "budget_measure": "housing_assistance", + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "id": "housing_assistance_take_up_neutralization", + "issue": "PolicyEngine/populace#312", + "min_abs_effect": 100000000.0, + "name": "Measured housing-assistance take-up neutralization", + "neutralized_variable": "takes_up_housing_assistance_if_eligible", + "parameter_changes": {}, + "period": 2024, + "reason": "PolicyEngine-US multiplies HUD HAP by the restored SPM-unit take-up leaf after eligibility. Populace keeps that leaf exactly equal to source-backed housing-assistance receipt, so neutralizing it must remove the assistance paid to measured/imputed recipients. A 6,000-household production-ingredient smoke scored $202.795 million baseline-minus-neutralized; the $100 million floor is below that observed subset effect but far above numerical noise. A default-only or absent carry makes the source reconciliation or this uniquely isolating probe fail." } ], "schema_version": 1, diff --git a/packages/populace-build/src/populace/build/us/source_stages.json b/packages/populace-build/src/populace/build/us/source_stages.json index b278790f..dbf4be55 100644 --- a/packages/populace-build/src/populace/build/us/source_stages.json +++ b/packages/populace-build/src/populace/build/us/source_stages.json @@ -13,7 +13,28 @@ "kind": "public_microdata", "format": "fixed_width_or_csv", "vintage": "2015", - "locator": "IRS SOI Public Use File" + "locator": "IRS SOI Public Use File, including state and local tax refunds E00700, educator expense E03220, alimony income E00800, alimony expense E03500, domestic-production deduction E03240, casualty-loss E20500, unreimbursed employee business expense E20400, farm operations E02100, farm rent E27200, Form 4952 elected investment income E58990, collectibles gain E24518, and unrecaptured section 1250 gain E24515" + }, + { + "kind": "versioned_derived_microdata", + "format": "hdf5_arrays", + "vintage": "2024 (2015-based)", + "locator": "release://policyengine/irs-soi-puf/1.8.0/puf_2024.h5; qbi_simulation_version=1 materialized at archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe" + }, + { + "kind": "archived_derivation_evidence", + "repository_owner": "PolicyEngine", + "repository_name_parts": ["policyengine-", "us-data"], + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "datasets", "puf", "puf.py"], + "lines": "654-708,804-850", + "outputs": [ + "long_term_capital_gains_on_collectibles", + "salt_refund_income", + "unrecaptured_section_1250_gain" + ], + "imputation_path_parts": ["policyengine_", "us_data", "calibration", "puf_impute.py"], + "imputation_lines": "90-198,940-1075" } ], "operations": [ @@ -27,7 +48,34 @@ "ordinary_dividend_source": "E00600", "qualified_dividend_source": "E00650", "qualified_dividend_output": "qualified_dividend_income", - "non_qualified_dividend_output": "non_qualified_dividend_income" + "non_qualified_dividend_output": "non_qualified_dividend_income", + "qualified_tuition_primary_source": "E03230", + "qualified_tuition_optional_source": "E87530", + "qualified_tuition_output": "qualified_tuition_expenses", + "alimony_income_source": "E00800", + "alimony_income_output": "alimony_income", + "alimony_expense_source": "E03500", + "alimony_expense_output": "alimony_expense", + "casualty_loss_source": "E20500", + "casualty_loss_output": "casualty_loss", + "domestic_production_ald_source": "E03240", + "domestic_production_ald_output": "domestic_production_ald", + "educator_expense_source": "E03220", + "educator_expense_output": "educator_expense", + "unreimbursed_business_employee_expenses_source": "E20400", + "unreimbursed_business_employee_expenses_output": "unreimbursed_business_employee_expenses", + "farm_operations_income_source": "E02100", + "farm_operations_income_output": "farm_operations_income", + "farm_rent_income_source": "E27200", + "farm_rent_income_output": "farm_rent_income", + "investment_income_elected_form_4952_source": "E58990", + "investment_income_elected_form_4952_output": "investment_income_elected_form_4952", + "salt_refund_income_source": "E00700", + "salt_refund_income_output": "salt_refund_income", + "collectibles_capital_gain_source": "E24518", + "collectibles_capital_gain_output": "long_term_capital_gains_on_collectibles", + "unrecaptured_section_1250_gain_source": "E24515", + "unrecaptured_section_1250_gain_output": "unrecaptured_section_1250_gain" }, { "kind": "disaggregate_aggregate_records", @@ -74,6 +122,8 @@ "tax_exempt_interest_income", "short_term_capital_gains", "long_term_capital_gains_before_response", + "long_term_capital_gains_on_collectibles", + "unrecaptured_section_1250_gain", "non_sch_d_capital_gains", "taxable_private_pension_income", "taxable_ira_distributions", @@ -81,16 +131,27 @@ "social_security_disability", "social_security_dependents", "social_security_survivors", + "alimony_income", + "alimony_expense", + "salt_refund_income", "charitable_cash_donations", "charitable_non_cash_donations", "real_estate_taxes", "home_mortgage_interest", + "investment_income_elected_form_4952", "student_loan_interest", + "educator_expense", + "qualified_tuition_expenses", + "casualty_loss", + "domestic_production_ald", + "unreimbursed_business_employee_expenses", "traditional_ira_contributions_desired", "self_employed_pension_contributions_desired", "rental_income", "estate_income", "farm_income", + "farm_operations_income", + "farm_rent_income", "miscellaneous_income", "partnership_income", "s_corp_income", @@ -101,7 +162,22 @@ "second_home_mortgage_interest", "first_home_mortgage_origination_year", "second_home_mortgage_origination_year", - "health_savings_account_ald" + "health_savings_account_ald", + "estate_income_would_be_qualified", + "farm_operations_income_would_be_qualified", + "farm_rent_income_would_be_qualified", + "partnership_s_corp_income_would_be_qualified", + "rental_income_would_be_qualified", + "self_employment_income_would_be_qualified", + "sstb_self_employment_income_would_be_qualified", + "business_is_sstb", + "qualified_bdc_income", + "qualified_reit_and_ptp_income", + "sstb_self_employment_income_before_lsr", + "sstb_unadjusted_basis_qualified_property", + "sstb_w2_wages_from_qualified_business", + "unadjusted_basis_qualified_property", + "w2_wages_from_qualified_business" ], "nonnegative_outputs": [ "employment_income_before_lsr", @@ -111,17 +187,28 @@ "non_qualified_dividend_income", "tax_exempt_interest_income", "long_term_capital_gains_before_response", + "long_term_capital_gains_on_collectibles", + "unrecaptured_section_1250_gain", "taxable_private_pension_income", "taxable_ira_distributions", "social_security_retirement", "social_security_disability", "social_security_dependents", "social_security_survivors", + "alimony_income", + "alimony_expense", + "salt_refund_income", "charitable_cash_donations", "charitable_non_cash_donations", "real_estate_taxes", "home_mortgage_interest", + "investment_income_elected_form_4952", "student_loan_interest", + "educator_expense", + "qualified_tuition_expenses", + "casualty_loss", + "domestic_production_ald", + "unreimbursed_business_employee_expenses", "traditional_ira_contributions_desired", "self_employed_pension_contributions_desired", "health_savings_account_ald", @@ -130,9 +217,507 @@ "first_home_mortgage_interest", "second_home_mortgage_interest", "first_home_mortgage_origination_year", - "second_home_mortgage_origination_year" + "second_home_mortgage_origination_year", + "qualified_bdc_income", + "qualified_reit_and_ptp_income", + "sstb_unadjusted_basis_qualified_property", + "sstb_w2_wages_from_qualified_business", + "unadjusted_basis_qualified_property", + "w2_wages_from_qualified_business" + ], + "notes": "PUF detail is a primary-source donor and must stay separate from benchmark comparisons. IRS disclosure aggregate rows are replaced from raw PUF aggregate totals before uprating; Forbes top-tail synthesis is not enabled. Qualified tuition reproduces the retired source transform as max(E03230, E87530), falling back to E03230 when the optional field is absent. Educator expense reproduces the retired direct mapping educator_expense = E03220 at archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe, datasets/puf/puf.py lines 636-649, with filer/spouse allocation by earnings share at lines 617-620 and export at lines 804-815; the processed 1.8.0 artifact carries the leaf. Populace places this PUF-only detail on its dedicated PUF support channel and sparsifies to the donor positive rate, rather than reproducing the retired both-half override. Form 4952 elected investment income reproduces the archived direct mapping investment_income_elected_form_4952 = E58990 at puf.py line 708, its export at lines 804-850, and its retired QRF target/override entries at calibration/puf_impute.py lines 90-198; the pinned processed 1.8.0 artifact carries nondefault source signal. SALT refund income reproduces the archived direct mapping salt_refund_income = E00700 at puf.py line 707, its export at lines 804-850, and its retired QRF target/override entries at calibration/puf_impute.py lines 90-198; the pinned processed 1.8.0 artifact carries nondefault source signal. The retired processed person array embodies a randomized EARNSPLIT filer/spouse allocation (puf.py lines 477-546 and 1513-1601), but the hermetic ASEC target has no EARNSPLIT. Populace therefore aggregates each processed donor leaf to its identifiable tax-unit total, places that total on the unit's first person, and sparsifies only the PUF support channel to the donor positive rate rather than inventing an earnings-based split or reproducing the retired both-half override. This within-unit placement is policy-neutral in PolicyEngine-US 1.764.6 because the relevant formulas sum these leaves to the tax unit. Alimony reproduces the retired direct PUF mappings alimony_income = E00800 and alimony_expense = E03500 at archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe, datasets/puf/puf.py lines 640-641, and is re-derived after aggregate-row disaggregation; reported ASEC OI_VAL with OI_OFF = 20 remains on the ASEC half while codes 20 (alimony) and 12 (strike benefits) are excluded from miscellaneous_income per datasets/cps/cps.py lines 1481-1492. Domestic production reproduces the retired direct mapping domestic_production_ald = E03240 at archived puf.py line 646, appears in the exported field list at lines 808-815, and is re-derived after aggregate-row disaggregation; it is the former Section 199 deduction, not Section 199A QBI. The retired weighted-QRF target and override lists include it at calibration/puf_impute.py lines 90-198, with the fit at lines 940-1075. Casualty loss reproduces the retired direct mapping casualty_loss = E20500 at the same commit, datasets/puf/puf.py lines 636-643, and is re-derived after aggregate-row disaggregation. Unreimbursed employee business expenses reproduce the archived E20400 proxy for all deductible miscellaneous expenses at the same commit, datasets/puf/puf.py lines 663-665, and are likewise re-derived after disaggregation. Signed farm business inputs reproduce farm_operations_income = E02100 and farm_rent_income = E27200 at archived puf.py lines 636-704 and are re-derived after disaggregation without clipping; ASEC farm_operations_income = FRSE_VAL at archived cps.py lines 1363-1382 remains measured on that channel, while PUF farm_income = T27800 is a distinct Schedule J leaf and a forbidden substitute. The retired farm QRF overwrote both clone halves (calibration/puf_impute.py lines 513-672); Populace deliberately preserves measured ASEC operations income and replaces only the PUF support channel. Section 199A inputs are carried from the versioned processed PUF artifact rather than redrawn: archived puf.py lines 105-405 define the seeded qualification, investment-income, SSTB, W-2, and UBIA simulations; lines 748-787 enforce the all-or-nothing SSTB split; lines 860-879 export the 15 leaves; qbi_assumptions.yaml lines 1-118 pin the assumptions; calibration/puf_impute.py lines 90-198 and 940-1075 pin the retired QRF targets and fit. Populace source-aligns boolean person counts and restores the archived split/exposure identities after its weighted PUF QRF." + }, + { + "stage": "education_inputs", + "survey": "IRS PUF 2015 (uprated) + Census CPS ASEC", + "source": "https://www.irs.gov/statistics/soi-tax-stats-individual-public-use-microdata-files", + "grain": "person", + "artifacts": [ + { + "kind": "public_microdata", + "format": "fixed_width_or_csv", + "vintage": "2015", + "source": "https://www.irs.gov/statistics/soi-tax-stats-individual-public-use-microdata-files", + "locator": "IRS SOI Public Use File E03230 and optional E87530 tuition fields" + }, + { + "kind": "public_microdata", + "format": "census_asec", + "vintage": "build_year", + "source": "https://www.census.gov/programs-surveys/cps/data/datasets.html", + "locator": "Census CPS ASEC person ED_VAL educational-assistance field" + } + ], + "operations": [ + { + "kind": "read_table", + "table": "person", + "weight": "person_weight" + }, + { + "kind": "derive_education_inputs" + } + ], + "outputs": [ + "qualified_tuition_expenses", + "educational_assistance", + "is_pursuing_credential_for_american_opportunity_credit", + "attends_eligible_educational_institution_for_american_opportunity_credit", + "is_enrolled_at_least_half_time_for_american_opportunity_credit", + "has_american_opportunity_credit_1098_t_or_exception", + "has_american_opportunity_credit_institution_ein" + ], + "nonnegative_outputs": [ + "qualified_tuition_expenses", + "educational_assistance" + ], + "notes": "Retired published-eCPS construction: the PUF publication path drops the reported american_opportunity_credit tax output before AOTC factual-input postprocessing, so archived extended_cps.py takes its positive-qualified-tuition fallback and sets all five affirmative AOTC inputs on that mask (publication candidate 9f8db56e9ed1ca25f0ed16285fcc928343d2e474). educational_assistance carries directly from measured ASEC ED_VAL (archived cps.py at 9a823603e6b5fb916d65ec45d74c9c7eb0043db1). The pinned export independently confirms that the five flags and tuition share one support mask." + }, + { + "stage": "retirement_contributions", + "survey": "Census CPS ASEC + published retirement-contribution shares", + "source": "https://www.census.gov/programs-surveys/cps.html", + "grain": "person", + "artifacts": [ + { + "kind": "public_microdata", + "format": "census_asec", + "vintage": "build_year", + "source": "https://www.census.gov/programs-surveys/cps/data/datasets.html", + "locator": "Census CPS ASEC person RETCB_VAL, WSAL_VAL, and SEMP_VAL fields" + }, + { + "kind": "administrative_table", + "format": "published_estimates", + "vintage": "2022-2024", + "source": "https://www.irs.gov/statistics/soi-tax-stats-accumulation-and-distribution-of-individual-retirement-arrangements", + "locator": "IRS SOI IRA contribution totals and Publication 1304 Table 1.4 Keogh-plan payments" + }, + { + "kind": "administrative_table", + "format": "published_time_series", + "vintage": "2023", + "source": "https://fred.stlouisfed.org/series/Y351RC1A027NBEA", + "locator": "BEA/FRED employee and employer defined-contribution pension-plan contribution series Y351RC1A027NBEA and W351RC0A144NBEA" + }, + { + "kind": "research_report", + "format": "pdf", + "vintage": "2024", + "source": "https://corporate.vanguard.com/content/dam/corp/research/pdf/how_america_saves_report_2024.pdf", + "locator": "Vanguard How America Saves 2024 Roth defined-contribution participation evidence" + }, + { + "kind": "research_report", + "format": "web", + "vintage": "2024", + "source": "https://www.psca.org/news/psca-news/2024/12/401k-savings-and-participation-rates-rise/", + "locator": "Plan Sponsor Council of America 67th Annual Survey Roth participation evidence" + } + ], + "operations": [ + { + "kind": "read_table", + "table": "person", + "weight": "person_weight" + }, + { + "kind": "derive_retirement_contributions", + "se_pension_share": 0.046, + "dc_share_of_remainder": 0.908, + "roth_dc_share": 0.15, + "traditional_ira_share": 0.392 + }, + { + "kind": "impute_retirement_contributions_to_puf_support", + "predictors": [ + "age", + "is_male", + "has_esi", + "tax_unit_is_joint", + "tax_unit_count_dependents", + "employment_income", + "self_employment_income", + "social_security" + ], + "max_train_samples": 5000, + "n_estimators": 100, + "seed_from_build_config": true, + "weight": "person_weight" + } + ], + "outputs": [ + "traditional_401k_contributions_desired", + "roth_401k_contributions_desired", + "traditional_ira_contributions_desired", + "roth_ira_contributions_desired", + "self_employed_pension_contributions_desired" + ], + "nonnegative_outputs": [ + "traditional_401k_contributions_desired", + "roth_401k_contributions_desired", + "traditional_ira_contributions_desired", + "roth_ira_contributions_desired", + "self_employed_pension_contributions_desired" + ], + "notes": "Exact port of the retired published-eCPS allocation at 9a823603e6b5fb916d65ec45d74c9c7eb0043db1 (datasets/cps/cps.py lines 1500-1552; imputation_parameters.yaml lines 24-55). Measured ASEC RETCB_VAL is allocated across self-employed pension, defined-contribution, and IRA pools only where ASEC reports the corresponding earned-income source. After PUF support expansion, the retired second-stage treatment (extended_cps.py lines 639-780 and 1014-1073) QRF-imputes all five leaves onto the PUF half from the ASEC rows and PUF-imputed income. PolicyEngine-US applies statutory limits to these five desired-contribution leaves; this source stage does not cap or synthesize the measured total." + }, + { + "stage": "childcare_inputs", + "survey": "Census CPS ASEC", + "source": "https://www.census.gov/programs-surveys/cps.html", + "grain": "person", + "artifacts": [ + { + "kind": "public_microdata", + "format": "census_asec", + "vintage": "build_year", + "source": "https://www.census.gov/programs-surveys/cps/data/datasets.html", + "locator": "Census CPS ASEC replicated SPM-unit SPM_CHILDCAREXPNS field" + } + ], + "operations": [ + { + "kind": "read_table", + "table": "person", + "weight": "person_weight" + }, + { + "kind": "derive_childcare_inputs" + }, + { + "kind": "impute_childcare_to_puf_support", + "predictors": [ + "age", + "is_male", + "has_esi", + "tax_unit_is_joint", + "tax_unit_count_dependents", + "employment_income", + "self_employment_income", + "social_security" + ], + "max_train_samples": 5000, + "n_estimators": 100, + "seed_from_build_config": true, + "weight": "person_weight", + "reduction": "value_from_first_person" + } + ], + "outputs": [ + "spm_unit_pre_subsidy_childcare_expenses" + ], + "nonnegative_outputs": [ + "spm_unit_pre_subsidy_childcare_expenses" + ], + "notes": "Measured ASEC SPM_CHILDCAREXPNS carries directly to spm_unit_pre_subsidy_childcare_expenses (archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe, datasets/cps/cps.py lines 1612-1622). After support expansion, the archived second-stage method uses the eight declared predictors and at most 5,000 ASEC rows to QRF-impute the PUF half, then reduces non-person output with value_from_first_person (extended_cps.py lines 234-248, 639-739, and 1014-1069). Populace deliberately strengthens the archived unweighted fit by using typed person weights." + }, + { + "stage": "energy_subsidy", + "survey": "Census CPS ASEC", + "source": "https://www.census.gov/programs-surveys/cps.html", + "grain": "person", + "artifacts": [ + { + "kind": "public_microdata", + "format": "census_asec", + "vintage": "build_year", + "source": "https://www.census.gov/programs-surveys/cps/data/datasets.html", + "locator": "Census CPS ASEC replicated SPM-unit SPM_ENGVAL field" + } + ], + "operations": [ + { + "kind": "read_table", + "table": "person", + "weight": "person_weight" + }, + { + "kind": "derive_energy_subsidy" + }, + { + "kind": "impute_energy_subsidy_to_puf_support", + "predictors": [ + "age", + "is_male", + "has_esi", + "tax_unit_is_joint", + "tax_unit_count_dependents", + "employment_income", + "self_employment_income", + "social_security" + ], + "max_train_samples": 5000, + "n_estimators": 100, + "seed_from_build_config": true, + "weight": "person_weight", + "reduction": "value_from_first_person" + } + ], + "outputs": [ + "spm_unit_energy_subsidy" + ], + "nonnegative_outputs": [ + "spm_unit_energy_subsidy" + ], + "notes": "Measured ASEC SPM_ENGVAL carries directly to spm_unit_energy_subsidy (archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe, datasets/cps/cps.py lines 1612-1622). After support expansion, the archived second-stage method includes this leaf in CPS_ONLY_IMPUTED_VARIABLES, uses the eight declared predictors and at most 5,000 ASEC rows to QRF-impute the PUF half, then reduces this non-person output with value_from_first_person (extended_cps.py lines 135-194, 234-248, 639-739, and 1014-1073). Populace deliberately strengthens the archived unweighted fit by using typed person weights." + }, + { + "stage": "child_support_inputs", + "survey": "Census CPS ASEC", + "source": "https://www.census.gov/programs-surveys/cps.html", + "grain": "person", + "artifacts": [ + { + "kind": "public_microdata", + "format": "census_asec", + "vintage": "build_year and prior year", + "source": "https://www.census.gov/programs-surveys/cps/data/datasets.html", + "locator": "Census CPS ASEC person CSP_VAL child support received and CHSP_VAL annual child support paid fields" + } + ], + "operations": [ + { + "kind": "read_table", + "table": "person", + "weight": "person_weight" + }, + { + "kind": "derive_child_support_inputs", + "received_source": "CSP_VAL", + "received_output": "child_support_received", + "expense_source": "CHSP_VAL", + "expense_output": "child_support_expense" + }, + { + "kind": "impute_child_support_to_puf_support", + "predictors": [ + "age", + "is_male", + "has_esi", + "tax_unit_is_joint", + "tax_unit_count_dependents", + "employment_income", + "self_employment_income", + "social_security" + ], + "max_train_samples": 5000, + "n_estimators": 100, + "seed_from_build_config": true, + "weight": "person_weight" + } + ], + "outputs": [ + "child_support_received", + "child_support_expense" + ], + "nonnegative_outputs": [ + "child_support_received", + "child_support_expense" + ], + "notes": "Measured annual ASEC CSP_VAL carries directly to child_support_received and positive annual CHSP_VAL carries directly to child_support_expense (archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe, datasets/cps/cps.py lines 1493-1496 and 1572-1574). After PUF tax-detail imputation, one joint person-grain QRF uses the exact archived eight-predictor subset and at most 5,000 ASEC training people to replace both leaves on the PUF support half (extended_cps.py lines 135-194, 234-248, 639-745, and 1014-1076). Populace deliberately strengthens the archived unweighted fit by using typed person weights." + }, + { + "stage": "disability_benefits_input", + "survey": "Census CPS ASEC", + "source": "https://www.census.gov/programs-surveys/cps.html", + "grain": "person", + "artifacts": [ + { + "kind": "public_microdata", + "format": "census_asec", + "vintage": "build_year and prior year", + "source": "https://www.census.gov/programs-surveys/cps/data/datasets.html", + "locator": "Census CPS ASEC person DIS_VAL1, DIS_SC1, DIS_VAL2, and DIS_SC2 fields" + } + ], + "operations": [ + { + "kind": "read_table", + "table": "person", + "weight": "person_weight" + }, + { + "kind": "derive_disability_benefits", + "first_amount_source": "DIS_VAL1", + "first_code_source": "DIS_SC1", + "second_amount_source": "DIS_VAL2", + "second_code_source": "DIS_SC2", + "workers_compensation_code": 1, + "output": "disability_benefits" + }, + { + "kind": "impute_disability_benefits_to_puf_support", + "predictors": [ + "age", + "is_male", + "has_esi", + "tax_unit_is_joint", + "tax_unit_count_dependents", + "employment_income", + "self_employment_income", + "social_security" + ], + "max_train_samples": 5000, + "n_estimators": 100, + "seed_from_build_config": true, + "weight": "person_weight" + } + ], + "outputs": [ + "disability_benefits" + ], + "nonnegative_outputs": [ + "disability_benefits" + ], + "notes": "The retired ASEC derivation sums annual DIS_VAL1 and DIS_VAL2 only when the corresponding DIS_SC source code is not 1 (workers' compensation), keeping Social Security disability and workers' compensation in their separate leaves (archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe, datasets/cps/cps.py lines 1561-1571). After PUF tax-detail imputation, the CPS-only output is replaced on the PUF support half using the archived eight predictors and at most 5,000 ASEC training people (extended_cps.py lines 135-194, 234-248, 639-745, and 1014-1076). The archive did not explicitly pin the QRF tree count; Populace pins 100 trees and typed person weights as reproducibility hardening." + }, + { + "stage": "workers_compensation_input", + "survey": "Census CPS ASEC", + "source": "https://www.census.gov/programs-surveys/cps.html", + "grain": "person", + "artifacts": [ + { + "kind": "public_microdata", + "format": "census_asec", + "vintage": "2022-2024 pooled", + "source": "https://www.census.gov/programs-surveys/cps/data/datasets.html", + "locator": "Census CPS ASEC person annual workers' compensation amount WC_VAL" + }, + { + "kind": "archived_derivation_evidence", + "repository_owner": "PolicyEngine", + "repository_name_parts": ["policyengine-", "us-data"], + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "datasets", "cps", "cps.py"], + "lines": "1559-1571", + "source_declaration_path_parts": ["policyengine_", "us_data", "datasets", "cps", "census_cps.py"], + "source_declaration_lines": "306-381", + "puf_output_path_parts": ["policyengine_", "us_data", "datasets", "cps", "extended_cps.py"], + "puf_output_lines": "135-194", + "puf_predictor_lines": "234-248", + "puf_imputation_lines": "639-745,1014-1073" + } + ], + "operations": [ + { + "kind": "read_table", + "table": "person", + "weight": "person_weight" + }, + { + "kind": "derive_workers_compensation", + "source": "WC_VAL", + "output": "workers_compensation" + }, + { + "kind": "impute_workers_compensation_to_puf_support", + "predictors": [ + "age", + "is_male", + "has_esi", + "tax_unit_is_joint", + "tax_unit_count_dependents", + "employment_income", + "self_employment_income", + "social_security" + ], + "max_train_samples": 5000, + "n_estimators": 100, + "seed_from_build_config": true, + "weight": "person_weight" + } + ], + "outputs": [ + "workers_compensation" + ], + "nonnegative_outputs": [ + "workers_compensation" + ], + "notes": "Exact annual person-level WC_VAL carry from archived cps.py lines 1559-1571; code-1 DIS_VAL slots remain excluded from the separate disability_benefits leaf rather than added here. The CPS-only workers_compensation output is replaced on the PUF support half with the archived eight predictors and at most 5,000 ASEC training people (extended_cps.py lines 135-194, 234-248, 639-745, and 1014-1073). Populace pins 100 trees and typed person weights as reproducibility hardening." + }, + { + "stage": "weeks_unemployed_input", + "survey": "Census CPS ASEC", + "source": "https://www.census.gov/programs-surveys/cps.html", + "grain": "person", + "artifacts": [ + { + "kind": "public_microdata", + "format": "zip_csv", + "vintage": "2023 ASEC / 2022 income reference year", + "locator": "https://www2.census.gov/programs-surveys/cps/datasets/2023/march/asecpub23csv.zip", + "sha256": "d2e000250782adfbdd7f29c82b66d866591a30f0d330496698ec19f9c784ce11", + "size_bytes": 150165063, + "member": "pppub23.csv", + "member_size_bytes": 281065733, + "member_crc32": "49c09e5f", + "member_sha256": "19b56537e50e7663f954361ef2bb5ce9cef8d9d45f156fe1a69a99b654198ffe" + }, + { + "kind": "official_data_dictionary", + "format": "pdf", + "vintage": "2023", + "locator": "https://www2.census.gov/programs-surveys/cps/datasets/2023/march/asec2023_ddl_pub_full.pdf", + "lines": "person LKWEEKS entry, position 305 length 2" + }, + { + "kind": "archived_derivation_evidence", + "repository_owner": "PolicyEngine", + "repository_name_parts": ["policyengine-", "us-data"], + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "datasets", "cps", "cps.py"], + "lines": "1438-1442", + "source_path_parts": ["policyengine_", "us_data", "datasets", "cps", "census_cps.py"], + "source_lines": "39-58,89-133,245,306-351", + "puf_path_parts": ["policyengine_", "us_data", "calibration", "puf_impute.py"], + "puf_lines": "703-785" + } + ], + "operations": [ + { + "kind": "read_table", + "table": "person", + "weight": "person_weight" + }, + { + "kind": "derive_weeks_unemployed", + "source": "LKWEEKS", + "niu_value": -1, + "niu_replacement": 0, + "output": "weeks_unemployed" + }, + { + "kind": "impute_weeks_unemployed_to_puf_support", + "predictors": [ + "age", + "is_male", + "tax_unit_is_joint", + "is_tax_unit_head", + "is_tax_unit_spouse", + "is_tax_unit_dependent", + "unemployment_compensation" + ], + "max_train_samples": 5000, + "n_estimators": 100, + "seed_from_build_config": true, + "weight": "person_weight", + "clip": [ + 0, + 52 + ], + "round": "nearest_integer", + "zero_unless": "unemployment_compensation > 0" + } + ], + "outputs": [ + "weeks_unemployed" + ], + "nonnegative_outputs": [ + "weeks_unemployed" ], - "notes": "PUF detail is a primary-source donor and must stay separate from benchmark comparisons. IRS disclosure aggregate rows are replaced from raw PUF aggregate totals before uprating; Forbes top-tail synthesis is not enabled." + "notes": "Restores the final retired eCPS person-year weeks_unemployed input. Archived cps.py lines 1438-1442 carries Census ASEC LKWEEKS directly, mapping only the historical -1 not-in-universe sentinel to zero; the SHA-pinned 2023 ASEC public-use archive supplies the omitted 2022-income-year source rows exactly. Its 146,133 person records match the locked 2022 H5 one-for-one on PH_SEQ, P_SEQ, A_LINENO, A_AGE, A_FNLWGT, UC_VAL, and WKSWORK. The source has 4,200 positive LKWEEKS rows, range 0--51, 3.1456094276% A_FNLWGT-weighted positive support, and 172,563,622.04 weighted person-weeks; Census allocation flags are retained because the retired direct carry did not filter them. Archived calibration/puf_impute.py lines 703-785 replaces only the PUF support half with a QRF using age, sex, three tax-unit roles, joint filing, and unemployment compensation, caps training at 5,000 ASEC people, clips predictions to 0--52, and forces predictions to zero without positive unemployment compensation. Populace pins 100 trees and typed person weights as reproducibility hardening, and rounds its interpolating QRF draws back to the source's integer-week domain; no 2022 source value is filled statistically." }, { "stage": "capital_gain_distributions", @@ -186,15 +771,30 @@ "artifacts": [ { "kind": "public_microdata", - "format": "sas_or_stata", + "format": "stata_in_zip", + "vintage": "2022", + "locator": "https://www.federalreserve.gov/econres/files/scfp2022s.zip", + "member": "rscfp2022.dta", + "sha256": "3bb4d890ae2463ff6039ec7692e375f544dd98a55a37ca2cb2340354b9cc9d80", + "member_sha256": "6b8dd2d935a76ed225ddebc80fb2db22a467f0c80d9a1acaa67b4584aa4bafd1" + }, + { + "kind": "public_microdata", + "format": "stata_in_zip", "vintage": "2022", - "locator": "Federal Reserve Survey of Consumer Finances public extract" + "locator": "https://www.federalreserve.gov/econres/files/scf2022s.zip", + "member": "p22i6.dta", + "expected_rows": 22975, + "integrity_note": "Full-file SHA-256 pending one network-enabled provisioning fetch; do not reuse the unrelated summary-extract hashes. Runtime validates the exact member plus a complete one-to-one (y1, yy1) implicate join." } ], "operations": [ { - "kind": "read_table", - "table": "scf_household", + "kind": "read_tables", + "tables": [ + "scf_summary_extract", + "scf_full_extract" + ], "weight": "wgt" }, { @@ -207,14 +807,50 @@ ], "scope": "not_applicable_or_missing_fields" }, + { + "kind": "join", + "left": "scf_summary_extract", + "right": "scf_full_extract", + "keys": [ + "y1", + "yy1" + ], + "validation": "one_to_one_all_five_implicates" + }, { "kind": "derive", "outputs": [ "auto_loan_interest", + "auto_loan_balance", + "net_worth_anchor", "net_worth_components", "financial_asset_components", "household_asset_components" - ] + ], + "net_worth_source_column": "networth", + "auto_loan_amount_columns": [ + "x2209", + "x2309", + "x2409", + "x7158" + ], + "auto_loan_rate_columns": [ + "x2219", + "x2319", + "x2419", + "x7170" + ], + "rate_divisor": 10000, + "floor_negative_source_values": true + }, + { + "kind": "derive", + "outputs": [ + "qualified_passenger_vehicle_loan_interest" + ], + "method": "expected_share_of_auto_loan_interest", + "annual_qualifying_loan_target": 6000000, + "source": "https://www.irs.gov/irb/2026-05_IRB" }, { "kind": "fit_weighted_qrf", @@ -227,13 +863,20 @@ "employment_income", "interest_dividend_income", "social_security_pension_income" - ] + ], + "net_worth_target": "networth", + "net_worth_signed": true }, { "kind": "head_carry", "from_entity": "household", "to_entity": "person", - "condition": "is_household_head" + "condition": "is_household_head", + "outputs": [ + "bank_account_assets", + "bond_assets", + "stock_assets" + ] }, { "kind": "support_clip", @@ -263,14 +906,418 @@ "scf_other_debt", "auto_loan_balance", "auto_loan_interest", + "qualified_passenger_vehicle_loan_interest", "bank_account_assets", "bond_assets", "stock_assets" ], "nonnegative_outputs": [ - "auto_loan_interest" + "auto_loan_balance", + "auto_loan_interest", + "qualified_passenger_vehicle_loan_interest" + ], + "notes": "The full SCF joins all five implicates to the summary extract on (y1, yy1). Archived commit 42ed5d45 utils/asset_imputation.py lines 15-19 maps summary-extract networth directly to the scf_net_worth QRF anchor; calibration/source_impute.py lines 1146-1164 fits that anchor with these eight predictors and calibration/source_impute.py lines 1324-1338 forces its construction-only balance-sheet components to sum back to it before exporting net_worth. Populace therefore persists that signed household anchor directly, including indebted households, without manufacturing or exporting the internal scf_* components. Legacy auto_loan_balance preserves the retired pipeline's name for original financed amounts, not remaining principal. Auto outputs and net_worth stay Household-grain; only the three SSI liquid-asset leaves are head-carried to Person. The OBBBA qualifying-interest proxy uses Treasury/IRS's approximately six million qualifying loans issued annually and is bounded by total auto-loan interest. Negative amount/rate codes are floored at the source boundary for the nonnegative auto outputs; signed net_worth is never clipped." + }, + { + "stage": "ssi_disability_criteria", + "survey": "Census SIPP", + "source": "https://www.census.gov/programs-surveys/sipp.html", + "grain": "person", + "artifacts": [ + { + "kind": "public_microdata", + "format": "pipe_delimited_csv", + "vintage": "2023", + "locator": "Census SIPP 2023 public-use file; immutable mirror revision 21280dca5995e978d706740a8a4b9b7860cfd7b6", + "sha256": "5c30439e365fc26483318ef61d1d8f4bb2f0e9d6bb47c22c06756a7698733ee2", + "size_bytes": 3726010471 + } + ], + "operations": [ + { + "kind": "read_table", + "table": "sipp_person", + "delimiter": "|", + "month_column": "MONTHCODE", + "month": 12, + "source_columns": [ + "AJOCDVAL", + "AJOCHKVAL", + "AJOGOVSVAL", + "AJOMCBDVAL", + "AJOMFVAL", + "AJOMMVAL", + "AJOSAVVAL", + "AJOSTVAL", + "AJSCDVAL", + "AJSCHKVAL", + "AJSGOVSVAL", + "AJSMCBDVAL", + "AJSMFVAL", + "AJSMMVAL", + "AJSSAVVAL", + "AJSSTVAL", + "AOCDVAL", + "AOCHKVAL", + "AOGOVSVAL", + "AOMCBDVAL", + "AOMFVAL", + "AOMMVAL", + "AOSAVVAL", + "AOSTVAL", + "ASSI_BRSN", + "ASSI_YRYN", + "EAMBULAT", + "ECOGNIT", + "EDISABL", + "EDISANY", + "EERRANDS", + "EHEARING", + "EHLTHCOND", + "EMS", + "ENJ_NOWRK3", + "ESEEING", + "ESELFCARE", + "ESEX", + "ESSI_BRSN", + "ESSRSN2YN", + "MONTHCODE", + "PNUM", + "RDIS", + "RDIS_ALT", + "RSSI_YRYN", + "SPANEL", + "SSUID", + "SWAVE", + "TAGE", + "TDIS10AMT", + "TDIS1AMT", + "TDIS2AMT", + "TDIS3AMT", + "TDIS4AMT", + "TDIS5AMT", + "TDIS6AMT", + "TDIS7AMT", + "TDIS8AMT", + "TDIS9AMT", + "TINC_BANK", + "TINC_BOND", + "TINC_RENT", + "TINC_STMF", + "TJB1_MSUM", + "TJB2_MSUM", + "TJB3_MSUM", + "TJB4_MSUM", + "TJB5_MSUM", + "TJB6_MSUM", + "TJB7_MSUM", + "TPTOTINC", + "TRETINCAMT", + "TSSSAMT", + "TVAL_BANK", + "TVAL_BOND", + "TVAL_STMF", + "WPFINWGT" + ] + }, + { + "kind": "fit_weighted_qrf", + "predictors": [ + "age", + "is_female", + "is_married", + "employment_income", + "interest_income", + "dividend_income", + "rental_income", + "bank_account_assets", + "stock_assets", + "bond_assets", + "count_under_18", + "difficulty_dressing_or_bathing", + "difficulty_hearing", + "difficulty_seeing", + "difficulty_doing_errands", + "difficulty_walking_or_climbing_stairs", + "difficulty_remembering_or_making_decisions", + "social_security_disability", + "has_disability_income" + ], + "target": "meets_ssi_disability_criteria", + "weight": "household_weight", + "training_candidate": "ssi_disability_training_candidate", + "label_source_columns": [ + "RSSI_YRYN", + "ESSI_BRSN" + ], + "label_allocation_columns": [ + "ASSI_YRYN", + "ASSI_BRSN" + ], + "max_train_samples": 20000, + "sample_with_replacement": true, + "training_sample_seed_name": "sipp_ssi_disability_model_training_sample", + "training_sample_seed": 8386123572872638692, + "n_estimators": 100, + "model_seed": 42, + "seed_from_build_config": false, + "postprediction_signal_predictors": [ + "difficulty_dressing_or_bathing", + "difficulty_hearing", + "difficulty_seeing", + "difficulty_doing_errands", + "difficulty_walking_or_climbing_stairs", + "difficulty_remembering_or_making_decisions", + "social_security_disability", + "has_disability_income" + ], + "preserve_under_65_asec_ssi_reporters": true + } + ], + "outputs": [ + "meets_ssi_disability_criteria" + ], + "notes": "Exact final retired eCPS SSI disability-criterion method at archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe. datasets/sipp/sipp.py lines 63-105 and 404-428 define the output, nineteen predictors, source columns, and allocation flags; lines 541-562 require observed receipt and, for recipients, observed benefit reason; lines 565-630 apply the source-time SSI resource, countable-income, and SGA candidate screen; lines 633-703 construct December labels and predictors; lines 724-795 require a measured difficulty, SSDI, or disability-income signal after prediction; and lines 892-949 perform the fixed weighted bootstrap and QRF fit. The immutable donor yields 9,346 candidates (577 positive, 8,769 negative), 88,690,359.47893329 of total source weight, and a 0.05566746987074994 weighted true share. Its archived fixed WPFINWGT bootstrap seed 8386123572872638692 resamples 9,346 rows with replacement (5,314 unique source rows; 524 positive) before an unweighted 100-tree MicroImpute QRF with model seed 42. datasets/cps/cps.py lines 2853-2886 and calibration/source_impute.py lines 869-990 predict the ASEC receiver and preserve direct under-65 ASEC SSI reporters. datasets/cps/extended_cps.py lines 140-170, 392-424, and 981-992 separately predict PUF-support people from their imputed income, asset, and disability-signal inputs; the ASEC reporter anchor is never copied to PUF support." + }, + { + "stage": "sipp_head_start", + "survey": "Census SIPP", + "source": "https://www.census.gov/programs-surveys/sipp.html", + "grain": "person", + "artifacts": [ + { + "kind": "public_microdata", + "format": "pipe_delimited_csv", + "vintage": "2023", + "locator": "Census SIPP 2023 public-use file; immutable mirror revision 21280dca5995e978d706740a8a4b9b7860cfd7b6", + "sha256": "5c30439e365fc26483318ef61d1d8f4bb2f0e9d6bb47c22c06756a7698733ee2", + "size_bytes": 3726010471 + }, + { + "kind": "official_data_dictionary", + "format": "pdf", + "vintage": "2023", + "locator": "https://www2.census.gov/programs-surveys/sipp/tech-documentation/data-dictionaries/2023/2023_SIPP_Data_Dictionary.pdf", + "lines": "pages 7, 1028, and 1031" + }, + { + "kind": "official_survey_instrument", + "format": "pdf", + "vintage": "2023", + "locator": "https://www2.census.gov/programs-surveys/sipp/questionnaires/2023/2023_SIPP_PU_Instrument_Specifications.pdf", + "lines": "page 742" + }, + { + "kind": "administrative_validation_only", + "format": "pdf", + "vintage": "2023-2024", + "locator": "https://headstart.gov/sites/default/files/pdf/service-snapshot-hs-2023-2024.pdf", + "measure": "cumulative enrollment by age; not used for calibration or assignment" + }, + { + "kind": "archived_derivation_evidence", + "repository_owner": "PolicyEngine", + "repository_name_parts": [ + "policyengine-", + "us-data" + ], + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": [ + "policyengine_", + "us_data", + "datasets", + "cps", + "cps.py" + ], + "lines": "582-583,640-648", + "rate_path_parts": [ + "policyengine_", + "us_data", + "parameters", + "take_up", + "head_start.yaml" + ], + "rate_lines": "1-10", + "randomness_path_parts": [ + "policyengine_", + "us_data", + "utils", + "randomness.py" + ], + "randomness_lines": "5-28" + } + ], + "operations": [ + { + "kind": "read_table", + "table": "sipp_person", + "delimiter": "|", + "month_column": "MONTHCODE", + "month": 12, + "source_columns": [ + "SSUID", + "PNUM", + "MONTHCODE", + "WPFINWGT", + "TAGE", + "ESEX", + "EED_SCRNR", + "AED_SCRNR", + "EEDGRADE", + "AEDGRADE", + "EEDEMONTH", + "AEDMONTH", + "EEDHEADST", + "AEDHEADST", + "TJB1_MSUM", + "TJB2_MSUM", + "TJB3_MSUM", + "TJB4_MSUM", + "TJB5_MSUM", + "TJB6_MSUM", + "TJB7_MSUM" + ] + }, + { + "kind": "fit_weighted_qrf", + "predictors": [ + "age", + "is_female", + "household_size", + "count_under_18", + "count_under_6", + "household_employment_income" + ], + "target": "takes_up_head_start_if_eligible", + "weight": "sipp_weight", + "age_domain": [ + 3, + 5 + ], + "direct_response_filter": "AEDHEADST == 1 and EEDHEADST in [1, 2]", + "structural_no_filters": [ + "AEDHEADST == 0 and AED_SCRNR == 1 and EED_SCRNR == 2", + "AEDHEADST == 0 and AED_SCRNR == 1 and EED_SCRNR == 1 and AEDGRADE == 1 and EEDGRADE != 21" + ], + "excluded_statuses": "AEDHEADST == 2 or AED_SCRNR == 4", + "n_estimators": 100, + "seed_from_build_config": true, + "assignment_unit": "person_source_id", + "fan_to_support_clones": true + } + ], + "outputs": [ + "takes_up_head_start_if_eligible" + ], + "notes": "Restores the retired Head Start take-up export from a measured person-month response instead of replaying its unsourced NIEER scalar or stochastic draw. The official 2023 SIPP dictionary defines EEDHEADST for nursery/preschool enrollees and records AEDHEADST allocation status; the survey instrument describes a federally sponsored preschool and names Head Start, Even Start, and Fair Start as examples, so this is documented as a measured Head Start proxy rather than a program-only identifier. The immutable December (December 2022 reference-month) donor keeps only direct age-3--5 answers and strict structural negatives whose upstream enrollment or grade screen is also reported, excluding hot-decked target answers and imputed upstream screens. The reviewed transform yields 785 training people (45 positive, 740 negative), 7,978,494.5412483 of total weight, 491,970.1041311 positive weight, and a 0.06166202177461505 weighted true share. The positive weighted count is close to the independent 2023--24 PIR age-matched cumulative count, but cumulative enrollment includes turnover and funded enrollment is capacity, so neither aggregate is used for calibration or person assignment. A weighted QRF predicts once per stable source-person identity on the age-3--5 receiver domain, fans the same decision to both support clones, and leaves every out-of-domain person false; modeled eligibility remains solely in PolicyEngine-US." + }, + { + "stage": "ssi_take_up", + "survey": "CPS ASEC reported SSI + SSA SSI Monthly Statistics December 2024", + "source": "https://www.ssa.gov/policy/docs/statcomps/ssi_monthly/2024-12/table01.html", + "grain": "person", + "artifacts": [ + { + "kind": "public_microdata", + "format": "census_asec", + "vintage": "build_year", + "locator": "CPS ASEC annual SSI amount SSI_VAL, used only as the reported-recipient true anchor" + }, + { + "kind": "administrative_table", + "format": "html_table", + "vintage": "2024-12", + "locator": "SSA SSI Monthly Statistics, December 2024, Table 1, row Total with—Federal payment, by age", + "source": "https://www.ssa.gov/policy/docs/statcomps/ssi_monthly/2024-12/table01.html", + "measure": "Total with—Federal payment", + "target_values": { + "under_18": 1001922, + "18_64": 3905779, + "65_plus": 2382142, + "total": 7289843 + } + }, + { + "kind": "archived_derivation_evidence", + "repository_owner": "PolicyEngine", + "repository_name_parts": [ + "policyengine-", + "us-data" + ], + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": [ + "policyengine_", + "us_data", + "datasets", + "cps", + "cps.py" + ], + "lines": "584,650-657,1497-1499", + "randomness_path_parts": [ + "policyengine_", + "us_data", + "datasets", + "cps", + "takeup.py" + ], + "randomness_lines": "10-35", + "targets_path_parts": [ + "policyengine_", + "us_data", + "utils", + "ssi_targets.py" + ], + "targets_lines": "41-74" + } ], - "notes": "Sentinel replacement is declared here so impossible negative auto-loan interest is rejected at the source boundary and again at export gates." + "operations": [ + { + "kind": "read_table", + "table": "person", + "weight": "person_weight" + }, + { + "kind": "assign_binary_from_rate", + "output": "takes_up_ssi_if_eligible", + "draw": "stable_source_person_draw", + "rate_key": "ssi_age_band_count_prior", + "rate_column": "ssi_take_up_assignment_prior", + "reported_true_anchor": "SSI_VAL > 0", + "assignment_unit": "person_source_id", + "fan_to_support_clones": true + }, + { + "kind": "calibrate_binary_assignment", + "variable": "takes_up_ssi_if_eligible", + "targets": [ + "ssa_ssi_federal_payment_recipients_by_age" + ], + "preserve_true_anchors": true, + "preserve_true_anchor": "SSI_VAL > 0", + "domain": "uncapped_ssi > 0", + "weight": "person_weight", + "draw": "stable_source_person_draw", + "calibration_unit": "person_source_id", + "age_bands": { + "under_18": "age < 18", + "18_64": "18 <= age < 65", + "65_plus": "age >= 65" + }, + "target_source": "https://www.ssa.gov/policy/docs/statcomps/ssi_monthly/2024-12/table01.html", + "target_period": "2024-12", + "target_measure": "Total with—Federal payment", + "target_values": { + "under_18": 1001922, + "18_64": 3905779, + "65_plus": 2382142 + }, + "aggregate_target": 7289843 + } + ], + "outputs": [ + "takes_up_ssi_if_eligible" + ], + "notes": "Restores the final retired eCPS SSI take-up export from source-backed evidence at archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe. datasets/cps/cps.py line 584 maps the direct ASEC SSI_VAL reporter signal, lines 650-657 derive the take-up leaf, and lines 1497-1499 export it; datasets/cps/takeup.py lines 10-35 define the retired stable random assignment. The retired all-age 0.50 rate has no full-population administrative source and is not reused. Instead, utils/ssi_targets.py lines 41-74 preserve the archived age-target method and cite SSA SSI Monthly Statistics, December 2024, Table 1, row Total with—Federal payment: 1,001,922 recipients under 18, 3,905,779 ages 18-64, and 2,382,142 ages 65+, totaling 7,289,843. Direct SSI_VAL reporters remain true. Remaining source-person identities are assigned deterministically within the uncapped_ssi > 0 domain and fanned identically to both support clones; an age band whose modeled eligible support cannot reach its SSA count saturates without assigning outside the eligibility domain, and the diagnostic records the irreducible shortfall." }, { "stage": "sipp_tips", @@ -280,9 +1327,10 @@ "artifacts": [ { "kind": "public_microdata", - "format": "census_extract", - "vintage": "latest_available", - "locator": "Census Survey of Income and Program Participation" + "format": "csv", + "vintage": "2023", + "locator": "Census SIPP 2023 public-use slim tip extract; immutable mirror revision 21280dca5995e978d706740a8a4b9b7860cfd7b6", + "sha256": "1f0bcb8e045ef1118e8eba4b4a2997bdaaf947bd0dd09d41fa7c7d5657a3d7d5" } ], "operations": [ @@ -294,65 +1342,154 @@ "kind": "fit_tip_income_model", "predictors": [ "employment_income", - "is_tipped_occupation", - "age" + "age", + "count_under_18", + "count_under_6", + "is_tipped_occupation" + ], + "source_columns": [ + "TJB1_TXAMT", + "TJB2_TXAMT", + "TJB3_TXAMT", + "TJB4_TXAMT", + "TJB5_TXAMT", + "TJB6_TXAMT", + "TJB7_TXAMT" + ], + "allocation_flag_columns": [ + "AJB1_TXAMT", + "AJB2_TXAMT", + "AJB3_TXAMT", + "AJB4_TXAMT", + "AJB5_TXAMT", + "AJB6_TXAMT", + "AJB7_TXAMT" + ], + "observed_status_values": [ + 0, + 1, + 9 ] - }, - { - "kind": "zero_when_false", - "condition": "is_tipped_occupation" } ], "outputs": [ - "tip_income" + "tip_income", + "treasury_tipped_occupation_code" ], "nonnegative_outputs": [ - "tip_income" - ] + "tip_income", + "treasury_tipped_occupation_code" + ], + "notes": "Tip income is the weighted-QRF port pinned to the retired implementation at 9a823603e6b5fb916d65ec45d74c9c7eb0043db1: observed-source December SIPP monthly tips across jobs 1-7 times 12, conditioned on employment income, age, child counts, and Treasury tipped occupation. SIPP allocation statuses 0, 1, and 9 are retained; imputed statuses 2-8 are excluded. The occupation indicator is a predictor rather than a hard domain mask: the pinned reference eCPS carries positive tips outside Treasury-listed occupations, while the OBBBA formula applies its own occupation-qualification test. treasury_tipped_occupation_code is mapped directly from CPS ASEC PEIOOCC using the IRS/Treasury 2025-42 tipped-occupation list and Census 2018 occupation/SOC crosswalk. The final retired downloader moved to an unpinned SIPP 2024 zip; this stage follows the final observed-source method while pinning the archived repository's immutable 2023 donor artifact, the artifact underlying the reference eCPS." }, { "stage": "org_wages", "survey": "CPS ORG", - "source": "https://www.bls.gov/cps/earnings.htm", + "source": "https://www2.census.gov/programs-surveys/cps/datasets/2024/basic/", "grain": "person", "artifacts": [ { "kind": "public_microdata", - "format": "cps_basic_monthly", - "vintage": "build_year", - "locator": "CPS Outgoing Rotation Group" + "format": "cps_basic_monthly_csv_or_zip", + "vintage": "2024", + "locator": "jan24pub through dec24pub; HRMIS 4 and 8", + "generated_cache": "census_cps_org_2024_wages.csv.gz", + "generated_cache_content_sha256": "d74600236cfdd34033d487cd9d82f6eb00b1858ba28de17f4f985f2aec516f86", + "generated_cache_rows": 119237 + }, + { + "kind": "administrative_table", + "format": "published_table", + "vintage": "2024", + "locator": "https://www.bls.gov/news.release/archives/union2_01282025.htm#union2.t01.1" + }, + { + "kind": "administrative_table", + "format": "published_table", + "vintage": "2024", + "locator": "https://www.bls.gov/news.release/archives/union2_01282025.htm#union2.t05.1" } ], "operations": [ { "kind": "read_table", - "table": "cps_org_person" + "table": "twelve_month_cps_org_person" }, { - "kind": "fit_labor_market_models", + "kind": "derive", + "method": "filter_outgoing_rotation_wage_salary_workers", + "rotation_groups": [ + 4, + 8 + ] + }, + { + "kind": "fit_weighted_qrf", "predictors": [ + "employment_income_before_lsr", + "weekly_hours_worked_before_lsr", "age", "is_female", - "cps_race", "is_hispanic", - "state_fips", - "employment_income", + "race_wbho", + "state_fips" + ], + "targets": [ + "hourly_wage", + "is_paid_hourly" + ] + }, + { + "kind": "assign_binary_from_rate", + "method": "assign_union_coverage", + "source": "BLS 2024 represented-by-union state and demographic rates" + }, + { + "kind": "derive", + "method": "derive_from_asec", + "inputs": [ + "PRDTRACE", + "PRDTHSP", + "POCCU2" + ] + }, + { + "kind": "derive", + "method": "derive_flsa_overtime_premium", + "inputs": [ + "employment_income_before_lsr", "hours_worked_last_week", - "weeks_worked" + "weeks_worked", + "is_paid_hourly", + "has_never_worked", + "is_military", + "is_computer_scientist", + "is_executive_administrative_professional", + "is_farmer_fisher" ] } ], "outputs": [ + "cps_race", + "is_hispanic", + "detailed_occupation_recode", + "has_never_worked", + "is_military", + "is_computer_scientist", + "is_executive_administrative_professional", + "is_farmer_fisher", "hourly_wage", "is_paid_hourly", "is_union_member_or_covered", "fsla_overtime_premium" ], "nonnegative_outputs": [ + "cps_race", + "detailed_occupation_recode", "hourly_wage", "fsla_overtime_premium" ], - "notes": "Usual weekly hours are owned by the hours_worked stage (ASEC HRSWK), matching the retired enhanced-CPS pipeline mapping; this stage consumes them as predictors." + "notes": "Exact port of the final retired ORG/FLSA method at 9a823603e6b5fb916d65ec45d74c9c7eb0043db1. The twelve official 2024 CPS basic-month files are transformed into the upstream 119,237-row cache; its canonical CSV content hash is pinned so a source reissue fails closed. Hourly wage and hourly-pay status use the weighted QRF; union coverage uses the archived deterministic BLS 2024 state-rate assignment; CPS race, Hispanic status, and FLSA occupation groups carry directly from ASEC. Usual hours remain owned by hours_worked. fsla_overtime_premium uses the upstream variable spelling and archived annual-wage-share formula." }, { "stage": "meps_esi_premiums", @@ -360,6 +1497,12 @@ "source": "https://meps.ahrq.gov/mepsweb/survey_comp/Insurance.jsp", "grain": "person", "artifacts": [ + { + "kind": "public_microdata", + "format": "census_asec", + "vintage": "hermetic ASEC pool years (2022-2024 for target 2024)", + "locator": "Census CPS ASEC NOW_OWNGRP, NOW_HIPAID, NOW_GRPFTYP, and PHIP_VAL fields; the three NOW_* fields are absent from the current hermetic HDF inputs" + }, { "kind": "administrative_table", "format": "published_table", @@ -375,10 +1518,20 @@ { "kind": "assign_by_plan_type", "inputs": [ - "has_esi", - "is_esi_dependent", - "tax_unit_size", - "state_fips" + "NOW_OWNGRP", + "NOW_HIPAID", + "NOW_GRPFTYP", + "PHIP_VAL" + ], + "owner_code": 1, + "employer_pays_all_code": 1, + "employer_pays_some_code": 2, + "family_plan_code": 1, + "self_only_plan_code": 2, + "source_unavailable_when_missing": [ + "NOW_OWNGRP", + "NOW_HIPAID", + "NOW_GRPFTYP" ] } ], @@ -387,7 +1540,8 @@ ], "nonnegative_outputs": [ "employer_sponsored_insurance_premiums" - ] + ], + "notes": "The retired derivation identifies the policy owner with NOW_OWNGRP, employer payment status with NOW_HIPAID, family/self-only plan type with NOW_GRPFTYP, and reported employee-paid premium with PHIP_VAL, then applies MEPS-IC family/self-only total and employee contribution priors (archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe, datasets/cps/cps.py lines 197-271 and 1575-1581). datasets/cps/census_cps.py lines 13-55 declares the three NOW_* fields optional. All three SHA-recorded 2022-2024 HDF inputs used by the hermetic build omit the three NOW_* fields, so this stage is intentionally source-unavailable; has_esi and NOW_GRP are forbidden proxies." }, { "stage": "prior_year_income", @@ -398,32 +1552,52 @@ { "kind": "public_microdata", "format": "census_asec", - "vintage": "build_year_minus_one", - "locator": "Census CPS ASEC person file" + "vintage": "pooled_adjacent_years", + "locator": "Census CPS ASEC person files identified by source_year and PERIDNUM" } ], "operations": [ { "kind": "read_table", - "table": "asec_person", - "columns": [ - "PERIDNUM", - "WSAL_VAL", - "SEMP_VAL" - ] - }, - { - "kind": "join", - "left": "current_person.PERIDNUM", - "right": "prior_person.PERIDNUM" + "table": "person" }, { - "kind": "replace_sentinels", + "kind": "derive_prior_year_income", + "person_id": "PERIDNUM", + "source_year": "source_year", + "prior_year_offset": -1, + "employment_source": "WSAL_VAL", + "self_employment_source": "SEMP_VAL", + "employment_allocation_flag": "I_ERNVAL", + "self_employment_allocation_flag": "I_SEVAL", + "unallocated_flag": 0, "sentinels": [ -1, -9999 ], - "value": null + "fallback_to_current": true, + "no_prior_artifact": "leave_defaults" + }, + { + "kind": "impute_prior_year_income_to_puf_support", + "predictors": [ + "age", + "is_male", + "has_esi", + "tax_unit_is_joint", + "tax_unit_count_dependents", + "employment_income", + "self_employment_income", + "social_security" + ], + "outputs": [ + "employment_income_last_year", + "self_employment_income_last_year" + ], + "max_train_samples": 5000, + "n_estimators": 100, + "seed_from_build_config": true, + "weight": "person_weight" } ], "outputs": [ @@ -432,9 +1606,9 @@ "previous_year_income_available" ], "nonnegative_outputs": [ - "employment_income_last_year", - "self_employment_income_last_year" - ] + "employment_income_last_year" + ], + "notes": "Archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe datasets/cps/cps.py lines 1680-1783 returns without creating values when no prior-year artifact exists; otherwise it joins adjacent ASEC files on PERIDNUM, requires I_ERNVAL == I_SEVAL == 0 on the prior record, treats only -1/-9999 as unavailable, records availability before falling back to current WSAL_VAL/SEMP_VAL, and preserves signed self-employment losses. After support cloning, datasets/cps/extended_cps.py lines 140-194, 639-745, and 1014-1073 jointly QRF-imputes both earnings leaves onto the PUF half from the exact eight predictors with a 5,000-person cap. Populace adds typed design weights. The archived finalizer drops formula-owned employment_income_last_year; the signed self-employment leaf and availability flag persist." }, { "stage": "immigration_status", @@ -608,6 +1782,144 @@ ], "notes": "Anchored count-calibration to FNS state caseloads (populace #372), following the medicaid_take_up pattern at SPM-unit grain. Reported recipients always take up; the fill draws at an in-build state rate (FNS average-monthly household count over weighted modeled-eligible units, an assignment prior, never cited as provenance) and is greedily calibrated to the FNS FY2024 state counts among eligible non-anchored units. Caseload semantics are fiscal-year average-monthly stocks - the same rows the snap_households weight-calibration targets compile from, so the seed and the calibration objective agree. CPS receipt underreporting is strongly state-dependent, so the national snap_take_up fill leaves the worst-underreporting states short of taker households no reweighting can recover. The assign operation deliberately carries no eligibility mask: off-domain units keep a draw-based propensity so eligibility-expanding reforms do not inherit a hard-coded zero response. States whose FNS count meets or exceeds modeled eligible weight saturate (all eligible units take up) and are recorded as saturated in diagnostics, not failed. Persons caseloads are deliberately untargeted: the SNAP assistance unit is often a subset of the SPM unit, so member counts overcount FNS participants by roughly half." }, + { + "stage": "relationship_inputs", + "survey": "Census CPS ASEC", + "source": "https://www.census.gov/programs-surveys/cps.html", + "grain": "person", + "artifacts": [ + { + "kind": "public_microdata", + "format": "census_asec", + "vintage": "build_year", + "locator": "Census CPS ASEC person household sequence (PH_SEQ), within-household person sequence (P_SEQ), and marital status (A_MARITL)" + } + ], + "operations": [ + { + "kind": "read_table", + "table": "person" + }, + { + "kind": "derive_relationship_inputs" + } + ], + "outputs": [ + "is_household_head", + "is_separated", + "is_surviving_spouse" + ], + "notes": "Exact measured ASEC mappings from archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe datasets/cps/cps.py lines 1069-1075 and 1209-1221: is_household_head <- P_SEQ == 1; is_separated <- A_MARITL == 6; is_surviving_spouse <- A_MARITL == 4. Nothing is imputed. The stage fails unless every frame household has exactly one P_SEQ == 1 record (PH_SEQ before support cloning, person_household_id after cloning)." + }, + { + "stage": "medicare_take_up_input", + "survey": "Census CPS ASEC", + "source": "https://www.census.gov/programs-surveys/cps.html", + "grain": "person", + "artifacts": [ + { + "kind": "public_microdata", + "format": "census_asec", + "vintage": "2022-2024 pooled", + "locator": "SHA-locked Census CPS ASEC person Medicare coverage last year (MCARE)" + }, + { + "kind": "archived_derivation_evidence", + "repository_owner": "PolicyEngine", + "repository_name_parts": ["policyengine-", "us-data"], + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "datasets", "cps", "cps.py"], + "lines": "1579-1585", + "source_declaration_path_parts": ["policyengine_", "us_data", "datasets", "cps", "census_cps.py"], + "source_declaration_lines": "39-58,306-406", + "clone_path_parts": ["policyengine_", "us_data", "calibration", "puf_impute.py"], + "clone_lines": "608-629", + "export_path_parts": ["policyengine_", "us_data", "datasets", "cps", "extended_cps.py"], + "export_lines": "1747-1754" + } + ], + "operations": [ + { + "kind": "read_table", + "table": "person" + }, + { + "kind": "derive_medicare_take_up", + "source": "MCARE", + "enrolled_code": 1, + "output": "takes_up_medicare_if_eligible" + } + ], + "outputs": [ + "takes_up_medicare_if_eligible" + ], + "notes": "Exact measured ASEC mapping from archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe datasets/cps/cps.py line 1584: medicare_enrolled <- MCARE == 1. Archived calibration/puf_impute.py lines 608-629 duplicates non-imputed CPS variables onto the PUF support half, and datasets/cps/extended_cps.py lines 1747-1754 renames medicare_enrolled to the leaf takes_up_medicare_if_eligible before export. No take-up rate or stochastic draw is used." + }, + { + "stage": "retirement_distributions", + "survey": "Census CPS ASEC", + "source": "https://www.census.gov/programs-surveys/cps.html", + "grain": "person", + "artifacts": [ + { + "kind": "public_microdata", + "format": "census_asec", + "vintage": "build_year", + "locator": "Census CPS ASEC person retirement-distribution account codes and amounts (DST_SC1/DST_VAL1, DST_SC2/DST_VAL2, DST_SC1_YNG/DST_VAL1_YNG, DST_SC2_YNG/DST_VAL2_YNG)" + } + ], + "operations": [ + { + "kind": "read_table", + "table": "person" + }, + { + "kind": "derive_retirement_distributions", + "output_by_account_code": { + "1": "taxable_401k_distributions", + "2": "taxable_403b_distributions", + "3": "tax_exempt_ira_distributions", + "4": "taxable_ira_distributions", + "5": "keogh_distributions", + "6": "taxable_sep_distributions" + } + }, + { + "kind": "impute_retirement_distributions_to_puf_support", + "predictors": [ + "age", + "is_male", + "has_esi", + "tax_unit_is_joint", + "tax_unit_count_dependents", + "employment_income", + "self_employment_income", + "social_security" + ], + "max_train_samples": 5000, + "n_estimators": 100, + "seed_from_build_config": true, + "weight": "person_weight" + } + ], + "outputs": [ + "taxable_401k_distributions", + "taxable_403b_distributions", + "tax_exempt_ira_distributions", + "taxable_ira_distributions", + "keogh_distributions", + "taxable_sep_distributions" + ], + "nonnegative_outputs": [ + "taxable_401k_distributions", + "taxable_403b_distributions", + "tax_exempt_ira_distributions", + "taxable_ira_distributions", + "keogh_distributions", + "taxable_sep_distributions" + ], + "notes": "Direct measured ASEC mappings from archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe datasets/cps/cps.py lines 1448-1481. Four account-code/amount pairs are summed by person: code 1 -> taxable 401(k), 2 -> taxable 403(b), 3 -> tax-exempt Roth IRA, 4 -> taxable regular IRA, 5 -> Keogh, and 6 -> taxable SEP. Archived imputation_parameters.yaml lines 10-15 sets the 401(k), 403(b), regular-IRA, and SEP taxable fractions to 1.0; Keogh is directly taxable in the engine aggregate. Code 7 has no PolicyEngine-US leaf and is deliberately not allocated. After support expansion, the retired second stage (extended_cps.py lines 140-148, 639-745, and 1014-1073) QRF-imputes, among these six populated leaves, taxable 401(k), taxable 403(b), Keogh, and taxable SEP onto the PUF half. Its three tax-exempt employer-account companions remain zero under the 1.0 taxable shares and lie outside the populated coverage surface. The measured ASEC half remains exact, PUF-sourced taxable IRA is preserved, and tax-exempt Roth IRA remains a direct copied ASEC mapping. No uncited allocation or redraw is introduced." + }, { "stage": "eligibility_inputs", "survey": "Census CPS ASEC", @@ -681,6 +1993,60 @@ ], "notes": "The CPS ASEC does not ask about pregnancy, so is_pregnant is seeded rather than mapped, matching the retired enhanced-CPS pipeline: women aged 15-44 receive stable blake2b draws against the pregnancy rate (point-in-time rate = annual births x 39/52 over the female 15-44 population; 4.1% nationally, the retired pipeline's national rate). The retired pipeline varied the rate by state using live CDC VSRR + Census ACS fetches; populace builds are hermetic, so the national rate applies uniformly and state-level variation is follow-up work under populace #351. Without this stage is_pregnant defaults to False for everyone and the SNAP ABAWD pregnancy exemption (7 U.S.C. 2015(o)(3)(E)) never fires." }, + { + "stage": "wic_claim_input", + "survey": "USDA FNS WIC Eligibility and Enrollment Estimates + Census CPS ASEC", + "source": "https://www.fns.usda.gov/research/wic/eligibility-and-coverage-rates-2022", + "grain": "person", + "artifacts": [ + { + "kind": "administrative_table", + "format": "published_estimate", + "vintage": "CY2022", + "source": "https://fns-prod.azureedge.us/sites/default/files/resource-files/wic-eer-2022-summary.pdf", + "locator": "USDA FNS WIC coverage rates by participant category, Summary Table 1" + }, + { + "kind": "archived_derivation_evidence", + "repository_owner": "PolicyEngine", + "repository_name_parts": ["policyengine-", "us-data"], + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "datasets", "cps", "cps.py"], + "lines": "684-691", + "parameter_path_parts": ["policyengine_", "us_data", "parameters", "take_up", "wic_takeup.yaml"], + "parameter_lines": "1-33", + "randomness_path_parts": ["policyengine_", "us_data", "utils", "randomness.py"], + "randomness_lines": "5-28" + } + ], + "operations": [ + { + "kind": "read_table", + "table": "person", + "weight": "person_weight" + }, + { + "kind": "derive_wic_claim", + "seed_from_build_config": true, + "category_rates": { + "source": "https://fns-prod.azureedge.us/sites/default/files/resource-files/wic-eer-2022-summary.pdf", + "vintage": "CY2022", + "values": { + "pregnant": 0.456, + "postpartum": 0.689, + "breastfeeding": 0.663, + "infant": 0.784, + "child": 0.46, + "none": 0.0 + } + } + } + ], + "outputs": [ + "would_claim_wic" + ], + "notes": "The retired eCPS derives PolicyEngine-US wic_category_str and assigns a stable seeded Bernoulli at category-specific rates (archived cps.py lines 684-691, wic_takeup.yaml lines 1-33, and randomness.py lines 5-28). The archived rates exactly match USDA FNS CY2022 average-month coverage: pregnant 45.6%, all postpartum 68.9%, breastfeeding 66.3%, infants 78.4%, and children 46.0%. Populace deliberately runs this stage after pregnancy; the retired file computed WIC categories before its later pregnancy seed, which accidentally made the pregnant rate unreachable. The hermetic sources carry no breastfeeding assessment, so breastfeeding remains unavailable and mothers of infants use FNS's all-postpartum rate rather than an invented breastfeeding status. Stable source-identity hashes preserve paired support-clone outcomes. The separate is_wic_at_nutritional_risk input remains a reviewed source-unavailability exclusion." + }, { "stage": "snap_abawd_discretionary_exemption", "survey": "Census CPS ASEC + statutory exemption cap (7 U.S.C. 2015(o)(6))", @@ -767,7 +2133,7 @@ "mortgage_principal", "investment_interest_expense" ], - "notes": "The conversion is a declarative structural transform over PUF and SCF outputs, not a benchmark data loader." + "notes": "The archived conversion maps the sole interest field used by the retired derivation, E19200, to interest_deduction, assigns the identical amount as deductible mortgage interest (archived puf.py lines 654 and 1548-1550), and derives investment_interest_expense only as their nonnegative residual (utils/mortgage_interest.py lines 268-320). The retired enhanced-CPS QRF predicted those source-identical aliases separately before conversion, so positive residuals are stochastic prediction divergence rather than observed detail. The SHA-pinned PUF 1.8.0 artifact contains 499,045 all-zero investment_interest_expense values and omits both aliases and raw E19200; all three locked ASEC inputs omit them too. A nonzero hermetic restoration would synthesize an unobserved mortgage-versus-investment split, so investment_interest_expense remains a SOURCE UNAVAILABILITY WITH EVIDENCE exclusion." }, { "stage": "acs_rent", @@ -777,52 +2143,140 @@ "artifacts": [ { "kind": "public_microdata", - "format": "acs_pums", + "format": "census_asec_hdf5", + "vintage": "2022-2024 pooled", + "locator": "SHA-locked Census CPS ASEC person and household tables", + "source_columns": [ + "H_TENURE", + "SPM_CAPHOUSESUB", + "SPM_TENMORTSTATUS", + "SPM_ID" + ] + }, + { + "kind": "versioned_derived_microdata", + "format": "hdf5_arrays", "vintage": "2022", - "locator": "Census ACS PUMS household and person files" + "locator": "Census ACS 2022 PUMS processed person/household arrays (acs_2022.h5)", + "sha256": "0b319b496f19a6913066f9c5ea572edfda3d78a187be6f375846617d0b441bd4", + "official_person_source": "https://www2.census.gov/programs-surveys/acs/data/pums/2022/1-Year/csv_pus.zip", + "official_household_source": "https://www2.census.gov/programs-surveys/acs/data/pums/2022/1-Year/csv_hus.zip", + "integrity_note": "The runtime validates the immutable artifact hash, entity alignment, archived head-only annual-rent construction, tenure collapse, both target-allocation flags, nonnegative weights, and positive total household support. Zero-WGTP group-quarters heads remain in the deterministic archived sample and receive zero mass in Populace's weighted fit; the runtime does not import or trust a local generating checkout." + }, + { + "kind": "archived_derivation_evidence", + "repository_owner": "PolicyEngine", + "repository_name_parts": ["policyengine-", "us-data"], + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "datasets", "cps", "cps.py"], + "lines": "417-535,1612-1635", + "take_up_lines": "585,664-682", + "take_up_parameter_path_parts": ["policyengine_", "us_data", "parameters", "take_up", "housing_assistance.yaml"], + "take_up_parameter_lines": "1-15", + "hud_etl_path_parts": ["policyengine_", "us_data", "db", "etl_housing_assistance.py"], + "hud_etl_lines": "35-169", + "acs_path_parts": ["policyengine_", "us_data", "datasets", "acs", "acs.py"], + "acs_lines": "82-135", + "census_acs_path_parts": ["policyengine_", "us_data", "datasets", "acs", "census_acs.py"], + "census_acs_lines": "13-68,95-125", + "puf_imputation_path_parts": ["policyengine_", "us_data", "datasets", "cps", "extended_cps.py"], + "puf_imputation_lines": "135-194,234-248,639-745,1014-1073" } ], "operations": [ { - "kind": "read_tables", - "tables": [ - "acs_household", - "acs_person" - ] + "kind": "derive_housing_tenure_inputs", + "household_tenure_source": "H_TENURE", + "household_tenure_map": { + "0": "NONE", + "1": "OWNED_WITH_MORTGAGE", + "2": "RENTED", + "3": "NONE" + }, + "housing_assistance_source": "SPM_CAPHOUSESUB", + "housing_take_up_output": "takes_up_housing_assistance_if_eligible", + "housing_take_up_assignment": "equal_to_measured_receipt", + "spm_tenure_source": "SPM_TENMORTSTATUS", + "spm_tenure_map": { + "1": "OWNER_WITH_MORTGAGE", + "2": "OWNER_WITHOUT_MORTGAGE", + "3": "RENTER" + } }, { - "kind": "aggregate_person_to_household", - "aggregates": [ - "employment_income", - "household_size" - ] + "kind": "read_acs_rent_donor", + "table": "acs_2022_arrays", + "target": "rent", + "allocation_flag": "rent_is_allocated", + "source_unit": "annual_usd", + "source_already_annualized": true, + "household_head_only": true }, { - "kind": "fit_weighted_imputer", + "kind": "fit_weighted_acs_rent_qrf", "predictors": [ - "state_fips", - "household_employment_income", + "is_household_head", + "age", + "is_male", + "tenure_type", + "employment_income", + "self_employment_income", + "social_security", + "pension_income", + "state_code_str", "household_size" ], - "weight": "household_weight" - }, - { - "kind": "zero_when_false", - "condition": "tenure_type == RENTED" + "target": "rent", + "shared_sample_targets": [ + "rent", + "real_estate_taxes" + ], + "target_allocation_flags": { + "rent": "rent_is_allocated", + "real_estate_taxes": "real_estate_taxes_is_allocated" + }, + "per_target_initial_cap": 5000, + "sample_fill_salt": "fill", + "weight": "household_weight", + "max_train_samples": 10000, + "n_estimators": 100, + "training_seed_name": "legacy_acs_rent_training_sample", + "seed_from_build_config": true, + "head_carry_output": "pre_subsidy_rent" }, { - "kind": "head_carry", - "from_entity": "household", - "to_entity": "person", - "condition": "is_household_head" + "kind": "impute_housing_assistance_to_puf_support", + "predictors": [ + "age", + "is_male", + "has_esi", + "tax_unit_is_joint", + "tax_unit_count_dependents", + "employment_income", + "self_employment_income", + "social_security" + ], + "target": "receives_housing_assistance", + "take_up_output": "takes_up_housing_assistance_if_eligible", + "take_up_assignment": "equal_to_imputed_receipt", + "weight": "person_weight", + "max_train_samples": 5000, + "n_estimators": 100, + "seed_from_build_config": true, + "reduction": "value_from_first_person" } ], "outputs": [ - "pre_subsidy_rent" + "pre_subsidy_rent", + "receives_housing_assistance", + "takes_up_housing_assistance_if_eligible", + "spm_unit_tenure_type", + "tenure_type" ], "nonnegative_outputs": [ "pre_subsidy_rent" - ] + ], + "notes": "Exact CPS H_TENURE and SPM carries plus the archived 10-predictor ACS rent QRF. receives_housing_assistance receives the retired second-stage CPS-to-PUF QRF, and takes_up_housing_assistance_if_eligible is kept exactly equal to that measured/imputed receipt anchor. Archived cps.py lines 664-682 preserved reporters before adding eligible nonreporters from a 0.50 draw; parameters/take_up/housing_assistance.yaml lines 12-15 say 0.50 was an ad hoc eCPS benchmark adjustment pending fuller calibration. Archived db/etl_housing_assistance.py lines 35-169 use a live HUD workbook to produce household-grain state counts, not SPM-unit recipient identities; that workbook is not pinned in the hermetic build and allocating its counts to SPM units would synthesize recipients. Populace never turns a reporter off and does not add unobserved take-up. Rent and both tenure enums are cloned unchanged." }, { "stage": "vehicle_assets", @@ -832,32 +2286,66 @@ "artifacts": [ { "kind": "public_microdata", - "format": "census_extract", - "vintage": "latest_available", - "locator": "Census Survey of Income and Program Participation" + "format": "pipe_delimited_csv", + "vintage": "2023", + "locator": "Census SIPP 2023 public-use file; immutable mirror revision 21280dca5995e978d706740a8a4b9b7860cfd7b6", + "sha256": "5c30439e365fc26483318ef61d1d8f4bb2f0e9d6bb47c22c06756a7698733ee2", + "size_bytes": 3726010471 } ], "operations": [ { "kind": "read_table", - "table": "sipp_household" + "table": "sipp_person", + "delimiter": "|", + "month_column": "MONTHCODE", + "month": 12, + "source_columns": [ + "SSUID", + "PNUM", + "MONTHCODE", + "WPFINWGT", + "TAGE", + "ESEX", + "EMS", + "TPTOTINC", + "TINC_BANK", + "TINC_STMF", + "TINC_BOND", + "TINC_RENT", + "TVEH_NUM", + "THVAL_VEH", + "THVAL_HOME", + "AVEH_NUM", + "AHVAL_VEH", + "AVEH1VAL", + "AVEH2VAL", + "AVEH3VAL" + ] }, { "kind": "fit_vehicle_model", "predictors": [ - "employment_income", - "interest_dividend_income", - "rental_income", - "age", - "is_female", - "is_married", - "tenure_type" - ] - }, - { - "kind": "fold_into", - "output": "net_worth", - "component": "household_vehicles_value" + "household_employment_income", + "household_interest_income", + "household_dividend_income", + "household_rental_income", + "reference_age", + "reference_is_female", + "reference_is_married", + "count_under_18", + "household_size", + "is_homeowner" + ], + "weight": "household_weight", + "observed_status_values": [0, 1, 9], + "owned_allocation_flag": "AVEH_NUM", + "value_allocation_flags": ["AVEH1VAL", "AVEH2VAL", "AVEH3VAL"], + "training_sample_cap": 20000, + "count_model": "weighted_random_forest_classifier", + "value_model": "weighted_qrf_conditioned_on_predicted_count", + "n_estimators": 100, + "seed": 42 } ], "outputs": [ @@ -867,7 +2355,75 @@ "nonnegative_outputs": [ "household_vehicles_owned", "household_vehicles_value" - ] + ], + "notes": "Restores the retired household-grain SIPP vehicle family from the immutable full 2023 public-use donor at archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe. datasets/sipp/sipp.py lines 430-456 and 953-1059 pin the exact 20 columns, December household transform, target-specific observed-source masks, positive survey weights, and 20,000-row cap; calibration/source_impute.py lines 992-1084 and utils/asset_imputation.py lines 503-603 pin the CPS receiver and output types. Vehicle count is classified first and vehicle value is then drawn conditionally; both are clipped nonnegative, count is rounded int32, and value is float32. The 65.6 MB tip-only slim SIPP file lacks this family. Populace deliberately does not execute the old placeholder fold into net_worth: the retired pipeline blended SIPP/SCF vehicle value before reconciling construction-only balance-sheet components, while Populace now persists the independently source-backed direct SCF networth anchor. Keeping the vehicle leaves and signed SCF anchor as separate policy inputs avoids inventing a partial component formula." + }, + { + "stage": "voluntary_filing_input", + "survey": "Census SIPP", + "source": "https://www.census.gov/programs-surveys/sipp.html", + "grain": "tax_unit", + "artifacts": [ + { + "kind": "public_microdata", + "format": "pipe_delimited_csv", + "vintage": "2023", + "locator": "Census SIPP 2023 public-use file; immutable mirror revision 21280dca5995e978d706740a8a4b9b7860cfd7b6", + "sha256": "5c30439e365fc26483318ef61d1d8f4bb2f0e9d6bb47c22c06756a7698733ee2", + "size_bytes": 3726010471 + } + ], + "operations": [ + { + "kind": "read_table", + "table": "sipp_person", + "delimiter": "|", + "month_column": "MONTHCODE", + "month": 12, + "source_columns": [ + "SSUID", + "PNUM", + "MONTHCODE", + "WPFINWGT", + "TAGE", + "ESEX", + "EPNSPOUSE", + "AFILING", + "EFILING", + "AWILLFILE", + "EWILLFILE", + "EDEPCLM", + "TJB1_MSUM", + "TJB2_MSUM", + "TJB3_MSUM", + "TJB4_MSUM", + "TJB5_MSUM", + "TJB6_MSUM", + "TJB7_MSUM" + ] + }, + { + "kind": "fit_weighted_qrf", + "predictors": [ + "employment_income", + "reference_age", + "reference_is_female", + "reference_is_married", + "count_under_18" + ], + "target": "would_file_taxes_voluntarily", + "weight": "tax_unit_weight", + "response_filter": "AFILING == 1 and (EFILING == 1 or (EFILING == 2 and AWILLFILE == 1))", + "dependent_exclusion": "EDEPCLM == 1", + "canonical_unit": "SSUID plus the sorted PNUM/EPNSPOUSE pair for reciprocal spouses; otherwise PNUM", + "n_estimators": 100, + "seed_from_build_config": true + } + ], + "outputs": [ + "would_file_taxes_voluntarily" + ], + "notes": "Restores a measured voluntary-filing response from the pinned full 2023 SIPP donor instead of replaying the retired uncited demographic probability table. Archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe, datasets/cps/cps.py lines 726-747, defines the retired output, predictors, and random draw; parameters/take_up/voluntary_filing.yaml lines 1-43 contains the uncited child-by-wage-by-age rates being replaced. Census SIPP 2023 public-use dictionary fields anchor the observed response: retain December records with AFILING == 1 and, when EFILING == 2, AWILLFILE == 1; set the target when EFILING == 1 or EWILLFILE == 1; and exclude EDEPCLM == 1 people because they report being claimed as another return's dependent. Reciprocal spouses form one canonical unit through EPNSPOUSE; their filing responses agree in the donor. Reference attributes and survey weight come from the complete unit's minimum-PNUM member before response filtering. Annual employment income is the canonical unit sum of December TJB1_MSUM through TJB7_MSUM times 12. count_under_18 is the full December SSUID household count, computed before response filtering and repeated as household composition context for each canonical source unit; it is deliberately not a claimed-dependent count because SIPP does not identify the claimant and separated-parent pointers can double-assign children. The source transformation yields 22,313 observed canonical units: 16,820 positive and 5,493 negative; 22,296 have positive finite reference weights, including 16,817 positive units, at a 0.7603084563 weighted true share. The weighted QRF is fit once per source tax unit and predicted once per unique receiver tax_unit_source_id so support clones remain identical." }, { "stage": "aca_marketplace_inputs", @@ -1027,6 +2583,62 @@ "takes_up_medicaid_if_eligible" ], "notes": "Anchored count-calibration (take-up contract treatment: count_calibrated; populace #331). Eligible reporters always take up; the fill draws at an in-build state rate (CMS count over weighted modeled eligibles, an assignment prior, never cited as provenance) and is greedily calibrated to the CMS December 2024 state snapshot among eligible non-anchored persons. Point-in-time (average-month) enrollment semantics per #332 - never ever-enrolled-in-year counts. The assign operation deliberately carries no eligibility mask: off-domain persons keep a draw-based propensity so eligibility-expanding reforms do not inherit a hard-coded zero response; the engine's is_medicaid_eligible gate hides the flag at baseline. States whose CMS count meets or exceeds modeled eligible weight saturate (all eligibles enroll) and are recorded as saturated in diagnostics, not failed (#170's eligibility-undercount shortfall is out of scope). CHIP deliberately excluded pending the #321 M-CHIP/separate-CHIP ledger concept split." + }, + { + "stage": "other_health_insurance_premiums", + "survey": "Census CPS ASEC", + "source": "https://www.census.gov/programs-surveys/cps.html", + "grain": "person", + "artifacts": [ + { + "kind": "public_microdata", + "format": "census_asec", + "vintage": "hermetic ASEC pool years (2022-2024 for target 2024)", + "source": "https://www.census.gov/programs-surveys/cps/data/datasets.html", + "locator": "Census CPS ASEC person PHIP_VAL reported premium, carried as health_insurance_premiums_without_medicare_part_b" + } + ], + "operations": [ + { + "kind": "read_table", + "table": "person", + "weight": "person_weight" + }, + { + "kind": "derive_other_health_insurance_premiums", + "reported_source": "health_insurance_premiums_without_medicare_part_b", + "chip_premium_source": "chip_premium", + "marketplace_net_premium_source": "marketplace_net_premium", + "medicaid_premium_source": "medicaid_premium", + "output": "other_health_insurance_premiums" + }, + { + "kind": "impute_other_health_insurance_premiums_to_puf_support", + "predictors": [ + "age", + "is_male", + "has_esi", + "tax_unit_is_joint", + "tax_unit_count_dependents", + "employment_income", + "self_employment_income", + "social_security" + ], + "max_train_samples": 5000, + "n_estimators": 100, + "seed_from_build_config": true, + "weight": "person_weight" + } + ], + "outputs": [ + "health_insurance_premiums_without_medicare_part_b", + "other_health_insurance_premiums" + ], + "nonnegative_outputs": [ + "health_insurance_premiums_without_medicare_part_b", + "other_health_insurance_premiums" + ], + "notes": "After ACA and Medicaid take-up inputs are materialized, the retired annual residual is max(PHIP_VAL - chip_premium - marketplace_net_premium - medicaid_premium, 0), with tax-unit modeled premiums assigned to the first person (archived commit 42ed5d45c56df80d754fbe24cce21cfeb8d05cbe, datasets/cps/cps.py lines 828-944). The archived CPS-only output list and exact eight-predictor, at-most-5,000-person joint QRF are extended_cps.py lines 135-194, 234-248, and 639-745, with only the PUF support half replaced at lines 1014-1076. Because the archive predicted the two PUF leaves jointly without clipping one to the other, PUF-only residual-order exceedances remain diagnostics; measured ASEC rows must preserve the exact residual identity. The archive did not explicitly pin tree count or weights; Populace pins 100 trees and typed person weights as reproducibility hardening." } ] } diff --git a/packages/populace-build/src/populace/build/us/take_up_contract.json b/packages/populace-build/src/populace/build/us/take_up_contract.json index 666fe8bb..0c75e49a 100644 --- a/packages/populace-build/src/populace/build/us/take_up_contract.json +++ b/packages/populace-build/src/populace/build/us/take_up_contract.json @@ -111,10 +111,11 @@ "entity": "person", "value_type": "bool", "default": true, - "populace_treatment": "near_universal", + "populace_treatment": "out_of_scope", + "scope_owner": "medicare_take_up_input source stage (measured ASEC MCARE == 1)", "consumed_via": "medicare_enrolled (adds), gated by is_medicare_eligible", - "rate": {"status": "near_universal_by_design", "source": "https://www.cms.gov/newsroom/fact-sheets", "note": "Part A near-automatic (~99%+ premium-free via 40+ work quarters); Part B ~90-93% among Part-A holders. Auto-enrollment at 65 for SSA claimants makes eligibility approximately equal enrollment."}, - "notes": "Take-up seeding is not meaningful for Medicare: the administrative reality is ~100% and CMS publishes enrollment counts rather than a take-up-among-eligibles rate. The engine default (True) is approximately correct; forcing sub-100% would be the error. Documented, not seeded." + "rate": {"status": "not_used_measured_source"}, + "notes": "The dedicated source stage restores the retired measured mapping MCARE == 1 from every locked ASEC year and copies it with source-identity clones onto the PUF support half. out_of_scope means out of scope for generic Bernoulli seeding, not unseeded." }, { "variable": "takes_up_ssi_if_eligible", @@ -122,11 +123,31 @@ "entity": "person", "value_type": "bool", "default": true, - "populace_treatment": "rate_unsourced", + "populace_treatment": "count_calibrated", "consumed_via": "ssi", - "rate": {"status": "scope_mismatch", "us_data_value": 0.50, "verified_source_value": {"value": 0.49, "scope": "adults 65+ only", "citation": "Urban Institute Giannarelli et al. 2024 (ATTIS, 2018), 49.0% aged SSI participation"}}, - "followup": "Source an SSI participation rate that covers the full eligible population (aged + blind + disabled), ideally split, as a Ledger fact; SSA disability-recipient take-up is not covered by the Urban aged-only estimate.", - "notes": "us-data's 0.50 was applied to ALL SSI eligibles, but the cited Urban estimate (49%) is explicitly adults 65+ only, while the majority of the SSI caseload is blind/disabled. Applying an aged-only rate to the whole caseload is unsupported. Left unseeded until a full-scope rate is sourced." + "calibration": { + "anchor": "SSI_VAL", + "targets": ["ssa_ssi_federal_payment_recipients_by_age"], + "target_table": "ssa_ssi_federal_payment_recipients_by_age", + "target_source": "https://www.ssa.gov/policy/docs/statcomps/ssi_monthly/2024-12/table01.html", + "target_period": "2024-12", + "target_measure": "Total with—Federal payment", + "target_values": { + "under_18": 1001922, + "18_64": 3905779, + "65_plus": 2382142 + }, + "aggregate_target": 7289843, + "age_bands": { + "under_18": "age < 18", + "18_64": "18 <= age < 65", + "65_plus": "age >= 65" + }, + "semantics": "SSA SSI Monthly Statistics December 2024 Table 1 recipients in the Total with—Federal payment row, calibrated within uncapped_ssi > 0 by source-person identity; unreachable age bands saturate without assigning outside modeled eligibility" + }, + "rate": {"status": "rejected_scope_mismatch", "retired_value": 0.50, "verified_source_value": {"value": 0.49, "scope": "adults 65+ only", "citation": "Urban Institute Giannarelli et al. 2024 (ATTIS, 2018), 49.0% aged SSI participation"}}, + "scope_owner": "ssi_take_up source stage (eCPS exported-input coverage)", + "notes": "The retired 0.50 rate was applied to all SSI eligibles even though its cited 49% estimate covers adults 65+ only; it is rejected, not copied. The ssi_take_up stage instead preserves every direct ASEC SSI_VAL reporter and calibrates eligible source-person identities to SSA SSI Monthly Statistics December 2024 Table 1 counts in the Total with—Federal payment row (under 18: 1,001,922; ages 18-64: 3,905,779; ages 65+: 2,382,142; total: 7,289,843). This is administrative count calibration, not a participation-rate claim. If modeled eligible support cannot reach a band target, the stage saturates that band and reports the shortfall rather than assigning an ineligible person." }, { "variable": "takes_up_dc_ptc", @@ -146,11 +167,11 @@ "entity": "person", "value_type": "bool", "default": true, - "populace_treatment": "rate_unsourced", + "populace_treatment": "out_of_scope", + "scope_owner": "sipp_head_start source stage (measured SIPP enrollment response)", "consumed_via": "head_start", - "rate": {"status": "research_only_not_administrative", "us_data_value": "0.40 (2020), 0.30 (2021)", "detail": "NIEER (Rutgers) State(s) of Head Start chart values; 40% is a pre-pandemic plateau and 30% the single pandemic program-year, so the 2020->2021 annual split is a framing artifact. NIEER is a research institute, not ACF/HHS."}, - "followup": "Construct a Head Start participation rate from ACF Head Start Program Facts enrollment over a Census poverty-eligible denominator (Ledger fact).", - "notes": "The spec requires an administrative (FNS/CMS/IRS/ACF) source; NIEER does not clear that bar and the year split is an artifact. Left unseeded pending an ACF-sourced rate." + "rate": {"status": "not_used_measured_source"}, + "notes": "The sipp_head_start source stage fits a weighted QRF to direct December age-3--5 SIPP EEDHEADST responses and strict reported structural negatives, then predicts once per source-person identity and fans the decision to both support clones. The survey instrument describes a federally sponsored preschool and names Head Start, Even Start, and Fair Start as examples, so the response is documented as a measured Head Start proxy rather than a program-only identifier. The retired NIEER scalar is not reused. ACF cumulative enrollment includes turnover and funded enrollment is capacity; neither aggregate is converted into a participation rate or allocated to people. out_of_scope means out of scope for generic Bernoulli seeding, not unseeded." }, { "variable": "takes_up_early_head_start_if_eligible", @@ -160,9 +181,9 @@ "default": true, "populace_treatment": "rate_unsourced", "consumed_via": "early_head_start", - "rate": {"status": "research_only_not_administrative", "us_data_value": 0.09, "detail": "NIEER (Rutgers) chart value, ~9% throughout including 2020-2021. Research institute, not ACF/HHS."}, - "followup": "Construct an Early Head Start participation rate from ACF Program Facts over an eligible-infant/toddler denominator (Ledger fact).", - "notes": "Same as Head Start: NIEER is research-only, below the administrative-source bar. Left unseeded pending an ACF-sourced rate." + "rate": {"status": "source_unavailable", "retired_value": 0.09, "detail": "The retired NIEER research-chart scalar was not an observed individual-level source and is rejected."}, + "followup": "Require a locked individual-level source that observes Early Head Start enrollment across infants, toddlers, and pregnant participants before restoring this person flag.", + "notes": "SOURCE UNAVAILABILITY WITH EVIDENCE is recorded structurally in ecps_parity_known_gaps.json. The pinned SIPP nursery/preschool item has no reported positives under age 3, its child-care items begin at age 2 and do not distinguish Early Head Start, and it has no pregnancy branch. All locked ASEC and PUF inputs omit an individual enrollment signal. ACF cumulative enrollment includes turnover and funded enrollment is capacity, so allocating either aggregate would synthesize person identities." }, { "variable": "takes_up_housing_assistance_if_eligible", @@ -170,11 +191,11 @@ "entity": "spm_unit", "value_type": "bool", "default": true, - "populace_treatment": "rate_unsourced", + "populace_treatment": "out_of_scope", + "scope_owner": "acs_rent source stage (measured/imputed housing-assistance receipt anchor)", "consumed_via": "housing_assistance", - "rate": {"status": "no_source", "detail": "No housing-assistance take-up rate parameter has ever existed in us-data. Housing assistance is rationed (waitlists, funding caps), so 'take-up among eligibles' is not the right model - assistance is capped by appropriation, not declined by eligibles."}, - "followup": "Housing assistance is supply-constrained; a HUD-sourced coverage rate (assisted households over eligible/waitlisted households) would be the Ledger fact, and the semantics differ from voluntary take-up.", - "notes": "No sourced rate and the take-up abstraction is a poor fit for a rationed program. Left unseeded; flagged for guardian review of the right modeling abstraction." + "rate": {"status": "not_used_measured_source"}, + "notes": "The source stage preserves every ASEC SPM_CAPHOUSESUB reporter and sets take-up exactly equal to receipt on both support channels. The retired 0.50 parameter was an ad hoc benchmark adjustment. The archived HUD ETL instead reads household-grain state counts from a live workbook, which cannot identify SPM-unit recipients in the hermetic build; Populace therefore does not allocate extra recipients. out_of_scope means out of scope for generic Bernoulli seeding, not unseeded." }, { "variable": "takes_up_aca_if_eligible", diff --git a/packages/populace-build/src/populace/build/us_runtime/__init__.py b/packages/populace-build/src/populace/build/us_runtime/__init__.py index d677677a..0d79dde7 100644 --- a/packages/populace-build/src/populace/build/us_runtime/__init__.py +++ b/packages/populace-build/src/populace/build/us_runtime/__init__.py @@ -35,12 +35,76 @@ load_source_manifest, load_support_spine_manifest, ) +from populace.build.us_runtime.alimony import ( + ALIMONY_ASEC_ARCHIVED_DERIVATION_URL, + ALIMONY_PUF_ARCHIVED_DERIVATION_URL, + US_ALIMONY_NONCONSTANT_PERSON_COLUMNS, + US_ALIMONY_OUTPUT_COLUMNS, + US_ALIMONY_STAGE_NAME, + derive_us_alimony_from_asec, + derive_us_alimony_from_puf, + us_alimony_signal_gate, + us_alimony_stage_spec, + us_alimony_summary, +) from populace.build.us_runtime.asec_pool import ( AsecSource, build_pooled_asec_unit_frame, load_asec_h5_tables, pool_asec_sources, ) +from populace.build.us_runtime.capital_gain_details import ( + CAPITAL_GAIN_DETAILS_ARCHIVED_DERIVATION_URL, + CAPITAL_GAIN_DETAILS_ARCHIVED_EXPORT_URL, + CAPITAL_GAIN_DETAILS_ARCHIVED_IMPUTATION_URL, + CAPITAL_GAIN_DETAILS_ARCHIVED_PERSON_ALLOCATION_URL, + CAPITAL_GAIN_DETAILS_ARCHIVED_PUF_ARTIFACT_URL, + US_CAPITAL_GAIN_DETAILS_NONCONSTANT_PERSON_COLUMNS, + US_CAPITAL_GAIN_DETAILS_NONCONSTANT_TAX_UNIT_COLUMNS, + US_CAPITAL_GAIN_DETAILS_OUTPUT_COLUMNS, + US_CAPITAL_GAIN_DETAILS_STAGE_NAME, + derive_us_capital_gain_details_from_puf, + us_capital_gain_details_signal_gate, + us_capital_gain_details_stage_spec, + us_capital_gain_details_summary, +) +from populace.build.us_runtime.casualty_losses import ( + US_CASUALTY_LOSS_NONCONSTANT_PERSON_COLUMNS, + US_CASUALTY_LOSS_OUTPUT_COLUMNS, + US_CASUALTY_LOSS_STAGE_NAME, + derive_us_casualty_loss_from_puf, + us_casualty_loss_signal_gate, + us_casualty_loss_stage_spec, + us_casualty_loss_summary, +) +from populace.build.us_runtime.child_support import ( + CHILD_SUPPORT_ARCHIVED_PUF_IMPUTATION_URL, + CHILD_SUPPORT_ARCHIVED_PUF_OUTPUTS_URL, + CHILD_SUPPORT_EXPENSE_ARCHIVED_DERIVATION_URL, + CHILD_SUPPORT_RECEIVED_ARCHIVED_DERIVATION_URL, + US_CHILD_SUPPORT_NONCONSTANT_PERSON_COLUMNS, + US_CHILD_SUPPORT_OUTPUT_COLUMNS, + US_CHILD_SUPPORT_REQUIRED_SOURCE_COLUMNS, + US_CHILD_SUPPORT_STAGE_NAME, + derive_us_child_support_from_asec, + derive_us_child_support_from_manifest, + impute_us_child_support_to_puf_support_from_manifest, + us_child_support_signal_gate, + us_child_support_stage_spec, + us_child_support_summary, + with_us_child_support_inputs, +) +from populace.build.us_runtime.childcare import ( + US_CHILDCARE_OUTPUT_COLUMNS, + US_CHILDCARE_REQUIRED_SOURCE_COLUMNS, + US_CHILDCARE_STAGE_NAME, + derive_us_childcare_from_manifest, + impute_us_childcare_to_puf_support_from_manifest, + us_childcare_signal_gate, + us_childcare_stage_spec, + us_childcare_summary, + with_us_childcare_inputs, +) from populace.build.us_runtime.congressional_district_geography import ( CONGRESSIONAL_DISTRICT_GEOID_COLUMN, SOI_CONGRESSIONAL_DISTRICT_RECORD_SET_ID, @@ -72,6 +136,61 @@ demographics_payload, write_demographics, ) +from populace.build.us_runtime.disability_benefits import ( + DISABILITY_BENEFITS_ARCHIVED_DERIVATION_URL, + DISABILITY_BENEFITS_ARCHIVED_PUF_IMPUTATION_URL, + DISABILITY_BENEFITS_ARCHIVED_PUF_OUTPUTS_URL, + DISABILITY_BENEFITS_ARCHIVED_SOURCE_COLUMNS_URL, + US_DISABILITY_BENEFITS_NONCONSTANT_PERSON_COLUMNS, + US_DISABILITY_BENEFITS_OUTPUT_COLUMNS, + US_DISABILITY_BENEFITS_REQUIRED_SOURCE_COLUMNS, + US_DISABILITY_BENEFITS_STAGE_NAME, + derive_us_disability_benefits_from_asec, + derive_us_disability_benefits_from_manifest, + impute_us_disability_benefits_to_puf_support_from_manifest, + us_disability_benefits_signal_gate, + us_disability_benefits_stage_spec, + us_disability_benefits_summary, + with_us_disability_benefits, +) +from populace.build.us_runtime.domestic_production import ( + DOMESTIC_PRODUCTION_ALD_ARCHIVED_DERIVATION_URL, + DOMESTIC_PRODUCTION_ALD_ARCHIVED_EXPORT_URL, + DOMESTIC_PRODUCTION_ALD_ARCHIVED_IMPUTATION_URL, + DOMESTIC_PRODUCTION_ALD_ARCHIVED_PUF_ARTIFACT_URL, + US_DOMESTIC_PRODUCTION_ALD_NONCONSTANT_TAX_UNIT_COLUMNS, + US_DOMESTIC_PRODUCTION_ALD_OUTPUT_COLUMNS, + US_DOMESTIC_PRODUCTION_ALD_STAGE_NAME, + derive_us_domestic_production_ald_from_puf, + us_domestic_production_ald_signal_gate, + us_domestic_production_ald_stage_spec, + us_domestic_production_ald_summary, +) +from populace.build.us_runtime.education_inputs import ( + US_AOTC_ELIGIBILITY_OUTPUT_COLUMNS, + US_EDUCATION_INPUTS_NONCONSTANT_PERSON_COLUMNS, + US_EDUCATION_INPUTS_OUTPUT_COLUMNS, + US_EDUCATION_INPUTS_REQUIRED_SOURCE_COLUMNS, + US_EDUCATION_INPUTS_STAGE_NAME, + derive_us_education_inputs_from_manifest, + us_education_inputs_signal_gate, + us_education_inputs_stage_spec, + us_education_inputs_summary, + with_us_education_inputs, +) +from populace.build.us_runtime.educator_expenses import ( + EDUCATOR_EXPENSE_ARCHIVED_ALLOCATION_URL, + EDUCATOR_EXPENSE_ARCHIVED_DERIVATION_URL, + EDUCATOR_EXPENSE_ARCHIVED_EXPORT_URL, + EDUCATOR_EXPENSE_ARCHIVED_PUF_IMPUTATION_URL, + US_EDUCATOR_EXPENSE_NONCONSTANT_PERSON_COLUMNS, + US_EDUCATOR_EXPENSE_OUTPUT_COLUMNS, + US_EDUCATOR_EXPENSE_STAGE_NAME, + derive_us_educator_expense_from_puf, + us_educator_expense_signal_gate, + us_educator_expense_stage_spec, + us_educator_expense_summary, +) from populace.build.us_runtime.eligibility_inputs import ( US_ELIGIBILITY_INPUTS_NONCONSTANT_PERSON_COLUMNS, US_ELIGIBILITY_INPUTS_OUTPUT_COLUMNS, @@ -83,6 +202,34 @@ us_eligibility_inputs_summary, with_us_eligibility_inputs, ) +from populace.build.us_runtime.energy_subsidy import ( + ENERGY_SUBSIDY_ARCHIVED_CPS_DERIVATION_URL, + ENERGY_SUBSIDY_ARCHIVED_PUF_IMPUTATION_URL, + US_ENERGY_SUBSIDY_OUTPUT_COLUMNS, + US_ENERGY_SUBSIDY_REQUIRED_SOURCE_COLUMNS, + US_ENERGY_SUBSIDY_STAGE_NAME, + derive_us_energy_subsidy_from_manifest, + impute_us_energy_subsidy_to_puf_support_from_manifest, + us_energy_subsidy_signal_gate, + us_energy_subsidy_stage_spec, + us_energy_subsidy_summary, + with_us_energy_subsidy_input, +) +from populace.build.us_runtime.farm_business_income import ( + FARM_BUSINESS_INCOME_ARCHIVED_CPS_FARM_INCOME_URL, + FARM_BUSINESS_INCOME_ARCHIVED_DERIVATION_URL, + FARM_BUSINESS_INCOME_ARCHIVED_EXPORT_URL, + FARM_BUSINESS_INCOME_ARCHIVED_IMPUTATION_URL, + FARM_BUSINESS_INCOME_ARCHIVED_OVERRIDE_URL, + FARM_BUSINESS_INCOME_ARCHIVED_PUF_ARTIFACT_URL, + US_FARM_BUSINESS_INCOME_NONCONSTANT_PERSON_COLUMNS, + US_FARM_BUSINESS_INCOME_OUTPUT_COLUMNS, + US_FARM_BUSINESS_INCOME_STAGE_NAME, + derive_us_farm_business_income_from_puf, + us_farm_business_income_signal_gate, + us_farm_business_income_stage_spec, + us_farm_business_income_summary, +) from populace.build.us_runtime.fiscal_targets import ( SOI_VARIABLE_MAP, US_FISCAL_LEDGER_PARITY_REGISTRY, @@ -104,6 +251,20 @@ SimpleTaxExpenditureReform, compile_us_fiscal_target_registry, ) +from populace.build.us_runtime.form_4952 import ( + FORM_4952_ARCHIVED_DERIVATION_URL, + FORM_4952_ARCHIVED_EXPORT_URL, + FORM_4952_ARCHIVED_IMPUTATION_URL, + FORM_4952_ARCHIVED_PERSON_ALLOCATION_URL, + FORM_4952_ARCHIVED_PUF_ARTIFACT_URL, + US_FORM_4952_NONCONSTANT_PERSON_COLUMNS, + US_FORM_4952_OUTPUT_COLUMNS, + US_FORM_4952_STAGE_NAME, + derive_us_form_4952_election_from_puf, + us_form_4952_election_signal_gate, + us_form_4952_election_stage_spec, + us_form_4952_election_summary, +) from populace.build.us_runtime.geography_ladder import ( GEOGRAPHY_LADDER_ARTIFACT_SHA256_ATTR, GEOGRAPHY_LADDER_VINTAGES_ATTR, @@ -130,6 +291,34 @@ us_hours_worked_summary, with_us_hours_worked_inputs, ) +from populace.build.us_runtime.housing_inputs import ( + ACS_2022_RENT_ARTIFACT_SHA256, + HOUSING_INPUTS_ARCHIVED_ACS_DERIVATION_URL, + HOUSING_INPUTS_ARCHIVED_CPS_RENT_URL, + HOUSING_INPUTS_ARCHIVED_CPS_SPM_URL, + HOUSING_INPUTS_ARCHIVED_PUF_IMPUTATION_URL, + HOUSING_TAKE_UP_ARCHIVED_DERIVATION_URL, + HOUSING_TAKE_UP_ARCHIVED_HUD_ETL_URL, + HOUSING_TAKE_UP_ARCHIVED_PARAMETER_URL, + US_HOUSING_HOUSEHOLD_OUTPUT_COLUMNS, + US_HOUSING_INPUTS_OUTPUT_COLUMNS, + US_HOUSING_INPUTS_STAGE_NAME, + US_HOUSING_NONCONSTANT_HOUSEHOLD_COLUMNS, + US_HOUSING_NONCONSTANT_PERSON_COLUMNS, + US_HOUSING_NONCONSTANT_SPM_UNIT_COLUMNS, + US_HOUSING_PERSON_OUTPUT_COLUMNS, + US_HOUSING_REQUIRED_HOUSEHOLD_SOURCE_COLUMNS, + US_HOUSING_REQUIRED_PERSON_SOURCE_COLUMNS, + US_HOUSING_SPM_UNIT_OUTPUT_COLUMNS, + derive_us_housing_inputs, + impute_us_housing_assistance_to_puf_support, + impute_us_pre_subsidy_rent, + load_acs_2022_rent_donor, + us_housing_inputs_signal_gate, + us_housing_inputs_stage_spec, + us_housing_inputs_summary, + with_us_housing_inputs, +) from populace.build.us_runtime.immigration import ( IMMIGRATION_STATUS_VALUES, SSN_CARD_TYPE_VALUES, @@ -157,15 +346,80 @@ US_MEDICAID_TAKE_UP_VARIABLE, MedicaidEnrollmentSubstitution, apply_us_medicaid_enrollment_substitutions, + us_medicaid_source_person_table, us_medicaid_take_up_diagnostics, us_medicaid_take_up_gate, with_us_medicaid_take_up, write_us_medicaid_take_up_diagnostics, ) +from populace.build.us_runtime.medicare_take_up import ( + MEDICARE_TAKE_UP_ARCHIVED_CLONE_URL, + MEDICARE_TAKE_UP_ARCHIVED_DERIVATION_URL, + MEDICARE_TAKE_UP_ARCHIVED_EXPORT_URL, + MEDICARE_TAKE_UP_ARCHIVED_SOURCE_COLUMNS_URL, + US_MEDICARE_TAKE_UP_NONCONSTANT_PERSON_COLUMNS, + US_MEDICARE_TAKE_UP_OUTPUT_COLUMNS, + US_MEDICARE_TAKE_UP_REQUIRED_SOURCE_COLUMNS, + US_MEDICARE_TAKE_UP_STAGE_NAME, + derive_us_medicare_take_up_from_manifest, + us_medicare_take_up_signal_gate, + us_medicare_take_up_stage_spec, + us_medicare_take_up_summary, + with_us_medicare_take_up_input, +) +from populace.build.us_runtime.misc_itemized import ( + US_MISC_ITEMIZED_NONCONSTANT_PERSON_COLUMNS, + US_MISC_ITEMIZED_OUTPUT_COLUMNS, + US_MISC_ITEMIZED_STAGE_NAME, + derive_us_misc_itemized_from_puf, + us_misc_itemized_signal_gate, + us_misc_itemized_stage_spec, + us_misc_itemized_summary, +) from populace.build.us_runtime.nonzero_shares import ( nonzero_share, us_nonzero_shares, ) +from populace.build.us_runtime.org_wages import ( + BLS_STATE_UNION_REPRESENTATION_RATE_2024, + FLSA_EXECUTIVE_ADMINISTRATIVE_PROFESSIONAL_OCCUPATION_CODES, + FLSA_OVERTIME_OCCUPATION_CODES, + ORG_2024_DONOR_CONTENT_SHA256, + ORG_2024_DONOR_FILENAME, + ORG_PREDICTORS, + US_ORG_WAGES_NONCONSTANT_PERSON_COLUMNS, + US_ORG_WAGES_OUTPUT_COLUMNS, + US_ORG_WAGES_REQUIRED_SOURCE_COLUMNS, + US_ORG_WAGES_STAGE_NAME, + derive_flsa_overtime_premium, + derive_us_org_occupation_inputs, + fetch_org_2024_donor, + impute_us_org_wages, + load_org_2024_donor, + us_org_wages_signal_gate, + us_org_wages_stage_spec, + us_org_wages_summary, + with_us_org_wages_inputs, +) +from populace.build.us_runtime.other_health_insurance import ( + OTHER_HEALTH_INSURANCE_ARCHIVED_DERIVATION_URL, + OTHER_HEALTH_INSURANCE_ARCHIVED_PUF_IMPUTATION_URL, + OTHER_HEALTH_INSURANCE_ARCHIVED_PUF_OUTPUTS_URL, + OTHER_HEALTH_INSURANCE_ARCHIVED_PUF_PREDICTORS_URL, + OTHER_HEALTH_INSURANCE_ARCHIVED_PUF_SPLICE_URL, + US_OTHER_HEALTH_INSURANCE_MODELED_PREMIUM_VARIABLES, + US_OTHER_HEALTH_INSURANCE_NONCONSTANT_PERSON_COLUMNS, + US_OTHER_HEALTH_INSURANCE_OUTPUT_COLUMNS, + US_OTHER_HEALTH_INSURANCE_REQUIRED_SOURCE_COLUMNS, + US_OTHER_HEALTH_INSURANCE_STAGE_NAME, + derive_us_other_health_insurance_from_asec, + derive_us_other_health_insurance_from_manifest, + impute_us_other_health_insurance_to_puf_support_from_manifest, + us_other_health_insurance_signal_gate, + us_other_health_insurance_stage_spec, + us_other_health_insurance_summary, + with_us_other_health_insurance_inputs, +) from populace.build.us_runtime.parity_reference import ( ECPS_PARITY_KNOWN_GAPS_RESOURCE, ECPS_PARITY_REFERENCE_RESOURCE, @@ -186,6 +440,26 @@ us_pregnancy_summary, with_us_pregnancy_inputs, ) +from populace.build.us_runtime.prior_year_income import ( + PRIOR_YEAR_INCOME_ARCHIVED_DERIVATION_URL, + PRIOR_YEAR_INCOME_ARCHIVED_FINALIZER_URL, + PRIOR_YEAR_INCOME_ARCHIVED_FORMULA_OUTPUT_URL, + PRIOR_YEAR_INCOME_ARCHIVED_PUF_IMPUTATION_URL, + PRIOR_YEAR_INCOME_ARCHIVED_PUF_OUTPUTS_URL, + PRIOR_YEAR_INCOME_ARCHIVED_PUF_SPLICE_URL, + US_PRIOR_YEAR_INCOME_NONCONSTANT_PERSON_COLUMNS, + US_PRIOR_YEAR_INCOME_OUTPUT_COLUMNS, + US_PRIOR_YEAR_INCOME_PERSISTED_OUTPUT_COLUMNS, + US_PRIOR_YEAR_INCOME_REQUIRED_SOURCE_COLUMNS, + US_PRIOR_YEAR_INCOME_STAGE_NAME, + derive_us_prior_year_income_from_manifest, + impute_us_prior_year_income_to_puf_support_from_manifest, + us_prior_year_income_signal_gate, + us_prior_year_income_source_reconciliation_gate, + us_prior_year_income_stage_spec, + us_prior_year_income_summary, + with_us_prior_year_income_inputs, +) from populace.build.us_runtime.puf_support import ( BASE_ASEC_SUPPORT_CHANNEL, PUF_TAX_DETAIL_DEFAULT_PERSON_OUTPUTS, @@ -219,6 +493,24 @@ assemble_us_puma_ladder, parse_tract_to_puma_relationship, ) +from populace.build.us_runtime.qbi_inputs import ( + QBI_ARCHIVED_ASSUMPTIONS_URL, + QBI_ARCHIVED_CLONE_URL, + QBI_ARCHIVED_DERIVATION_URL, + QBI_ARCHIVED_EXPORT_URL, + QBI_ARCHIVED_IMPUTATION_URL, + QBI_ARCHIVED_PUF_ARTIFACT_URL, + QBI_ARCHIVED_SIMULATION_URL, + US_QBI_BOOLEAN_OUTPUT_COLUMNS, + US_QBI_NONCONSTANT_PERSON_COLUMNS, + US_QBI_NONNEGATIVE_OUTPUT_COLUMNS, + US_QBI_OUTPUT_COLUMNS, + US_QBI_STAGE_NAME, + us_qbi_inputs_signal_gate, + us_qbi_inputs_stage_spec, + us_qbi_inputs_summary, + with_us_qbi_input_reconciliation, +) from populace.build.us_runtime.reform_coverage_smoke import ( us_reform_coverage_smoke_gate, ) @@ -235,7 +527,19 @@ us_register_consistency_gate, us_register_contradictions, ) +from populace.build.us_runtime.relationship_inputs import ( + US_RELATIONSHIP_INPUTS_NONCONSTANT_PERSON_COLUMNS, + US_RELATIONSHIP_INPUTS_OUTPUT_COLUMNS, + US_RELATIONSHIP_INPUTS_REQUIRED_SOURCE_COLUMNS, + US_RELATIONSHIP_INPUTS_STAGE_NAME, + derive_us_relationship_inputs_from_manifest, + us_relationship_inputs_signal_gate, + us_relationship_inputs_stage_spec, + us_relationship_inputs_summary, + with_us_relationship_inputs, +) from populace.build.us_runtime.release_input_coverage import ( + POST_REFERENCE_ECPS_REQUIRED_INPUTS, SSI_COUNTABLE_RESOURCE_ASSETS, US_RELEASE_INPUT_COVERAGE_RESOURCE, ReformCoverageProbe, @@ -248,20 +552,139 @@ us_release_input_coverage_reviewed_exclusions, us_release_reform_coverage_probes, ) +from populace.build.us_runtime.retirement_contributions import ( + US_RETIREMENT_CONTRIBUTION_NONCONSTANT_PERSON_COLUMNS, + US_RETIREMENT_CONTRIBUTION_OUTPUT_COLUMNS, + US_RETIREMENT_CONTRIBUTION_REQUIRED_SOURCE_COLUMNS, + US_RETIREMENT_CONTRIBUTION_STAGE_NAME, + derive_us_retirement_contributions_from_manifest, + impute_us_retirement_contributions_to_puf_support_from_manifest, + us_retirement_contributions_signal_gate, + us_retirement_contributions_stage_spec, + us_retirement_contributions_summary, + with_us_retirement_contribution_inputs, +) +from populace.build.us_runtime.retirement_distributions import ( + RETIREMENT_DISTRIBUTIONS_ARCHIVED_DERIVATION_URL, + RETIREMENT_DISTRIBUTIONS_ARCHIVED_PARAMETERS_URL, + US_RETIREMENT_DISTRIBUTION_NONCONSTANT_PERSON_COLUMNS, + US_RETIREMENT_DISTRIBUTION_OUTPUT_COLUMNS, + US_RETIREMENT_DISTRIBUTION_REQUIRED_SOURCE_COLUMNS, + US_RETIREMENT_DISTRIBUTION_STAGE_NAME, + derive_us_retirement_distributions_from_manifest, + impute_us_retirement_distributions_to_puf_support_from_manifest, + us_retirement_distributions_signal_gate, + us_retirement_distributions_stage_spec, + us_retirement_distributions_summary, + with_us_retirement_distribution_inputs, +) +from populace.build.us_runtime.salt_refund_income import ( + SALT_REFUND_ARCHIVED_DERIVATION_URL, + SALT_REFUND_ARCHIVED_EXPORT_URL, + SALT_REFUND_ARCHIVED_IMPUTATION_URL, + SALT_REFUND_ARCHIVED_PERSON_ALLOCATION_URL, + SALT_REFUND_ARCHIVED_PUF_ARTIFACT_URL, + US_SALT_REFUND_NONCONSTANT_PERSON_COLUMNS, + US_SALT_REFUND_OUTPUT_COLUMNS, + US_SALT_REFUND_STAGE_NAME, + derive_us_salt_refund_income_from_puf, + us_salt_refund_income_signal_gate, + us_salt_refund_income_stage_spec, + us_salt_refund_income_summary, +) +from populace.build.us_runtime.scf_auto_loans import ( + QUALIFIED_AUTO_LOAN_ANNUAL_ISSUANCE_TARGET, + SCF_2022_FULL_EXTRACT_MEMBER, + SCF_2022_FULL_EXTRACT_MEMBER_SHA256, + SCF_2022_FULL_EXTRACT_URL, + SCF_2022_FULL_EXTRACT_ZIP_SHA256, + SCF_AUTO_LOAN_AMOUNT_COLUMNS, + SCF_AUTO_LOAN_RATE_COLUMNS, + US_SCF_AUTO_LOAN_NONCONSTANT_HOUSEHOLD_COLUMNS, + US_SCF_AUTO_LOAN_OUTPUT_COLUMNS, + fetch_scf_2022_full_extract, + impute_us_scf_auto_loans, + load_scf_2022_auto_loan_donor, + qualified_auto_loan_interest_proxy, + us_scf_auto_loans_signal_gate, + us_scf_auto_loans_stage_spec, + us_scf_auto_loans_summary, + with_us_scf_auto_loan_inputs, +) from populace.build.us_runtime.scf_wealth import ( SCF_FINANCIAL_ASSET_TARGET_COMPONENTS, + SCF_NET_WORTH_TARGET_COMPONENTS, SCF_WEALTH_PREDICTORS, US_SCF_FINANCIAL_ASSET_OUTPUT_COLUMNS, + US_SCF_NET_WORTH_OUTPUT_COLUMNS, + US_SCF_WEALTH_NONCONSTANT_HOUSEHOLD_COLUMNS, US_SCF_WEALTH_NONCONSTANT_PERSON_COLUMNS, US_SCF_WEALTH_STAGE_NAME, fetch_scf_2022_summary_extract, impute_us_scf_financial_assets, + impute_us_scf_net_worth, load_scf_2022_financial_asset_donor, us_scf_wealth_signal_gate, us_scf_wealth_stage_spec, us_scf_wealth_summary, with_us_scf_wealth_inputs, ) +from populace.build.us_runtime.sipp_head_start import ( + HEAD_START_SIPP_DICTIONARY_URL, + SIPP_2023_HEAD_START_DONOR_REVISION, + SIPP_2023_HEAD_START_DONOR_SHA256, + SIPP_2023_HEAD_START_DONOR_SIZE_BYTES, + SIPP_2023_HEAD_START_DONOR_URL, + SIPP_HEAD_START_FIT_PARAMETERS, + SIPP_HEAD_START_MODEL_PREDICTORS, + SIPP_HEAD_START_READ_PARAMETERS, + SIPP_HEAD_START_SOURCE_COLUMNS, + US_SIPP_HEAD_START_NONCONSTANT_PERSON_COLUMNS, + US_SIPP_HEAD_START_OUTPUT_COLUMNS, + US_SIPP_HEAD_START_REQUIRED_SOURCE_COLUMNS, + US_SIPP_HEAD_START_STAGE_NAME, + fetch_sipp_2023_head_start_donor, + impute_us_sipp_head_start, + load_sipp_2023_head_start_donor, + us_sipp_head_start_signal_gate, + us_sipp_head_start_stage_spec, + us_sipp_head_start_summary, + with_us_sipp_head_start_input, +) +from populace.build.us_runtime.sipp_tips import ( + CENSUS_OCCUPATION_CODE_TO_TTOC, + SIPP_2023_TIP_DONOR_REVISION, + SIPP_2023_TIP_DONOR_SHA256, + SIPP_2023_TIP_DONOR_URL, + SIPP_TIP_OUTPUT_COLUMNS, + SIPP_TIP_PREDICTORS, + US_SIPP_TIPS_NONCONSTANT_PERSON_COLUMNS, + US_SIPP_TIPS_OUTPUT_COLUMNS, + US_SIPP_TIPS_REQUIRED_SOURCE_COLUMNS, + US_SIPP_TIPS_STAGE_NAME, + derive_treasury_tipped_occupation_code, + fetch_sipp_2023_tip_donor, + impute_us_sipp_tips, + load_sipp_2023_tip_donor, + us_sipp_tips_signal_gate, + us_sipp_tips_stage_spec, + us_sipp_tips_summary, + with_us_sipp_tip_inputs, +) +from populace.build.us_runtime.sipp_vehicles import ( + SIPP_2023_VEHICLE_DONOR_REVISION, + SIPP_2023_VEHICLE_DONOR_SHA256, + SIPP_2023_VEHICLE_DONOR_SIZE_BYTES, + SIPP_2023_VEHICLE_DONOR_URL, + US_SIPP_VEHICLE_NONCONSTANT_HOUSEHOLD_COLUMNS, + US_SIPP_VEHICLE_OUTPUT_COLUMNS, + fetch_sipp_2023_vehicle_donor, + load_sipp_2023_vehicle_donor, + us_sipp_vehicles_signal_gate, + us_sipp_vehicles_stage_spec, + us_sipp_vehicles_summary, + with_us_sipp_vehicle_inputs, +) from populace.build.us_runtime.snap_discretionary_exemption import ( US_SNAP_DISCRETIONARY_EXEMPTION_NONCONSTANT_PERSON_COLUMNS, US_SNAP_DISCRETIONARY_EXEMPTION_OUTPUT_COLUMN, @@ -307,6 +730,51 @@ disaggregate_us_puf_aggregate_records_from_manifest, us_source_operation_handlers, ) +from populace.build.us_runtime.ssi_disability_criteria import ( + SIPP_2023_SSI_DISABILITY_DONOR_REVISION, + SIPP_2023_SSI_DISABILITY_DONOR_SHA256, + SIPP_2023_SSI_DISABILITY_DONOR_SIZE_BYTES, + SIPP_2023_SSI_DISABILITY_DONOR_URL, + SIPP_SSI_DISABILITY_DIFFICULTY_PREDICTORS, + SIPP_SSI_DISABILITY_FIT_PARAMETERS, + SIPP_SSI_DISABILITY_MODEL_PREDICTORS, + SIPP_SSI_DISABILITY_READ_PARAMETERS, + SIPP_SSI_DISABILITY_SOURCE_COLUMNS, + SSI_DISABILITY_ARCHIVED_CPS_URL, + SSI_DISABILITY_ARCHIVED_EXTENDED_CPS_URL, + SSI_DISABILITY_ARCHIVED_SIPP_URL, + SSI_DISABILITY_ARCHIVED_SOURCE_IMPUTE_URL, + SSI_DISABILITY_SIPP_DICTIONARY_URL, + US_SSI_DISABILITY_CRITERIA_NONCONSTANT_PERSON_COLUMNS, + US_SSI_DISABILITY_CRITERIA_OUTPUT_COLUMNS, + US_SSI_DISABILITY_CRITERIA_STAGE_NAME, + fetch_sipp_2023_ssi_disability_donor, + impute_us_ssi_disability_criteria, + load_sipp_2023_ssi_disability_donor, + us_ssi_disability_criteria_signal_gate, + us_ssi_disability_criteria_stage_spec, + us_ssi_disability_criteria_summary, + with_us_ssi_disability_criteria, +) +from populace.build.us_runtime.ssi_take_up import ( + SSI_TAKE_UP_ARCHIVED_DERIVATION_URL, + SSI_TAKE_UP_ARCHIVED_EXPORT_URL, + SSI_TAKE_UP_ARCHIVED_RANDOMNESS_URL, + SSI_TAKE_UP_ARCHIVED_REPORTER_URL, + SSI_TAKE_UP_ARCHIVED_TARGETS_URL, + SSI_TAKE_UP_SSA_SOURCE_URL, + US_SSI_TAKE_UP_ANCHOR, + US_SSI_TAKE_UP_NONCONSTANT_PERSON_COLUMNS, + US_SSI_TAKE_UP_OUTPUT_COLUMNS, + US_SSI_TAKE_UP_REQUIRED_SOURCE_COLUMNS, + US_SSI_TAKE_UP_STAGE_NAME, + US_SSI_TAKE_UP_TARGET_TABLE_NAME, + us_ssi_take_up_diagnostics, + us_ssi_take_up_gate, + us_ssi_take_up_reporter_source_ids, + with_us_ssi_take_up, + write_us_ssi_take_up_diagnostics, +) from populace.build.us_runtime.take_up import ( US_TAKE_UP_SHARE_BAND, SeededTakeUpResult, @@ -332,6 +800,95 @@ us_source_stage_outputs, us_validation_input_coverage_gate, ) +from populace.build.us_runtime.voluntary_filing import ( + SIPP_2023_VOLUNTARY_FILING_DONOR_REVISION, + SIPP_2023_VOLUNTARY_FILING_DONOR_SHA256, + SIPP_2023_VOLUNTARY_FILING_DONOR_SIZE_BYTES, + SIPP_2023_VOLUNTARY_FILING_DONOR_URL, + SIPP_VOLUNTARY_FILING_MODEL_PREDICTORS, + SIPP_VOLUNTARY_FILING_SOURCE_COLUMNS, + US_VOLUNTARY_FILING_NONCONSTANT_TAX_UNIT_COLUMNS, + US_VOLUNTARY_FILING_OUTPUT_COLUMNS, + US_VOLUNTARY_FILING_STAGE_NAME, + VOLUNTARY_FILING_ARCHIVED_DERIVATION_URL, + VOLUNTARY_FILING_ARCHIVED_PARAMETERS_URL, + VOLUNTARY_FILING_SIPP_DICTIONARY_URL, + fetch_sipp_2023_voluntary_filing_donor, + impute_us_voluntary_filing, + load_sipp_2023_voluntary_filing_donor, + us_voluntary_filing_signal_gate, + us_voluntary_filing_stage_spec, + us_voluntary_filing_summary, + with_us_voluntary_filing_input, +) +from populace.build.us_runtime.weeks_unemployed import ( + ASEC_2023_WEEKS_UNEMPLOYED_MEMBER, + ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_CRC32, + ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_SHA256, + ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_SIZE_BYTES, + ASEC_2023_WEEKS_UNEMPLOYED_RAW_ROWS, + ASEC_2023_WEEKS_UNEMPLOYED_SOURCE_COLUMNS, + ASEC_2023_WEEKS_UNEMPLOYED_SOURCE_SHA256, + ASEC_2023_WEEKS_UNEMPLOYED_SOURCE_YEAR, + ASEC_2023_WEEKS_UNEMPLOYED_UNIQUE_KEYS, + ASEC_2023_WEEKS_UNEMPLOYED_WEIGHTED_SOURCE_SHARE, + ASEC_2023_WEEKS_UNEMPLOYED_WEIGHTED_WEEKS, + ASEC_2023_WEEKS_UNEMPLOYED_ZIP_SHA256, + ASEC_2023_WEEKS_UNEMPLOYED_ZIP_SIZE_BYTES, + ASEC_2023_WEEKS_UNEMPLOYED_ZIP_URL, + US_WEEKS_UNEMPLOYED_NONCONSTANT_PERSON_COLUMNS, + US_WEEKS_UNEMPLOYED_OUTPUT_COLUMNS, + US_WEEKS_UNEMPLOYED_REQUIRED_SOURCE_COLUMNS, + US_WEEKS_UNEMPLOYED_STAGE_NAME, + WEEKS_UNEMPLOYED_ARCHIVED_DERIVATION_URL, + WEEKS_UNEMPLOYED_ARCHIVED_PUF_IMPUTATION_URL, + WEEKS_UNEMPLOYED_ARCHIVED_SOURCE_URL, + WEEKS_UNEMPLOYED_DERIVE_PARAMETERS, + WEEKS_UNEMPLOYED_PUF_IMPUTATION_PARAMETERS, + WEEKS_UNEMPLOYED_PUF_PREDICTORS, + WEEKS_UNEMPLOYED_READ_PARAMETERS, + derive_us_weeks_unemployed_from_manifest, + fetch_asec_2023_weeks_unemployed_source, + fill_asec_2022_weeks_unemployed_source, + impute_us_weeks_unemployed_to_puf_support_from_manifest, + load_asec_2023_weeks_unemployed_source, + us_weeks_unemployed_signal_gate, + us_weeks_unemployed_stage_spec, + us_weeks_unemployed_summary, + with_us_weeks_unemployed, +) +from populace.build.us_runtime.wic_claim import ( + US_WIC_CLAIM_NONCONSTANT_PERSON_COLUMNS, + US_WIC_CLAIM_OUTPUT_COLUMNS, + US_WIC_CLAIM_REQUIRED_SOURCE_COLUMNS, + US_WIC_CLAIM_STAGE_NAME, + WIC_CLAIM_ARCHIVED_DERIVATION_URL, + WIC_CLAIM_ARCHIVED_PARAMETERS_URL, + WIC_CLAIM_ARCHIVED_RANDOMNESS_URL, + WIC_CLAIM_FNS_SOURCE_URL, + derive_us_wic_claim_from_manifest, + us_wic_claim_signal_gate, + us_wic_claim_stage_spec, + us_wic_claim_summary, + with_us_wic_claim_input, +) +from populace.build.us_runtime.workers_compensation import ( + US_WORKERS_COMPENSATION_NONCONSTANT_PERSON_COLUMNS, + US_WORKERS_COMPENSATION_OUTPUT_COLUMNS, + US_WORKERS_COMPENSATION_REQUIRED_SOURCE_COLUMNS, + US_WORKERS_COMPENSATION_STAGE_NAME, + WORKERS_COMPENSATION_ARCHIVED_DERIVATION_URL, + WORKERS_COMPENSATION_ARCHIVED_PUF_IMPUTATION_URL, + WORKERS_COMPENSATION_ARCHIVED_PUF_OUTPUTS_URL, + WORKERS_COMPENSATION_ARCHIVED_SOURCE_COLUMNS_URL, + derive_us_workers_compensation_from_asec, + derive_us_workers_compensation_from_manifest, + impute_us_workers_compensation_to_puf_support_from_manifest, + us_workers_compensation_signal_gate, + us_workers_compensation_stage_spec, + us_workers_compensation_summary, + with_us_workers_compensation, +) from populace.frame import Frame __all__ = [ @@ -401,6 +958,32 @@ "us_hours_worked_stage_spec", "us_hours_worked_summary", "with_us_hours_worked_inputs", + "ACS_2022_RENT_ARTIFACT_SHA256", + "HOUSING_INPUTS_ARCHIVED_ACS_DERIVATION_URL", + "HOUSING_INPUTS_ARCHIVED_CPS_RENT_URL", + "HOUSING_INPUTS_ARCHIVED_CPS_SPM_URL", + "HOUSING_INPUTS_ARCHIVED_PUF_IMPUTATION_URL", + "HOUSING_TAKE_UP_ARCHIVED_DERIVATION_URL", + "HOUSING_TAKE_UP_ARCHIVED_HUD_ETL_URL", + "HOUSING_TAKE_UP_ARCHIVED_PARAMETER_URL", + "US_HOUSING_HOUSEHOLD_OUTPUT_COLUMNS", + "US_HOUSING_INPUTS_OUTPUT_COLUMNS", + "US_HOUSING_INPUTS_STAGE_NAME", + "US_HOUSING_NONCONSTANT_HOUSEHOLD_COLUMNS", + "US_HOUSING_NONCONSTANT_PERSON_COLUMNS", + "US_HOUSING_NONCONSTANT_SPM_UNIT_COLUMNS", + "US_HOUSING_PERSON_OUTPUT_COLUMNS", + "US_HOUSING_REQUIRED_HOUSEHOLD_SOURCE_COLUMNS", + "US_HOUSING_REQUIRED_PERSON_SOURCE_COLUMNS", + "US_HOUSING_SPM_UNIT_OUTPUT_COLUMNS", + "derive_us_housing_inputs", + "impute_us_housing_assistance_to_puf_support", + "impute_us_pre_subsidy_rent", + "load_acs_2022_rent_donor", + "us_housing_inputs_signal_gate", + "us_housing_inputs_stage_spec", + "us_housing_inputs_summary", + "with_us_housing_inputs", "US_SNAP_TAKE_UP_OUTPUT_COLUMN", "US_SNAP_TAKE_UP_RAW_COLUMN", "US_SNAP_TAKE_UP_STAGE_NAME", @@ -444,18 +1027,468 @@ "us_eligibility_inputs_stage_spec", "us_eligibility_inputs_summary", "with_us_eligibility_inputs", + "US_RELATIONSHIP_INPUTS_NONCONSTANT_PERSON_COLUMNS", + "US_RELATIONSHIP_INPUTS_OUTPUT_COLUMNS", + "US_RELATIONSHIP_INPUTS_REQUIRED_SOURCE_COLUMNS", + "US_RELATIONSHIP_INPUTS_STAGE_NAME", + "derive_us_relationship_inputs_from_manifest", + "us_relationship_inputs_signal_gate", + "us_relationship_inputs_stage_spec", + "us_relationship_inputs_summary", + "with_us_relationship_inputs", + "ALIMONY_ASEC_ARCHIVED_DERIVATION_URL", + "ALIMONY_PUF_ARCHIVED_DERIVATION_URL", + "US_ALIMONY_NONCONSTANT_PERSON_COLUMNS", + "US_ALIMONY_OUTPUT_COLUMNS", + "US_ALIMONY_STAGE_NAME", + "derive_us_alimony_from_asec", + "derive_us_alimony_from_puf", + "us_alimony_signal_gate", + "us_alimony_stage_spec", + "us_alimony_summary", + "US_CASUALTY_LOSS_NONCONSTANT_PERSON_COLUMNS", + "US_CASUALTY_LOSS_OUTPUT_COLUMNS", + "US_CASUALTY_LOSS_STAGE_NAME", + "derive_us_casualty_loss_from_puf", + "us_casualty_loss_signal_gate", + "us_casualty_loss_stage_spec", + "us_casualty_loss_summary", + "CAPITAL_GAIN_DETAILS_ARCHIVED_DERIVATION_URL", + "CAPITAL_GAIN_DETAILS_ARCHIVED_EXPORT_URL", + "CAPITAL_GAIN_DETAILS_ARCHIVED_IMPUTATION_URL", + "CAPITAL_GAIN_DETAILS_ARCHIVED_PERSON_ALLOCATION_URL", + "CAPITAL_GAIN_DETAILS_ARCHIVED_PUF_ARTIFACT_URL", + "US_CAPITAL_GAIN_DETAILS_NONCONSTANT_PERSON_COLUMNS", + "US_CAPITAL_GAIN_DETAILS_NONCONSTANT_TAX_UNIT_COLUMNS", + "US_CAPITAL_GAIN_DETAILS_OUTPUT_COLUMNS", + "US_CAPITAL_GAIN_DETAILS_STAGE_NAME", + "derive_us_capital_gain_details_from_puf", + "us_capital_gain_details_signal_gate", + "us_capital_gain_details_stage_spec", + "us_capital_gain_details_summary", + "DOMESTIC_PRODUCTION_ALD_ARCHIVED_DERIVATION_URL", + "DOMESTIC_PRODUCTION_ALD_ARCHIVED_EXPORT_URL", + "DOMESTIC_PRODUCTION_ALD_ARCHIVED_IMPUTATION_URL", + "DOMESTIC_PRODUCTION_ALD_ARCHIVED_PUF_ARTIFACT_URL", + "US_DOMESTIC_PRODUCTION_ALD_NONCONSTANT_TAX_UNIT_COLUMNS", + "US_DOMESTIC_PRODUCTION_ALD_OUTPUT_COLUMNS", + "US_DOMESTIC_PRODUCTION_ALD_STAGE_NAME", + "derive_us_domestic_production_ald_from_puf", + "us_domestic_production_ald_signal_gate", + "us_domestic_production_ald_stage_spec", + "us_domestic_production_ald_summary", + "US_CHILDCARE_OUTPUT_COLUMNS", + "US_CHILDCARE_REQUIRED_SOURCE_COLUMNS", + "US_CHILDCARE_STAGE_NAME", + "derive_us_childcare_from_manifest", + "impute_us_childcare_to_puf_support_from_manifest", + "us_childcare_signal_gate", + "us_childcare_stage_spec", + "us_childcare_summary", + "with_us_childcare_inputs", + "ENERGY_SUBSIDY_ARCHIVED_CPS_DERIVATION_URL", + "ENERGY_SUBSIDY_ARCHIVED_PUF_IMPUTATION_URL", + "US_ENERGY_SUBSIDY_OUTPUT_COLUMNS", + "US_ENERGY_SUBSIDY_REQUIRED_SOURCE_COLUMNS", + "US_ENERGY_SUBSIDY_STAGE_NAME", + "derive_us_energy_subsidy_from_manifest", + "impute_us_energy_subsidy_to_puf_support_from_manifest", + "us_energy_subsidy_signal_gate", + "us_energy_subsidy_stage_spec", + "us_energy_subsidy_summary", + "with_us_energy_subsidy_input", + "CHILD_SUPPORT_ARCHIVED_PUF_IMPUTATION_URL", + "CHILD_SUPPORT_ARCHIVED_PUF_OUTPUTS_URL", + "CHILD_SUPPORT_EXPENSE_ARCHIVED_DERIVATION_URL", + "CHILD_SUPPORT_RECEIVED_ARCHIVED_DERIVATION_URL", + "US_CHILD_SUPPORT_NONCONSTANT_PERSON_COLUMNS", + "US_CHILD_SUPPORT_OUTPUT_COLUMNS", + "US_CHILD_SUPPORT_REQUIRED_SOURCE_COLUMNS", + "US_CHILD_SUPPORT_STAGE_NAME", + "derive_us_child_support_from_asec", + "derive_us_child_support_from_manifest", + "impute_us_child_support_to_puf_support_from_manifest", + "us_child_support_signal_gate", + "us_child_support_stage_spec", + "us_child_support_summary", + "with_us_child_support_inputs", + "DISABILITY_BENEFITS_ARCHIVED_DERIVATION_URL", + "DISABILITY_BENEFITS_ARCHIVED_PUF_IMPUTATION_URL", + "DISABILITY_BENEFITS_ARCHIVED_PUF_OUTPUTS_URL", + "DISABILITY_BENEFITS_ARCHIVED_SOURCE_COLUMNS_URL", + "US_DISABILITY_BENEFITS_NONCONSTANT_PERSON_COLUMNS", + "US_DISABILITY_BENEFITS_OUTPUT_COLUMNS", + "US_DISABILITY_BENEFITS_REQUIRED_SOURCE_COLUMNS", + "US_DISABILITY_BENEFITS_STAGE_NAME", + "derive_us_disability_benefits_from_asec", + "derive_us_disability_benefits_from_manifest", + "impute_us_disability_benefits_to_puf_support_from_manifest", + "us_disability_benefits_signal_gate", + "us_disability_benefits_stage_spec", + "us_disability_benefits_summary", + "with_us_disability_benefits", + "US_WORKERS_COMPENSATION_NONCONSTANT_PERSON_COLUMNS", + "US_WORKERS_COMPENSATION_OUTPUT_COLUMNS", + "US_WORKERS_COMPENSATION_REQUIRED_SOURCE_COLUMNS", + "US_WORKERS_COMPENSATION_STAGE_NAME", + "WORKERS_COMPENSATION_ARCHIVED_DERIVATION_URL", + "WORKERS_COMPENSATION_ARCHIVED_PUF_IMPUTATION_URL", + "WORKERS_COMPENSATION_ARCHIVED_PUF_OUTPUTS_URL", + "WORKERS_COMPENSATION_ARCHIVED_SOURCE_COLUMNS_URL", + "derive_us_workers_compensation_from_asec", + "derive_us_workers_compensation_from_manifest", + "impute_us_workers_compensation_to_puf_support_from_manifest", + "us_workers_compensation_signal_gate", + "us_workers_compensation_stage_spec", + "us_workers_compensation_summary", + "with_us_workers_compensation", + "ASEC_2023_WEEKS_UNEMPLOYED_MEMBER", + "ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_CRC32", + "ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_SHA256", + "ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_SIZE_BYTES", + "ASEC_2023_WEEKS_UNEMPLOYED_RAW_ROWS", + "ASEC_2023_WEEKS_UNEMPLOYED_SOURCE_COLUMNS", + "ASEC_2023_WEEKS_UNEMPLOYED_SOURCE_SHA256", + "ASEC_2023_WEEKS_UNEMPLOYED_SOURCE_YEAR", + "ASEC_2023_WEEKS_UNEMPLOYED_UNIQUE_KEYS", + "ASEC_2023_WEEKS_UNEMPLOYED_WEIGHTED_SOURCE_SHARE", + "ASEC_2023_WEEKS_UNEMPLOYED_WEIGHTED_WEEKS", + "ASEC_2023_WEEKS_UNEMPLOYED_ZIP_SHA256", + "ASEC_2023_WEEKS_UNEMPLOYED_ZIP_SIZE_BYTES", + "ASEC_2023_WEEKS_UNEMPLOYED_ZIP_URL", + "US_WEEKS_UNEMPLOYED_NONCONSTANT_PERSON_COLUMNS", + "US_WEEKS_UNEMPLOYED_OUTPUT_COLUMNS", + "US_WEEKS_UNEMPLOYED_REQUIRED_SOURCE_COLUMNS", + "US_WEEKS_UNEMPLOYED_STAGE_NAME", + "WEEKS_UNEMPLOYED_ARCHIVED_DERIVATION_URL", + "WEEKS_UNEMPLOYED_ARCHIVED_PUF_IMPUTATION_URL", + "WEEKS_UNEMPLOYED_ARCHIVED_SOURCE_URL", + "WEEKS_UNEMPLOYED_DERIVE_PARAMETERS", + "WEEKS_UNEMPLOYED_PUF_IMPUTATION_PARAMETERS", + "WEEKS_UNEMPLOYED_PUF_PREDICTORS", + "WEEKS_UNEMPLOYED_READ_PARAMETERS", + "derive_us_weeks_unemployed_from_manifest", + "fetch_asec_2023_weeks_unemployed_source", + "fill_asec_2022_weeks_unemployed_source", + "impute_us_weeks_unemployed_to_puf_support_from_manifest", + "load_asec_2023_weeks_unemployed_source", + "us_weeks_unemployed_signal_gate", + "us_weeks_unemployed_stage_spec", + "us_weeks_unemployed_summary", + "with_us_weeks_unemployed", + "US_WIC_CLAIM_NONCONSTANT_PERSON_COLUMNS", + "US_WIC_CLAIM_OUTPUT_COLUMNS", + "US_WIC_CLAIM_REQUIRED_SOURCE_COLUMNS", + "US_WIC_CLAIM_STAGE_NAME", + "WIC_CLAIM_ARCHIVED_DERIVATION_URL", + "WIC_CLAIM_ARCHIVED_PARAMETERS_URL", + "WIC_CLAIM_ARCHIVED_RANDOMNESS_URL", + "WIC_CLAIM_FNS_SOURCE_URL", + "derive_us_wic_claim_from_manifest", + "us_wic_claim_signal_gate", + "us_wic_claim_stage_spec", + "us_wic_claim_summary", + "with_us_wic_claim_input", + "US_MISC_ITEMIZED_NONCONSTANT_PERSON_COLUMNS", + "US_MISC_ITEMIZED_OUTPUT_COLUMNS", + "US_MISC_ITEMIZED_STAGE_NAME", + "derive_us_misc_itemized_from_puf", + "us_misc_itemized_signal_gate", + "us_misc_itemized_stage_spec", + "us_misc_itemized_summary", + "FORM_4952_ARCHIVED_DERIVATION_URL", + "FORM_4952_ARCHIVED_EXPORT_URL", + "FORM_4952_ARCHIVED_IMPUTATION_URL", + "FORM_4952_ARCHIVED_PERSON_ALLOCATION_URL", + "FORM_4952_ARCHIVED_PUF_ARTIFACT_URL", + "US_FORM_4952_NONCONSTANT_PERSON_COLUMNS", + "US_FORM_4952_OUTPUT_COLUMNS", + "US_FORM_4952_STAGE_NAME", + "derive_us_form_4952_election_from_puf", + "us_form_4952_election_signal_gate", + "us_form_4952_election_stage_spec", + "us_form_4952_election_summary", + "SALT_REFUND_ARCHIVED_DERIVATION_URL", + "SALT_REFUND_ARCHIVED_EXPORT_URL", + "SALT_REFUND_ARCHIVED_IMPUTATION_URL", + "SALT_REFUND_ARCHIVED_PERSON_ALLOCATION_URL", + "SALT_REFUND_ARCHIVED_PUF_ARTIFACT_URL", + "US_SALT_REFUND_NONCONSTANT_PERSON_COLUMNS", + "US_SALT_REFUND_OUTPUT_COLUMNS", + "US_SALT_REFUND_STAGE_NAME", + "derive_us_salt_refund_income_from_puf", + "us_salt_refund_income_signal_gate", + "us_salt_refund_income_stage_spec", + "us_salt_refund_income_summary", + "QBI_ARCHIVED_ASSUMPTIONS_URL", + "QBI_ARCHIVED_CLONE_URL", + "QBI_ARCHIVED_DERIVATION_URL", + "QBI_ARCHIVED_EXPORT_URL", + "QBI_ARCHIVED_IMPUTATION_URL", + "QBI_ARCHIVED_PUF_ARTIFACT_URL", + "QBI_ARCHIVED_SIMULATION_URL", + "US_QBI_BOOLEAN_OUTPUT_COLUMNS", + "US_QBI_NONCONSTANT_PERSON_COLUMNS", + "US_QBI_NONNEGATIVE_OUTPUT_COLUMNS", + "US_QBI_OUTPUT_COLUMNS", + "US_QBI_STAGE_NAME", + "us_qbi_inputs_signal_gate", + "us_qbi_inputs_stage_spec", + "us_qbi_inputs_summary", + "with_us_qbi_input_reconciliation", + "US_AOTC_ELIGIBILITY_OUTPUT_COLUMNS", + "US_EDUCATION_INPUTS_NONCONSTANT_PERSON_COLUMNS", + "US_EDUCATION_INPUTS_OUTPUT_COLUMNS", + "US_EDUCATION_INPUTS_REQUIRED_SOURCE_COLUMNS", + "US_EDUCATION_INPUTS_STAGE_NAME", + "derive_us_education_inputs_from_manifest", + "us_education_inputs_signal_gate", + "us_education_inputs_stage_spec", + "us_education_inputs_summary", + "with_us_education_inputs", + "EDUCATOR_EXPENSE_ARCHIVED_ALLOCATION_URL", + "EDUCATOR_EXPENSE_ARCHIVED_DERIVATION_URL", + "EDUCATOR_EXPENSE_ARCHIVED_EXPORT_URL", + "EDUCATOR_EXPENSE_ARCHIVED_PUF_IMPUTATION_URL", + "US_EDUCATOR_EXPENSE_NONCONSTANT_PERSON_COLUMNS", + "US_EDUCATOR_EXPENSE_OUTPUT_COLUMNS", + "US_EDUCATOR_EXPENSE_STAGE_NAME", + "derive_us_educator_expense_from_puf", + "us_educator_expense_signal_gate", + "us_educator_expense_stage_spec", + "us_educator_expense_summary", + "FARM_BUSINESS_INCOME_ARCHIVED_CPS_FARM_INCOME_URL", + "FARM_BUSINESS_INCOME_ARCHIVED_DERIVATION_URL", + "FARM_BUSINESS_INCOME_ARCHIVED_EXPORT_URL", + "FARM_BUSINESS_INCOME_ARCHIVED_IMPUTATION_URL", + "FARM_BUSINESS_INCOME_ARCHIVED_OVERRIDE_URL", + "FARM_BUSINESS_INCOME_ARCHIVED_PUF_ARTIFACT_URL", + "US_FARM_BUSINESS_INCOME_NONCONSTANT_PERSON_COLUMNS", + "US_FARM_BUSINESS_INCOME_OUTPUT_COLUMNS", + "US_FARM_BUSINESS_INCOME_STAGE_NAME", + "derive_us_farm_business_income_from_puf", + "us_farm_business_income_signal_gate", + "us_farm_business_income_stage_spec", + "us_farm_business_income_summary", + "US_RETIREMENT_CONTRIBUTION_NONCONSTANT_PERSON_COLUMNS", + "US_RETIREMENT_CONTRIBUTION_OUTPUT_COLUMNS", + "US_RETIREMENT_CONTRIBUTION_REQUIRED_SOURCE_COLUMNS", + "US_RETIREMENT_CONTRIBUTION_STAGE_NAME", + "derive_us_retirement_contributions_from_manifest", + "impute_us_retirement_contributions_to_puf_support_from_manifest", + "us_retirement_contributions_signal_gate", + "us_retirement_contributions_stage_spec", + "us_retirement_contributions_summary", + "with_us_retirement_contribution_inputs", + "RETIREMENT_DISTRIBUTIONS_ARCHIVED_DERIVATION_URL", + "RETIREMENT_DISTRIBUTIONS_ARCHIVED_PARAMETERS_URL", + "US_RETIREMENT_DISTRIBUTION_NONCONSTANT_PERSON_COLUMNS", + "US_RETIREMENT_DISTRIBUTION_OUTPUT_COLUMNS", + "US_RETIREMENT_DISTRIBUTION_REQUIRED_SOURCE_COLUMNS", + "US_RETIREMENT_DISTRIBUTION_STAGE_NAME", + "derive_us_retirement_distributions_from_manifest", + "impute_us_retirement_distributions_to_puf_support_from_manifest", + "us_retirement_distributions_signal_gate", + "us_retirement_distributions_stage_spec", + "us_retirement_distributions_summary", + "with_us_retirement_distribution_inputs", + "QUALIFIED_AUTO_LOAN_ANNUAL_ISSUANCE_TARGET", + "SCF_2022_FULL_EXTRACT_MEMBER", + "SCF_2022_FULL_EXTRACT_MEMBER_SHA256", + "SCF_2022_FULL_EXTRACT_URL", + "SCF_2022_FULL_EXTRACT_ZIP_SHA256", + "SCF_AUTO_LOAN_AMOUNT_COLUMNS", + "SCF_AUTO_LOAN_RATE_COLUMNS", + "US_SCF_AUTO_LOAN_NONCONSTANT_HOUSEHOLD_COLUMNS", + "US_SCF_AUTO_LOAN_OUTPUT_COLUMNS", + "fetch_scf_2022_full_extract", + "impute_us_scf_auto_loans", + "load_scf_2022_auto_loan_donor", + "qualified_auto_loan_interest_proxy", + "us_scf_auto_loans_signal_gate", + "us_scf_auto_loans_stage_spec", + "us_scf_auto_loans_summary", + "with_us_scf_auto_loan_inputs", "SCF_FINANCIAL_ASSET_TARGET_COMPONENTS", + "SCF_NET_WORTH_TARGET_COMPONENTS", "SCF_WEALTH_PREDICTORS", "US_SCF_FINANCIAL_ASSET_OUTPUT_COLUMNS", + "US_SCF_NET_WORTH_OUTPUT_COLUMNS", + "US_SCF_WEALTH_NONCONSTANT_HOUSEHOLD_COLUMNS", "US_SCF_WEALTH_NONCONSTANT_PERSON_COLUMNS", "US_SCF_WEALTH_STAGE_NAME", "fetch_scf_2022_summary_extract", "impute_us_scf_financial_assets", + "impute_us_scf_net_worth", "load_scf_2022_financial_asset_donor", "us_scf_wealth_signal_gate", "us_scf_wealth_stage_spec", "us_scf_wealth_summary", "with_us_scf_wealth_inputs", + "HEAD_START_SIPP_DICTIONARY_URL", + "SIPP_2023_HEAD_START_DONOR_REVISION", + "SIPP_2023_HEAD_START_DONOR_SHA256", + "SIPP_2023_HEAD_START_DONOR_SIZE_BYTES", + "SIPP_2023_HEAD_START_DONOR_URL", + "SIPP_HEAD_START_FIT_PARAMETERS", + "SIPP_HEAD_START_MODEL_PREDICTORS", + "SIPP_HEAD_START_READ_PARAMETERS", + "SIPP_HEAD_START_SOURCE_COLUMNS", + "US_SIPP_HEAD_START_NONCONSTANT_PERSON_COLUMNS", + "US_SIPP_HEAD_START_OUTPUT_COLUMNS", + "US_SIPP_HEAD_START_REQUIRED_SOURCE_COLUMNS", + "US_SIPP_HEAD_START_STAGE_NAME", + "fetch_sipp_2023_head_start_donor", + "impute_us_sipp_head_start", + "load_sipp_2023_head_start_donor", + "us_sipp_head_start_signal_gate", + "us_sipp_head_start_stage_spec", + "us_sipp_head_start_summary", + "with_us_sipp_head_start_input", + "SIPP_2023_SSI_DISABILITY_DONOR_REVISION", + "SIPP_2023_SSI_DISABILITY_DONOR_SHA256", + "SIPP_2023_SSI_DISABILITY_DONOR_SIZE_BYTES", + "SIPP_2023_SSI_DISABILITY_DONOR_URL", + "SIPP_SSI_DISABILITY_DIFFICULTY_PREDICTORS", + "SIPP_SSI_DISABILITY_FIT_PARAMETERS", + "SIPP_SSI_DISABILITY_MODEL_PREDICTORS", + "SIPP_SSI_DISABILITY_READ_PARAMETERS", + "SIPP_SSI_DISABILITY_SOURCE_COLUMNS", + "SSI_DISABILITY_ARCHIVED_CPS_URL", + "SSI_DISABILITY_ARCHIVED_EXTENDED_CPS_URL", + "SSI_DISABILITY_ARCHIVED_SIPP_URL", + "SSI_DISABILITY_ARCHIVED_SOURCE_IMPUTE_URL", + "SSI_DISABILITY_SIPP_DICTIONARY_URL", + "US_SSI_DISABILITY_CRITERIA_NONCONSTANT_PERSON_COLUMNS", + "US_SSI_DISABILITY_CRITERIA_OUTPUT_COLUMNS", + "US_SSI_DISABILITY_CRITERIA_STAGE_NAME", + "fetch_sipp_2023_ssi_disability_donor", + "impute_us_ssi_disability_criteria", + "load_sipp_2023_ssi_disability_donor", + "us_ssi_disability_criteria_signal_gate", + "us_ssi_disability_criteria_stage_spec", + "us_ssi_disability_criteria_summary", + "with_us_ssi_disability_criteria", + "SSI_TAKE_UP_ARCHIVED_DERIVATION_URL", + "SSI_TAKE_UP_ARCHIVED_EXPORT_URL", + "SSI_TAKE_UP_ARCHIVED_RANDOMNESS_URL", + "SSI_TAKE_UP_ARCHIVED_REPORTER_URL", + "SSI_TAKE_UP_ARCHIVED_TARGETS_URL", + "SSI_TAKE_UP_SSA_SOURCE_URL", + "US_SSI_TAKE_UP_ANCHOR", + "US_SSI_TAKE_UP_NONCONSTANT_PERSON_COLUMNS", + "US_SSI_TAKE_UP_OUTPUT_COLUMNS", + "US_SSI_TAKE_UP_REQUIRED_SOURCE_COLUMNS", + "US_SSI_TAKE_UP_STAGE_NAME", + "US_SSI_TAKE_UP_TARGET_TABLE_NAME", + "us_ssi_take_up_diagnostics", + "us_ssi_take_up_gate", + "us_ssi_take_up_reporter_source_ids", + "with_us_ssi_take_up", + "write_us_ssi_take_up_diagnostics", + "CENSUS_OCCUPATION_CODE_TO_TTOC", + "SIPP_2023_TIP_DONOR_REVISION", + "SIPP_2023_TIP_DONOR_SHA256", + "SIPP_2023_TIP_DONOR_URL", + "SIPP_TIP_OUTPUT_COLUMNS", + "SIPP_TIP_PREDICTORS", + "US_SIPP_TIPS_NONCONSTANT_PERSON_COLUMNS", + "US_SIPP_TIPS_OUTPUT_COLUMNS", + "US_SIPP_TIPS_REQUIRED_SOURCE_COLUMNS", + "US_SIPP_TIPS_STAGE_NAME", + "derive_treasury_tipped_occupation_code", + "fetch_sipp_2023_tip_donor", + "impute_us_sipp_tips", + "load_sipp_2023_tip_donor", + "us_sipp_tips_signal_gate", + "us_sipp_tips_stage_spec", + "us_sipp_tips_summary", + "with_us_sipp_tip_inputs", + "SIPP_2023_VEHICLE_DONOR_REVISION", + "SIPP_2023_VEHICLE_DONOR_SHA256", + "SIPP_2023_VEHICLE_DONOR_SIZE_BYTES", + "SIPP_2023_VEHICLE_DONOR_URL", + "US_SIPP_VEHICLE_NONCONSTANT_HOUSEHOLD_COLUMNS", + "US_SIPP_VEHICLE_OUTPUT_COLUMNS", + "fetch_sipp_2023_vehicle_donor", + "load_sipp_2023_vehicle_donor", + "us_sipp_vehicles_signal_gate", + "us_sipp_vehicles_stage_spec", + "us_sipp_vehicles_summary", + "with_us_sipp_vehicle_inputs", + "SIPP_2023_VOLUNTARY_FILING_DONOR_REVISION", + "SIPP_2023_VOLUNTARY_FILING_DONOR_SHA256", + "SIPP_2023_VOLUNTARY_FILING_DONOR_SIZE_BYTES", + "SIPP_2023_VOLUNTARY_FILING_DONOR_URL", + "SIPP_VOLUNTARY_FILING_MODEL_PREDICTORS", + "SIPP_VOLUNTARY_FILING_SOURCE_COLUMNS", + "US_VOLUNTARY_FILING_NONCONSTANT_TAX_UNIT_COLUMNS", + "US_VOLUNTARY_FILING_OUTPUT_COLUMNS", + "US_VOLUNTARY_FILING_STAGE_NAME", + "VOLUNTARY_FILING_ARCHIVED_DERIVATION_URL", + "VOLUNTARY_FILING_ARCHIVED_PARAMETERS_URL", + "VOLUNTARY_FILING_SIPP_DICTIONARY_URL", + "fetch_sipp_2023_voluntary_filing_donor", + "impute_us_voluntary_filing", + "load_sipp_2023_voluntary_filing_donor", + "us_voluntary_filing_signal_gate", + "us_voluntary_filing_stage_spec", + "us_voluntary_filing_summary", + "with_us_voluntary_filing_input", + "BLS_STATE_UNION_REPRESENTATION_RATE_2024", + "FLSA_EXECUTIVE_ADMINISTRATIVE_PROFESSIONAL_OCCUPATION_CODES", + "FLSA_OVERTIME_OCCUPATION_CODES", + "ORG_2024_DONOR_CONTENT_SHA256", + "ORG_2024_DONOR_FILENAME", + "ORG_PREDICTORS", + "US_ORG_WAGES_NONCONSTANT_PERSON_COLUMNS", + "US_ORG_WAGES_OUTPUT_COLUMNS", + "US_ORG_WAGES_REQUIRED_SOURCE_COLUMNS", + "US_ORG_WAGES_STAGE_NAME", + "derive_flsa_overtime_premium", + "derive_us_org_occupation_inputs", + "fetch_org_2024_donor", + "impute_us_org_wages", + "load_org_2024_donor", + "us_org_wages_signal_gate", + "us_org_wages_stage_spec", + "us_org_wages_summary", + "with_us_org_wages_inputs", + "OTHER_HEALTH_INSURANCE_ARCHIVED_DERIVATION_URL", + "OTHER_HEALTH_INSURANCE_ARCHIVED_PUF_IMPUTATION_URL", + "OTHER_HEALTH_INSURANCE_ARCHIVED_PUF_OUTPUTS_URL", + "OTHER_HEALTH_INSURANCE_ARCHIVED_PUF_PREDICTORS_URL", + "OTHER_HEALTH_INSURANCE_ARCHIVED_PUF_SPLICE_URL", + "US_OTHER_HEALTH_INSURANCE_MODELED_PREMIUM_VARIABLES", + "US_OTHER_HEALTH_INSURANCE_NONCONSTANT_PERSON_COLUMNS", + "US_OTHER_HEALTH_INSURANCE_OUTPUT_COLUMNS", + "US_OTHER_HEALTH_INSURANCE_REQUIRED_SOURCE_COLUMNS", + "US_OTHER_HEALTH_INSURANCE_STAGE_NAME", + "derive_us_other_health_insurance_from_asec", + "derive_us_other_health_insurance_from_manifest", + "impute_us_other_health_insurance_to_puf_support_from_manifest", + "us_other_health_insurance_signal_gate", + "us_other_health_insurance_stage_spec", + "us_other_health_insurance_summary", + "with_us_other_health_insurance_inputs", + "PRIOR_YEAR_INCOME_ARCHIVED_DERIVATION_URL", + "PRIOR_YEAR_INCOME_ARCHIVED_FINALIZER_URL", + "PRIOR_YEAR_INCOME_ARCHIVED_FORMULA_OUTPUT_URL", + "PRIOR_YEAR_INCOME_ARCHIVED_PUF_IMPUTATION_URL", + "PRIOR_YEAR_INCOME_ARCHIVED_PUF_OUTPUTS_URL", + "PRIOR_YEAR_INCOME_ARCHIVED_PUF_SPLICE_URL", + "US_PRIOR_YEAR_INCOME_NONCONSTANT_PERSON_COLUMNS", + "US_PRIOR_YEAR_INCOME_OUTPUT_COLUMNS", + "US_PRIOR_YEAR_INCOME_PERSISTED_OUTPUT_COLUMNS", + "US_PRIOR_YEAR_INCOME_REQUIRED_SOURCE_COLUMNS", + "US_PRIOR_YEAR_INCOME_STAGE_NAME", + "derive_us_prior_year_income_from_manifest", + "impute_us_prior_year_income_to_puf_support_from_manifest", + "us_prior_year_income_signal_gate", + "us_prior_year_income_source_reconciliation_gate", + "us_prior_year_income_stage_spec", + "us_prior_year_income_summary", + "with_us_prior_year_income_inputs", "derive_us_immigration_status_from_manifest", "us_immigration_composition_gate", "us_immigration_composition_summary", @@ -474,6 +1507,7 @@ "apply_us_medicaid_enrollment_substitutions", "count_calibrated_take_up_programs", "us_medicaid_take_up_diagnostics", + "us_medicaid_source_person_table", "us_medicaid_take_up_gate", "us_take_up_participation_diagnostics", "us_take_up_signal_gate", @@ -482,6 +1516,19 @@ "with_us_take_up_inputs", "write_us_medicaid_take_up_diagnostics", "write_us_take_up_participation_diagnostics", + "MEDICARE_TAKE_UP_ARCHIVED_CLONE_URL", + "MEDICARE_TAKE_UP_ARCHIVED_DERIVATION_URL", + "MEDICARE_TAKE_UP_ARCHIVED_EXPORT_URL", + "MEDICARE_TAKE_UP_ARCHIVED_SOURCE_COLUMNS_URL", + "US_MEDICARE_TAKE_UP_NONCONSTANT_PERSON_COLUMNS", + "US_MEDICARE_TAKE_UP_OUTPUT_COLUMNS", + "US_MEDICARE_TAKE_UP_REQUIRED_SOURCE_COLUMNS", + "US_MEDICARE_TAKE_UP_STAGE_NAME", + "derive_us_medicare_take_up_from_manifest", + "us_medicare_take_up_signal_gate", + "us_medicare_take_up_stage_spec", + "us_medicare_take_up_summary", + "with_us_medicare_take_up_input", "TakeUpContract", "TakeUpProgram", "assert_take_up_contract_current", @@ -544,6 +1591,7 @@ "ValidationInputLeaf", "assert_validation_leaf_registry_current", "SSI_COUNTABLE_RESOURCE_ASSETS", + "POST_REFERENCE_ECPS_REQUIRED_INPUTS", "US_RELEASE_INPUT_COVERAGE_RESOURCE", "ReformCoverageProbe", "ReleaseInputColumn", @@ -641,8 +1689,38 @@ def to_manifest(self) -> dict[str, object]: survey="Fed SCF 2022", source="https://www.federalreserve.gov/econres/scfindex.htm", notes=( - "Wealth components (accounts, stocks, bonds, debts) and net " - "worth; household grain, head-carried person assets." + "Wealth components and debts; household-grain auto loans use the " + "full SCF while liquid assets are head-carried to person." + ), + ), + US_SSI_DISABILITY_CRITERIA_STAGE_NAME: DonorSpec( + survey="Census SIPP", + source="https://www.census.gov/programs-surveys/sipp.html", + notes=( + "Latent under-65 SSI disability/blindness criterion from the " + "pinned full 2023 public-use donor. ASEC and PUF-support people " + "are predicted separately, and only direct under-65 ASEC SSI " + "reporters receive the observed-reporter anchor." + ), + ), + US_SIPP_HEAD_START_STAGE_NAME: DonorSpec( + survey="Census SIPP", + source="https://www.census.gov/programs-surveys/sipp.html", + notes=( + "Direct December nursery/preschool federally sponsored-program " + "responses train a weighted Head Start take-up proxy for ages " + "3--5; strict reported structural negatives exclude hot-decked " + "answers and the prediction is shared by support clones." + ), + ), + US_SSI_TAKE_UP_STAGE_NAME: DonorSpec( + survey="CPS ASEC reported SSI + SSA SSI Monthly Statistics December 2024", + source=SSI_TAKE_UP_SSA_SOURCE_URL, + notes=( + "Direct ASEC SSI_VAL reporters anchor person-level take-up; " + "eligible source-person identities are count-calibrated by age to " + "SSA December 2024 Federal-payment recipient counts and fanned to " + "both support clones." ), ), "sipp_tips": DonorSpec( @@ -652,7 +1730,7 @@ def to_manifest(self) -> dict[str, object]: ), "org_wages": DonorSpec( survey="CPS ORG", - source="https://www.bls.gov/cps/earnings.htm", + source=("https://www2.census.gov/programs-surveys/cps/datasets/2024/basic/"), notes=( "Hourly-wage labor-market inputs. Donor load failures abort the " "build — the silent zero-fallback this stage once had is " @@ -701,10 +1779,25 @@ def to_manifest(self) -> dict[str, object]: "state-dependent CPS underreporting." ), ), - "prior_year_income": DonorSpec( + US_OTHER_HEALTH_INSURANCE_STAGE_NAME: DonorSpec( + survey="Census CPS ASEC", + source="https://www.census.gov/programs-surveys/cps.html", + notes=( + "Measured PHIP_VAL is reduced by PolicyEngine-calculated CHIP, " + "Marketplace, and Medicaid premiums after take-up stages; an " + "ASEC-trained weighted QRF replaces both premium leaves on the " + "PUF support half." + ), + ), + US_PRIOR_YEAR_INCOME_STAGE_NAME: DonorSpec( survey="CPS ASEC (prior year)", source="https://www.census.gov/programs-surveys/cps.html", - notes="PERIDNUM longitudinal join for prior-year earnings.", + notes=( + "Adjacent-year PERIDNUM join for measured prior-year earnings, " + "with Census allocation flags and sentinels enforced. A joint " + "eight-predictor weighted QRF replaces both earnings leaves on " + "the PUF support half; signed self-employment losses survive." + ), ), US_IMMIGRATION_STAGE_NAME: DonorSpec( survey="CPS ASEC + published unauthorized-population estimates", @@ -750,6 +1843,32 @@ def to_manifest(self) -> dict[str, object]: "exemption channels default to False/0." ), ), + US_RELATIONSHIP_INPUTS_STAGE_NAME: DonorSpec( + survey="Census CPS ASEC", + source="https://www.census.gov/programs-surveys/cps.html", + notes=( + "Measured household-head and marital-status flags mapped exactly " + "from P_SEQ and A_MARITL; nothing is imputed." + ), + ), + US_MEDICARE_TAKE_UP_STAGE_NAME: DonorSpec( + survey="Census CPS ASEC", + source="https://www.census.gov/programs-surveys/cps.html", + notes=( + "Measured Medicare enrollment mapped exactly from MCARE == 1 and " + "copied onto the PUF support clone; no take-up rate is applied." + ), + ), + US_RETIREMENT_DISTRIBUTION_STAGE_NAME: DonorSpec( + survey="Census CPS ASEC", + source="https://www.census.gov/programs-surveys/cps.html", + notes=( + "ASEC rows map exactly from all four DST_SC*/DST_VAL* pairs; the " + "archived CPS-only QRF replaces the four populated non-IRA leaves " + "on PUF support while preserving IRA channel ownership. Nothing " + "is allocated across accounts." + ), + ), US_PREGNANCY_STAGE_NAME: DonorSpec( survey="Census CPS ASEC + CDC natality-derived national pregnancy rate", source="https://www.cdc.gov/nchs/nvss/births.htm", @@ -761,6 +1880,16 @@ def to_manifest(self) -> dict[str, object]: "The ASEC does not measure pregnancy." ), ), + US_WIC_CLAIM_STAGE_NAME: DonorSpec( + survey="USDA FNS WIC Eligibility and Enrollment Estimates + Census CPS ASEC", + source="https://www.fns.usda.gov/research/wic/eligibility-and-coverage-rates-2022", + notes=( + "Stable person-level claim draws use official CY2022 FNS category " + "coverage rates after the pregnancy and parent-input stages. The " + "all-postpartum rate is used because no hermetic source identifies " + "breastfeeding; nutritional risk remains separately excluded." + ), + ), US_SNAP_DISCRETIONARY_EXEMPTION_STAGE_NAME: DonorSpec( survey="Census CPS ASEC + statutory exemption cap (7 U.S.C. 2015(o)(6))", source="https://www.law.cornell.edu/uscode/text/7/2015#o_6", @@ -771,17 +1900,100 @@ def to_manifest(self) -> dict[str, object]: "state usage of the cap (#323)." ), ), + US_CHILDCARE_STAGE_NAME: DonorSpec( + survey="Census CPS ASEC", + source="https://www.census.gov/programs-surveys/cps.html", + notes=( + "Measured replicated SPM_CHILDCAREXPNS is validated and carried " + "to the SPM-unit childcare leaf. After support expansion, an " + "ASEC-trained weighted QRF replaces only the PUF half and the " + "archived first-person reduction places predictions on SPM units." + ), + ), + US_ENERGY_SUBSIDY_STAGE_NAME: DonorSpec( + survey="Census CPS ASEC", + source="https://www.census.gov/programs-surveys/cps.html", + notes=( + "Measured replicated SPM_ENGVAL is validated and carried to the " + "SPM-unit energy-subsidy leaf. After support expansion, an " + "ASEC-trained weighted QRF replaces only the PUF half and the " + "archived first-person reduction places predictions on SPM units." + ), + ), + US_CHILD_SUPPORT_STAGE_NAME: DonorSpec( + survey="Census CPS ASEC", + source="https://www.census.gov/programs-surveys/cps.html", + notes=( + "Measured annual CSP_VAL receipts and positive CHSP_VAL expenses " + "are carried directly. After PUF tax-detail imputation, one " + "ASEC-trained weighted QRF jointly replaces both person leaves " + "on the PUF support half using the archived predictor subset." + ), + ), + US_DISABILITY_BENEFITS_STAGE_NAME: DonorSpec( + survey="Census CPS ASEC", + source="https://www.census.gov/programs-surveys/cps.html", + notes=( + "Measured annual DIS_VAL1/DIS_VAL2 benefits are retained only " + "when their source code is not workers' compensation. After PUF " + "tax-detail imputation, an ASEC-trained weighted QRF replaces the " + "CPS-only person leaf on the PUF support half." + ), + ), + US_WORKERS_COMPENSATION_STAGE_NAME: DonorSpec( + survey="Census CPS ASEC", + source="https://www.census.gov/programs-surveys/cps.html", + notes=( + "Measured annual WC_VAL is carried directly. After PUF tax-detail " + "imputation, an ASEC-trained weighted QRF replaces the CPS-only " + "person leaf on the PUF support half." + ), + ), + US_WEEKS_UNEMPLOYED_STAGE_NAME: DonorSpec( + survey="Census CPS ASEC", + source="https://www.census.gov/programs-surveys/cps.html", + notes=( + "Measured LKWEEKS is carried directly, including an exact " + "identity-keyed repair of the omitted 2022-income-year column " + "from the pinned official 2023 ASEC archive. An ASEC-trained " + "QRF then replaces only the PUF support half, with the archived " + "unemployment-compensation zero rule." + ), + ), "puf_tax_detail": DonorSpec( survey="IRS PUF 2015 (uprated)", source="https://www.irs.gov/statistics/soi-tax-stats-individual-public-use-microdata-files", notes=( - "Itemized-deduction detail, QBI components, partnership SE, " - "mortgage-interest split; IRS disclosure aggregate rows are " + "Itemized-deduction detail, versioned processed-PUF Section 199A " + "simulation leaves (carried without redrawing), partnership SE, " + "mortgage-interest split, direct E00800/E03500 alimony, direct " + "E20500 casualty loss, and the E20400 miscellaneous-expense proxy; IRS " + "disclosure aggregate rows are " "disaggregated from raw PUF totals before uprating, with Forbes " "top-tail synthesis disabled; support clipped to the PUF's own " "realized ranges." ), ), + US_EDUCATION_INPUTS_STAGE_NAME: DonorSpec( + survey="IRS PUF 2015 (uprated) + Census CPS ASEC", + source="https://www.irs.gov/statistics/soi-tax-stats-individual-public-use-microdata-files", + notes=( + "Qualified tuition comes from the PUF E03230/E87530 maximum; " + "the published retired path drops the reported AOTC output and its " + "five affirmative factual inputs therefore follow positive tuition; " + "educational assistance carries directly from ASEC ED_VAL." + ), + ), + US_RETIREMENT_CONTRIBUTION_STAGE_NAME: DonorSpec( + survey="Census CPS ASEC + published retirement-contribution shares", + source="https://www.census.gov/programs-surveys/cps.html", + notes=( + "Measured ASEC RETCB_VAL is allocated across five desired " + "retirement-contribution leaves using archived IRS/BEA/Vanguard/" + "PSCA shares, then CPS-trained QRF predictions replace the PUF " + "support half. PolicyEngine-US owns all statutory caps." + ), + ), "capital_gain_distributions": DonorSpec( survey="IRS SOI Sales of Capital Assets (TY2015) + Pub 1304 Table 1.4", source=( @@ -795,17 +2007,32 @@ def to_manifest(self) -> dict[str, object]: "the two routes are mutually exclusive on a real return." ), ), - "acs_rent": DonorSpec( + US_HOUSING_INPUTS_STAGE_NAME: DonorSpec( survey="Census ACS 2022", source="https://www.census.gov/programs-surveys/acs", - notes="Rent for renter households.", + notes=( + "Exact ASEC H_TENURE and SPM housing carries plus annual ACS PUMS " + "rent imputation on household heads; reported housing assistance " + "alone receives the archived second-stage PUF-support QRF." + ), ), "vehicle_assets": DonorSpec( survey="Census SIPP", source="https://www.census.gov/programs-surveys/sipp.html", notes=( - "Household vehicle count and value; vehicle value folds into " - "household net worth." + "Household vehicle count and value from the pinned full 2023 " + "public-use donor. They remain independent policy inputs until " + "the full mixed-source net-worth reconciliation is restored." + ), + ), + US_VOLUNTARY_FILING_STAGE_NAME: DonorSpec( + survey="Census SIPP", + source="https://www.census.gov/programs-surveys/sipp.html", + notes=( + "Measured 2023 SIPP filing and expected-filing responses replace " + "the retired uncited demographic probability table. Reciprocal " + "spouses form one source unit, reported dependents are excluded, " + "and one weighted-QRF prediction is shared across support clones." ), ), } @@ -816,27 +2043,44 @@ def to_manifest(self) -> dict[str, object]: "asec_load", "unit_assignment", "derive_cps_carried", + US_PRIOR_YEAR_INCOME_STAGE_NAME, US_IMMIGRATION_STAGE_NAME, US_HOURS_WORKED_STAGE_NAME, US_SNAP_TAKE_UP_STAGE_NAME, + US_RELATIONSHIP_INPUTS_STAGE_NAME, + US_MEDICARE_TAKE_UP_STAGE_NAME, + US_HOUSING_INPUTS_STAGE_NAME, + US_RETIREMENT_DISTRIBUTION_STAGE_NAME, US_ELIGIBILITY_INPUTS_STAGE_NAME, US_PREGNANCY_STAGE_NAME, + US_WIC_CLAIM_STAGE_NAME, US_SNAP_DISCRETIONARY_EXEMPTION_STAGE_NAME, + US_RETIREMENT_CONTRIBUTION_STAGE_NAME, + US_CHILDCARE_STAGE_NAME, + US_ENERGY_SUBSIDY_STAGE_NAME, US_PUF_SUPPORT_STAGE_NAME, "puf_tax_detail", + US_CHILD_SUPPORT_STAGE_NAME, + US_DISABILITY_BENEFITS_STAGE_NAME, + US_WORKERS_COMPENSATION_STAGE_NAME, + US_WEEKS_UNEMPLOYED_STAGE_NAME, + US_EDUCATION_INPUTS_STAGE_NAME, "capital_gain_distributions", "scf_wealth", + US_SSI_DISABILITY_CRITERIA_STAGE_NAME, + US_SIPP_HEAD_START_STAGE_NAME, + US_SSI_TAKE_UP_STAGE_NAME, "sipp_tips", "org_wages", "meps_esi_premiums", - "prior_year_income", "mortgage_conversion", - "acs_rent", "vehicle_assets", + US_VOLUNTARY_FILING_STAGE_NAME, "entity_placement", "aca_marketplace_inputs", "medicaid_take_up", US_SNAP_STATE_TAKE_UP_STAGE, + US_OTHER_HEALTH_INSURANCE_STAGE_NAME, "export", ) diff --git a/packages/populace-build/src/populace/build/us_runtime/alimony.py b/packages/populace-build/src/populace/build/us_runtime/alimony.py new file mode 100644 index 00000000..f55c74a2 --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/alimony.py @@ -0,0 +1,308 @@ +"""ASEC/IRS PUF alimony inputs for the US build. + +The retired eCPS pipeline measured recipient income from Census ASEC other- +income records and tax-return income/expense from two IRS PUF fields. The +ASEC half keeps reported ``OI_VAL`` only when ``OI_OFF == 20``; the PUF support +half is populated from the direct ``E00800`` / ``E03500`` mappings through the +shared weighted PUF QRF. ``miscellaneous_income`` must simultaneously exclude +alimony (and the retired pipeline's strike-benefit code 12), otherwise the ASEC +amount is counted twice in gross income. + +This module owns only factual input leaves. PolicyEngine-US owns taxable-income +and above-the-line-deduction formulas. +""" + +from __future__ import annotations + +from importlib.resources import files + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import SourceStageSpec, load_source_manifest +from populace.frame import Frame + +__all__ = [ + "ALIMONY_ASEC_ARCHIVED_DERIVATION_URL", + "ALIMONY_PUF_ARCHIVED_DERIVATION_URL", + "US_ALIMONY_NONCONSTANT_PERSON_COLUMNS", + "US_ALIMONY_OUTPUT_COLUMNS", + "US_ALIMONY_STAGE_NAME", + "derive_us_alimony_from_asec", + "derive_us_alimony_from_puf", + "us_alimony_signal_gate", + "us_alimony_stage_spec", + "us_alimony_summary", +] + +_ARCHIVED_DATA_REPOSITORY = "policyengine-" + "us-data" +_ARCHIVED_COMMIT = "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe" +ALIMONY_ASEC_ARCHIVED_DERIVATION_URL = ( + "https://github.com/PolicyEngine/" + f"{_ARCHIVED_DATA_REPOSITORY}/blob/{_ARCHIVED_COMMIT}/" + "policyengine_" + "us_data/datasets/cps/cps.py#L1481-L1492" +) +ALIMONY_PUF_ARCHIVED_DERIVATION_URL = ( + "https://github.com/PolicyEngine/" + f"{_ARCHIVED_DATA_REPOSITORY}/blob/{_ARCHIVED_COMMIT}/" + "policyengine_" + "us_data/datasets/puf/puf.py#L640-L641" +) + +US_ALIMONY_STAGE_NAME = "puf_tax_detail" +US_ALIMONY_OUTPUT_COLUMNS: tuple[str, ...] = ( + "alimony_income", + "alimony_expense", +) +US_ALIMONY_NONCONSTANT_PERSON_COLUMNS = US_ALIMONY_OUTPUT_COLUMNS + +_ASEC_ALIMONY_OTHER_INCOME_CODE = 20 +_ASEC_SEPARATE_OTHER_INCOME_CODES = frozenset({12, 20}) +_NONZERO_SHARE_BAND = (0.00001, 0.02) + + +def us_alimony_stage_spec() -> SourceStageSpec: + """Load and validate the shared PUF tax-detail stage declaration.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + spec = manifest.stage_map()[US_ALIMONY_STAGE_NAME] + missing = sorted(set(US_ALIMONY_OUTPUT_COLUMNS) - set(spec.outputs)) + if missing: + raise ValueError( + f"{US_ALIMONY_STAGE_NAME!r} manifest stage does not declare " + f"alimony output(s) {missing}." + ) + return spec + + +def _strict_numeric_source( + table: pd.DataFrame, + column: str, + *, + label: str, + nonnegative: bool, +) -> np.ndarray: + if column not in table.columns: + raise ValueError(f"{label} requires source column {column!r}.") + values = pd.to_numeric(table[column], errors="coerce").to_numpy(dtype=np.float64) + nonfinite = ~np.isfinite(values) + if bool(nonfinite.any()): + raise ValueError( + f"{label} source column {column!r} contains " + f"{int(np.count_nonzero(nonfinite))} nonnumeric or nonfinite value(s)." + ) + if nonnegative: + negative = values < 0.0 + if bool(negative.any()): + raise ValueError( + f"{label} source column {column!r} contains " + f"{int(np.count_nonzero(negative))} negative value(s)." + ) + return values + + +def derive_us_alimony_from_asec(person: pd.DataFrame) -> pd.DataFrame: + """Carry reported ASEC alimony and remove it from miscellaneous income. + + An already materialized pair is preserved, which makes the CPS-carried + transform idempotent on a staged base artifact. If either leaf is missing, + both raw ASEC fields are mandatory; a missing source must not be healed with + fabricated zeros. + """ + + raw_sources = {"OI_VAL", "OI_OFF"} + present_sources = raw_sources.intersection(person.columns) + if not present_sources: + materialized = {"alimony_income", "miscellaneous_income"} + if materialized.issubset(person.columns): + return person.copy(deep=True) + missing = sorted(materialized - set(person.columns)) + raise ValueError( + "ASEC alimony derivation cannot reuse an unstaged table without " + f"raw OI_VAL/OI_OFF; missing materialized column(s) {missing}." + ) + if present_sources != raw_sources: + missing = sorted(raw_sources - set(person.columns)) + raise ValueError( + "ASEC alimony derivation requires both raw source columns; " + f"missing {missing}." + ) + + label = "ASEC alimony derivation" + amounts = _strict_numeric_source( + person, + "OI_VAL", + label=label, + nonnegative=True, + ) + codes = _strict_numeric_source( + person, + "OI_OFF", + label=label, + nonnegative=True, + ) + noninteger = codes != np.floor(codes) + if bool(noninteger.any()): + raise ValueError( + "ASEC alimony derivation source column 'OI_OFF' contains " + f"{int(np.count_nonzero(noninteger))} noninteger code(s)." + ) + integer_codes = codes.astype(np.int64) + + result = person.copy(deep=True) + result["alimony_income"] = np.where( + integer_codes == _ASEC_ALIMONY_OTHER_INCOME_CODE, + amounts, + 0.0, + ) + result["miscellaneous_income"] = np.where( + np.isin(integer_codes, list(_ASEC_SEPARATE_OTHER_INCOME_CODES)), + 0.0, + amounts, + ) + return result + + +def derive_us_alimony_from_puf( + puf: pd.DataFrame, + *, + income_source_column: str = "E00800", + income_output_column: str = "alimony_income", + expense_source_column: str = "E03500", + expense_output_column: str = "alimony_expense", +) -> pd.DataFrame: + """Carry the archived PUF alimony income and expense fields exactly.""" + + income = _strict_numeric_source( + puf, + income_source_column, + label="PUF alimony derivation", + nonnegative=True, + ) + expense = _strict_numeric_source( + puf, + expense_source_column, + label="PUF alimony derivation", + nonnegative=True, + ) + result = puf.copy(deep=True) + result[income_output_column] = income + result[expense_output_column] = expense + return result + + +def us_alimony_summary(frame: Frame) -> dict[str, object]: + """Return weighted signal and validity diagnostics for both leaves.""" + + person = frame.table("person") + weights = np.asarray(frame.resolve_weights("person").values, dtype=np.float64) + total_weight = float(weights.sum()) + columns: dict[str, object] = {} + for column in US_ALIMONY_OUTPUT_COLUMNS: + values = pd.to_numeric(person[column], errors="coerce").to_numpy( + dtype=np.float64 + ) + finite = np.isfinite(values) + positive = finite & (values > 0.0) + positive_share = ( + float(weights[positive].sum()) / total_weight if total_weight > 0.0 else 0.0 + ) + columns[column] = { + "positive_share": positive_share, + "positive_share_band": list(_NONZERO_SHARE_BAND), + "unweighted_total": float(np.nansum(values)), + "nonfinite": int(np.count_nonzero(~finite)), + "negative": int(np.count_nonzero(finite & (values < 0.0))), + } + return {"columns": columns} + + +def us_alimony_signal_gate(frame: Frame) -> GateResult: + """Require finite, nonnegative, plausibly sparse signal in both leaves.""" + + person = frame.table("person") + missing = [column for column in US_ALIMONY_OUTPUT_COLUMNS if column not in person] + if missing: + return GateResult( + name="alimony_inputs_signal", + passed=False, + failures=(f"person columns missing: {missing}.",), + details={"missing": missing}, + ) + + summary = us_alimony_summary(frame) + failures: list[str] = [] + column_summaries = summary["columns"] + for column in US_ALIMONY_OUTPUT_COLUMNS: + details = column_summaries[column] + if details["nonfinite"]: + failures.append( + f"{column}: {int(details['nonfinite'])} nonfinite value(s)." + ) + if details["negative"]: + failures.append(f"{column}: {int(details['negative'])} negative value(s).") + share = float(details["positive_share"]) + low, high = details["positive_share_band"] + if not (low <= share <= high): + failures.append( + f"{column}: positive share {share:.6f} outside plausibility " + f"band [{low}, {high}]." + ) + raw_sources = {"OI_VAL", "OI_OFF"} + if raw_sources.issubset(person.columns) and "miscellaneous_income" not in person: + failures.append( + "ASEC raw OI_VAL/OI_OFF are present but miscellaneous_income is missing." + ) + elif raw_sources.issubset(person.columns): + source_mask = np.ones(len(person), dtype=bool) + if "person_support_channel" in person: + source_mask = ( + person["person_support_channel"].to_numpy(dtype=object) == "asec" + ) + amounts = pd.to_numeric(person["OI_VAL"], errors="coerce").to_numpy( + dtype=np.float64 + )[source_mask] + codes = pd.to_numeric(person["OI_OFF"], errors="coerce").to_numpy( + dtype=np.float64 + )[source_mask] + if bool(np.isfinite(amounts).all() and np.isfinite(codes).all()): + integer_codes = codes.astype(np.int64) + expected_alimony = np.where( + integer_codes == _ASEC_ALIMONY_OTHER_INCOME_CODE, + amounts, + 0.0, + ) + expected_miscellaneous = np.where( + np.isin(integer_codes, list(_ASEC_SEPARATE_OTHER_INCOME_CODES)), + 0.0, + amounts, + ) + actual_alimony = pd.to_numeric( + person.loc[source_mask, "alimony_income"], errors="coerce" + ).to_numpy(dtype=np.float64) + actual_miscellaneous = pd.to_numeric( + person.loc[source_mask, "miscellaneous_income"], errors="coerce" + ).to_numpy(dtype=np.float64) + alimony_mismatch = ~np.isclose(actual_alimony, expected_alimony) + miscellaneous_mismatch = ~np.isclose( + actual_miscellaneous, + expected_miscellaneous, + ) + if bool(alimony_mismatch.any()): + failures.append( + "ASEC alimony_income disagrees with OI_OFF == 20 / OI_VAL " + f"on {int(np.count_nonzero(alimony_mismatch))} row(s)." + ) + if bool(miscellaneous_mismatch.any()): + failures.append( + "ASEC miscellaneous_income does not exclude alimony/strike " + f"codes on {int(np.count_nonzero(miscellaneous_mismatch))} row(s)." + ) + return GateResult( + name="alimony_inputs_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) diff --git a/packages/populace-build/src/populace/build/us_runtime/asec_pool.py b/packages/populace-build/src/populace/build/us_runtime/asec_pool.py index 1a8c5442..2979c6f7 100644 --- a/packages/populace-build/src/populace/build/us_runtime/asec_pool.py +++ b/packages/populace-build/src/populace/build/us_runtime/asec_pool.py @@ -177,7 +177,10 @@ def build_pooled_asec_unit_frame( tax_unit_mode=tax_unit_mode, strata=strata, ) - frame = _attach_household_attributes(frame, ("state_fips",)) + # Keep raw household tenure through unit construction. The archived eCPS + # housing stage maps H_TENURE directly; SPM_TENMORTSTATUS is a distinct + # concept and cannot reconstruct rent-free occupancy (H_TENURE == 3). + frame = _attach_household_attributes(frame, ("state_fips", "H_TENURE")) return frame, pooled.metadata @@ -292,6 +295,22 @@ def _remap_source_year( .astype("int64") .to_numpy() ) + if "H_TENURE" in household: + tenure_by_household = ( + household[["H_SEQ", "H_TENURE"]] + .drop_duplicates("H_SEQ") + .set_index("H_SEQ")["H_TENURE"] + ) + result["H_TENURE"] = result["source_household_id"].map(tenure_by_household) + if result["H_TENURE"].isna().any(): + missing = ( + result.loc[result["H_TENURE"].isna(), "source_household_id"] + .head() + .tolist() + ) + raise ValueError( + f"ASEC H_TENURE is missing for source household id(s): {missing}." + ) result["PH_SEQ"] = result["PH_SEQ"].map(household_map).astype("int64") result["household_id"] = result["PH_SEQ"] result[ASEC_PERSON_WEIGHT_COLUMN] = ( diff --git a/packages/populace-build/src/populace/build/us_runtime/capital_gain_details.py b/packages/populace-build/src/populace/build/us_runtime/capital_gain_details.py new file mode 100644 index 00000000..b6ff57ec --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/capital_gain_details.py @@ -0,0 +1,270 @@ +"""IRS PUF collectibles and unrecaptured-section-1250 gain inputs. + +The retired eCPS pipeline mapped ``E24518`` directly to +``long_term_capital_gains_on_collectibles`` and ``E24515`` directly to +``unrecaptured_section_1250_gain``. The processed PUF carries the first leaf at +person grain and the second at tax-unit grain. Populace reduces the person leaf +to its identifiable tax-unit total for the shared PUF fit, then uses +first-person placement on the PUF support channel because PolicyEngine-US sums +it back to the tax unit. This avoids inventing the retired randomized +filer/spouse earnings split. + +Both leaves are primary-source amounts. PolicyEngine-US owns their federal and +Massachusetts tax formulas; this module only carries and validates the inputs. +""" + +from __future__ import annotations + +from importlib.resources import files + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import SourceStageSpec, load_source_manifest +from populace.frame import Frame + +__all__ = [ + "CAPITAL_GAIN_DETAILS_ARCHIVED_DERIVATION_URL", + "CAPITAL_GAIN_DETAILS_ARCHIVED_EXPORT_URL", + "CAPITAL_GAIN_DETAILS_ARCHIVED_IMPUTATION_URL", + "CAPITAL_GAIN_DETAILS_ARCHIVED_PERSON_ALLOCATION_URL", + "CAPITAL_GAIN_DETAILS_ARCHIVED_PUF_ARTIFACT_URL", + "US_CAPITAL_GAIN_DETAILS_NONCONSTANT_PERSON_COLUMNS", + "US_CAPITAL_GAIN_DETAILS_NONCONSTANT_TAX_UNIT_COLUMNS", + "US_CAPITAL_GAIN_DETAILS_OUTPUT_COLUMNS", + "US_CAPITAL_GAIN_DETAILS_STAGE_NAME", + "derive_us_capital_gain_details_from_puf", + "us_capital_gain_details_signal_gate", + "us_capital_gain_details_stage_spec", + "us_capital_gain_details_summary", +] + +_ARCHIVED_DATA_REPOSITORY = "policyengine-" + "us-data" +_ARCHIVED_ROOT = ( + "https://github.com/PolicyEngine/" + f"{_ARCHIVED_DATA_REPOSITORY}/blob/" + "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe/" + "policyengine_" + "us_data/" +) +CAPITAL_GAIN_DETAILS_ARCHIVED_DERIVATION_URL = ( + _ARCHIVED_ROOT + "datasets/puf/puf.py#L636-L702" +) +CAPITAL_GAIN_DETAILS_ARCHIVED_EXPORT_URL = ( + _ARCHIVED_ROOT + "datasets/puf/puf.py#L804-L850" +) +CAPITAL_GAIN_DETAILS_ARCHIVED_IMPUTATION_URL = ( + _ARCHIVED_ROOT + "calibration/puf_impute.py#L90-L198" +) +CAPITAL_GAIN_DETAILS_ARCHIVED_PERSON_ALLOCATION_URL = ( + _ARCHIVED_ROOT + "datasets/puf/puf.py#L1513-L1601" +) +CAPITAL_GAIN_DETAILS_ARCHIVED_PUF_ARTIFACT_URL = ( + _ARCHIVED_ROOT + "datasets/puf/puf.py#L1655-L1660" +) + +US_CAPITAL_GAIN_DETAILS_STAGE_NAME = "puf_tax_detail" +US_CAPITAL_GAIN_DETAILS_NONCONSTANT_PERSON_COLUMNS: tuple[str, ...] = ( + "long_term_capital_gains_on_collectibles", +) +US_CAPITAL_GAIN_DETAILS_NONCONSTANT_TAX_UNIT_COLUMNS: tuple[str, ...] = ( + "unrecaptured_section_1250_gain", +) +US_CAPITAL_GAIN_DETAILS_OUTPUT_COLUMNS = ( + *US_CAPITAL_GAIN_DETAILS_NONCONSTANT_PERSON_COLUMNS, + *US_CAPITAL_GAIN_DETAILS_NONCONSTANT_TAX_UNIT_COLUMNS, +) + +_BASE_ASEC_SUPPORT_CHANNEL = "asec" +_PUF_TAX_DETAIL_SUPPORT_CHANNEL = "puf_tax_detail" +_SIGNAL_BANDS = { + "long_term_capital_gains_on_collectibles": { + "overall": (0.00001, 0.005), + "puf": (0.00002, 0.01), + }, + "unrecaptured_section_1250_gain": { + "overall": (0.0005, 0.03), + "puf": (0.001, 0.06), + }, +} + + +def us_capital_gain_details_stage_spec() -> SourceStageSpec: + """Load and validate the shared PUF tax-detail stage declaration.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + spec = manifest.stage_map()[US_CAPITAL_GAIN_DETAILS_STAGE_NAME] + missing = sorted(set(US_CAPITAL_GAIN_DETAILS_OUTPUT_COLUMNS) - set(spec.outputs)) + if missing: + raise ValueError( + f"{US_CAPITAL_GAIN_DETAILS_STAGE_NAME!r} manifest stage does not " + f"declare capital-gain detail output(s) {missing}." + ) + return spec + + +def derive_us_capital_gain_details_from_puf( + puf: pd.DataFrame, + *, + collectibles_source_column: str = "E24518", + collectibles_output_column: str = "long_term_capital_gains_on_collectibles", + unrecaptured_source_column: str = "E24515", + unrecaptured_output_column: str = "unrecaptured_section_1250_gain", +) -> pd.DataFrame: + """Carry the archived E24518 and E24515 fields without alteration.""" + + sources = { + collectibles_output_column: collectibles_source_column, + unrecaptured_output_column: unrecaptured_source_column, + } + missing = sorted(set(sources.values()) - set(puf.columns)) + if missing: + raise ValueError( + f"PUF capital-gain detail derivation requires source columns {missing}." + ) + + result = puf.copy(deep=True) + for output, source in sources.items(): + values = pd.to_numeric(puf[source], errors="coerce").to_numpy(dtype=np.float64) + nonfinite = ~np.isfinite(values) + if bool(nonfinite.any()): + raise ValueError( + f"PUF capital-gain source column {source!r} contains " + f"{int(np.count_nonzero(nonfinite))} nonnumeric or nonfinite " + "value(s)." + ) + negative = values < 0.0 + if bool(negative.any()): + raise ValueError( + f"PUF capital-gain source column {source!r} contains " + f"{int(np.count_nonzero(negative))} negative value(s)." + ) + result[output] = values + return result + + +def _column_summary( + frame: Frame, + *, + entity: str, + output: str, +) -> dict[str, object]: + table = frame.table(entity) + values = pd.to_numeric(table[output], errors="coerce").to_numpy(dtype=np.float64) + weights = np.asarray(frame.resolve_weights(entity).values, dtype=np.float64) + finite = np.isfinite(values) + positive = finite & (values > 0.0) + total_weight = float(weights.sum()) + summary: dict[str, object] = { + "positive_share": ( + float(weights[positive].sum()) / total_weight if total_weight > 0.0 else 0.0 + ), + "positive_share_band": list(_SIGNAL_BANDS[output]["overall"]), + "weighted_total": float((np.nan_to_num(values) * weights).sum()), + "nonfinite": int(np.count_nonzero(~finite)), + "negative": int(np.count_nonzero(finite & (values < 0.0))), + } + support_channel = f"{entity}_support_channel" + if support_channel in table.columns: + channel_values = table[support_channel].to_numpy() + channels: dict[str, dict[str, float | int]] = {} + for channel in table[support_channel].dropna().unique(): + channel_mask = channel_values == channel + channel_weight = float(weights[channel_mask].sum()) + channel_positive = channel_mask & positive + channels[str(channel)] = { + "positive_rows": int(np.count_nonzero(channel_positive)), + "positive_share": ( + float(weights[channel_positive].sum()) / channel_weight + if channel_weight > 0.0 + else 0.0 + ), + "weighted_total": float( + (np.nan_to_num(values[channel_mask]) * weights[channel_mask]).sum() + ), + } + summary["channels"] = channels + return summary + + +def us_capital_gain_details_summary(frame: Frame) -> dict[str, dict[str, object]]: + """Return weighted signal, support-channel, and validity diagnostics.""" + + return { + "long_term_capital_gains_on_collectibles": _column_summary( + frame, + entity="person", + output="long_term_capital_gains_on_collectibles", + ), + "unrecaptured_section_1250_gain": _column_summary( + frame, + entity="tax_unit", + output="unrecaptured_section_1250_gain", + ), + } + + +def us_capital_gain_details_signal_gate(frame: Frame) -> GateResult: + """Require finite, nonnegative, source-aligned signal in both PUF leaves.""" + + entity_outputs = { + "person": US_CAPITAL_GAIN_DETAILS_NONCONSTANT_PERSON_COLUMNS, + "tax_unit": US_CAPITAL_GAIN_DETAILS_NONCONSTANT_TAX_UNIT_COLUMNS, + } + missing = { + entity: sorted(set(outputs) - set(frame.table(entity).columns)) + for entity, outputs in entity_outputs.items() + } + missing = {entity: outputs for entity, outputs in missing.items() if outputs} + if missing: + return GateResult( + name="capital_gain_details_signal", + passed=False, + failures=(f"entity columns missing: {missing}.",), + details={"missing": missing}, + ) + + summary = us_capital_gain_details_summary(frame) + failures: list[str] = [] + for output, details in summary.items(): + if details["nonfinite"]: + failures.append(f"{output} nonfinite values: {int(details['nonfinite'])}.") + if details["negative"]: + failures.append(f"{output} negative values: {int(details['negative'])}.") + share = float(details["positive_share"]) + low, high = details["positive_share_band"] + if not (low <= share <= high): + failures.append( + f"{output} positive share {share:.6f} outside plausibility " + f"band [{low}, {high}]." + ) + channels = details.get("channels") + if not isinstance(channels, dict): + continue + asec = channels.get(_BASE_ASEC_SUPPORT_CHANNEL) + puf = channels.get(_PUF_TAX_DETAIL_SUPPORT_CHANNEL) + if asec is None: + failures.append(f"{output} is missing the ASEC support channel.") + elif float(asec["weighted_total"]) != 0.0: + failures.append( + f"{output} must remain zero on the source-unobserved ASEC channel." + ) + if puf is None: + failures.append(f"{output} is missing the PUF tax-detail channel.") + else: + puf_share = float(puf["positive_share"]) + puf_low, puf_high = _SIGNAL_BANDS[output]["puf"] + if not (puf_low <= puf_share <= puf_high): + failures.append( + f"PUF {output} positive share {puf_share:.6f} outside " + f"plausibility band [{puf_low}, {puf_high}]." + ) + + return GateResult( + name="capital_gain_details_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) diff --git a/packages/populace-build/src/populace/build/us_runtime/casualty_losses.py b/packages/populace-build/src/populace/build/us_runtime/casualty_losses.py new file mode 100644 index 00000000..586a26c5 --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/casualty_losses.py @@ -0,0 +1,166 @@ +"""IRS PUF casualty-loss inputs for the US build. + +The retired eCPS pipeline carried ``casualty_loss`` directly from IRS PUF +field ``E20500``. The immutable archived source coordinate is exposed as +``CASUALTY_LOSS_ARCHIVED_DERIVATION_URL`` below. + +There is no CPS analogue and no synthetic fallback. The PUF tax-detail stage +therefore derives the source value exactly, then its existing weighted QRF +places that source-backed detail on the dedicated PUF support channel. Because +casualty losses are rare, the shared PUF runtime sparsifies the predictions to +the donor's weighted positive rate while preserving their weighted total. + +PolicyEngine-US owns the casualty-loss deduction formula and its AGI floor. +This module persists only the factual person input, never the computed +deduction. +""" + +from __future__ import annotations + +from importlib.resources import files + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import SourceStageSpec, load_source_manifest +from populace.frame import Frame + +__all__ = [ + "CASUALTY_LOSS_ARCHIVED_DERIVATION_URL", + "US_CASUALTY_LOSS_NONCONSTANT_PERSON_COLUMNS", + "US_CASUALTY_LOSS_OUTPUT_COLUMNS", + "US_CASUALTY_LOSS_STAGE_NAME", + "derive_us_casualty_loss_from_puf", + "us_casualty_loss_signal_gate", + "us_casualty_loss_stage_spec", + "us_casualty_loss_summary", +] + +_ARCHIVED_DATA_REPOSITORY = "policyengine-" + "us-data" +CASUALTY_LOSS_ARCHIVED_DERIVATION_URL = ( + "https://github.com/PolicyEngine/" + f"{_ARCHIVED_DATA_REPOSITORY}/blob/" + "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe/" + "policyengine_" + "us_data/datasets/puf/puf.py#L636-L643" +) + +US_CASUALTY_LOSS_STAGE_NAME = "puf_tax_detail" +US_CASUALTY_LOSS_OUTPUT_COLUMNS: tuple[str, ...] = ("casualty_loss",) +US_CASUALTY_LOSS_NONCONSTANT_PERSON_COLUMNS = US_CASUALTY_LOSS_OUTPUT_COLUMNS + +_CASUALTY_LOSS_NONZERO_SHARE_BAND = (0.0001, 0.02) + + +def us_casualty_loss_stage_spec() -> SourceStageSpec: + """Load and validate the shared PUF tax-detail stage declaration.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + spec = manifest.stage_map()[US_CASUALTY_LOSS_STAGE_NAME] + missing = sorted(set(US_CASUALTY_LOSS_OUTPUT_COLUMNS) - set(spec.outputs)) + if missing: + raise ValueError( + f"{US_CASUALTY_LOSS_STAGE_NAME!r} manifest stage does not declare " + f"casualty-loss output(s) {missing}." + ) + return spec + + +def derive_us_casualty_loss_from_puf( + puf: pd.DataFrame, + *, + source_column: str = "E20500", + output_column: str = "casualty_loss", +) -> pd.DataFrame: + """Carry the archived IRS PUF casualty-loss field without alteration. + + Invalid or negative source values fail closed. Replacing them with zeros + would invent observations and could silently erase the reform input this + stage exists to restore. + """ + + if source_column not in puf.columns: + raise ValueError( + f"PUF casualty-loss derivation requires source column {source_column!r}." + ) + values = pd.to_numeric(puf[source_column], errors="coerce") + numeric = values.to_numpy(dtype=np.float64) + nonfinite = ~np.isfinite(numeric) + if bool(nonfinite.any()): + raise ValueError( + f"PUF casualty-loss source column {source_column!r} contains " + f"{int(np.count_nonzero(nonfinite))} nonnumeric or nonfinite value(s)." + ) + negative = numeric < 0.0 + if bool(negative.any()): + raise ValueError( + f"PUF casualty-loss source column {source_column!r} contains " + f"{int(np.count_nonzero(negative))} negative value(s)." + ) + + result = puf.copy(deep=True) + result[output_column] = numeric + return result + + +def us_casualty_loss_summary(frame: Frame) -> dict[str, object]: + """Return weighted signal and validity diagnostics for casualty loss.""" + + person = frame.table("person") + values = pd.to_numeric(person["casualty_loss"], errors="coerce").to_numpy( + dtype=np.float64 + ) + weights = np.asarray(frame.resolve_weights("person").values, dtype=np.float64) + finite = np.isfinite(values) + positive = finite & (values > 0.0) + total_weight = float(weights.sum()) + positive_share = ( + float(weights[positive].sum()) / total_weight if total_weight > 0.0 else 0.0 + ) + return { + "positive_share": positive_share, + "positive_share_band": list(_CASUALTY_LOSS_NONZERO_SHARE_BAND), + "unweighted_total": float(np.nansum(values)), + "nonfinite": int(np.count_nonzero(~finite)), + "negative": int(np.count_nonzero(finite & (values < 0.0))), + } + + +def us_casualty_loss_signal_gate(frame: Frame) -> GateResult: + """Require a finite, nonnegative, plausibly sparse casualty-loss input.""" + + person = frame.table("person") + missing = [ + column + for column in US_CASUALTY_LOSS_OUTPUT_COLUMNS + if column not in person.columns + ] + if missing: + return GateResult( + name="casualty_loss_signal", + passed=False, + failures=(f"person columns missing: {missing}.",), + details={"missing": missing}, + ) + + summary = us_casualty_loss_summary(frame) + failures: list[str] = [] + if summary["nonfinite"]: + failures.append(f"casualty_loss nonfinite values: {int(summary['nonfinite'])}.") + if summary["negative"]: + failures.append(f"casualty_loss negative values: {int(summary['negative'])}.") + share = float(summary["positive_share"]) + low, high = summary["positive_share_band"] + if not (low <= share <= high): + failures.append( + f"casualty-loss positive share {share:.6f} outside plausibility " + f"band [{low}, {high}]." + ) + return GateResult( + name="casualty_loss_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) diff --git a/packages/populace-build/src/populace/build/us_runtime/child_support.py b/packages/populace-build/src/populace/build/us_runtime/child_support.py new file mode 100644 index 00000000..b55be74d --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/child_support.py @@ -0,0 +1,653 @@ +"""CPS ASEC child-support inputs with retired PUF-clone QRF treatment. + +The retired eCPS pipeline carried annual person-level child support received +directly from ASEC ``CSP_VAL`` and annual child support paid from ``CHSP_VAL``. +After its PUF support clone received PUF-imputed income, one second-stage QRF +jointly replaced both child-support leaves on the clone half using eight +documented predictors and at most 5,000 ASEC training people. Immutable +archived source coordinates are exposed below. + +This port preserves the annual positive-dollar semantics, observed top-coded +values, and person grain. It does not balance paid against received, negate +expenses, impose mutual exclusivity, or invent a child/household eligibility +mask. The QRF fit uses typed person weights, deliberately strengthening the +archived unweighted fit under Populace's build-wide weighting contract. +""" + +from __future__ import annotations + +from importlib.resources import files +from typing import Any + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import ( + SourceOperationSpec, + SourceStageSpec, + load_source_manifest, +) +from populace.build.source_runtime import ( + SourceRuntimeConfig, + SourceRuntimeContext, + SourceRuntimeError, + run_source_stage, +) +from populace.frame import Frame +from populace.frame.units import US_SCHEMA + +__all__ = [ + "CHILD_SUPPORT_ARCHIVED_PUF_IMPUTATION_URL", + "CHILD_SUPPORT_ARCHIVED_PUF_OUTPUTS_URL", + "CHILD_SUPPORT_EXPENSE_ARCHIVED_DERIVATION_URL", + "CHILD_SUPPORT_RECEIVED_ARCHIVED_DERIVATION_URL", + "US_CHILD_SUPPORT_NONCONSTANT_PERSON_COLUMNS", + "US_CHILD_SUPPORT_OUTPUT_COLUMNS", + "US_CHILD_SUPPORT_REQUIRED_SOURCE_COLUMNS", + "US_CHILD_SUPPORT_STAGE_NAME", + "derive_us_child_support_from_asec", + "derive_us_child_support_from_manifest", + "impute_us_child_support_to_puf_support_from_manifest", + "us_child_support_signal_gate", + "us_child_support_stage_spec", + "us_child_support_summary", + "with_us_child_support_inputs", +] + +QRF: Any | None = None + +_ARCHIVED_DATA_REPOSITORY = "policyengine-" + "us-data" +_ARCHIVED_ROOT = ( + "https://github.com/PolicyEngine/" + f"{_ARCHIVED_DATA_REPOSITORY}/blob/" + "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe/" + "policyengine_" + "us_data/" +) +CHILD_SUPPORT_RECEIVED_ARCHIVED_DERIVATION_URL = ( + _ARCHIVED_ROOT + "datasets/cps/cps.py#L1493-L1496" +) +CHILD_SUPPORT_EXPENSE_ARCHIVED_DERIVATION_URL = ( + _ARCHIVED_ROOT + "datasets/cps/cps.py#L1572-L1574" +) +CHILD_SUPPORT_ARCHIVED_PUF_OUTPUTS_URL = ( + _ARCHIVED_ROOT + "datasets/cps/extended_cps.py#L135-L194" +) +CHILD_SUPPORT_ARCHIVED_PUF_IMPUTATION_URL = ( + _ARCHIVED_ROOT + "datasets/cps/extended_cps.py#L639-L745" +) + +US_CHILD_SUPPORT_STAGE_NAME = "child_support_inputs" +US_CHILD_SUPPORT_OUTPUT_COLUMNS: tuple[str, ...] = ( + "child_support_received", + "child_support_expense", +) +US_CHILD_SUPPORT_NONCONSTANT_PERSON_COLUMNS = US_CHILD_SUPPORT_OUTPUT_COLUMNS +US_CHILD_SUPPORT_REQUIRED_SOURCE_COLUMNS: tuple[str, ...] = ( + "CSP_VAL", + "CHSP_VAL", +) + +_PERSON_WEIGHT_COLUMN = "person_weight" +_PERSON_SUPPORT_CHANNEL_COLUMN = "person_support_channel" +_BASE_ASEC_SUPPORT_CHANNEL = "asec" +_PUF_TAX_DETAIL_SUPPORT_CHANNEL = "puf_tax_detail" +_PREDICTORS: tuple[str, ...] = ( + "age", + "is_male", + "has_esi", + "tax_unit_is_joint", + "tax_unit_count_dependents", + "employment_income", + "self_employment_income", + "social_security", +) +_PREDICTOR_PREFIX = "child_support_predictor_" +_EXPECTED_DIRECT_PARAMETERS = { + "received_source": "CSP_VAL", + "received_output": "child_support_received", + "expense_source": "CHSP_VAL", + "expense_output": "child_support_expense", +} +_PUF_IMPUTATION_PARAMETER_KEYS = frozenset( + { + "predictors", + "max_train_samples", + "n_estimators", + "seed_from_build_config", + "weight", + } +) +_NONZERO_SHARE_BANDS: dict[str, tuple[float, float]] = { + "child_support_received": (0.005, 0.15), + "child_support_expense": (0.005, 0.15), +} +_CHANNEL_NONZERO_SHARE_BAND = (0.002, 0.15) + + +def us_child_support_stage_spec() -> SourceStageSpec: + """Load and validate the packaged child-support stage declaration.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + stage_map = manifest.stage_map() + if US_CHILD_SUPPORT_STAGE_NAME not in stage_map: + raise ValueError( + f"US source manifest declares no {US_CHILD_SUPPORT_STAGE_NAME!r} stage." + ) + spec = stage_map[US_CHILD_SUPPORT_STAGE_NAME] + missing = sorted(set(US_CHILD_SUPPORT_OUTPUT_COLUMNS) - set(spec.outputs)) + if missing: + raise ValueError( + f"{US_CHILD_SUPPORT_STAGE_NAME!r} manifest stage does not declare " + f"output(s) {missing}; the runtime and manifest have drifted." + ) + return spec + + +def _strict_nonnegative_source(person: pd.DataFrame, column: str) -> np.ndarray: + if column not in person.columns: + raise SourceRuntimeError( + f"US child-support derivation requires ASEC source column {column!r}." + ) + values = pd.to_numeric(person[column], errors="coerce").to_numpy(dtype=np.float64) + nonfinite = int(np.count_nonzero(~np.isfinite(values))) + if nonfinite: + raise SourceRuntimeError( + f"US child-support source {column!r} contains {nonfinite} nonnumeric " + "or nonfinite value(s)." + ) + negative = int(np.count_nonzero(values < 0.0)) + if negative: + raise SourceRuntimeError( + f"US child-support source {column!r} contains {negative} negative value(s)." + ) + return values + + +def derive_us_child_support_from_asec( + person: pd.DataFrame, + *, + received_source_column: str = "CSP_VAL", + received_output_column: str = "child_support_received", + expense_source_column: str = "CHSP_VAL", + expense_output_column: str = "child_support_expense", +) -> pd.DataFrame: + """Carry both measured annual ASEC child-support fields exactly.""" + + received = _strict_nonnegative_source(person, received_source_column) + expense = _strict_nonnegative_source(person, expense_source_column) + result = person.copy(deep=True) + result[received_output_column] = received + result[expense_output_column] = expense + return result + + +def derive_us_child_support_from_manifest( + frame: pd.DataFrame | None, + operation: SourceOperationSpec, + _context: SourceRuntimeContext | None, +) -> pd.DataFrame: + """Interpret the manifest's exact ASEC direct-carry operation.""" + + if operation.kind != "derive_child_support_inputs": + raise SourceRuntimeError( + "US child-support derivation received unexpected operation " + f"{operation.kind!r}." + ) + if frame is None: + raise SourceRuntimeError( + "US child-support derivation requires the person table first." + ) + parameters = dict(operation.parameters) + if parameters != _EXPECTED_DIRECT_PARAMETERS: + raise SourceRuntimeError( + "US child-support direct mapping drifted from the archived method: " + f"expected {_EXPECTED_DIRECT_PARAMETERS}, got {parameters}." + ) + return derive_us_child_support_from_asec( + frame, + received_source_column=parameters["received_source"], + received_output_column=parameters["received_output"], + expense_source_column=parameters["expense_source"], + expense_output_column=parameters["expense_output"], + ) + + +def impute_us_child_support_to_puf_support_from_manifest( + frame: pd.DataFrame | None, + operation: SourceOperationSpec, + context: SourceRuntimeContext | None, +) -> pd.DataFrame: + """Jointly QRF-impute the two leaves onto PUF-clone people.""" + + if operation.kind != "impute_child_support_to_puf_support": + raise SourceRuntimeError( + "US child-support PUF imputation received unexpected operation " + f"{operation.kind!r}." + ) + if frame is None: + raise SourceRuntimeError( + "US child-support PUF imputation requires the person table first." + ) + unexpected = sorted(set(operation.parameters) - _PUF_IMPUTATION_PARAMETER_KEYS) + missing_parameters = sorted( + _PUF_IMPUTATION_PARAMETER_KEYS - set(operation.parameters) + ) + if unexpected or missing_parameters: + raise SourceRuntimeError( + "US child-support PUF imputation parameters must match the archived " + f"method; missing={missing_parameters}, unexpected={unexpected}." + ) + if _PERSON_SUPPORT_CHANNEL_COLUMN not in frame.columns: + return frame.copy(deep=True) + + predictors = tuple(str(value) for value in operation.parameters["predictors"]) + if predictors != _PREDICTORS: + raise SourceRuntimeError( + "US child-support PUF predictors drifted from the archived method: " + f"expected {list(_PREDICTORS)}, got {list(predictors)}." + ) + if operation.parameters["weight"] != _PERSON_WEIGHT_COLUMN: + raise SourceRuntimeError( + "US child-support PUF imputation must use typed person weights." + ) + if operation.parameters["seed_from_build_config"] is not True: + raise SourceRuntimeError( + "US child-support PUF imputation seed must come from the build config." + ) + max_train_samples = int(operation.parameters["max_train_samples"]) + n_estimators = int(operation.parameters["n_estimators"]) + if max_train_samples <= 0 or n_estimators <= 0: + raise SourceRuntimeError( + "US child-support PUF max_train_samples and n_estimators must be positive." + ) + + predictor_columns = [_PREDICTOR_PREFIX + name for name in predictors] + required = [ + _PERSON_WEIGHT_COLUMN, + *predictor_columns, + *US_CHILD_SUPPORT_OUTPUT_COLUMNS, + ] + missing = [column for column in required if column not in frame.columns] + if missing: + raise SourceRuntimeError( + f"US child-support PUF imputation is missing column(s): {missing}." + ) + + channel = frame[_PERSON_SUPPORT_CHANNEL_COLUMN].astype(str) + asec_mask = channel == _BASE_ASEC_SUPPORT_CHANNEL + puf_mask = channel == _PUF_TAX_DETAIL_SUPPORT_CHANNEL + if not asec_mask.any() or not puf_mask.any(): + raise SourceRuntimeError( + "US child-support PUF imputation requires nonempty ASEC and " + "PUF-tax-detail support channels." + ) + + training = frame.loc[ + asec_mask, + [*predictor_columns, *US_CHILD_SUPPORT_OUTPUT_COLUMNS], + ].copy() + training.columns = [*predictors, *US_CHILD_SUPPORT_OUTPUT_COLUMNS] + test = frame.loc[puf_mask, predictor_columns].copy() + test.columns = list(predictors) + weights = pd.to_numeric( + frame.loc[asec_mask, _PERSON_WEIGHT_COLUMN], errors="coerce" + ) + numeric_weights = weights.to_numpy(dtype=np.float64) + if not np.isfinite(numeric_weights).all() or bool((numeric_weights < 0.0).any()): + raise SourceRuntimeError( + "US child-support QRF person weights must be finite and nonnegative." + ) + if float(numeric_weights.sum()) <= 0.0: + raise SourceRuntimeError("US child-support QRF person weights sum to zero.") + + if len(training) > max_train_samples: + sample = training.sample( + n=max_train_samples, + random_state=(context.config.seed if context is not None else 0), + ).index + training = training.loc[sample] + weights = weights.loc[sample] + + for column in (*predictors, *US_CHILD_SUPPORT_OUTPUT_COLUMNS): + training[column] = pd.to_numeric(training[column], errors="coerce") + if not np.isfinite(training[column].to_numpy(dtype=np.float64)).all(): + raise SourceRuntimeError( + f"US child-support QRF training column {column!r} contains " + "nonfinite values." + ) + for column in predictors: + test[column] = pd.to_numeric(test[column], errors="coerce") + if not np.isfinite(test[column].to_numpy(dtype=np.float64)).all(): + raise SourceRuntimeError( + f"US child-support QRF prediction column {column!r} contains " + "nonfinite values." + ) + + global QRF + if QRF is None: + from importlib import import_module + + QRF = import_module("populace.fit").QRF + seed = context.config.seed if context is not None else 0 + fitted = QRF(n_estimators=n_estimators, seed=seed).fit( + training, + list(predictors), + list(US_CHILD_SUPPORT_OUTPUT_COLUMNS), + weights=weights.to_numpy(dtype=np.float64), + ) + predictions = fitted.predict(test) + missing_outputs = [ + output + for output in US_CHILD_SUPPORT_OUTPUT_COLUMNS + if output not in predictions + ] + if missing_outputs: + raise SourceRuntimeError( + f"US child-support QRF prediction is missing output(s): {missing_outputs}." + ) + + result = frame.copy(deep=True) + for output in US_CHILD_SUPPORT_OUTPUT_COLUMNS: + predicted = pd.to_numeric(predictions[output], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not np.isfinite(predicted).all(): + raise SourceRuntimeError( + f"US child-support QRF produced nonfinite {output} values." + ) + if bool((predicted < 0.0).any()): + raise SourceRuntimeError( + f"US child-support QRF produced negative {output} values." + ) + result.loc[puf_mask, output] = predicted + return result + + +def _person_child_support_predictors(frame: Frame) -> pd.DataFrame: + """Build the retired eight second-stage QRF predictors on people.""" + + person = frame.table("person") + tax_unit = frame.table("tax_unit") + predictors = pd.DataFrame(index=person.index) + + def _numeric(*columns: str) -> np.ndarray: + for column in columns: + if column in person: + values = pd.to_numeric(person[column], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not np.isfinite(values).all(): + raise SourceRuntimeError( + "US child-support QRF predictor source " + f"{column!r} contains nonfinite values." + ) + return values + raise SourceRuntimeError( + "US child-support PUF imputation cannot construct a predictor from " + f"any of {list(columns)}." + ) + + predictors["age"] = _numeric("age", "A_AGE") + if "is_male" in person: + predictors["is_male"] = person["is_male"].fillna(False).astype(bool) + elif "is_female" in person: + predictors["is_male"] = ~person["is_female"].fillna(False).astype(bool) + elif "A_SEX" in person: + predictors["is_male"] = _numeric("A_SEX") == 1 + else: + raise SourceRuntimeError( + "US child-support PUF imputation requires is_male, is_female, or " + "measured A_SEX." + ) + if "has_esi" not in person: + raise SourceRuntimeError("US child-support PUF imputation requires has_esi.") + predictors["has_esi"] = person["has_esi"].fillna(False).astype(bool) + + filing_status_column = next( + ( + column + for column in ("filing_status_input", "filing_status") + if column in tax_unit + ), + None, + ) + if filing_status_column is None: + raise SourceRuntimeError( + "US child-support PUF imputation requires tax-unit filing status." + ) + filing_status = tax_unit[filing_status_column].map( + lambda value: ( + value.decode() if isinstance(value, (bytes, np.bytes_)) else str(value) + ) + ) + joint_by_id = pd.Series( + filing_status.eq("JOINT").to_numpy(), + index=tax_unit["tax_unit_id"].to_numpy(), + ) + predictors["tax_unit_is_joint"] = ( + person["person_tax_unit_id"].map(joint_by_id).fillna(False).astype(bool) + ) + + if "tax_unit_role_input" not in person: + raise SourceRuntimeError( + "US child-support PUF imputation requires tax_unit_role_input to " + "count dependents." + ) + dependent = ( + person["tax_unit_role_input"] + .map( + lambda value: ( + value.decode() if isinstance(value, (bytes, np.bytes_)) else str(value) + ) + ) + .eq("DEPENDENT") + ) + predictors["tax_unit_count_dependents"] = ( + dependent.groupby(person["person_tax_unit_id"]) + .transform("sum") + .to_numpy(dtype=np.float64) + ) + predictors["employment_income"] = _numeric( + "employment_income_before_lsr", "WSAL_VAL" + ) + predictors["self_employment_income"] = _numeric( + "self_employment_income_before_lsr", "SEMP_VAL" + ) + social_security_columns = ( + "social_security_retirement", + "social_security_disability", + "social_security_dependents", + "social_security_survivors", + ) + if all(column in person for column in social_security_columns): + predictors["social_security"] = np.sum( + np.column_stack([_numeric(column) for column in social_security_columns]), + axis=1, + ) + else: + predictors["social_security"] = _numeric("SS_VAL") + return predictors + + +def with_us_child_support_inputs( + frame: Frame, + *, + seed: int, + time_period: int, + allow_existing_without_source: bool = False, +) -> Frame: + """Materialize direct ASEC and post-PUF-clone child-support inputs.""" + + if frame.schema != US_SCHEMA: + raise ValueError("US child-support inputs require the US schema.") + person = frame.table("person") + source_available = all( + column in person for column in US_CHILD_SUPPORT_REQUIRED_SOURCE_COLUMNS + ) + if not source_available: + if allow_existing_without_source and _child_support_surface_carries_signal( + frame + ): + return frame + missing = [ + column + for column in US_CHILD_SUPPORT_REQUIRED_SOURCE_COLUMNS + if column not in person + ] + raise ValueError( + "US child-support stage cannot heal a default surface without " + f"measured ASEC source column(s): {missing}." + ) + + stage_person = person.copy(deep=True) + stage_person[_PERSON_WEIGHT_COLUMN] = frame.resolve_weights("person").values + if _PERSON_SUPPORT_CHANNEL_COLUMN in person: + predictors = _person_child_support_predictors(frame) + for column in _PREDICTORS: + stage_person[_PREDICTOR_PREFIX + column] = predictors[column].to_numpy() + output = run_source_stage( + us_child_support_stage_spec(), + tables={"person": stage_person}, + operation_handlers={ + "derive_child_support_inputs": derive_us_child_support_from_manifest, + "impute_child_support_to_puf_support": ( + impute_us_child_support_to_puf_support_from_manifest + ), + }, + config=SourceRuntimeConfig(seed=int(seed), target_year=int(time_period)), + ) + aligned = output.set_index("person_id").reindex(person["person_id"]) + for column in US_CHILD_SUPPORT_OUTPUT_COLUMNS: + if aligned[column].isna().any(): + raise ValueError( + f"US child-support stage output {column!r} does not cover every person." + ) + + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + for column in US_CHILD_SUPPORT_OUTPUT_COLUMNS: + tables["person"][column] = aligned[column].to_numpy(dtype=np.float64) + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + ) + + +def us_child_support_summary(frame: Frame) -> dict[str, object]: + """Return weighted signal and validity diagnostics for both leaves.""" + + person = frame.table("person") + weights = np.asarray(frame.resolve_weights("person").values, dtype=np.float64) + total_weight = float(weights.sum()) + channel = ( + person[_PERSON_SUPPORT_CHANNEL_COLUMN].astype(str).to_numpy() + if _PERSON_SUPPORT_CHANNEL_COLUMN in person + else None + ) + columns: dict[str, dict[str, object]] = {} + for output in US_CHILD_SUPPORT_OUTPUT_COLUMNS: + values = pd.to_numeric(person[output], errors="coerce").to_numpy( + dtype=np.float64 + ) + finite = np.isfinite(values) + positive = finite & (values > 0.0) + detail: dict[str, object] = { + "positive_share": ( + float(weights[positive].sum()) / total_weight + if total_weight > 0.0 + else 0.0 + ), + "positive_share_band": list(_NONZERO_SHARE_BANDS[output]), + "weighted_total": float((np.nan_to_num(values) * weights).sum()), + "nonfinite": int(np.count_nonzero(~finite)), + "negative": int(np.count_nonzero(finite & (values < 0.0))), + } + if channel is not None: + channels: dict[str, dict[str, float | int]] = {} + for name in (_BASE_ASEC_SUPPORT_CHANNEL, _PUF_TAX_DETAIL_SUPPORT_CHANNEL): + mask = channel == name + channel_weight = float(weights[mask].sum()) + channels[name] = { + "positive_rows": int(np.count_nonzero(mask & positive)), + "positive_share": ( + float(weights[mask & positive].sum()) / channel_weight + if channel_weight > 0.0 + else 0.0 + ), + "weighted_total": float( + (np.nan_to_num(values[mask]) * weights[mask]).sum() + ), + } + detail["channels"] = channels + columns[output] = detail + return {"columns": columns} + + +def us_child_support_signal_gate(frame: Frame) -> GateResult: + """Require finite, nonnegative, nondefault signal on both support halves.""" + + person = frame.table("person") + missing = [ + output for output in US_CHILD_SUPPORT_OUTPUT_COLUMNS if output not in person + ] + if missing: + return GateResult( + name="child_support_inputs_signal", + passed=False, + failures=(f"person columns missing: {missing}.",), + details={"missing": missing}, + ) + + summary = us_child_support_summary(frame) + failures: list[str] = [] + columns = summary["columns"] + for output in US_CHILD_SUPPORT_OUTPUT_COLUMNS: + detail = columns[output] + if detail["nonfinite"]: + failures.append(f"{output}: {int(detail['nonfinite'])} nonfinite values.") + if detail["negative"]: + failures.append(f"{output}: {int(detail['negative'])} negative values.") + share = float(detail["positive_share"]) + low, high = detail["positive_share_band"] + if not (low <= share <= high): + failures.append( + f"{output}: nonzero share {share:.4f} outside plausibility " + f"band [{low}, {high}]." + ) + channels = detail.get("channels") + if isinstance(channels, dict): + channel_low, channel_high = _CHANNEL_NONZERO_SHARE_BAND + for name in ( + _BASE_ASEC_SUPPORT_CHANNEL, + _PUF_TAX_DETAIL_SUPPORT_CHANNEL, + ): + channel_detail = channels.get(name) + if not isinstance(channel_detail, dict): + failures.append(f"{output}: missing {name} channel diagnostics.") + continue + channel_share = float(channel_detail["positive_share"]) + if not (channel_low <= channel_share <= channel_high): + failures.append( + f"{output}: {name} nonzero share {channel_share:.4f} " + f"outside plausibility band [{channel_low}, {channel_high}]." + ) + return GateResult( + name="child_support_inputs_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) + + +def _child_support_surface_carries_signal(frame: Frame) -> bool: + if any( + output not in frame.table("person") + for output in US_CHILD_SUPPORT_OUTPUT_COLUMNS + ): + return False + return us_child_support_signal_gate(frame).passed diff --git a/packages/populace-build/src/populace/build/us_runtime/childcare.py b/packages/populace-build/src/populace/build/us_runtime/childcare.py new file mode 100644 index 00000000..43d55414 --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/childcare.py @@ -0,0 +1,590 @@ +"""ASEC-measured childcare expenses with retired PUF-half QRF treatment. + +The retired eCPS pipeline carried measured annual SPM-unit childcare expenses +from ASEC ``SPM_CHILDCAREXPNS`` and then QRF-imputed that CPS-only input onto +the PUF clone half. The QRF used eight person predictors, at most 5,000 ASEC +training rows, and reduced non-person outputs with ``value_from_first_person``. +Immutable archived source coordinates are exposed below. + +The ASEC value is replicated on every member of an SPM unit. This port refuses +missing, nonfinite, negative, or internally inconsistent replicas rather than +silently choosing a maximum or filling zeros. It also uses typed person weights +for the QRF fit: the archived call omitted weights, but weighted fitting is a +hard Populace build contract. PolicyEngine-US owns all downstream CDCC and SPM +formulas; this stage persists only the factual SPM-unit input. +""" + +from __future__ import annotations + +from importlib.resources import files +from typing import Any + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import ( + SourceOperationSpec, + SourceStageSpec, + load_source_manifest, +) +from populace.build.source_runtime import ( + SourceRuntimeConfig, + SourceRuntimeContext, + SourceRuntimeError, + run_source_stage, +) +from populace.frame import Frame +from populace.frame.units import US_SCHEMA + +__all__ = [ + "CHILDCARE_ARCHIVED_CPS_DERIVATION_URL", + "CHILDCARE_ARCHIVED_PUF_IMPUTATION_URL", + "US_CHILDCARE_OUTPUT_COLUMNS", + "US_CHILDCARE_REQUIRED_SOURCE_COLUMNS", + "US_CHILDCARE_STAGE_NAME", + "derive_us_childcare_from_manifest", + "impute_us_childcare_to_puf_support_from_manifest", + "us_childcare_signal_gate", + "us_childcare_stage_spec", + "us_childcare_summary", + "with_us_childcare_inputs", +] + +QRF: Any | None = None + +_ARCHIVED_DATA_REPOSITORY = "policyengine-" + "us-data" +_ARCHIVED_ROOT = ( + "https://github.com/PolicyEngine/" + f"{_ARCHIVED_DATA_REPOSITORY}/blob/" + "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe/" + "policyengine_" + "us_data/" +) +CHILDCARE_ARCHIVED_CPS_DERIVATION_URL = ( + _ARCHIVED_ROOT + "datasets/cps/cps.py#L1612-L1622" +) +CHILDCARE_ARCHIVED_PUF_IMPUTATION_URL = ( + _ARCHIVED_ROOT + "datasets/cps/extended_cps.py#L639-L739" +) + +US_CHILDCARE_STAGE_NAME = "childcare_inputs" +US_CHILDCARE_OUTPUT_COLUMNS: tuple[str, ...] = ( + "spm_unit_pre_subsidy_childcare_expenses", +) +US_CHILDCARE_REQUIRED_SOURCE_COLUMNS: tuple[str, ...] = ( + "person_spm_unit_id", + "SPM_CHILDCAREXPNS", +) + +_OUTPUT = US_CHILDCARE_OUTPUT_COLUMNS[0] +_PERSON_WEIGHT_COLUMN = "person_weight" +_PERSON_SUPPORT_CHANNEL_COLUMN = "person_support_channel" +_BASE_ASEC_SUPPORT_CHANNEL = "asec" +_PUF_TAX_DETAIL_SUPPORT_CHANNEL = "puf_tax_detail" +_PUF_PREDICTORS: tuple[str, ...] = ( + "age", + "is_male", + "has_esi", + "tax_unit_is_joint", + "tax_unit_count_dependents", + "employment_income", + "self_employment_income", + "social_security", +) +_PUF_PREDICTOR_PREFIX = "childcare_predictor_" +_DIRECT_PARAMETER_KEYS = frozenset() +_PUF_IMPUTATION_PARAMETER_KEYS = frozenset( + { + "predictors", + "max_train_samples", + "n_estimators", + "seed_from_build_config", + "weight", + "reduction", + } +) +_NONZERO_SHARE_BAND = (0.005, 0.50) +_REPLICA_ATOL = 1e-6 + + +def us_childcare_stage_spec() -> SourceStageSpec: + """Load the packaged childcare source-stage declaration.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + stage_map = manifest.stage_map() + if US_CHILDCARE_STAGE_NAME not in stage_map: + raise ValueError( + f"US source manifest declares no {US_CHILDCARE_STAGE_NAME!r} stage." + ) + spec = stage_map[US_CHILDCARE_STAGE_NAME] + missing = sorted(set(US_CHILDCARE_OUTPUT_COLUMNS) - set(spec.outputs)) + if missing: + raise ValueError( + f"{US_CHILDCARE_STAGE_NAME!r} manifest stage does not declare " + f"output(s) {missing}; the runtime and manifest have drifted." + ) + return spec + + +def _numeric_source(frame: pd.DataFrame, column: str) -> np.ndarray: + values = pd.to_numeric(frame[column], errors="coerce").to_numpy(dtype=np.float64) + nonfinite = int(np.count_nonzero(~np.isfinite(values))) + if nonfinite: + raise SourceRuntimeError( + f"US childcare source {column!r} contains {nonfinite} nonfinite " + "value(s); the measured source must not be silently replaced." + ) + return values + + +def derive_us_childcare_from_manifest( + frame: pd.DataFrame | None, + operation: SourceOperationSpec, + _context: SourceRuntimeContext | None, +) -> pd.DataFrame: + """Carry measured, replicated ASEC childcare values on person rows.""" + + if operation.kind != "derive_childcare_inputs": + raise SourceRuntimeError( + f"US childcare derivation received unexpected operation {operation.kind!r}." + ) + if frame is None: + raise SourceRuntimeError( + "US childcare derivation requires the person table to be read first." + ) + unexpected = sorted(set(operation.parameters) - _DIRECT_PARAMETER_KEYS) + if unexpected: + raise SourceRuntimeError( + f"US childcare derivation received unsupported parameter(s): {unexpected}." + ) + missing = [ + column + for column in US_CHILDCARE_REQUIRED_SOURCE_COLUMNS + if column not in frame.columns + ] + if missing: + raise SourceRuntimeError( + "US childcare derivation requires measured ASEC source column(s): " + f"{missing}." + ) + + values = _numeric_source(frame, "SPM_CHILDCAREXPNS") + negative = int(np.count_nonzero(values < 0.0)) + if negative: + raise SourceRuntimeError( + "US childcare source 'SPM_CHILDCAREXPNS' contains " + f"{negative} negative value(s)." + ) + spm_ids = pd.to_numeric(frame["person_spm_unit_id"], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not np.isfinite(spm_ids).all(): + raise SourceRuntimeError( + "US childcare source person_spm_unit_id contains nonfinite values." + ) + replicas = pd.DataFrame({"spm_unit_id": spm_ids, "value": values}) + bounds = replicas.groupby("spm_unit_id", sort=False)["value"].agg(["min", "max"]) + inconsistent = ~np.isclose( + bounds["min"].to_numpy(dtype=np.float64), + bounds["max"].to_numpy(dtype=np.float64), + rtol=0.0, + atol=_REPLICA_ATOL, + ) + if bool(inconsistent.any()): + bad_ids = bounds.index[inconsistent].astype("int64").tolist() + raise SourceRuntimeError( + "US childcare source SPM_CHILDCAREXPNS disagrees within replicated " + f"SPM unit(s): {bad_ids[:10]}." + ) + + result = frame.copy(deep=True) + result[_OUTPUT] = values + return result + + +def impute_us_childcare_to_puf_support_from_manifest( + frame: pd.DataFrame | None, + operation: SourceOperationSpec, + context: SourceRuntimeContext | None, +) -> pd.DataFrame: + """QRF-impute childcare onto PUF people before SPM-unit reduction.""" + + if operation.kind != "impute_childcare_to_puf_support": + raise SourceRuntimeError( + "US childcare PUF imputation received unexpected operation " + f"{operation.kind!r}." + ) + if frame is None: + raise SourceRuntimeError( + "US childcare PUF imputation requires the person table to be read first." + ) + unexpected = sorted(set(operation.parameters) - _PUF_IMPUTATION_PARAMETER_KEYS) + missing_parameters = sorted( + _PUF_IMPUTATION_PARAMETER_KEYS - set(operation.parameters) + ) + if unexpected or missing_parameters: + raise SourceRuntimeError( + "US childcare PUF imputation parameters must match the archived " + f"method; missing={missing_parameters}, unexpected={unexpected}." + ) + if _PERSON_SUPPORT_CHANNEL_COLUMN not in frame.columns: + return frame.copy(deep=True) + + predictors = tuple(str(value) for value in operation.parameters["predictors"]) + if predictors != _PUF_PREDICTORS: + raise SourceRuntimeError( + "US childcare PUF predictors drifted from the archived method: " + f"expected {list(_PUF_PREDICTORS)}, got {list(predictors)}." + ) + if operation.parameters["weight"] != _PERSON_WEIGHT_COLUMN: + raise SourceRuntimeError( + "US childcare PUF imputation must use the typed person-weight " + f"column {_PERSON_WEIGHT_COLUMN!r}." + ) + if operation.parameters["seed_from_build_config"] is not True: + raise SourceRuntimeError( + "US childcare PUF imputation seed must come from the build config." + ) + if operation.parameters["reduction"] != "value_from_first_person": + raise SourceRuntimeError( + "US childcare PUF imputation must use the archived " + "value_from_first_person reduction." + ) + max_train_samples = int(operation.parameters["max_train_samples"]) + n_estimators = int(operation.parameters["n_estimators"]) + if max_train_samples <= 0 or n_estimators <= 0: + raise SourceRuntimeError( + "US childcare PUF max_train_samples and n_estimators must be positive." + ) + + predictor_columns = [_PUF_PREDICTOR_PREFIX + name for name in predictors] + required = [_PERSON_WEIGHT_COLUMN, *predictor_columns, _OUTPUT] + missing = [column for column in required if column not in frame.columns] + if missing: + raise SourceRuntimeError( + f"US childcare PUF imputation is missing source column(s): {missing}." + ) + + channel = frame[_PERSON_SUPPORT_CHANNEL_COLUMN].astype(str) + asec_mask = channel == _BASE_ASEC_SUPPORT_CHANNEL + puf_mask = channel == _PUF_TAX_DETAIL_SUPPORT_CHANNEL + if not asec_mask.any() or not puf_mask.any(): + raise SourceRuntimeError( + "US childcare PUF imputation requires nonempty ASEC and " + "PUF-tax-detail support channels." + ) + + training = frame.loc[asec_mask, [*predictor_columns, _OUTPUT]].copy() + training.columns = [*predictors, _OUTPUT] + test = frame.loc[puf_mask, predictor_columns].copy() + test.columns = list(predictors) + weights = pd.to_numeric( + frame.loc[asec_mask, _PERSON_WEIGHT_COLUMN], errors="coerce" + ) + numeric_weights = weights.to_numpy(dtype=np.float64) + if not np.isfinite(numeric_weights).all() or bool((numeric_weights < 0.0).any()): + raise SourceRuntimeError( + "US childcare QRF person weights must be finite and nonnegative." + ) + if float(numeric_weights.sum()) <= 0.0: + raise SourceRuntimeError("US childcare QRF person weights sum to zero.") + + if len(training) > max_train_samples: + sample = training.sample( + n=max_train_samples, + random_state=(context.config.seed if context is not None else 0), + ).index + training = training.loc[sample] + weights = weights.loc[sample] + + for column in (*predictors, _OUTPUT): + training[column] = pd.to_numeric(training[column], errors="coerce") + if not np.isfinite(training[column].to_numpy(dtype=np.float64)).all(): + raise SourceRuntimeError( + f"US childcare QRF training column {column!r} contains " + "nonfinite values." + ) + for column in predictors: + test[column] = pd.to_numeric(test[column], errors="coerce") + if not np.isfinite(test[column].to_numpy(dtype=np.float64)).all(): + raise SourceRuntimeError( + f"US childcare QRF prediction column {column!r} contains " + "nonfinite values." + ) + + global QRF + if QRF is None: + from importlib import import_module + + QRF = import_module("populace.fit").QRF + seed = context.config.seed if context is not None else 0 + fitted = QRF(n_estimators=n_estimators, seed=seed).fit( + training, + list(predictors), + [_OUTPUT], + weights=weights.to_numpy(dtype=np.float64), + ) + predictions = fitted.predict(test) + if _OUTPUT not in predictions: + raise SourceRuntimeError( + f"US childcare QRF prediction is missing output {_OUTPUT!r}." + ) + predicted = pd.to_numeric(predictions[_OUTPUT], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not np.isfinite(predicted).all(): + raise SourceRuntimeError("US childcare QRF produced nonfinite predictions.") + if bool((predicted < 0.0).any()): + raise SourceRuntimeError("US childcare QRF produced negative predictions.") + + result = frame.copy(deep=True) + result.loc[puf_mask, _OUTPUT] = predicted + return result + + +def _person_childcare_predictors(frame: Frame) -> pd.DataFrame: + """Build the retired eight QRF predictors on person rows.""" + + person = frame.table("person") + tax_unit = frame.table("tax_unit") + predictors = pd.DataFrame(index=person.index) + + def _numeric(*columns: str) -> np.ndarray: + for column in columns: + if column in person: + values = pd.to_numeric(person[column], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not np.isfinite(values).all(): + raise SourceRuntimeError( + "US childcare QRF predictor source " + f"{column!r} contains nonfinite values." + ) + return values + raise SourceRuntimeError( + "US childcare PUF imputation cannot construct predictor from any " + f"of {list(columns)}." + ) + + predictors["age"] = _numeric("age", "A_AGE") + if "is_male" in person: + predictors["is_male"] = person["is_male"].fillna(False).astype(bool) + elif "is_female" in person: + predictors["is_male"] = ~person["is_female"].fillna(False).astype(bool) + elif "A_SEX" in person: + predictors["is_male"] = _numeric("A_SEX") == 1 + else: + raise SourceRuntimeError( + "US childcare PUF imputation requires is_male, is_female, or " + "measured A_SEX." + ) + if "has_esi" not in person: + raise SourceRuntimeError("US childcare PUF imputation requires has_esi.") + predictors["has_esi"] = person["has_esi"].fillna(False).astype(bool) + + filing_status_column = next( + ( + column + for column in ("filing_status_input", "filing_status") + if column in tax_unit + ), + None, + ) + if filing_status_column is None: + raise SourceRuntimeError( + "US childcare PUF imputation requires tax-unit filing status." + ) + filing_status = tax_unit[filing_status_column].map( + lambda value: ( + value.decode() if isinstance(value, (bytes, np.bytes_)) else str(value) + ) + ) + joint_by_id = pd.Series( + filing_status.eq("JOINT").to_numpy(), + index=tax_unit["tax_unit_id"].to_numpy(), + ) + predictors["tax_unit_is_joint"] = ( + person["person_tax_unit_id"].map(joint_by_id).fillna(False).astype(bool) + ) + + if "tax_unit_role_input" not in person: + raise SourceRuntimeError( + "US childcare PUF imputation requires tax_unit_role_input to count " + "dependents." + ) + dependent = ( + person["tax_unit_role_input"] + .map( + lambda value: ( + value.decode() if isinstance(value, (bytes, np.bytes_)) else str(value) + ) + ) + .eq("DEPENDENT") + ) + predictors["tax_unit_count_dependents"] = ( + dependent.groupby(person["person_tax_unit_id"]) + .transform("sum") + .to_numpy(dtype=np.float64) + ) + predictors["employment_income"] = _numeric( + "employment_income_before_lsr", "WSAL_VAL" + ) + predictors["self_employment_income"] = _numeric( + "self_employment_income_before_lsr", "SEMP_VAL" + ) + social_security_columns = ( + "social_security_retirement", + "social_security_disability", + "social_security_dependents", + "social_security_survivors", + ) + if all(column in person for column in social_security_columns): + predictors["social_security"] = np.sum( + np.column_stack([_numeric(column) for column in social_security_columns]), + axis=1, + ) + else: + predictors["social_security"] = _numeric("SS_VAL") + return predictors + + +def with_us_childcare_inputs( + frame: Frame, + *, + seed: int, + time_period: int, + allow_existing_without_source: bool = False, +) -> Frame: + """Materialize measured/QRF childcare on the SPM-unit input leaf. + + ``allow_existing_without_source`` is only for a downstream release build + consuming a base artifact that already passed this stage's signal gate. + Source/base construction remains strict and cannot reuse clone duplicates. + """ + + if frame.schema != US_SCHEMA: + raise ValueError("US childcare inputs require the US schema.") + person = frame.table("person") + spm_unit = frame.table("spm_unit") + source_available = all( + column in person for column in US_CHILDCARE_REQUIRED_SOURCE_COLUMNS + ) + if not source_available: + if allow_existing_without_source and _childcare_surface_carries_signal(frame): + return frame + missing = [ + column + for column in US_CHILDCARE_REQUIRED_SOURCE_COLUMNS + if column not in person + ] + raise ValueError( + "US childcare stage cannot heal a default surface without measured " + f"ASEC source column(s): {missing}." + ) + + stage_person = person.copy(deep=True) + stage_person[_PERSON_WEIGHT_COLUMN] = frame.resolve_weights("person").values + if _PERSON_SUPPORT_CHANNEL_COLUMN in person: + predictors = _person_childcare_predictors(frame) + for column in _PUF_PREDICTORS: + stage_person[_PUF_PREDICTOR_PREFIX + column] = predictors[column].to_numpy() + output = run_source_stage( + us_childcare_stage_spec(), + tables={"person": stage_person}, + operation_handlers={ + "derive_childcare_inputs": derive_us_childcare_from_manifest, + "impute_childcare_to_puf_support": ( + impute_us_childcare_to_puf_support_from_manifest + ), + }, + config=SourceRuntimeConfig(seed=int(seed), target_year=int(time_period)), + ) + aligned_people = output.set_index("person_id").reindex(person["person_id"]) + if aligned_people[_OUTPUT].isna().any(): + raise ValueError( + "US childcare stage output does not cover every person before " + "SPM-unit reduction." + ) + unit_values = ( + aligned_people.assign( + person_spm_unit_id=person["person_spm_unit_id"].to_numpy() + ) + .groupby("person_spm_unit_id", sort=False)[_OUTPUT] + .first() + ) + aligned_units = unit_values.reindex(spm_unit["spm_unit_id"]) + if aligned_units.isna().any(): + raise ValueError("US childcare stage output does not cover every SPM unit.") + + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + tables["spm_unit"][_OUTPUT] = aligned_units.to_numpy(dtype=np.float64) + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + ) + + +def us_childcare_summary(frame: Frame) -> dict[str, object]: + """Return weighted signal and validity diagnostics for childcare.""" + + values = pd.to_numeric(frame.table("spm_unit")[_OUTPUT], errors="coerce").to_numpy( + dtype=np.float64 + ) + weights = np.asarray(frame.resolve_weights("spm_unit").values, dtype=np.float64) + finite = np.isfinite(values) + positive = finite & (values > 0.0) + total_weight = float(weights.sum()) + positive_share = ( + float(weights[positive].sum()) / total_weight if total_weight > 0.0 else 0.0 + ) + return { + "positive_share": positive_share, + "positive_share_band": list(_NONZERO_SHARE_BAND), + "weighted_total": float(np.sum(np.nan_to_num(values) * weights)), + "nonfinite": int(np.count_nonzero(~finite)), + "negative": int(np.count_nonzero(finite & (values < 0.0))), + } + + +def us_childcare_signal_gate(frame: Frame) -> GateResult: + """Require finite, nonnegative, nondefault SPM-unit childcare signal.""" + + spm_unit = frame.table("spm_unit") + if _OUTPUT not in spm_unit: + return GateResult( + name="childcare_inputs_signal", + passed=False, + failures=(f"spm_unit columns missing: [{_OUTPUT!r}].",), + details={"missing": [_OUTPUT]}, + ) + + summary = us_childcare_summary(frame) + failures: list[str] = [] + if summary["nonfinite"]: + failures.append(f"{_OUTPUT}: {int(summary['nonfinite'])} nonfinite values.") + if summary["negative"]: + failures.append(f"{_OUTPUT}: {int(summary['negative'])} negative values.") + share = float(summary["positive_share"]) + low, high = summary["positive_share_band"] + if not (low <= share <= high): + failures.append( + f"{_OUTPUT}: nonzero share {share:.4f} outside plausibility band " + f"[{low}, {high}]." + ) + return GateResult( + name="childcare_inputs_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) + + +def _childcare_surface_carries_signal(frame: Frame) -> bool: + if _OUTPUT not in frame.table("spm_unit"): + return False + return us_childcare_signal_gate(frame).passed diff --git a/packages/populace-build/src/populace/build/us_runtime/cps_carried.py b/packages/populace-build/src/populace/build/us_runtime/cps_carried.py index 0b4d1f65..21283ec6 100644 --- a/packages/populace-build/src/populace/build/us_runtime/cps_carried.py +++ b/packages/populace-build/src/populace/build/us_runtime/cps_carried.py @@ -13,6 +13,7 @@ import numpy as np import pandas as pd +from populace.build.us_runtime.alimony import derive_us_alimony_from_asec from populace.frame import US_SCHEMA, Frame __all__ = [ @@ -63,7 +64,7 @@ "other_medical_expenses", "over_the_counter_health_expenses", "rental_income", - "farm_income", + "farm_operations_income", "has_champva_health_coverage_at_interview", "has_esi", "has_indian_health_service_coverage_at_interview", @@ -76,6 +77,7 @@ "has_va_health_coverage_at_interview", "is_female", "unemployment_compensation", + "alimony_income", "miscellaneous_income", } ) @@ -154,11 +156,13 @@ def derive_us_cps_carried_inputs(frame: Frame) -> Frame: ) _fill_missing(person, "taxable_ira_distributions", _ira_distributions(person)) + person = derive_us_alimony_from_asec(person) + tables["person"] = person + direct_sources: Mapping[str, str] = { "rental_income": "RNT_VAL", - "farm_income": "FRSE_VAL", + "farm_operations_income": "FRSE_VAL", "unemployment_compensation": "UC_VAL", - "miscellaneous_income": "OI_VAL", "health_insurance_premiums_without_medicare_part_b": "PHIP_VAL", "medicare_part_b_premiums": "PEMCPREM", "other_medical_expenses": "PMED_VAL", diff --git a/packages/populace-build/src/populace/build/us_runtime/disability_benefits.py b/packages/populace-build/src/populace/build/us_runtime/disability_benefits.py new file mode 100644 index 00000000..716b469b --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/disability_benefits.py @@ -0,0 +1,657 @@ +"""CPS ASEC non-SSA, non-workers-comp disability benefit input. + +The retired eCPS pipeline summed the two annual ASEC disability-income slots +only when their source code was not 1 (workers' compensation). It then used +the common eight-predictor, at-most-5,000-person QRF to replace this CPS-only +leaf on the PUF support half. Immutable archived coordinates are exposed +below. + +This port preserves positive annual dollars at person grain. It does not fold +Social Security disability or workers' compensation into this leaf, and it +does not invent eligibility masks or rebalance the two source slots. The QRF +uses typed person weights under Populace's build-wide weighting contract. +""" + +from __future__ import annotations + +from importlib.resources import files +from typing import Any + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import ( + SourceOperationSpec, + SourceStageSpec, + load_source_manifest, +) +from populace.build.source_runtime import ( + SourceRuntimeConfig, + SourceRuntimeContext, + SourceRuntimeError, + run_source_stage, +) +from populace.frame import Frame +from populace.frame.units import US_SCHEMA + +__all__ = [ + "DISABILITY_BENEFITS_ARCHIVED_DERIVATION_URL", + "DISABILITY_BENEFITS_ARCHIVED_PUF_IMPUTATION_URL", + "DISABILITY_BENEFITS_ARCHIVED_PUF_OUTPUTS_URL", + "DISABILITY_BENEFITS_ARCHIVED_SOURCE_COLUMNS_URL", + "US_DISABILITY_BENEFITS_NONCONSTANT_PERSON_COLUMNS", + "US_DISABILITY_BENEFITS_OUTPUT_COLUMNS", + "US_DISABILITY_BENEFITS_REQUIRED_SOURCE_COLUMNS", + "US_DISABILITY_BENEFITS_STAGE_NAME", + "derive_us_disability_benefits_from_asec", + "derive_us_disability_benefits_from_manifest", + "impute_us_disability_benefits_to_puf_support_from_manifest", + "us_disability_benefits_signal_gate", + "us_disability_benefits_stage_spec", + "us_disability_benefits_summary", + "with_us_disability_benefits", +] + +QRF: Any | None = None + +_ARCHIVED_DATA_REPOSITORY = "policyengine-" + "us-data" +_ARCHIVED_ROOT = ( + "https://github.com/PolicyEngine/" + f"{_ARCHIVED_DATA_REPOSITORY}/blob/" + "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe/" + "policyengine_" + "us_data/" +) +DISABILITY_BENEFITS_ARCHIVED_DERIVATION_URL = ( + _ARCHIVED_ROOT + "datasets/cps/cps.py#L1561-L1571" +) +DISABILITY_BENEFITS_ARCHIVED_SOURCE_COLUMNS_URL = ( + _ARCHIVED_ROOT + "datasets/cps/census_cps.py#L306-L381" +) +DISABILITY_BENEFITS_ARCHIVED_PUF_OUTPUTS_URL = ( + _ARCHIVED_ROOT + "datasets/cps/extended_cps.py#L135-L194" +) +DISABILITY_BENEFITS_ARCHIVED_PUF_IMPUTATION_URL = ( + _ARCHIVED_ROOT + "datasets/cps/extended_cps.py#L639-L745" +) + +US_DISABILITY_BENEFITS_STAGE_NAME = "disability_benefits_input" +US_DISABILITY_BENEFITS_OUTPUT_COLUMNS: tuple[str, ...] = ("disability_benefits",) +US_DISABILITY_BENEFITS_NONCONSTANT_PERSON_COLUMNS = ( + US_DISABILITY_BENEFITS_OUTPUT_COLUMNS +) +US_DISABILITY_BENEFITS_REQUIRED_SOURCE_COLUMNS: tuple[str, ...] = ( + "DIS_VAL1", + "DIS_SC1", + "DIS_VAL2", + "DIS_SC2", +) + +_OUTPUT = US_DISABILITY_BENEFITS_OUTPUT_COLUMNS[0] +_PERSON_WEIGHT_COLUMN = "person_weight" +_PERSON_SUPPORT_CHANNEL_COLUMN = "person_support_channel" +_BASE_ASEC_SUPPORT_CHANNEL = "asec" +_PUF_TAX_DETAIL_SUPPORT_CHANNEL = "puf_tax_detail" +_PREDICTORS: tuple[str, ...] = ( + "age", + "is_male", + "has_esi", + "tax_unit_is_joint", + "tax_unit_count_dependents", + "employment_income", + "self_employment_income", + "social_security", +) +_PREDICTOR_PREFIX = "disability_benefits_predictor_" +_EXPECTED_DIRECT_PARAMETERS = { + "first_amount_source": "DIS_VAL1", + "first_code_source": "DIS_SC1", + "second_amount_source": "DIS_VAL2", + "second_code_source": "DIS_SC2", + "workers_compensation_code": 1, + "output": _OUTPUT, +} +_PUF_IMPUTATION_PARAMETER_KEYS = frozenset( + { + "predictors", + "max_train_samples", + "n_estimators", + "seed_from_build_config", + "weight", + } +) +# Broad source-plausibility bounds: reject default/near-universal surfaces, +# while leaving the exact raw-source and support-channel tests to pin semantics. +_NONZERO_SHARE_BAND = (0.003, 0.15) +_CHANNEL_NONZERO_SHARE_BANDS = { + _BASE_ASEC_SUPPORT_CHANNEL: (0.003, 0.02), + _PUF_TAX_DETAIL_SUPPORT_CHANNEL: (0.002, 0.15), +} + + +def us_disability_benefits_stage_spec() -> SourceStageSpec: + """Load and validate the packaged disability-benefits stage.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + stage_map = manifest.stage_map() + if US_DISABILITY_BENEFITS_STAGE_NAME not in stage_map: + raise ValueError( + "US source manifest declares no " + f"{US_DISABILITY_BENEFITS_STAGE_NAME!r} stage." + ) + spec = stage_map[US_DISABILITY_BENEFITS_STAGE_NAME] + missing = sorted(set(US_DISABILITY_BENEFITS_OUTPUT_COLUMNS) - set(spec.outputs)) + if missing: + raise ValueError( + f"{US_DISABILITY_BENEFITS_STAGE_NAME!r} manifest stage does not " + f"declare output(s) {missing}; the runtime and manifest have drifted." + ) + return spec + + +def _strict_numeric_source( + person: pd.DataFrame, + column: str, + *, + nonnegative: bool, +) -> np.ndarray: + if column not in person.columns: + raise SourceRuntimeError( + f"US disability-benefits derivation requires ASEC source column {column!r}." + ) + values = pd.to_numeric(person[column], errors="coerce").to_numpy(dtype=np.float64) + nonfinite = int(np.count_nonzero(~np.isfinite(values))) + if nonfinite: + raise SourceRuntimeError( + f"US disability-benefits source {column!r} contains {nonfinite} " + "nonnumeric or nonfinite value(s)." + ) + if nonnegative: + negative = int(np.count_nonzero(values < 0.0)) + if negative: + raise SourceRuntimeError( + f"US disability-benefits source {column!r} contains {negative} " + "negative value(s)." + ) + return values + + +def derive_us_disability_benefits_from_asec( + person: pd.DataFrame, + *, + first_amount_source: str = "DIS_VAL1", + first_code_source: str = "DIS_SC1", + second_amount_source: str = "DIS_VAL2", + second_code_source: str = "DIS_SC2", + workers_compensation_code: int = 1, + output_column: str = _OUTPUT, +) -> pd.DataFrame: + """Apply the retired two-slot, non-workers-compensation annual sum.""" + + first_amount = _strict_numeric_source( + person, + first_amount_source, + nonnegative=True, + ) + first_code = _strict_numeric_source( + person, + first_code_source, + nonnegative=False, + ) + second_amount = _strict_numeric_source( + person, + second_amount_source, + nonnegative=True, + ) + second_code = _strict_numeric_source( + person, + second_code_source, + nonnegative=False, + ) + result = person.copy(deep=True) + result[output_column] = first_amount * ( + first_code != workers_compensation_code + ) + second_amount * (second_code != workers_compensation_code) + return result + + +def derive_us_disability_benefits_from_manifest( + frame: pd.DataFrame | None, + operation: SourceOperationSpec, + _context: SourceRuntimeContext | None, +) -> pd.DataFrame: + """Interpret the manifest's exact ASEC two-slot derivation.""" + + if operation.kind != "derive_disability_benefits": + raise SourceRuntimeError( + "US disability-benefits derivation received unexpected operation " + f"{operation.kind!r}." + ) + if frame is None: + raise SourceRuntimeError( + "US disability-benefits derivation requires the person table first." + ) + parameters = dict(operation.parameters) + if parameters != _EXPECTED_DIRECT_PARAMETERS: + raise SourceRuntimeError( + "US disability-benefits derivation drifted from the archived method: " + f"expected {_EXPECTED_DIRECT_PARAMETERS}, got {parameters}." + ) + return derive_us_disability_benefits_from_asec( + frame, + first_amount_source=parameters["first_amount_source"], + first_code_source=parameters["first_code_source"], + second_amount_source=parameters["second_amount_source"], + second_code_source=parameters["second_code_source"], + workers_compensation_code=int(parameters["workers_compensation_code"]), + output_column=parameters["output"], + ) + + +def impute_us_disability_benefits_to_puf_support_from_manifest( + frame: pd.DataFrame | None, + operation: SourceOperationSpec, + context: SourceRuntimeContext | None, +) -> pd.DataFrame: + """QRF-impute the CPS-only leaf onto PUF-clone people.""" + + if operation.kind != "impute_disability_benefits_to_puf_support": + raise SourceRuntimeError( + "US disability-benefits PUF imputation received unexpected operation " + f"{operation.kind!r}." + ) + if frame is None: + raise SourceRuntimeError( + "US disability-benefits PUF imputation requires the person table first." + ) + unexpected = sorted(set(operation.parameters) - _PUF_IMPUTATION_PARAMETER_KEYS) + missing_parameters = sorted( + _PUF_IMPUTATION_PARAMETER_KEYS - set(operation.parameters) + ) + if unexpected or missing_parameters: + raise SourceRuntimeError( + "US disability-benefits PUF imputation parameters must match the " + f"archived method; missing={missing_parameters}, " + f"unexpected={unexpected}." + ) + if _PERSON_SUPPORT_CHANNEL_COLUMN not in frame.columns: + return frame.copy(deep=True) + + predictors = tuple(str(value) for value in operation.parameters["predictors"]) + if predictors != _PREDICTORS: + raise SourceRuntimeError( + "US disability-benefits PUF predictors drifted from the archived " + f"method: expected {list(_PREDICTORS)}, got {list(predictors)}." + ) + if operation.parameters["weight"] != _PERSON_WEIGHT_COLUMN: + raise SourceRuntimeError( + "US disability-benefits PUF imputation must use typed person weights." + ) + if operation.parameters["seed_from_build_config"] is not True: + raise SourceRuntimeError( + "US disability-benefits PUF imputation seed must come from the build " + "config." + ) + max_train_samples = int(operation.parameters["max_train_samples"]) + n_estimators = int(operation.parameters["n_estimators"]) + if max_train_samples <= 0 or n_estimators <= 0: + raise SourceRuntimeError( + "US disability-benefits PUF max_train_samples and n_estimators must " + "be positive." + ) + + predictor_columns = [_PREDICTOR_PREFIX + name for name in predictors] + required = [_PERSON_WEIGHT_COLUMN, *predictor_columns, _OUTPUT] + missing = [column for column in required if column not in frame.columns] + if missing: + raise SourceRuntimeError( + f"US disability-benefits PUF imputation is missing column(s): {missing}." + ) + + channel = frame[_PERSON_SUPPORT_CHANNEL_COLUMN].astype(str) + asec_mask = channel == _BASE_ASEC_SUPPORT_CHANNEL + puf_mask = channel == _PUF_TAX_DETAIL_SUPPORT_CHANNEL + if not asec_mask.any() or not puf_mask.any(): + raise SourceRuntimeError( + "US disability-benefits PUF imputation requires nonempty ASEC and " + "PUF-tax-detail support channels." + ) + + training = frame.loc[asec_mask, [*predictor_columns, _OUTPUT]].copy() + training.columns = [*predictors, _OUTPUT] + test = frame.loc[puf_mask, predictor_columns].copy() + test.columns = list(predictors) + weights = pd.to_numeric( + frame.loc[asec_mask, _PERSON_WEIGHT_COLUMN], errors="coerce" + ) + numeric_weights = weights.to_numpy(dtype=np.float64) + if not np.isfinite(numeric_weights).all() or bool((numeric_weights < 0.0).any()): + raise SourceRuntimeError( + "US disability-benefits QRF person weights must be finite and nonnegative." + ) + if float(numeric_weights.sum()) <= 0.0: + raise SourceRuntimeError( + "US disability-benefits QRF person weights sum to zero." + ) + + if len(training) > max_train_samples: + sample = training.sample( + n=max_train_samples, + random_state=(context.config.seed if context is not None else 0), + ).index + training = training.loc[sample] + weights = weights.loc[sample] + + for column in (*predictors, _OUTPUT): + training[column] = pd.to_numeric(training[column], errors="coerce") + if not np.isfinite(training[column].to_numpy(dtype=np.float64)).all(): + raise SourceRuntimeError( + f"US disability-benefits QRF training column {column!r} " + "contains nonfinite values." + ) + for column in predictors: + test[column] = pd.to_numeric(test[column], errors="coerce") + if not np.isfinite(test[column].to_numpy(dtype=np.float64)).all(): + raise SourceRuntimeError( + f"US disability-benefits QRF prediction column {column!r} " + "contains nonfinite values." + ) + + global QRF + if QRF is None: + from importlib import import_module + + QRF = import_module("populace.fit").QRF + seed = context.config.seed if context is not None else 0 + fitted = QRF(n_estimators=n_estimators, seed=seed).fit( + training, + list(predictors), + [_OUTPUT], + weights=weights.to_numpy(dtype=np.float64), + ) + predictions = fitted.predict(test) + if _OUTPUT not in predictions: + raise SourceRuntimeError( + f"US disability-benefits QRF prediction is missing {_OUTPUT!r}." + ) + predicted = pd.to_numeric(predictions[_OUTPUT], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not np.isfinite(predicted).all(): + raise SourceRuntimeError( + f"US disability-benefits QRF produced nonfinite {_OUTPUT} values." + ) + if bool((predicted < 0.0).any()): + raise SourceRuntimeError( + f"US disability-benefits QRF produced negative {_OUTPUT} values." + ) + result = frame.copy(deep=True) + result.loc[puf_mask, _OUTPUT] = predicted + return result + + +def _person_disability_benefits_predictors(frame: Frame) -> pd.DataFrame: + """Build the retired eight second-stage QRF predictors on people.""" + + person = frame.table("person") + tax_unit = frame.table("tax_unit") + predictors = pd.DataFrame(index=person.index) + + def _numeric(*columns: str) -> np.ndarray: + for column in columns: + if column in person: + values = pd.to_numeric(person[column], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not np.isfinite(values).all(): + raise SourceRuntimeError( + "US disability-benefits QRF predictor source " + f"{column!r} contains nonfinite values." + ) + return values + raise SourceRuntimeError( + "US disability-benefits PUF imputation cannot construct a predictor " + f"from any of {list(columns)}." + ) + + predictors["age"] = _numeric("age", "A_AGE") + if "is_male" in person: + predictors["is_male"] = person["is_male"].fillna(False).astype(bool) + elif "is_female" in person: + predictors["is_male"] = ~person["is_female"].fillna(False).astype(bool) + elif "A_SEX" in person: + predictors["is_male"] = _numeric("A_SEX") == 1 + else: + raise SourceRuntimeError( + "US disability-benefits PUF imputation requires is_male, is_female, " + "or measured A_SEX." + ) + if "has_esi" not in person: + raise SourceRuntimeError( + "US disability-benefits PUF imputation requires has_esi." + ) + predictors["has_esi"] = person["has_esi"].fillna(False).astype(bool) + + filing_status_column = next( + ( + column + for column in ("filing_status_input", "filing_status") + if column in tax_unit + ), + None, + ) + if filing_status_column is None: + raise SourceRuntimeError( + "US disability-benefits PUF imputation requires tax-unit filing status." + ) + filing_status = tax_unit[filing_status_column].map( + lambda value: ( + value.decode() if isinstance(value, (bytes, np.bytes_)) else str(value) + ) + ) + joint_by_id = pd.Series( + filing_status.eq("JOINT").to_numpy(), + index=tax_unit["tax_unit_id"].to_numpy(), + ) + predictors["tax_unit_is_joint"] = ( + person["person_tax_unit_id"].map(joint_by_id).fillna(False).astype(bool) + ) + + if "tax_unit_role_input" not in person: + raise SourceRuntimeError( + "US disability-benefits PUF imputation requires tax_unit_role_input " + "to count dependents." + ) + dependent = ( + person["tax_unit_role_input"] + .map( + lambda value: ( + value.decode() if isinstance(value, (bytes, np.bytes_)) else str(value) + ) + ) + .eq("DEPENDENT") + ) + predictors["tax_unit_count_dependents"] = ( + dependent.groupby(person["person_tax_unit_id"]) + .transform("sum") + .to_numpy(dtype=np.float64) + ) + predictors["employment_income"] = _numeric( + "employment_income_before_lsr", "WSAL_VAL" + ) + predictors["self_employment_income"] = _numeric( + "self_employment_income_before_lsr", "SEMP_VAL" + ) + social_security_columns = ( + "social_security_retirement", + "social_security_disability", + "social_security_dependents", + "social_security_survivors", + ) + if all(column in person for column in social_security_columns): + predictors["social_security"] = np.sum( + np.column_stack([_numeric(column) for column in social_security_columns]), + axis=1, + ) + else: + predictors["social_security"] = _numeric("SS_VAL") + return predictors + + +def with_us_disability_benefits( + frame: Frame, + *, + seed: int, + time_period: int, + allow_existing_without_source: bool = False, +) -> Frame: + """Materialize direct ASEC and post-PUF-clone disability benefits.""" + + if frame.schema != US_SCHEMA: + raise ValueError("US disability benefits require the US schema.") + person = frame.table("person") + source_available = all( + column in person for column in US_DISABILITY_BENEFITS_REQUIRED_SOURCE_COLUMNS + ) + if not source_available: + if allow_existing_without_source and _surface_carries_signal(frame): + return frame + missing = [ + column + for column in US_DISABILITY_BENEFITS_REQUIRED_SOURCE_COLUMNS + if column not in person + ] + raise ValueError( + "US disability-benefits stage cannot heal a default surface without " + f"measured ASEC source column(s): {missing}." + ) + + stage_person = person.copy(deep=True) + stage_person[_PERSON_WEIGHT_COLUMN] = frame.resolve_weights("person").values + if _PERSON_SUPPORT_CHANNEL_COLUMN in person: + predictors = _person_disability_benefits_predictors(frame) + for column in _PREDICTORS: + stage_person[_PREDICTOR_PREFIX + column] = predictors[column].to_numpy() + output = run_source_stage( + us_disability_benefits_stage_spec(), + tables={"person": stage_person}, + operation_handlers={ + "derive_disability_benefits": (derive_us_disability_benefits_from_manifest), + "impute_disability_benefits_to_puf_support": ( + impute_us_disability_benefits_to_puf_support_from_manifest + ), + }, + config=SourceRuntimeConfig(seed=int(seed), target_year=int(time_period)), + ) + aligned = output.set_index("person_id").reindex(person["person_id"]) + if aligned[_OUTPUT].isna().any(): + raise ValueError( + f"US disability-benefits stage output {_OUTPUT!r} does not cover " + "every person." + ) + + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + tables["person"][_OUTPUT] = aligned[_OUTPUT].to_numpy(dtype=np.float64) + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + ) + + +def us_disability_benefits_summary(frame: Frame) -> dict[str, object]: + """Return weighted signal and validity diagnostics.""" + + person = frame.table("person") + weights = np.asarray(frame.resolve_weights("person").values, dtype=np.float64) + values = pd.to_numeric(person[_OUTPUT], errors="coerce").to_numpy(dtype=np.float64) + finite = np.isfinite(values) + positive = finite & (values > 0.0) + total_weight = float(weights.sum()) + summary: dict[str, object] = { + "positive_share": ( + float(weights[positive].sum()) / total_weight if total_weight > 0.0 else 0.0 + ), + "positive_share_band": list(_NONZERO_SHARE_BAND), + "weighted_total": float((np.nan_to_num(values) * weights).sum()), + "nonfinite": int(np.count_nonzero(~finite)), + "negative": int(np.count_nonzero(finite & (values < 0.0))), + } + if _PERSON_SUPPORT_CHANNEL_COLUMN in person: + channel = person[_PERSON_SUPPORT_CHANNEL_COLUMN].astype(str).to_numpy() + channels: dict[str, dict[str, float | int]] = {} + for name in (_BASE_ASEC_SUPPORT_CHANNEL, _PUF_TAX_DETAIL_SUPPORT_CHANNEL): + mask = channel == name + channel_weight = float(weights[mask].sum()) + channels[name] = { + "positive_rows": int(np.count_nonzero(mask & positive)), + "positive_share": ( + float(weights[mask & positive].sum()) / channel_weight + if channel_weight > 0.0 + else 0.0 + ), + "weighted_total": float( + (np.nan_to_num(values[mask]) * weights[mask]).sum() + ), + } + summary["channels"] = channels + return summary + + +def us_disability_benefits_signal_gate(frame: Frame) -> GateResult: + """Require finite, nonnegative, nondefault signal on both support halves.""" + + person = frame.table("person") + if _OUTPUT not in person: + return GateResult( + name="disability_benefits_signal", + passed=False, + failures=(f"person column missing: {_OUTPUT}.",), + details={"missing": [_OUTPUT]}, + ) + + summary = us_disability_benefits_summary(frame) + failures: list[str] = [] + if summary["nonfinite"]: + failures.append(f"{_OUTPUT}: {int(summary['nonfinite'])} nonfinite values.") + if summary["negative"]: + failures.append(f"{_OUTPUT}: {int(summary['negative'])} negative values.") + share = float(summary["positive_share"]) + low, high = summary["positive_share_band"] + if not (low <= share <= high): + failures.append( + f"{_OUTPUT}: nonzero share {share:.4f} outside plausibility band " + f"[{low}, {high}]." + ) + channels = summary.get("channels") + if isinstance(channels, dict): + for name in (_BASE_ASEC_SUPPORT_CHANNEL, _PUF_TAX_DETAIL_SUPPORT_CHANNEL): + detail = channels.get(name) + if not isinstance(detail, dict): + failures.append(f"{_OUTPUT}: missing {name} channel diagnostics.") + continue + channel_low, channel_high = _CHANNEL_NONZERO_SHARE_BANDS[name] + channel_share = float(detail["positive_share"]) + if not (channel_low <= channel_share <= channel_high): + failures.append( + f"{_OUTPUT}: {name} nonzero share {channel_share:.4f} " + "outside plausibility band " + f"[{channel_low}, {channel_high}]." + ) + return GateResult( + name="disability_benefits_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) + + +def _surface_carries_signal(frame: Frame) -> bool: + return ( + _OUTPUT in frame.table("person") + and us_disability_benefits_signal_gate(frame).passed + ) diff --git a/packages/populace-build/src/populace/build/us_runtime/domestic_production.py b/packages/populace-build/src/populace/build/us_runtime/domestic_production.py new file mode 100644 index 00000000..7a356621 --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/domestic_production.py @@ -0,0 +1,235 @@ +"""IRS PUF domestic-production deduction input for the US build. + +The retired eCPS pipeline carried ``domestic_production_ald`` directly from +IRS PUF field ``E03240``. This is the former Section 199 domestic production +activities deduction, not the later Section 199A qualified-business-income +deduction. The immutable archived derivation, export list, and versioned PUF +artifact coordinates are exposed below. + +There is no CPS analogue and no synthetic fallback. The PUF tax-detail stage +derives the source value exactly, then its weighted QRF places that source-backed +tax-unit input on the dedicated PUF support channel. The support runtime +preserves the donor's sparse positive rate rather than smearing a rare PUF +amount across the population. The versioned 2024 PUF is 2015-based and uprates +this legacy field; it is a frozen counterfactual source value, not observed +current-law 2024 deduction activity. + +PolicyEngine-US owns which years include the input in above-the-line +deductions. Populace persists the archived input even though current-law years +after 2017 exclude the former deduction from adjusted gross income. +""" + +from __future__ import annotations + +from importlib.resources import files + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import SourceStageSpec, load_source_manifest +from populace.frame import Frame + +__all__ = [ + "DOMESTIC_PRODUCTION_ALD_ARCHIVED_DERIVATION_URL", + "DOMESTIC_PRODUCTION_ALD_ARCHIVED_EXPORT_URL", + "DOMESTIC_PRODUCTION_ALD_ARCHIVED_IMPUTATION_URL", + "DOMESTIC_PRODUCTION_ALD_ARCHIVED_PUF_ARTIFACT_URL", + "US_DOMESTIC_PRODUCTION_ALD_NONCONSTANT_TAX_UNIT_COLUMNS", + "US_DOMESTIC_PRODUCTION_ALD_OUTPUT_COLUMNS", + "US_DOMESTIC_PRODUCTION_ALD_STAGE_NAME", + "derive_us_domestic_production_ald_from_puf", + "us_domestic_production_ald_signal_gate", + "us_domestic_production_ald_stage_spec", + "us_domestic_production_ald_summary", +] + +_ARCHIVED_DATA_REPOSITORY = "policyengine-" + "us-data" +_ARCHIVED_ROOT = ( + "https://github.com/PolicyEngine/" + f"{_ARCHIVED_DATA_REPOSITORY}/blob/" + "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe/" + "policyengine_" + "us_data/" +) +DOMESTIC_PRODUCTION_ALD_ARCHIVED_DERIVATION_URL = ( + _ARCHIVED_ROOT + "datasets/puf/puf.py#L646" +) +DOMESTIC_PRODUCTION_ALD_ARCHIVED_EXPORT_URL = ( + _ARCHIVED_ROOT + "datasets/puf/puf.py#L808-L815" +) +DOMESTIC_PRODUCTION_ALD_ARCHIVED_IMPUTATION_URL = ( + _ARCHIVED_ROOT + "calibration/puf_impute.py#L90-L198" +) +DOMESTIC_PRODUCTION_ALD_ARCHIVED_PUF_ARTIFACT_URL = ( + _ARCHIVED_ROOT + "datasets/puf/puf.py#L1655-L1660" +) + +US_DOMESTIC_PRODUCTION_ALD_STAGE_NAME = "puf_tax_detail" +US_DOMESTIC_PRODUCTION_ALD_OUTPUT_COLUMNS: tuple[str, ...] = ( + "domestic_production_ald", +) +US_DOMESTIC_PRODUCTION_ALD_NONCONSTANT_TAX_UNIT_COLUMNS = ( + US_DOMESTIC_PRODUCTION_ALD_OUTPUT_COLUMNS +) + +# The pinned full PUF donor has a 0.4808% weighted positive rate. The lower +# bound rejects an all-default export; the upper bound catches QRF smearing +# while allowing sampling and support-channel composition differences. +_DOMESTIC_PRODUCTION_ALD_POSITIVE_SHARE_BAND = (0.0001, 0.02) + + +def us_domestic_production_ald_stage_spec() -> SourceStageSpec: + """Load and validate the shared PUF tax-detail stage declaration.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + spec = manifest.stage_map()[US_DOMESTIC_PRODUCTION_ALD_STAGE_NAME] + missing = sorted(set(US_DOMESTIC_PRODUCTION_ALD_OUTPUT_COLUMNS) - set(spec.outputs)) + if missing: + raise ValueError( + f"{US_DOMESTIC_PRODUCTION_ALD_STAGE_NAME!r} manifest stage does not " + f"declare domestic-production output(s) {missing}." + ) + return spec + + +def derive_us_domestic_production_ald_from_puf( + puf: pd.DataFrame, + *, + source_column: str = "E03240", + output_column: str = "domestic_production_ald", +) -> pd.DataFrame: + """Carry the archived IRS PUF E03240 field without alteration. + + Invalid or negative source values fail closed. Replacing them with zeros + would invent observations and could silently erase the input this stage + exists to restore. + """ + + if source_column not in puf.columns: + raise ValueError( + "PUF domestic-production-ALD derivation requires source column " + f"{source_column!r}." + ) + values = pd.to_numeric(puf[source_column], errors="coerce") + numeric = values.to_numpy(dtype=np.float64) + nonfinite = ~np.isfinite(numeric) + if bool(nonfinite.any()): + raise ValueError( + "PUF domestic-production-ALD source column " + f"{source_column!r} contains " + f"{int(np.count_nonzero(nonfinite))} nonnumeric or nonfinite value(s)." + ) + negative = numeric < 0.0 + if bool(negative.any()): + raise ValueError( + "PUF domestic-production-ALD source column " + f"{source_column!r} contains " + f"{int(np.count_nonzero(negative))} negative value(s)." + ) + + result = puf.copy(deep=True) + result[output_column] = numeric + return result + + +def us_domestic_production_ald_summary(frame: Frame) -> dict[str, object]: + """Return weighted signal and validity diagnostics for the tax-unit input.""" + + tax_unit = frame.table("tax_unit") + values = pd.to_numeric( + tax_unit["domestic_production_ald"], errors="coerce" + ).to_numpy(dtype=np.float64) + weights = np.asarray( + frame.resolve_weights("tax_unit").values, + dtype=np.float64, + ) + finite = np.isfinite(values) + positive = finite & (values > 0.0) + total_weight = float(weights.sum()) + positive_share = ( + float(weights[positive].sum()) / total_weight if total_weight > 0.0 else 0.0 + ) + summary: dict[str, object] = { + "positive_share": positive_share, + "positive_share_band": list(_DOMESTIC_PRODUCTION_ALD_POSITIVE_SHARE_BAND), + "unweighted_total": float(np.nansum(values)), + "nonfinite": int(np.count_nonzero(~finite)), + "negative": int(np.count_nonzero(finite & (values < 0.0))), + } + support_channel = "tax_unit_support_channel" + if support_channel in tax_unit.columns: + channels: dict[str, dict[str, float | int]] = {} + for channel in tax_unit[support_channel].dropna().unique(): + channel_mask = tax_unit[support_channel].to_numpy() == channel + channel_weight = float(weights[channel_mask].sum()) + channel_positive = channel_mask & positive + channels[str(channel)] = { + "positive_rows": int(np.count_nonzero(channel_positive)), + "positive_share": ( + float(weights[channel_positive].sum()) / channel_weight + if channel_weight > 0.0 + else 0.0 + ), + "weighted_total": float( + (weights[channel_mask] * values[channel_mask]).sum() + ), + } + summary["channels"] = channels + return summary + + +def us_domestic_production_ald_signal_gate(frame: Frame) -> GateResult: + """Require a finite, nonnegative, plausibly sparse tax-unit input.""" + + tax_unit = frame.table("tax_unit") + missing = [ + column + for column in US_DOMESTIC_PRODUCTION_ALD_OUTPUT_COLUMNS + if column not in tax_unit.columns + ] + if missing: + return GateResult( + name="domestic_production_ald_signal", + passed=False, + failures=(f"tax_unit columns missing: {missing}.",), + details={"missing": missing}, + ) + + summary = us_domestic_production_ald_summary(frame) + failures: list[str] = [] + if summary["nonfinite"]: + failures.append( + f"domestic_production_ald nonfinite values: {int(summary['nonfinite'])}." + ) + if summary["negative"]: + failures.append( + f"domestic_production_ald negative values: {int(summary['negative'])}." + ) + share = float(summary["positive_share"]) + low, high = summary["positive_share_band"] + if not (low <= share <= high): + failures.append( + f"domestic-production-ALD positive share {share:.6f} outside " + f"plausibility band [{low}, {high}]." + ) + channels = summary.get("channels") + if isinstance(channels, dict): + asec = channels.get("asec") + puf = channels.get("puf_tax_detail") + if isinstance(asec, dict) and int(asec.get("positive_rows", 0)): + failures.append( + "domestic_production_ald has positive values on the ASEC support " + "channel, which has no E03240 source." + ) + if not isinstance(puf, dict) or not int(puf.get("positive_rows", 0)): + failures.append( + "domestic_production_ald has no positive PUF tax-detail support values." + ) + return GateResult( + name="domestic_production_ald_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) diff --git a/packages/populace-build/src/populace/build/us_runtime/education_inputs.py b/packages/populace-build/src/populace/build/us_runtime/education_inputs.py new file mode 100644 index 00000000..751fe261 --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/education_inputs.py @@ -0,0 +1,304 @@ +"""Education-credit inputs from IRS PUF tuition and CPS ASEC assistance. + +The retired eCPS build derived person qualified tuition as the larger of IRS +PUF fields E03230 and E87530 (falling back to E03230) and carried educational +assistance directly from CPS ASEC ``ED_VAL``. In the archived publication +path at ``9f8db56e9ed1ca25f0ed16285fcc928343d2e474``, +``calibration/puf_impute.py`` drops the reported PUF +``american_opportunity_credit`` output before +``datasets/cps/extended_cps.py::_impute_aotc_eligibility_inputs`` runs. That +method therefore takes its documented fallback: every positive-tuition person +gets the five affirmative American Opportunity Tax Credit factual inputs. The +pinned reference independently confirms the published result: qualified +tuition and all five flags have exactly the same support mask. + +This stage is intentionally factual-input-only. It does not persist computed +education credits or eligibility formulas owned by PolicyEngine-US. +""" + +from __future__ import annotations + +from importlib.resources import files + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import ( + SourceOperationSpec, + SourceStageSpec, + load_source_manifest, +) +from populace.build.source_runtime import ( + SourceRuntimeConfig, + SourceRuntimeContext, + SourceRuntimeError, + run_source_stage, +) +from populace.frame import Frame +from populace.frame.units import US_SCHEMA + +__all__ = [ + "US_AOTC_ELIGIBILITY_OUTPUT_COLUMNS", + "US_EDUCATION_INPUTS_NONCONSTANT_PERSON_COLUMNS", + "US_EDUCATION_INPUTS_OUTPUT_COLUMNS", + "US_EDUCATION_INPUTS_REQUIRED_SOURCE_COLUMNS", + "US_EDUCATION_INPUTS_STAGE_NAME", + "derive_us_education_inputs_from_manifest", + "us_education_inputs_signal_gate", + "us_education_inputs_stage_spec", + "us_education_inputs_summary", + "with_us_education_inputs", +] + +US_EDUCATION_INPUTS_STAGE_NAME = "education_inputs" + +US_AOTC_ELIGIBILITY_OUTPUT_COLUMNS: tuple[str, ...] = ( + "is_pursuing_credential_for_american_opportunity_credit", + "attends_eligible_educational_institution_for_american_opportunity_credit", + "is_enrolled_at_least_half_time_for_american_opportunity_credit", + "has_american_opportunity_credit_1098_t_or_exception", + "has_american_opportunity_credit_institution_ein", +) + +US_EDUCATION_INPUTS_OUTPUT_COLUMNS: tuple[str, ...] = ( + "qualified_tuition_expenses", + "educational_assistance", + *US_AOTC_ELIGIBILITY_OUTPUT_COLUMNS, +) + +US_EDUCATION_INPUTS_NONCONSTANT_PERSON_COLUMNS = US_EDUCATION_INPUTS_OUTPUT_COLUMNS + +US_EDUCATION_INPUTS_REQUIRED_SOURCE_COLUMNS: tuple[str, ...] = ( + "ED_VAL", + "qualified_tuition_expenses", +) + +_PERSON_WEIGHT_COLUMN = "person_weight" +_TUITION_SHARE_BAND = (0.003, 0.05) +_ASSISTANCE_SHARE_BAND = (0.005, 0.08) +_DERIVE_EDUCATION_INPUTS_PARAMETER_KEYS = frozenset() + + +def us_education_inputs_stage_spec() -> SourceStageSpec: + """Load the packaged ``education_inputs`` source-stage declaration.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + stage_map = manifest.stage_map() + if US_EDUCATION_INPUTS_STAGE_NAME not in stage_map: + raise ValueError( + f"US source manifest declares no {US_EDUCATION_INPUTS_STAGE_NAME!r} stage." + ) + spec = stage_map[US_EDUCATION_INPUTS_STAGE_NAME] + missing = sorted(set(US_EDUCATION_INPUTS_OUTPUT_COLUMNS) - set(spec.outputs)) + if missing: + raise ValueError( + f"{US_EDUCATION_INPUTS_STAGE_NAME!r} manifest stage does not declare " + f"output(s) {missing}; the runtime and manifest have drifted." + ) + return spec + + +def derive_us_education_inputs_from_manifest( + frame: pd.DataFrame | None, + operation: SourceOperationSpec, + _context: SourceRuntimeContext | None, +) -> pd.DataFrame: + """Derive the seven education inputs from a person source table.""" + + if operation.kind != "derive_education_inputs": + raise SourceRuntimeError( + "US education-input derivation received unexpected operation " + f"{operation.kind!r}." + ) + if frame is None: + raise SourceRuntimeError( + "US education-input derivation requires the person table to be read first." + ) + unexpected = sorted( + set(operation.parameters) - _DERIVE_EDUCATION_INPUTS_PARAMETER_KEYS + ) + if unexpected: + raise SourceRuntimeError( + "US education-input derivation received unsupported parameter(s): " + f"{unexpected}." + ) + missing = [ + column + for column in US_EDUCATION_INPUTS_REQUIRED_SOURCE_COLUMNS + if column not in frame.columns + ] + if missing: + raise SourceRuntimeError( + f"US education-input derivation requires source column(s): {missing}." + ) + + result = frame.copy(deep=True) + tuition = ( + pd.to_numeric(result["qualified_tuition_expenses"], errors="coerce") + .fillna(0.0) + .clip(lower=0.0) + ) + assistance = ( + pd.to_numeric(result["ED_VAL"], errors="coerce").fillna(0.0).clip(lower=0.0) + ) + aotc_student = tuition > 0.0 + result["qualified_tuition_expenses"] = tuition.to_numpy(dtype=np.float64) + result["educational_assistance"] = assistance.to_numpy(dtype=np.float64) + for column in US_AOTC_ELIGIBILITY_OUTPUT_COLUMNS: + result[column] = aotc_student.to_numpy(dtype=bool) + return result + + +def with_us_education_inputs( + frame: Frame, + *, + seed: int, + time_period: int, +) -> Frame: + """Materialize education inputs on a US frame, healing default surfaces.""" + + if frame.schema != US_SCHEMA: + raise ValueError("US education inputs require the US schema.") + if _education_surface_carries_signal(frame): + return frame + + person = frame.table("person") + stage_person = person.copy(deep=True) + stage_person[_PERSON_WEIGHT_COLUMN] = frame.resolve_weights("person").values + output = run_source_stage( + us_education_inputs_stage_spec(), + tables={"person": stage_person}, + operation_handlers={ + "derive_education_inputs": derive_us_education_inputs_from_manifest, + }, + config=SourceRuntimeConfig(seed=int(seed), target_year=int(time_period)), + ) + aligned = output.set_index("person_id").reindex(person["person_id"]) + for column in US_EDUCATION_INPUTS_OUTPUT_COLUMNS: + if aligned[column].isna().any(): + raise ValueError( + "US education-input stage output does not cover every person for " + f"{column!r}." + ) + + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + for column in US_EDUCATION_INPUTS_OUTPUT_COLUMNS: + if column in US_AOTC_ELIGIBILITY_OUTPUT_COLUMNS: + tables["person"][column] = aligned[column].to_numpy(dtype=bool) + else: + tables["person"][column] = aligned[column].to_numpy(dtype=np.float64) + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + ) + + +def us_education_inputs_summary(frame: Frame) -> dict[str, object]: + """Return weighted signal and coherence diagnostics for education inputs.""" + + person = frame.table("person") + weights = np.asarray(frame.resolve_weights("person").values, dtype=np.float64) + total_weight = float(weights.sum()) + tuition = pd.to_numeric( + person["qualified_tuition_expenses"], errors="coerce" + ).to_numpy(dtype=np.float64) + assistance = pd.to_numeric( + person["educational_assistance"], errors="coerce" + ).to_numpy(dtype=np.float64) + tuition_positive = np.isfinite(tuition) & (tuition > 0.0) + assistance_positive = np.isfinite(assistance) & (assistance > 0.0) + + def _share(mask: np.ndarray) -> float: + return float(weights[mask].sum()) / total_weight if total_weight > 0 else 0.0 + + incoherent: dict[str, int] = {} + for column in US_AOTC_ELIGIBILITY_OUTPUT_COLUMNS: + values = person[column].fillna(False).astype(bool).to_numpy() + incoherent[column] = int(np.count_nonzero(values != tuition_positive)) + return { + "qualified_tuition_share": _share(tuition_positive), + "educational_assistance_share": _share(assistance_positive), + "qualified_tuition_total": float(np.nansum(tuition)), + "educational_assistance_total": float(np.nansum(assistance)), + "tuition_share_band": list(_TUITION_SHARE_BAND), + "assistance_share_band": list(_ASSISTANCE_SHARE_BAND), + "nonfinite_tuition": int(np.count_nonzero(~np.isfinite(tuition))), + "nonfinite_assistance": int(np.count_nonzero(~np.isfinite(assistance))), + "negative_tuition": int(np.count_nonzero(tuition < 0.0)), + "negative_assistance": int(np.count_nonzero(assistance < 0.0)), + "incoherent_aotc_flags": incoherent, + } + + +def us_education_inputs_signal_gate(frame: Frame) -> GateResult: + """Require nonzero, plausible, coherent education-input signal.""" + + person = frame.table("person") + missing = [ + column + for column in US_EDUCATION_INPUTS_OUTPUT_COLUMNS + if column not in person.columns + ] + if missing: + return GateResult( + name="education_inputs_signal", + passed=False, + failures=(f"person columns missing: {missing}.",), + details={"missing": missing}, + ) + + summary = us_education_inputs_summary(frame) + failures: list[str] = [] + for count_key, label in ( + ("nonfinite_tuition", "qualified_tuition_expenses nonfinite values"), + ("nonfinite_assistance", "educational_assistance nonfinite values"), + ("negative_tuition", "qualified_tuition_expenses negative values"), + ("negative_assistance", "educational_assistance negative values"), + ): + count = int(summary[count_key]) + if count: + failures.append(f"{label}: {count}.") + + for share_key, band_key, label in ( + ( + "qualified_tuition_share", + "tuition_share_band", + "qualified-tuition share", + ), + ( + "educational_assistance_share", + "assistance_share_band", + "educational-assistance share", + ), + ): + share = float(summary[share_key]) + low, high = summary[band_key] + if not (low <= share <= high): + failures.append( + f"{label} {share:.3f} outside plausibility band [{low}, {high}]." + ) + + for column, count in summary["incoherent_aotc_flags"].items(): + if count: + failures.append( + f"{column}: {count} rows disagree with positive qualified tuition." + ) + return GateResult( + name="education_inputs_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) + + +def _education_surface_carries_signal(frame: Frame) -> bool: + person = frame.table("person") + if not all(column in person for column in US_EDUCATION_INPUTS_OUTPUT_COLUMNS): + return False + return us_education_inputs_signal_gate(frame).passed diff --git a/packages/populace-build/src/populace/build/us_runtime/educator_expenses.py b/packages/populace-build/src/populace/build/us_runtime/educator_expenses.py new file mode 100644 index 00000000..4e5aec59 --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/educator_expenses.py @@ -0,0 +1,212 @@ +"""IRS PUF educator-expense input for the US build. + +The retired eCPS pipeline carried ``educator_expense`` directly from IRS PUF +field ``E03220`` and allocated the tax-unit amount between filer and spouse by +earnings share. Immutable archived coordinates for the derivation, allocation, +export, and PUF-QRF treatment are exposed below. + +There is no CPS analogue and no synthetic fallback. Populace's established +PUF-detail topology therefore carries the processed, source-backed leaf into +the shared weighted QRF, places it only on the dedicated PUF support channel, +and sparsifies it to the donor's weighted positive rate. This differs openly +from the retired pipeline's both-half override, while preserving population +mass and the build's source-channel contract. PolicyEngine-US owns the +above-the-line deduction formula; this module persists only the factual input. +""" + +from __future__ import annotations + +from importlib.resources import files + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import SourceStageSpec, load_source_manifest +from populace.frame import Frame + +__all__ = [ + "EDUCATOR_EXPENSE_ARCHIVED_ALLOCATION_URL", + "EDUCATOR_EXPENSE_ARCHIVED_DERIVATION_URL", + "EDUCATOR_EXPENSE_ARCHIVED_EXPORT_URL", + "EDUCATOR_EXPENSE_ARCHIVED_PUF_IMPUTATION_URL", + "US_EDUCATOR_EXPENSE_NONCONSTANT_PERSON_COLUMNS", + "US_EDUCATOR_EXPENSE_OUTPUT_COLUMNS", + "US_EDUCATOR_EXPENSE_STAGE_NAME", + "derive_us_educator_expense_from_puf", + "us_educator_expense_signal_gate", + "us_educator_expense_stage_spec", + "us_educator_expense_summary", +] + +_ARCHIVED_DATA_REPOSITORY = "policyengine-" + "us-data" +_ARCHIVED_ROOT = ( + "https://github.com/PolicyEngine/" + f"{_ARCHIVED_DATA_REPOSITORY}/blob/" + "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe/" + "policyengine_" + "us_data/" +) +EDUCATOR_EXPENSE_ARCHIVED_DERIVATION_URL = ( + _ARCHIVED_ROOT + "datasets/puf/puf.py#L636-L649" +) +EDUCATOR_EXPENSE_ARCHIVED_ALLOCATION_URL = ( + _ARCHIVED_ROOT + "datasets/puf/puf.py#L617-L620" +) +EDUCATOR_EXPENSE_ARCHIVED_EXPORT_URL = _ARCHIVED_ROOT + "datasets/puf/puf.py#L804-L815" +EDUCATOR_EXPENSE_ARCHIVED_PUF_IMPUTATION_URL = ( + _ARCHIVED_ROOT + "calibration/puf_impute.py#L940-L1075" +) + +US_EDUCATOR_EXPENSE_STAGE_NAME = "puf_tax_detail" +US_EDUCATOR_EXPENSE_OUTPUT_COLUMNS: tuple[str, ...] = ("educator_expense",) +US_EDUCATOR_EXPENSE_NONCONSTANT_PERSON_COLUMNS = US_EDUCATOR_EXPENSE_OUTPUT_COLUMNS + +_OUTPUT = US_EDUCATOR_EXPENSE_OUTPUT_COLUMNS[0] +_PERSON_SUPPORT_CHANNEL_COLUMN = "person_support_channel" +_BASE_ASEC_SUPPORT_CHANNEL = "asec" +_PUF_TAX_DETAIL_SUPPORT_CHANNEL = "puf_tax_detail" +_OVERALL_POSITIVE_SHARE_BAND = (0.002, 0.04) +_PUF_POSITIVE_SHARE_BAND = (0.005, 0.08) + + +def us_educator_expense_stage_spec() -> SourceStageSpec: + """Load and validate the shared PUF tax-detail stage declaration.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + spec = manifest.stage_map()[US_EDUCATOR_EXPENSE_STAGE_NAME] + missing = sorted(set(US_EDUCATOR_EXPENSE_OUTPUT_COLUMNS) - set(spec.outputs)) + if missing: + raise ValueError( + f"{US_EDUCATOR_EXPENSE_STAGE_NAME!r} manifest stage does not declare " + f"educator-expense output(s) {missing}." + ) + return spec + + +def derive_us_educator_expense_from_puf( + puf: pd.DataFrame, + *, + source_column: str = "E03220", + output_column: str = _OUTPUT, +) -> pd.DataFrame: + """Carry the archived IRS PUF educator-expense field exactly.""" + + if source_column not in puf.columns: + raise ValueError( + f"PUF educator-expense derivation requires source column {source_column!r}." + ) + values = pd.to_numeric(puf[source_column], errors="coerce").to_numpy( + dtype=np.float64 + ) + nonfinite = ~np.isfinite(values) + if bool(nonfinite.any()): + raise ValueError( + f"PUF educator-expense source column {source_column!r} contains " + f"{int(np.count_nonzero(nonfinite))} nonnumeric or nonfinite value(s)." + ) + negative = values < 0.0 + if bool(negative.any()): + raise ValueError( + f"PUF educator-expense source column {source_column!r} contains " + f"{int(np.count_nonzero(negative))} negative value(s)." + ) + + result = puf.copy(deep=True) + result[output_column] = values + return result + + +def us_educator_expense_summary(frame: Frame) -> dict[str, object]: + """Return weighted signal, channel, and validity diagnostics.""" + + person = frame.table("person") + values = pd.to_numeric(person[_OUTPUT], errors="coerce").to_numpy(dtype=np.float64) + weights = np.asarray(frame.resolve_weights("person").values, dtype=np.float64) + finite = np.isfinite(values) + positive = finite & (values > 0.0) + total_weight = float(weights.sum()) + summary: dict[str, object] = { + "positive_share": ( + float(weights[positive].sum()) / total_weight if total_weight > 0.0 else 0.0 + ), + "positive_share_band": list(_OVERALL_POSITIVE_SHARE_BAND), + "weighted_total": float((np.nan_to_num(values) * weights).sum()), + "nonfinite": int(np.count_nonzero(~finite)), + "negative": int(np.count_nonzero(finite & (values < 0.0))), + } + if _PERSON_SUPPORT_CHANNEL_COLUMN in person: + channel = person[_PERSON_SUPPORT_CHANNEL_COLUMN].astype(str).to_numpy() + channels: dict[str, dict[str, float | int]] = {} + for name in (_BASE_ASEC_SUPPORT_CHANNEL, _PUF_TAX_DETAIL_SUPPORT_CHANNEL): + mask = channel == name + channel_weight = float(weights[mask].sum()) + channels[name] = { + "positive_rows": int(np.count_nonzero(mask & positive)), + "positive_share": ( + float(weights[mask & positive].sum()) / channel_weight + if channel_weight > 0.0 + else 0.0 + ), + "weighted_total": float( + (np.nan_to_num(values[mask]) * weights[mask]).sum() + ), + } + summary["channels"] = channels + return summary + + +def us_educator_expense_signal_gate(frame: Frame) -> GateResult: + """Require source-backed, sparse PUF signal and a zero ASEC channel.""" + + person = frame.table("person") + if _OUTPUT not in person: + return GateResult( + name="educator_expense_signal", + passed=False, + failures=(f"person column missing: {_OUTPUT}.",), + details={"missing": [_OUTPUT]}, + ) + + summary = us_educator_expense_summary(frame) + failures: list[str] = [] + if summary["nonfinite"]: + failures.append(f"{_OUTPUT}: {int(summary['nonfinite'])} nonfinite values.") + if summary["negative"]: + failures.append(f"{_OUTPUT}: {int(summary['negative'])} negative values.") + share = float(summary["positive_share"]) + low, high = summary["positive_share_band"] + if not (low <= share <= high): + failures.append( + f"{_OUTPUT}: positive share {share:.4f} outside plausibility band " + f"[{low}, {high}]." + ) + + channels = summary.get("channels") + if isinstance(channels, dict): + asec = channels.get(_BASE_ASEC_SUPPORT_CHANNEL) + puf = channels.get(_PUF_TAX_DETAIL_SUPPORT_CHANNEL) + if not isinstance(asec, dict) or not isinstance(puf, dict): + failures.append(f"{_OUTPUT}: missing support-channel diagnostics.") + else: + asec_share = float(asec["positive_share"]) + if asec_share != 0.0: + failures.append( + f"{_OUTPUT}: ASEC support channel must remain source-zero; " + f"positive share is {asec_share:.4f}." + ) + puf_share = float(puf["positive_share"]) + puf_low, puf_high = _PUF_POSITIVE_SHARE_BAND + if not (puf_low <= puf_share <= puf_high): + failures.append( + f"{_OUTPUT}: PUF support-channel positive share " + f"{puf_share:.4f} outside plausibility band " + f"[{puf_low}, {puf_high}]." + ) + return GateResult( + name="educator_expense_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) diff --git a/packages/populace-build/src/populace/build/us_runtime/energy_subsidy.py b/packages/populace-build/src/populace/build/us_runtime/energy_subsidy.py new file mode 100644 index 00000000..c70167c4 --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/energy_subsidy.py @@ -0,0 +1,659 @@ +"""ASEC-measured energy subsidies with retired PUF-half QRF treatment. + +The retired eCPS pipeline carried measured annual SPM-unit energy subsidies +from ASEC ``SPM_ENGVAL`` and then QRF-imputed that CPS-only input onto the PUF +clone half. The QRF used eight person predictors, at most 5,000 ASEC training +rows, and reduced non-person outputs with ``value_from_first_person``. +Immutable archived source coordinates are exposed below. + +The ASEC value is replicated on every member of an SPM unit. This port refuses +missing, nonfinite, negative, or internally inconsistent replicas rather than +silently choosing one or filling zeros. It also uses typed person weights for +the QRF fit: the archived call omitted weights, but weighted fitting is a hard +Populace build contract. PolicyEngine-US owns the downstream SPM-resource and +poverty formulas; this stage persists only the factual SPM-unit input. +""" + +from __future__ import annotations + +from importlib.resources import files +from typing import Any + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import ( + SourceOperationSpec, + SourceStageSpec, + load_source_manifest, +) +from populace.build.source_runtime import ( + SourceRuntimeConfig, + SourceRuntimeContext, + SourceRuntimeError, + run_source_stage, +) +from populace.frame import Frame +from populace.frame.units import US_SCHEMA + +__all__ = [ + "ENERGY_SUBSIDY_ARCHIVED_CPS_DERIVATION_URL", + "ENERGY_SUBSIDY_ARCHIVED_PUF_IMPUTATION_URL", + "US_ENERGY_SUBSIDY_OUTPUT_COLUMNS", + "US_ENERGY_SUBSIDY_REQUIRED_SOURCE_COLUMNS", + "US_ENERGY_SUBSIDY_STAGE_NAME", + "derive_us_energy_subsidy_from_manifest", + "impute_us_energy_subsidy_to_puf_support_from_manifest", + "us_energy_subsidy_signal_gate", + "us_energy_subsidy_stage_spec", + "us_energy_subsidy_summary", + "with_us_energy_subsidy_input", +] + +QRF: Any | None = None + +_ARCHIVED_DATA_REPOSITORY = "policyengine-" + "us-data" +_ARCHIVED_ROOT = ( + "https://github.com/PolicyEngine/" + f"{_ARCHIVED_DATA_REPOSITORY}/blob/" + "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe/" + "policyengine_" + "us_data/" +) +ENERGY_SUBSIDY_ARCHIVED_CPS_DERIVATION_URL = ( + _ARCHIVED_ROOT + "datasets/cps/cps.py#L1612-L1622" +) +ENERGY_SUBSIDY_ARCHIVED_PUF_IMPUTATION_URL = ( + _ARCHIVED_ROOT + "datasets/cps/extended_cps.py#L639-L739" +) + +US_ENERGY_SUBSIDY_STAGE_NAME = "energy_subsidy" +US_ENERGY_SUBSIDY_OUTPUT_COLUMNS: tuple[str, ...] = ("spm_unit_energy_subsidy",) +US_ENERGY_SUBSIDY_REQUIRED_SOURCE_COLUMNS: tuple[str, ...] = ( + "person_spm_unit_id", + "SPM_ENGVAL", +) + +_OUTPUT = US_ENERGY_SUBSIDY_OUTPUT_COLUMNS[0] +_PERSON_WEIGHT_COLUMN = "person_weight" +_PERSON_SUPPORT_CHANNEL_COLUMN = "person_support_channel" +_SPM_UNIT_SUPPORT_CHANNEL_COLUMN = "spm_unit_support_channel" +_BASE_ASEC_SUPPORT_CHANNEL = "asec" +_PUF_TAX_DETAIL_SUPPORT_CHANNEL = "puf_tax_detail" +_PUF_PREDICTORS: tuple[str, ...] = ( + "age", + "is_male", + "has_esi", + "tax_unit_is_joint", + "tax_unit_count_dependents", + "employment_income", + "self_employment_income", + "social_security", +) +_PUF_PREDICTOR_PREFIX = "energy_subsidy_predictor_" +_DIRECT_PARAMETER_KEYS = frozenset() +_PUF_IMPUTATION_PARAMETER_KEYS = frozenset( + { + "predictors", + "max_train_samples", + "n_estimators", + "seed_from_build_config", + "weight", + "reduction", + } +) +# The three locked ASEC vintages carry a weighted SPM-unit positive share of +# 3.26%-3.34%. The PUF QRF bound is deliberately wider because it is a modelled +# support channel; the overall band rejects both a dead and a near-universal +# surface while accommodating small test frames and vintage drift. +_OVERALL_POSITIVE_SHARE_BAND = (0.01, 0.20) +_CHANNEL_POSITIVE_SHARE_BANDS = { + _BASE_ASEC_SUPPORT_CHANNEL: (0.025, 0.045), + _PUF_TAX_DETAIL_SUPPORT_CHANNEL: (0.002, 0.20), +} +_REPLICA_ATOL = 1e-6 +_EXPECTED_OPERATION_KINDS = ( + "read_table", + "derive_energy_subsidy", + "impute_energy_subsidy_to_puf_support", +) + + +def us_energy_subsidy_stage_spec() -> SourceStageSpec: + """Load and strictly validate the packaged energy-subsidy stage.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + stage_map = manifest.stage_map() + if US_ENERGY_SUBSIDY_STAGE_NAME not in stage_map: + raise ValueError( + f"US source manifest declares no {US_ENERGY_SUBSIDY_STAGE_NAME!r} stage." + ) + spec = stage_map[US_ENERGY_SUBSIDY_STAGE_NAME] + operation_kinds = tuple(operation.kind for operation in spec.operations) + failures: list[str] = [] + if spec.stage != US_ENERGY_SUBSIDY_STAGE_NAME: + failures.append(f"stage={spec.stage!r}") + if spec.grain != "person": + failures.append(f"grain={spec.grain!r}") + if tuple(spec.outputs) != US_ENERGY_SUBSIDY_OUTPUT_COLUMNS: + failures.append(f"outputs={list(spec.outputs)!r}") + if tuple(spec.nonnegative_outputs) != US_ENERGY_SUBSIDY_OUTPUT_COLUMNS: + failures.append(f"nonnegative_outputs={list(spec.nonnegative_outputs)!r}") + if operation_kinds != _EXPECTED_OPERATION_KINDS: + failures.append(f"operations={list(operation_kinds)!r}") + if failures: + raise ValueError( + f"{US_ENERGY_SUBSIDY_STAGE_NAME!r} manifest contract drifted: " + + "; ".join(failures) + + "." + ) + return spec + + +def _numeric_source(frame: pd.DataFrame, column: str) -> np.ndarray: + values = pd.to_numeric(frame[column], errors="coerce").to_numpy(dtype=np.float64) + nonfinite = int(np.count_nonzero(~np.isfinite(values))) + if nonfinite: + raise SourceRuntimeError( + f"US energy-subsidy source {column!r} contains {nonfinite} " + "nonfinite value(s); the measured source must not be silently " + "replaced." + ) + return values + + +def derive_us_energy_subsidy_from_manifest( + frame: pd.DataFrame | None, + operation: SourceOperationSpec, + _context: SourceRuntimeContext | None, +) -> pd.DataFrame: + """Carry exact measured, replicated ASEC energy-subsidy values.""" + + if operation.kind != "derive_energy_subsidy": + raise SourceRuntimeError( + "US energy-subsidy derivation received unexpected operation " + f"{operation.kind!r}." + ) + if frame is None: + raise SourceRuntimeError( + "US energy-subsidy derivation requires the person table to be read first." + ) + unexpected = sorted(set(operation.parameters) - _DIRECT_PARAMETER_KEYS) + if unexpected: + raise SourceRuntimeError( + "US energy-subsidy derivation received unsupported parameter(s): " + f"{unexpected}." + ) + missing = [ + column + for column in US_ENERGY_SUBSIDY_REQUIRED_SOURCE_COLUMNS + if column not in frame.columns + ] + if missing: + raise SourceRuntimeError( + "US energy-subsidy derivation requires measured ASEC source " + f"column(s): {missing}." + ) + + values = _numeric_source(frame, "SPM_ENGVAL") + negative = int(np.count_nonzero(values < 0.0)) + if negative: + raise SourceRuntimeError( + "US energy-subsidy source 'SPM_ENGVAL' contains " + f"{negative} negative value(s)." + ) + spm_ids = pd.to_numeric(frame["person_spm_unit_id"], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not np.isfinite(spm_ids).all(): + raise SourceRuntimeError( + "US energy-subsidy source person_spm_unit_id contains nonfinite values." + ) + replicas = pd.DataFrame({"spm_unit_id": spm_ids, "value": values}) + bounds = replicas.groupby("spm_unit_id", sort=False)["value"].agg(["min", "max"]) + inconsistent = ~np.isclose( + bounds["min"].to_numpy(dtype=np.float64), + bounds["max"].to_numpy(dtype=np.float64), + rtol=0.0, + atol=_REPLICA_ATOL, + ) + if bool(inconsistent.any()): + bad_ids = bounds.index[inconsistent].astype("int64").tolist() + raise SourceRuntimeError( + "US energy-subsidy source SPM_ENGVAL disagrees within replicated " + f"SPM unit(s): {bad_ids[:10]}." + ) + + result = frame.copy(deep=True) + result[_OUTPUT] = values + return result + + +def impute_us_energy_subsidy_to_puf_support_from_manifest( + frame: pd.DataFrame | None, + operation: SourceOperationSpec, + context: SourceRuntimeContext | None, +) -> pd.DataFrame: + """QRF-impute energy subsidies onto PUF people before unit reduction.""" + + if operation.kind != "impute_energy_subsidy_to_puf_support": + raise SourceRuntimeError( + "US energy-subsidy PUF imputation received unexpected operation " + f"{operation.kind!r}." + ) + if frame is None: + raise SourceRuntimeError( + "US energy-subsidy PUF imputation requires the person table to be " + "read first." + ) + unexpected = sorted(set(operation.parameters) - _PUF_IMPUTATION_PARAMETER_KEYS) + missing_parameters = sorted( + _PUF_IMPUTATION_PARAMETER_KEYS - set(operation.parameters) + ) + if unexpected or missing_parameters: + raise SourceRuntimeError( + "US energy-subsidy PUF imputation parameters must match the " + "archived method; " + f"missing={missing_parameters}, unexpected={unexpected}." + ) + if _PERSON_SUPPORT_CHANNEL_COLUMN not in frame.columns: + return frame.copy(deep=True) + + predictors = tuple(str(value) for value in operation.parameters["predictors"]) + if predictors != _PUF_PREDICTORS: + raise SourceRuntimeError( + "US energy-subsidy PUF predictors drifted from the archived " + f"method: expected {list(_PUF_PREDICTORS)}, got {list(predictors)}." + ) + if operation.parameters["weight"] != _PERSON_WEIGHT_COLUMN: + raise SourceRuntimeError( + "US energy-subsidy PUF imputation must use the typed " + f"person-weight column {_PERSON_WEIGHT_COLUMN!r}." + ) + if operation.parameters["seed_from_build_config"] is not True: + raise SourceRuntimeError( + "US energy-subsidy PUF imputation seed must come from the build config." + ) + if operation.parameters["reduction"] != "value_from_first_person": + raise SourceRuntimeError( + "US energy-subsidy PUF imputation must use the archived " + "value_from_first_person reduction." + ) + max_train_samples = int(operation.parameters["max_train_samples"]) + n_estimators = int(operation.parameters["n_estimators"]) + if max_train_samples <= 0 or n_estimators <= 0: + raise SourceRuntimeError( + "US energy-subsidy PUF max_train_samples and n_estimators must be positive." + ) + + predictor_columns = [_PUF_PREDICTOR_PREFIX + name for name in predictors] + required = [_PERSON_WEIGHT_COLUMN, *predictor_columns, _OUTPUT] + missing = [column for column in required if column not in frame.columns] + if missing: + raise SourceRuntimeError( + f"US energy-subsidy PUF imputation is missing source column(s): {missing}." + ) + + channel = frame[_PERSON_SUPPORT_CHANNEL_COLUMN].astype(str) + asec_mask = channel == _BASE_ASEC_SUPPORT_CHANNEL + puf_mask = channel == _PUF_TAX_DETAIL_SUPPORT_CHANNEL + if not asec_mask.any() or not puf_mask.any(): + raise SourceRuntimeError( + "US energy-subsidy PUF imputation requires nonempty ASEC and " + "PUF-tax-detail support channels." + ) + + training = frame.loc[asec_mask, [*predictor_columns, _OUTPUT]].copy() + training.columns = [*predictors, _OUTPUT] + test = frame.loc[puf_mask, predictor_columns].copy() + test.columns = list(predictors) + weights = pd.to_numeric( + frame.loc[asec_mask, _PERSON_WEIGHT_COLUMN], errors="coerce" + ) + numeric_weights = weights.to_numpy(dtype=np.float64) + if not np.isfinite(numeric_weights).all() or bool((numeric_weights < 0.0).any()): + raise SourceRuntimeError( + "US energy-subsidy QRF person weights must be finite and nonnegative." + ) + if float(numeric_weights.sum()) <= 0.0: + raise SourceRuntimeError("US energy-subsidy QRF person weights sum to zero.") + + if len(training) > max_train_samples: + sample = training.sample( + n=max_train_samples, + random_state=(context.config.seed if context is not None else 0), + ).index + training = training.loc[sample] + weights = weights.loc[sample] + + for column in (*predictors, _OUTPUT): + training[column] = pd.to_numeric(training[column], errors="coerce") + if not np.isfinite(training[column].to_numpy(dtype=np.float64)).all(): + raise SourceRuntimeError( + f"US energy-subsidy QRF training column {column!r} contains " + "nonfinite values." + ) + for column in predictors: + test[column] = pd.to_numeric(test[column], errors="coerce") + if not np.isfinite(test[column].to_numpy(dtype=np.float64)).all(): + raise SourceRuntimeError( + f"US energy-subsidy QRF prediction column {column!r} contains " + "nonfinite values." + ) + + global QRF + if QRF is None: + from importlib import import_module + + QRF = import_module("populace.fit").QRF + seed = context.config.seed if context is not None else 0 + fitted = QRF(n_estimators=n_estimators, seed=seed).fit( + training, + list(predictors), + [_OUTPUT], + weights=weights.to_numpy(dtype=np.float64), + ) + predictions = fitted.predict(test) + if _OUTPUT not in predictions: + raise SourceRuntimeError( + f"US energy-subsidy QRF prediction is missing output {_OUTPUT!r}." + ) + predicted = pd.to_numeric(predictions[_OUTPUT], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not np.isfinite(predicted).all(): + raise SourceRuntimeError( + "US energy-subsidy QRF produced nonfinite predictions." + ) + if bool((predicted < 0.0).any()): + raise SourceRuntimeError("US energy-subsidy QRF produced negative predictions.") + + result = frame.copy(deep=True) + result.loc[puf_mask, _OUTPUT] = predicted + return result + + +def _person_energy_subsidy_predictors(frame: Frame) -> pd.DataFrame: + """Build the retired eight QRF predictors on person rows.""" + + person = frame.table("person") + tax_unit = frame.table("tax_unit") + predictors = pd.DataFrame(index=person.index) + + def _numeric(*columns: str) -> np.ndarray: + for column in columns: + if column in person: + values = pd.to_numeric(person[column], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not np.isfinite(values).all(): + raise SourceRuntimeError( + "US energy-subsidy QRF predictor source " + f"{column!r} contains nonfinite values." + ) + return values + raise SourceRuntimeError( + "US energy-subsidy PUF imputation cannot construct predictor from " + f"any of {list(columns)}." + ) + + predictors["age"] = _numeric("age", "A_AGE") + if "is_male" in person: + predictors["is_male"] = person["is_male"].fillna(False).astype(bool) + elif "is_female" in person: + predictors["is_male"] = ~person["is_female"].fillna(False).astype(bool) + elif "A_SEX" in person: + predictors["is_male"] = _numeric("A_SEX") == 1 + else: + raise SourceRuntimeError( + "US energy-subsidy PUF imputation requires is_male, is_female, or " + "measured A_SEX." + ) + if "has_esi" not in person: + raise SourceRuntimeError("US energy-subsidy PUF imputation requires has_esi.") + predictors["has_esi"] = person["has_esi"].fillna(False).astype(bool) + + filing_status_column = next( + ( + column + for column in ("filing_status_input", "filing_status") + if column in tax_unit + ), + None, + ) + if filing_status_column is None: + raise SourceRuntimeError( + "US energy-subsidy PUF imputation requires tax-unit filing status." + ) + filing_status = tax_unit[filing_status_column].map( + lambda value: ( + value.decode() if isinstance(value, (bytes, np.bytes_)) else str(value) + ) + ) + joint_by_id = pd.Series( + filing_status.eq("JOINT").to_numpy(), + index=tax_unit["tax_unit_id"].to_numpy(), + ) + predictors["tax_unit_is_joint"] = ( + person["person_tax_unit_id"].map(joint_by_id).fillna(False).astype(bool) + ) + + if "tax_unit_role_input" not in person: + raise SourceRuntimeError( + "US energy-subsidy PUF imputation requires tax_unit_role_input to " + "count dependents." + ) + dependent = ( + person["tax_unit_role_input"] + .map( + lambda value: ( + value.decode() if isinstance(value, (bytes, np.bytes_)) else str(value) + ) + ) + .eq("DEPENDENT") + ) + predictors["tax_unit_count_dependents"] = ( + dependent.groupby(person["person_tax_unit_id"]) + .transform("sum") + .to_numpy(dtype=np.float64) + ) + predictors["employment_income"] = _numeric( + "employment_income_before_lsr", "WSAL_VAL" + ) + predictors["self_employment_income"] = _numeric( + "self_employment_income_before_lsr", "SEMP_VAL" + ) + social_security_columns = ( + "social_security_retirement", + "social_security_disability", + "social_security_dependents", + "social_security_survivors", + ) + if all(column in person for column in social_security_columns): + predictors["social_security"] = np.sum( + np.column_stack([_numeric(column) for column in social_security_columns]), + axis=1, + ) + else: + predictors["social_security"] = _numeric("SS_VAL") + return predictors + + +def with_us_energy_subsidy_input( + frame: Frame, + *, + seed: int, + time_period: int, + allow_existing_without_source: bool = False, +) -> Frame: + """Materialize measured/QRF energy subsidies on the SPM-unit leaf. + + ``allow_existing_without_source`` is only for a downstream release build + consuming a base artifact that already passed this stage's signal gate. + Source/base construction remains strict and cannot reuse clone duplicates. + """ + + if frame.schema != US_SCHEMA: + raise ValueError("US energy-subsidy input requires the US schema.") + person = frame.table("person") + spm_unit = frame.table("spm_unit") + source_available = all( + column in person for column in US_ENERGY_SUBSIDY_REQUIRED_SOURCE_COLUMNS + ) + if not source_available: + if allow_existing_without_source and _surface_carries_signal(frame): + return frame + missing = [ + column + for column in US_ENERGY_SUBSIDY_REQUIRED_SOURCE_COLUMNS + if column not in person + ] + raise ValueError( + "US energy-subsidy stage cannot heal a default surface without " + f"measured ASEC source column(s): {missing}." + ) + + stage_person = person.copy(deep=True) + stage_person[_PERSON_WEIGHT_COLUMN] = frame.resolve_weights("person").values + if _PERSON_SUPPORT_CHANNEL_COLUMN in person: + predictors = _person_energy_subsidy_predictors(frame) + for column in _PUF_PREDICTORS: + stage_person[_PUF_PREDICTOR_PREFIX + column] = predictors[column].to_numpy() + output = run_source_stage( + us_energy_subsidy_stage_spec(), + tables={"person": stage_person}, + operation_handlers={ + "derive_energy_subsidy": derive_us_energy_subsidy_from_manifest, + "impute_energy_subsidy_to_puf_support": ( + impute_us_energy_subsidy_to_puf_support_from_manifest + ), + }, + config=SourceRuntimeConfig(seed=int(seed), target_year=int(time_period)), + ) + aligned_people = output.set_index("person_id").reindex(person["person_id"]) + if aligned_people[_OUTPUT].isna().any(): + raise ValueError( + "US energy-subsidy stage output does not cover every person before " + "SPM-unit reduction." + ) + unit_values = ( + aligned_people.assign( + person_spm_unit_id=person["person_spm_unit_id"].to_numpy() + ) + .groupby("person_spm_unit_id", sort=False)[_OUTPUT] + .first() + ) + aligned_units = unit_values.reindex(spm_unit["spm_unit_id"]) + if aligned_units.isna().any(): + raise ValueError( + "US energy-subsidy stage output does not cover every SPM unit." + ) + + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + tables["spm_unit"][_OUTPUT] = aligned_units.to_numpy(dtype=np.float64) + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + ) + + +def us_energy_subsidy_summary(frame: Frame) -> dict[str, object]: + """Return weighted signal and validity diagnostics by support channel.""" + + spm_unit = frame.table("spm_unit") + values = pd.to_numeric(spm_unit[_OUTPUT], errors="coerce").to_numpy( + dtype=np.float64 + ) + weights = np.asarray(frame.resolve_weights("spm_unit").values, dtype=np.float64) + finite = np.isfinite(values) + positive = finite & (values > 0.0) + total_weight = float(weights.sum()) + summary: dict[str, object] = { + "positive_share": ( + float(weights[positive].sum()) / total_weight if total_weight > 0.0 else 0.0 + ), + "positive_share_band": list(_OVERALL_POSITIVE_SHARE_BAND), + "weighted_total": float(np.sum(np.nan_to_num(values) * weights)), + "nonfinite": int(np.count_nonzero(~finite)), + "negative": int(np.count_nonzero(finite & (values < 0.0))), + } + if _SPM_UNIT_SUPPORT_CHANNEL_COLUMN in spm_unit: + channel = spm_unit[_SPM_UNIT_SUPPORT_CHANNEL_COLUMN].astype(str).to_numpy() + channels: dict[str, dict[str, float | int | list[float]]] = {} + for name in (_BASE_ASEC_SUPPORT_CHANNEL, _PUF_TAX_DETAIL_SUPPORT_CHANNEL): + mask = channel == name + channel_weight = float(weights[mask].sum()) + channels[name] = { + "positive_rows": int(np.count_nonzero(mask & positive)), + "positive_share": ( + float(weights[mask & positive].sum()) / channel_weight + if channel_weight > 0.0 + else 0.0 + ), + "positive_share_band": list(_CHANNEL_POSITIVE_SHARE_BANDS[name]), + "weighted_total": float( + np.sum(np.nan_to_num(values[mask]) * weights[mask]) + ), + } + summary["channels"] = channels + return summary + + +def us_energy_subsidy_signal_gate(frame: Frame) -> GateResult: + """Require finite, nonnegative, plausible signal on both support halves.""" + + spm_unit = frame.table("spm_unit") + if _OUTPUT not in spm_unit: + return GateResult( + name="energy_subsidy_signal", + passed=False, + failures=(f"spm_unit columns missing: [{_OUTPUT!r}].",), + details={"missing": [_OUTPUT]}, + ) + + summary = us_energy_subsidy_summary(frame) + failures: list[str] = [] + if summary["nonfinite"]: + failures.append(f"{_OUTPUT}: {int(summary['nonfinite'])} nonfinite values.") + if summary["negative"]: + failures.append(f"{_OUTPUT}: {int(summary['negative'])} negative values.") + share = float(summary["positive_share"]) + low, high = summary["positive_share_band"] + if not (low <= share <= high): + failures.append( + f"{_OUTPUT}: positive share {share:.4f} outside plausibility band " + f"[{low}, {high}]." + ) + channels = summary.get("channels") + if isinstance(channels, dict): + for name in (_BASE_ASEC_SUPPORT_CHANNEL, _PUF_TAX_DETAIL_SUPPORT_CHANNEL): + detail = channels.get(name) + if not isinstance(detail, dict): + failures.append(f"{_OUTPUT}: missing {name} channel diagnostics.") + continue + channel_share = float(detail["positive_share"]) + channel_low, channel_high = detail["positive_share_band"] + if not (channel_low <= channel_share <= channel_high): + failures.append( + f"{_OUTPUT}: {name} positive share {channel_share:.4f} " + "outside plausibility band " + f"[{channel_low}, {channel_high}]." + ) + return GateResult( + name="energy_subsidy_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) + + +def _surface_carries_signal(frame: Frame) -> bool: + return ( + _OUTPUT in frame.table("spm_unit") + and us_energy_subsidy_signal_gate(frame).passed + ) diff --git a/packages/populace-build/src/populace/build/us_runtime/farm_business_income.py b/packages/populace-build/src/populace/build/us_runtime/farm_business_income.py new file mode 100644 index 00000000..0aa9a0e8 --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/farm_business_income.py @@ -0,0 +1,289 @@ +"""Signed IRS PUF farm-operations and farm-rent income inputs. + +The retired eCPS PUF pipeline mapped Schedule F net income ``E02100`` to +``farm_operations_income`` and farm-rent net income ``E27200`` to +``farm_rent_income``. Both sources are signed: losses are factual inputs and +must never be clipped or converted to positive-only support. CPS ASEC +``FRSE_VAL`` independently measures ``farm_operations_income`` on the ASEC +channel. The PUF's separate ``farm_income`` field comes from ``T27800`` and is +not a substitute for either Section 199A income-definition input. + +The processed, versioned PUF artifact already carries the two archived leaves. +Populace's shared weighted QRF replaces them only on the dedicated PUF support +channel, preserving measured ``FRSE_VAL`` operations income and the default-zero +rent leaf on ASEC. The qualification flags are a separate QBI family and do +not create or suppress either signed amount here. This is deliberate Populace +hardening: the retired override replaced both extended-CPS halves, whereas the +source-preserving support design never overwrites measured ASEC observations. +""" + +from __future__ import annotations + +from importlib.resources import files + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import SourceStageSpec, load_source_manifest +from populace.frame import Frame + +__all__ = [ + "FARM_BUSINESS_INCOME_ARCHIVED_CPS_FARM_INCOME_URL", + "FARM_BUSINESS_INCOME_ARCHIVED_DERIVATION_URL", + "FARM_BUSINESS_INCOME_ARCHIVED_EXPORT_URL", + "FARM_BUSINESS_INCOME_ARCHIVED_IMPUTATION_URL", + "FARM_BUSINESS_INCOME_ARCHIVED_OVERRIDE_URL", + "FARM_BUSINESS_INCOME_ARCHIVED_PUF_ARTIFACT_URL", + "US_FARM_BUSINESS_INCOME_NONCONSTANT_PERSON_COLUMNS", + "US_FARM_BUSINESS_INCOME_OUTPUT_COLUMNS", + "US_FARM_BUSINESS_INCOME_STAGE_NAME", + "derive_us_farm_business_income_from_puf", + "us_farm_business_income_signal_gate", + "us_farm_business_income_stage_spec", + "us_farm_business_income_summary", +] + +_ARCHIVED_DATA_REPOSITORY = "policyengine-" + "us-data" +_ARCHIVED_ROOT = ( + "https://github.com/PolicyEngine/" + f"{_ARCHIVED_DATA_REPOSITORY}/blob/" + "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe/" + "policyengine_" + "us_data/" +) +FARM_BUSINESS_INCOME_ARCHIVED_DERIVATION_URL = ( + _ARCHIVED_ROOT + "datasets/puf/puf.py#L636-L704" +) +FARM_BUSINESS_INCOME_ARCHIVED_CPS_FARM_INCOME_URL = ( + _ARCHIVED_ROOT + "datasets/cps/cps.py#L1363-L1382" +) +FARM_BUSINESS_INCOME_ARCHIVED_EXPORT_URL = ( + _ARCHIVED_ROOT + "datasets/puf/puf.py#L804-L875" +) +FARM_BUSINESS_INCOME_ARCHIVED_IMPUTATION_URL = ( + _ARCHIVED_ROOT + "calibration/puf_impute.py#L80-L198" +) +FARM_BUSINESS_INCOME_ARCHIVED_OVERRIDE_URL = ( + _ARCHIVED_ROOT + "calibration/puf_impute.py#L513-L672" +) +FARM_BUSINESS_INCOME_ARCHIVED_PUF_ARTIFACT_URL = ( + _ARCHIVED_ROOT + "datasets/puf/puf.py#L1655-L1660" +) + +US_FARM_BUSINESS_INCOME_STAGE_NAME = "puf_tax_detail" +US_FARM_BUSINESS_INCOME_OUTPUT_COLUMNS: tuple[str, ...] = ( + "farm_operations_income", + "farm_rent_income", +) +US_FARM_BUSINESS_INCOME_NONCONSTANT_PERSON_COLUMNS = ( + US_FARM_BUSINESS_INCOME_OUTPUT_COLUMNS +) + +_OPERATIONS_OUTPUT, _RENT_OUTPUT = US_FARM_BUSINESS_INCOME_OUTPUT_COLUMNS +_SOURCE_COLUMNS = ("E02100", "E27200") +_PERSON_SUPPORT_CHANNEL_COLUMN = "person_support_channel" +_BASE_ASEC_SUPPORT_CHANNEL = "asec" +_PUF_TAX_DETAIL_SUPPORT_CHANNEL = "puf_tax_detail" +_NONZERO_SHARE_BANDS: dict[str, tuple[float, float]] = { + _OPERATIONS_OUTPUT: (0.0001, 0.20), + _RENT_OUTPUT: (0.00001, 0.20), +} + + +def us_farm_business_income_stage_spec() -> SourceStageSpec: + """Load and validate the shared PUF tax-detail output declaration.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + spec = manifest.stage_map()[US_FARM_BUSINESS_INCOME_STAGE_NAME] + missing = sorted(set(US_FARM_BUSINESS_INCOME_OUTPUT_COLUMNS) - set(spec.outputs)) + if missing: + raise ValueError( + f"{US_FARM_BUSINESS_INCOME_STAGE_NAME!r} manifest stage does not " + f"declare farm-business output(s) {missing}." + ) + forbidden_nonnegative = sorted( + set(US_FARM_BUSINESS_INCOME_OUTPUT_COLUMNS) & set(spec.nonnegative_outputs) + ) + if forbidden_nonnegative: + raise ValueError( + "Signed farm-business inputs must not be declared nonnegative: " + f"{forbidden_nonnegative}." + ) + return spec + + +def _strict_signed_values(puf: pd.DataFrame, column: str) -> np.ndarray: + if column not in puf.columns: + raise ValueError( + f"PUF farm-business derivation requires source column {column!r}." + ) + numeric = pd.to_numeric(puf[column], errors="coerce").to_numpy(dtype=np.float64) + nonfinite = int(np.count_nonzero(~np.isfinite(numeric))) + if nonfinite: + raise ValueError( + f"PUF farm-business source column {column!r} contains {nonfinite} " + "nonnumeric or nonfinite value(s)." + ) + return numeric + + +def derive_us_farm_business_income_from_puf( + puf: pd.DataFrame, + *, + operations_source_column: str = _SOURCE_COLUMNS[0], + operations_output_column: str = _OPERATIONS_OUTPUT, + rent_source_column: str = _SOURCE_COLUMNS[1], + rent_output_column: str = _RENT_OUTPUT, +) -> pd.DataFrame: + """Carry both archived signed PUF fields without clipping or mutation.""" + + operations = _strict_signed_values(puf, operations_source_column) + rent = _strict_signed_values(puf, rent_source_column) + result = puf.copy(deep=True) + result[operations_output_column] = operations + result[rent_output_column] = rent + return result + + +def us_farm_business_income_summary(frame: Frame) -> dict[str, object]: + """Return signed, weighted, and support-channel signal diagnostics.""" + + person = frame.table("person") + weights = np.asarray(frame.resolve_weights("person").values, dtype=np.float64) + total_weight = float(weights.sum()) + channel = ( + person[_PERSON_SUPPORT_CHANNEL_COLUMN].astype(str).to_numpy() + if _PERSON_SUPPORT_CHANNEL_COLUMN in person + else None + ) + columns: dict[str, dict[str, object]] = {} + for output in US_FARM_BUSINESS_INCOME_OUTPUT_COLUMNS: + values = pd.to_numeric(person[output], errors="coerce").to_numpy( + dtype=np.float64 + ) + finite = np.isfinite(values) + nonzero = finite & (values != 0.0) + positive = finite & (values > 0.0) + negative = finite & (values < 0.0) + detail: dict[str, object] = { + "nonzero_share": ( + float(weights[nonzero].sum()) / total_weight + if total_weight > 0.0 + else 0.0 + ), + "nonzero_share_band": list(_NONZERO_SHARE_BANDS[output]), + "positive_share": ( + float(weights[positive].sum()) / total_weight + if total_weight > 0.0 + else 0.0 + ), + "negative_share": ( + float(weights[negative].sum()) / total_weight + if total_weight > 0.0 + else 0.0 + ), + "weighted_total": float((np.nan_to_num(values) * weights).sum()), + "weighted_absolute_total": float( + (np.abs(np.nan_to_num(values)) * weights).sum() + ), + "nonfinite": int(np.count_nonzero(~finite)), + } + if channel is not None: + channels: dict[str, dict[str, float | int]] = {} + for name in ( + _BASE_ASEC_SUPPORT_CHANNEL, + _PUF_TAX_DETAIL_SUPPORT_CHANNEL, + ): + mask = channel == name + channel_weight = float(weights[mask].sum()) + channels[name] = { + "nonzero_rows": int(np.count_nonzero(mask & nonzero)), + "nonzero_share": ( + float(weights[mask & nonzero].sum()) / channel_weight + if channel_weight > 0.0 + else 0.0 + ), + "positive_rows": int(np.count_nonzero(mask & positive)), + "negative_rows": int(np.count_nonzero(mask & negative)), + "weighted_total": float( + (np.nan_to_num(values[mask]) * weights[mask]).sum() + ), + "weighted_absolute_total": float( + (np.abs(np.nan_to_num(values[mask])) * weights[mask]).sum() + ), + } + detail["channels"] = channels + columns[output] = detail + return {"columns": columns} + + +def us_farm_business_income_signal_gate(frame: Frame) -> GateResult: + """Require finite signed signal on each source-backed support channel.""" + + person = frame.table("person") + missing = [ + output + for output in US_FARM_BUSINESS_INCOME_OUTPUT_COLUMNS + if output not in person.columns + ] + if missing: + return GateResult( + name="farm_business_income_signal", + passed=False, + failures=(f"person columns missing: {missing}.",), + details={"missing": missing}, + ) + + summary = us_farm_business_income_summary(frame) + failures: list[str] = [] + columns = summary["columns"] + for output in US_FARM_BUSINESS_INCOME_OUTPUT_COLUMNS: + detail = columns[output] + if detail["nonfinite"]: + failures.append(f"{output}: {int(detail['nonfinite'])} nonfinite value(s).") + share = float(detail["nonzero_share"]) + low, high = detail["nonzero_share_band"] + if not (low <= share <= high): + failures.append( + f"{output}: nonzero share {share:.6f} outside plausibility " + f"band [{low}, {high}]." + ) + if float(detail["weighted_absolute_total"]) <= 0.0: + failures.append(f"{output}: weighted absolute total is zero.") + channels = detail.get("channels") + if isinstance(channels, dict): + asec = channels.get(_BASE_ASEC_SUPPORT_CHANNEL) + puf = channels.get(_PUF_TAX_DETAIL_SUPPORT_CHANNEL) + asec_nonzero = ( + int(asec.get("nonzero_rows", 0)) if isinstance(asec, dict) else 0 + ) + if output == _OPERATIONS_OUTPUT and not asec_nonzero: + failures.append( + f"{output}: ASEC support has no measured FRSE_VAL signal." + ) + if output == _OPERATIONS_OUTPUT and isinstance(asec, dict): + if not int(asec.get("positive_rows", 0)): + failures.append(f"{output}: ASEC support has no positive values.") + if not int(asec.get("negative_rows", 0)): + failures.append(f"{output}: ASEC support has no farm losses.") + if output == _RENT_OUTPUT and asec_nonzero: + failures.append( + f"{output}: ASEC support carries nonzero values despite no " + "farm-rent source." + ) + if not isinstance(puf, dict) or not int(puf.get("nonzero_rows", 0)): + failures.append(f"{output}: PUF tax-detail support has no signal.") + elif not int(puf.get("positive_rows", 0)): + failures.append( + f"{output}: PUF tax-detail support has no positive values." + ) + elif not int(puf.get("negative_rows", 0)): + failures.append(f"{output}: PUF tax-detail support has no farm losses.") + return GateResult( + name="farm_business_income_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) diff --git a/packages/populace-build/src/populace/build/us_runtime/form_4952.py b/packages/populace-build/src/populace/build/us_runtime/form_4952.py new file mode 100644 index 00000000..01fc009c --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/form_4952.py @@ -0,0 +1,220 @@ +"""IRS PUF Form 4952 elected-investment-income input for the US build. + +The retired eCPS pipeline carried +``investment_income_elected_form_4952`` directly from IRS PUF field +``E58990``. Immutable archived coordinates for that derivation, the exported +field, and its weighted PUF imputation are exposed below. + +There is no CPS analogue and no synthetic fallback. The PUF tax-detail stage +carries the source exactly, then its shared weighted QRF places the input only +on the dedicated PUF support channel and sparsifies it to the donor's weighted +positive rate. PolicyEngine-US aggregates the person input to the tax unit and +subtracts it from net capital gain; this module persists only the factual +election amount. +""" + +from __future__ import annotations + +from importlib.resources import files + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import SourceStageSpec, load_source_manifest +from populace.frame import Frame + +__all__ = [ + "FORM_4952_ARCHIVED_DERIVATION_URL", + "FORM_4952_ARCHIVED_EXPORT_URL", + "FORM_4952_ARCHIVED_IMPUTATION_URL", + "FORM_4952_ARCHIVED_PERSON_ALLOCATION_URL", + "FORM_4952_ARCHIVED_PUF_ARTIFACT_URL", + "US_FORM_4952_NONCONSTANT_PERSON_COLUMNS", + "US_FORM_4952_OUTPUT_COLUMNS", + "US_FORM_4952_STAGE_NAME", + "derive_us_form_4952_election_from_puf", + "us_form_4952_election_signal_gate", + "us_form_4952_election_stage_spec", + "us_form_4952_election_summary", +] + +_ARCHIVED_DATA_REPOSITORY = "policyengine-" + "us-data" +_ARCHIVED_ROOT = ( + "https://github.com/PolicyEngine/" + f"{_ARCHIVED_DATA_REPOSITORY}/blob/" + "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe/" + "policyengine_" + "us_data/" +) +FORM_4952_ARCHIVED_DERIVATION_URL = _ARCHIVED_ROOT + "datasets/puf/puf.py#L708" +FORM_4952_ARCHIVED_EXPORT_URL = _ARCHIVED_ROOT + "datasets/puf/puf.py#L804-L850" +FORM_4952_ARCHIVED_IMPUTATION_URL = ( + _ARCHIVED_ROOT + "calibration/puf_impute.py#L90-L198" +) +FORM_4952_ARCHIVED_PERSON_ALLOCATION_URL = ( + _ARCHIVED_ROOT + "datasets/puf/puf.py#L477-L546" +) +FORM_4952_ARCHIVED_PUF_ARTIFACT_URL = _ARCHIVED_ROOT + "datasets/puf/puf.py#L1655-L1660" + +US_FORM_4952_STAGE_NAME = "puf_tax_detail" +US_FORM_4952_OUTPUT_COLUMNS: tuple[str, ...] = ("investment_income_elected_form_4952",) +US_FORM_4952_NONCONSTANT_PERSON_COLUMNS = US_FORM_4952_OUTPUT_COLUMNS + +_OUTPUT = US_FORM_4952_OUTPUT_COLUMNS[0] +_PERSON_SUPPORT_CHANNEL_COLUMN = "person_support_channel" +_BASE_ASEC_SUPPORT_CHANNEL = "asec" +_PUF_TAX_DETAIL_SUPPORT_CHANNEL = "puf_tax_detail" +_OVERALL_POSITIVE_SHARE_BAND = (0.00005, 0.02) +_PUF_POSITIVE_SHARE_BAND = (0.0001, 0.04) + + +def us_form_4952_election_stage_spec() -> SourceStageSpec: + """Load and validate the shared PUF tax-detail stage declaration.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + spec = manifest.stage_map()[US_FORM_4952_STAGE_NAME] + missing = sorted(set(US_FORM_4952_OUTPUT_COLUMNS) - set(spec.outputs)) + if missing: + raise ValueError( + f"{US_FORM_4952_STAGE_NAME!r} manifest stage does not declare " + f"Form 4952 output(s) {missing}." + ) + return spec + + +def derive_us_form_4952_election_from_puf( + puf: pd.DataFrame, + *, + source_column: str = "E58990", + output_column: str = _OUTPUT, +) -> pd.DataFrame: + """Carry the archived IRS PUF E58990 field without alteration.""" + + if source_column not in puf.columns: + raise ValueError( + f"PUF Form 4952 derivation requires source column {source_column!r}." + ) + values = pd.to_numeric(puf[source_column], errors="coerce").to_numpy( + dtype=np.float64 + ) + nonfinite = ~np.isfinite(values) + if bool(nonfinite.any()): + raise ValueError( + f"PUF Form 4952 source column {source_column!r} contains " + f"{int(np.count_nonzero(nonfinite))} nonnumeric or nonfinite value(s)." + ) + negative = values < 0.0 + if bool(negative.any()): + raise ValueError( + f"PUF Form 4952 source column {source_column!r} contains " + f"{int(np.count_nonzero(negative))} negative value(s)." + ) + + result = puf.copy(deep=True) + result[output_column] = values + return result + + +def us_form_4952_election_summary(frame: Frame) -> dict[str, object]: + """Return weighted signal, channel, and validity diagnostics.""" + + person = frame.table("person") + values = pd.to_numeric(person[_OUTPUT], errors="coerce").to_numpy(dtype=np.float64) + weights = np.asarray(frame.resolve_weights("person").values, dtype=np.float64) + finite = np.isfinite(values) + positive = finite & (values > 0.0) + total_weight = float(weights.sum()) + summary: dict[str, object] = { + "positive_share": ( + float(weights[positive].sum()) / total_weight if total_weight > 0.0 else 0.0 + ), + "positive_share_band": list(_OVERALL_POSITIVE_SHARE_BAND), + "weighted_total": float((np.nan_to_num(values) * weights).sum()), + "nonfinite": int(np.count_nonzero(~finite)), + "negative": int(np.count_nonzero(finite & (values < 0.0))), + } + if _PERSON_SUPPORT_CHANNEL_COLUMN in person.columns: + channels: dict[str, dict[str, float | int]] = {} + channel_values = person[_PERSON_SUPPORT_CHANNEL_COLUMN].to_numpy() + for channel in person[_PERSON_SUPPORT_CHANNEL_COLUMN].dropna().unique(): + channel_mask = channel_values == channel + channel_weight = float(weights[channel_mask].sum()) + channel_positive = channel_mask & positive + channels[str(channel)] = { + "positive_rows": int(np.count_nonzero(channel_positive)), + "positive_share": ( + float(weights[channel_positive].sum()) / channel_weight + if channel_weight > 0.0 + else 0.0 + ), + "weighted_total": float( + (np.nan_to_num(values[channel_mask]) * weights[channel_mask]).sum() + ), + } + summary["channels"] = channels + return summary + + +def us_form_4952_election_signal_gate(frame: Frame) -> GateResult: + """Require a finite, sparse, source-aligned Form 4952 input.""" + + person = frame.table("person") + missing = [ + column for column in US_FORM_4952_OUTPUT_COLUMNS if column not in person.columns + ] + if missing: + return GateResult( + name="form_4952_election_signal", + passed=False, + failures=(f"person columns missing: {missing}.",), + details={"missing": missing}, + ) + + summary = us_form_4952_election_summary(frame) + failures: list[str] = [] + if summary["nonfinite"]: + failures.append(f"{_OUTPUT} nonfinite values: {int(summary['nonfinite'])}.") + if summary["negative"]: + failures.append(f"{_OUTPUT} negative values: {int(summary['negative'])}.") + share = float(summary["positive_share"]) + low, high = summary["positive_share_band"] + if not (low <= share <= high): + failures.append( + f"Form 4952 positive share {share:.6f} outside plausibility " + f"band [{low}, {high}]." + ) + + channels = summary.get("channels") + if isinstance(channels, dict): + asec = channels.get(_BASE_ASEC_SUPPORT_CHANNEL) + puf = channels.get(_PUF_TAX_DETAIL_SUPPORT_CHANNEL) + if asec is None: + failures.append( + "Form 4952 signal gate is missing the ASEC support channel." + ) + elif float(asec["weighted_total"]) != 0.0: + failures.append( + "Form 4952 input must remain zero on the source-unobserved " + "ASEC support channel." + ) + if puf is None: + failures.append( + "Form 4952 signal gate is missing the PUF-tax-detail support channel." + ) + else: + puf_share = float(puf["positive_share"]) + puf_low, puf_high = _PUF_POSITIVE_SHARE_BAND + if not (puf_low <= puf_share <= puf_high): + failures.append( + f"PUF Form 4952 positive share {puf_share:.6f} outside " + f"plausibility band [{puf_low}, {puf_high}]." + ) + + return GateResult( + name="form_4952_election_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) diff --git a/packages/populace-build/src/populace/build/us_runtime/housing_inputs.py b/packages/populace-build/src/populace/build/us_runtime/housing_inputs.py new file mode 100644 index 00000000..a7af2486 --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/housing_inputs.py @@ -0,0 +1,1407 @@ +"""CPS/ACS housing and tenure inputs carried by the retired eCPS build. + +The archived pipeline used two primary sources for this family: + +* CPS ASEC household ``H_TENURE`` mapped directly to household + ``tenure_type``. ASEC SPM ``SPM_CAPHOUSESUB`` and + ``SPM_TENMORTSTATUS`` mapped directly to ``receives_housing_assistance`` + and ``spm_unit_tenure_type``. Measured assistance receipt also anchors + ``takes_up_housing_assistance_if_eligible`` without adding unobserved + recipients; and +* the processed 2022 ACS public-use artifact supplied a 10-predictor rent + donor. The archived QRF selected household heads, built a shared 10,000-row + sample for rent and real-estate taxes with target-specific allocation masks, + and placed annual ``pre_subsidy_rent`` on the CPS household head. + +Populace preserves those source semantics and uses its shared design-weighted +QRF. Weighting the ACS fit is the sole deliberate strengthening of the +retired implementation: the archived fit loaded ``household_weight`` but did +not pass it through to microimpute. No housing values are manufactured from +defaults or from the retired eCPS output. +""" + +from __future__ import annotations + +import hashlib +import warnings +from importlib.resources import files +from pathlib import Path +from typing import Any + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import SourceStageSpec, load_source_manifest +from populace.frame import US_SCHEMA, Frame + +__all__ = [ + "ACS_2022_RENT_ARTIFACT_SHA256", + "HOUSING_INPUTS_ARCHIVED_ACS_DERIVATION_URL", + "HOUSING_INPUTS_ARCHIVED_CPS_RENT_URL", + "HOUSING_INPUTS_ARCHIVED_CPS_SPM_URL", + "HOUSING_INPUTS_ARCHIVED_PUF_IMPUTATION_URL", + "HOUSING_TAKE_UP_ARCHIVED_DERIVATION_URL", + "HOUSING_TAKE_UP_ARCHIVED_HUD_ETL_URL", + "HOUSING_TAKE_UP_ARCHIVED_PARAMETER_URL", + "US_HOUSING_HOUSEHOLD_OUTPUT_COLUMNS", + "US_HOUSING_INPUTS_OUTPUT_COLUMNS", + "US_HOUSING_INPUTS_STAGE_NAME", + "US_HOUSING_NONCONSTANT_HOUSEHOLD_COLUMNS", + "US_HOUSING_NONCONSTANT_PERSON_COLUMNS", + "US_HOUSING_NONCONSTANT_SPM_UNIT_COLUMNS", + "US_HOUSING_PERSON_OUTPUT_COLUMNS", + "US_HOUSING_REQUIRED_HOUSEHOLD_SOURCE_COLUMNS", + "US_HOUSING_REQUIRED_PERSON_SOURCE_COLUMNS", + "US_HOUSING_SPM_UNIT_OUTPUT_COLUMNS", + "derive_us_housing_inputs", + "impute_us_pre_subsidy_rent", + "impute_us_housing_assistance_to_puf_support", + "load_acs_2022_rent_donor", + "us_housing_inputs_signal_gate", + "us_housing_inputs_stage_spec", + "us_housing_inputs_summary", + "with_us_housing_inputs", +] + +_ARCHIVED_COMMIT = "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe" +_ARCHIVED_ROOT = ( + "https://github.com/PolicyEngine/" + + "policyengine-" + + "us-data/blob/" + + _ARCHIVED_COMMIT +) +HOUSING_INPUTS_ARCHIVED_CPS_RENT_URL = ( + _ARCHIVED_ROOT + "/policyengine_" + "us_data/datasets/cps/cps.py#L417-L535" +) +HOUSING_INPUTS_ARCHIVED_CPS_SPM_URL = ( + _ARCHIVED_ROOT + "/policyengine_" + "us_data/datasets/cps/cps.py#L1612-L1635" +) +HOUSING_INPUTS_ARCHIVED_ACS_DERIVATION_URL = ( + _ARCHIVED_ROOT + "/policyengine_" + "us_data/datasets/acs/acs.py#L82-L135" +) +HOUSING_INPUTS_ARCHIVED_PUF_IMPUTATION_URL = ( + _ARCHIVED_ROOT + "/policyengine_" + "us_data/datasets/cps/extended_cps.py#L135-L194" +) +HOUSING_TAKE_UP_ARCHIVED_DERIVATION_URL = ( + _ARCHIVED_ROOT + "/policyengine_" + "us_data/datasets/cps/cps.py#L664-L682" +) +HOUSING_TAKE_UP_ARCHIVED_PARAMETER_URL = ( + _ARCHIVED_ROOT + + "/policyengine_" + + "us_data/parameters/take_up/housing_assistance.yaml#L1-L15" +) +HOUSING_TAKE_UP_ARCHIVED_HUD_ETL_URL = ( + _ARCHIVED_ROOT + "/policyengine_" + "us_data/db/etl_housing_assistance.py#L35-L169" +) + +US_HOUSING_INPUTS_STAGE_NAME = "acs_rent" +US_HOUSING_PERSON_OUTPUT_COLUMNS: tuple[str, ...] = ("pre_subsidy_rent",) +US_HOUSING_HOUSEHOLD_OUTPUT_COLUMNS: tuple[str, ...] = ("tenure_type",) +US_HOUSING_SPM_UNIT_OUTPUT_COLUMNS: tuple[str, ...] = ( + "receives_housing_assistance", + "takes_up_housing_assistance_if_eligible", + "spm_unit_tenure_type", +) +US_HOUSING_INPUTS_OUTPUT_COLUMNS: tuple[str, ...] = ( + *US_HOUSING_PERSON_OUTPUT_COLUMNS, + *US_HOUSING_SPM_UNIT_OUTPUT_COLUMNS, + *US_HOUSING_HOUSEHOLD_OUTPUT_COLUMNS, +) +US_HOUSING_NONCONSTANT_PERSON_COLUMNS = US_HOUSING_PERSON_OUTPUT_COLUMNS +US_HOUSING_NONCONSTANT_SPM_UNIT_COLUMNS = US_HOUSING_SPM_UNIT_OUTPUT_COLUMNS +US_HOUSING_NONCONSTANT_HOUSEHOLD_COLUMNS = US_HOUSING_HOUSEHOLD_OUTPUT_COLUMNS + +US_HOUSING_REQUIRED_PERSON_SOURCE_COLUMNS: tuple[str, ...] = ( + "SPM_CAPHOUSESUB", + "SPM_TENMORTSTATUS", + "person_spm_unit_id", + "person_household_id", + "is_household_head", + "age", + "is_female", + "employment_income_before_lsr", + "self_employment_income_before_lsr", + "social_security_retirement", + "social_security_disability", + "social_security_survivors", + "social_security_dependents", + "taxable_private_pension_income", + "tax_exempt_private_pension_income", +) +US_HOUSING_REQUIRED_HOUSEHOLD_SOURCE_COLUMNS: tuple[str, ...] = ( + "H_TENURE", + "state_fips", +) + +# SHA-256 of the hermetic processed ACS_2022 ARRAYS artifact used by Build J. +# The loader below does not trust the local generating checkout: it validates +# the exact arrays and entity relationships documented by the immutable +# archived implementation before exposing a donor. +ACS_2022_RENT_ARTIFACT_SHA256 = ( + "0b319b496f19a6913066f9c5ea572edfda3d78a187be6f375846617d0b441bd4" +) + +ACS_RENT_PREDICTORS: tuple[str, ...] = ( + "is_household_head", + "age", + "is_male", + "tenure_type", + "employment_income", + "self_employment_income", + "social_security", + "pension_income", + "state_code_str", + "household_size", +) +_DONOR_WEIGHT_COLUMN = "household_weight" +_DONOR_ALLOCATION_COLUMN = "rent_is_allocated" +_DONOR_REAL_ESTATE_TAX_COLUMN = "real_estate_taxes" +_DONOR_REAL_ESTATE_TAX_ALLOCATION_COLUMN = "real_estate_taxes_is_allocated" +_MAX_TRAIN_SAMPLES = 10_000 +_DEFAULT_N_ESTIMATORS = 100 +_PUF_MAX_TRAIN_SAMPLES = 5_000 +_RENT_SHARE_BAND = (0.05, 0.25) +_HOUSING_ASSISTANCE_SHARE_BAND = (0.005, 0.08) + +_HOUSEHOLD_TENURE_MAP = { + 0: "NONE", + 1: "OWNED_WITH_MORTGAGE", + 2: "RENTED", + 3: "NONE", +} +_HOUSEHOLD_TENURE_CODES = { + "NONE": 0.0, + "OWNED_WITH_MORTGAGE": 1.0, + "OWNED_OUTRIGHT": 1.0, + "RENTED": 2.0, +} +_ACS_CATEGORICAL_PREDICTORS = ("tenure_type", "state_code_str") +_SPM_TENURE_MAP = { + 1: "OWNER_WITH_MORTGAGE", + 2: "OWNER_WITHOUT_MORTGAGE", + 3: "RENTER", +} +_HOUSEHOLD_TENURE_VALUES = frozenset(_HOUSEHOLD_TENURE_CODES) +_SPM_TENURE_VALUES = frozenset(_SPM_TENURE_MAP.values()) +_EXPECTED_HOUSEHOLD_TENURE_VALUES = frozenset({"NONE", "OWNED_WITH_MORTGAGE", "RENTED"}) +_EXPECTED_SPM_TENURE_VALUES = frozenset( + {"OWNER_WITH_MORTGAGE", "OWNER_WITHOUT_MORTGAGE", "RENTER"} +) + +QRF: Any | None = None + + +def us_housing_inputs_stage_spec() -> SourceStageSpec: + """Load and validate the packaged ACS/CPS housing-stage declaration.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + stage_map = manifest.stage_map() + if US_HOUSING_INPUTS_STAGE_NAME not in stage_map: + raise ValueError( + f"US source manifest declares no {US_HOUSING_INPUTS_STAGE_NAME!r} stage." + ) + spec = stage_map[US_HOUSING_INPUTS_STAGE_NAME] + missing = sorted(set(US_HOUSING_INPUTS_OUTPUT_COLUMNS) - set(spec.outputs)) + if missing: + raise ValueError( + f"{US_HOUSING_INPUTS_STAGE_NAME!r} manifest stage does not declare " + f"output(s) {missing}; the runtime and manifest have drifted." + ) + return spec + + +def _sha256(path: Path) -> str: + digest = hashlib.sha256() + with path.open("rb") as stream: + for block in iter(lambda: stream.read(1024 * 1024), b""): + digest.update(block) + return digest.hexdigest() + + +def _decode_strings(values: np.ndarray) -> np.ndarray: + raw = np.asarray(values) + if raw.dtype.kind == "S": + return np.char.decode(raw, "utf-8") + return raw.astype(str) + + +def _numeric_array(values: Any, *, name: str) -> np.ndarray: + result = pd.to_numeric(pd.Series(np.asarray(values)), errors="coerce").to_numpy( + dtype=np.float64 + ) + if not np.isfinite(result).all(): + raise ValueError(f"Housing source column {name!r} contains nonfinite values.") + return result + + +def load_acs_2022_rent_donor( + path: str | Path, + *, + expected_sha256: str | None = ACS_2022_RENT_ARTIFACT_SHA256, +) -> pd.DataFrame: + """Load the archived processed ACS 2022 household-head rent donor. + + The loader reads entity arrays directly, aligns household variables through + ``person_household_id``, collapses ACS owned-outright tenure exactly as the + retired rent stage did, and retains only household heads. Both archived + target-allocation flags remain attached so the exact joint target sample + can be replayed before the rent-specific fit. + """ + + import h5py + + source = Path(path) + if not source.exists(): + raise FileNotFoundError(f"ACS 2022 rent donor not found: {source}") + if expected_sha256 is not None: + actual = _sha256(source) + if actual != expected_sha256: + raise ValueError( + f"ACS 2022 rent donor SHA-256 mismatch: {actual} != {expected_sha256}." + ) + + person_columns = ( + "person_id", + "person_household_id", + "is_household_head", + "age", + "is_male", + "employment_income", + "self_employment_income", + "social_security", + "taxable_private_pension_income", + "rent", + "rent_is_allocated", + "real_estate_taxes", + "real_estate_taxes_is_allocated", + ) + household_columns = ( + "household_id", + "household_weight", + "state_fips", + "tenure_type", + ) + with h5py.File(source, mode="r") as h5: + missing = [ + column + for column in (*person_columns, *household_columns) + if column not in h5 + ] + if missing: + raise ValueError(f"ACS 2022 rent donor missing array(s): {missing}.") + arrays = { + column: np.asarray(h5[column]) + for column in (*person_columns, *household_columns) + } + + person_n = len(arrays["person_id"]) + household_n = len(arrays["household_id"]) + bad_person_lengths = { + column: len(arrays[column]) + for column in person_columns + if len(arrays[column]) != person_n + } + bad_household_lengths = { + column: len(arrays[column]) + for column in household_columns + if len(arrays[column]) != household_n + } + if bad_person_lengths or bad_household_lengths: + raise ValueError( + "ACS 2022 rent donor entity-array lengths disagree: " + f"person={bad_person_lengths}, household={bad_household_lengths}." + ) + + household_ids = np.asarray(arrays["household_id"]) + if pd.Index(household_ids).duplicated().any(): + raise ValueError("ACS 2022 household_id must be unique.") + person_household_ids = np.asarray(arrays["person_household_id"]) + household_index = pd.Index(household_ids) + positions = household_index.get_indexer(person_household_ids) + if (positions < 0).any(): + bad = np.unique(person_household_ids[positions < 0])[:5].tolist() + raise ValueError( + f"ACS 2022 people reference missing household_id value(s): {bad}." + ) + + head_mask = np.asarray(arrays["is_household_head"], dtype=bool) + head_household_ids = person_household_ids[head_mask] + if pd.Index(head_household_ids).duplicated().any(): + raise ValueError("ACS 2022 rent donor has multiple heads in a household.") + household_size = pd.Series(person_household_ids).value_counts(sort=False) + + household_tenure = _decode_strings(arrays["tenure_type"]) + normalized_tenure = np.asarray( + [ + "OWNED_WITH_MORTGAGE" if value == "OWNED_OUTRIGHT" else value + for value in household_tenure + ], + dtype=object, + ) + unknown_tenure = sorted(set(normalized_tenure) - _HOUSEHOLD_TENURE_VALUES) + if unknown_tenure: + raise ValueError( + f"ACS 2022 rent donor has unknown tenure value(s): {unknown_tenure}." + ) + + head_positions = positions[head_mask] + donor = pd.DataFrame( + { + "is_household_head": np.ones(int(head_mask.sum()), dtype=np.float64), + "age": _numeric_array(arrays["age"][head_mask], name="age"), + "is_male": np.asarray(arrays["is_male"][head_mask], dtype=np.float64), + "tenure_type": normalized_tenure[head_positions].astype(str), + "employment_income": _numeric_array( + arrays["employment_income"][head_mask], name="employment_income" + ), + "self_employment_income": _numeric_array( + arrays["self_employment_income"][head_mask], + name="self_employment_income", + ), + "social_security": _numeric_array( + arrays["social_security"][head_mask], name="social_security" + ), + "pension_income": _numeric_array( + arrays["taxable_private_pension_income"][head_mask], + name="taxable_private_pension_income", + ), + "state_code_str": np.asarray( + [f"{int(value):02d}" for value in arrays["state_fips"][head_positions]], + dtype=object, + ), + "household_size": household_size.reindex(head_household_ids).to_numpy( + dtype=np.float64 + ), + "rent": _numeric_array(arrays["rent"][head_mask], name="rent"), + _DONOR_ALLOCATION_COLUMN: np.asarray( + arrays["rent_is_allocated"][head_mask], dtype=bool + ), + _DONOR_REAL_ESTATE_TAX_COLUMN: _numeric_array( + arrays["real_estate_taxes"][head_mask], name="real_estate_taxes" + ), + _DONOR_REAL_ESTATE_TAX_ALLOCATION_COLUMN: np.asarray( + arrays["real_estate_taxes_is_allocated"][head_mask], dtype=bool + ), + _DONOR_WEIGHT_COLUMN: _numeric_array( + arrays["household_weight"][head_positions], name="household_weight" + ), + } + ) + if (donor["rent"] < 0.0).any(): + raise ValueError("ACS 2022 rent donor contains negative rent values.") + if (donor[_DONOR_REAL_ESTATE_TAX_COLUMN] < 0.0).any(): + raise ValueError( + "ACS 2022 rent donor contains negative real-estate-tax values." + ) + # Keep zero-WGTP group-quarters heads through the archived joint-target + # sampler. The retired unweighted fit retained them; Populace's deliberate + # design-weighting strengthening gives them zero modeling mass without + # changing which deterministic 10,000-row sample was selected. + if (donor[_DONOR_WEIGHT_COLUMN] < 0.0).any(): + raise ValueError("ACS 2022 rent donor contains negative weights.") + if float(donor[_DONOR_WEIGHT_COLUMN].sum()) <= 0.0: + raise ValueError("ACS 2022 rent donor has no positive household weight.") + if not (~donor[_DONOR_ALLOCATION_COLUMN]).any(): + raise ValueError("ACS 2022 rent donor has no unallocated rent observations.") + if not (~donor[_DONOR_REAL_ESTATE_TAX_ALLOCATION_COLUMN]).any(): + raise ValueError( + "ACS 2022 rent donor has no unallocated real-estate-tax observations." + ) + return donor + + +def _required_columns( + table: pd.DataFrame, columns: tuple[str, ...], label: str +) -> None: + missing = [column for column in columns if column not in table.columns] + if missing: + raise ValueError(f"US housing inputs require {label} column(s): {missing}.") + + +def _constant_source_by_unit( + person: pd.DataFrame, + *, + membership_column: str, + source_column: str, + unit_ids: pd.Series, +) -> np.ndarray: + source = pd.to_numeric(person[source_column], errors="coerce") + if source.isna().any(): + raise ValueError( + f"US housing source {source_column!r} contains missing values." + ) + replicas = pd.DataFrame( + { + "unit_id": person[membership_column].to_numpy(), + "value": source.to_numpy(dtype=np.float64), + } + ) + bounds = replicas.groupby("unit_id", sort=False)["value"].agg(["min", "max"]) + unequal = bounds["min"].to_numpy() != bounds["max"].to_numpy() + if unequal.any(): + bad = bounds.index.to_numpy()[unequal][:5].tolist() + raise ValueError( + f"US housing source {source_column!r} must be constant within its " + f"SPM unit; unit id(s) {bad} disagree." + ) + values = bounds["min"].reindex(unit_ids) + if values.isna().any(): + bad = unit_ids.loc[values.isna()].head().tolist() + raise ValueError( + f"US housing source {source_column!r} does not cover SPM unit id(s) {bad}." + ) + return values.to_numpy(dtype=np.float64) + + +def derive_us_housing_inputs(frame: Frame) -> Frame: + """Carry the three exact ASEC housing/tenure inputs onto their entities.""" + + if frame.schema != US_SCHEMA: + raise ValueError("US housing inputs require the US schema.") + person = frame.table("person") + household = frame.table("household") + spm_unit = frame.table("spm_unit") + _required_columns( + person, + ("SPM_CAPHOUSESUB", "SPM_TENMORTSTATUS", "person_spm_unit_id"), + "person", + ) + _required_columns(household, ("H_TENURE",), "household") + + raw_household_tenure = pd.to_numeric( + household["H_TENURE"], errors="coerce" + ).to_numpy(dtype=np.float64) + if not np.isfinite(raw_household_tenure).all(): + raise ValueError("US housing source H_TENURE contains nonfinite values.") + raw_household_codes = raw_household_tenure.astype(np.int64) + if not np.array_equal(raw_household_tenure, raw_household_codes): + raise ValueError("US housing source H_TENURE contains non-integer codes.") + unknown_household_codes = sorted( + set(raw_household_codes) - set(_HOUSEHOLD_TENURE_MAP) + ) + if unknown_household_codes: + raise ValueError( + "US housing source H_TENURE contains unknown code(s): " + f"{unknown_household_codes}." + ) + + subsidy = _constant_source_by_unit( + person, + membership_column="person_spm_unit_id", + source_column="SPM_CAPHOUSESUB", + unit_ids=spm_unit["spm_unit_id"], + ) + if (subsidy < 0.0).any(): + raise ValueError("US housing source SPM_CAPHOUSESUB contains negative values.") + raw_spm_tenure = _constant_source_by_unit( + person, + membership_column="person_spm_unit_id", + source_column="SPM_TENMORTSTATUS", + unit_ids=spm_unit["spm_unit_id"], + ) + raw_spm_codes = raw_spm_tenure.astype(np.int64) + if not np.array_equal(raw_spm_tenure, raw_spm_codes): + raise ValueError("US housing source SPM_TENMORTSTATUS has non-integer codes.") + unknown_spm_codes = sorted(set(raw_spm_codes) - set(_SPM_TENURE_MAP)) + if unknown_spm_codes: + raise ValueError( + "US housing source SPM_TENMORTSTATUS contains unknown code(s): " + f"{unknown_spm_codes}." + ) + + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + tables["household"]["tenure_type"] = np.asarray( + [_HOUSEHOLD_TENURE_MAP[code] for code in raw_household_codes], dtype=object + ) + tables["spm_unit"]["receives_housing_assistance"] = subsidy > 0.0 + tables["spm_unit"]["takes_up_housing_assistance_if_eligible"] = subsidy > 0.0 + tables["spm_unit"]["spm_unit_tenure_type"] = np.asarray( + [_SPM_TENURE_MAP[code] for code in raw_spm_codes], dtype=object + ) + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + ) + + +def _person_numeric(person: pd.DataFrame, *columns: str) -> np.ndarray: + for column in columns: + if column in person.columns: + values = pd.to_numeric(person[column], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not np.isfinite(values).all(): + raise ValueError( + f"US housing recipient column {column!r} contains nonfinite values." + ) + return values + raise ValueError( + f"US housing recipient is missing every alternative column: {list(columns)}." + ) + + +def _person_sum(person: pd.DataFrame, columns: tuple[str, ...]) -> np.ndarray: + present = [column for column in columns if column in person.columns] + if not present: + raise ValueError( + f"US housing recipient is missing component column(s): {list(columns)}." + ) + values = np.zeros(len(person), dtype=np.float64) + for column in present: + values += _person_numeric(person, column) + return values + + +def _strict_boolean_signal( + values: pd.Series, +) -> tuple[np.ndarray, int, int]: + """Normalize boolean/0/1 values while retaining missing/invalid counts.""" + + raw = values.to_numpy(dtype=object) + normalized = np.zeros(len(raw), dtype=bool) + missing = 0 + invalid = 0 + for index, value in enumerate(raw): + if pd.isna(value): + missing += 1 + continue + if isinstance(value, (bool, np.bool_)): + normalized[index] = bool(value) + continue + if isinstance(value, (int, np.integer, float, np.floating)): + numeric = float(value) + if np.isfinite(numeric) and numeric in (0.0, 1.0): + normalized[index] = numeric == 1.0 + continue + invalid += 1 + return normalized, missing, invalid + + +def _recipient_head_features(frame: Frame) -> tuple[pd.DataFrame, np.ndarray]: + person = frame.table("person") + household = frame.table("household") + _required_columns( + person, + ( + "person_household_id", + "is_household_head", + "age", + "is_female", + "employment_income_before_lsr", + "self_employment_income_before_lsr", + ), + "person", + ) + _required_columns( + household, + ("household_id", "state_fips", "tenure_type"), + "household", + ) + + head_mask = person["is_household_head"].fillna(False).astype(bool).to_numpy() + heads = person.loc[head_mask] + head_ids = heads["person_household_id"].to_numpy() + if pd.Index(head_ids).duplicated().any(): + raise ValueError("US housing recipient selected multiple heads per household.") + household_ids = household["household_id"].to_numpy() + missing = sorted(set(household_ids) - set(head_ids)) + extra = sorted(set(head_ids) - set(household_ids)) + if missing or extra: + raise ValueError( + "US housing household-head alignment failed: " + f"missing={missing[:5]}, extra={extra[:5]}." + ) + + tenure = household["tenure_type"].astype(str) + unknown = sorted(set(tenure) - _HOUSEHOLD_TENURE_VALUES) + if unknown: + raise ValueError( + f"US housing recipient has unknown tenure value(s): {unknown}." + ) + tenure_by_id = pd.Series(tenure.to_numpy(), index=household_ids) + state_by_id = pd.Series( + pd.to_numeric(household["state_fips"], errors="coerce").to_numpy(), + index=household_ids, + ) + size_by_id = person["person_household_id"].value_counts(sort=False) + + social_security = _person_sum( + person, + ( + "social_security_retirement", + "social_security_disability", + "social_security_survivors", + "social_security_dependents", + ), + ) + pension_income = _person_sum( + person, + ( + "taxable_private_pension_income", + "tax_exempt_private_pension_income", + "taxable_public_pension_income", + "tax_exempt_public_pension_income", + ), + ) + all_features = pd.DataFrame( + { + "is_household_head": np.ones(len(person), dtype=np.float64), + "age": _person_numeric(person, "age", "A_AGE"), + "is_male": (~person["is_female"].fillna(False).astype(bool)).to_numpy( + dtype=np.float64 + ), + "employment_income": _person_numeric( + person, "employment_income_before_lsr" + ), + "self_employment_income": _person_numeric( + person, "self_employment_income_before_lsr" + ), + "social_security": social_security, + "pension_income": pension_income, + }, + index=person.index, + ).loc[head_mask] + all_features["tenure_type"] = pd.Series(head_ids).map(tenure_by_id).to_numpy() + all_features["state_code_str"] = np.asarray( + [f"{int(value):02d}" for value in pd.Series(head_ids).map(state_by_id)], + dtype=object, + ) + all_features["household_size"] = pd.Series(head_ids).map(size_by_id).to_numpy() + all_features.index = head_ids + aligned = all_features.reindex(household_ids) + for column in set(ACS_RENT_PREDICTORS) - set(_ACS_CATEGORICAL_PREDICTORS): + numeric = pd.to_numeric(aligned[column], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not np.isfinite(numeric).all(): + raise ValueError( + f"US housing recipient predictor {column!r} contains nonfinite values." + ) + aligned[column] = numeric + return aligned.loc[:, list(ACS_RENT_PREDICTORS)], head_mask + + +def _encode_acs_predictors( + training: pd.DataFrame, + prediction: pd.DataFrame, +) -> tuple[pd.DataFrame, pd.DataFrame, tuple[str, ...]]: + """Dummy-encode the retired QRF's string predictors without ordinality.""" + + numeric_predictors = tuple( + predictor + for predictor in ACS_RENT_PREDICTORS + if predictor not in _ACS_CATEGORICAL_PREDICTORS + ) + encoded_training = training.loc[:, numeric_predictors].copy() + encoded_prediction = prediction.loc[:, numeric_predictors].copy() + encoded_columns = list(numeric_predictors) + for column in _ACS_CATEGORICAL_PREDICTORS: + training_values = training[column].astype(str) + prediction_values = prediction[column].astype(str) + levels = tuple(sorted(training_values.unique())) + unknown = sorted(set(prediction_values.unique()) - set(levels)) + if unknown: + raise ValueError( + f"ACS rent recipient {column!r} has donor-unsupported value(s): " + f"{unknown}." + ) + if len(levels) < 2: + raise ValueError( + f"ACS rent donor categorical predictor {column!r} is constant." + ) + for level in levels[1:]: + dummy = f"{column}__{level}" + encoded_training[dummy] = training_values.eq(level).to_numpy( + dtype=np.float64 + ) + encoded_prediction[dummy] = prediction_values.eq(level).to_numpy( + dtype=np.float64 + ) + encoded_columns.append(dummy) + return encoded_training, encoded_prediction, tuple(encoded_columns) + + +def _stable_string_hash(value: str) -> np.uint64: + with warnings.catch_warnings(): + warnings.filterwarnings("ignore", "overflow encountered", RuntimeWarning) + result = np.uint64(0) + for byte in value.encode("utf-8"): + result = result * np.uint64(31) + np.uint64(byte) + result = result ^ (result >> np.uint64(33)) + result = result * np.uint64(0xFF51AFD7ED558CCD) + result = result ^ (result >> np.uint64(33)) + return result + + +def _archived_training_rng(*, salt: str | None = None) -> np.random.Generator: + key = "legacy_acs_rent_training_sample" + if salt is not None: + key = f"{key}:{salt}" + return np.random.default_rng(int(_stable_string_hash(key)) % (2**63)) + + +def _archived_joint_training_sample( + donor: pd.DataFrame, +) -> tuple[pd.DataFrame, dict[str, np.ndarray]]: + """Replay the retired rent/real-estate-tax target-filtered sample cap.""" + + target_masks = { + "rent": ( + np.isfinite(donor["rent"].to_numpy(dtype=np.float64)) + & ~donor[_DONOR_ALLOCATION_COLUMN].astype(bool).to_numpy() + ), + _DONOR_REAL_ESTATE_TAX_COLUMN: ( + np.isfinite(donor[_DONOR_REAL_ESTATE_TAX_COLUMN].to_numpy(dtype=np.float64)) + & ~donor[_DONOR_REAL_ESTATE_TAX_ALLOCATION_COLUMN].astype(bool).to_numpy() + ), + } + for target, mask in target_masks.items(): + if not mask.any(): + raise ValueError(f"ACS rent donor has no observed rows for {target}.") + + union_mask = np.logical_or.reduce(tuple(target_masks.values())) + union_positions = np.flatnonzero(union_mask) + if len(union_positions) <= _MAX_TRAIN_SAMPLES: + sample_positions = union_positions + else: + selected: list[int] = [] + selected_set: set[int] = set() + per_target_cap = _MAX_TRAIN_SAMPLES // len(target_masks) + for target, mask in target_masks.items(): + target_positions = np.flatnonzero(mask) + target_sample = _archived_training_rng(salt=target).choice( + target_positions, + size=min(per_target_cap, len(target_positions)), + replace=False, + ) + for position in target_sample: + integer_position = int(position) + if integer_position not in selected_set: + selected.append(integer_position) + selected_set.add(integer_position) + + remaining_n = _MAX_TRAIN_SAMPLES - len(selected) + if remaining_n > 0: + remaining_positions = np.asarray( + [ + position + for position in union_positions + if int(position) not in selected_set + ], + dtype=np.int64, + ) + if len(remaining_positions): + fill_sample = _archived_training_rng(salt="fill").choice( + remaining_positions, + size=min(remaining_n, len(remaining_positions)), + replace=False, + ) + selected.extend(int(position) for position in fill_sample) + sample_positions = np.asarray(selected, dtype=np.int64) + + sampled = donor.iloc[sample_positions].copy().reset_index(drop=True) + sampled_masks = { + target: mask[sample_positions] for target, mask in target_masks.items() + } + return sampled, sampled_masks + + +def impute_us_pre_subsidy_rent( + frame: Frame, + donor: pd.DataFrame, + *, + seed: int, + n_estimators: int = _DEFAULT_N_ESTIMATORS, +) -> np.ndarray: + """Draw annual ACS rent once per CPS household and place it on the head.""" + + required = ( + *ACS_RENT_PREDICTORS, + "rent", + _DONOR_ALLOCATION_COLUMN, + _DONOR_REAL_ESTATE_TAX_COLUMN, + _DONOR_REAL_ESTATE_TAX_ALLOCATION_COLUMN, + _DONOR_WEIGHT_COLUMN, + ) + missing = [column for column in required if column not in donor.columns] + if missing: + raise ValueError(f"ACS rent donor table missing column(s): {missing}.") + training_sample, target_masks = _archived_joint_training_sample( + donor.loc[:, required].copy() + ) + fit_frame = training_sample.loc[target_masks["rent"]].copy() + if fit_frame.empty: + raise ValueError("ACS rent donor has no observed training rows after sampling.") + for column in ( + *( + predictor + for predictor in ACS_RENT_PREDICTORS + if predictor not in _ACS_CATEGORICAL_PREDICTORS + ), + "rent", + _DONOR_WEIGHT_COLUMN, + ): + fit_frame[column] = pd.to_numeric(fit_frame[column], errors="coerce") + if not np.isfinite(fit_frame[column].to_numpy(dtype=np.float64)).all(): + raise ValueError( + f"ACS rent donor column {column!r} contains nonfinite values." + ) + if (fit_frame["rent"] < 0.0).any(): + raise ValueError("ACS rent donor contains negative rent values.") + if (fit_frame[_DONOR_WEIGHT_COLUMN] < 0.0).any(): + raise ValueError("ACS rent donor weights must be nonnegative.") + if float(fit_frame[_DONOR_WEIGHT_COLUMN].sum()) <= 0.0: + raise ValueError("ACS rent donor sampled weights sum to zero.") + + global QRF + if QRF is None: + from importlib import import_module + + QRF = import_module("populace.fit").QRF + features, head_mask = _recipient_head_features(frame) + encoded_training, encoded_features, encoded_predictors = _encode_acs_predictors( + fit_frame, + features, + ) + encoded_training["rent"] = fit_frame["rent"].to_numpy(dtype=np.float64) + encoded_training[_DONOR_WEIGHT_COLUMN] = fit_frame[_DONOR_WEIGHT_COLUMN].to_numpy( + dtype=np.float64 + ) + fitted = QRF(n_estimators=int(n_estimators), seed=int(seed)).fit( + encoded_training, + predictors=list(encoded_predictors), + targets=["rent"], + weights=_DONOR_WEIGHT_COLUMN, + ) + predicted = pd.to_numeric( + fitted.predict(encoded_features)["rent"], errors="coerce" + ).to_numpy(dtype=np.float64) + if not np.isfinite(predicted).all(): + raise ValueError("ACS rent QRF produced nonfinite predictions.") + predicted = np.maximum(predicted, 0.0) + person_rent = np.zeros(frame.n("person"), dtype=np.float64) + person_rent[head_mask] = predicted + return person_rent + + +def _person_puf_predictors(frame: Frame) -> pd.DataFrame: + """Build the retired eight CPS-to-PUF receiver predictors on people.""" + + person = frame.table("person") + tax_unit = frame.table("tax_unit") + predictors = pd.DataFrame(index=person.index) + predictors["age"] = _person_numeric(person, "age", "A_AGE") + if "is_male" in person: + predictors["is_male"] = person["is_male"].fillna(False).astype(bool) + elif "is_female" in person: + predictors["is_male"] = ~person["is_female"].fillna(False).astype(bool) + elif "A_SEX" in person: + predictors["is_male"] = _person_numeric(person, "A_SEX") == 1 + else: + raise ValueError( + "US housing PUF imputation requires is_male, is_female, or A_SEX." + ) + if "has_esi" not in person: + raise ValueError("US housing PUF imputation requires person.has_esi.") + predictors["has_esi"] = person["has_esi"].fillna(False).astype(bool) + + filing_status_column = next( + ( + column + for column in ("filing_status_input", "filing_status") + if column in tax_unit + ), + None, + ) + if filing_status_column is None: + raise ValueError("US housing PUF imputation requires tax-unit filing status.") + filing_status = tax_unit[filing_status_column].map( + lambda value: ( + value.decode() if isinstance(value, (bytes, np.bytes_)) else str(value) + ) + ) + joint_by_id = pd.Series( + filing_status.eq("JOINT").to_numpy(), + index=tax_unit["tax_unit_id"].to_numpy(), + ) + predictors["tax_unit_is_joint"] = ( + person["person_tax_unit_id"].map(joint_by_id).fillna(False).astype(bool) + ) + if "tax_unit_role_input" not in person: + raise ValueError( + "US housing PUF imputation requires tax_unit_role_input to count " + "dependents." + ) + dependent = ( + person["tax_unit_role_input"] + .map( + lambda value: ( + value.decode() if isinstance(value, (bytes, np.bytes_)) else str(value) + ) + ) + .eq("DEPENDENT") + ) + predictors["tax_unit_count_dependents"] = ( + dependent.groupby(person["person_tax_unit_id"]) + .transform("sum") + .to_numpy(dtype=np.float64) + ) + predictors["employment_income"] = _person_numeric( + person, "employment_income_before_lsr", "WSAL_VAL" + ) + predictors["self_employment_income"] = _person_numeric( + person, "self_employment_income_before_lsr", "SEMP_VAL" + ) + social_security_columns = ( + "social_security_retirement", + "social_security_disability", + "social_security_dependents", + "social_security_survivors", + ) + if all(column in person for column in social_security_columns): + predictors["social_security"] = _person_sum(person, social_security_columns) + else: + predictors["social_security"] = _person_numeric(person, "SS_VAL") + return predictors + + +def impute_us_housing_assistance_to_puf_support( + frame: Frame, + *, + seed: int, + n_estimators: int = _DEFAULT_N_ESTIMATORS, + max_train_samples: int = _PUF_MAX_TRAIN_SAMPLES, +) -> Frame: + """Replace only the PUF clone's housing-assistance receipt flag by QRF. + + The archived extended-CPS stage included this one housing leaf in its + common eight-predictor CPS-to-PUF fit, then reduced person predictions to + the SPM unit by first-person value. The measured ASEC half remains exact; + rent and both tenure enums are cloned unchanged. + """ + + from populace.build.us_runtime.puf_support import ( + BASE_ASEC_SUPPORT_CHANNEL, + PUF_TAX_DETAIL_SUPPORT_CHANNEL, + support_channel_column, + ) + + if frame.schema != US_SCHEMA: + raise ValueError("US housing inputs require the US schema.") + person = frame.table("person") + spm_unit = frame.table("spm_unit") + person_channel_column = support_channel_column("person") + spm_channel_column = support_channel_column("spm_unit") + _required_columns( + person, + (person_channel_column, "person_spm_unit_id", "person_id"), + "person", + ) + _required_columns( + spm_unit, + ( + spm_channel_column, + "spm_unit_id", + "receives_housing_assistance", + "takes_up_housing_assistance_if_eligible", + ), + "spm_unit", + ) + predictors = _person_puf_predictors(frame) + receipt_values, receipt_missing, receipt_invalid = _strict_boolean_signal( + spm_unit["receives_housing_assistance"] + ) + take_up_values, take_up_missing, take_up_invalid = _strict_boolean_signal( + spm_unit["takes_up_housing_assistance_if_eligible"] + ) + if receipt_missing or receipt_invalid or take_up_missing or take_up_invalid: + raise ValueError( + "US housing PUF imputation requires complete boolean receipt/take-up " + "anchors; " + f"receipt_missing={receipt_missing}, receipt_invalid={receipt_invalid}, " + f"take_up_missing={take_up_missing}, take_up_invalid={take_up_invalid}." + ) + mismatches = int(np.count_nonzero(receipt_values != take_up_values)) + if mismatches: + raise ValueError( + "US housing PUF imputation requires take-up to equal measured receipt " + f"before fitting; mismatches={mismatches}." + ) + receipt_by_unit = pd.Series( + receipt_values, + index=spm_unit["spm_unit_id"].to_numpy(), + ) + person_target = person["person_spm_unit_id"].map(receipt_by_unit) + if person_target.isna().any(): + bad = person.loc[person_target.isna(), "person_spm_unit_id"].head().tolist() + raise ValueError( + f"US housing assistance target does not cover person SPM unit id(s) {bad}." + ) + + channel = person[person_channel_column].astype(str) + asec_mask = channel == BASE_ASEC_SUPPORT_CHANNEL + puf_mask = channel == PUF_TAX_DETAIL_SUPPORT_CHANNEL + if not asec_mask.any() or not puf_mask.any(): + raise ValueError( + "US housing PUF imputation requires nonempty ASEC and PUF-tax-detail " + "support channels." + ) + predictor_names = tuple(predictors.columns) + training = predictors.loc[asec_mask].copy() + training["receives_housing_assistance"] = person_target.loc[asec_mask].to_numpy( + dtype=np.float64 + ) + weights = pd.Series( + np.asarray(frame.resolve_weights("person").values, dtype=np.float64), + index=person.index, + ).loc[asec_mask] + if not np.isfinite(weights.to_numpy()).all() or (weights < 0.0).any(): + raise ValueError( + "US housing PUF QRF person weights must be finite/nonnegative." + ) + if float(weights.sum()) <= 0.0: + raise ValueError("US housing PUF QRF person weights sum to zero.") + if len(training) > int(max_train_samples): + selected = training.sample( + n=int(max_train_samples), random_state=int(seed) + ).index + training = training.loc[selected] + weights = weights.loc[selected] + test = predictors.loc[puf_mask].copy() + for column in (*predictor_names, "receives_housing_assistance"): + training[column] = pd.to_numeric(training[column], errors="coerce") + if not np.isfinite(training[column].to_numpy(dtype=np.float64)).all(): + raise ValueError( + f"US housing PUF QRF training column {column!r} is nonfinite." + ) + for column in predictor_names: + test[column] = pd.to_numeric(test[column], errors="coerce") + if not np.isfinite(test[column].to_numpy(dtype=np.float64)).all(): + raise ValueError( + f"US housing PUF QRF prediction column {column!r} is nonfinite." + ) + + global QRF + if QRF is None: + from importlib import import_module + + QRF = import_module("populace.fit").QRF + fitted = QRF(n_estimators=int(n_estimators), seed=int(seed)).fit( + training, + predictors=list(predictor_names), + targets=["receives_housing_assistance"], + weights=weights.to_numpy(dtype=np.float64), + ) + predicted = pd.to_numeric( + fitted.predict(test)["receives_housing_assistance"], errors="coerce" + ).to_numpy(dtype=np.float64) + if not np.isfinite(predicted).all(): + raise ValueError("US housing PUF QRF produced nonfinite predictions.") + + person_result = person_target.astype(bool).copy() + person_result.loc[puf_mask] = predicted >= 0.5 + unit_values = ( + pd.DataFrame( + { + "spm_unit_id": person["person_spm_unit_id"].to_numpy(), + "value": person_result.to_numpy(dtype=bool), + } + ) + .groupby("spm_unit_id", sort=False)["value"] + .first() + ) + aligned = unit_values.reindex(spm_unit["spm_unit_id"]) + if aligned.isna().any(): + raise ValueError("US housing PUF QRF output does not cover every SPM unit.") + puf_units = ( + spm_unit[spm_channel_column].astype(str).eq(PUF_TAX_DETAIL_SUPPORT_CHANNEL) + ) + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + tables["spm_unit"].loc[puf_units, "receives_housing_assistance"] = aligned.to_numpy( + dtype=bool + )[puf_units.to_numpy()] + tables["spm_unit"].loc[puf_units, "takes_up_housing_assistance_if_eligible"] = ( + aligned.to_numpy(dtype=bool)[puf_units.to_numpy()] + ) + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + ) + + +def with_us_housing_inputs( + frame: Frame, + *, + seed: int, + time_period: int, + acs_rent_donor: pd.DataFrame, + n_estimators: int = _DEFAULT_N_ESTIMATORS, +) -> Frame: + """Materialize source-backed housing and tenure inputs on a US frame.""" + + if frame.schema != US_SCHEMA: + raise ValueError("US housing inputs require the US schema.") + del time_period # The archived donor vintage is pinned in the stage spec. + us_housing_inputs_stage_spec() + if us_housing_inputs_signal_gate(frame).passed: + return frame + carried = derive_us_housing_inputs(frame) + rent = impute_us_pre_subsidy_rent( + carried, + acs_rent_donor, + seed=int(seed), + n_estimators=int(n_estimators), + ) + tables = {entity: carried.table(entity).copy() for entity in carried.entities} + tables["person"]["pre_subsidy_rent"] = rent + return Frame( + tables, + carried.schema, + {entity: carried.weights_for(entity) for entity in carried.weighted_entities}, + carried.strata, + mass_log=carried.mass_log, + ) + + +def us_housing_inputs_summary(frame: Frame) -> dict[str, object]: + """Return weighted signal and entity-coherence diagnostics.""" + + person = frame.table("person") + household = frame.table("household") + spm_unit = frame.table("spm_unit") + rent = pd.to_numeric(person["pre_subsidy_rent"], errors="coerce").to_numpy( + dtype=np.float64 + ) + person_weights = np.asarray( + frame.resolve_weights("person").values, dtype=np.float64 + ) + total_person_weight = float(person_weights.sum()) + positive_rent = np.isfinite(rent) & (rent > 0.0) + rent_share = ( + float(person_weights[positive_rent].sum()) / total_person_weight + if total_person_weight > 0.0 + else 0.0 + ) + + head = person["is_household_head"].fillna(False).astype(bool).to_numpy() + tenure_by_household = pd.Series( + household["tenure_type"].astype(str).to_numpy(), + index=household["household_id"].to_numpy(), + ) + person_tenure = person["person_household_id"].map(tenure_by_household).astype(str) + nonrenter_positive = positive_rent & person_tenure.ne("RENTED").to_numpy() + + receives, receives_missing, receives_invalid = _strict_boolean_signal( + spm_unit["receives_housing_assistance"] + ) + takes_up, takes_up_missing, takes_up_invalid = _strict_boolean_signal( + spm_unit["takes_up_housing_assistance_if_eligible"] + ) + spm_weights = np.asarray(frame.resolve_weights("spm_unit").values, dtype=np.float64) + total_spm_weight = float(spm_weights.sum()) + receives_share = ( + float(spm_weights[receives].sum()) / total_spm_weight + if total_spm_weight > 0.0 + else 0.0 + ) + from populace.build.us_runtime.puf_support import ( + BASE_ASEC_SUPPORT_CHANNEL, + PUF_TAX_DETAIL_SUPPORT_CHANNEL, + support_channel_column, + ) + + assistance_share_by_channel: dict[str, float] = {} + assistance_positive_by_channel: dict[str, int] = {} + spm_channel_column = support_channel_column("spm_unit") + if spm_channel_column in spm_unit: + channels = spm_unit[spm_channel_column].astype(str).to_numpy() + for channel_name in ( + BASE_ASEC_SUPPORT_CHANNEL, + PUF_TAX_DETAIL_SUPPORT_CHANNEL, + ): + channel_mask = channels == channel_name + if not channel_mask.any(): + continue + channel_weight = float(spm_weights[channel_mask].sum()) + assistance_share_by_channel[channel_name] = ( + float(spm_weights[channel_mask & receives].sum()) / channel_weight + if channel_weight > 0.0 + else 0.0 + ) + assistance_positive_by_channel[channel_name] = int( + np.count_nonzero(channel_mask & receives) + ) + household_values = sorted(set(household["tenure_type"].dropna().astype(str))) + spm_values = sorted(set(spm_unit["spm_unit_tenure_type"].dropna().astype(str))) + return { + "pre_subsidy_rent_share": rent_share, + "pre_subsidy_rent_share_band": list(_RENT_SHARE_BAND), + "pre_subsidy_rent_total": float(np.nansum(rent * person_weights)), + "pre_subsidy_rent_unweighted_total": float(np.nansum(rent)), + "housing_assistance_share": receives_share, + "housing_assistance_share_band": list(_HOUSING_ASSISTANCE_SHARE_BAND), + "housing_assistance_receipt_missing_count": receives_missing, + "housing_assistance_receipt_invalid_count": receives_invalid, + "housing_assistance_take_up_missing_count": takes_up_missing, + "housing_assistance_take_up_invalid_count": takes_up_invalid, + "housing_assistance_take_up_source_mismatch_count": int( + np.count_nonzero(takes_up != receives) + ), + "housing_assistance_share_by_support_channel": assistance_share_by_channel, + "housing_assistance_positive_by_support_channel": ( + assistance_positive_by_channel + ), + "has_spm_support_channel_metadata": spm_channel_column in spm_unit, + "nonfinite_rent": int(np.count_nonzero(~np.isfinite(rent))), + "negative_rent": int(np.count_nonzero(rent < 0.0)), + "positive_rent_nonhead": int(np.count_nonzero(positive_rent & ~head)), + "positive_rent_nonrenter": int(np.count_nonzero(nonrenter_positive)), + "household_tenure_values": household_values, + "spm_tenure_values": spm_values, + "unknown_household_tenure_values": sorted( + set(household_values) - _HOUSEHOLD_TENURE_VALUES + ), + "unknown_spm_tenure_values": sorted(set(spm_values) - _SPM_TENURE_VALUES), + } + + +def us_housing_inputs_signal_gate(frame: Frame) -> GateResult: + """Require plausible non-default signal for all five restored inputs.""" + + missing: list[str] = [] + for entity, columns in ( + ("person", US_HOUSING_PERSON_OUTPUT_COLUMNS), + ("spm_unit", US_HOUSING_SPM_UNIT_OUTPUT_COLUMNS), + ("household", US_HOUSING_HOUSEHOLD_OUTPUT_COLUMNS), + ): + table = frame.table(entity) + missing.extend( + f"{entity}.{column}" for column in columns if column not in table.columns + ) + if missing: + return GateResult( + name="housing_inputs_signal", + passed=False, + failures=(f"columns missing: {missing}.",), + details={"missing": missing}, + ) + person = frame.table("person") + required_person = ("is_household_head", "person_household_id") + missing_support = [column for column in required_person if column not in person] + if missing_support: + return GateResult( + name="housing_inputs_signal", + passed=False, + failures=(f"person support columns missing: {missing_support}.",), + details={"missing": missing_support}, + ) + + summary = us_housing_inputs_summary(frame) + failures: list[str] = [] + for key, label in ( + ( + "housing_assistance_receipt_missing_count", + "receives_housing_assistance missing values", + ), + ( + "housing_assistance_receipt_invalid_count", + "receives_housing_assistance invalid non-boolean values", + ), + ( + "housing_assistance_take_up_missing_count", + "takes_up_housing_assistance_if_eligible missing values", + ), + ( + "housing_assistance_take_up_invalid_count", + "takes_up_housing_assistance_if_eligible invalid non-boolean values", + ), + ): + count = int(summary[key]) + if count: + failures.append(f"{label}: {count}.") + take_up_mismatches = int( + summary["housing_assistance_take_up_source_mismatch_count"] + ) + if take_up_mismatches: + failures.append( + "takes_up_housing_assistance_if_eligible differs from measured " + f"receives_housing_assistance on {take_up_mismatches} SPM unit(s)." + ) + for key, label in ( + ("nonfinite_rent", "nonfinite pre_subsidy_rent"), + ("negative_rent", "negative pre_subsidy_rent"), + ("positive_rent_nonhead", "positive rent on non-head people"), + ): + count = int(summary[key]) + if count: + failures.append(f"{label}: {count}.") + for share_key, band_key, label in ( + ( + "pre_subsidy_rent_share", + "pre_subsidy_rent_share_band", + "positive pre-subsidy-rent share", + ), + ( + "housing_assistance_share", + "housing_assistance_share_band", + "housing-assistance receipt share", + ), + ): + share = float(summary[share_key]) + low, high = summary[band_key] + if not (low <= share <= high): + failures.append( + f"{label} {share:.3f} outside plausibility band [{low}, {high}]." + ) + from populace.build.us_runtime.puf_support import PUF_TAX_DETAIL_SUPPORT_CHANNEL + + channel_shares = summary["housing_assistance_share_by_support_channel"] + if summary["has_spm_support_channel_metadata"]: + if PUF_TAX_DETAIL_SUPPORT_CHANNEL not in channel_shares: + failures.append( + "housing assistance support metadata has no PUF-tax-detail channel." + ) + else: + puf_share = float(channel_shares[PUF_TAX_DETAIL_SUPPORT_CHANNEL]) + low, high = _HOUSING_ASSISTANCE_SHARE_BAND + if not (low <= puf_share <= high): + failures.append( + "PUF-tax-detail housing-assistance receipt share " + f"{puf_share:.3f} outside plausibility band [{low}, {high}]." + ) + household_values = summary["household_tenure_values"] + spm_values = summary["spm_tenure_values"] + if set(household_values) != _EXPECTED_HOUSEHOLD_TENURE_VALUES: + failures.append( + "household tenure_type does not carry the three locked ASEC " + f"categories: {household_values}." + ) + if set(spm_values) != _EXPECTED_SPM_TENURE_VALUES: + failures.append( + "spm_unit_tenure_type does not carry the three locked ASEC " + f"categories: {spm_values}." + ) + if summary["unknown_household_tenure_values"]: + failures.append( + "unknown household tenure value(s): " + f"{summary['unknown_household_tenure_values']}." + ) + if summary["unknown_spm_tenure_values"]: + failures.append( + f"unknown SPM tenure value(s): {summary['unknown_spm_tenure_values']}." + ) + return GateResult( + name="housing_inputs_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) diff --git a/packages/populace-build/src/populace/build/us_runtime/l0_refit_export.py b/packages/populace-build/src/populace/build/us_runtime/l0_refit_export.py index 77ece2cb..ba0ffe2e 100644 --- a/packages/populace-build/src/populace/build/us_runtime/l0_refit_export.py +++ b/packages/populace-build/src/populace/build/us_runtime/l0_refit_export.py @@ -13,12 +13,43 @@ import numpy as np from populace.build.gates import input_mass_parity_gate +from populace.build.us_runtime.alimony import US_ALIMONY_NONCONSTANT_PERSON_COLUMNS +from populace.build.us_runtime.capital_gain_details import ( + US_CAPITAL_GAIN_DETAILS_NONCONSTANT_PERSON_COLUMNS, + US_CAPITAL_GAIN_DETAILS_NONCONSTANT_TAX_UNIT_COLUMNS, +) +from populace.build.us_runtime.casualty_losses import ( + US_CASUALTY_LOSS_NONCONSTANT_PERSON_COLUMNS, +) +from populace.build.us_runtime.child_support import ( + US_CHILD_SUPPORT_NONCONSTANT_PERSON_COLUMNS, +) +from populace.build.us_runtime.childcare import US_CHILDCARE_OUTPUT_COLUMNS from populace.build.us_runtime.congressional_district_geography import ( CONGRESSIONAL_DISTRICT_GEOID_COLUMN, ) +from populace.build.us_runtime.disability_benefits import ( + US_DISABILITY_BENEFITS_NONCONSTANT_PERSON_COLUMNS, +) +from populace.build.us_runtime.domestic_production import ( + US_DOMESTIC_PRODUCTION_ALD_NONCONSTANT_TAX_UNIT_COLUMNS, +) +from populace.build.us_runtime.education_inputs import ( + US_EDUCATION_INPUTS_NONCONSTANT_PERSON_COLUMNS, +) +from populace.build.us_runtime.educator_expenses import ( + US_EDUCATOR_EXPENSE_NONCONSTANT_PERSON_COLUMNS, +) from populace.build.us_runtime.eligibility_inputs import ( US_ELIGIBILITY_INPUTS_NONCONSTANT_PERSON_COLUMNS, ) +from populace.build.us_runtime.energy_subsidy import US_ENERGY_SUBSIDY_OUTPUT_COLUMNS +from populace.build.us_runtime.farm_business_income import ( + US_FARM_BUSINESS_INCOME_NONCONSTANT_PERSON_COLUMNS, +) +from populace.build.us_runtime.form_4952 import ( + US_FORM_4952_NONCONSTANT_PERSON_COLUMNS, +) from populace.build.us_runtime.geography_ladder import ( US_GEOGRAPHY_LADDER_COLUMNS, us_geography_ladder_gate, @@ -26,34 +57,136 @@ from populace.build.us_runtime.hours_worked import ( US_HOURS_WORKED_NONCONSTANT_PERSON_COLUMNS, ) +from populace.build.us_runtime.housing_inputs import ( + US_HOUSING_NONCONSTANT_HOUSEHOLD_COLUMNS, + US_HOUSING_NONCONSTANT_PERSON_COLUMNS, + US_HOUSING_NONCONSTANT_SPM_UNIT_COLUMNS, +) from populace.build.us_runtime.immigration import ( US_IMMIGRATION_NONCONSTANT_PERSON_COLUMNS, ) from populace.build.us_runtime.input_mass import us_input_mass_totals +from populace.build.us_runtime.medicare_take_up import ( + US_MEDICARE_TAKE_UP_NONCONSTANT_PERSON_COLUMNS, +) +from populace.build.us_runtime.misc_itemized import ( + US_MISC_ITEMIZED_NONCONSTANT_PERSON_COLUMNS, +) +from populace.build.us_runtime.org_wages import ( + US_ORG_WAGES_NONCONSTANT_PERSON_COLUMNS, +) +from populace.build.us_runtime.other_health_insurance import ( + US_OTHER_HEALTH_INSURANCE_NONCONSTANT_PERSON_COLUMNS, +) from populace.build.us_runtime.pregnancy import ( US_PREGNANCY_NONCONSTANT_PERSON_COLUMNS, ) +from populace.build.us_runtime.prior_year_income import ( + US_PRIOR_YEAR_INCOME_NONCONSTANT_PERSON_COLUMNS, +) +from populace.build.us_runtime.qbi_inputs import ( + US_QBI_NONCONSTANT_PERSON_COLUMNS, +) +from populace.build.us_runtime.relationship_inputs import ( + US_RELATIONSHIP_INPUTS_NONCONSTANT_PERSON_COLUMNS, +) +from populace.build.us_runtime.retirement_contributions import ( + US_RETIREMENT_CONTRIBUTION_NONCONSTANT_PERSON_COLUMNS, +) +from populace.build.us_runtime.retirement_distributions import ( + US_RETIREMENT_DISTRIBUTION_NONCONSTANT_PERSON_COLUMNS, +) +from populace.build.us_runtime.salt_refund_income import ( + US_SALT_REFUND_NONCONSTANT_PERSON_COLUMNS, +) +from populace.build.us_runtime.scf_auto_loans import ( + US_SCF_AUTO_LOAN_NONCONSTANT_HOUSEHOLD_COLUMNS, +) from populace.build.us_runtime.scf_wealth import ( + US_SCF_WEALTH_NONCONSTANT_HOUSEHOLD_COLUMNS, US_SCF_WEALTH_NONCONSTANT_PERSON_COLUMNS, ) +from populace.build.us_runtime.sipp_head_start import ( + US_SIPP_HEAD_START_NONCONSTANT_PERSON_COLUMNS, +) +from populace.build.us_runtime.sipp_tips import ( + US_SIPP_TIPS_NONCONSTANT_PERSON_COLUMNS, +) +from populace.build.us_runtime.sipp_vehicles import ( + US_SIPP_VEHICLE_NONCONSTANT_HOUSEHOLD_COLUMNS, +) from populace.build.us_runtime.snap_discretionary_exemption import ( US_SNAP_DISCRETIONARY_EXEMPTION_NONCONSTANT_PERSON_COLUMNS, ) +from populace.build.us_runtime.ssi_disability_criteria import ( + US_SSI_DISABILITY_CRITERIA_NONCONSTANT_PERSON_COLUMNS, +) +from populace.build.us_runtime.ssi_take_up import ( + US_SSI_TAKE_UP_NONCONSTANT_PERSON_COLUMNS, +) +from populace.build.us_runtime.voluntary_filing import ( + US_VOLUNTARY_FILING_NONCONSTANT_TAX_UNIT_COLUMNS, +) +from populace.build.us_runtime.weeks_unemployed import ( + US_WEEKS_UNEMPLOYED_NONCONSTANT_PERSON_COLUMNS, +) +from populace.build.us_runtime.wic_claim import ( + US_WIC_CLAIM_NONCONSTANT_PERSON_COLUMNS, +) +from populace.build.us_runtime.workers_compensation import ( + US_WORKERS_COMPENSATION_NONCONSTANT_PERSON_COLUMNS, +) from populace.frame import US_SCHEMA, Frame, MassChange, WeightKind, Weights from populace.frame.adapters.policyengine_us import PolicyEngineUSEngine US_RELEASE_REQUIRED_TAX_UNIT_SOURCE_COLUMNS = ( "takes_up_aca_if_eligible", "selected_marketplace_plan_benchmark_ratio", + *US_DOMESTIC_PRODUCTION_ALD_NONCONSTANT_TAX_UNIT_COLUMNS, + *US_CAPITAL_GAIN_DETAILS_NONCONSTANT_TAX_UNIT_COLUMNS, + *US_VOLUNTARY_FILING_NONCONSTANT_TAX_UNIT_COLUMNS, ) US_RELEASE_REQUIRED_PERSON_SOURCE_COLUMNS = ( *US_IMMIGRATION_NONCONSTANT_PERSON_COLUMNS, *US_HOURS_WORKED_NONCONSTANT_PERSON_COLUMNS, *US_ELIGIBILITY_INPUTS_NONCONSTANT_PERSON_COLUMNS, + *US_RELATIONSHIP_INPUTS_NONCONSTANT_PERSON_COLUMNS, + *US_MEDICARE_TAKE_UP_NONCONSTANT_PERSON_COLUMNS, *US_PREGNANCY_NONCONSTANT_PERSON_COLUMNS, + *US_PRIOR_YEAR_INCOME_NONCONSTANT_PERSON_COLUMNS, *US_SNAP_DISCRETIONARY_EXEMPTION_NONCONSTANT_PERSON_COLUMNS, *US_SCF_WEALTH_NONCONSTANT_PERSON_COLUMNS, + *US_SSI_DISABILITY_CRITERIA_NONCONSTANT_PERSON_COLUMNS, + *US_SIPP_HEAD_START_NONCONSTANT_PERSON_COLUMNS, + *US_SSI_TAKE_UP_NONCONSTANT_PERSON_COLUMNS, + *US_SIPP_TIPS_NONCONSTANT_PERSON_COLUMNS, + *US_ALIMONY_NONCONSTANT_PERSON_COLUMNS, + *US_CASUALTY_LOSS_NONCONSTANT_PERSON_COLUMNS, + *US_CHILD_SUPPORT_NONCONSTANT_PERSON_COLUMNS, + *US_DISABILITY_BENEFITS_NONCONSTANT_PERSON_COLUMNS, + *US_WORKERS_COMPENSATION_NONCONSTANT_PERSON_COLUMNS, + *US_WEEKS_UNEMPLOYED_NONCONSTANT_PERSON_COLUMNS, + *US_WIC_CLAIM_NONCONSTANT_PERSON_COLUMNS, + *US_EDUCATOR_EXPENSE_NONCONSTANT_PERSON_COLUMNS, + *US_MISC_ITEMIZED_NONCONSTANT_PERSON_COLUMNS, + *US_EDUCATION_INPUTS_NONCONSTANT_PERSON_COLUMNS, + *US_RETIREMENT_CONTRIBUTION_NONCONSTANT_PERSON_COLUMNS, + *US_RETIREMENT_DISTRIBUTION_NONCONSTANT_PERSON_COLUMNS, + *US_QBI_NONCONSTANT_PERSON_COLUMNS, + *US_ORG_WAGES_NONCONSTANT_PERSON_COLUMNS, + *US_OTHER_HEALTH_INSURANCE_NONCONSTANT_PERSON_COLUMNS, + *US_FARM_BUSINESS_INCOME_NONCONSTANT_PERSON_COLUMNS, + *US_FORM_4952_NONCONSTANT_PERSON_COLUMNS, + *US_CAPITAL_GAIN_DETAILS_NONCONSTANT_PERSON_COLUMNS, + *US_SALT_REFUND_NONCONSTANT_PERSON_COLUMNS, + *US_HOUSING_NONCONSTANT_PERSON_COLUMNS, +) + +US_RELEASE_REQUIRED_SPM_UNIT_SOURCE_COLUMNS = ( + *US_CHILDCARE_OUTPUT_COLUMNS, + *US_ENERGY_SUBSIDY_OUTPUT_COLUMNS, + *US_HOUSING_NONCONSTANT_SPM_UNIT_COLUMNS, ) #: The geography spine a US release carries by default: state and district, @@ -66,6 +199,17 @@ *US_GEOGRAPHY_LADDER_COLUMNS, ) +# Unlike the geography spine above, these household inputs must carry signal. +# Presence-only would accept broadcast-zero engine defaults, silently zeroing +# the OBBBA auto-loan provision and erasing the restored vehicle/net-worth +# distributions while still satisfying a key-only export check. +US_RELEASE_REQUIRED_HOUSEHOLD_NONCONSTANT_SOURCE_COLUMNS = ( + *US_SCF_AUTO_LOAN_NONCONSTANT_HOUSEHOLD_COLUMNS, + *US_SCF_WEALTH_NONCONSTANT_HOUSEHOLD_COLUMNS, + *US_SIPP_VEHICLE_NONCONSTANT_HOUSEHOLD_COLUMNS, + *US_HOUSING_NONCONSTANT_HOUSEHOLD_COLUMNS, +) + @dataclass(frozen=True) class L0RefitWeights: @@ -268,12 +412,16 @@ def assert_required_us_release_source_columns( *, columns: tuple[str, ...] = US_RELEASE_REQUIRED_TAX_UNIT_SOURCE_COLUMNS, person_columns: tuple[str, ...] = US_RELEASE_REQUIRED_PERSON_SOURCE_COLUMNS, + spm_unit_columns: tuple[str, ...] = US_RELEASE_REQUIRED_SPM_UNIT_SOURCE_COLUMNS, household_columns: tuple[str, ...] = (US_RELEASE_REQUIRED_HOUSEHOLD_SOURCE_COLUMNS), + household_nonconstant_columns: tuple[str, ...] = ( + US_RELEASE_REQUIRED_HOUSEHOLD_NONCONSTANT_SOURCE_COLUMNS + ), ) -> None: """Require source-stage columns needed by US release gates. - Tax-unit columns come from the ACA Marketplace source stage; person - columns are the SSN/immigration surface (a missing or constant + Tax-unit columns come from the ACA Marketplace and PUF tax-detail stages; + person columns are the SSN/immigration surface (a missing or constant ``ssn_card_type`` reproduces the everyone-is-a-citizen failure of populace issue #225); household columns are the geography spine (a release without the block-anchored ladder of populace #275 cannot be @@ -286,7 +434,9 @@ def assert_required_us_release_source_columns( for entity, required, check_nonconstant in ( ("tax_unit", columns, True), ("person", person_columns, True), + ("spm_unit", spm_unit_columns, True), ("household", household_columns, False), + ("household", household_nonconstant_columns, True), ): table = frame.table(entity) for column in required: @@ -296,7 +446,27 @@ def assert_required_us_release_source_columns( if not check_nonconstant: continue unique = table[column].dropna().unique() - if len(unique) < 2: + # A one-household test/export can still carry real (non-default) + # auto-loan signal even though two distinct values are impossible. + # For normal multi-household releases these columns must be truly + # nonconstant; an all-zero broadcast always fails. + single_nondefault_auto_value = ( + entity == "household" + and column in household_nonconstant_columns + and len(table) == 1 + and len(unique) == 1 + and bool(unique[0]) + ) + single_nondefault_spm_value = ( + entity == "spm_unit" + and column in spm_unit_columns + and len(table) == 1 + and len(unique) == 1 + and bool(unique[0]) + ) + if len(unique) < 2 and not ( + single_nondefault_auto_value or single_nondefault_spm_value + ): failures.append(f"{entity}.{column}: not nonconstant") if failures: raise ValueError( @@ -434,9 +604,15 @@ def export_us_l0_refit_h5( "required_person_source_columns": list( US_RELEASE_REQUIRED_PERSON_SOURCE_COLUMNS ), + "required_spm_unit_source_columns": list( + US_RELEASE_REQUIRED_SPM_UNIT_SOURCE_COLUMNS + ), "required_household_source_columns": list( US_RELEASE_REQUIRED_HOUSEHOLD_SOURCE_COLUMNS ), + "required_household_nonconstant_source_columns": list( + US_RELEASE_REQUIRED_HOUSEHOLD_NONCONSTANT_SOURCE_COLUMNS + ), "required_source_columns_checked": bool(require_source_columns), "geography_ladder_gate_enforced": bool(require_geography_ladder), "geography_ladder_gate": { @@ -582,7 +758,9 @@ def main(argv: list[str] | None = None) -> None: __all__ = [ "L0RefitWeights", "US_RELEASE_REQUIRED_HOUSEHOLD_SOURCE_COLUMNS", + "US_RELEASE_REQUIRED_HOUSEHOLD_NONCONSTANT_SOURCE_COLUMNS", "US_RELEASE_REQUIRED_PERSON_SOURCE_COLUMNS", + "US_RELEASE_REQUIRED_SPM_UNIT_SOURCE_COLUMNS", "US_RELEASE_REQUIRED_TAX_UNIT_SOURCE_COLUMNS", "attach_l0_refit_entity_weights", "attach_l0_refit_weights", diff --git a/packages/populace-build/src/populace/build/us_runtime/medicaid_take_up.py b/packages/populace-build/src/populace/build/us_runtime/medicaid_take_up.py index 82fa0681..f8d20176 100644 --- a/packages/populace-build/src/populace/build/us_runtime/medicaid_take_up.py +++ b/packages/populace-build/src/populace/build/us_runtime/medicaid_take_up.py @@ -344,6 +344,7 @@ def us_medicaid_take_up_diagnostics( state_targets: pd.DataFrame, *, substitutions: Sequence[dict[str, object]] = (), + weights_basis: str = "pre_calibration_design_weights", ) -> dict[str, object]: """The #170 eligibility-to-enrollment surface, by state and nationally. @@ -416,11 +417,7 @@ def us_medicaid_take_up_diagnostics( "anchor": US_MEDICAID_TAKE_UP_ANCHOR, "enrollment_semantics": "point_in_time_monthly_snapshot", "target_table": US_MEDICAID_ENROLLMENT_TARGET_TABLE, - # Stage-time surface: weights are the pre-calibration design weights. - # Post-calibration weighted enrollment is pulled to the same CMS - # counts by the medicaid_enrollment weight-calibration targets, but - # eligible/anchored weights here have no post-calibration counterpart. - "weights_basis": "pre_calibration_design_weights", + "weights_basis": str(weights_basis), "national": { "eligible_weight": total_eligible, "enrolled_weight": total_enrolled, diff --git a/packages/populace-build/src/populace/build/us_runtime/medicare_take_up.py b/packages/populace-build/src/populace/build/us_runtime/medicare_take_up.py new file mode 100644 index 00000000..ae6e87e8 --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/medicare_take_up.py @@ -0,0 +1,303 @@ +"""Measured CPS ASEC Medicare enrollment exported as the take-up input leaf. + +The retired eCPS pipeline maps ``MCARE == 1`` to the formula-level +``medicare_enrolled`` value, duplicates non-PUF-imputed CPS variables onto the +PUF support half, and finally renames that measured value to the input leaf +``takes_up_medicare_if_eligible``. This module restores the same source-backed +boolean directly, without applying a participation-rate draw. +""" + +from __future__ import annotations + +from importlib.resources import files + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import ( + SourceOperationSpec, + SourceStageSpec, + load_source_manifest, +) +from populace.build.source_runtime import ( + SourceRuntimeConfig, + SourceRuntimeContext, + SourceRuntimeError, + run_source_stage, +) +from populace.frame import Frame +from populace.frame.units import US_SCHEMA + +__all__ = [ + "MEDICARE_TAKE_UP_ARCHIVED_CLONE_URL", + "MEDICARE_TAKE_UP_ARCHIVED_DERIVATION_URL", + "MEDICARE_TAKE_UP_ARCHIVED_EXPORT_URL", + "MEDICARE_TAKE_UP_ARCHIVED_SOURCE_COLUMNS_URL", + "US_MEDICARE_TAKE_UP_NONCONSTANT_PERSON_COLUMNS", + "US_MEDICARE_TAKE_UP_OUTPUT_COLUMNS", + "US_MEDICARE_TAKE_UP_REQUIRED_SOURCE_COLUMNS", + "US_MEDICARE_TAKE_UP_STAGE_NAME", + "derive_us_medicare_take_up_from_manifest", + "us_medicare_take_up_signal_gate", + "us_medicare_take_up_stage_spec", + "us_medicare_take_up_summary", + "with_us_medicare_take_up_input", +] + +_ARCHIVED_DATA_REPOSITORY = "policyengine-" + "us-data" +_ARCHIVED_ROOT = ( + "https://github.com/PolicyEngine/" + f"{_ARCHIVED_DATA_REPOSITORY}/blob/" + "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe/" + "policyengine_" + "us_data/" +) +MEDICARE_TAKE_UP_ARCHIVED_DERIVATION_URL = ( + _ARCHIVED_ROOT + "datasets/cps/cps.py#L1579-L1585" +) +MEDICARE_TAKE_UP_ARCHIVED_SOURCE_COLUMNS_URL = ( + _ARCHIVED_ROOT + "datasets/cps/census_cps.py#L39-L58" +) +MEDICARE_TAKE_UP_ARCHIVED_CLONE_URL = ( + _ARCHIVED_ROOT + "calibration/puf_impute.py#L608-L629" +) +MEDICARE_TAKE_UP_ARCHIVED_EXPORT_URL = ( + _ARCHIVED_ROOT + "datasets/cps/extended_cps.py#L1747-L1754" +) + +US_MEDICARE_TAKE_UP_STAGE_NAME = "medicare_take_up_input" +US_MEDICARE_TAKE_UP_OUTPUT_COLUMNS: tuple[str, ...] = ("takes_up_medicare_if_eligible",) +US_MEDICARE_TAKE_UP_NONCONSTANT_PERSON_COLUMNS = US_MEDICARE_TAKE_UP_OUTPUT_COLUMNS +US_MEDICARE_TAKE_UP_REQUIRED_SOURCE_COLUMNS: tuple[str, ...] = ("MCARE",) + +_OUTPUT = US_MEDICARE_TAKE_UP_OUTPUT_COLUMNS[0] +_SOURCE = US_MEDICARE_TAKE_UP_REQUIRED_SOURCE_COLUMNS[0] +_PERSON_WEIGHT_COLUMN = "person_weight" +_PERSON_SUPPORT_CHANNEL_COLUMN = "person_support_channel" +_VALID_SOURCE_CODES = frozenset({0, 1, 2}) +_ENROLLED_CODE = 1 +_EXPECTED_PARAMETERS = { + "source": _SOURCE, + "enrolled_code": _ENROLLED_CODE, + "output": _OUTPUT, +} +_WEIGHTED_ENROLLED_SHARE_BAND = (0.15, 0.24) +_CHANNEL_ENROLLED_SHARE_BAND = (0.15, 0.24) + + +def us_medicare_take_up_stage_spec() -> SourceStageSpec: + """Load and validate the packaged Medicare take-up stage.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + stage_map = manifest.stage_map() + if US_MEDICARE_TAKE_UP_STAGE_NAME not in stage_map: + raise ValueError( + f"US source manifest declares no {US_MEDICARE_TAKE_UP_STAGE_NAME!r} stage." + ) + spec = stage_map[US_MEDICARE_TAKE_UP_STAGE_NAME] + if tuple(spec.outputs) != US_MEDICARE_TAKE_UP_OUTPUT_COLUMNS: + raise ValueError( + f"{US_MEDICARE_TAKE_UP_STAGE_NAME!r} manifest outputs do not " + "match the runtime-owned Medicare take-up family." + ) + return spec + + +def _source_codes(person: pd.DataFrame, source: str) -> np.ndarray: + if source not in person.columns: + raise SourceRuntimeError( + f"US Medicare take-up derivation requires ASEC source column {source!r}." + ) + numeric = pd.to_numeric(person[source], errors="coerce").to_numpy(dtype=np.float64) + valid = ( + np.isfinite(numeric) + & (numeric == np.floor(numeric)) + & np.isin(numeric, np.fromiter(_VALID_SOURCE_CODES, dtype=np.int64)) + ) + if not valid.all(): + rows = np.flatnonzero(~valid)[:5].tolist() + raise SourceRuntimeError( + f"US Medicare take-up source {source!r} must contain only integer " + f"codes {sorted(_VALID_SOURCE_CODES)}; invalid row(s): {rows}." + ) + return numeric.astype(np.int8) + + +def derive_us_medicare_take_up_from_manifest( + frame: pd.DataFrame | None, + operation: SourceOperationSpec, + _context: SourceRuntimeContext | None, +) -> pd.DataFrame: + """Map the exact retired ``MCARE == 1`` enrollment observation.""" + + if operation.kind != "derive_medicare_take_up": + raise SourceRuntimeError( + "US Medicare take-up derivation received unexpected operation " + f"{operation.kind!r}." + ) + if frame is None: + raise SourceRuntimeError( + "US Medicare take-up derivation requires the person table first." + ) + parameters = dict(operation.parameters) + if parameters != _EXPECTED_PARAMETERS: + raise SourceRuntimeError( + "US Medicare take-up derivation drifted from the archived method: " + f"expected {_EXPECTED_PARAMETERS}, got {parameters}." + ) + result = frame.copy(deep=True) + result[_OUTPUT] = _source_codes(result, _SOURCE) == _ENROLLED_CODE + return result + + +def _surface_matches_source(frame: Frame) -> bool: + person = frame.table("person") + if _OUTPUT not in person or person[_OUTPUT].dropna().nunique() <= 1: + return False + if _SOURCE not in person: + return True + try: + expected = _source_codes(person, _SOURCE) == _ENROLLED_CODE + except SourceRuntimeError: + return False + observed = person[_OUTPUT].fillna(False).astype(bool).to_numpy() + return bool(np.array_equal(observed, expected)) + + +def with_us_medicare_take_up_input( + frame: Frame, + *, + seed: int, + time_period: int, +) -> Frame: + """Materialize the measured Medicare take-up leaf on a US frame.""" + + if frame.schema != US_SCHEMA: + raise ValueError("US Medicare take-up input requires the US schema.") + if _surface_matches_source(frame): + return frame + + person = frame.table("person") + stage_person = person.copy(deep=True) + stage_person[_PERSON_WEIGHT_COLUMN] = frame.resolve_weights("person").values + output = run_source_stage( + us_medicare_take_up_stage_spec(), + tables={"person": stage_person}, + operation_handlers={ + "derive_medicare_take_up": derive_us_medicare_take_up_from_manifest + }, + config=SourceRuntimeConfig(seed=int(seed), target_year=int(time_period)), + ) + aligned = output.set_index("person_id").reindex(person["person_id"]) + if aligned[_OUTPUT].isna().any(): + raise ValueError( + "US Medicare take-up stage output does not cover every person." + ) + + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + tables["person"][_OUTPUT] = aligned[_OUTPUT].to_numpy(dtype=bool) + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + ) + + +def us_medicare_take_up_summary(frame: Frame) -> dict[str, object]: + """Return weighted signal and exact source-reconciliation diagnostics.""" + + person = frame.table("person") + weights = np.asarray(frame.resolve_weights("person").values, dtype=np.float64) + values = person[_OUTPUT].fillna(False).astype(bool).to_numpy() + total_weight = float(weights.sum()) + summary: dict[str, object] = { + "weighted_enrolled_share": ( + float(weights[values].sum()) / total_weight if total_weight > 0 else 0.0 + ), + "weighted_enrolled_share_band": list(_WEIGHTED_ENROLLED_SHARE_BAND), + "positive_count": int(np.count_nonzero(values)), + "unique_count": int(person[_OUTPUT].dropna().nunique()), + "missing_count": int(person[_OUTPUT].isna().sum()), + } + if _SOURCE in person: + source_values = _source_codes(person, _SOURCE) == _ENROLLED_CODE + summary["source_mismatch_count"] = int( + np.count_nonzero(values != source_values) + ) + if _PERSON_SUPPORT_CHANNEL_COLUMN in person: + channel_shares: dict[str, float] = {} + channels = ( + person[_PERSON_SUPPORT_CHANNEL_COLUMN] + .fillna("") + .astype(str) + .to_numpy() + ) + for channel in sorted(set(channels.tolist())): + mask = channels == channel + channel_weight = float(weights[mask].sum()) + channel_shares[str(channel)] = ( + float(weights[mask & values].sum()) / channel_weight + if channel_weight > 0 + else 0.0 + ) + summary["channel_weighted_enrolled_shares"] = channel_shares + summary["channel_weighted_enrolled_share_band"] = list( + _CHANNEL_ENROLLED_SHARE_BAND + ) + return summary + + +def us_medicare_take_up_signal_gate(frame: Frame) -> GateResult: + """Require measured, nonconstant, source-consistent Medicare enrollment.""" + + person = frame.table("person") + if _OUTPUT not in person: + return GateResult( + name="medicare_take_up_input_signal", + passed=False, + failures=(f"{_OUTPUT}: missing",), + details={}, + ) + try: + summary = us_medicare_take_up_summary(frame) + except SourceRuntimeError as exc: + return GateResult( + name="medicare_take_up_input_signal", + passed=False, + failures=(str(exc),), + details={}, + ) + failures: list[str] = [] + if int(summary["missing_count"]): + failures.append(f"{_OUTPUT}: missing values") + if int(summary["unique_count"]) < 2: + failures.append(f"{_OUTPUT}: constant") + share = float(summary["weighted_enrolled_share"]) + low, high = _WEIGHTED_ENROLLED_SHARE_BAND + if not low <= share <= high: + failures.append( + f"{_OUTPUT}: weighted enrolled share {share:.6f} outside " + f"[{low:.3f}, {high:.3f}]" + ) + mismatches = int(summary.get("source_mismatch_count", 0)) + if mismatches: + failures.append(f"{_OUTPUT}: {mismatches} MCARE reconciliation mismatch(es)") + channel_shares = summary.get("channel_weighted_enrolled_shares", {}) + channel_low, channel_high = _CHANNEL_ENROLLED_SHARE_BAND + for channel, channel_share in dict(channel_shares).items(): + if not channel_low <= float(channel_share) <= channel_high: + failures.append( + f"{_OUTPUT}: {channel} weighted enrolled share " + f"{float(channel_share):.6f} outside " + f"[{channel_low:.3f}, {channel_high:.3f}]" + ) + return GateResult( + name="medicare_take_up_input_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) diff --git a/packages/populace-build/src/populace/build/us_runtime/misc_itemized.py b/packages/populace-build/src/populace/build/us_runtime/misc_itemized.py new file mode 100644 index 00000000..947a0851 --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/misc_itemized.py @@ -0,0 +1,172 @@ +"""IRS PUF miscellaneous-itemized input for the US build. + +The retired eCPS pipeline used IRS PUF ``E20400`` (unreimbursed employee +business expenses) as its explicit proxy for all expenses eligible for the +miscellaneous itemized deduction. The immutable archived source coordinate is +exposed as ``MISC_ITEMIZED_ARCHIVED_DERIVATION_URL`` below. + +There is no CPS amount and no synthetic fallback. The PUF tax-detail stage +carries the source exactly, then its existing weighted QRF places that +source-backed detail on the dedicated PUF support channel. PolicyEngine-US owns +the two-percent-of-AGI floor and deduction formula; this module persists only +the factual person input. +""" + +from __future__ import annotations + +from importlib.resources import files + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import SourceStageSpec, load_source_manifest +from populace.frame import Frame + +__all__ = [ + "MISC_ITEMIZED_ARCHIVED_DERIVATION_URL", + "US_MISC_ITEMIZED_NONCONSTANT_PERSON_COLUMNS", + "US_MISC_ITEMIZED_OUTPUT_COLUMNS", + "US_MISC_ITEMIZED_STAGE_NAME", + "derive_us_misc_itemized_from_puf", + "us_misc_itemized_signal_gate", + "us_misc_itemized_stage_spec", + "us_misc_itemized_summary", +] + +_ARCHIVED_DATA_REPOSITORY = "policyengine-" + "us-data" +MISC_ITEMIZED_ARCHIVED_DERIVATION_URL = ( + "https://github.com/PolicyEngine/" + f"{_ARCHIVED_DATA_REPOSITORY}/blob/" + "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe/" + "policyengine_" + "us_data/datasets/puf/puf.py#L663-L665" +) + +US_MISC_ITEMIZED_STAGE_NAME = "puf_tax_detail" +US_MISC_ITEMIZED_OUTPUT_COLUMNS: tuple[str, ...] = ( + "unreimbursed_business_employee_expenses", +) +US_MISC_ITEMIZED_NONCONSTANT_PERSON_COLUMNS = US_MISC_ITEMIZED_OUTPUT_COLUMNS + +_MISC_ITEMIZED_NONZERO_SHARE_BAND = (0.05, 0.55) + + +def us_misc_itemized_stage_spec() -> SourceStageSpec: + """Load and validate the shared PUF tax-detail stage declaration.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + spec = manifest.stage_map()[US_MISC_ITEMIZED_STAGE_NAME] + missing = sorted(set(US_MISC_ITEMIZED_OUTPUT_COLUMNS) - set(spec.outputs)) + if missing: + raise ValueError( + f"{US_MISC_ITEMIZED_STAGE_NAME!r} manifest stage does not declare " + f"miscellaneous-itemized output(s) {missing}." + ) + return spec + + +def derive_us_misc_itemized_from_puf( + puf: pd.DataFrame, + *, + source_column: str = "E20400", + output_column: str = "unreimbursed_business_employee_expenses", +) -> pd.DataFrame: + """Carry the archived IRS PUF miscellaneous-expense proxy exactly. + + Invalid or negative source values fail closed. Replacing them with zeros + would invent observations and could silently erase the reform input this + stage exists to restore. + """ + + if source_column not in puf.columns: + raise ValueError( + f"PUF miscellaneous-itemized derivation requires source column " + f"{source_column!r}." + ) + values = pd.to_numeric(puf[source_column], errors="coerce") + numeric = values.to_numpy(dtype=np.float64) + nonfinite = ~np.isfinite(numeric) + if bool(nonfinite.any()): + raise ValueError( + f"PUF miscellaneous-itemized source column {source_column!r} contains " + f"{int(np.count_nonzero(nonfinite))} nonnumeric or nonfinite value(s)." + ) + negative = numeric < 0.0 + if bool(negative.any()): + raise ValueError( + f"PUF miscellaneous-itemized source column {source_column!r} contains " + f"{int(np.count_nonzero(negative))} negative value(s)." + ) + + result = puf.copy(deep=True) + result[output_column] = numeric + return result + + +def us_misc_itemized_summary(frame: Frame) -> dict[str, object]: + """Return weighted signal and validity diagnostics for the proxy input.""" + + person = frame.table("person") + values = pd.to_numeric( + person["unreimbursed_business_employee_expenses"], errors="coerce" + ).to_numpy(dtype=np.float64) + weights = np.asarray(frame.resolve_weights("person").values, dtype=np.float64) + finite = np.isfinite(values) + positive = finite & (values > 0.0) + total_weight = float(weights.sum()) + positive_share = ( + float(weights[positive].sum()) / total_weight if total_weight > 0.0 else 0.0 + ) + return { + "positive_share": positive_share, + "positive_share_band": list(_MISC_ITEMIZED_NONZERO_SHARE_BAND), + "unweighted_total": float(np.nansum(values)), + "nonfinite": int(np.count_nonzero(~finite)), + "negative": int(np.count_nonzero(finite & (values < 0.0))), + } + + +def us_misc_itemized_signal_gate(frame: Frame) -> GateResult: + """Require a finite, nonnegative, nondefault miscellaneous-expense input.""" + + person = frame.table("person") + missing = [ + column + for column in US_MISC_ITEMIZED_OUTPUT_COLUMNS + if column not in person.columns + ] + if missing: + return GateResult( + name="misc_itemized_signal", + passed=False, + failures=(f"person columns missing: {missing}.",), + details={"missing": missing}, + ) + + summary = us_misc_itemized_summary(frame) + failures: list[str] = [] + if summary["nonfinite"]: + failures.append( + "unreimbursed_business_employee_expenses nonfinite values: " + f"{int(summary['nonfinite'])}." + ) + if summary["negative"]: + failures.append( + "unreimbursed_business_employee_expenses negative values: " + f"{int(summary['negative'])}." + ) + share = float(summary["positive_share"]) + low, high = summary["positive_share_band"] + if not (low <= share <= high): + failures.append( + f"miscellaneous-itemized positive share {share:.6f} outside " + f"plausibility band [{low}, {high}]." + ) + return GateResult( + name="misc_itemized_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) diff --git a/packages/populace-build/src/populace/build/us_runtime/org_wages.py b/packages/populace-build/src/populace/build/us_runtime/org_wages.py new file mode 100644 index 00000000..71d5fc48 --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/org_wages.py @@ -0,0 +1,905 @@ +"""CPS ORG labor-market inputs and the FLSA overtime-premium proxy. + +This stage ports the final retired eCPS implementation at +``9a823603e6b5fb916d65ec45d74c9c7eb0043db1``. It builds a person-level +training donor from the twelve official 2024 CPS basic-month outgoing- +rotation-group files, fits the weighted Populace QRF for hourly wage and +hourly-pay status, assigns union coverage from the published 2024 BLS state +rates, carries the ASEC occupation fields, and derives the intentionally +misspelled PolicyEngine input ``fsla_overtime_premium``. + +The upstream implementation cached the transformed monthly files as +``census_cps_org_2024_wages.csv.gz`` but did not publish that cache. Populace +pins the canonical *uncompressed CSV content* of the generated 119,237-row +donor. A cached file is accepted only when that digest matches; rebuilding +from a silently reissued Census file therefore fails rather than changing a +release. +""" + +from __future__ import annotations + +import gzip +import hashlib +import io +import urllib.request +import zipfile +from importlib.resources import files +from pathlib import Path + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import SourceStageSpec, load_source_manifest +from populace.frame import Frame +from populace.frame.units import US_SCHEMA + +__all__ = [ + "BLS_STATE_UNION_REPRESENTATION_RATE_2024", + "FLSA_EXECUTIVE_ADMINISTRATIVE_PROFESSIONAL_OCCUPATION_CODES", + "FLSA_OVERTIME_OCCUPATION_CODES", + "ORG_2024_DONOR_CONTENT_SHA256", + "ORG_2024_DONOR_FILENAME", + "ORG_PREDICTORS", + "US_ORG_WAGES_NONCONSTANT_PERSON_COLUMNS", + "US_ORG_WAGES_OUTPUT_COLUMNS", + "US_ORG_WAGES_REQUIRED_SOURCE_COLUMNS", + "US_ORG_WAGES_STAGE_NAME", + "derive_flsa_overtime_premium", + "derive_us_org_occupation_inputs", + "fetch_org_2024_donor", + "impute_us_org_wages", + "load_org_2024_donor", + "us_org_wages_signal_gate", + "us_org_wages_stage_spec", + "us_org_wages_summary", + "with_us_org_wages_inputs", +] + +US_ORG_WAGES_STAGE_NAME = "org_wages" +ORG_2024_DONOR_FILENAME = "census_cps_org_2024_wages.csv.gz" +ORG_2024_DONOR_CONTENT_SHA256 = ( + "d74600236cfdd34033d487cd9d82f6eb00b1858ba28de17f4f985f2aec516f86" +) +ORG_2024_DONOR_COMPRESSED_SHA256 = ( + "66fa5b6aa4087413b691038767b51f603281ff55411b58259922f78e67460372" +) +ORG_YEAR = 2024 +ORG_MONTHS = ( + "jan", + "feb", + "mar", + "apr", + "may", + "jun", + "jul", + "aug", + "sep", + "oct", + "nov", + "dec", +) + +ORG_PREDICTORS: tuple[str, ...] = ( + "employment_income", + "weekly_hours_worked", + "age", + "is_female", + "is_hispanic", + "race_wbho", + "state_fips", +) +ORG_QRF_TARGETS: tuple[str, ...] = ("hourly_wage", "is_paid_hourly") +_DONOR_WEIGHT_COLUMN = "sample_weight" +_DEFAULT_N_ESTIMATORS = 100 + +US_ORG_WAGES_OUTPUT_COLUMNS: tuple[str, ...] = ( + "cps_race", + "is_hispanic", + "detailed_occupation_recode", + "has_never_worked", + "is_military", + "is_computer_scientist", + "is_executive_administrative_professional", + "is_farmer_fisher", + "hourly_wage", + "is_paid_hourly", + "is_union_member_or_covered", + "fsla_overtime_premium", +) +US_ORG_WAGES_NONCONSTANT_PERSON_COLUMNS = US_ORG_WAGES_OUTPUT_COLUMNS +US_ORG_WAGES_REQUIRED_SOURCE_COLUMNS: tuple[str, ...] = ( + "person_household_id", + "age", + "is_female", + "PRDTRACE", + "PRDTHSP", + "POCCU2", + "employment_income_before_lsr", + "weekly_hours_worked_before_lsr", + "hours_worked_last_week", + "weeks_worked", +) + +FLSA_OVERTIME_OCCUPATION_CODES: dict[str, int] = { + "has_never_worked": 53, + "is_military": 52, + "is_computer_scientist": 8, + "is_farmer_fisher": 41, +} +FLSA_EXECUTIVE_ADMINISTRATIVE_PROFESSIONAL_OCCUPATION_CODES = frozenset( + ( + 1, + 2, + 3, + 5, + 6, + 7, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 18, + 19, + 25, + 26, + 27, + 28, + 29, + 34, + 36, + 38, + 39, + 40, + 42, + 50, + ) +) + +_CPS_BASIC_MONTHLY_ORG_COLUMNS = ( + "HRMIS", + "gestfips", + "prtage", + "pesex", + "ptdtrace", + "pehspnon", + "pworwgt", + "pternwa", + "pternhly", + "peernhry", + "pehruslt", + "prerelg", + "pemlr", + "peio1cow", +) +_CPS_BASIC_MONTHLY_ORG_FWF_COLUMNS = ( + "HRMIS", + "gestfips", + "prtage", + "pesex", + "ptdtrace", + "pehspnon", + "pemlr", + "pehruslt", + "peio1cow", + "prerelg", + "peernhry", + "pternhly", + "pternwa", + "pworwgt", +) +_CPS_BASIC_MONTHLY_ORG_COLSPECS = ( + (62, 64), + (92, 94), + (121, 123), + (128, 130), + (138, 140), + (156, 158), + (179, 181), + (223, 226), + (430, 433), + (497, 499), + (505, 507), + (519, 523), + (526, 534), + (602, 612), +) + +BLS_UNION_REPRESENTATION_RATE_2024 = np.float32(0.111) +_BLS_SEX_RATES = {False: np.float32(0.113), True: np.float32(0.108)} +_BLS_AGE_RATES = { + (16, 24): np.float32(0.054), + (25, 34): np.float32(0.099), + (35, 44): np.float32(0.122), + (45, 54): np.float32(0.138), + (55, 64): np.float32(0.130), + (65, 150): np.float32(0.102), +} +_BLS_RACE_RATES = { + 1: np.float32(0.108), + 2: np.float32(0.132), + 3: np.float32(0.097), + 4: BLS_UNION_REPRESENTATION_RATE_2024, +} +_BLS_FULL_TIME_RATE = np.float32(0.120) +_BLS_PART_TIME_RATE = np.float32(0.066) +BLS_STATE_UNION_REPRESENTATION_RATE_2024: dict[int, np.float32] = { + 1: np.float32(0.078), + 2: np.float32(0.195), + 4: np.float32(0.045), + 5: np.float32(0.044), + 6: np.float32(0.163), + 8: np.float32(0.080), + 9: np.float32(0.178), + 10: np.float32(0.089), + 11: np.float32(0.117), + 12: np.float32(0.063), + 13: np.float32(0.044), + 15: np.float32(0.275), + 16: np.float32(0.059), + 17: np.float32(0.142), + 18: np.float32(0.104), + 19: np.float32(0.083), + 20: np.float32(0.080), + 21: np.float32(0.112), + 22: np.float32(0.050), + 23: np.float32(0.153), + 24: np.float32(0.134), + 25: np.float32(0.156), + 26: np.float32(0.147), + 27: np.float32(0.148), + 28: np.float32(0.079), + 29: np.float32(0.093), + 30: np.float32(0.131), + 31: np.float32(0.081), + 32: np.float32(0.134), + 33: np.float32(0.106), + 34: np.float32(0.174), + 35: np.float32(0.088), + 36: np.float32(0.219), + 37: np.float32(0.031), + 38: np.float32(0.063), + 39: np.float32(0.133), + 40: np.float32(0.062), + 41: np.float32(0.175), + 42: np.float32(0.124), + 44: np.float32(0.153), + 45: np.float32(0.041), + 46: np.float32(0.037), + 47: np.float32(0.056), + 48: np.float32(0.054), + 49: np.float32(0.078), + 50: np.float32(0.158), + 51: np.float32(0.057), + 53: np.float32(0.183), + 54: np.float32(0.100), + 55: np.float32(0.069), + 56: np.float32(0.067), +} + +_SHARE_BANDS: dict[str, tuple[float, float]] = { + "cps_race": (0.95, 1.0), + "is_hispanic": (0.05, 0.30), + "detailed_occupation_recode": (0.65, 0.95), + "has_never_worked": (0.15, 0.35), + "is_military": (0.0005, 0.01), + "is_computer_scientist": (0.005, 0.05), + "is_executive_administrative_professional": (0.20, 0.40), + "is_farmer_fisher": (0.001, 0.015), + "hourly_wage": (0.40, 0.70), + "is_paid_hourly": (0.15, 0.40), + "is_union_member_or_covered": (0.02, 0.12), + "fsla_overtime_premium": (0.01, 0.15), +} + + +def us_org_wages_stage_spec() -> SourceStageSpec: + """Load and validate the packaged ``org_wages`` declaration.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + spec = manifest.stage_map().get(US_ORG_WAGES_STAGE_NAME) + if spec is None: + raise ValueError("US source manifest declares no 'org_wages' stage.") + missing = sorted(set(US_ORG_WAGES_OUTPUT_COLUMNS) - set(spec.outputs)) + if missing: + raise ValueError( + f"'org_wages' manifest stage does not declare output(s) {missing}." + ) + return spec + + +def _sha256(payload: bytes) -> str: + return hashlib.sha256(payload).hexdigest() + + +def _canonical_donor_bytes(path: str | Path) -> bytes: + source = Path(path) + if source.suffix == ".gz": + return gzip.decompress(source.read_bytes()) + return source.read_bytes() + + +def _month_url(month: str, suffix: str) -> str: + return ( + "https://www2.census.gov/programs-surveys/cps/datasets/" + f"{ORG_YEAR}/basic/{month}24pub.{suffix}" + ) + + +def _normalized_month(frame: pd.DataFrame) -> pd.DataFrame: + lookup = {str(column).lower(): column for column in frame.columns} + missing = [ + column + for column in _CPS_BASIC_MONTHLY_ORG_COLUMNS + if column.lower() not in lookup + ] + if missing: + raise ValueError(f"CPS basic ORG month missing column(s): {missing}.") + selected = frame[ + [lookup[column.lower()] for column in _CPS_BASIC_MONTHLY_ORG_COLUMNS] + ].copy() + selected.columns = _CPS_BASIC_MONTHLY_ORG_COLUMNS + return selected.apply(pd.to_numeric, errors="coerce") + + +def _load_month_from_network(month: str) -> pd.DataFrame: + """Load one official month, preferring CSV and falling back to its ZIP.""" + + csv_error: Exception | None = None + try: + with urllib.request.urlopen( # noqa: S310 + _month_url(month, "csv"), timeout=60 + ) as response: + return _normalized_month(pd.read_csv(io.BytesIO(response.read()))) + except Exception as error: # pragma: no cover - live source fallback + csv_error = error + try: + with urllib.request.urlopen( # noqa: S310 + _month_url(month, "zip"), timeout=60 + ) as response: + payload = response.read() + with zipfile.ZipFile(io.BytesIO(payload)) as archive: + member = next( + name for name in archive.namelist() if name.lower().endswith(".dat") + ) + with archive.open(member) as stream: + frame = pd.read_fwf( + stream, + colspecs=_CPS_BASIC_MONTHLY_ORG_COLSPECS, + names=_CPS_BASIC_MONTHLY_ORG_FWF_COLUMNS, + header=None, + ) + return _normalized_month(frame) + except Exception as zip_error: # pragma: no cover - live source failure + raise ValueError( + f"Could not load CPS ORG {month} {ORG_YEAR} from CSV or ZIP: " + f"csv={csv_error!r}; zip={zip_error!r}." + ) from zip_error + + +def _derive_wbho(cps_race: np.ndarray, is_hispanic: np.ndarray) -> np.ndarray: + hispanic = np.asarray(is_hispanic, dtype=bool) + race = np.asarray(cps_race) + return np.select( + [hispanic, (race == 1) & ~hispanic, (race == 2) & ~hispanic], + [3, 1, 2], + default=4, + ).astype(np.float32) + + +def _transform_month(raw: pd.DataFrame) -> pd.DataFrame: + org = raw.loc[ + raw["HRMIS"].isin([4, 8]) + & (raw["pworwgt"] > 0) + & (raw["prerelg"] == 1) + & raw["pemlr"].isin([1, 2]) + & raw["peio1cow"].isin([1, 2, 3, 4, 5]) + & raw["peernhry"].isin([1, 2]) + & (raw["gestfips"] > 0) + & (raw["prtage"] >= 16) + & (raw["pehruslt"] > 0) + & (raw["pternwa"] > 0) + ].copy() + weekly = org["pternwa"].to_numpy(dtype=np.float32) / np.float32(100.0) + hours = org["pehruslt"].to_numpy(dtype=np.float32) + direct_hourly = org["pternhly"].to_numpy(dtype=np.float32) / np.float32(100.0) + hispanic = (org["pehspnon"].to_numpy() == 1).astype(np.float32) + result = pd.DataFrame( + { + "employment_income": weekly * np.float32(52.0), + "weekly_hours_worked": hours, + "age": org["prtage"].to_numpy(dtype=np.float32), + "is_female": (org["pesex"].to_numpy() == 2).astype(np.float32), + "is_hispanic": hispanic, + "race_wbho": _derive_wbho(org["ptdtrace"].to_numpy(), hispanic), + "state_fips": org["gestfips"].to_numpy(dtype=np.float32), + "hourly_wage": np.where( + (org["peernhry"].to_numpy() == 1) & (direct_hourly > 0), + direct_hourly, + weekly / hours, + ), + "is_paid_hourly": (org["peernhry"].to_numpy() == 1).astype(np.float32), + _DONOR_WEIGHT_COLUMN: org["pworwgt"].to_numpy(dtype=np.float32), + } + ) + return result.loc[ + (result["employment_income"] > 0) + & (result["weekly_hours_worked"] > 0) + & (result["hourly_wage"] > 0) + ].reset_index(drop=True) + + +def fetch_org_2024_donor( + cache_dir: str | Path | None = None, + *, + expected_content_sha256: str | None = ORG_2024_DONOR_CONTENT_SHA256, +) -> Path: + """Build/cache the exact twelve-month ORG donor and verify its content.""" + + root = ( + Path(cache_dir) + if cache_dir is not None + else Path.home() / ".cache" / "populace" / "org" + ) + root.mkdir(parents=True, exist_ok=True) + target = root / ORG_2024_DONOR_FILENAME + if target.exists(): + try: + content = _canonical_donor_bytes(target) + except OSError: + content = b"" + if content and ( + expected_content_sha256 is None + or _sha256(content) == expected_content_sha256 + ): + return target + + donor = pd.concat( + [_transform_month(_load_month_from_network(month)) for month in ORG_MONTHS], + ignore_index=True, + ) + content = donor.to_csv(index=False).encode("utf-8") + digest = _sha256(content) + if expected_content_sha256 is not None and digest != expected_content_sha256: + raise ValueError( + "CPS ORG donor failed canonical sha-256 verification after rebuilding: " + f"expected {expected_content_sha256}, got {digest}." + ) + target.write_bytes(gzip.compress(content, compresslevel=9, mtime=0)) + return target + + +def load_org_2024_donor( + path: str | Path, + *, + expected_content_sha256: str | None = None, +) -> pd.DataFrame: + """Load a transformed ORG donor, optionally verifying canonical bytes.""" + + content = _canonical_donor_bytes(path) + if expected_content_sha256 is not None: + digest = _sha256(content) + if digest != expected_content_sha256: + raise ValueError( + "CPS ORG donor failed canonical sha-256 verification: " + f"expected {expected_content_sha256}, got {digest}." + ) + donor = pd.read_csv(io.BytesIO(content)) + required = [*ORG_PREDICTORS, *ORG_QRF_TARGETS, _DONOR_WEIGHT_COLUMN] + missing = [column for column in required if column not in donor] + if missing: + raise ValueError(f"CPS ORG donor missing column(s): {missing}.") + donor = donor.loc[:, required].apply(pd.to_numeric, errors="coerce") + finite = np.isfinite(donor.to_numpy(dtype=np.float64)).all(axis=1) + valid = ( + finite + & (donor[_DONOR_WEIGHT_COLUMN].to_numpy() > 0) + & (donor["employment_income"].to_numpy() > 0) + & (donor["weekly_hours_worked"].to_numpy() > 0) + & (donor["hourly_wage"].to_numpy() > 0) + ) + result = donor.loc[valid].reset_index(drop=True) + if result.empty: + raise ValueError("CPS ORG donor has no valid training rows.") + return result + + +def derive_us_org_occupation_inputs(person: pd.DataFrame) -> pd.DataFrame: + """Carry CPS race/ethnicity and derive the exact POCCU2 FLSA flags.""" + + missing = [ + column for column in ("PRDTRACE", "PRDTHSP", "POCCU2") if column not in person + ] + if missing: + raise ValueError( + f"ORG occupation derivation needs person column(s): {missing}." + ) + occupation = pd.to_numeric(person["POCCU2"], errors="coerce").fillna(0).astype(int) + result = pd.DataFrame(index=person.index) + result["cps_race"] = ( + pd.to_numeric(person["PRDTRACE"], errors="coerce").fillna(0).astype(np.int16) + ) + result["is_hispanic"] = ( + pd.to_numeric(person["PRDTHSP"], errors="coerce").fillna(0).ne(0) + ) + result["detailed_occupation_recode"] = occupation.astype(np.int16) + for variable, code in FLSA_OVERTIME_OCCUPATION_CODES.items(): + result[variable] = occupation.eq(code) + result["is_executive_administrative_professional"] = occupation.isin( + FLSA_EXECUTIVE_ADMINISTRATIVE_PROFESSIONAL_OCCUPATION_CODES + ) + return result + + +def _person_state_fips(frame: Frame) -> np.ndarray: + person = frame.table("person") + if "state_fips" in person: + return pd.to_numeric(person["state_fips"], errors="coerce").fillna(0).to_numpy() + household = frame.table("household") + if "state_fips" not in household: + raise ValueError("ORG imputation needs household.state_fips.") + mapping = pd.Series( + household["state_fips"].to_numpy(), + index=household["household_id"].to_numpy(), + ) + state = person["person_household_id"].map(mapping) + if state.isna().any(): + raise ValueError( + "ORG imputation could not map every person to household state_fips." + ) + return pd.to_numeric(state, errors="coerce").fillna(0).to_numpy() + + +def _recipient_features(frame: Frame, carried: pd.DataFrame) -> pd.DataFrame: + person = frame.table("person") + missing = [ + column + for column in US_ORG_WAGES_REQUIRED_SOURCE_COLUMNS + if column not in person + ] + if missing: + raise ValueError(f"ORG imputation needs person column(s): {missing}.") + features = pd.DataFrame(index=person.index) + features["employment_income"] = ( + pd.to_numeric(person["employment_income_before_lsr"], errors="coerce") + .fillna(0) + .clip(lower=0) + ) + features["weekly_hours_worked"] = ( + pd.to_numeric(person["weekly_hours_worked_before_lsr"], errors="coerce") + .fillna(0) + .clip(lower=0) + ) + features["age"] = pd.to_numeric(person["age"], errors="coerce").fillna(0) + features["is_female"] = pd.to_numeric(person["is_female"], errors="coerce").fillna( + 0 + ) + features["is_hispanic"] = carried["is_hispanic"].astype(np.float64) + features["race_wbho"] = _derive_wbho( + carried["cps_race"].to_numpy(), carried["is_hispanic"].to_numpy() + ) + features["state_fips"] = _person_state_fips(frame) + return features.loc[:, list(ORG_PREDICTORS)] + + +def _union_priority(features: pd.DataFrame) -> np.ndarray: + base = float(BLS_UNION_REPRESENTATION_RATE_2024) + age = features["age"].to_numpy(dtype=np.float64) + age_rate = np.full(len(features), base) + for (low, high), rate in _BLS_AGE_RATES.items(): + age_rate[(age >= low) & (age <= high)] = rate + sex_rate = np.where( + features["is_female"].to_numpy() >= 0.5, + _BLS_SEX_RATES[True], + _BLS_SEX_RATES[False], + ) + race = features["race_wbho"].to_numpy(dtype=int) + race_rate = np.full(len(features), base) + for code, rate in _BLS_RACE_RATES.items(): + race_rate[race == code] = rate + hours_rate = np.where( + features["weekly_hours_worked"].to_numpy() >= 35, + _BLS_FULL_TIME_RATE, + _BLS_PART_TIME_RATE, + ) + return np.clip( + (age_rate / base) + * (sex_rate / base) + * (race_rate / base) + * (hours_rate / base), + 1e-6, + None, + ) + + +def _assign_union(features: pd.DataFrame) -> np.ndarray: + result = np.zeros(len(features), dtype=bool) + eligible = ( + (features["employment_income"].to_numpy() > 0) + & (features["weekly_hours_worked"].to_numpy() > 0) + & (features["age"].to_numpy() >= 16) + ) + if not eligible.any(): + return result + hashes = pd.util.hash_pandas_object( + pd.DataFrame( + { + column: np.round(features[column].to_numpy(dtype=np.float64), 4) + for column in ORG_PREDICTORS + } + ), + index=False, + ).to_numpy(dtype=np.uint64) + uniform = (hashes.astype(np.float64) + 0.5) / (np.iinfo(np.uint64).max + 1.0) + key = -np.log(np.clip(uniform, 1e-12, 1 - 1e-12)) / _union_priority(features) + states = np.nan_to_num(features["state_fips"].to_numpy(), nan=-1).astype(int) + for state in np.unique(states[eligible]): + positions = np.flatnonzero(eligible & (states == state)) + rate = float( + BLS_STATE_UNION_REPRESENTATION_RATE_2024.get( + state, BLS_UNION_REPRESENTATION_RATE_2024 + ) + ) + count = min(len(positions), max(0, int(np.rint(rate * len(positions))))) + if count: + selected = np.argpartition(key[positions], count - 1)[:count] + result[positions[selected]] = True + return result + + +def impute_us_org_wages( + frame: Frame, + donor: pd.DataFrame, + *, + seed: int, + n_estimators: int = _DEFAULT_N_ESTIMATORS, +) -> tuple[pd.DataFrame, pd.DataFrame]: + """QRF-impute hourly inputs and return them with direct ASEC carries.""" + + from populace.fit import QRF + + carried = derive_us_org_occupation_inputs(frame.table("person")) + features = _recipient_features(frame, carried) + required = [*ORG_PREDICTORS, *ORG_QRF_TARGETS, _DONOR_WEIGHT_COLUMN] + missing = [column for column in required if column not in donor] + if missing: + raise ValueError(f"CPS ORG donor missing column(s): {missing}.") + fitted = QRF(n_estimators=int(n_estimators), seed=int(seed)).fit( + donor.loc[:, required], + predictors=list(ORG_PREDICTORS), + targets=list(ORG_QRF_TARGETS), + weights=_DONOR_WEIGHT_COLUMN, + ) + predicted = fitted.predict(features) + inactive = (features["employment_income"].to_numpy() <= 0) | ( + features["weekly_hours_worked"].to_numpy() <= 0 + ) + wages = pd.DataFrame(index=features.index) + wages["hourly_wage"] = np.maximum( + pd.to_numeric(predicted["hourly_wage"], errors="coerce").fillna(0).to_numpy(), 0 + ) + wages["is_paid_hourly"] = ( + pd.to_numeric(predicted["is_paid_hourly"], errors="coerce").fillna(0).to_numpy() + >= 0.5 + ) + wages["is_union_member_or_covered"] = _assign_union(features) + wages.loc[inactive, "hourly_wage"] = 0.0 + wages.loc[inactive, ["is_paid_hourly", "is_union_member_or_covered"]] = False + return carried, wages + + +def _flsa_policy(year: int) -> tuple[float, float, float, float, float]: + from policyengine_us import CountryTaxBenefitSystem + + overtime = ( + CountryTaxBenefitSystem() + .parameters(f"{int(year)}-01-01") + .gov.irs.income.exemption.overtime + ) + hours = float(overtime.hours_threshold) + return ( + float(overtime.hce_salary_threshold), + float(overtime.salary_basis_threshold) * 52.0, + float(overtime.computer_salary_threshold) * hours * 52.0, + hours, + float(overtime.rate_multiplier), + ) + + +def derive_flsa_overtime_premium( + *, + time_period: int, + employment_income: np.ndarray | pd.Series, + hours_worked_last_week: np.ndarray | pd.Series, + weeks_worked: np.ndarray | pd.Series, + is_paid_hourly: np.ndarray | pd.Series, + has_never_worked: np.ndarray | pd.Series, + is_military: np.ndarray | pd.Series, + is_executive_administrative_professional: np.ndarray | pd.Series, + is_farmer_fisher: np.ndarray | pd.Series, + is_computer_scientist: np.ndarray | pd.Series, + policy: tuple[float, float, float, float, float] | None = None, +) -> np.ndarray: + """Port the retired annual-wage-share FLSA overtime proxy exactly.""" + + income = np.maximum( + np.nan_to_num(np.asarray(employment_income, dtype=np.float64), nan=0), 0 + ) + hours = np.maximum( + np.nan_to_num(np.asarray(hours_worked_last_week, dtype=np.float64), nan=0), 0 + ) + weeks = np.maximum( + np.nan_to_num(np.asarray(weeks_worked, dtype=np.float64), nan=0), 0 + ) + paid_hourly = np.asarray(is_paid_hourly, dtype=bool) + never = np.asarray(has_never_worked, dtype=bool) + military = np.asarray(is_military, dtype=bool) + eap = np.asarray(is_executive_administrative_professional, dtype=bool) + farmer = np.asarray(is_farmer_fisher, dtype=bool) + computer = np.asarray(is_computer_scientist, dtype=bool) + hce, salary_basis, computer_salary, threshold_hours, multiplier = ( + _flsa_policy(time_period) if policy is None else policy + ) + overtime_hours = np.maximum(hours - threshold_hours, 0) + straight_equivalent = ( + np.minimum(hours, threshold_hours) + overtime_hours * multiplier + ) + premium_share = np.divide( + (multiplier - 1) * overtime_hours, + straight_equivalent, + out=np.zeros_like(income), + where=straight_equivalent > 0, + ) + salary_threshold = np.full_like(income, hce) + salary_threshold = np.where(computer, min(computer_salary, hce), salary_threshold) + salary_threshold = np.where(eap | farmer, min(salary_basis, hce), salary_threshold) + always_exempt = never | military + exempt = always_exempt | ((income >= salary_threshold) & ~paid_hourly) + premium = np.where(~exempt & (weeks > 0), income * premium_share, 0) + return np.minimum(premium, income).astype(np.float32) + + +def _surface_has_signal(person: pd.DataFrame) -> bool: + return all( + column in person and person[column].dropna().nunique() > 1 + for column in US_ORG_WAGES_OUTPUT_COLUMNS + ) + + +def with_us_org_wages_inputs( + frame: Frame, + *, + seed: int, + time_period: int, + org_donor: pd.DataFrame, +) -> Frame: + """Apply the complete ORG/occupation/FLSA family before calibration.""" + + if frame.schema != US_SCHEMA: + raise ValueError("US ORG inputs require the US schema.") + us_org_wages_stage_spec() + person = frame.table("person") + if _surface_has_signal(person): + return frame + carried, wages = impute_us_org_wages(frame, org_donor, seed=seed) + premium = derive_flsa_overtime_premium( + time_period=time_period, + employment_income=person["employment_income_before_lsr"], + hours_worked_last_week=person["hours_worked_last_week"], + weeks_worked=person["weeks_worked"], + is_paid_hourly=wages["is_paid_hourly"], + has_never_worked=carried["has_never_worked"], + is_military=carried["is_military"], + is_executive_administrative_professional=carried[ + "is_executive_administrative_professional" + ], + is_farmer_fisher=carried["is_farmer_fisher"], + is_computer_scientist=carried["is_computer_scientist"], + ) + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + for column in carried: + tables["person"][column] = carried[column].to_numpy() + for column in wages: + tables["person"][column] = wages[column].to_numpy() + tables["person"]["fsla_overtime_premium"] = premium + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + ) + + +def us_org_wages_summary(frame: Frame) -> dict[str, object]: + """Weighted incidence plus structural-invariant diagnostics.""" + + person = frame.table("person") + weights = np.asarray(frame.resolve_weights("person").values, dtype=np.float64) + total_weight = float(weights.sum()) + + def share(column: str) -> float: + values = pd.to_numeric(person[column], errors="coerce").fillna(0).to_numpy() + return float(weights[values != 0].sum() / total_weight) if total_weight else 0.0 + + premium = pd.to_numeric(person["fsla_overtime_premium"], errors="coerce").to_numpy( + dtype=np.float64 + ) + income = ( + pd.to_numeric(person["employment_income_before_lsr"], errors="coerce") + .fillna(0) + .to_numpy(dtype=np.float64) + ) + hours = ( + pd.to_numeric(person["hours_worked_last_week"], errors="coerce") + .fillna(0) + .to_numpy(dtype=np.float64) + ) + weeks = ( + pd.to_numeric(person["weeks_worked"], errors="coerce") + .fillna(0) + .to_numpy(dtype=np.float64) + ) + never = person["has_never_worked"].astype(bool).to_numpy() + military = person["is_military"].astype(bool).to_numpy() + violations = { + "nonfinite": int((~np.isfinite(premium)).sum()), + "negative": int((premium < 0).sum()), + "exceeds_employment_income": int( + (premium > np.maximum(income, 0) + 1e-4).sum() + ), + "positive_without_overtime": int(((premium > 0) & (hours <= 40)).sum()), + "positive_without_weeks": int(((premium > 0) & (weeks <= 0)).sum()), + "positive_always_exempt": int(((premium > 0) & (never | military)).sum()), + } + return { + "nonzero_shares": {column: share(column) for column in _SHARE_BANDS}, + "share_bands": {column: list(band) for column, band in _SHARE_BANDS.items()}, + "unique_counts": { + column: int(person[column].dropna().nunique()) + for column in US_ORG_WAGES_OUTPUT_COLUMNS + }, + "fsla_overtime_premium_weighted_total": float(np.nansum(premium * weights)), + "constraint_violations": violations, + } + + +def us_org_wages_signal_gate(frame: Frame) -> GateResult: + """Require the full ORG/FLSA family to carry plausible, coherent signal.""" + + person = frame.table("person") + missing = [column for column in US_ORG_WAGES_OUTPUT_COLUMNS if column not in person] + if missing: + return GateResult( + name="org_wages_signal", + passed=False, + failures=(f"person columns missing: {missing}.",), + details={"missing": missing}, + ) + summary = us_org_wages_summary(frame) + failures: list[str] = [] + for column, count in summary["unique_counts"].items(): + if count < 2: + failures.append(f"{column}: constant column carries no signal.") + for column, band in _SHARE_BANDS.items(): + value = float(summary["nonzero_shares"][column]) + low, high = band + if not (low <= value <= high): + failures.append( + f"{column}: weighted nonzero share {value:.4f} outside [{low}, {high}]." + ) + for name, count in summary["constraint_violations"].items(): + if count: + failures.append(f"fsla_overtime_premium: {count} {name} violation(s).") + return GateResult( + name="org_wages_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) diff --git a/packages/populace-build/src/populace/build/us_runtime/other_health_insurance.py b/packages/populace-build/src/populace/build/us_runtime/other_health_insurance.py new file mode 100644 index 00000000..2ecc35a9 --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/other_health_insurance.py @@ -0,0 +1,879 @@ +"""CPS ASEC other-health-insurance premium inputs. + +The retired eCPS pipeline kept the CPS-reported non-Medicare premium and +derived ``other_health_insurance_premiums`` as its nonnegative residual after +subtracting baseline CHIP, Marketplace, and Medicaid premiums. Those three +modeled premiums live on tax units and were assigned wholly to the first +person in each tax unit before subtraction. The retired extended-CPS stage +then jointly QRF-imputed the reported and residual leaves onto its PUF clone +half using eight documented predictors and at most 5,000 ASEC people. + +This port preserves measured ASEC values exactly and replaces only the PUF +support channel. It pins 100 trees and supplies typed person weights as +Populace reproducibility hardening; the archive left the tree count implicit +and did not weight the fit. The separate employer-premium target is not +proxied here; ``has_esi`` remains only the archived predictor shared by the +restored targets. +""" + +from __future__ import annotations + +import gc +from importlib.resources import files +from typing import Any + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import ( + SourceOperationSpec, + SourceStageSpec, + load_source_manifest, +) +from populace.build.source_runtime import ( + SourceRuntimeConfig, + SourceRuntimeContext, + SourceRuntimeError, + run_source_stage, +) +from populace.frame import Frame +from populace.frame.adapters.policyengine_us import PolicyEngineUSEngine +from populace.frame.units import US_SCHEMA + +__all__ = [ + "OTHER_HEALTH_INSURANCE_ARCHIVED_DERIVATION_URL", + "OTHER_HEALTH_INSURANCE_ARCHIVED_PUF_IMPUTATION_URL", + "OTHER_HEALTH_INSURANCE_ARCHIVED_PUF_OUTPUTS_URL", + "OTHER_HEALTH_INSURANCE_ARCHIVED_PUF_PREDICTORS_URL", + "OTHER_HEALTH_INSURANCE_ARCHIVED_PUF_SPLICE_URL", + "US_OTHER_HEALTH_INSURANCE_MODELED_PREMIUM_VARIABLES", + "US_OTHER_HEALTH_INSURANCE_NONCONSTANT_PERSON_COLUMNS", + "US_OTHER_HEALTH_INSURANCE_OUTPUT_COLUMNS", + "US_OTHER_HEALTH_INSURANCE_REQUIRED_SOURCE_COLUMNS", + "US_OTHER_HEALTH_INSURANCE_STAGE_NAME", + "derive_us_other_health_insurance_from_asec", + "derive_us_other_health_insurance_from_manifest", + "impute_us_other_health_insurance_to_puf_support_from_manifest", + "us_other_health_insurance_signal_gate", + "us_other_health_insurance_stage_spec", + "us_other_health_insurance_summary", + "with_us_other_health_insurance_inputs", +] + +QRF: Any | None = None + +_ARCHIVED_DATA_REPOSITORY = "policyengine-" + "us-data" +_ARCHIVED_ROOT = ( + "https://github.com/PolicyEngine/" + f"{_ARCHIVED_DATA_REPOSITORY}/blob/" + "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe/" + "policyengine_" + "us_data/" +) +OTHER_HEALTH_INSURANCE_ARCHIVED_DERIVATION_URL = ( + _ARCHIVED_ROOT + "datasets/cps/cps.py#L828-L944" +) +OTHER_HEALTH_INSURANCE_ARCHIVED_PUF_OUTPUTS_URL = ( + _ARCHIVED_ROOT + "datasets/cps/extended_cps.py#L135-L194" +) +OTHER_HEALTH_INSURANCE_ARCHIVED_PUF_PREDICTORS_URL = ( + _ARCHIVED_ROOT + "datasets/cps/extended_cps.py#L234-L248" +) +OTHER_HEALTH_INSURANCE_ARCHIVED_PUF_IMPUTATION_URL = ( + _ARCHIVED_ROOT + "datasets/cps/extended_cps.py#L639-L745" +) +OTHER_HEALTH_INSURANCE_ARCHIVED_PUF_SPLICE_URL = ( + _ARCHIVED_ROOT + "datasets/cps/extended_cps.py#L1014-L1076" +) + +US_OTHER_HEALTH_INSURANCE_STAGE_NAME = "other_health_insurance_premiums" +US_OTHER_HEALTH_INSURANCE_OUTPUT_COLUMNS: tuple[str, ...] = ( + "health_insurance_premiums_without_medicare_part_b", + "other_health_insurance_premiums", +) +US_OTHER_HEALTH_INSURANCE_NONCONSTANT_PERSON_COLUMNS: tuple[str, ...] = ( + US_OTHER_HEALTH_INSURANCE_OUTPUT_COLUMNS[1], +) +US_OTHER_HEALTH_INSURANCE_REQUIRED_SOURCE_COLUMNS: tuple[str, ...] = ( + US_OTHER_HEALTH_INSURANCE_OUTPUT_COLUMNS[0], +) +US_OTHER_HEALTH_INSURANCE_MODELED_PREMIUM_VARIABLES: tuple[str, ...] = ( + "chip_premium", + "marketplace_net_premium", + "medicaid_premium", +) + +_REPORTED_OUTPUT, _OTHER_OUTPUT = US_OTHER_HEALTH_INSURANCE_OUTPUT_COLUMNS +_PERSON_WEIGHT_COLUMN = "person_weight" +_PERSON_SUPPORT_CHANNEL_COLUMN = "person_support_channel" +_BASE_ASEC_SUPPORT_CHANNEL = "asec" +_PUF_TAX_DETAIL_SUPPORT_CHANNEL = "puf_tax_detail" +_PREDICTORS: tuple[str, ...] = ( + "age", + "is_male", + "has_esi", + "tax_unit_is_joint", + "tax_unit_count_dependents", + "employment_income", + "self_employment_income", + "social_security", +) +_PREDICTOR_PREFIX = "other_health_insurance_predictor_" +_EXPECTED_DIRECT_PARAMETERS = { + "reported_source": _REPORTED_OUTPUT, + "chip_premium_source": "chip_premium", + "marketplace_net_premium_source": "marketplace_net_premium", + "medicaid_premium_source": "medicaid_premium", + "output": _OTHER_OUTPUT, +} +_PUF_IMPUTATION_PARAMETER_KEYS = frozenset( + { + "predictors", + "max_train_samples", + "n_estimators", + "seed_from_build_config", + "weight", + } +) +_MAX_TRAIN_SAMPLES = 5_000 +_N_ESTIMATORS = 100 +_POSITIVE_SHARE_BAND = (0.15, 0.65) + + +def us_other_health_insurance_stage_spec() -> SourceStageSpec: + """Load and validate the packaged other-premium stage declaration.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + stage_map = manifest.stage_map() + if US_OTHER_HEALTH_INSURANCE_STAGE_NAME not in stage_map: + raise ValueError( + "US source manifest declares no " + f"{US_OTHER_HEALTH_INSURANCE_STAGE_NAME!r} stage." + ) + spec = stage_map[US_OTHER_HEALTH_INSURANCE_STAGE_NAME] + if tuple(spec.outputs) != US_OTHER_HEALTH_INSURANCE_OUTPUT_COLUMNS: + raise ValueError( + f"{US_OTHER_HEALTH_INSURANCE_STAGE_NAME!r} outputs must preserve " + "the archived target order " + f"{list(US_OTHER_HEALTH_INSURANCE_OUTPUT_COLUMNS)}; got " + f"{list(spec.outputs)}." + ) + return spec + + +def _strict_nonnegative_values(person: pd.DataFrame, column: str) -> np.ndarray: + if column not in person.columns: + raise SourceRuntimeError( + f"US other-health-insurance derivation requires source column {column!r}." + ) + values = pd.to_numeric(person[column], errors="coerce").to_numpy(dtype=np.float64) + nonfinite = int(np.count_nonzero(~np.isfinite(values))) + if nonfinite: + raise SourceRuntimeError( + f"US other-health-insurance source {column!r} contains {nonfinite} " + "nonnumeric or nonfinite value(s)." + ) + negative = int(np.count_nonzero(values < 0.0)) + if negative: + raise SourceRuntimeError( + f"US other-health-insurance source {column!r} contains {negative} " + "negative value(s)." + ) + return values + + +def derive_us_other_health_insurance_from_asec( + person: pd.DataFrame, + *, + reported_source_column: str = _REPORTED_OUTPUT, + chip_premium_source_column: str = "chip_premium", + marketplace_net_premium_source_column: str = "marketplace_net_premium", + medicaid_premium_source_column: str = "medicaid_premium", + output_column: str = _OTHER_OUTPUT, +) -> pd.DataFrame: + """Derive the nonnegative residual after baseline modeled premiums.""" + + reported = _strict_nonnegative_values(person, reported_source_column) + modeled = np.zeros(len(person), dtype=np.float64) + for column in ( + chip_premium_source_column, + marketplace_net_premium_source_column, + medicaid_premium_source_column, + ): + modeled += _strict_nonnegative_values(person, column) + + result = person.copy(deep=True) + result[output_column] = np.clip(reported - modeled, 0.0, None) + return result + + +def derive_us_other_health_insurance_from_manifest( + frame: pd.DataFrame | None, + operation: SourceOperationSpec, + _context: SourceRuntimeContext | None, +) -> pd.DataFrame: + """Interpret the manifest's archived premium-residual operation.""" + + if operation.kind != "derive_other_health_insurance_premiums": + raise SourceRuntimeError( + "US other-health-insurance derivation received unexpected operation " + f"{operation.kind!r}." + ) + if frame is None: + raise SourceRuntimeError( + "US other-health-insurance derivation requires the person table first." + ) + parameters = dict(operation.parameters) + if parameters != _EXPECTED_DIRECT_PARAMETERS: + raise SourceRuntimeError( + "US other-health-insurance direct mapping drifted from the archived " + f"method: expected {_EXPECTED_DIRECT_PARAMETERS}, got {parameters}." + ) + return derive_us_other_health_insurance_from_asec( + frame, + reported_source_column=parameters["reported_source"], + chip_premium_source_column=parameters["chip_premium_source"], + marketplace_net_premium_source_column=parameters[ + "marketplace_net_premium_source" + ], + medicaid_premium_source_column=parameters["medicaid_premium_source"], + output_column=parameters["output"], + ) + + +def impute_us_other_health_insurance_to_puf_support_from_manifest( + frame: pd.DataFrame | None, + operation: SourceOperationSpec, + context: SourceRuntimeContext | None, +) -> pd.DataFrame: + """Jointly QRF-impute both premium leaves onto PUF-support people.""" + + if operation.kind != "impute_other_health_insurance_premiums_to_puf_support": + raise SourceRuntimeError( + "US other-health-insurance PUF imputation received unexpected " + f"operation {operation.kind!r}." + ) + if frame is None: + raise SourceRuntimeError( + "US other-health-insurance PUF imputation requires the person table first." + ) + unexpected = sorted(set(operation.parameters) - _PUF_IMPUTATION_PARAMETER_KEYS) + missing_parameters = sorted( + _PUF_IMPUTATION_PARAMETER_KEYS - set(operation.parameters) + ) + if unexpected or missing_parameters: + raise SourceRuntimeError( + "US other-health-insurance PUF imputation parameters must match the " + f"archived method; missing={missing_parameters}, " + f"unexpected={unexpected}." + ) + if _PERSON_SUPPORT_CHANNEL_COLUMN not in frame.columns: + return frame.copy(deep=True) + + predictors = tuple(str(value) for value in operation.parameters["predictors"]) + if predictors != _PREDICTORS: + raise SourceRuntimeError( + "US other-health-insurance PUF predictors drifted from the archived " + f"method: expected {list(_PREDICTORS)}, got {list(predictors)}." + ) + if operation.parameters["weight"] != _PERSON_WEIGHT_COLUMN: + raise SourceRuntimeError( + "US other-health-insurance PUF imputation must use typed person weights." + ) + if operation.parameters["seed_from_build_config"] is not True: + raise SourceRuntimeError( + "US other-health-insurance PUF imputation seed must come from the " + "build config." + ) + max_train_samples = int(operation.parameters["max_train_samples"]) + n_estimators = int(operation.parameters["n_estimators"]) + if not 0 < max_train_samples <= _MAX_TRAIN_SAMPLES: + raise SourceRuntimeError( + "US other-health-insurance PUF max_train_samples must be between " + f"1 and {_MAX_TRAIN_SAMPLES}." + ) + if n_estimators != _N_ESTIMATORS: + raise SourceRuntimeError( + "US other-health-insurance PUF imputation must pin exactly " + f"{_N_ESTIMATORS} trees; got {n_estimators}." + ) + + predictor_columns = [_PREDICTOR_PREFIX + name for name in predictors] + required = [ + _PERSON_WEIGHT_COLUMN, + *predictor_columns, + *US_OTHER_HEALTH_INSURANCE_OUTPUT_COLUMNS, + ] + missing = [column for column in required if column not in frame.columns] + if missing: + raise SourceRuntimeError( + f"US other-health-insurance PUF imputation is missing column(s): {missing}." + ) + + channel = frame[_PERSON_SUPPORT_CHANNEL_COLUMN].astype(str) + asec_mask = channel == _BASE_ASEC_SUPPORT_CHANNEL + puf_mask = channel == _PUF_TAX_DETAIL_SUPPORT_CHANNEL + if not asec_mask.any() or not puf_mask.any(): + raise SourceRuntimeError( + "US other-health-insurance PUF imputation requires nonempty ASEC " + "and PUF-tax-detail support channels." + ) + + training = frame.loc[ + asec_mask, + [*predictor_columns, *US_OTHER_HEALTH_INSURANCE_OUTPUT_COLUMNS], + ].copy() + training.columns = [*predictors, *US_OTHER_HEALTH_INSURANCE_OUTPUT_COLUMNS] + test = frame.loc[puf_mask, predictor_columns].copy() + test.columns = list(predictors) + weights = pd.to_numeric( + frame.loc[asec_mask, _PERSON_WEIGHT_COLUMN], errors="coerce" + ) + numeric_weights = weights.to_numpy(dtype=np.float64) + if not np.isfinite(numeric_weights).all() or bool((numeric_weights < 0.0).any()): + raise SourceRuntimeError( + "US other-health-insurance QRF person weights must be finite and " + "nonnegative." + ) + if float(numeric_weights.sum()) <= 0.0: + raise SourceRuntimeError( + "US other-health-insurance QRF person weights sum to zero." + ) + + if len(training) > max_train_samples: + sample = training.sample( + n=max_train_samples, + random_state=(context.config.seed if context is not None else 0), + ).index + training = training.loc[sample] + weights = weights.loc[sample] + + for column in (*predictors, *US_OTHER_HEALTH_INSURANCE_OUTPUT_COLUMNS): + training[column] = pd.to_numeric(training[column], errors="coerce") + if not np.isfinite(training[column].to_numpy(dtype=np.float64)).all(): + raise SourceRuntimeError( + "US other-health-insurance QRF training column " + f"{column!r} contains nonfinite values." + ) + for column in predictors: + test[column] = pd.to_numeric(test[column], errors="coerce") + if not np.isfinite(test[column].to_numpy(dtype=np.float64)).all(): + raise SourceRuntimeError( + "US other-health-insurance QRF prediction column " + f"{column!r} contains nonfinite values." + ) + + global QRF + if QRF is None: + from importlib import import_module + + QRF = import_module("populace.fit").QRF + seed = context.config.seed if context is not None else 0 + fitted = QRF(n_estimators=n_estimators, seed=seed).fit( + training, + list(predictors), + list(US_OTHER_HEALTH_INSURANCE_OUTPUT_COLUMNS), + weights=weights.to_numpy(dtype=np.float64), + ) + predictions = fitted.predict(test) + missing_outputs = [ + output + for output in US_OTHER_HEALTH_INSURANCE_OUTPUT_COLUMNS + if output not in predictions + ] + if missing_outputs: + raise SourceRuntimeError( + "US other-health-insurance QRF prediction is missing output(s): " + f"{missing_outputs}." + ) + + result = frame.copy(deep=True) + for output in US_OTHER_HEALTH_INSURANCE_OUTPUT_COLUMNS: + predicted = pd.to_numeric(predictions[output], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not np.isfinite(predicted).all(): + raise SourceRuntimeError( + f"US other-health-insurance QRF produced nonfinite {output} values." + ) + if bool((predicted < 0.0).any()): + raise SourceRuntimeError( + f"US other-health-insurance QRF produced negative {output} values." + ) + result.loc[puf_mask, output] = predicted + return result + + +def _person_other_health_insurance_predictors(frame: Frame) -> pd.DataFrame: + """Build the retired eight second-stage QRF predictors on people.""" + + person = frame.table("person") + tax_unit = frame.table("tax_unit") + predictors = pd.DataFrame(index=person.index) + + def _numeric(*columns: str) -> np.ndarray: + for column in columns: + if column in person: + values = pd.to_numeric(person[column], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not np.isfinite(values).all(): + raise SourceRuntimeError( + "US other-health-insurance QRF predictor source " + f"{column!r} contains nonfinite values." + ) + return values + raise SourceRuntimeError( + "US other-health-insurance PUF imputation cannot construct a " + f"predictor from any of {list(columns)}." + ) + + predictors["age"] = _numeric("age", "A_AGE") + if "is_male" in person: + predictors["is_male"] = person["is_male"].fillna(False).astype(bool) + elif "is_female" in person: + predictors["is_male"] = ~person["is_female"].fillna(False).astype(bool) + elif "A_SEX" in person: + predictors["is_male"] = _numeric("A_SEX") == 1 + else: + raise SourceRuntimeError( + "US other-health-insurance PUF imputation requires is_male, " + "is_female, or measured A_SEX." + ) + if "has_esi" not in person: + raise SourceRuntimeError( + "US other-health-insurance PUF imputation requires has_esi." + ) + predictors["has_esi"] = person["has_esi"].fillna(False).astype(bool) + + filing_status_column = next( + ( + column + for column in ("filing_status_input", "filing_status") + if column in tax_unit + ), + None, + ) + if filing_status_column is None: + raise SourceRuntimeError( + "US other-health-insurance PUF imputation requires tax-unit filing status." + ) + filing_status = tax_unit[filing_status_column].map( + lambda value: ( + value.decode() if isinstance(value, (bytes, np.bytes_)) else str(value) + ) + ) + joint_by_id = pd.Series( + filing_status.eq("JOINT").to_numpy(), + index=tax_unit["tax_unit_id"].to_numpy(), + ) + predictors["tax_unit_is_joint"] = ( + person["person_tax_unit_id"].map(joint_by_id).fillna(False).astype(bool) + ) + + if "tax_unit_role_input" not in person: + raise SourceRuntimeError( + "US other-health-insurance PUF imputation requires " + "tax_unit_role_input to count dependents." + ) + dependent = ( + person["tax_unit_role_input"] + .map( + lambda value: ( + value.decode() if isinstance(value, (bytes, np.bytes_)) else str(value) + ) + ) + .eq("DEPENDENT") + ) + predictors["tax_unit_count_dependents"] = ( + dependent.groupby(person["person_tax_unit_id"]) + .transform("sum") + .to_numpy(dtype=np.float64) + ) + predictors["employment_income"] = _numeric( + "employment_income_before_lsr", "WSAL_VAL" + ) + predictors["self_employment_income"] = _numeric( + "self_employment_income_before_lsr", "SEMP_VAL" + ) + social_security_columns = ( + "social_security_retirement", + "social_security_disability", + "social_security_dependents", + "social_security_survivors", + ) + if all(column in person for column in social_security_columns): + predictors["social_security"] = np.sum( + np.column_stack([_numeric(column) for column in social_security_columns]), + axis=1, + ) + else: + predictors["social_security"] = _numeric("SS_VAL") + return predictors + + +def _tax_unit_values_on_first_person( + frame: Frame, + values: np.ndarray, + *, + variable: str, +) -> np.ndarray: + """Allocate each tax-unit formula amount only to its first person.""" + + tax_unit = frame.table("tax_unit") + person = frame.table("person") + numeric = np.asarray(values, dtype=np.float64) + if numeric.shape != (len(tax_unit),): + raise ValueError( + f"Materialized {variable!r} has shape {numeric.shape}; expected " + f"{(len(tax_unit),)} for tax units." + ) + if not np.isfinite(numeric).all() or bool((numeric < 0.0).any()): + raise ValueError(f"Materialized {variable!r} must be finite and nonnegative.") + + by_id = pd.Series(numeric, index=tax_unit["tax_unit_id"].to_numpy()) + memberships = person["person_tax_unit_id"] + mapped = memberships.map(by_id) + if mapped.isna().any(): + missing = sorted(set(memberships[mapped.isna()].tolist())) + raise ValueError( + f"Cannot allocate {variable!r}; unknown person tax-unit IDs {missing}." + ) + allocated = np.zeros(len(person), dtype=np.float64) + first_person = ~memberships.duplicated(keep="first") + allocated[first_person.to_numpy()] = mapped[first_person].to_numpy(dtype=np.float64) + represented = set(memberships.tolist()) + missing_tax_units = sorted(set(tax_unit["tax_unit_id"].tolist()) - represented) + if missing_tax_units: + raise ValueError( + f"Cannot allocate {variable!r}; tax units without people " + f"{missing_tax_units}." + ) + return allocated + + +def _materialize_modeled_premiums( + frame: Frame, + *, + time_period: int, + maximum_microsim_batch_size: int | None, +) -> dict[str, np.ndarray]: + """Materialize tax-unit premiums in household batches and realign by ID.""" + + household = frame.table("household") + tax_unit = frame.table("tax_unit") + person = frame.table("person") + household_ids = household["household_id"].to_numpy() + tax_unit_ids = tax_unit["tax_unit_id"].to_numpy() + tax_unit_positions = pd.Series( + np.arange(len(tax_unit_ids), dtype=np.int64), + index=tax_unit_ids, + ) + modeled = { + variable: np.zeros(len(tax_unit_ids), dtype=np.float64) + for variable in US_OTHER_HEALTH_INSURANCE_MODELED_PREMIUM_VARIABLES + } + assigned = np.zeros(len(tax_unit_ids), dtype=bool) + batch_size = ( + len(household_ids) + if maximum_microsim_batch_size is None or maximum_microsim_batch_size <= 0 + else min(int(maximum_microsim_batch_size), len(household_ids)) + ) + if batch_size == 0: + raise ValueError( + "US other-health-insurance premium materialization requires at " + "least one household." + ) + + engine = PolicyEngineUSEngine() + for start in range(0, len(household_ids), batch_size): + positions = np.arange( + start, + min(start + batch_size, len(household_ids)), + dtype=np.int64, + ) + full_batch = len(positions) == len(household_ids) + batch_frame = ( + frame + if full_batch + else frame.select( + person["person_household_id"].isin(household_ids[positions]) + ) + ) + batch_tax_unit_ids = batch_frame.table("tax_unit")["tax_unit_id"].to_numpy() + full_positions = tax_unit_positions.reindex(batch_tax_unit_ids).to_numpy() + if np.isnan(full_positions).any(): + raise ValueError( + "US other-health-insurance premium batch produced tax-unit IDs " + "outside the full frame." + ) + full_positions = full_positions.astype(np.int64) + if assigned[full_positions].any(): + raise ValueError( + "US other-health-insurance premium batches overlap tax units; " + "tax units must not span household batches." + ) + materialized = engine.materialize( + batch_frame, + list(US_OTHER_HEALTH_INSURANCE_MODELED_PREMIUM_VARIABLES), + period=int(time_period), + ) + for variable in US_OTHER_HEALTH_INSURANCE_MODELED_PREMIUM_VARIABLES: + if variable not in materialized: + raise ValueError( + "PolicyEngine-US did not materialize required formula " + f"{variable!r}." + ) + values = np.asarray(materialized[variable], dtype=np.float64) + if values.shape != (len(batch_tax_unit_ids),): + raise ValueError( + f"Materialized {variable!r} has shape {values.shape}; expected " + f"{(len(batch_tax_unit_ids),)} for the tax-unit batch." + ) + modeled[variable][full_positions] = values + assigned[full_positions] = True + del materialized, batch_frame + gc.collect() + + if not assigned.all(): + missing_ids = tax_unit_ids[~assigned].tolist() + raise ValueError( + "US other-health-insurance premium batches did not cover tax-unit " + f"IDs {missing_ids}." + ) + return modeled + + +def with_us_other_health_insurance_inputs( + frame: Frame, + *, + seed: int, + time_period: int, + maximum_microsim_batch_size: int | None = None, + allow_existing_without_source: bool = False, +) -> Frame: + """Materialize modeled premiums, derive ASEC residuals, and impute PUF.""" + + if frame.schema != US_SCHEMA: + raise ValueError("US other-health-insurance inputs require the US schema.") + person = frame.table("person") + source_available = all( + column in person for column in US_OTHER_HEALTH_INSURANCE_REQUIRED_SOURCE_COLUMNS + ) + if not source_available: + if ( + allow_existing_without_source + and _other_health_insurance_surface_carries_signal(frame) + ): + return frame + missing = [ + column + for column in US_OTHER_HEALTH_INSURANCE_REQUIRED_SOURCE_COLUMNS + if column not in person + ] + raise ValueError( + "US other-health-insurance stage cannot heal a default surface " + f"without measured ASEC source column(s): {missing}." + ) + + modeled = _materialize_modeled_premiums( + frame, + time_period=int(time_period), + maximum_microsim_batch_size=maximum_microsim_batch_size, + ) + stage_person = person.copy(deep=True) + for variable in US_OTHER_HEALTH_INSURANCE_MODELED_PREMIUM_VARIABLES: + stage_person[variable] = _tax_unit_values_on_first_person( + frame, + np.asarray(modeled[variable]), + variable=variable, + ) + stage_person[_PERSON_WEIGHT_COLUMN] = frame.resolve_weights("person").values + if _PERSON_SUPPORT_CHANNEL_COLUMN in person: + predictors = _person_other_health_insurance_predictors(frame) + for column in _PREDICTORS: + stage_person[_PREDICTOR_PREFIX + column] = predictors[column].to_numpy() + + output = run_source_stage( + us_other_health_insurance_stage_spec(), + tables={"person": stage_person}, + operation_handlers={ + "derive_other_health_insurance_premiums": ( + derive_us_other_health_insurance_from_manifest + ), + "impute_other_health_insurance_premiums_to_puf_support": ( + impute_us_other_health_insurance_to_puf_support_from_manifest + ), + }, + config=SourceRuntimeConfig(seed=int(seed), target_year=int(time_period)), + ) + aligned = output.set_index("person_id").reindex(person["person_id"]) + for column in US_OTHER_HEALTH_INSURANCE_OUTPUT_COLUMNS: + if aligned[column].isna().any(): + raise ValueError( + "US other-health-insurance stage output " + f"{column!r} does not cover every person." + ) + + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + for column in US_OTHER_HEALTH_INSURANCE_OUTPUT_COLUMNS: + tables["person"][column] = aligned[column].to_numpy(dtype=np.float64) + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + ) + + +def us_other_health_insurance_summary(frame: Frame) -> dict[str, object]: + """Return weighted signal, validity, channel, and residual diagnostics.""" + + person = frame.table("person") + weights = np.asarray(frame.resolve_weights("person").values, dtype=np.float64) + total_weight = float(weights.sum()) + channel = ( + person[_PERSON_SUPPORT_CHANNEL_COLUMN].astype(str).to_numpy() + if _PERSON_SUPPORT_CHANNEL_COLUMN in person + else None + ) + columns: dict[str, dict[str, object]] = {} + numeric_outputs: dict[str, np.ndarray] = {} + for output in US_OTHER_HEALTH_INSURANCE_OUTPUT_COLUMNS: + values = pd.to_numeric(person[output], errors="coerce").to_numpy( + dtype=np.float64 + ) + numeric_outputs[output] = values + finite = np.isfinite(values) + positive = finite & (values > 0.0) + detail: dict[str, object] = { + "positive_share": ( + float(weights[positive].sum()) / total_weight + if total_weight > 0.0 + else 0.0 + ), + "positive_share_band": list(_POSITIVE_SHARE_BAND), + "weighted_total": float((np.nan_to_num(values) * weights).sum()), + "nonfinite": int(np.count_nonzero(~finite)), + "negative": int(np.count_nonzero(finite & (values < 0.0))), + } + if channel is not None: + channels: dict[str, dict[str, float | int | list[float]]] = {} + for name in (_BASE_ASEC_SUPPORT_CHANNEL, _PUF_TAX_DETAIL_SUPPORT_CHANNEL): + mask = channel == name + channel_weight = float(weights[mask].sum()) + channels[name] = { + "positive_rows": int(np.count_nonzero(mask & positive)), + "positive_share": ( + float(weights[mask & positive].sum()) / channel_weight + if channel_weight > 0.0 + else 0.0 + ), + "positive_share_band": list(_POSITIVE_SHARE_BAND), + "weighted_total": float( + (np.nan_to_num(values[mask]) * weights[mask]).sum() + ), + } + detail["channels"] = channels + columns[output] = detail + + reported = numeric_outputs[_REPORTED_OUTPUT] + other = numeric_outputs[_OTHER_OUTPUT] + comparable = np.isfinite(reported) & np.isfinite(other) + residual_violation = comparable & (other > reported) + result: dict[str, object] = { + "columns": columns, + "other_exceeds_reported": int(np.count_nonzero(residual_violation)), + "weighted_other_excess": float( + (np.maximum(other - reported, 0.0)[comparable] * weights[comparable]).sum() + ), + } + if channel is not None: + result["other_exceeds_reported_by_channel"] = { + name: int(np.count_nonzero(residual_violation & (channel == name))) + for name in (_BASE_ASEC_SUPPORT_CHANNEL, _PUF_TAX_DETAIL_SUPPORT_CHANNEL) + } + return result + + +def us_other_health_insurance_signal_gate(frame: Frame) -> GateResult: + """Require plausible finite premiums and the residual ordering identity.""" + + person = frame.table("person") + missing = [ + output + for output in US_OTHER_HEALTH_INSURANCE_OUTPUT_COLUMNS + if output not in person + ] + if missing: + return GateResult( + name="other_health_insurance_premiums_signal", + passed=False, + failures=(f"person columns missing: {missing}.",), + details={"missing": missing}, + ) + + summary = us_other_health_insurance_summary(frame) + failures: list[str] = [] + columns = summary["columns"] + for output in US_OTHER_HEALTH_INSURANCE_OUTPUT_COLUMNS: + detail = columns[output] + if detail["nonfinite"]: + failures.append(f"{output}: {int(detail['nonfinite'])} nonfinite values.") + if detail["negative"]: + failures.append(f"{output}: {int(detail['negative'])} negative values.") + share = float(detail["positive_share"]) + low, high = detail["positive_share_band"] + if not (low <= share <= high): + failures.append( + f"{output}: positive share {share:.4f} outside plausibility " + f"band [{low}, {high}]." + ) + channels = detail.get("channels") + if isinstance(channels, dict): + for name in (_BASE_ASEC_SUPPORT_CHANNEL, _PUF_TAX_DETAIL_SUPPORT_CHANNEL): + channel_detail = channels.get(name) + if not isinstance(channel_detail, dict): + failures.append(f"{output}: missing {name} channel diagnostics.") + continue + channel_share = float(channel_detail["positive_share"]) + channel_low, channel_high = channel_detail["positive_share_band"] + if not (channel_low <= channel_share <= channel_high): + failures.append( + f"{output}: {name} positive share {channel_share:.4f} " + "outside plausibility band " + f"[{channel_low}, {channel_high}]." + ) + channel_violations = summary.get("other_exceeds_reported_by_channel") + # The measured ASEC residual must preserve its exact source identity. The + # archived joint QRF did not reconcile its two independently predicted + # premium leaves, so PUF-only exceedances remain diagnostics rather than a + # post-hoc clipping rule the source never applied. + identity_violations = ( + int(channel_violations.get(_BASE_ASEC_SUPPORT_CHANNEL, 0)) + if isinstance(channel_violations, dict) + else int(summary["other_exceeds_reported"]) + ) + if identity_violations: + failures.append( + f"{_OTHER_OUTPUT}: {identity_violations} measured ASEC value(s) " + f"exceed {_REPORTED_OUTPUT}." + ) + return GateResult( + name="other_health_insurance_premiums_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) + + +def _other_health_insurance_surface_carries_signal(frame: Frame) -> bool: + if any( + output not in frame.table("person") + for output in US_OTHER_HEALTH_INSURANCE_OUTPUT_COLUMNS + ): + return False + return us_other_health_insurance_signal_gate(frame).passed diff --git a/packages/populace-build/src/populace/build/us_runtime/prior_year_income.py b/packages/populace-build/src/populace/build/us_runtime/prior_year_income.py new file mode 100644 index 00000000..ab009b24 --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/prior_year_income.py @@ -0,0 +1,964 @@ +"""Adjacent-year CPS ASEC earnings carried by the retired eCPS build. + +The archived pipeline joins each current ASEC person to the preceding ASEC +file by ``PERIDNUM``. It accepts a prior observation only when both Census +allocation flags are zero, treats ``-1`` and ``-9999`` as unavailable, records +whether both prior earnings amounts were observed, and otherwise falls back to +the current ASEC values. Self-employment losses are source values, not errors, +and remain signed. + +After the PUF support clone is created, the retired second-stage CPS-only QRF +jointly replaces the wage and self-employment prior-year amounts on the PUF +half using eight demographic/income predictors and a 5,000-person training +cap. Populace preserves that treatment and strengthens it with typed person +design weights. ``employment_income_last_year`` is formula-owned in +PolicyEngine-US 1.764.6 and the retired finalizer explicitly dropped it, so it +is retained only through the joint fit and removed from the export-ready +support frame. The two persisted input leaves are +``self_employment_income_last_year`` and ``previous_year_income_available``. +""" + +from __future__ import annotations + +from importlib.resources import files +from typing import Any + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import ( + SourceOperationSpec, + SourceStageSpec, + load_source_manifest, +) +from populace.build.source_runtime import ( + SourceRuntimeConfig, + SourceRuntimeContext, + SourceRuntimeError, + run_source_stage, +) +from populace.frame import Frame +from populace.frame.units import US_SCHEMA + +__all__ = [ + "PRIOR_YEAR_INCOME_ARCHIVED_DERIVATION_URL", + "PRIOR_YEAR_INCOME_ARCHIVED_PUF_IMPUTATION_URL", + "PRIOR_YEAR_INCOME_ARCHIVED_PUF_OUTPUTS_URL", + "PRIOR_YEAR_INCOME_ARCHIVED_PUF_SPLICE_URL", + "PRIOR_YEAR_INCOME_ARCHIVED_FORMULA_OUTPUT_URL", + "PRIOR_YEAR_INCOME_ARCHIVED_FINALIZER_URL", + "US_PRIOR_YEAR_INCOME_NONCONSTANT_PERSON_COLUMNS", + "US_PRIOR_YEAR_INCOME_OUTPUT_COLUMNS", + "US_PRIOR_YEAR_INCOME_PERSISTED_OUTPUT_COLUMNS", + "US_PRIOR_YEAR_INCOME_REQUIRED_SOURCE_COLUMNS", + "US_PRIOR_YEAR_INCOME_STAGE_NAME", + "derive_us_prior_year_income_from_manifest", + "impute_us_prior_year_income_to_puf_support_from_manifest", + "us_prior_year_income_signal_gate", + "us_prior_year_income_source_reconciliation_gate", + "us_prior_year_income_stage_spec", + "us_prior_year_income_summary", + "with_us_prior_year_income_inputs", +] + +QRF: Any | None = None + +_ARCHIVED_COMMIT = "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe" +_ARCHIVED_REPOSITORY = "policyengine-" + "us-data" +_ARCHIVED_PACKAGE = "policyengine_" + "us_data" +_ARCHIVED_ROOT = ( + f"https://github.com/PolicyEngine/{_ARCHIVED_REPOSITORY}/blob/" + f"{_ARCHIVED_COMMIT}/{_ARCHIVED_PACKAGE}/datasets/cps" +) +PRIOR_YEAR_INCOME_ARCHIVED_DERIVATION_URL = _ARCHIVED_ROOT + "/cps.py#L1680-L1783" +PRIOR_YEAR_INCOME_ARCHIVED_PUF_OUTPUTS_URL = ( + _ARCHIVED_ROOT + "/extended_cps.py#L140-L194" +) +PRIOR_YEAR_INCOME_ARCHIVED_PUF_IMPUTATION_URL = ( + _ARCHIVED_ROOT + "/extended_cps.py#L639-L745" +) +PRIOR_YEAR_INCOME_ARCHIVED_PUF_SPLICE_URL = ( + _ARCHIVED_ROOT + "/extended_cps.py#L1014-L1073" +) +PRIOR_YEAR_INCOME_ARCHIVED_FORMULA_OUTPUT_URL = ( + _ARCHIVED_ROOT + "/extended_cps.py#L837-L848" +) +PRIOR_YEAR_INCOME_ARCHIVED_FINALIZER_URL = ( + _ARCHIVED_ROOT + "/extended_cps.py#L1463-L1499" +) + +US_PRIOR_YEAR_INCOME_STAGE_NAME = "prior_year_income" +US_PRIOR_YEAR_INCOME_OUTPUT_COLUMNS: tuple[str, ...] = ( + "employment_income_last_year", + "self_employment_income_last_year", + "previous_year_income_available", +) +US_PRIOR_YEAR_INCOME_PERSISTED_OUTPUT_COLUMNS: tuple[str, ...] = ( + "self_employment_income_last_year", + "previous_year_income_available", +) +US_PRIOR_YEAR_INCOME_NONCONSTANT_PERSON_COLUMNS = ( + US_PRIOR_YEAR_INCOME_PERSISTED_OUTPUT_COLUMNS +) +US_PRIOR_YEAR_INCOME_REQUIRED_SOURCE_COLUMNS: tuple[str, ...] = ( + "source_year", + "PERIDNUM", + "WSAL_VAL", + "SEMP_VAL", + "I_ERNVAL", + "I_SEVAL", +) + +_PERSON_WEIGHT_COLUMN = "person_weight" +_PERSON_SUPPORT_CHANNEL_COLUMN = "person_support_channel" +_PERSON_SUPPORT_SOURCE_ID_COLUMN = "person_source_id" +_BASE_ASEC_SUPPORT_CHANNEL = "asec" +_PUF_TAX_DETAIL_SUPPORT_CHANNEL = "puf_tax_detail" +_FORMULA_OWNED_OUTPUT = "employment_income_last_year" +_PUF_QRF_OUTPUT_COLUMNS: tuple[str, ...] = ( + "employment_income_last_year", + "self_employment_income_last_year", +) +_PUF_PREDICTORS: tuple[str, ...] = ( + "age", + "is_male", + "has_esi", + "tax_unit_is_joint", + "tax_unit_count_dependents", + "employment_income", + "self_employment_income", + "social_security", +) +_PUF_PREDICTOR_PREFIX = "prior_year_income_predictor_" +_DERIVE_PARAMETER_KEYS = frozenset( + { + "person_id", + "source_year", + "prior_year_offset", + "employment_source", + "self_employment_source", + "employment_allocation_flag", + "self_employment_allocation_flag", + "unallocated_flag", + "sentinels", + "fallback_to_current", + "no_prior_artifact", + } +) +_EXPECTED_DERIVE_PARAMETERS: dict[str, Any] = { + "person_id": "PERIDNUM", + "source_year": "source_year", + "prior_year_offset": -1, + "employment_source": "WSAL_VAL", + "self_employment_source": "SEMP_VAL", + "employment_allocation_flag": "I_ERNVAL", + "self_employment_allocation_flag": "I_SEVAL", + "unallocated_flag": 0, + "sentinels": [-1, -9999], + "fallback_to_current": True, + "no_prior_artifact": "leave_defaults", +} +_PUF_IMPUTATION_PARAMETER_KEYS = frozenset( + { + "predictors", + "outputs", + "max_train_samples", + "n_estimators", + "seed_from_build_config", + "weight", + } +) +_PREVIOUS_YEAR_AVAILABLE_SHARE_BAND = (0.05, 0.50) +_SELF_EMPLOYMENT_NONZERO_SHARE_BAND = (0.01, 0.25) + + +def us_prior_year_income_stage_spec() -> SourceStageSpec: + """Load and validate the packaged prior-year-income stage declaration.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + stage_map = manifest.stage_map() + if US_PRIOR_YEAR_INCOME_STAGE_NAME not in stage_map: + raise ValueError( + f"US source manifest declares no {US_PRIOR_YEAR_INCOME_STAGE_NAME!r} stage." + ) + spec = stage_map[US_PRIOR_YEAR_INCOME_STAGE_NAME] + if tuple(spec.outputs) != US_PRIOR_YEAR_INCOME_OUTPUT_COLUMNS: + raise ValueError( + "prior_year_income manifest outputs do not match the runtime-owned family." + ) + if "self_employment_income_last_year" in spec.nonnegative_outputs: + raise ValueError( + "self_employment_income_last_year must remain signed; the archived " + "ASEC derivation preserves reported losses." + ) + return spec + + +def _validated_derive_parameters(operation: SourceOperationSpec) -> None: + unexpected = sorted(set(operation.parameters) - _DERIVE_PARAMETER_KEYS) + missing = sorted(_DERIVE_PARAMETER_KEYS - set(operation.parameters)) + if unexpected or missing: + raise SourceRuntimeError( + "US prior-year-income derivation parameters drifted from the " + f"archived method; missing={missing}, unexpected={unexpected}." + ) + if dict(operation.parameters) != _EXPECTED_DERIVE_PARAMETERS: + raise SourceRuntimeError( + "US prior-year-income derivation parameters must exactly match " + f"the archived method: expected {_EXPECTED_DERIVE_PARAMETERS}, " + f"got {dict(operation.parameters)}." + ) + + +def _numeric_source(frame: pd.DataFrame, column: str) -> pd.Series: + return pd.to_numeric(frame[column], errors="coerce").astype("float64") + + +def derive_us_prior_year_income_from_manifest( + frame: pd.DataFrame | None, + operation: SourceOperationSpec, + _context: SourceRuntimeContext | None, +) -> pd.DataFrame: + """Reproduce the archived adjacent-ASEC ``PERIDNUM`` earnings join.""" + + if operation.kind != "derive_prior_year_income": + raise SourceRuntimeError( + "US prior-year-income derivation received unexpected operation " + f"{operation.kind!r}." + ) + if frame is None: + raise SourceRuntimeError( + "US prior-year-income derivation requires the pooled person table." + ) + _validated_derive_parameters(operation) + missing = [ + column + for column in US_PRIOR_YEAR_INCOME_REQUIRED_SOURCE_COLUMNS + if column not in frame.columns + ] + if missing: + raise SourceRuntimeError( + "US prior-year-income derivation requires measured ASEC source " + f"column(s): {missing}." + ) + if _PERSON_SUPPORT_CHANNEL_COLUMN in frame.columns: + absent = [ + column + for column in US_PRIOR_YEAR_INCOME_OUTPUT_COLUMNS + if column not in frame.columns + ] + if absent: + raise SourceRuntimeError( + "US prior-year income must be derived before support cloning; " + f"the expanded frame is missing {absent}." + ) + return frame.copy(deep=True) + + result = frame.copy(deep=True) + source_year = pd.to_numeric(result["source_year"], errors="coerce") + valid_year = np.isfinite(source_year.to_numpy(dtype=np.float64)) & ( + source_year.to_numpy(dtype=np.float64) + == np.floor(source_year.to_numpy(dtype=np.float64)) + ) + if not valid_year.all(): + rows = np.flatnonzero(~valid_year)[:5].tolist() + raise SourceRuntimeError( + f"US prior-year-income source_year is invalid at row(s) {rows}." + ) + if result["PERIDNUM"].isna().any(): + rows = result.index[result["PERIDNUM"].isna()].tolist()[:5] + raise SourceRuntimeError( + f"US prior-year-income PERIDNUM is missing at row(s) {rows}." + ) + key = pd.MultiIndex.from_arrays( + [source_year.astype("int64"), result["PERIDNUM"].astype(str)], + names=["source_year", "PERIDNUM"], + ) + duplicated = key.duplicated(keep=False) + if duplicated.any(): + examples = list(dict.fromkeys(key[duplicated].tolist()))[:5] + raise SourceRuntimeError( + "US prior-year-income join requires unique (source_year, PERIDNUM) " + f"keys; duplicate key(s): {examples}." + ) + + employment = _numeric_source(result, "WSAL_VAL") + self_employment = _numeric_source(result, "SEMP_VAL") + employment_flag = _numeric_source(result, "I_ERNVAL") + self_employment_flag = _numeric_source(result, "I_SEVAL") + eligible_prior = employment_flag.eq(0) & self_employment_flag.eq(0) + + prior = pd.DataFrame( + { + "source_year": source_year.loc[eligible_prior].astype("int64") + 1, + "PERIDNUM": result.loc[eligible_prior, "PERIDNUM"].astype(str), + "employment_income_last_year": employment.loc[eligible_prior], + "self_employment_income_last_year": self_employment.loc[eligible_prior], + } + ).set_index(["source_year", "PERIDNUM"]) + current_key = pd.MultiIndex.from_arrays( + [source_year.astype("int64"), result["PERIDNUM"].astype(str)], + names=["source_year", "PERIDNUM"], + ) + aligned = prior.reindex(current_key).reset_index(drop=True) + + sentinels = {-1.0, -9999.0} + invalid_prior = aligned["employment_income_last_year"].isin(sentinels) | aligned[ + "self_employment_income_last_year" + ].isin(sentinels) + aligned.loc[ + invalid_prior, + ["employment_income_last_year", "self_employment_income_last_year"], + ] = np.nan + available = ( + aligned["employment_income_last_year"].notna() + & aligned["self_employment_income_last_year"].notna() + ) + + current_employment = employment.mask(employment.isin(sentinels)) + current_self_employment = self_employment.mask(self_employment.isin(sentinels)) + source_year_values = set(source_year.astype("int64").tolist()) + cohorts_with_prior_artifact = {prior_year + 1 for prior_year in source_year_values} + can_fallback = source_year.astype("int64").isin(cohorts_with_prior_artifact) + current_employment = current_employment.where(can_fallback) + current_self_employment = current_self_employment.where(can_fallback) + result["employment_income_last_year"] = ( + aligned["employment_income_last_year"] + .fillna(current_employment.reset_index(drop=True)) + .fillna(0.0) + .to_numpy(dtype=np.float64) + ) + result["self_employment_income_last_year"] = ( + aligned["self_employment_income_last_year"] + .fillna(current_self_employment.reset_index(drop=True)) + .fillna(0.0) + .to_numpy(dtype=np.float64) + ) + result["previous_year_income_available"] = available.to_numpy(dtype=bool) + if (result["employment_income_last_year"] < 0).any(): + rows = result.index[result["employment_income_last_year"] < 0].tolist()[:5] + raise SourceRuntimeError( + "US prior-year wage income must be nonnegative after sentinel " + f"replacement; negative value(s) at row(s) {rows}." + ) + return result + + +def impute_us_prior_year_income_to_puf_support_from_manifest( + frame: pd.DataFrame | None, + operation: SourceOperationSpec, + context: SourceRuntimeContext | None, +) -> pd.DataFrame: + """Jointly QRF-impute prior-year earnings on the PUF support half.""" + + if operation.kind != "impute_prior_year_income_to_puf_support": + raise SourceRuntimeError( + "US prior-year-income PUF imputation received unexpected operation " + f"{operation.kind!r}." + ) + if frame is None: + raise SourceRuntimeError( + "US prior-year-income PUF imputation requires the person table." + ) + unexpected = sorted(set(operation.parameters) - _PUF_IMPUTATION_PARAMETER_KEYS) + missing_parameters = sorted( + _PUF_IMPUTATION_PARAMETER_KEYS - set(operation.parameters) + ) + if unexpected or missing_parameters: + raise SourceRuntimeError( + "US prior-year-income PUF imputation parameters must match the " + f"archived method; missing={missing_parameters}, " + f"unexpected={unexpected}." + ) + if _PERSON_SUPPORT_CHANNEL_COLUMN not in frame.columns: + return frame.copy(deep=True) + + predictors = tuple(str(value) for value in operation.parameters["predictors"]) + outputs = tuple(str(value) for value in operation.parameters["outputs"]) + if predictors != _PUF_PREDICTORS or outputs != _PUF_QRF_OUTPUT_COLUMNS: + raise SourceRuntimeError( + "US prior-year-income PUF predictors/outputs drifted from the " + f"archived joint fit: predictors={list(predictors)}, " + f"outputs={list(outputs)}." + ) + if operation.parameters["weight"] != _PERSON_WEIGHT_COLUMN: + raise SourceRuntimeError( + "US prior-year-income PUF imputation must use typed person weights." + ) + if operation.parameters["seed_from_build_config"] is not True: + raise SourceRuntimeError( + "US prior-year-income PUF imputation seed must come from the build." + ) + max_train_samples = int(operation.parameters["max_train_samples"]) + n_estimators = int(operation.parameters["n_estimators"]) + if max_train_samples <= 0 or n_estimators <= 0: + raise SourceRuntimeError("US prior-year-income PUF fit sizes must be positive.") + + predictor_columns = [_PUF_PREDICTOR_PREFIX + value for value in predictors] + required = [ + _PERSON_WEIGHT_COLUMN, + *predictor_columns, + *_PUF_QRF_OUTPUT_COLUMNS, + ] + missing = [column for column in required if column not in frame.columns] + if missing: + raise SourceRuntimeError( + "US prior-year-income PUF imputation is missing source column(s): " + f"{missing}." + ) + + channel = frame[_PERSON_SUPPORT_CHANNEL_COLUMN].astype(str) + asec_mask = channel.eq(_BASE_ASEC_SUPPORT_CHANNEL) + puf_mask = channel.eq(_PUF_TAX_DETAIL_SUPPORT_CHANNEL) + if not asec_mask.any() or not puf_mask.any(): + raise SourceRuntimeError( + "US prior-year-income PUF imputation requires nonempty ASEC and " + "PUF-tax-detail support channels." + ) + + training = frame.loc[ + asec_mask, + [*predictor_columns, *_PUF_QRF_OUTPUT_COLUMNS], + ].copy() + training.columns = [*predictors, *_PUF_QRF_OUTPUT_COLUMNS] + weights = pd.to_numeric( + frame.loc[asec_mask, _PERSON_WEIGHT_COLUMN], errors="coerce" + ) + weight_values = weights.to_numpy(dtype=np.float64) + if ( + not np.isfinite(weight_values).all() + or (weight_values < 0).any() + or float(weight_values.sum()) <= 0 + ): + raise SourceRuntimeError( + "US prior-year-income QRF requires finite, nonnegative typed " + "person weights with positive total mass." + ) + if len(training) > max_train_samples: + sample = training.sample( + n=max_train_samples, + random_state=(context.config.seed if context is not None else 0), + ).index + training = training.loc[sample] + weights = weights.loc[sample] + sampled_weight_values = weights.to_numpy(dtype=np.float64) + if ( + not np.isfinite(sampled_weight_values).all() + or (sampled_weight_values < 0).any() + or float(sampled_weight_values.sum()) <= 0 + ): + raise SourceRuntimeError( + "US prior-year-income QRF sampled training weights must be " + "finite and nonnegative with positive total mass." + ) + + test = frame.loc[puf_mask, predictor_columns].copy() + test.columns = list(predictors) + for column in (*predictors, *_PUF_QRF_OUTPUT_COLUMNS): + training[column] = pd.to_numeric(training[column], errors="coerce") + if not np.isfinite(training[column].to_numpy(dtype=np.float64)).all(): + raise SourceRuntimeError( + f"US prior-year-income QRF training column {column!r} " + "contains nonfinite values." + ) + for column in predictors: + test[column] = pd.to_numeric(test[column], errors="coerce") + if not np.isfinite(test[column].to_numpy(dtype=np.float64)).all(): + raise SourceRuntimeError( + f"US prior-year-income QRF prediction column {column!r} " + "contains nonfinite values." + ) + + global QRF + if QRF is None: + from importlib import import_module + + QRF = import_module("populace.fit").QRF + seed = context.config.seed if context is not None else 0 + fitted = QRF(n_estimators=n_estimators, seed=seed).fit( + training, + list(predictors), + list(_PUF_QRF_OUTPUT_COLUMNS), + weights=weights.to_numpy(dtype=np.float64), + ) + predictions = fitted.predict(test) + if not isinstance(predictions, pd.DataFrame): + raise SourceRuntimeError( + "US prior-year-income QRF must return a pandas DataFrame." + ) + missing_predictions = [ + column for column in _PUF_QRF_OUTPUT_COLUMNS if column not in predictions + ] + if missing_predictions: + raise SourceRuntimeError( + f"US prior-year-income QRF omitted output column(s): {missing_predictions}." + ) + expected_predictions = int(puf_mask.sum()) + if len(predictions) != expected_predictions: + raise SourceRuntimeError( + "US prior-year-income QRF returned " + f"{len(predictions)} rows; expected {expected_predictions}." + ) + for column in _PUF_QRF_OUTPUT_COLUMNS: + values = pd.to_numeric(predictions[column], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not np.isfinite(values).all(): + raise SourceRuntimeError( + f"US prior-year-income QRF output {column!r} contains nonfinite values." + ) + if column == _FORMULA_OWNED_OUTPUT and (values < 0).any(): + rows = np.flatnonzero(values < 0)[:5].tolist() + raise SourceRuntimeError( + "US prior-year-income QRF predicted negative wage income at " + f"PUF prediction row(s) {rows}." + ) + + result = frame.copy(deep=True) + for column in _PUF_QRF_OUTPUT_COLUMNS: + result.loc[puf_mask, column] = pd.to_numeric( + predictions[column], errors="coerce" + ).to_numpy(dtype=np.float64) + return result + + +def _person_prior_year_income_predictors(frame: Frame) -> pd.DataFrame: + """Build the archived eight second-stage predictors on person rows.""" + + person = frame.table("person") + tax_unit = frame.table("tax_unit") + predictors = pd.DataFrame(index=person.index) + + def _numeric(*columns: str) -> np.ndarray: + for column in columns: + if column in person: + values = pd.to_numeric(person[column], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not np.isfinite(values).all(): + raise SourceRuntimeError( + "US prior-year-income QRF predictor source " + f"{column!r} contains nonfinite values." + ) + return values + raise SourceRuntimeError( + "US prior-year-income PUF imputation cannot construct a predictor " + f"from any of {list(columns)}." + ) + + predictors["age"] = _numeric("age", "A_AGE") + if "is_male" in person: + predictors["is_male"] = person["is_male"].fillna(False).astype(bool) + elif "is_female" in person: + predictors["is_male"] = ~person["is_female"].fillna(False).astype(bool) + elif "A_SEX" in person: + predictors["is_male"] = _numeric("A_SEX") == 1 + else: + raise SourceRuntimeError( + "US prior-year-income PUF imputation requires is_male, is_female, " + "or measured A_SEX." + ) + if "has_esi" not in person: + raise SourceRuntimeError( + "US prior-year-income PUF imputation requires has_esi." + ) + predictors["has_esi"] = person["has_esi"].fillna(False).astype(bool) + + filing_status_column = next( + ( + column + for column in ("filing_status_input", "filing_status") + if column in tax_unit + ), + None, + ) + if filing_status_column is None: + raise SourceRuntimeError( + "US prior-year-income PUF imputation requires tax-unit filing status." + ) + filing_status = tax_unit[filing_status_column].map( + lambda value: ( + value.decode() if isinstance(value, (bytes, np.bytes_)) else str(value) + ) + ) + joint_by_id = pd.Series( + filing_status.eq("JOINT").to_numpy(), + index=tax_unit["tax_unit_id"].to_numpy(), + ) + predictors["tax_unit_is_joint"] = ( + person["person_tax_unit_id"].map(joint_by_id).fillna(False).astype(bool) + ) + + if "tax_unit_role_input" not in person: + raise SourceRuntimeError( + "US prior-year-income PUF imputation requires tax_unit_role_input." + ) + dependent = ( + person["tax_unit_role_input"] + .map( + lambda value: ( + value.decode() if isinstance(value, (bytes, np.bytes_)) else str(value) + ) + ) + .eq("DEPENDENT") + ) + predictors["tax_unit_count_dependents"] = dependent.groupby( + person["person_tax_unit_id"] + ).transform("sum") + predictors["employment_income"] = _numeric( + "employment_income_before_lsr", "WSAL_VAL" + ) + predictors["self_employment_income"] = _numeric( + "self_employment_income_before_lsr", "SEMP_VAL" + ) + social_security_columns = ( + "social_security_retirement", + "social_security_disability", + "social_security_dependents", + "social_security_survivors", + ) + if all(column in person for column in social_security_columns): + predictors["social_security"] = np.sum( + np.column_stack([_numeric(column) for column in social_security_columns]), + axis=1, + ) + else: + predictors["social_security"] = _numeric("SS_VAL") + return predictors + + +def _replace_person_table(frame: Frame, person: pd.DataFrame) -> Frame: + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + tables["person"] = person + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + ) + + +def with_us_prior_year_income_inputs( + frame: Frame, + *, + seed: int, + time_period: int, +) -> Frame: + """Materialize adjacent-year earnings and PUF-support replacements.""" + + if frame.schema != US_SCHEMA: + raise ValueError("US prior-year income requires the US schema.") + person = frame.table("person") + has_support_channels = _PERSON_SUPPORT_CHANNEL_COLUMN in person.columns + has_raw_sources = all( + column in person for column in US_PRIOR_YEAR_INCOME_REQUIRED_SOURCE_COLUMNS + ) + if ( + not has_support_channels + and all(column in person for column in US_PRIOR_YEAR_INCOME_OUTPUT_COLUMNS) + and not has_raw_sources + ): + return frame + + stage_person = person.copy(deep=True) + stage_person[_PERSON_WEIGHT_COLUMN] = frame.resolve_weights("person").values + if has_support_channels: + predictors = _person_prior_year_income_predictors(frame) + for column in _PUF_PREDICTORS: + stage_person[_PUF_PREDICTOR_PREFIX + column] = predictors[column].to_numpy() + output = run_source_stage( + us_prior_year_income_stage_spec(), + tables={"person": stage_person}, + operation_handlers={ + "derive_prior_year_income": derive_us_prior_year_income_from_manifest, + "impute_prior_year_income_to_puf_support": ( + impute_us_prior_year_income_to_puf_support_from_manifest + ), + }, + config=SourceRuntimeConfig(seed=seed, target_year=time_period), + ).drop( + columns=[ + _PERSON_WEIGHT_COLUMN, + *(_PUF_PREDICTOR_PREFIX + value for value in _PUF_PREDICTORS), + ], + errors="ignore", + ) + if has_support_channels: + output = output.drop(columns=[_FORMULA_OWNED_OUTPUT]) + return _replace_person_table(frame, output) + + +def us_prior_year_income_summary(frame: Frame) -> dict[str, object]: + """Return weighted signal, signed mass, and support-channel diagnostics.""" + + person = frame.table("person") + weights = np.asarray(frame.resolve_weights("person").values, dtype=np.float64) + total_weight = float(weights.sum()) + availability = person["previous_year_income_available"].astype(bool).to_numpy() + self_employment = pd.to_numeric( + person["self_employment_income_last_year"], errors="coerce" + ).to_numpy(dtype=np.float64) + nonzero = self_employment != 0 + + def _share(mask: np.ndarray) -> float: + return float(weights[mask].sum()) / total_weight if total_weight > 0 else 0.0 + + channels: dict[str, dict[str, float | int]] = {} + if _PERSON_SUPPORT_CHANNEL_COLUMN in person: + support_channel = person[_PERSON_SUPPORT_CHANNEL_COLUMN].astype(str).to_numpy() + for channel in (_BASE_ASEC_SUPPORT_CHANNEL, _PUF_TAX_DETAIL_SUPPORT_CHANNEL): + mask = support_channel == channel + channel_weight = float(weights[mask].sum()) + channels[channel] = { + "rows": int(mask.sum()), + "weighted_population": channel_weight, + "availability_share": ( + float(weights[mask & availability].sum()) / channel_weight + if channel_weight > 0 + else 0.0 + ), + "self_employment_nonzero_share": ( + float(weights[mask & nonzero].sum()) / channel_weight + if channel_weight > 0 + else 0.0 + ), + "self_employment_negative_rows": int( + np.count_nonzero(mask & (self_employment < 0)) + ), + } + + clone_availability_mismatches = 0 + if _PERSON_SUPPORT_SOURCE_ID_COLUMN in person: + clone_availability_mismatches = int( + person.assign(_availability=availability) + .groupby(_PERSON_SUPPORT_SOURCE_ID_COLUMN, sort=False)["_availability"] + .nunique() + .gt(1) + .sum() + ) + + return { + "rows": int(len(person)), + "weighted_population": total_weight, + "previous_year_income_available_share": _share(availability), + "previous_year_income_available_share_band": list( + _PREVIOUS_YEAR_AVAILABLE_SHARE_BAND + ), + "self_employment_income_last_year_nonzero_share": _share(nonzero), + "self_employment_income_last_year_nonzero_share_band": list( + _SELF_EMPLOYMENT_NONZERO_SHARE_BAND + ), + "self_employment_income_last_year_positive_rows": int( + np.count_nonzero(self_employment > 0) + ), + "self_employment_income_last_year_negative_rows": int( + np.count_nonzero(self_employment < 0) + ), + "self_employment_income_last_year_weighted_total": float( + np.dot(weights, self_employment) + ), + "unique_counts": { + column: int(person[column].dropna().nunique()) + for column in US_PRIOR_YEAR_INCOME_PERSISTED_OUTPUT_COLUMNS + }, + "channels": channels, + "clone_availability_mismatches": clone_availability_mismatches, + } + + +def us_prior_year_income_signal_gate(frame: Frame) -> GateResult: + """Fail closed on missing, default-only, or implausible prior-year inputs.""" + + person = frame.table("person") + missing = [ + column + for column in US_PRIOR_YEAR_INCOME_PERSISTED_OUTPUT_COLUMNS + if column not in person + ] + if missing: + return GateResult( + name="prior_year_income_signal", + passed=False, + failures=(f"person columns missing: {missing}.",), + details={"missing": missing}, + ) + availability_source = person["previous_year_income_available"] + if availability_source.isna().any() or not pd.api.types.is_bool_dtype( + availability_source.dtype + ): + return GateResult( + name="prior_year_income_signal", + passed=False, + failures=( + "previous_year_income_available must be a complete boolean " + "source column.", + ), + details={ + "dtype": str(availability_source.dtype), + "missing": int(availability_source.isna().sum()), + }, + ) + self_employment = pd.to_numeric( + person["self_employment_income_last_year"], errors="coerce" + ).to_numpy(dtype=np.float64) + if not np.isfinite(self_employment).all(): + rows = np.flatnonzero(~np.isfinite(self_employment))[:5].tolist() + return GateResult( + name="prior_year_income_signal", + passed=False, + failures=( + "self_employment_income_last_year contains nonfinite value(s) " + f"at row(s) {rows}.", + ), + details={"nonfinite_rows": rows}, + ) + + summary = us_prior_year_income_summary(frame) + failures: list[str] = [] + checks = ( + ( + "previous_year_income_available_share", + "previous_year_income_available_share_band", + "previous-year availability weighted share", + ), + ( + "self_employment_income_last_year_nonzero_share", + "self_employment_income_last_year_nonzero_share_band", + "prior-year self-employment nonzero weighted share", + ), + ) + for share_key, band_key, label in checks: + share = float(summary[share_key]) + lower, upper = summary[band_key] + if not lower <= share <= upper: + failures.append(f"{label} {share:.6f} outside [{lower:.6f}, {upper:.6f}].") + for column, count in summary["unique_counts"].items(): + if int(count) < 2: + failures.append(f"{column} is degenerate with {count} distinct value(s).") + channels = summary["channels"] + if channels: + for channel in (_BASE_ASEC_SUPPORT_CHANNEL, _PUF_TAX_DETAIL_SUPPORT_CHANNEL): + channel_summary = channels.get(channel) + if not channel_summary or int(channel_summary["rows"]) == 0: + failures.append(f"{channel} prior-year-income support is empty.") + continue + if float(channel_summary["self_employment_nonzero_share"]) <= 0: + failures.append( + f"{channel} self_employment_income_last_year is default-only." + ) + if float(channel_summary["availability_share"]) <= 0: + failures.append( + f"{channel} previous_year_income_available is default-only." + ) + asec = channels.get(_BASE_ASEC_SUPPORT_CHANNEL) + if asec and int(asec["self_employment_negative_rows"]) == 0: + failures.append( + "ASEC self_employment_income_last_year lost all signed loss signal." + ) + elif int(summary["self_employment_income_last_year_negative_rows"]) == 0: + failures.append("self_employment_income_last_year lost all signed loss signal.") + mismatches = int(summary["clone_availability_mismatches"]) + if mismatches: + failures.append( + f"{mismatches} source person(s) disagree on prior-year availability " + "across support clones." + ) + return GateResult( + name="prior_year_income_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) + + +def us_prior_year_income_source_reconciliation_gate(frame: Frame) -> GateResult: + """Recompute the ASEC carry and require exact persisted-value agreement. + + This runs on the reusable base before sparse support selection, while all + adjacent-year source rows are still present. Only the ASEC channel is + reconciled: PUF-support amounts are deliberately owned by the joint QRF. + """ + + person = frame.table("person") + if _PERSON_SUPPORT_CHANNEL_COLUMN in person: + asec_mask = ( + person[_PERSON_SUPPORT_CHANNEL_COLUMN] + .astype(str) + .eq(_BASE_ASEC_SUPPORT_CHANNEL) + ) + actual = person.loc[asec_mask].reset_index(drop=True).copy() + else: + actual = person.reset_index(drop=True).copy() + missing_sources = [ + column + for column in US_PRIOR_YEAR_INCOME_REQUIRED_SOURCE_COLUMNS + if column not in actual + ] + compare_columns = [ + column for column in US_PRIOR_YEAR_INCOME_OUTPUT_COLUMNS if column in actual + ] + missing_outputs = sorted( + set(US_PRIOR_YEAR_INCOME_PERSISTED_OUTPUT_COLUMNS) - set(compare_columns) + ) + if missing_sources or missing_outputs: + failures = [] + if missing_sources: + failures.append(f"ASEC source columns missing: {missing_sources}.") + if missing_outputs: + failures.append(f"ASEC persisted outputs missing: {missing_outputs}.") + return GateResult( + name="prior_year_income_source_reconciliation", + passed=False, + failures=tuple(failures), + details={ + "missing_sources": missing_sources, + "missing_outputs": missing_outputs, + }, + ) + + source = actual.drop( + columns=[ + *US_PRIOR_YEAR_INCOME_OUTPUT_COLUMNS, + _PERSON_SUPPORT_CHANNEL_COLUMN, + ], + errors="ignore", + ) + operation = next( + operation + for operation in us_prior_year_income_stage_spec().operations + if operation.kind == "derive_prior_year_income" + ) + expected = derive_us_prior_year_income_from_manifest(source, operation, None) + mismatch_counts: dict[str, int] = {} + for column in compare_columns: + if column == "previous_year_income_available": + mismatches = actual[column].to_numpy(dtype=bool) != expected[ + column + ].to_numpy(dtype=bool) + else: + observed = pd.to_numeric(actual[column], errors="coerce").to_numpy( + dtype=np.float64 + ) + target = expected[column].to_numpy(dtype=np.float64) + mismatches = ~np.isclose(observed, target, rtol=0.0, atol=0.0) + mismatch_counts[column] = int(np.count_nonzero(mismatches)) + failures = tuple( + f"ASEC {column} differs from the archived adjacent-year derivation on " + f"{count} row(s)." + for column, count in mismatch_counts.items() + if count + ) + return GateResult( + name="prior_year_income_source_reconciliation", + passed=not failures, + failures=failures, + details={ + "asec_rows": int(len(actual)), + "compared_columns": compare_columns, + "mismatch_counts": mismatch_counts, + }, + ) diff --git a/packages/populace-build/src/populace/build/us_runtime/puf_aggregate_records.py b/packages/populace-build/src/populace/build/us_runtime/puf_aggregate_records.py index 57a138d2..96e23c11 100644 --- a/packages/populace-build/src/populace/build/us_runtime/puf_aggregate_records.py +++ b/packages/populace-build/src/populace/build/us_runtime/puf_aggregate_records.py @@ -17,6 +17,27 @@ import numpy as np import pandas as pd +from populace.build.us_runtime.alimony import derive_us_alimony_from_puf +from populace.build.us_runtime.capital_gain_details import ( + derive_us_capital_gain_details_from_puf, +) +from populace.build.us_runtime.casualty_losses import ( + derive_us_casualty_loss_from_puf, +) +from populace.build.us_runtime.domestic_production import ( + derive_us_domestic_production_ald_from_puf, +) +from populace.build.us_runtime.educator_expenses import ( + derive_us_educator_expense_from_puf, +) +from populace.build.us_runtime.farm_business_income import ( + derive_us_farm_business_income_from_puf, +) +from populace.build.us_runtime.form_4952 import derive_us_form_4952_election_from_puf +from populace.build.us_runtime.misc_itemized import derive_us_misc_itemized_from_puf +from populace.build.us_runtime.salt_refund_income import ( + derive_us_salt_refund_income_from_puf, +) from populace.calibrate import relative_error_loss __all__ = [ @@ -69,6 +90,7 @@ "E00400": "tax_exempt_interest_income", "E00600": "ordinary_dividends", "E00650": "qualified_dividends", + "E00700": "state_and_local_tax_refund_income", "E00900": "business_net_profits", "E01100": "capital_gains_distributions", "E01400": "ira_distributions", @@ -78,6 +100,7 @@ "E02400": "total_social_security", "E02500": "taxable_social_security", "E03210": "student_loan_interest", + "E03220": "educator_expense", "E17500": "medical_expense_deduction", "E18400": "state_income_tax_paid", "E18500": "real_estate_taxes_paid", @@ -89,7 +112,7 @@ "E24515": "unrecaptured_section_1250_gain", "E24518": "collectibles_capital_gains", "E26270": "partnership_and_s_corp_income", - "E87521": "net_investment_income_tax", + "E87521": "american_opportunity_credit", } _COMBINED_SOURCE_FIELDS = { "capital_gains_proxy": ("P22250", "P23250", "E01100"), @@ -209,6 +232,37 @@ def derive_puf_policyengine_variables( qualified_dividend_source: str = "E00650", qualified_dividend_output: str = "qualified_dividend_income", non_qualified_dividend_output: str = "non_qualified_dividend_income", + qualified_tuition_primary_source: str | None = None, + qualified_tuition_optional_source: str | None = None, + qualified_tuition_output: str = "qualified_tuition_expenses", + alimony_income_source: str | None = None, + alimony_income_output: str = "alimony_income", + alimony_expense_source: str | None = None, + alimony_expense_output: str = "alimony_expense", + casualty_loss_source: str | None = None, + casualty_loss_output: str = "casualty_loss", + domestic_production_ald_source: str | None = None, + domestic_production_ald_output: str = "domestic_production_ald", + educator_expense_source: str | None = None, + educator_expense_output: str = "educator_expense", + unreimbursed_business_employee_expenses_source: str | None = None, + unreimbursed_business_employee_expenses_output: str = ( + "unreimbursed_business_employee_expenses" + ), + farm_operations_income_source: str | None = None, + farm_operations_income_output: str = "farm_operations_income", + farm_rent_income_source: str | None = None, + farm_rent_income_output: str = "farm_rent_income", + investment_income_elected_form_4952_source: str | None = None, + investment_income_elected_form_4952_output: str = ( + "investment_income_elected_form_4952" + ), + salt_refund_income_source: str | None = None, + salt_refund_income_output: str = "salt_refund_income", + collectibles_capital_gain_source: str | None = None, + collectibles_capital_gain_output: str = "long_term_capital_gains_on_collectibles", + unrecaptured_section_1250_gain_source: str | None = None, + unrecaptured_section_1250_gain_output: str = "unrecaptured_section_1250_gain", ) -> pd.DataFrame: """Translate raw IRS PUF columns into PolicyEngine input variables.""" @@ -225,6 +279,106 @@ def derive_puf_policyengine_variables( result[qualified_dividend_output] = qualified result[non_qualified_dividend_output] = ordinary - qualified + if qualified_tuition_primary_source is not None: + _require_columns(result, [qualified_tuition_primary_source]) + tuition = _numeric_series(result[qualified_tuition_primary_source]).clip( + lower=0.0 + ) + if ( + qualified_tuition_optional_source is not None + and qualified_tuition_optional_source in result + ): + optional = _numeric_series(result[qualified_tuition_optional_source]).clip( + lower=0.0 + ) + tuition = pd.Series( + np.maximum(tuition.to_numpy(), optional.to_numpy()), + index=result.index, + dtype="float64", + ) + result[qualified_tuition_output] = tuition + alimony_sources = (alimony_income_source, alimony_expense_source) + if any(source is not None for source in alimony_sources): + if not all(source is not None for source in alimony_sources): + raise ValueError( + "PUF alimony derivation requires both income and expense source " + "columns when either is configured." + ) + result = derive_us_alimony_from_puf( + result, + income_source_column=alimony_income_source, + income_output_column=alimony_income_output, + expense_source_column=alimony_expense_source, + expense_output_column=alimony_expense_output, + ) + if casualty_loss_source is not None: + result = derive_us_casualty_loss_from_puf( + result, + source_column=casualty_loss_source, + output_column=casualty_loss_output, + ) + if domestic_production_ald_source is not None: + result = derive_us_domestic_production_ald_from_puf( + result, + source_column=domestic_production_ald_source, + output_column=domestic_production_ald_output, + ) + if educator_expense_source is not None: + result = derive_us_educator_expense_from_puf( + result, + source_column=educator_expense_source, + output_column=educator_expense_output, + ) + if unreimbursed_business_employee_expenses_source is not None: + result = derive_us_misc_itemized_from_puf( + result, + source_column=unreimbursed_business_employee_expenses_source, + output_column=unreimbursed_business_employee_expenses_output, + ) + if investment_income_elected_form_4952_source is not None: + result = derive_us_form_4952_election_from_puf( + result, + source_column=investment_income_elected_form_4952_source, + output_column=investment_income_elected_form_4952_output, + ) + if salt_refund_income_source is not None: + result = derive_us_salt_refund_income_from_puf( + result, + source_column=salt_refund_income_source, + output_column=salt_refund_income_output, + ) + capital_gain_detail_sources = ( + collectibles_capital_gain_source, + unrecaptured_section_1250_gain_source, + ) + if any(source is not None for source in capital_gain_detail_sources): + if not all(source is not None for source in capital_gain_detail_sources): + raise ValueError( + "PUF capital-gain detail derivation requires both collectibles " + "and unrecaptured-section-1250 source columns when either is " + "configured." + ) + result = derive_us_capital_gain_details_from_puf( + result, + collectibles_source_column=collectibles_capital_gain_source, + collectibles_output_column=collectibles_capital_gain_output, + unrecaptured_source_column=unrecaptured_section_1250_gain_source, + unrecaptured_output_column=unrecaptured_section_1250_gain_output, + ) + farm_sources = (farm_operations_income_source, farm_rent_income_source) + if any(source is not None for source in farm_sources): + if not all(source is not None for source in farm_sources): + raise ValueError( + "PUF farm-business derivation requires both operations and rent " + "source columns when either is configured." + ) + result = derive_us_farm_business_income_from_puf( + result, + operations_source_column=farm_operations_income_source, + operations_output_column=farm_operations_income_output, + rent_source_column=farm_rent_income_source, + rent_output_column=farm_rent_income_output, + ) return result @@ -281,7 +435,17 @@ def disaggregate_puf_aggregate_records( synthetic_df = pd.concat(pieces, ignore_index=True) result = pd.concat([regular, synthetic_df], ignore_index=True) - return _reconcile_puf_dividend_columns_from_components(result) + result = _reconcile_puf_dividend_columns_from_components(result) + result = _reconcile_puf_qualified_tuition_from_sources(result) + result = _reconcile_puf_alimony_from_sources(result) + result = _reconcile_puf_casualty_loss_from_source(result) + result = _reconcile_puf_domestic_production_ald_from_source(result) + result = _reconcile_puf_educator_expense_from_source(result) + result = _reconcile_puf_misc_itemized_from_source(result) + result = _reconcile_puf_form_4952_election_from_source(result) + result = _reconcile_puf_salt_refund_income_from_source(result) + result = _reconcile_puf_capital_gain_details_from_sources(result) + return _reconcile_puf_farm_business_income_from_sources(result) def audit_puf_aggregate_disaggregation( @@ -848,6 +1012,147 @@ def _reconcile_puf_dividend_columns_from_components(puf: pd.DataFrame) -> pd.Dat return result +def _reconcile_puf_qualified_tuition_from_sources( + puf: pd.DataFrame, +) -> pd.DataFrame: + """Recompute tuition after aggregate-row source amounts are replaced. + + ``derive_puf_policyengine_variables`` runs before aggregate-record + disaggregation. The disaggregator subsequently reallocates the raw + ``E03230``/``E87530`` amounts, so carrying the earlier derived column would + leave donor-template values that no longer agree with either source field. + Re-deriving here preserves the retired ``max(E03230, E87530)`` contract on + every regular and synthetic row. + """ + + output = "qualified_tuition_expenses" + primary = "E03230" + optional = "E87530" + if output not in puf.columns or primary not in puf.columns: + return puf + + result = puf.copy() + tuition = _numeric_series(result[primary]).clip(lower=0.0) + if optional in result.columns: + optional_tuition = _numeric_series(result[optional]).clip(lower=0.0) + tuition = pd.Series( + np.maximum(tuition.to_numpy(), optional_tuition.to_numpy()), + index=result.index, + dtype="float64", + ) + result[output] = tuition + return result + + +def _reconcile_puf_casualty_loss_from_source( + puf: pd.DataFrame, +) -> pd.DataFrame: + """Recompute casualty loss after aggregate-row amounts are replaced. + + The raw ``E20500`` amount participates in PUF aggregate-record + disaggregation. Reusing the pre-disaggregation derived column would leave + donor-template values on synthetic rows, so restore the archived direct + mapping after the source amounts reach their final rows. + """ + + if "casualty_loss" not in puf.columns or "E20500" not in puf.columns: + return puf + return derive_us_casualty_loss_from_puf(puf) + + +def _reconcile_puf_alimony_from_sources(puf: pd.DataFrame) -> pd.DataFrame: + """Recompute both alimony leaves after raw source amounts are replaced.""" + + required = {"alimony_income", "alimony_expense", "E00800", "E03500"} + if not required.issubset(puf.columns): + return puf + return derive_us_alimony_from_puf(puf) + + +def _reconcile_puf_domestic_production_ald_from_source( + puf: pd.DataFrame, +) -> pd.DataFrame: + """Recompute the E03240 carry after aggregate-row amounts are replaced.""" + + output = "domestic_production_ald" + if output not in puf.columns or "E03240" not in puf.columns: + return puf + return derive_us_domestic_production_ald_from_puf(puf) + + +def _reconcile_puf_educator_expense_from_source(puf: pd.DataFrame) -> pd.DataFrame: + """Recompute the E03220 carry after aggregate-row amounts are replaced.""" + + if "educator_expense" not in puf.columns or "E03220" not in puf.columns: + return puf + return derive_us_educator_expense_from_puf(puf) + + +def _reconcile_puf_misc_itemized_from_source( + puf: pd.DataFrame, +) -> pd.DataFrame: + """Recompute the E20400 proxy after aggregate-row amounts are replaced.""" + + output = "unreimbursed_business_employee_expenses" + if output not in puf.columns or "E20400" not in puf.columns: + return puf + return derive_us_misc_itemized_from_puf(puf) + + +def _reconcile_puf_form_4952_election_from_source( + puf: pd.DataFrame, +) -> pd.DataFrame: + """Recompute the E58990 carry after aggregate-row amounts are replaced.""" + + output = "investment_income_elected_form_4952" + if output not in puf.columns or "E58990" not in puf.columns: + return puf + return derive_us_form_4952_election_from_puf(puf) + + +def _reconcile_puf_salt_refund_income_from_source( + puf: pd.DataFrame, +) -> pd.DataFrame: + """Recompute the E00700 carry after aggregate-row amounts are replaced.""" + + output = "salt_refund_income" + if output not in puf.columns or "E00700" not in puf.columns: + return puf + return derive_us_salt_refund_income_from_puf(puf) + + +def _reconcile_puf_capital_gain_details_from_sources( + puf: pd.DataFrame, +) -> pd.DataFrame: + """Recompute E24518/E24515 carries after source amounts are replaced.""" + + required = { + "long_term_capital_gains_on_collectibles", + "unrecaptured_section_1250_gain", + "E24518", + "E24515", + } + if not required.issubset(puf.columns): + return puf + return derive_us_capital_gain_details_from_puf(puf) + + +def _reconcile_puf_farm_business_income_from_sources( + puf: pd.DataFrame, +) -> pd.DataFrame: + """Recompute signed E02100/E27200 carries after disaggregation.""" + + required = { + "farm_operations_income", + "farm_rent_income", + "E02100", + "E27200", + } + if not required.issubset(puf.columns): + return puf + return derive_us_farm_business_income_from_puf(puf) + + def _assert_nonnegative_dividend_component( values: pd.Series, *, diff --git a/packages/populace-build/src/populace/build/us_runtime/puf_support.py b/packages/populace-build/src/populace/build/us_runtime/puf_support.py index 9597a31c..f2542a47 100644 --- a/packages/populace-build/src/populace/build/us_runtime/puf_support.py +++ b/packages/populace-build/src/populace/build/us_runtime/puf_support.py @@ -16,6 +16,11 @@ import pandas as pd from populace.build.gates import FitWeightRecord +from populace.build.us_runtime.qbi_inputs import ( + US_QBI_BOOLEAN_OUTPUT_COLUMNS, + US_QBI_NONNEGATIVE_OUTPUT_COLUMNS, + US_QBI_OUTPUT_COLUMNS, +) from populace.frame import US_SCHEMA, Frame, WeightKind, Weights from populace.frame.schema import EntitySchema @@ -71,6 +76,7 @@ "tax_exempt_interest_income", "short_term_capital_gains", "long_term_capital_gains_before_response", + "long_term_capital_gains_on_collectibles", "non_sch_d_capital_gains", "taxable_private_pension_income", "taxable_ira_distributions", @@ -78,11 +84,19 @@ "social_security_disability", "social_security_dependents", "social_security_survivors", + "alimony_income", + "alimony_expense", + "salt_refund_income", "charitable_cash_donations", "charitable_non_cash_donations", "real_estate_taxes", "home_mortgage_interest", + "investment_income_elected_form_4952", "student_loan_interest", + "educator_expense", + "qualified_tuition_expenses", + "casualty_loss", + "unreimbursed_business_employee_expenses", # The engine owns the realized contribution amounts through the # IRA-limit scale and self-employment caps; the persistable leaves are # the desired contributions, equal to the PUF's observed deductions at @@ -92,10 +106,13 @@ "rental_income", "estate_income", "farm_income", + "farm_operations_income", + "farm_rent_income", "miscellaneous_income", "partnership_income", "s_corp_income", "partnership_self_employment_net_earnings", + *US_QBI_OUTPUT_COLUMNS, ) PUF_TAX_DETAIL_SOCIAL_SECURITY_COMPONENT_OUTPUTS = ( @@ -106,6 +123,8 @@ ) PUF_TAX_DETAIL_DEFAULT_TAX_UNIT_OUTPUTS: tuple[str, ...] = ( + "domestic_production_ald", + "unrecaptured_section_1250_gain", "first_home_mortgage_balance", "second_home_mortgage_balance", "first_home_mortgage_interest", @@ -121,12 +140,33 @@ "second_home_mortgage_origination_year", } ) +_PUF_TAX_DETAIL_BOOLEAN_PERSON_OUTPUTS = frozenset(US_QBI_BOOLEAN_OUTPUT_COLUMNS) +_PUF_TAX_DETAIL_SPARSE_TAX_UNIT_OUTPUTS = frozenset( + { + "domestic_production_ald", + "unrecaptured_section_1250_gain", + } +) _PUF_TAX_DETAIL_SPARSE_PERSON_OUTPUTS = frozenset( { "taxable_interest_income", + "qualified_tuition_expenses", + "educator_expense", + "casualty_loss", + "investment_income_elected_form_4952", + "long_term_capital_gains_on_collectibles", + "alimony_income", + "alimony_expense", + "salt_refund_income", } ) +# ASEC directly measures recipient alimony. The PUF QRF therefore sparsifies +# only the cloned PUF half for this leaf; pruning the ASEC half would discard +# reported source observations. Expense has no ASEC analogue, so its zero ASEC +# half is unaffected by the ordinary all-channel sparsification loop. +_PUF_TAX_DETAIL_PRESERVE_BASE_ASEC_OUTPUTS = frozenset({"alimony_income"}) + # Known formula-owned outputs the PUF tax-detail donor must never carry as # persistable leaves. This is a documented *seed* set, not the whole story: # :func:`resolve_formula_owned_outputs` unions it with the set derived live @@ -159,6 +199,7 @@ "non_qualified_dividend_income", "tax_exempt_interest_income", "long_term_capital_gains_before_response", + "long_term_capital_gains_on_collectibles", "taxable_pension_income", "taxable_private_pension_income", "taxable_ira_distributions", @@ -166,11 +207,21 @@ "social_security_disability", "social_security_dependents", "social_security_survivors", + "alimony_income", + "alimony_expense", + "salt_refund_income", "charitable_cash_donations", "charitable_non_cash_donations", "real_estate_taxes", "home_mortgage_interest", + "investment_income_elected_form_4952", "student_loan_interest", + "educator_expense", + "qualified_tuition_expenses", + "casualty_loss", + "unreimbursed_business_employee_expenses", + "domestic_production_ald", + "unrecaptured_section_1250_gain", "traditional_ira_contributions_desired", "self_employed_pension_contributions_desired", "health_savings_account_ald", @@ -180,6 +231,7 @@ "second_home_mortgage_interest", "first_home_mortgage_origination_year", "second_home_mortgage_origination_year", + *US_QBI_NONNEGATIVE_OUTPUT_COLUMNS, } ) # Person-grain PE input leaves the processed PUF artifact only observes as @@ -197,6 +249,11 @@ # Contributions require compensation, so an earnings basis keeps per-person # contribution limits binding the way they do on the donor records. _PERSON_OUTPUT_DISTRIBUTION_BASIS: Mapping[str, tuple[str, ...]] = { + "casualty_loss": ( + "employment_income_before_lsr", + "self_employment_income_before_lsr", + ), + "educator_expense": ("employment_income_before_lsr",), "traditional_ira_contributions_desired": ( "employment_income_before_lsr", "self_employment_income_before_lsr", @@ -204,6 +261,56 @@ "self_employed_pension_contributions_desired": ( "self_employment_income_before_lsr", ), + "qualified_tuition_expenses": ("is_full_time_college_student",), + "estate_income_would_be_qualified": ("estate_income",), + "farm_operations_income_would_be_qualified": ("farm_operations_income",), + "farm_rent_income_would_be_qualified": ("farm_rent_income",), + "partnership_s_corp_income_would_be_qualified": ( + "partnership_income", + "s_corp_income", + ), + "rental_income_would_be_qualified": ("rental_income",), + "self_employment_income_would_be_qualified": ( + "self_employment_income_before_lsr", + "sstb_self_employment_income_before_lsr", + ), + "sstb_self_employment_income_would_be_qualified": ( + "sstb_self_employment_income_before_lsr", + "self_employment_income_before_lsr", + ), + "business_is_sstb": ( + "sstb_self_employment_income_before_lsr", + "self_employment_income_before_lsr", + "partnership_income", + "s_corp_income", + "estate_income", + ), + "qualified_bdc_income": ("non_qualified_dividend_income",), + "qualified_reit_and_ptp_income": ( + "non_qualified_dividend_income", + "partnership_income", + "s_corp_income", + ), + "sstb_self_employment_income_before_lsr": ( + "business_is_sstb", + "self_employment_income_before_lsr", + ), + "sstb_unadjusted_basis_qualified_property": ("business_is_sstb",), + "sstb_w2_wages_from_qualified_business": ("business_is_sstb",), + "unadjusted_basis_qualified_property": ( + "self_employment_income_before_lsr", + "partnership_income", + "s_corp_income", + "rental_income", + "estate_income", + ), + "w2_wages_from_qualified_business": ( + "self_employment_income_before_lsr", + "partnership_income", + "s_corp_income", + "rental_income", + "estate_income", + ), } _PREDICTOR_LEAF_ALIASES: Mapping[str, tuple[str, ...]] = { "employment_income": ("employment_income_before_lsr",), @@ -410,7 +517,10 @@ def impute_us_puf_tax_detail_support( tax-unit predictions from the PUF donor; person-grain predicted tax-unit totals are distributed over the cloned people using their copied ASEC within-tax-unit shares, falling back to the first person in the unit when - the copied support has no mass for a variable. + the copied support has no mass for a variable. Boolean QBI targets are + modeled as tax-unit person counts, snapped to observed integer counts, and + placed on source-aligned people instead of collapsing every positive count + onto one person. Args: fit_records: An optional sink for the build-level weights audit (populace @@ -472,9 +582,7 @@ def impute_us_puf_tax_detail_support( # Record the kind the fit *resolved* to (not the "design" spec above): # the build-level weights audit reads this back to prove the production # fit did not silently resolve unweighted (populace #300). - fit_records.append( - FitWeightRecord(US_PUF_SUPPORT_FIT_NAME, fitted.weight_kind) - ) + fit_records.append(FitWeightRecord(US_PUF_SUPPORT_FIT_NAME, fitted.weight_kind)) features = _tax_unit_feature_frame(frame, predictors) puf_mask = ( @@ -492,6 +600,16 @@ def impute_us_puf_tax_detail_support( predictions[column], donor[column], ) + if column in _PUF_TAX_DETAIL_SPARSE_TAX_UNIT_OUTPUTS: + predictions[column] = _snap_to_observed_values( + predictions[column], + donor[column], + ) + if column in _PUF_TAX_DETAIL_BOOLEAN_PERSON_OUTPUTS: + predictions[column] = _snap_to_observed_values( + predictions[column], + donor[column], + ) if column in _PUF_TAX_DETAIL_DISCRETE_TAX_UNIT_OUTPUTS: predictions[column] = _snap_to_observed_values( predictions[column], @@ -510,18 +628,44 @@ def impute_us_puf_tax_detail_support( for column in tax_unit_outputs: _ensure_float_output_column(tables["tax_unit"], column) tables["tax_unit"].loc[puf_mask, column] = predictions[column].to_numpy() + for column in tax_unit_outputs: + if column in _PUF_TAX_DETAIL_SPARSE_TAX_UNIT_OUTPUTS: + _sparsify_tax_unit_output_to_donor_positive_rate( + tables, + column=column, + donor_positive_rate=_weighted_positive_rate( + donor[column], + donor["weight"], + ), + household_weights=frame.weights_for("household").values, + tax_unit_channel=tax_unit_channel, + ) person_puf_mask = tables["person"][person_channel] == PUF_TAX_DETAIL_SUPPORT_CHANNEL for column in person_outputs: _ensure_float_output_column(tables["person"], column) - _write_person_tax_unit_totals( - tables["person"], - mask=person_puf_mask, - column=column, - totals=pd.Series(predictions[column].to_numpy(), index=tax_unit_ids), - nonnegative=column in _PUF_TAX_DETAIL_NONNEGATIVE_OUTPUTS, - fallback_basis_columns=_PERSON_OUTPUT_DISTRIBUTION_BASIS.get(column, ()), - ) + totals = pd.Series(predictions[column].to_numpy(), index=tax_unit_ids) + if column in _PUF_TAX_DETAIL_BOOLEAN_PERSON_OUTPUTS: + _write_person_tax_unit_boolean_counts( + tables["person"], + mask=person_puf_mask, + column=column, + totals=totals, + fallback_basis_columns=_PERSON_OUTPUT_DISTRIBUTION_BASIS.get( + column, () + ), + ) + else: + _write_person_tax_unit_totals( + tables["person"], + mask=person_puf_mask, + column=column, + totals=totals, + nonnegative=column in _PUF_TAX_DETAIL_NONNEGATIVE_OUTPUTS, + fallback_basis_columns=_PERSON_OUTPUT_DISTRIBUTION_BASIS.get( + column, () + ), + ) for column in person_outputs: if column in _PUF_TAX_DETAIL_SPARSE_PERSON_OUTPUTS: _sparsify_tax_unit_person_output_to_donor_positive_rate( @@ -971,6 +1115,71 @@ def _snap_to_observed_values( return observed_array[positions] +def _write_person_tax_unit_boolean_counts( + person: pd.DataFrame, + *, + mask: pd.Series, + column: str, + totals: pd.Series, + fallback_basis_columns: tuple[str, ...] = (), +) -> None: + """Place a predicted number of true people within each tax unit. + + PUF QBI flags are person inputs, but the shared support model trains at + tax-unit grain. Their donor targets are therefore integer counts. Treating + a count as an amount would put (say) ``2`` on one person and ``0`` on the + spouse; a later boolean cast would silently turn two true people into one. + This placement preserves the snapped count and ranks people by the source + columns the flag qualifies, with stable first-person fallback for ties. + """ + + row_ids = person.loc[mask, "person_tax_unit_id"] + if row_ids.empty: + return + score = np.zeros(len(row_ids), dtype=np.float64) + for basis_column in fallback_basis_columns: + if basis_column not in person.columns: + continue + score += ( + pd.to_numeric(person.loc[mask, basis_column], errors="coerce") + .fillna(0.0) + .clip(lower=0.0) + .to_numpy(dtype=np.float64) + ) + + placement = pd.DataFrame( + { + "tax_unit_id": row_ids.to_numpy(), + "score": score, + "source_position": np.arange(len(row_ids), dtype=np.int64), + }, + index=row_ids.index, + ).sort_values( + ["tax_unit_id", "score", "source_position"], + ascending=[True, False, True], + kind="mergesort", + ) + placement["rank"] = placement.groupby("tax_unit_id", sort=False).cumcount() + placement["unit_size"] = placement.groupby("tax_unit_id", sort=False)[ + "tax_unit_id" + ].transform("size") + desired = ( + placement["tax_unit_id"] + .map(totals) + .fillna(0.0) + .round() + .clip(lower=0.0) + .to_numpy(dtype=np.float64) + ) + desired = np.minimum( + desired, + placement["unit_size"].to_numpy(dtype=np.float64), + ) + placement["selected"] = placement["rank"].to_numpy() < desired + selected = placement["selected"].reindex(row_ids.index).fillna(False) + person.loc[mask, column] = selected.to_numpy(dtype=np.float64) + + def _write_person_tax_unit_totals( person: pd.DataFrame, *, @@ -1032,6 +1241,72 @@ def _weighted_positive_rate(values: pd.Series, weights: pd.Series) -> float: return float((numeric_weights * (numeric_values > 0.0)).sum() / total_weight) +def _sparsify_tax_unit_output_to_donor_positive_rate( + tables: Mapping[str, pd.DataFrame], + *, + column: str, + donor_positive_rate: float, + household_weights: np.ndarray, + tax_unit_channel: str, +) -> None: + """Prune a sparse tax-unit amount to the donor's weighted positive rate.""" + + positive_rate = float(np.clip(donor_positive_rate, 0.0, 1.0)) + person = tables["person"] + household = tables["household"] + tax_unit = tables["tax_unit"] + if len(household_weights) != len(household): + raise ValueError( + "household_weights must align with household rows, got " + f"{len(household_weights)} weights for {len(household)} households." + ) + household_weight = pd.Series( + np.asarray(household_weights, dtype=np.float64), + index=household["household_id"], + ) + tax_unit_household_id = ( + person.groupby("person_tax_unit_id", sort=False)["person_household_id"] + .first() + .astype("int64") + ) + tax_unit_weight = tax_unit_household_id.map(household_weight).fillna(0.0) + + for channel in tax_unit[tax_unit_channel].dropna().unique(): + channel_mask = tax_unit[tax_unit_channel] == channel + channel_rows = tax_unit.loc[channel_mask] + amounts = pd.Series( + pd.to_numeric(channel_rows[column], errors="coerce") + .fillna(0.0) + .to_numpy(dtype=np.float64), + index=channel_rows["tax_unit_id"].to_numpy(), + ) + weights = tax_unit_weight.reindex(amounts.index).fillna(0.0) + positive = amounts > 0.0 + positive_weight = float(weights[positive].sum()) + desired_positive_weight = positive_rate * float(weights.sum()) + if positive_weight <= desired_positive_weight or positive_weight <= 0.0: + continue + + ranked = ( + pd.DataFrame({"amount": amounts[positive], "weight": weights[positive]}) + .sort_values("amount", ascending=False) + .copy() + ) + cumulative = ranked["weight"].cumsum() + keep = cumulative <= desired_positive_weight + if not keep.any() and len(keep) > 0: + keep.iloc[0] = True + kept_ids = set(ranked.index[keep]) + + sparse_amounts = amounts.copy() + sparse_amounts.loc[positive & ~amounts.index.isin(kept_ids)] = 0.0 + original_total = float((amounts * weights).sum()) + sparse_total = float((sparse_amounts * weights).sum()) + if original_total != 0.0 and sparse_total != 0.0: + sparse_amounts *= original_total / sparse_total + tax_unit.loc[channel_mask, column] = sparse_amounts.to_numpy() + + def _sparsify_tax_unit_person_output_to_donor_positive_rate( tables: Mapping[str, pd.DataFrame], *, @@ -1067,7 +1342,16 @@ def _sparsify_tax_unit_person_output_to_donor_positive_rate( .sum() ) - for channel in tax_unit[tax_unit_channel].dropna().unique(): + channels = tax_unit[tax_unit_channel].dropna().unique() + if column in _PUF_TAX_DETAIL_PRESERVE_BASE_ASEC_OUTPUTS: + channels = np.asarray( + [ + channel + for channel in channels + if channel == PUF_TAX_DETAIL_SUPPORT_CHANNEL + ] + ) + for channel in channels: channel_tax_unit_ids = tax_unit.loc[ tax_unit[tax_unit_channel] == channel, "tax_unit_id", @@ -1192,6 +1476,7 @@ def _person_source_values( source_aliases = { "employment_income_before_lsr": ("employment_income",), "self_employment_income_before_lsr": ("self_employment_income",), + "sstb_self_employment_income_before_lsr": ("sstb_self_employment_income",), "long_term_capital_gains_before_response": ("long_term_capital_gains",), "taxable_private_pension_income": ("taxable_pension_income",), # The PUF observes realized IRA deductions; at baseline the engine's @@ -1227,6 +1512,11 @@ def _person_source_values( "taxable_unemployment_compensation" in arrays ): return _numeric_array(arrays["taxable_unemployment_compensation"]) + if output == "qualified_tuition_expenses" and "E03230" in arrays: + tuition = _numeric_array(arrays["E03230"]) + if "E87530" in arrays: + tuition = np.maximum(tuition, _numeric_array(arrays["E87530"])) + return np.maximum(tuition, 0.0) if output in PUF_TAX_DETAIL_SOCIAL_SECURITY_COMPONENT_OUTPUTS: if output != "social_security_retirement": for source in ("social_security", "total_social_security", "E02400"): diff --git a/packages/populace-build/src/populace/build/us_runtime/qbi_inputs.py b/packages/populace-build/src/populace/build/us_runtime/qbi_inputs.py new file mode 100644 index 00000000..fdfc93cd --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/qbi_inputs.py @@ -0,0 +1,415 @@ +"""Archived PUF Section 199A input family for the US build. + +The retired eCPS PUF pipeline used a pinned, seeded QBI simulation to create +source-level qualification flags, SSTB classification and self-employment +splits, allocable W-2 wages / UBIA, and qualified REIT/PTP and BDC income. The +frozen processed-PUF artifact consumed by the hermetic build carries those +materialized simulated leaves; Populace does not redraw them. The shared +weighted PUF QRF places them on the PUF support channel. + +This module restores the cross-column identities after imputation. In +particular, the archived PUF model is all-or-nothing at tax-record grain: the +SSTB flag routes the entire predicted Schedule C amount to the SSTB leaf and +duplicates the total W-2 wage and UBIA pools into their SSTB-allocable leaves. +The base W-2/UBIA leaves remain total pools, not non-SSTB complements. +PolicyEngine-US owns the QBI deduction formulas and statutory limits. +""" + +from __future__ import annotations + +from importlib.resources import files + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import SourceStageSpec, load_source_manifest +from populace.frame import US_SCHEMA, Frame + +__all__ = [ + "QBI_ARCHIVED_ASSUMPTIONS_URL", + "QBI_ARCHIVED_CLONE_URL", + "QBI_ARCHIVED_DERIVATION_URL", + "QBI_ARCHIVED_EXPORT_URL", + "QBI_ARCHIVED_IMPUTATION_URL", + "QBI_ARCHIVED_PUF_ARTIFACT_URL", + "QBI_ARCHIVED_SIMULATION_URL", + "US_QBI_BOOLEAN_OUTPUT_COLUMNS", + "US_QBI_NONCONSTANT_PERSON_COLUMNS", + "US_QBI_NONNEGATIVE_OUTPUT_COLUMNS", + "US_QBI_OUTPUT_COLUMNS", + "US_QBI_STAGE_NAME", + "us_qbi_inputs_signal_gate", + "us_qbi_inputs_stage_spec", + "us_qbi_inputs_summary", + "with_us_qbi_input_reconciliation", +] + +_ARCHIVED_DATA_REPOSITORY = "policyengine-" + "us-data" +_ARCHIVED_COMMIT = "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe" +_ARCHIVED_ROOT = ( + "https://github.com/PolicyEngine/" + f"{_ARCHIVED_DATA_REPOSITORY}/blob/{_ARCHIVED_COMMIT}/" + "policyengine_" + "us_data/" +) +QBI_ARCHIVED_DERIVATION_URL = _ARCHIVED_ROOT + "datasets/puf/puf.py#L748-L787" +QBI_ARCHIVED_SIMULATION_URL = _ARCHIVED_ROOT + "datasets/puf/puf.py#L105-L405" +QBI_ARCHIVED_EXPORT_URL = _ARCHIVED_ROOT + "datasets/puf/puf.py#L860-L879" +QBI_ARCHIVED_ASSUMPTIONS_URL = ( + _ARCHIVED_ROOT + "datasets/puf/qbi_assumptions.yaml#L1-L118" +) +QBI_ARCHIVED_IMPUTATION_URL = _ARCHIVED_ROOT + "calibration/puf_impute.py#L99-L198" +QBI_ARCHIVED_CLONE_URL = _ARCHIVED_ROOT + "calibration/puf_impute.py#L513-L685" +QBI_ARCHIVED_PUF_ARTIFACT_URL = _ARCHIVED_ROOT + "datasets/puf/puf.py#L1655-L1660" + +US_QBI_STAGE_NAME = "puf_tax_detail" + +_GENERAL_QUALIFICATION_FLAGS: tuple[str, ...] = ( + "estate_income_would_be_qualified", + "farm_operations_income_would_be_qualified", + "farm_rent_income_would_be_qualified", + "partnership_s_corp_income_would_be_qualified", + "rental_income_would_be_qualified", + "self_employment_income_would_be_qualified", +) +_SSTB_QUALIFICATION_FLAG = "sstb_self_employment_income_would_be_qualified" +US_QBI_BOOLEAN_OUTPUT_COLUMNS: tuple[str, ...] = ( + *_GENERAL_QUALIFICATION_FLAGS, + _SSTB_QUALIFICATION_FLAG, + # Keep the classifier last so the chained QRF can condition the SSTB draw + # on the qualification flags it must agree with. + "business_is_sstb", +) +US_QBI_NONNEGATIVE_OUTPUT_COLUMNS: tuple[str, ...] = ( + "qualified_bdc_income", + "qualified_reit_and_ptp_income", + "sstb_unadjusted_basis_qualified_property", + "sstb_w2_wages_from_qualified_business", + "unadjusted_basis_qualified_property", + "w2_wages_from_qualified_business", +) +US_QBI_OUTPUT_COLUMNS: tuple[str, ...] = ( + *US_QBI_BOOLEAN_OUTPUT_COLUMNS, + "qualified_bdc_income", + "qualified_reit_and_ptp_income", + "sstb_self_employment_income_before_lsr", + "sstb_unadjusted_basis_qualified_property", + "sstb_w2_wages_from_qualified_business", + "unadjusted_basis_qualified_property", + "w2_wages_from_qualified_business", +) +US_QBI_NONCONSTANT_PERSON_COLUMNS = US_QBI_OUTPUT_COLUMNS + +_PERSON_SUPPORT_CHANNEL_COLUMN = "person_support_channel" +_BASE_ASEC_SUPPORT_CHANNEL = "asec" +_SELF_EMPLOYMENT_COLUMN = "self_employment_income_before_lsr" +_SSTB_SELF_EMPLOYMENT_COLUMN = "sstb_self_employment_income_before_lsr" +_BOOLEAN_SHARE_BANDS: dict[str, tuple[float, float]] = { + "business_is_sstb": (0.001, 0.25), + **{column: (0.01, 0.9999) for column in _GENERAL_QUALIFICATION_FLAGS}, + _SSTB_QUALIFICATION_FLAG: (0.0001, 0.25), +} +_NUMERIC_NONZERO_SHARE_BANDS: dict[str, tuple[float, float]] = { + "qualified_bdc_income": (0.0001, 0.15), + "qualified_reit_and_ptp_income": (0.001, 0.35), + _SSTB_SELF_EMPLOYMENT_COLUMN: (0.0001, 0.25), + "sstb_unadjusted_basis_qualified_property": (0.0001, 0.25), + "sstb_w2_wages_from_qualified_business": (0.0001, 0.20), + "unadjusted_basis_qualified_property": (0.001, 0.45), + "w2_wages_from_qualified_business": (0.001, 0.35), +} +_INVARIANT_ATOL = 1e-8 + + +def us_qbi_inputs_stage_spec() -> SourceStageSpec: + """Load and validate the shared PUF stage's QBI output contract.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + spec = manifest.stage_map()[US_QBI_STAGE_NAME] + missing = sorted(set(US_QBI_OUTPUT_COLUMNS) - set(spec.outputs)) + if missing: + raise ValueError( + f"{US_QBI_STAGE_NAME!r} manifest stage does not declare QBI " + f"output(s) {missing}." + ) + return spec + + +def _numeric(person: pd.DataFrame, column: str) -> np.ndarray: + values = pd.to_numeric(person[column], errors="coerce").to_numpy(dtype=np.float64) + nonfinite = int(np.count_nonzero(~np.isfinite(values))) + if nonfinite: + raise ValueError( + f"US QBI input {column!r} contains {nonfinite} nonfinite value(s)." + ) + return values + + +def _optional_numeric(person: pd.DataFrame, column: str) -> np.ndarray: + if column not in person: + return np.zeros(len(person), dtype=np.float64) + return _numeric(person, column) + + +def with_us_qbi_input_reconciliation(frame: Frame) -> Frame: + """Restore archived all-or-nothing SSTB split identities after PUF QRF.""" + + if frame.schema != US_SCHEMA: + raise ValueError("US QBI input reconciliation requires the US schema.") + person = frame.table("person") + required = {*US_QBI_OUTPUT_COLUMNS, _SELF_EMPLOYMENT_COLUMN} + missing = sorted(required - set(person.columns)) + if missing: + raise ValueError( + "US QBI input reconciliation requires PUF-imputed source column(s): " + f"{missing}." + ) + + result = person.copy(deep=True) + asec_mask = np.zeros(len(result), dtype=bool) + if _PERSON_SUPPORT_CHANNEL_COLUMN in result: + asec_mask = ( + result[_PERSON_SUPPORT_CHANNEL_COLUMN].to_numpy(dtype=object) + == _BASE_ASEC_SUPPORT_CHANNEL + ) + + flags: dict[str, np.ndarray] = {} + for column in US_QBI_BOOLEAN_OUTPUT_COLUMNS: + values = _numeric(result, column) + flags[column] = values > 0.0 + # Populace's base ASEC support has no observations from the frozen PUF QBI + # simulation. Deliberately preserve PolicyEngine's ordinary qualification + # defaults on that channel and do not invent an SSTB classification there. + # This is a hermetic two-channel choice, not a claim about the retired + # clone implementation (which imputed absent PUF variables onto both halves). + for column in _GENERAL_QUALIFICATION_FLAGS: + flags[column][asec_mask] = True + flags["business_is_sstb"][asec_mask] = False + flags[_SSTB_QUALIFICATION_FLAG][asec_mask] = False + + non_sstb_self_employment = _numeric(result, _SELF_EMPLOYMENT_COLUMN) + sstb_self_employment = _numeric(result, _SSTB_SELF_EMPLOYMENT_COLUMN) + total_self_employment = non_sstb_self_employment + sstb_self_employment + + base_self_employment_qualified = ( + flags["self_employment_income_would_be_qualified"] + | flags[_SSTB_QUALIFICATION_FLAG] + ) + partnership_s_corp_income = _optional_numeric( + result, "partnership_income" + ) + _optional_numeric(result, "s_corp_income") + estate_income = _optional_numeric(result, "estate_income") + has_positive_qualified_mapped_source = ( + ((total_self_employment > 0.0) & base_self_employment_qualified) + | ( + (partnership_s_corp_income > 0.0) + & flags["partnership_s_corp_income_would_be_qualified"] + ) + | ((estate_income > 0.0) & flags["estate_income_would_be_qualified"]) + ) + business_is_sstb = flags["business_is_sstb"] & has_positive_qualified_mapped_source + flags["business_is_sstb"] = business_is_sstb + + flags["self_employment_income_would_be_qualified"] = ( + ~business_is_sstb & base_self_employment_qualified + ) + flags[_SSTB_QUALIFICATION_FLAG] = business_is_sstb & base_self_employment_qualified + + result[_SELF_EMPLOYMENT_COLUMN] = np.where( + business_is_sstb, + 0.0, + total_self_employment, + ) + result[_SSTB_SELF_EMPLOYMENT_COLUMN] = np.where( + business_is_sstb, + total_self_employment, + 0.0, + ) + + w2_wages = _numeric(result, "w2_wages_from_qualified_business") + ubia = _numeric(result, "unadjusted_basis_qualified_property") + result["sstb_w2_wages_from_qualified_business"] = np.where( + business_is_sstb, + w2_wages, + 0.0, + ) + result["sstb_unadjusted_basis_qualified_property"] = np.where( + business_is_sstb, + ubia, + 0.0, + ) + non_qualified_dividends = np.maximum( + _optional_numeric(result, "non_qualified_dividend_income"), + 0.0, + ) + result["qualified_bdc_income"] = np.minimum( + _numeric(result, "qualified_bdc_income"), + non_qualified_dividends, + ) + result["qualified_reit_and_ptp_income"] = np.minimum( + _numeric(result, "qualified_reit_and_ptp_income"), + non_qualified_dividends + np.maximum(partnership_s_corp_income, 0.0), + ) + for column, values in flags.items(): + result[column] = values.astype(bool) + + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + tables["person"] = result + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + ) + + +def us_qbi_inputs_summary(frame: Frame) -> dict[str, object]: + """Return weighted signal, validity, and split-identity diagnostics.""" + + person = frame.table("person") + weights = np.asarray(frame.resolve_weights("person").values, dtype=np.float64) + total_weight = float(weights.sum()) + columns: dict[str, dict[str, object]] = {} + for column in US_QBI_OUTPUT_COLUMNS: + values = pd.to_numeric(person[column], errors="coerce").to_numpy( + dtype=np.float64 + ) + finite = np.isfinite(values) + nonzero = finite & (values != 0.0) + nonzero_share = ( + float(weights[nonzero].sum()) / total_weight if total_weight > 0.0 else 0.0 + ) + band = ( + _BOOLEAN_SHARE_BANDS[column] + if column in _BOOLEAN_SHARE_BANDS + else _NUMERIC_NONZERO_SHARE_BANDS[column] + ) + columns[column] = { + "nonzero_share": nonzero_share, + "nonzero_share_band": list(band), + "nonfinite": int(np.count_nonzero(~finite)), + "negative": int(np.count_nonzero(finite & (values < 0.0))), + } + + business = person["business_is_sstb"].astype(bool).to_numpy() + self_employment = pd.to_numeric( + person[_SELF_EMPLOYMENT_COLUMN], errors="coerce" + ).to_numpy(dtype=np.float64) + sstb_self_employment = pd.to_numeric( + person[_SSTB_SELF_EMPLOYMENT_COLUMN], errors="coerce" + ).to_numpy(dtype=np.float64) + w2 = pd.to_numeric( + person["w2_wages_from_qualified_business"], errors="coerce" + ).to_numpy(dtype=np.float64) + sstb_w2 = pd.to_numeric( + person["sstb_w2_wages_from_qualified_business"], errors="coerce" + ).to_numpy(dtype=np.float64) + ubia = pd.to_numeric( + person["unadjusted_basis_qualified_property"], errors="coerce" + ).to_numpy(dtype=np.float64) + sstb_ubia = pd.to_numeric( + person["sstb_unadjusted_basis_qualified_property"], errors="coerce" + ).to_numpy(dtype=np.float64) + self_qualified = ( + person["self_employment_income_would_be_qualified"].astype(bool).to_numpy() + ) + sstb_qualified = person[_SSTB_QUALIFICATION_FLAG].astype(bool).to_numpy() + non_qualified_dividends = np.maximum( + _optional_numeric(person, "non_qualified_dividend_income"), + 0.0, + ) + partnership_s_corp_income = _optional_numeric( + person, "partnership_income" + ) + _optional_numeric(person, "s_corp_income") + qualified_bdc_income = _numeric(person, "qualified_bdc_income") + qualified_reit_and_ptp_income = _numeric(person, "qualified_reit_and_ptp_income") + invariants = { + "sstb_rows_with_non_sstb_income": int( + np.count_nonzero(business & ~np.isclose(self_employment, 0.0)) + ), + "non_sstb_rows_with_sstb_income": int( + np.count_nonzero(~business & ~np.isclose(sstb_self_employment, 0.0)) + ), + "sstb_w2_split_mismatches": int( + np.count_nonzero( + ~np.isclose(sstb_w2, np.where(business, w2, 0.0), atol=_INVARIANT_ATOL) + ) + ), + "sstb_ubia_split_mismatches": int( + np.count_nonzero( + ~np.isclose( + sstb_ubia, + np.where(business, ubia, 0.0), + atol=_INVARIANT_ATOL, + ) + ) + ), + "self_employment_qualification_overlap": int( + np.count_nonzero(self_qualified & sstb_qualified) + ), + "sstb_qualification_route_mismatches": int( + np.count_nonzero(sstb_qualified & ~business) + ), + "non_sstb_qualification_route_mismatches": int( + np.count_nonzero(self_qualified & business) + ), + "qualified_bdc_exposure_mismatches": int( + np.count_nonzero( + qualified_bdc_income > non_qualified_dividends + _INVARIANT_ATOL + ) + ), + "qualified_reit_ptp_exposure_mismatches": int( + np.count_nonzero( + qualified_reit_and_ptp_income + > non_qualified_dividends + + np.maximum(partnership_s_corp_income, 0.0) + + _INVARIANT_ATOL + ) + ), + } + return {"columns": columns, "invariants": invariants} + + +def us_qbi_inputs_signal_gate(frame: Frame) -> GateResult: + """Require nondefault QBI signal and archived SSTB split identities.""" + + person = frame.table("person") + missing = sorted( + {*US_QBI_OUTPUT_COLUMNS, _SELF_EMPLOYMENT_COLUMN} - set(person.columns) + ) + if missing: + return GateResult( + name="qbi_inputs_signal", + passed=False, + failures=(f"person columns missing: {missing}.",), + details={"missing": missing}, + ) + + summary = us_qbi_inputs_summary(frame) + failures: list[str] = [] + for column, details in summary["columns"].items(): + if details["nonfinite"]: + failures.append( + f"{column}: {int(details['nonfinite'])} nonfinite value(s)." + ) + if column in US_QBI_NONNEGATIVE_OUTPUT_COLUMNS and details["negative"]: + failures.append(f"{column}: {int(details['negative'])} negative value(s).") + share = float(details["nonzero_share"]) + low, high = details["nonzero_share_band"] + if not (low <= share <= high): + failures.append( + f"{column}: nonzero share {share:.6f} outside plausibility " + f"band [{low}, {high}]." + ) + for name, count in summary["invariants"].items(): + if count: + failures.append(f"{name}: {int(count)} violating row(s).") + return GateResult( + name="qbi_inputs_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) diff --git a/packages/populace-build/src/populace/build/us_runtime/reform_coverage_smoke.py b/packages/populace-build/src/populace/build/us_runtime/reform_coverage_smoke.py index 053013b1..ba304246 100644 --- a/packages/populace-build/src/populace/build/us_runtime/reform_coverage_smoke.py +++ b/packages/populace-build/src/populace/build/us_runtime/reform_coverage_smoke.py @@ -44,10 +44,22 @@ def _weighted_total(simulation: Any, measure: str, period: int) -> float: return float(simulation.calculate(measure, period).sum()) -def _build_reform(parameter_changes: dict[str, Any]) -> Any: +def _build_reform(probe: ReformCoverageProbe) -> Any: from policyengine_core.reforms import Reform - return Reform.from_dict(parameter_changes, country_id="us") + if probe.neutralized_variable: + variable = probe.neutralized_variable + + class _Neutralize(Reform): + def apply(self) -> None: + self.neutralize_variable(variable) + neutralized = self.variables[variable] + neutralized.default_value = ( + False if neutralized.value_type is bool else 0 + ) + + return _Neutralize + return Reform.from_dict(dict(probe.parameter_changes), country_id="us") def us_reform_coverage_smoke_gate( @@ -60,21 +72,21 @@ def us_reform_coverage_smoke_gate( """Score each pinned probe on the export; a ~$0 bound reform fails. For each probe, the budget-measure change (reform vs baseline, signed by - ``effect_direction``) must have magnitude at least ``min_abs_effect``. A - smaller magnitude means the reform did not bind — its ``binding_inputs`` are - absent/degenerate on the export — and fails the release. + ``effect_direction``) must have the declared ``expected_sign`` and magnitude + at least ``min_abs_effect``. A zero, undersized, or wrong-signed effect means + the reform did not bind as declared and fails the release. Args: simulate: ``simulate(None)`` builds the baseline; ``simulate(reform)`` builds the reformed simulation. Each result answers ``.calculate(measure, period).sum()`` as a weighted total. probes: Probes to run. Defaults to the shipped manifest's probe set. - period: The period to score at. + period: Default period for probes without a probe-specific period. name: Gate name for the manifest. Returns: The reform-coverage smoke gate result. Passes iff every probe scores at - least its ``min_abs_effect`` in magnitude. + its expected sign and at least its ``min_abs_effect`` in magnitude. """ probes = tuple( probes if probes is not None else us_release_reform_coverage_probes() @@ -89,30 +101,37 @@ def us_reform_coverage_smoke_gate( failures: list[str] = [] results: dict[str, Any] = {} for probe in probes: - reform = _build_reform(dict(probe.parameter_changes)) + reform = _build_reform(probe) reformed = simulate(reform) - baseline_total = _weighted_total(baseline, probe.budget_measure, period) - reform_total = _weighted_total(reformed, probe.budget_measure, period) + probe_period = int(probe.period if probe.period is not None else period) + baseline_total = _weighted_total(baseline, probe.budget_measure, probe_period) + reform_total = _weighted_total(reformed, probe.budget_measure, probe_period) if probe.effect_direction == "baseline_minus_reform": effect = baseline_total - reform_total else: effect = reform_total - baseline_total + signed_magnitude = effect if probe.expected_sign == "positive" else -effect + passed = signed_magnitude >= probe.min_abs_effect results[probe.id] = { "name": probe.name, + "period": probe_period, "budget_measure": probe.budget_measure, "baseline_total": baseline_total, "reform_total": reform_total, "effect": effect, + "expected_sign": probe.expected_sign, "min_abs_effect": probe.min_abs_effect, "binding_inputs": list(probe.binding_inputs), "issue": probe.issue, - "passed": abs(effect) >= probe.min_abs_effect, + "passed": passed, } - if abs(effect) < probe.min_abs_effect: + if not passed: failures.append( f"{probe.id}: '{probe.name}' scores {effect:+,.0f} on " - f"{probe.budget_measure} (|effect| < ${probe.min_abs_effect:,.0f}) " - "— the reform did not bind, so its input leaves " + f"{probe.budget_measure} for {probe_period}; expected a " + f"{probe.expected_sign} effect with magnitude at least " + f"${probe.min_abs_effect:,.0f}. The reform did not bind as " + "declared, so its input leaves " f"{list(probe.binding_inputs)} are absent or degenerate on the " f"export. {probe.reason} Restore them ({probe.issue})." ) @@ -122,7 +141,7 @@ def us_reform_coverage_smoke_gate( passed=not failures, failures=tuple(failures), details={ - "period": int(period), + "default_period": int(period), "probes": len(probes), "results": results, }, diff --git a/packages/populace-build/src/populace/build/us_runtime/relationship_inputs.py b/packages/populace-build/src/populace/build/us_runtime/relationship_inputs.py new file mode 100644 index 00000000..040b6732 --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/relationship_inputs.py @@ -0,0 +1,331 @@ +"""Measured CPS ASEC household-head and marital-status input leaves. + +The archived eCPS construction at commit +``42ed5d45c56df80d754fbe24cce21cfeb8d05cbe`` derives these inputs directly +in ``datasets/cps/cps.py``: + +- line 1074: ``is_household_head = P_SEQ == 1``; +- line 1212: ``is_surviving_spouse = A_MARITL == 4``; and +- line 1213: ``is_separated = A_MARITL == 6``. + +All three SHA-locked ASEC vintages retain the required raw columns. Nothing is +imputed, and the stage fails closed if a source column is missing, malformed, +or does not identify exactly one household head per source household. +""" + +from __future__ import annotations + +from importlib.resources import files + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import ( + SourceOperationSpec, + SourceStageSpec, + load_source_manifest, +) +from populace.build.source_runtime import ( + SourceRuntimeConfig, + SourceRuntimeContext, + SourceRuntimeError, + run_source_stage, +) +from populace.frame import Frame +from populace.frame.units import US_SCHEMA + +__all__ = [ + "US_RELATIONSHIP_INPUTS_NONCONSTANT_PERSON_COLUMNS", + "US_RELATIONSHIP_INPUTS_OUTPUT_COLUMNS", + "US_RELATIONSHIP_INPUTS_REQUIRED_SOURCE_COLUMNS", + "US_RELATIONSHIP_INPUTS_STAGE_NAME", + "derive_us_relationship_inputs_from_manifest", + "us_relationship_inputs_signal_gate", + "us_relationship_inputs_stage_spec", + "us_relationship_inputs_summary", + "with_us_relationship_inputs", +] + +US_RELATIONSHIP_INPUTS_STAGE_NAME = "relationship_inputs" + +US_RELATIONSHIP_INPUTS_OUTPUT_COLUMNS: tuple[str, ...] = ( + "is_household_head", + "is_separated", + "is_surviving_spouse", +) + +US_RELATIONSHIP_INPUTS_NONCONSTANT_PERSON_COLUMNS = ( + US_RELATIONSHIP_INPUTS_OUTPUT_COLUMNS +) + +US_RELATIONSHIP_INPUTS_REQUIRED_SOURCE_COLUMNS: tuple[str, ...] = ( + "PH_SEQ", + "P_SEQ", + "A_MARITL", +) + +_PERSON_WEIGHT_COLUMN = "person_weight" +_HOUSEHOLD_HEAD_SHARE_BAND = (0.30, 0.55) +_SEPARATED_SHARE_BAND = (0.003, 0.04) +_SURVIVING_SPOUSE_SHARE_BAND = (0.02, 0.08) +_VALID_A_MARITL_CODES = frozenset(range(1, 8)) +_DERIVE_RELATIONSHIP_INPUTS_PARAMETER_KEYS = frozenset() + + +def us_relationship_inputs_stage_spec() -> SourceStageSpec: + """Load the packaged ``relationship_inputs`` stage declaration.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + stage_map = manifest.stage_map() + if US_RELATIONSHIP_INPUTS_STAGE_NAME not in stage_map: + raise ValueError( + "US source manifest declares no " + f"{US_RELATIONSHIP_INPUTS_STAGE_NAME!r} stage." + ) + spec = stage_map[US_RELATIONSHIP_INPUTS_STAGE_NAME] + if tuple(spec.outputs) != US_RELATIONSHIP_INPUTS_OUTPUT_COLUMNS: + raise ValueError( + f"{US_RELATIONSHIP_INPUTS_STAGE_NAME!r} manifest outputs do not " + "match the runtime-owned relationship input family." + ) + return spec + + +def _strict_integer_source( + frame: pd.DataFrame, + column: str, + *, + minimum: int, + allowed: frozenset[int] | None = None, +) -> np.ndarray: + """Return a source column as integers or fail on missing/invalid values.""" + + numeric = pd.to_numeric(frame[column], errors="coerce").to_numpy(dtype=np.float64) + valid = np.isfinite(numeric) & (numeric == np.floor(numeric)) + valid &= numeric >= float(minimum) + if allowed is not None: + valid &= np.isin(numeric, np.fromiter(allowed, dtype=np.int64)) + if not valid.all(): + rows = np.flatnonzero(~valid)[:5].tolist() + raise SourceRuntimeError( + f"US relationship-input derivation requires valid integer {column}; " + f"invalid row(s): {rows}." + ) + return numeric.astype(np.int64) + + +def derive_us_relationship_inputs_from_manifest( + frame: pd.DataFrame | None, + operation: SourceOperationSpec, + _context: SourceRuntimeContext | None, +) -> pd.DataFrame: + """Map exact ASEC head and marital codes to PolicyEngine input leaves.""" + + if operation.kind != "derive_relationship_inputs": + raise SourceRuntimeError( + "US relationship-input derivation received unexpected operation " + f"{operation.kind!r}." + ) + if frame is None: + raise SourceRuntimeError( + "US relationship-input derivation requires the person table to be " + "read first." + ) + unexpected = sorted( + set(operation.parameters) - _DERIVE_RELATIONSHIP_INPUTS_PARAMETER_KEYS + ) + if unexpected: + raise SourceRuntimeError( + "US relationship-input derivation received unsupported " + f"parameter(s): {unexpected}." + ) + missing = [ + column + for column in US_RELATIONSHIP_INPUTS_REQUIRED_SOURCE_COLUMNS + if column not in frame.columns + ] + if missing: + raise SourceRuntimeError( + f"US relationship-input derivation requires raw ASEC column(s): {missing}." + ) + + household = _strict_integer_source(frame, "PH_SEQ", minimum=1) + person_sequence = _strict_integer_source(frame, "P_SEQ", minimum=1) + marital_status = _strict_integer_source( + frame, + "A_MARITL", + minimum=1, + allowed=_VALID_A_MARITL_CODES, + ) + is_head = person_sequence == 1 + grouping_household = ( + _strict_integer_source(frame, "person_household_id", minimum=1) + if "person_household_id" in frame + else household + ) + head_counts = pd.Series(is_head).groupby(grouping_household, sort=False).sum() + bad_households = head_counts.index[head_counts.to_numpy() != 1] + if len(bad_households): + examples = bad_households[:5].tolist() + raise SourceRuntimeError( + "US relationship-input derivation requires exactly one P_SEQ == 1 " + f"person per frame household; invalid household(s): {examples}." + ) + + result = frame.copy(deep=True) + result["is_household_head"] = is_head + result["is_separated"] = marital_status == 6 + result["is_surviving_spouse"] = marital_status == 4 + return result + + +def _relationship_surface_carries_signal(frame: Frame) -> bool: + person = frame.table("person") + if any(column not in person for column in US_RELATIONSHIP_INPUTS_OUTPUT_COLUMNS): + return False + return all( + person[column].dropna().nunique() > 1 + for column in US_RELATIONSHIP_INPUTS_OUTPUT_COLUMNS + ) + + +def with_us_relationship_inputs( + frame: Frame, + *, + seed: int, + time_period: int, +) -> Frame: + """Materialize measured ASEC relationship inputs on a US frame.""" + + if frame.schema != US_SCHEMA: + raise ValueError("US relationship inputs require the US schema.") + if _relationship_surface_carries_signal(frame): + return frame + + person = frame.table("person") + stage_person = person.copy(deep=True) + stage_person[_PERSON_WEIGHT_COLUMN] = frame.resolve_weights("person").values + output = run_source_stage( + us_relationship_inputs_stage_spec(), + tables={"person": stage_person}, + operation_handlers={ + "derive_relationship_inputs": (derive_us_relationship_inputs_from_manifest) + }, + config=SourceRuntimeConfig(seed=int(seed), target_year=int(time_period)), + ) + aligned = output.set_index("person_id").reindex(person["person_id"]) + for column in US_RELATIONSHIP_INPUTS_OUTPUT_COLUMNS: + if aligned[column].isna().any(): + raise ValueError( + "US relationship-input stage output does not cover every person " + f"for {column!r}." + ) + + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + for column in US_RELATIONSHIP_INPUTS_OUTPUT_COLUMNS: + tables["person"][column] = aligned[column].to_numpy(dtype=bool) + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + ) + + +def us_relationship_inputs_summary(frame: Frame) -> dict[str, object]: + """Return weighted relationship shares and one-head invariants.""" + + person = frame.table("person") + weights = np.asarray(frame.resolve_weights("person").values, dtype=np.float64) + total_weight = float(weights.sum()) + + def _share(column: str) -> float: + values = person[column].fillna(False).astype(bool).to_numpy() + return float(weights[values].sum()) / total_weight if total_weight > 0 else 0.0 + + household_column = ( + "person_household_id" if "person_household_id" in person else "PH_SEQ" + ) + head_counts = ( + person["is_household_head"] + .fillna(False) + .astype(bool) + .groupby(person[household_column], sort=False) + .sum() + ) + separated = person["is_separated"].fillna(False).astype(bool).to_numpy() + surviving = person["is_surviving_spouse"].fillna(False).astype(bool).to_numpy() + return { + "household_head_share": _share("is_household_head"), + "separated_share": _share("is_separated"), + "surviving_spouse_share": _share("is_surviving_spouse"), + "household_head_share_band": list(_HOUSEHOLD_HEAD_SHARE_BAND), + "separated_share_band": list(_SEPARATED_SHARE_BAND), + "surviving_spouse_share_band": list(_SURVIVING_SPOUSE_SHARE_BAND), + "households_without_exactly_one_head": int((head_counts != 1).sum()), + "separated_and_surviving": int(np.count_nonzero(separated & surviving)), + "unique_counts": { + column: int(person[column].dropna().nunique()) + for column in US_RELATIONSHIP_INPUTS_OUTPUT_COLUMNS + }, + } + + +def us_relationship_inputs_signal_gate(frame: Frame) -> GateResult: + """Require plausible signal and exactly one ASEC head per household.""" + + person = frame.table("person") + missing = [ + column + for column in US_RELATIONSHIP_INPUTS_OUTPUT_COLUMNS + if column not in person + ] + if missing: + return GateResult( + name="relationship_inputs_signal", + passed=False, + failures=(f"person columns missing: {missing}.",), + details={"missing": missing}, + ) + + summary = us_relationship_inputs_summary(frame) + failures: list[str] = [] + for share_key, band_key, label in ( + ( + "household_head_share", + "household_head_share_band", + "household-head weighted share", + ), + ("separated_share", "separated_share_band", "separated weighted share"), + ( + "surviving_spouse_share", + "surviving_spouse_share_band", + "surviving-spouse weighted share", + ), + ): + share = float(summary[share_key]) + lower, upper = summary[band_key] + if not lower <= share <= upper: + failures.append(f"{label} {share:.6f} outside [{lower:.6f}, {upper:.6f}].") + invalid_heads = int(summary["households_without_exactly_one_head"]) + if invalid_heads: + failures.append( + f"{invalid_heads} household(s) do not carry exactly one " + "is_household_head person." + ) + overlap = int(summary["separated_and_surviving"]) + if overlap: + failures.append(f"{overlap} person(s) are both separated and surviving spouse.") + for column, count in summary["unique_counts"].items(): + if int(count) < 2: + failures.append(f"{column} is degenerate with {count} distinct value(s).") + return GateResult( + name="relationship_inputs_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) diff --git a/packages/populace-build/src/populace/build/us_runtime/release_input_coverage.py b/packages/populace-build/src/populace/build/us_runtime/release_input_coverage.py index 9123ee99..04419268 100644 --- a/packages/populace-build/src/populace/build/us_runtime/release_input_coverage.py +++ b/packages/populace-build/src/populace/build/us_runtime/release_input_coverage.py @@ -47,9 +47,68 @@ from typing import Any from populace.build.gates import GateResult, input_column_coverage_gate +from populace.build.us_runtime.capital_gain_details import ( + US_CAPITAL_GAIN_DETAILS_OUTPUT_COLUMNS, +) +from populace.build.us_runtime.child_support import US_CHILD_SUPPORT_OUTPUT_COLUMNS +from populace.build.us_runtime.disability_benefits import ( + US_DISABILITY_BENEFITS_OUTPUT_COLUMNS, +) +from populace.build.us_runtime.domestic_production import ( + US_DOMESTIC_PRODUCTION_ALD_OUTPUT_COLUMNS, +) +from populace.build.us_runtime.educator_expenses import ( + US_EDUCATOR_EXPENSE_OUTPUT_COLUMNS, +) +from populace.build.us_runtime.energy_subsidy import US_ENERGY_SUBSIDY_OUTPUT_COLUMNS +from populace.build.us_runtime.farm_business_income import ( + US_FARM_BUSINESS_INCOME_OUTPUT_COLUMNS, +) +from populace.build.us_runtime.form_4952 import US_FORM_4952_OUTPUT_COLUMNS +from populace.build.us_runtime.housing_inputs import US_HOUSING_INPUTS_OUTPUT_COLUMNS +from populace.build.us_runtime.medicare_take_up import ( + US_MEDICARE_TAKE_UP_OUTPUT_COLUMNS, +) +from populace.build.us_runtime.other_health_insurance import ( + US_OTHER_HEALTH_INSURANCE_NONCONSTANT_PERSON_COLUMNS, +) +from populace.build.us_runtime.prior_year_income import ( + US_PRIOR_YEAR_INCOME_PERSISTED_OUTPUT_COLUMNS, +) +from populace.build.us_runtime.qbi_inputs import US_QBI_OUTPUT_COLUMNS +from populace.build.us_runtime.relationship_inputs import ( + US_RELATIONSHIP_INPUTS_OUTPUT_COLUMNS, +) +from populace.build.us_runtime.retirement_distributions import ( + US_RETIREMENT_DISTRIBUTION_OUTPUT_COLUMNS, +) +from populace.build.us_runtime.salt_refund_income import ( + US_SALT_REFUND_OUTPUT_COLUMNS, +) +from populace.build.us_runtime.scf_wealth import US_SCF_NET_WORTH_OUTPUT_COLUMNS +from populace.build.us_runtime.sipp_head_start import ( + US_SIPP_HEAD_START_OUTPUT_COLUMNS, +) +from populace.build.us_runtime.sipp_vehicles import US_SIPP_VEHICLE_OUTPUT_COLUMNS +from populace.build.us_runtime.ssi_disability_criteria import ( + US_SSI_DISABILITY_CRITERIA_OUTPUT_COLUMNS, +) +from populace.build.us_runtime.ssi_take_up import US_SSI_TAKE_UP_OUTPUT_COLUMNS +from populace.build.us_runtime.voluntary_filing import ( + US_VOLUNTARY_FILING_OUTPUT_COLUMNS, +) +from populace.build.us_runtime.weeks_unemployed import ( + US_WEEKS_UNEMPLOYED_OUTPUT_COLUMNS, +) +from populace.build.us_runtime.wic_claim import US_WIC_CLAIM_OUTPUT_COLUMNS +from populace.build.us_runtime.workers_compensation import ( + US_WORKERS_COMPENSATION_OUTPUT_COLUMNS, +) __all__ = [ "US_RELEASE_INPUT_COVERAGE_RESOURCE", + "POST_REFERENCE_ECPS_REQUIRED_INPUTS", + "RESTORED_REFERENCE_ECPS_REQUIRED_INPUTS", "ReformCoverageProbe", "ReleaseInputColumn", "ReleaseInputCoverageManifest", @@ -63,6 +122,64 @@ US_RELEASE_INPUT_COVERAGE_RESOURCE = "release_input_coverage_manifest.json" +# The frozen reference artifact predates the retired pipeline's export of the +# pure FLSA overtime-premium input, OBBBA's distinct qualifying passenger- +# vehicle interest leaf, and the final pipeline's five desired retirement- +# contribution inputs, and its final SIPP-imputed SSI disability criterion. +# These are hard requirements because their shipped validation rows otherwise +# become structural zeroes. +POST_REFERENCE_ECPS_REQUIRED_INPUTS = frozenset( + { + "fsla_overtime_premium", + "qualified_passenger_vehicle_loan_interest", + "traditional_401k_contributions_desired", + "roth_401k_contributions_desired", + "traditional_ira_contributions_desired", + "roth_ira_contributions_desired", + "self_employed_pension_contributions_desired", + *US_SSI_DISABILITY_CRITERIA_OUTPUT_COLUMNS, + } +) + +# Populated inputs in the frozen reference that have been restored from their +# primary-source derivations. Keeping this separate from the post-reference +# additions lets the anti-rot check reject any future attempt to put a completed +# family back behind a reviewed exclusion. +RESTORED_REFERENCE_ECPS_REQUIRED_INPUTS = frozenset( + { + "alimony_expense", + "alimony_income", + "casualty_loss", + *US_CHILD_SUPPORT_OUTPUT_COLUMNS, + *US_DISABILITY_BENEFITS_OUTPUT_COLUMNS, + *US_WORKERS_COMPENSATION_OUTPUT_COLUMNS, + *US_WEEKS_UNEMPLOYED_OUTPUT_COLUMNS, + *US_WIC_CLAIM_OUTPUT_COLUMNS, + *US_EDUCATOR_EXPENSE_OUTPUT_COLUMNS, + *US_DOMESTIC_PRODUCTION_ALD_OUTPUT_COLUMNS, + *US_OTHER_HEALTH_INSURANCE_NONCONSTANT_PERSON_COLUMNS, + *US_FARM_BUSINESS_INCOME_OUTPUT_COLUMNS, + *US_FORM_4952_OUTPUT_COLUMNS, + *US_CAPITAL_GAIN_DETAILS_OUTPUT_COLUMNS, + *US_SALT_REFUND_OUTPUT_COLUMNS, + *US_ENERGY_SUBSIDY_OUTPUT_COLUMNS, + *US_SCF_NET_WORTH_OUTPUT_COLUMNS, + *US_SIPP_VEHICLE_OUTPUT_COLUMNS, + *US_VOLUNTARY_FILING_OUTPUT_COLUMNS, + *US_SIPP_HEAD_START_OUTPUT_COLUMNS, + *US_SSI_TAKE_UP_OUTPUT_COLUMNS, + *US_HOUSING_INPUTS_OUTPUT_COLUMNS, + *US_MEDICARE_TAKE_UP_OUTPUT_COLUMNS, + *US_PRIOR_YEAR_INCOME_PERSISTED_OUTPUT_COLUMNS, + "spm_unit_pre_subsidy_childcare_expenses", + "household_weight", + "unreimbursed_business_employee_expenses", + *US_QBI_OUTPUT_COLUMNS, + *US_RELATIONSHIP_INPUTS_OUTPUT_COLUMNS, + *US_RETIREMENT_DISTRIBUTION_OUTPUT_COLUMNS, + } +) + _US_PACKAGE = "populace.build.us" _ECPS_PARITY_REFERENCE_RESOURCE = "ecps_parity_reference.json" @@ -136,14 +253,22 @@ class ReformCoverageProbe: id: Stable probe id. name: Human-readable reform description. parameter_changes: ``Reform.from_dict`` payload (country ``us``). + Exactly one of this mapping or ``neutralized_variable`` is + non-empty. + neutralized_variable: Input leaf neutralized by a + structural reform when no parameter change isolates the leaf. budget_measure: The variable whose weighted total change is scored. effect_direction: ``"reform_minus_baseline"`` or ``"baseline_minus_reform"``. + period: Optional probe-specific scoring year. OBBBA probes use 2026 + while the release's source-data period remains 2024. + expected_sign: Required sign of the direction-normalized effect. binding_inputs: The input leaves the reform binds through; named in the failure so the fix target is explicit. - min_abs_effect: The reform fails the smoke gate when - ``abs(effect) < min_abs_effect`` — a floor above simulation noise - and far below the plausible effect, so only a structural zero fails. + min_abs_effect: The reform fails the smoke gate when the effect in its + declared ``expected_sign`` direction is smaller than this floor. + The floor sits above simulation noise and far below the plausible + effect, so a structural zero or wrong-signed result fails. reason: Why a $0 here means a coverage hole (the mechanism). issue: Tracking issue for the restoration work. """ @@ -157,12 +282,27 @@ class ReformCoverageProbe: reason: str issue: str effect_direction: str = "reform_minus_baseline" + period: int | None = None + expected_sign: str = "positive" + neutralized_variable: str | None = None def __post_init__(self) -> None: if not self.id: raise ValueError("ReformCoverageProbe.id is required.") - if not self.parameter_changes: - raise ValueError(f"{self.id}: parameter_changes is required.") + has_params = bool(self.parameter_changes) + has_neutralize = bool(self.neutralized_variable) + if has_params == has_neutralize: + raise ValueError( + f"{self.id}: provide exactly one of parameter_changes or " + "neutralized_variable." + ) + if ( + self.neutralized_variable + and self.neutralized_variable not in self.binding_inputs + ): + raise ValueError( + f"{self.id}: neutralized_variable must be one of binding_inputs." + ) if not self.budget_measure: raise ValueError(f"{self.id}: budget_measure is required.") if self.effect_direction not in { @@ -175,6 +315,12 @@ def __post_init__(self) -> None: ) if not (self.min_abs_effect > 0): raise ValueError(f"{self.id}: min_abs_effect must be positive.") + if self.period is not None and self.period < 1990: + raise ValueError(f"{self.id}: period must be a plausible year.") + if self.expected_sign not in {"positive", "negative"}: + raise ValueError( + f"{self.id}: expected_sign must be 'positive' or 'negative'." + ) @dataclass(frozen=True) @@ -202,6 +348,11 @@ def __post_init__(self) -> None: if column.name in by_name: raise ValueError(f"Duplicate manifest column {column.name!r}.") by_name[column.name] = column + probe_ids: set[str] = set() + for probe in self.probes: + if probe.id in probe_ids: + raise ValueError(f"Duplicate reform coverage probe id {probe.id!r}.") + probe_ids.add(probe.id) object.__setattr__(self, "_by_name", by_name) @property @@ -303,6 +454,17 @@ def load_release_input_coverage_manifest( effect_direction=str( raw_probe.get("effect_direction", "reform_minus_baseline") ), + period=( + None + if raw_probe.get("period") is None + else int(raw_probe["period"]) + ), + expected_sign=str(raw_probe.get("expected_sign", "positive")), + neutralized_variable=( + None + if raw_probe.get("neutralized_variable") is None + else str(raw_probe["neutralized_variable"]) + ), ) ) @@ -385,6 +547,15 @@ def us_release_input_coverage_gate( if column in relevant and column not in present_values: present_values[column] = table[column].to_numpy() + # Frame weights are typed pipeline state, not ordinary table data. The US + # adapter materializes this authoritative vector as ``household_weight`` + # when writing PolicyEngine H5, overwriting any redundant table column. + # Mirror that export behavior here so the gate covers what is persisted. + if "household_weight" in relevant: + present_values.pop("household_weight", None) + if "household" in frame.weighted_entities: + present_values["household_weight"] = frame.weights_for("household").values + degenerate = _degenerate_columns(present_values, engine) return input_column_coverage_gate( present_values.keys(), @@ -435,8 +606,9 @@ def assert_release_input_coverage_manifest_current( (when available): - The declared columns must equal the reference eCPS populated input surface - (``ecps_parity_reference.json``): the manifest is the eCPS export surface, - no more, no less. A layer the incumbent gains/loses must be reflected here. + (``ecps_parity_reference.json``) plus the explicit post-reference inputs + required by shipped reform probes. A change to either surface must be + reflected here. - The three SSI countable-resource asset inputs must be ``required`` with no reviewed exclusion — the #368 red-gate guarantee cannot be quietly undone. - Every declared column must be a real PolicyEngine-US input leaf, and every @@ -453,7 +625,7 @@ def assert_release_input_coverage_manifest_current( failures: list[str] = [] declared = set(manifest.declared_columns) - surface = set(_ecps_populated_layers()) + surface = set(_ecps_populated_layers()) | set(POST_REFERENCE_ECPS_REQUIRED_INPUTS) missing_from_manifest = sorted(surface - declared) extra_in_manifest = sorted(declared - surface) if missing_from_manifest: @@ -464,7 +636,8 @@ def assert_release_input_coverage_manifest_current( ) if extra_in_manifest: failures.append( - "manifest declares column(s) the reference eCPS does not populate " + "manifest declares column(s) neither populated by the reference " + "eCPS nor documented as post-reference required inputs " f"{extra_in_manifest}; regenerate the manifest." ) @@ -483,6 +656,17 @@ def assert_release_input_coverage_manifest_current( "required manifest column (#368)." ) + for column in RESTORED_REFERENCE_ECPS_REQUIRED_INPUTS: + if column in reviewed: + failures.append( + f"{column}: restored reference eCPS input cannot return to a " + "reviewed exclusion." + ) + elif column not in required: + failures.append( + f"{column}: restored reference eCPS input must remain required." + ) + if engine is None: engine = _coverage_engine() if engine is not None: diff --git a/packages/populace-build/src/populace/build/us_runtime/retirement_contributions.py b/packages/populace-build/src/populace/build/us_runtime/retirement_contributions.py new file mode 100644 index 00000000..0ecb9908 --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/retirement_contributions.py @@ -0,0 +1,663 @@ +"""ASEC-reported retirement contributions split into PolicyEngine leaves. + +The retired eCPS pipeline read the measured annual contribution total from +ASEC ``RETCB_VAL`` and allocated it across five account-type inputs. The +allocation is pinned to the final archived implementation at +``9a823603e6b5fb916d65ec45d74c9c7eb0043db1``: + +* ``datasets/cps/cps.py`` lines 1504-1552 gates the self-employed share on + positive ``SEMP_VAL``, the defined-contribution share on positive + ``WSAL_VAL``, and the IRA remainder on either earnings source; and +* ``datasets/cps/imputation_parameters.yaml`` lines 24-55 records the + administrative sources and the four allocation shares. + +The outputs are the five ``*_desired`` input leaves. PolicyEngine-US 1.764.6 +owns the statutory IRA, elective-deferral, and self-employed-plan limits, so +this stage must not cap the measured source amount itself. +""" + +from __future__ import annotations + +from importlib.resources import files +from typing import Any + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import ( + SourceOperationSpec, + SourceStageSpec, + load_source_manifest, +) +from populace.build.source_runtime import ( + SourceRuntimeConfig, + SourceRuntimeContext, + SourceRuntimeError, + run_source_stage, +) +from populace.frame import Frame +from populace.frame.units import US_SCHEMA + +__all__ = [ + "US_RETIREMENT_CONTRIBUTION_NONCONSTANT_PERSON_COLUMNS", + "US_RETIREMENT_CONTRIBUTION_OUTPUT_COLUMNS", + "US_RETIREMENT_CONTRIBUTION_REQUIRED_SOURCE_COLUMNS", + "US_RETIREMENT_CONTRIBUTION_STAGE_NAME", + "derive_us_retirement_contributions_from_manifest", + "impute_us_retirement_contributions_to_puf_support_from_manifest", + "us_retirement_contributions_signal_gate", + "us_retirement_contributions_stage_spec", + "us_retirement_contributions_summary", + "with_us_retirement_contribution_inputs", +] + +QRF: Any | None = None + +US_RETIREMENT_CONTRIBUTION_STAGE_NAME = "retirement_contributions" + +US_RETIREMENT_CONTRIBUTION_OUTPUT_COLUMNS: tuple[str, ...] = ( + "traditional_401k_contributions_desired", + "roth_401k_contributions_desired", + "traditional_ira_contributions_desired", + "roth_ira_contributions_desired", + "self_employed_pension_contributions_desired", +) +US_RETIREMENT_CONTRIBUTION_NONCONSTANT_PERSON_COLUMNS = ( + US_RETIREMENT_CONTRIBUTION_OUTPUT_COLUMNS +) +US_RETIREMENT_CONTRIBUTION_REQUIRED_SOURCE_COLUMNS: tuple[str, ...] = ( + "RETCB_VAL", + "WSAL_VAL", + "SEMP_VAL", +) + +_PERSON_WEIGHT_COLUMN = "person_weight" +_PERSON_SUPPORT_CHANNEL_COLUMN = "person_support_channel" +_BASE_ASEC_SUPPORT_CHANNEL = "asec" +_PUF_TAX_DETAIL_SUPPORT_CHANNEL = "puf_tax_detail" +_PUF_PREDICTORS: tuple[str, ...] = ( + "age", + "is_male", + "has_esi", + "tax_unit_is_joint", + "tax_unit_count_dependents", + "employment_income", + "self_employment_income", + "social_security", +) +_PUF_PREDICTOR_PREFIX = "retirement_contribution_predictor_" +_SHARE_PARAMETER_KEYS = frozenset( + { + "se_pension_share", + "dc_share_of_remainder", + "roth_dc_share", + "traditional_ira_share", + } +) +_PUF_IMPUTATION_PARAMETER_KEYS = frozenset( + { + "predictors", + "max_train_samples", + "n_estimators", + "seed_from_build_config", + "weight", + } +) +# Deliberately broad source-plausibility bound. Its job is to reject a +# default/near-universal surface, while the exact allocation identities below +# protect the family semantics. +_NONZERO_SHARE_BAND = (0.0001, 0.75) + + +def us_retirement_contributions_stage_spec() -> SourceStageSpec: + """Load the packaged retirement-contribution source-stage declaration.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + stage_map = manifest.stage_map() + if US_RETIREMENT_CONTRIBUTION_STAGE_NAME not in stage_map: + raise ValueError( + "US source manifest declares no " + f"{US_RETIREMENT_CONTRIBUTION_STAGE_NAME!r} stage." + ) + spec = stage_map[US_RETIREMENT_CONTRIBUTION_STAGE_NAME] + missing = sorted(set(US_RETIREMENT_CONTRIBUTION_OUTPUT_COLUMNS) - set(spec.outputs)) + if missing: + raise ValueError( + f"{US_RETIREMENT_CONTRIBUTION_STAGE_NAME!r} manifest stage does not " + f"declare output(s) {missing}; the runtime and manifest have drifted." + ) + return spec + + +def _share_parameters(operation: SourceOperationSpec) -> dict[str, float]: + unexpected = sorted(set(operation.parameters) - _SHARE_PARAMETER_KEYS) + missing = sorted(_SHARE_PARAMETER_KEYS - set(operation.parameters)) + if unexpected or missing: + raise SourceRuntimeError( + "US retirement-contribution derivation parameters must match the " + f"archived allocation contract; missing={missing}, " + f"unexpected={unexpected}." + ) + shares = {key: float(operation.parameters[key]) for key in _SHARE_PARAMETER_KEYS} + invalid = {key: value for key, value in shares.items() if not 0 <= value <= 1} + if invalid: + raise SourceRuntimeError( + "US retirement-contribution allocation shares must be in [0, 1]: " + f"{invalid}." + ) + return shares + + +def _numeric_source(frame: pd.DataFrame, column: str) -> np.ndarray: + values = pd.to_numeric(frame[column], errors="coerce").to_numpy(dtype=np.float64) + nonfinite = int(np.count_nonzero(~np.isfinite(values))) + if nonfinite: + raise SourceRuntimeError( + f"US retirement-contribution source {column!r} contains " + f"{nonfinite} nonfinite value(s); the measured source must not be " + "silently replaced." + ) + return values + + +def derive_us_retirement_contributions_from_manifest( + frame: pd.DataFrame | None, + operation: SourceOperationSpec, + _context: SourceRuntimeContext | None, +) -> pd.DataFrame: + """Split measured ``RETCB_VAL`` across the five desired input leaves.""" + + if operation.kind != "derive_retirement_contributions": + raise SourceRuntimeError( + "US retirement-contribution derivation received unexpected " + f"operation {operation.kind!r}." + ) + if frame is None: + raise SourceRuntimeError( + "US retirement-contribution derivation requires the person table " + "to be read first." + ) + missing = [ + column + for column in US_RETIREMENT_CONTRIBUTION_REQUIRED_SOURCE_COLUMNS + if column not in frame.columns + ] + if missing: + raise SourceRuntimeError( + "US retirement-contribution derivation requires measured ASEC " + f"source column(s): {missing}." + ) + shares = _share_parameters(operation) + + result = frame.copy(deep=True) + retirement_contributions = _numeric_source(result, "RETCB_VAL") + negative_source = int(np.count_nonzero(retirement_contributions < 0)) + if negative_source: + raise SourceRuntimeError( + "US retirement-contribution source 'RETCB_VAL' contains " + f"{negative_source} negative value(s)." + ) + has_wages = _numeric_source(result, "WSAL_VAL") > 0 + has_self_employment = _numeric_source(result, "SEMP_VAL") > 0 + has_earned_income = has_wages | has_self_employment + + self_employed = np.where( + has_self_employment, + retirement_contributions * shares["se_pension_share"], + 0.0, + ) + remaining = np.maximum(retirement_contributions - self_employed, 0.0) + dc_pool = np.where( + has_wages, + remaining * shares["dc_share_of_remainder"], + 0.0, + ) + ira_pool = np.where(has_earned_income, remaining - dc_pool, 0.0) + + result["traditional_401k_contributions_desired"] = dc_pool * ( + 1.0 - shares["roth_dc_share"] + ) + result["roth_401k_contributions_desired"] = dc_pool * shares["roth_dc_share"] + result["traditional_ira_contributions_desired"] = ( + ira_pool * shares["traditional_ira_share"] + ) + result["roth_ira_contributions_desired"] = ira_pool * ( + 1.0 - shares["traditional_ira_share"] + ) + result["self_employed_pension_contributions_desired"] = self_employed + return result + + +def impute_us_retirement_contributions_to_puf_support_from_manifest( + frame: pd.DataFrame | None, + operation: SourceOperationSpec, + context: SourceRuntimeContext | None, +) -> pd.DataFrame: + """QRF-impute all five contributions onto the PUF support channel. + + This is the retired second-stage treatment: train on the measured ASEC + rows after their direct split, condition on the clone's PUF-imputed income, + replace all five PUF-half values, clip nonnegative, and zero employer-plan + contributions without wages and self-employed-plan contributions without + self-employment income. + + A non-expanded ASEC frame has no support-channel column; in that first pass + this operation is intentionally a no-op. The base builder runs the same + manifest again after PUF tax-detail imputation, when both support channels + exist and this operation becomes active. + """ + + if operation.kind != "impute_retirement_contributions_to_puf_support": + raise SourceRuntimeError( + "US retirement-contribution PUF imputation received unexpected " + f"operation {operation.kind!r}." + ) + if frame is None: + raise SourceRuntimeError( + "US retirement-contribution PUF imputation requires the person " + "table to be read first." + ) + unexpected = sorted(set(operation.parameters) - _PUF_IMPUTATION_PARAMETER_KEYS) + missing_parameters = sorted( + _PUF_IMPUTATION_PARAMETER_KEYS - set(operation.parameters) + ) + if unexpected or missing_parameters: + raise SourceRuntimeError( + "US retirement-contribution PUF imputation parameters must match " + f"the archived method; missing={missing_parameters}, " + f"unexpected={unexpected}." + ) + if _PERSON_SUPPORT_CHANNEL_COLUMN not in frame.columns: + return frame.copy(deep=True) + + predictors = tuple(str(value) for value in operation.parameters["predictors"]) + if predictors != _PUF_PREDICTORS: + raise SourceRuntimeError( + "US retirement-contribution PUF predictors drifted from the " + f"archived method: expected {list(_PUF_PREDICTORS)}, got " + f"{list(predictors)}." + ) + if operation.parameters["weight"] != _PERSON_WEIGHT_COLUMN: + raise SourceRuntimeError( + "US retirement-contribution PUF imputation must use the typed " + f"person-weight column {_PERSON_WEIGHT_COLUMN!r}." + ) + if operation.parameters["seed_from_build_config"] is not True: + raise SourceRuntimeError( + "US retirement-contribution PUF imputation seed must come from " + "the build config." + ) + max_train_samples = int(operation.parameters["max_train_samples"]) + n_estimators = int(operation.parameters["n_estimators"]) + if max_train_samples <= 0 or n_estimators <= 0: + raise SourceRuntimeError( + "US retirement-contribution PUF max_train_samples and n_estimators " + "must be positive." + ) + + required = [ + _PERSON_WEIGHT_COLUMN, + *(_PUF_PREDICTOR_PREFIX + predictor for predictor in predictors), + *US_RETIREMENT_CONTRIBUTION_OUTPUT_COLUMNS, + ] + missing = [column for column in required if column not in frame.columns] + if missing: + raise SourceRuntimeError( + "US retirement-contribution PUF imputation is missing source " + f"column(s): {missing}." + ) + + channel = frame[_PERSON_SUPPORT_CHANNEL_COLUMN].astype(str) + asec_mask = channel == _BASE_ASEC_SUPPORT_CHANNEL + puf_mask = channel == _PUF_TAX_DETAIL_SUPPORT_CHANNEL + if not asec_mask.any() or not puf_mask.any(): + raise SourceRuntimeError( + "US retirement-contribution PUF imputation requires nonempty ASEC " + "and PUF-tax-detail support channels." + ) + + predictor_columns = [_PUF_PREDICTOR_PREFIX + predictor for predictor in predictors] + training_columns = [ + *predictor_columns, + *US_RETIREMENT_CONTRIBUTION_OUTPUT_COLUMNS, + ] + training = frame.loc[asec_mask, training_columns].copy() + training.columns = [ + *(predictors), + *US_RETIREMENT_CONTRIBUTION_OUTPUT_COLUMNS, + ] + weights = pd.to_numeric( + frame.loc[asec_mask, _PERSON_WEIGHT_COLUMN], errors="coerce" + ).fillna(0.0) + if len(training) > max_train_samples: + sample = training.sample( + n=max_train_samples, + random_state=(context.config.seed if context is not None else 0), + ).index + training = training.loc[sample] + weights = weights.loc[sample] + + test = frame.loc[puf_mask, predictor_columns].copy() + test.columns = list(predictors) + for column in (*predictors, *US_RETIREMENT_CONTRIBUTION_OUTPUT_COLUMNS): + if column in training: + training[column] = pd.to_numeric(training[column], errors="coerce") + if not np.isfinite(training[column].to_numpy(dtype=np.float64)).all(): + raise SourceRuntimeError( + "US retirement-contribution QRF training column " + f"{column!r} contains nonfinite values." + ) + for column in predictors: + test[column] = pd.to_numeric(test[column], errors="coerce") + if not np.isfinite(test[column].to_numpy(dtype=np.float64)).all(): + raise SourceRuntimeError( + "US retirement-contribution QRF prediction column " + f"{column!r} contains nonfinite values." + ) + + global QRF + if QRF is None: + from importlib import import_module + + QRF = import_module("populace.fit").QRF + seed = context.config.seed if context is not None else 0 + fitted = QRF(n_estimators=n_estimators, seed=seed).fit( + training, + list(predictors), + list(US_RETIREMENT_CONTRIBUTION_OUTPUT_COLUMNS), + weights=weights.to_numpy(dtype=np.float64), + ) + predictions = fitted.predict(test).clip(lower=0.0) + no_wages = test["employment_income"].to_numpy(dtype=np.float64) == 0 + no_self_employment = test["self_employment_income"].to_numpy(dtype=np.float64) == 0 + predictions.loc[ + no_wages, + [ + "traditional_401k_contributions_desired", + "roth_401k_contributions_desired", + ], + ] = 0.0 + predictions.loc[ + no_self_employment, + "self_employed_pension_contributions_desired", + ] = 0.0 + + result = frame.copy(deep=True) + for column in US_RETIREMENT_CONTRIBUTION_OUTPUT_COLUMNS: + result.loc[puf_mask, column] = predictions[column].to_numpy(dtype=np.float64) + return result + + +def _person_retirement_predictors(frame: Frame) -> pd.DataFrame: + """Build the retired QRF predictors on person rows, failing closed.""" + + person = frame.table("person") + tax_unit = frame.table("tax_unit") + predictors = pd.DataFrame(index=person.index) + + def _numeric(*columns: str) -> np.ndarray: + for column in columns: + if column in person: + values = pd.to_numeric(person[column], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not np.isfinite(values).all(): + raise SourceRuntimeError( + "US retirement-contribution QRF predictor source " + f"{column!r} contains nonfinite values." + ) + return values + raise SourceRuntimeError( + "US retirement-contribution PUF imputation cannot construct " + f"predictor from any of {list(columns)}." + ) + + predictors["age"] = _numeric("age", "A_AGE") + if "is_male" in person: + predictors["is_male"] = person["is_male"].fillna(False).astype(bool) + elif "is_female" in person: + predictors["is_male"] = ~person["is_female"].fillna(False).astype(bool) + elif "A_SEX" in person: + predictors["is_male"] = _numeric("A_SEX") == 1 + else: + raise SourceRuntimeError( + "US retirement-contribution PUF imputation requires is_male, " + "is_female, or measured A_SEX." + ) + if "has_esi" not in person: + raise SourceRuntimeError( + "US retirement-contribution PUF imputation requires has_esi." + ) + predictors["has_esi"] = person["has_esi"].fillna(False).astype(bool) + + filing_status_column = next( + ( + column + for column in ("filing_status_input", "filing_status") + if column in tax_unit + ), + None, + ) + if filing_status_column is None: + raise SourceRuntimeError( + "US retirement-contribution PUF imputation requires tax-unit filing status." + ) + filing_status = tax_unit[filing_status_column].map( + lambda value: ( + value.decode() if isinstance(value, (bytes, np.bytes_)) else str(value) + ) + ) + joint_by_id = pd.Series( + filing_status.eq("JOINT").to_numpy(), + index=tax_unit["tax_unit_id"].to_numpy(), + ) + predictors["tax_unit_is_joint"] = ( + person["person_tax_unit_id"].map(joint_by_id).fillna(False).astype(bool) + ) + + if "tax_unit_role_input" not in person: + raise SourceRuntimeError( + "US retirement-contribution PUF imputation requires " + "tax_unit_role_input to count dependents." + ) + dependent = ( + person["tax_unit_role_input"] + .map( + lambda value: ( + value.decode() if isinstance(value, (bytes, np.bytes_)) else str(value) + ) + ) + .eq("DEPENDENT") + ) + dependent_count = dependent.groupby(person["person_tax_unit_id"]).transform("sum") + predictors["tax_unit_count_dependents"] = dependent_count.to_numpy(dtype=np.float64) + predictors["employment_income"] = _numeric( + "employment_income_before_lsr", + "WSAL_VAL", + ) + predictors["self_employment_income"] = _numeric( + "self_employment_income_before_lsr", + "SEMP_VAL", + ) + social_security_columns = ( + "social_security_retirement", + "social_security_disability", + "social_security_dependents", + "social_security_survivors", + ) + if all(column in person for column in social_security_columns): + predictors["social_security"] = np.sum( + np.column_stack([_numeric(column) for column in social_security_columns]), + axis=1, + ) + else: + predictors["social_security"] = _numeric("SS_VAL") + return predictors + + +def with_us_retirement_contribution_inputs( + frame: Frame, + *, + seed: int, + time_period: int, +) -> Frame: + """Materialize the five retirement-contribution inputs on a US frame.""" + + if frame.schema != US_SCHEMA: + raise ValueError("US retirement contributions require the US schema.") + person = frame.table("person") + has_support_channels = _PERSON_SUPPORT_CHANNEL_COLUMN in person.columns + if ( + _retirement_contribution_surface_carries_signal(frame) + and not has_support_channels + ): + return frame + + stage_person = person.copy(deep=True) + stage_person[_PERSON_WEIGHT_COLUMN] = frame.resolve_weights("person").values + if has_support_channels: + predictors = _person_retirement_predictors(frame) + for column in _PUF_PREDICTORS: + stage_person[_PUF_PREDICTOR_PREFIX + column] = predictors[column].to_numpy() + output = run_source_stage( + us_retirement_contributions_stage_spec(), + tables={"person": stage_person}, + operation_handlers={ + "derive_retirement_contributions": ( + derive_us_retirement_contributions_from_manifest + ), + "impute_retirement_contributions_to_puf_support": ( + impute_us_retirement_contributions_to_puf_support_from_manifest + ), + }, + config=SourceRuntimeConfig(seed=int(seed), target_year=int(time_period)), + ) + aligned = output.set_index("person_id").reindex(person["person_id"]) + for column in US_RETIREMENT_CONTRIBUTION_OUTPUT_COLUMNS: + if aligned[column].isna().any(): + raise ValueError( + "US retirement-contribution stage output does not cover every " + f"person for {column!r}." + ) + + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + for column in US_RETIREMENT_CONTRIBUTION_OUTPUT_COLUMNS: + tables["person"][column] = aligned[column].to_numpy(dtype=np.float64) + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + ) + + +def us_retirement_contributions_summary(frame: Frame) -> dict[str, object]: + """Return weighted signal and source-allocation diagnostics.""" + + person = frame.table("person") + weights = np.asarray(frame.resolve_weights("person").values, dtype=np.float64) + total_weight = float(weights.sum()) + contributions = { + column: pd.to_numeric(person[column], errors="coerce").to_numpy( + dtype=np.float64 + ) + for column in US_RETIREMENT_CONTRIBUTION_OUTPUT_COLUMNS + } + source = ( + _numeric_source(person, "RETCB_VAL") + if "RETCB_VAL" in person + else np.zeros(len(person), dtype=np.float64) + ) + combined = np.sum(np.column_stack(tuple(contributions.values())), axis=1) + + def _share(values: np.ndarray) -> float: + positive = np.isfinite(values) & (values > 0) + return ( + float(weights[positive].sum()) / total_weight if total_weight > 0 else 0.0 + ) + + nonfinite = { + column: int(np.count_nonzero(~np.isfinite(values))) + for column, values in contributions.items() + } + negative = { + column: int(np.count_nonzero(values < 0)) + for column, values in contributions.items() + } + source_positive = source > 0 + allocation_mismatch = source_positive & ~np.isclose( + combined, + source, + rtol=1e-9, + atol=1e-6, + ) + return { + "nonzero_shares": { + column: _share(values) for column, values in contributions.items() + }, + "weighted_totals": { + column: float(np.sum(np.nan_to_num(values) * weights)) + for column, values in contributions.items() + }, + "nonzero_share_band": list(_NONZERO_SHARE_BAND), + "nonfinite": nonfinite, + "negative": negative, + "source_positive_rows": int(np.count_nonzero(source_positive)), + "allocation_mismatch_rows": int(np.count_nonzero(allocation_mismatch)), + "source_total": float(np.sum(source * weights)), + "allocated_total": float(np.sum(np.nan_to_num(combined) * weights)), + } + + +def us_retirement_contributions_signal_gate(frame: Frame) -> GateResult: + """Require every retirement-contribution leaf to carry valid signal.""" + + person = frame.table("person") + missing = [ + column + for column in US_RETIREMENT_CONTRIBUTION_OUTPUT_COLUMNS + if column not in person.columns + ] + if missing: + return GateResult( + name="retirement_contributions_signal", + passed=False, + failures=(f"person columns missing: {missing}.",), + details={"missing": missing}, + ) + + summary = us_retirement_contributions_summary(frame) + failures: list[str] = [] + low, high = summary["nonzero_share_band"] + for column in US_RETIREMENT_CONTRIBUTION_OUTPUT_COLUMNS: + nonfinite = int(summary["nonfinite"][column]) + negative = int(summary["negative"][column]) + share = float(summary["nonzero_shares"][column]) + if nonfinite: + failures.append(f"{column}: {nonfinite} nonfinite values.") + if negative: + failures.append(f"{column}: {negative} negative values.") + if not (low <= share <= high): + failures.append( + f"{column}: nonzero share {share:.4f} outside plausibility " + f"band [{low}, {high}]." + ) + return GateResult( + name="retirement_contributions_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) + + +def _retirement_contribution_surface_carries_signal(frame: Frame) -> bool: + person = frame.table("person") + if not all( + column in person for column in US_RETIREMENT_CONTRIBUTION_OUTPUT_COLUMNS + ): + return False + return us_retirement_contributions_signal_gate(frame).passed diff --git a/packages/populace-build/src/populace/build/us_runtime/retirement_distributions.py b/packages/populace-build/src/populace/build/us_runtime/retirement_distributions.py new file mode 100644 index 00000000..ee99ccc7 --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/retirement_distributions.py @@ -0,0 +1,762 @@ +"""Measured ASEC retirement-account distributions by account type. + +The archived eCPS construction at commit +``42ed5d45c56df80d754fbe24cce21cfeb8d05cbe`` reads four paired ASEC +distribution slots in ``datasets/cps/cps.py`` lines 1448-1481. Each slot has +an account code (``DST_SC1``, ``DST_SC2``, ``DST_SC1_YNG``, or +``DST_SC2_YNG``) and its measured annual amount. Codes 1 through 6 identify +401(k), 403(b), Roth IRA, regular IRA, Keogh, and SEP accounts respectively. + +The same archived code and +``datasets/cps/imputation_parameters.yaml`` lines 10-15 treat 401(k), 403(b), +regular-IRA, SEP, and Keogh distributions as fully taxable, while Roth-IRA +distributions are tax exempt. After support cloning, archived +``datasets/cps/extended_cps.py`` lines 140-148, 639-745, and 1014-1073 replace, +among the populated leaves restored here, the 401(k), 403(b), Keogh, and SEP +values on the PUF half with CPS-trained QRF predictions. The archive also names +three tax-exempt employer-account companions; their taxable fractions are 1.0, +so they remain zero and are outside the populated coverage surface. The +measured ASEC half remains exact, the PUF taxable-IRA source remains untouched, +and the copied Roth-IRA value remains a direct ASEC mapping. The stage does not +allocate a total across account types or manufacture values for the ASEC code-7 +residual category. +""" + +from __future__ import annotations + +from importlib.resources import files +from typing import Any + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import ( + SourceOperationSpec, + SourceStageSpec, + load_source_manifest, +) +from populace.build.source_runtime import ( + SourceRuntimeConfig, + SourceRuntimeContext, + SourceRuntimeError, + run_source_stage, +) +from populace.frame import Frame +from populace.frame.units import US_SCHEMA + +__all__ = [ + "RETIREMENT_DISTRIBUTIONS_ARCHIVED_DERIVATION_URL", + "RETIREMENT_DISTRIBUTIONS_ARCHIVED_PARAMETERS_URL", + "US_RETIREMENT_DISTRIBUTION_NONCONSTANT_PERSON_COLUMNS", + "US_RETIREMENT_DISTRIBUTION_OUTPUT_COLUMNS", + "US_RETIREMENT_DISTRIBUTION_REQUIRED_SOURCE_COLUMNS", + "US_RETIREMENT_DISTRIBUTION_STAGE_NAME", + "derive_us_retirement_distributions_from_manifest", + "impute_us_retirement_distributions_to_puf_support_from_manifest", + "us_retirement_distributions_signal_gate", + "us_retirement_distributions_stage_spec", + "us_retirement_distributions_summary", + "with_us_retirement_distribution_inputs", +] + +QRF: Any | None = None + +_ARCHIVED_COMMIT = "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe" +_ARCHIVED_DATA_REPOSITORY = "policyengine-" + "us-data" +_ARCHIVED_PACKAGE_PATH = "policyengine_" + "us_data" +RETIREMENT_DISTRIBUTIONS_ARCHIVED_DERIVATION_URL = ( + f"https://github.com/PolicyEngine/{_ARCHIVED_DATA_REPOSITORY}/blob/" + f"{_ARCHIVED_COMMIT}/{_ARCHIVED_PACKAGE_PATH}/datasets/cps/" + "cps.py#L1448-L1481" +) +RETIREMENT_DISTRIBUTIONS_ARCHIVED_PARAMETERS_URL = ( + f"https://github.com/PolicyEngine/{_ARCHIVED_DATA_REPOSITORY}/blob/" + f"{_ARCHIVED_COMMIT}/{_ARCHIVED_PACKAGE_PATH}/datasets/cps/" + "imputation_parameters.yaml#L10-L15" +) + +US_RETIREMENT_DISTRIBUTION_STAGE_NAME = "retirement_distributions" + +US_RETIREMENT_DISTRIBUTION_OUTPUT_COLUMNS: tuple[str, ...] = ( + "taxable_401k_distributions", + "taxable_403b_distributions", + "tax_exempt_ira_distributions", + "taxable_ira_distributions", + "keogh_distributions", + "taxable_sep_distributions", +) +US_RETIREMENT_DISTRIBUTION_NONCONSTANT_PERSON_COLUMNS = ( + US_RETIREMENT_DISTRIBUTION_OUTPUT_COLUMNS +) + +_DISTRIBUTION_SLOT_SUFFIXES = ("1", "2", "1_YNG", "2_YNG") +US_RETIREMENT_DISTRIBUTION_REQUIRED_SOURCE_COLUMNS: tuple[str, ...] = tuple( + column + for suffix in _DISTRIBUTION_SLOT_SUFFIXES + for column in (f"DST_SC{suffix}", f"DST_VAL{suffix}") +) + +_EXPECTED_OUTPUT_BY_ACCOUNT_CODE: dict[int, str] = { + 1: "taxable_401k_distributions", + 2: "taxable_403b_distributions", + 3: "tax_exempt_ira_distributions", + 4: "taxable_ira_distributions", + 5: "keogh_distributions", + 6: "taxable_sep_distributions", +} +_VALID_ACCOUNT_CODES = frozenset(range(8)) +_DERIVE_PARAMETER_KEYS = frozenset({"output_by_account_code"}) +_PERSON_WEIGHT_COLUMN = "person_weight" +_PERSON_SUPPORT_CHANNEL_COLUMN = "person_support_channel" +_BASE_ASEC_SUPPORT_CHANNEL = "asec" +_PUF_TAX_DETAIL_SUPPORT_CHANNEL = "puf_tax_detail" +_PUF_PREDICTORS: tuple[str, ...] = ( + "age", + "is_male", + "has_esi", + "tax_unit_is_joint", + "tax_unit_count_dependents", + "employment_income", + "self_employment_income", + "social_security", +) +# Populated/restored subset of the archived stage-2 CPS-only retirement outputs. +# IRA distributions were not stage-2 QRF targets; the three archived tax-exempt +# employer-account companions remain default zero under the 1.0 taxable shares. +_PUF_QRF_OUTPUT_COLUMNS: tuple[str, ...] = ( + "taxable_401k_distributions", + "taxable_403b_distributions", + "keogh_distributions", + "taxable_sep_distributions", +) +_PUF_PREDICTOR_PREFIX = "retirement_distribution_predictor_" +_PUF_IMPUTATION_PARAMETER_KEYS = frozenset( + { + "predictors", + "max_train_samples", + "n_estimators", + "seed_from_build_config", + "weight", + } +) + +# Weighted shares are intentionally broad but family-specific. Their purpose +# is to reject a missing/default surface while accommodating source-year and +# L0-selection variation. The pinned eCPS shares range from 0.00002 (Keogh) to +# 0.0658 (regular IRA). +_NONZERO_SHARE_BANDS: dict[str, tuple[float, float]] = { + "taxable_401k_distributions": (0.001, 0.20), + "taxable_403b_distributions": (0.00005, 0.05), + "tax_exempt_ira_distributions": (0.0001, 0.05), + "taxable_ira_distributions": (0.001, 0.25), + "keogh_distributions": (0.0000001, 0.005), + "taxable_sep_distributions": (0.00001, 0.05), +} + + +def us_retirement_distributions_stage_spec() -> SourceStageSpec: + """Load the packaged retirement-distribution stage declaration.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + stage_map = manifest.stage_map() + if US_RETIREMENT_DISTRIBUTION_STAGE_NAME not in stage_map: + raise ValueError( + "US source manifest declares no " + f"{US_RETIREMENT_DISTRIBUTION_STAGE_NAME!r} stage." + ) + spec = stage_map[US_RETIREMENT_DISTRIBUTION_STAGE_NAME] + if tuple(spec.outputs) != US_RETIREMENT_DISTRIBUTION_OUTPUT_COLUMNS: + raise ValueError( + f"{US_RETIREMENT_DISTRIBUTION_STAGE_NAME!r} manifest outputs do " + "not match the runtime-owned retirement-distribution family." + ) + return spec + + +def _manifest_output_by_code(operation: SourceOperationSpec) -> dict[int, str]: + unexpected = sorted(set(operation.parameters) - _DERIVE_PARAMETER_KEYS) + if unexpected: + raise SourceRuntimeError( + "US retirement-distribution derivation received unsupported " + f"parameter(s): {unexpected}." + ) + raw = operation.parameters.get("output_by_account_code") + if not isinstance(raw, dict): + raise SourceRuntimeError( + "US retirement-distribution derivation requires an " + "output_by_account_code mapping." + ) + try: + parsed = {int(code): str(output) for code, output in raw.items()} + except (TypeError, ValueError) as exc: + raise SourceRuntimeError( + "US retirement-distribution account codes must be integers." + ) from exc + if parsed != _EXPECTED_OUTPUT_BY_ACCOUNT_CODE: + raise SourceRuntimeError( + "US retirement-distribution account-code mapping drifted from " + f"the archived source: expected {_EXPECTED_OUTPUT_BY_ACCOUNT_CODE}, " + f"got {parsed}." + ) + return parsed + + +def _strict_source_arrays( + frame: pd.DataFrame, +) -> tuple[dict[str, np.ndarray], dict[str, np.ndarray]]: + codes: dict[str, np.ndarray] = {} + amounts: dict[str, np.ndarray] = {} + for suffix in _DISTRIBUTION_SLOT_SUFFIXES: + code_column = f"DST_SC{suffix}" + amount_column = f"DST_VAL{suffix}" + numeric_code = pd.to_numeric(frame[code_column], errors="coerce").to_numpy( + dtype=np.float64 + ) + valid_code = np.isfinite(numeric_code) & ( + numeric_code == np.floor(numeric_code) + ) + valid_code &= np.isin( + numeric_code, + np.fromiter(_VALID_ACCOUNT_CODES, dtype=np.int64), + ) + if not valid_code.all(): + rows = np.flatnonzero(~valid_code)[:5].tolist() + raise SourceRuntimeError( + "US retirement-distribution source requires account codes in " + f"0..7; {code_column} has invalid row(s) {rows}." + ) + numeric_amount = pd.to_numeric(frame[amount_column], errors="coerce").to_numpy( + dtype=np.float64 + ) + valid_amount = np.isfinite(numeric_amount) & (numeric_amount >= 0) + if not valid_amount.all(): + rows = np.flatnonzero(~valid_amount)[:5].tolist() + raise SourceRuntimeError( + "US retirement-distribution source requires finite, " + f"nonnegative amounts; {amount_column} has invalid row(s) {rows}." + ) + code = numeric_code.astype(np.int64) + orphan_amount = (code == 0) & (numeric_amount != 0) + if orphan_amount.any(): + rows = np.flatnonzero(orphan_amount)[:5].tolist() + raise SourceRuntimeError( + f"US retirement-distribution source {amount_column} has a " + f"positive amount with NIU code 0 at row(s) {rows}." + ) + codes[suffix] = code + amounts[suffix] = numeric_amount + return codes, amounts + + +def _derived_outputs( + frame: pd.DataFrame, + output_by_code: dict[int, str], +) -> dict[str, np.ndarray]: + codes, amounts = _strict_source_arrays(frame) + outputs = { + output: np.zeros(len(frame), dtype=np.float64) + for output in US_RETIREMENT_DISTRIBUTION_OUTPUT_COLUMNS + } + for suffix in _DISTRIBUTION_SLOT_SUFFIXES: + for code, output in output_by_code.items(): + outputs[output] += np.where(codes[suffix] == code, amounts[suffix], 0.0) + return outputs + + +def derive_us_retirement_distributions_from_manifest( + frame: pd.DataFrame | None, + operation: SourceOperationSpec, + _context: SourceRuntimeContext | None, +) -> pd.DataFrame: + """Map the four measured ASEC account/amount pairs to six input leaves.""" + + if operation.kind != "derive_retirement_distributions": + raise SourceRuntimeError( + "US retirement-distribution derivation received unexpected " + f"operation {operation.kind!r}." + ) + if frame is None: + raise SourceRuntimeError( + "US retirement-distribution derivation requires the person table " + "to be read first." + ) + missing = [ + column + for column in US_RETIREMENT_DISTRIBUTION_REQUIRED_SOURCE_COLUMNS + if column not in frame.columns + ] + if missing: + raise SourceRuntimeError( + "US retirement-distribution derivation requires measured ASEC " + f"source column(s): {missing}." + ) + output_by_code = _manifest_output_by_code(operation) + result = frame.copy(deep=True) + derived = _derived_outputs(result, output_by_code) + preserved_puf_taxable_ira: np.ndarray | None = None + puf_mask: np.ndarray | None = None + if _PERSON_SUPPORT_CHANNEL_COLUMN in result: + if "taxable_ira_distributions" not in result: + raise SourceRuntimeError( + "US retirement-distribution support derivation requires the " + "PUF-sourced taxable_ira_distributions column." + ) + puf_mask = ( + result[_PERSON_SUPPORT_CHANNEL_COLUMN].astype(str).to_numpy() + == _PUF_TAX_DETAIL_SUPPORT_CHANNEL + ) + preserved_puf_taxable_ira = pd.to_numeric( + result["taxable_ira_distributions"], errors="coerce" + ).to_numpy(dtype=np.float64) + valid = np.isfinite(preserved_puf_taxable_ira[puf_mask]) & ( + preserved_puf_taxable_ira[puf_mask] >= 0 + ) + if not valid.all(): + rows = np.flatnonzero(puf_mask)[np.flatnonzero(~valid)[:5]].tolist() + raise SourceRuntimeError( + "US retirement-distribution support derivation requires " + "finite, nonnegative PUF taxable IRA values; invalid row(s) " + f"{rows}." + ) + + for output, values in derived.items(): + result[output] = values + if output == "taxable_ira_distributions" and puf_mask is not None: + assert preserved_puf_taxable_ira is not None + result.loc[puf_mask, output] = preserved_puf_taxable_ira[puf_mask] + return result + + +def impute_us_retirement_distributions_to_puf_support_from_manifest( + frame: pd.DataFrame | None, + operation: SourceOperationSpec, + context: SourceRuntimeContext | None, +) -> pd.DataFrame: + """Replace PUF-support copies with the retired CPS-trained QRF outputs. + + The first (unexpanded ASEC) pass is a no-op. Once the support spine has an + ASEC half and a PUF-tax-detail half, the common archived eight-predictor QRF + trains on measured ASEC distributions and replaces the four populated + non-IRA outputs restored here on the PUF half. PUF taxable IRA remains + sourced from the tax-detail stage, and Roth IRA remains the direct copied + ASEC mapping. A rare QRF output may legitimately remain zero on the PUF half + when the fixed 5,000-row training sample contains no positive donor; the + measured ASEC half remains authoritative and is never overwritten. + """ + + if operation.kind != "impute_retirement_distributions_to_puf_support": + raise SourceRuntimeError( + "US retirement-distribution PUF imputation received unexpected " + f"operation {operation.kind!r}." + ) + if frame is None: + raise SourceRuntimeError( + "US retirement-distribution PUF imputation requires the person " + "table to be read first." + ) + unexpected = sorted(set(operation.parameters) - _PUF_IMPUTATION_PARAMETER_KEYS) + missing_parameters = sorted( + _PUF_IMPUTATION_PARAMETER_KEYS - set(operation.parameters) + ) + if unexpected or missing_parameters: + raise SourceRuntimeError( + "US retirement-distribution PUF imputation parameters must match " + f"the archived method; missing={missing_parameters}, " + f"unexpected={unexpected}." + ) + if _PERSON_SUPPORT_CHANNEL_COLUMN not in frame.columns: + return frame.copy(deep=True) + + predictors = tuple(str(value) for value in operation.parameters["predictors"]) + if predictors != _PUF_PREDICTORS: + raise SourceRuntimeError( + "US retirement-distribution PUF predictors drifted from the " + f"archived method: expected {list(_PUF_PREDICTORS)}, got " + f"{list(predictors)}." + ) + if operation.parameters["weight"] != _PERSON_WEIGHT_COLUMN: + raise SourceRuntimeError( + "US retirement-distribution PUF imputation must use the typed " + f"person-weight column {_PERSON_WEIGHT_COLUMN!r}." + ) + if operation.parameters["seed_from_build_config"] is not True: + raise SourceRuntimeError( + "US retirement-distribution PUF imputation seed must come from " + "the build config." + ) + max_train_samples = int(operation.parameters["max_train_samples"]) + n_estimators = int(operation.parameters["n_estimators"]) + if max_train_samples <= 0 or n_estimators <= 0: + raise SourceRuntimeError( + "US retirement-distribution PUF max_train_samples and n_estimators " + "must be positive." + ) + + required = [ + _PERSON_WEIGHT_COLUMN, + *(_PUF_PREDICTOR_PREFIX + predictor for predictor in predictors), + *_PUF_QRF_OUTPUT_COLUMNS, + ] + missing = [column for column in required if column not in frame.columns] + if missing: + raise SourceRuntimeError( + "US retirement-distribution PUF imputation is missing source " + f"column(s): {missing}." + ) + + channel = frame[_PERSON_SUPPORT_CHANNEL_COLUMN].astype(str) + asec_mask = channel == _BASE_ASEC_SUPPORT_CHANNEL + puf_mask = channel == _PUF_TAX_DETAIL_SUPPORT_CHANNEL + if not asec_mask.any() or not puf_mask.any(): + raise SourceRuntimeError( + "US retirement-distribution PUF imputation requires nonempty ASEC " + "and PUF-tax-detail support channels." + ) + + predictor_columns = [_PUF_PREDICTOR_PREFIX + value for value in predictors] + training = frame.loc[ + asec_mask, + [*predictor_columns, *_PUF_QRF_OUTPUT_COLUMNS], + ].copy() + training.columns = [*predictors, *_PUF_QRF_OUTPUT_COLUMNS] + weights = pd.to_numeric( + frame.loc[asec_mask, _PERSON_WEIGHT_COLUMN], errors="coerce" + ).fillna(0.0) + if len(training) > max_train_samples: + sample = training.sample( + n=max_train_samples, + random_state=(context.config.seed if context is not None else 0), + ).index + training = training.loc[sample] + weights = weights.loc[sample] + + test = frame.loc[puf_mask, predictor_columns].copy() + test.columns = list(predictors) + for column in (*predictors, *_PUF_QRF_OUTPUT_COLUMNS): + training[column] = pd.to_numeric(training[column], errors="coerce") + if not np.isfinite(training[column].to_numpy(dtype=np.float64)).all(): + raise SourceRuntimeError( + "US retirement-distribution QRF training column " + f"{column!r} contains nonfinite values." + ) + for column in predictors: + test[column] = pd.to_numeric(test[column], errors="coerce") + if not np.isfinite(test[column].to_numpy(dtype=np.float64)).all(): + raise SourceRuntimeError( + "US retirement-distribution QRF prediction column " + f"{column!r} contains nonfinite values." + ) + + global QRF + if QRF is None: + from importlib import import_module + + QRF = import_module("populace.fit").QRF + seed = context.config.seed if context is not None else 0 + fitted = QRF(n_estimators=n_estimators, seed=seed).fit( + training, + list(predictors), + list(_PUF_QRF_OUTPUT_COLUMNS), + weights=weights.to_numpy(dtype=np.float64), + ) + predictions = fitted.predict(test).clip(lower=0.0) + + result = frame.copy(deep=True) + for column in _PUF_QRF_OUTPUT_COLUMNS: + result.loc[puf_mask, column] = predictions[column].to_numpy(dtype=np.float64) + return result + + +def _person_retirement_distribution_predictors(frame: Frame) -> pd.DataFrame: + """Build the archived eight QRF predictors on person rows.""" + + person = frame.table("person") + tax_unit = frame.table("tax_unit") + predictors = pd.DataFrame(index=person.index) + + def _numeric(*columns: str) -> np.ndarray: + for column in columns: + if column in person: + values = pd.to_numeric(person[column], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not np.isfinite(values).all(): + raise SourceRuntimeError( + "US retirement-distribution QRF predictor source " + f"{column!r} contains nonfinite values." + ) + return values + raise SourceRuntimeError( + "US retirement-distribution PUF imputation cannot construct " + f"predictor from any of {list(columns)}." + ) + + predictors["age"] = _numeric("age", "A_AGE") + if "is_male" in person: + predictors["is_male"] = person["is_male"].fillna(False).astype(bool) + elif "is_female" in person: + predictors["is_male"] = ~person["is_female"].fillna(False).astype(bool) + elif "A_SEX" in person: + predictors["is_male"] = _numeric("A_SEX") == 1 + else: + raise SourceRuntimeError( + "US retirement-distribution PUF imputation requires is_male, " + "is_female, or measured A_SEX." + ) + if "has_esi" not in person: + raise SourceRuntimeError( + "US retirement-distribution PUF imputation requires has_esi." + ) + predictors["has_esi"] = person["has_esi"].fillna(False).astype(bool) + + filing_status_column = next( + ( + column + for column in ("filing_status_input", "filing_status") + if column in tax_unit + ), + None, + ) + if filing_status_column is None: + raise SourceRuntimeError( + "US retirement-distribution PUF imputation requires tax-unit filing status." + ) + filing_status = tax_unit[filing_status_column].map( + lambda value: ( + value.decode() if isinstance(value, (bytes, np.bytes_)) else str(value) + ) + ) + joint_by_id = pd.Series( + filing_status.eq("JOINT").to_numpy(), + index=tax_unit["tax_unit_id"].to_numpy(), + ) + predictors["tax_unit_is_joint"] = ( + person["person_tax_unit_id"].map(joint_by_id).fillna(False).astype(bool) + ) + + if "tax_unit_role_input" not in person: + raise SourceRuntimeError( + "US retirement-distribution PUF imputation requires " + "tax_unit_role_input to count dependents." + ) + dependent = ( + person["tax_unit_role_input"] + .map( + lambda value: ( + value.decode() if isinstance(value, (bytes, np.bytes_)) else str(value) + ) + ) + .eq("DEPENDENT") + ) + predictors["tax_unit_count_dependents"] = dependent.groupby( + person["person_tax_unit_id"] + ).transform("sum") + predictors["employment_income"] = _numeric( + "employment_income_before_lsr", + "WSAL_VAL", + ) + predictors["self_employment_income"] = _numeric( + "self_employment_income_before_lsr", + "SEMP_VAL", + ) + social_security_columns = ( + "social_security_retirement", + "social_security_disability", + "social_security_dependents", + "social_security_survivors", + ) + if all(column in person for column in social_security_columns): + predictors["social_security"] = np.sum( + np.column_stack([_numeric(column) for column in social_security_columns]), + axis=1, + ) + else: + predictors["social_security"] = _numeric("SS_VAL") + return predictors + + +def _retirement_distribution_surface_carries_signal(frame: Frame) -> bool: + person = frame.table("person") + if any( + column not in person for column in US_RETIREMENT_DISTRIBUTION_OUTPUT_COLUMNS + ): + return False + return all( + person[column].dropna().nunique() > 1 + for column in US_RETIREMENT_DISTRIBUTION_OUTPUT_COLUMNS + ) + + +def with_us_retirement_distribution_inputs( + frame: Frame, + *, + seed: int, + time_period: int, +) -> Frame: + """Materialize measured retirement-distribution leaves on a US frame.""" + + if frame.schema != US_SCHEMA: + raise ValueError("US retirement distributions require the US schema.") + person = frame.table("person") + has_support_channels = _PERSON_SUPPORT_CHANNEL_COLUMN in person.columns + if ( + _retirement_distribution_surface_carries_signal(frame) + and not has_support_channels + ): + return frame + + stage_person = person.copy(deep=True) + stage_person[_PERSON_WEIGHT_COLUMN] = frame.resolve_weights("person").values + if has_support_channels: + predictors = _person_retirement_distribution_predictors(frame) + for column in _PUF_PREDICTORS: + stage_person[_PUF_PREDICTOR_PREFIX + column] = predictors[column].to_numpy() + output = run_source_stage( + us_retirement_distributions_stage_spec(), + tables={"person": stage_person}, + operation_handlers={ + "derive_retirement_distributions": ( + derive_us_retirement_distributions_from_manifest + ), + "impute_retirement_distributions_to_puf_support": ( + impute_us_retirement_distributions_to_puf_support_from_manifest + ), + }, + config=SourceRuntimeConfig(seed=int(seed), target_year=int(time_period)), + ) + aligned = output.set_index("person_id").reindex(person["person_id"]) + for column in US_RETIREMENT_DISTRIBUTION_OUTPUT_COLUMNS: + if aligned[column].isna().any(): + raise ValueError( + "US retirement-distribution stage output does not cover every " + f"person for {column!r}." + ) + + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + for column in US_RETIREMENT_DISTRIBUTION_OUTPUT_COLUMNS: + tables["person"][column] = aligned[column].to_numpy(dtype=np.float64) + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + ) + + +def us_retirement_distributions_summary(frame: Frame) -> dict[str, object]: + """Return weighted signal and exact source-reconciliation diagnostics.""" + + person = frame.table("person") + weights = np.asarray(frame.resolve_weights("person").values, dtype=np.float64) + total_weight = float(weights.sum()) + values = { + column: pd.to_numeric(person[column], errors="coerce").to_numpy( + dtype=np.float64 + ) + for column in US_RETIREMENT_DISTRIBUTION_OUTPUT_COLUMNS + } + + source_mismatches: dict[str, int] = {} + if all( + column in person + for column in US_RETIREMENT_DISTRIBUTION_REQUIRED_SOURCE_COLUMNS + ): + expected = _derived_outputs(person, _EXPECTED_OUTPUT_BY_ACCOUNT_CODE) + compare = np.ones(len(person), dtype=bool) + if _PERSON_SUPPORT_CHANNEL_COLUMN in person: + compare = ( + person[_PERSON_SUPPORT_CHANNEL_COLUMN].astype(str).to_numpy() + == _BASE_ASEC_SUPPORT_CHANNEL + ) + source_mismatches = { + column: int( + np.count_nonzero( + compare + & ~np.isclose(values[column], expected[column], rtol=0, atol=0) + ) + ) + for column in US_RETIREMENT_DISTRIBUTION_OUTPUT_COLUMNS + } + + def _share(array: np.ndarray) -> float: + positive = np.isfinite(array) & (array > 0) + return ( + float(weights[positive].sum()) / total_weight if total_weight > 0 else 0.0 + ) + + return { + "nonzero_shares": {column: _share(array) for column, array in values.items()}, + "weighted_totals": { + column: float(np.sum(np.nan_to_num(array) * weights)) + for column, array in values.items() + }, + "nonzero_share_bands": { + column: list(_NONZERO_SHARE_BANDS[column]) + for column in US_RETIREMENT_DISTRIBUTION_OUTPUT_COLUMNS + }, + "unique_counts": { + column: int(person[column].dropna().nunique()) + for column in US_RETIREMENT_DISTRIBUTION_OUTPUT_COLUMNS + }, + "nonfinite": { + column: int(np.count_nonzero(~np.isfinite(array))) + for column, array in values.items() + }, + "negative": { + column: int(np.count_nonzero(array < 0)) for column, array in values.items() + }, + "source_mismatches": source_mismatches, + } + + +def us_retirement_distributions_signal_gate(frame: Frame) -> GateResult: + """Require every measured account-type leaf to carry valid source signal.""" + + person = frame.table("person") + missing = [ + column + for column in US_RETIREMENT_DISTRIBUTION_OUTPUT_COLUMNS + if column not in person + ] + if missing: + return GateResult( + name="retirement_distributions_signal", + passed=False, + failures=(f"person columns missing: {missing}.",), + details={"missing": missing}, + ) + + summary = us_retirement_distributions_summary(frame) + failures: list[str] = [] + for column in US_RETIREMENT_DISTRIBUTION_OUTPUT_COLUMNS: + nonfinite = int(summary["nonfinite"][column]) + negative = int(summary["negative"][column]) + unique = int(summary["unique_counts"][column]) + share = float(summary["nonzero_shares"][column]) + low, high = summary["nonzero_share_bands"][column] + mismatch = int(summary["source_mismatches"].get(column, 0)) + if nonfinite: + failures.append(f"{column}: {nonfinite} nonfinite values.") + if negative: + failures.append(f"{column}: {negative} negative values.") + if unique < 2: + failures.append(f"{column}: degenerate with {unique} distinct value(s).") + if not low <= share <= high: + failures.append( + f"{column}: weighted nonzero share {share:.8f} outside " + f"[{low:.8f}, {high:.8f}]." + ) + if mismatch: + failures.append( + f"{column}: {mismatch} row(s) differ from measured ASEC slots." + ) + return GateResult( + name="retirement_distributions_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) diff --git a/packages/populace-build/src/populace/build/us_runtime/salt_refund_income.py b/packages/populace-build/src/populace/build/us_runtime/salt_refund_income.py new file mode 100644 index 00000000..0c14fa3e --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/salt_refund_income.py @@ -0,0 +1,225 @@ +"""IRS PUF state-and-local-tax-refund income for the US build. + +The retired eCPS pipeline carried ``salt_refund_income`` directly from IRS +PUF field ``E00700``. The immutable archived derivation, exported field, +PUF QRF target/override lists, person-allocation path, and versioned PUF +artifact coordinates are exposed below. + +There is no ASEC analogue and no synthetic fallback. The shared PUF +tax-detail stage reduces the processed person array to the identifiable +tax-unit total, imputes that source-backed total only onto the dedicated PUF +support channel, and places it on the unit's first person. PolicyEngine-US +sums this person input at tax-unit grain, so reproducing the retired randomized +filer/spouse ``EARNSPLIT`` allocation would add noise without changing policy +semantics. +""" + +from __future__ import annotations + +from importlib.resources import files + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import SourceStageSpec, load_source_manifest +from populace.frame import Frame + +__all__ = [ + "SALT_REFUND_ARCHIVED_DERIVATION_URL", + "SALT_REFUND_ARCHIVED_EXPORT_URL", + "SALT_REFUND_ARCHIVED_IMPUTATION_URL", + "SALT_REFUND_ARCHIVED_PERSON_ALLOCATION_URL", + "SALT_REFUND_ARCHIVED_PUF_ARTIFACT_URL", + "US_SALT_REFUND_NONCONSTANT_PERSON_COLUMNS", + "US_SALT_REFUND_OUTPUT_COLUMNS", + "US_SALT_REFUND_STAGE_NAME", + "derive_us_salt_refund_income_from_puf", + "us_salt_refund_income_signal_gate", + "us_salt_refund_income_stage_spec", + "us_salt_refund_income_summary", +] + +_ARCHIVED_DATA_REPOSITORY = "policyengine-" + "us-data" +_ARCHIVED_ROOT = ( + "https://github.com/PolicyEngine/" + f"{_ARCHIVED_DATA_REPOSITORY}/blob/" + "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe/" + "policyengine_" + "us_data/" +) +SALT_REFUND_ARCHIVED_DERIVATION_URL = _ARCHIVED_ROOT + "datasets/puf/puf.py#L707" +SALT_REFUND_ARCHIVED_EXPORT_URL = _ARCHIVED_ROOT + "datasets/puf/puf.py#L804-L850" +SALT_REFUND_ARCHIVED_IMPUTATION_URL = ( + _ARCHIVED_ROOT + "calibration/puf_impute.py#L90-L198" +) +SALT_REFUND_ARCHIVED_PERSON_ALLOCATION_URL = ( + _ARCHIVED_ROOT + "datasets/puf/puf.py#L1513-L1601" +) +SALT_REFUND_ARCHIVED_PUF_ARTIFACT_URL = ( + _ARCHIVED_ROOT + "datasets/puf/puf.py#L1655-L1660" +) + +US_SALT_REFUND_STAGE_NAME = "puf_tax_detail" +US_SALT_REFUND_OUTPUT_COLUMNS: tuple[str, ...] = ("salt_refund_income",) +US_SALT_REFUND_NONCONSTANT_PERSON_COLUMNS = US_SALT_REFUND_OUTPUT_COLUMNS + +_OUTPUT = US_SALT_REFUND_OUTPUT_COLUMNS[0] +_PERSON_SUPPORT_CHANNEL_COLUMN = "person_support_channel" +_BASE_ASEC_SUPPORT_CHANNEL = "asec" +_PUF_TAX_DETAIL_SUPPORT_CHANNEL = "puf_tax_detail" + +# The pinned processed PUF donor has a 10.74% weighted positive person share +# (13.58% at tax-unit grain). The broad bounds reject a default-only column +# and QRF smearing while allowing support sampling and composition changes. +_OVERALL_POSITIVE_SHARE_BAND = (0.005, 0.20) +_PUF_POSITIVE_SHARE_BAND = (0.01, 0.40) + + +def us_salt_refund_income_stage_spec() -> SourceStageSpec: + """Load and validate the shared PUF tax-detail stage declaration.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + spec = manifest.stage_map()[US_SALT_REFUND_STAGE_NAME] + missing = sorted(set(US_SALT_REFUND_OUTPUT_COLUMNS) - set(spec.outputs)) + if missing: + raise ValueError( + f"{US_SALT_REFUND_STAGE_NAME!r} manifest stage does not declare " + f"SALT-refund output(s) {missing}." + ) + return spec + + +def derive_us_salt_refund_income_from_puf( + puf: pd.DataFrame, + *, + source_column: str = "E00700", + output_column: str = _OUTPUT, +) -> pd.DataFrame: + """Carry the archived IRS PUF E00700 field without alteration.""" + + if source_column not in puf.columns: + raise ValueError( + f"PUF SALT-refund derivation requires source column {source_column!r}." + ) + values = pd.to_numeric(puf[source_column], errors="coerce").to_numpy( + dtype=np.float64 + ) + nonfinite = ~np.isfinite(values) + if bool(nonfinite.any()): + raise ValueError( + f"PUF SALT-refund source column {source_column!r} contains " + f"{int(np.count_nonzero(nonfinite))} nonnumeric or nonfinite value(s)." + ) + negative = values < 0.0 + if bool(negative.any()): + raise ValueError( + f"PUF SALT-refund source column {source_column!r} contains " + f"{int(np.count_nonzero(negative))} negative value(s)." + ) + + result = puf.copy(deep=True) + result[output_column] = values + return result + + +def us_salt_refund_income_summary(frame: Frame) -> dict[str, object]: + """Return weighted signal, support-channel, and validity diagnostics.""" + + person = frame.table("person") + values = pd.to_numeric(person[_OUTPUT], errors="coerce").to_numpy(dtype=np.float64) + weights = np.asarray(frame.resolve_weights("person").values, dtype=np.float64) + finite = np.isfinite(values) + positive = finite & (values > 0.0) + total_weight = float(weights.sum()) + summary: dict[str, object] = { + "positive_share": ( + float(weights[positive].sum()) / total_weight if total_weight > 0.0 else 0.0 + ), + "positive_share_band": list(_OVERALL_POSITIVE_SHARE_BAND), + "weighted_total": float((np.nan_to_num(values) * weights).sum()), + "nonfinite": int(np.count_nonzero(~finite)), + "negative": int(np.count_nonzero(finite & (values < 0.0))), + } + if _PERSON_SUPPORT_CHANNEL_COLUMN in person.columns: + channel_values = person[_PERSON_SUPPORT_CHANNEL_COLUMN].to_numpy() + channels: dict[str, dict[str, float | int]] = {} + for channel in person[_PERSON_SUPPORT_CHANNEL_COLUMN].dropna().unique(): + channel_mask = channel_values == channel + channel_weight = float(weights[channel_mask].sum()) + channel_positive = channel_mask & positive + channels[str(channel)] = { + "positive_rows": int(np.count_nonzero(channel_positive)), + "positive_share": ( + float(weights[channel_positive].sum()) / channel_weight + if channel_weight > 0.0 + else 0.0 + ), + "weighted_total": float( + (np.nan_to_num(values[channel_mask]) * weights[channel_mask]).sum() + ), + } + summary["channels"] = channels + return summary + + +def us_salt_refund_income_signal_gate(frame: Frame) -> GateResult: + """Require finite, nonnegative, source-aligned SALT-refund signal.""" + + person = frame.table("person") + missing = [ + column for column in US_SALT_REFUND_OUTPUT_COLUMNS if column not in person + ] + if missing: + return GateResult( + name="salt_refund_income_signal", + passed=False, + failures=(f"person columns missing: {missing}.",), + details={"missing": missing}, + ) + + summary = us_salt_refund_income_summary(frame) + failures: list[str] = [] + if summary["nonfinite"]: + failures.append(f"{_OUTPUT} nonfinite values: {int(summary['nonfinite'])}.") + if summary["negative"]: + failures.append(f"{_OUTPUT} negative values: {int(summary['negative'])}.") + share = float(summary["positive_share"]) + low, high = summary["positive_share_band"] + if not (low <= share <= high): + failures.append( + f"SALT-refund positive share {share:.6f} outside plausibility " + f"band [{low}, {high}]." + ) + + channels = summary.get("channels") + if isinstance(channels, dict): + asec = channels.get(_BASE_ASEC_SUPPORT_CHANNEL) + puf = channels.get(_PUF_TAX_DETAIL_SUPPORT_CHANNEL) + if asec is None: + failures.append("SALT-refund signal gate is missing the ASEC channel.") + elif float(asec["weighted_total"]) != 0.0: + failures.append( + "salt_refund_income must remain zero on the source-unobserved " + "ASEC support channel." + ) + if puf is None: + failures.append( + "SALT-refund signal gate is missing the PUF tax-detail channel." + ) + else: + puf_share = float(puf["positive_share"]) + puf_low, puf_high = _PUF_POSITIVE_SHARE_BAND + if not (puf_low <= puf_share <= puf_high): + failures.append( + f"PUF SALT-refund positive share {puf_share:.6f} outside " + f"plausibility band [{puf_low}, {puf_high}]." + ) + + return GateResult( + name="salt_refund_income_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) diff --git a/packages/populace-build/src/populace/build/us_runtime/scf_auto_loans.py b/packages/populace-build/src/populace/build/us_runtime/scf_auto_loans.py new file mode 100644 index 00000000..7a40e1a3 --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/scf_auto_loans.py @@ -0,0 +1,624 @@ +"""Household auto-loan inputs restored from the full 2022 SCF. + +The public SCF summary extract used by :mod:`scf_wealth` contains the common +demographic/income predictors but not the four vehicle-loan amount/rate pairs. +Those live in the full public extract. This module joins the two artifacts on +``(y1, yy1)`` (family and implicate), derives the retired eCPS targets, and fits +one weighted multi-target QRF draw per recipient household. + +The two legacy outputs preserve the retired implementation's semantics: +``auto_loan_balance`` is the sum of the four *original financed amounts*, not +the remaining principal balance, and ``auto_loan_interest`` is each amount +times its reported rate. Negative SCF codes are floored before rates are +divided by 10,000. + +OBBBA uses a distinct later pure input, +``qualified_passenger_vehicle_loan_interest``. SCF 2022 cannot observe a loan +originated after 2024 or final-assembly eligibility. Treasury and IRS estimate +roughly six million qualifying loans are issued annually (16m new light-vehicle +sales × 60% financed × 60% final assembly in the United States). The port uses +that official annual incidence as an expected-share proxy: six million divided +by weighted households with positive imputed interest, capped at one, times +each household's interest. This is deliberately transparent and conservative; +it restores a nonzero, correctly bounded reform input without pretending the +2022 donor directly observed a post-2024 statutory fact. +""" + +from __future__ import annotations + +from importlib.resources import files +from pathlib import Path + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import SourceStageSpec, load_source_manifest +from populace.build.us_runtime.eligibility_inputs import _own_children_in_household +from populace.build.us_runtime.scf_wealth import ( + SCF_WEALTH_PREDICTORS, + _recipient_cps_race, + _scf_summary_predictor_table, + _sum_present, +) +from populace.frame import Frame +from populace.frame.units import US_SCHEMA + +__all__ = [ + "QUALIFIED_AUTO_LOAN_ANNUAL_ISSUANCE_TARGET", + "SCF_2022_FULL_EXTRACT_MEMBER", + "SCF_2022_FULL_EXTRACT_MEMBER_SHA256", + "SCF_2022_FULL_EXTRACT_URL", + "SCF_2022_FULL_EXTRACT_ZIP_SHA256", + "SCF_AUTO_LOAN_AMOUNT_COLUMNS", + "SCF_AUTO_LOAN_RATE_COLUMNS", + "US_SCF_AUTO_LOAN_NONCONSTANT_HOUSEHOLD_COLUMNS", + "US_SCF_AUTO_LOAN_OUTPUT_COLUMNS", + "fetch_scf_2022_full_extract", + "impute_us_scf_auto_loans", + "load_scf_2022_auto_loan_donor", + "qualified_auto_loan_interest_proxy", + "us_scf_auto_loans_signal_gate", + "us_scf_auto_loans_stage_spec", + "us_scf_auto_loans_summary", + "with_us_scf_auto_loan_inputs", +] + +SCF_2022_FULL_EXTRACT_URL = "https://www.federalreserve.gov/econres/files/scf2022s.zip" +SCF_2022_FULL_EXTRACT_MEMBER = "p22i6.dta" + +# The source/member coordinates and 22,975-row public-file shape were verified +# against the Federal Reserve SCF download page and a successful retired build. +# The full archive itself was not retained locally. Keep these explicit rather +# than (incorrectly) reusing the unrelated summary-extract hashes; one +# network-enabled provisioning run can fill the two pins without changing any +# transformation semantics. +SCF_2022_FULL_EXTRACT_ZIP_SHA256: str | None = None +SCF_2022_FULL_EXTRACT_MEMBER_SHA256: str | None = None + +SCF_AUTO_LOAN_AMOUNT_COLUMNS: tuple[str, ...] = ( + "x2209", + "x2309", + "x2409", + "x7158", +) +SCF_AUTO_LOAN_RATE_COLUMNS: tuple[str, ...] = ( + "x2219", + "x2319", + "x2419", + "x7170", +) + +US_SCF_AUTO_LOAN_OUTPUT_COLUMNS: tuple[str, ...] = ( + "auto_loan_balance", + "auto_loan_interest", + "qualified_passenger_vehicle_loan_interest", +) +US_SCF_AUTO_LOAN_NONCONSTANT_HOUSEHOLD_COLUMNS = US_SCF_AUTO_LOAN_OUTPUT_COLUMNS + +QUALIFIED_AUTO_LOAN_ANNUAL_ISSUANCE_TARGET = 6_000_000.0 + +_SCF_STAGE_NAME = "scf_wealth" +_DONOR_WEIGHT_COLUMN = "scf_weight" +_HOUSEHOLD_ID_COLUMN = "person_household_id" +_DEFAULT_N_ESTIMATORS = 100 +_AUTO_LOAN_NONZERO_SHARE_BAND = (0.10, 0.60) +_SCF_HTTP_USER_AGENT = ( + "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) " + "AppleWebKit/537.36 (KHTML, like Gecko) Chrome/124.0 Safari/537.36" +) + +_RECIPIENT_INTEREST_DIVIDEND_COLUMNS = ( + "taxable_interest_income", + "tax_exempt_interest_income", + "qualified_dividend_income", + "non_qualified_dividend_income", +) +_RECIPIENT_SS_PENSION_COLUMNS = ( + "social_security_retirement", + "social_security_disability", + "social_security_survivors", + "social_security_dependents", + "taxable_private_pension_income", + "tax_exempt_private_pension_income", +) +_REQUIRED_RECIPIENT_PERSON_COLUMNS = ( + _HOUSEHOLD_ID_COLUMN, + "age", + "is_female", + "PRDTRACE", + "PRDTHSP", + "A_MARITL", + "A_LINENO", + "PH_SEQ", + "PEPAR1", + "PEPAR2", + "employment_income_before_lsr", +) + + +def _sha256_hexdigest(payload: bytes) -> str: + import hashlib + + return hashlib.sha256(payload).hexdigest() + + +def us_scf_auto_loans_stage_spec() -> SourceStageSpec: + """Load and validate the shared ``scf_wealth`` stage declaration.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + spec = manifest.stage_map()[_SCF_STAGE_NAME] + missing = sorted(set(US_SCF_AUTO_LOAN_OUTPUT_COLUMNS) - set(spec.outputs)) + if missing: + raise ValueError( + f"{_SCF_STAGE_NAME!r} manifest stage does not declare auto-loan " + f"output(s) {missing}." + ) + return spec + + +def fetch_scf_2022_full_extract( + cache_dir: str | Path | None = None, + *, + expected_member_sha256: str | None = SCF_2022_FULL_EXTRACT_MEMBER_SHA256, + expected_zip_sha256: str | None = SCF_2022_FULL_EXTRACT_ZIP_SHA256, +) -> Path: + """Download, optionally verify, extract, and cache the full 2022 SCF.""" + + import io + import urllib.request + import zipfile + + root = ( + Path(cache_dir).expanduser() + if cache_dir is not None + else Path.home() / ".cache" / "populace" / "scf" + ) + root.mkdir(parents=True, exist_ok=True) + target = root / SCF_2022_FULL_EXTRACT_MEMBER + if target.exists() and target.stat().st_size > 0: + if expected_member_sha256 is None: + return target + digest = _sha256_hexdigest(target.read_bytes()) + if digest == expected_member_sha256: + return target + + request = urllib.request.Request( + SCF_2022_FULL_EXTRACT_URL, + headers={"User-Agent": _SCF_HTTP_USER_AGENT}, + ) + with urllib.request.urlopen(request, timeout=120) as response: # noqa: S310 + payload = response.read() + if expected_zip_sha256 is not None: + digest = _sha256_hexdigest(payload) + if digest != expected_zip_sha256: + raise ValueError( + "SCF 2022 full-extract zip failed sha-256 verification: " + f"expected {expected_zip_sha256}, got {digest}." + ) + with zipfile.ZipFile(io.BytesIO(payload)) as archive: + if SCF_2022_FULL_EXTRACT_MEMBER not in archive.namelist(): + raise ValueError( + "SCF 2022 full-extract zip is missing expected member " + f"{SCF_2022_FULL_EXTRACT_MEMBER!r}." + ) + member_bytes = archive.read(SCF_2022_FULL_EXTRACT_MEMBER) + if expected_member_sha256 is not None: + digest = _sha256_hexdigest(member_bytes) + if digest != expected_member_sha256: + raise ValueError( + "SCF 2022 full-extract member failed sha-256 verification: " + f"expected {expected_member_sha256}, got {digest}." + ) + target.write_bytes(member_bytes) + return target + + +def load_scf_2022_auto_loan_donor( + summary_extract_path: str | Path, + full_extract_path: str | Path, +) -> pd.DataFrame: + """Join SCF summary/full implicates and derive the two retired targets.""" + + join_columns = ["y1", "yy1"] + summary = pd.read_stata(summary_extract_path, convert_categoricals=False) + try: + full = pd.read_stata( + full_extract_path, + convert_categoricals=False, + columns=[ + *join_columns, + *SCF_AUTO_LOAN_AMOUNT_COLUMNS, + *SCF_AUTO_LOAN_RATE_COLUMNS, + ], + ) + except ValueError as error: + if "not found in the Stata data set" not in str(error): + raise + raise ValueError( + f"SCF 2022 full extract missing required column(s): {error}." + ) from error + for label, table, required in ( + ( + "summary", + summary, + {*join_columns}, + ), + ( + "full", + full, + { + *join_columns, + *SCF_AUTO_LOAN_AMOUNT_COLUMNS, + *SCF_AUTO_LOAN_RATE_COLUMNS, + }, + ), + ): + missing = sorted(required - set(table.columns)) + if missing: + raise ValueError( + f"SCF 2022 {label} extract missing required column(s): {missing}." + ) + if table.duplicated(join_columns).any(): + raise ValueError( + f"SCF 2022 {label} extract does not have one-to-one (y1, yy1) keys." + ) + + predictors = _scf_summary_predictor_table(summary) + keyed_summary = summary.loc[:, join_columns].reset_index(drop=True) + keyed_summary["_summary_order"] = np.arange(len(keyed_summary)) + keyed_full = full.loc[ + :, + [ + *join_columns, + *SCF_AUTO_LOAN_AMOUNT_COLUMNS, + *SCF_AUTO_LOAN_RATE_COLUMNS, + ], + ] + joined = keyed_summary.merge( + keyed_full, + on=join_columns, + how="outer", + validate="one_to_one", + indicator=True, + sort=False, + ) + unmatched = joined.loc[joined["_merge"] != "both", join_columns + ["_merge"]] + if not unmatched.empty: + raise ValueError( + "SCF 2022 summary/full extracts have unmatched (y1, yy1) implicates; " + f"a one-to-one join is required ({len(unmatched)} unmatched row(s))." + ) + joined = joined.sort_values("_summary_order").reset_index(drop=True) + + amounts = joined.loc[:, list(SCF_AUTO_LOAN_AMOUNT_COLUMNS)].apply( + pd.to_numeric, errors="coerce" + ) + rates = joined.loc[:, list(SCF_AUTO_LOAN_RATE_COLUMNS)].apply( + pd.to_numeric, errors="coerce" + ) + amounts = amounts.fillna(0.0).clip(lower=0.0).to_numpy(dtype=np.float64) + rates = rates.fillna(0.0).clip(lower=0.0).to_numpy(dtype=np.float64) / 10_000.0 + + donor = predictors.reset_index(drop=True) + donor["auto_loan_balance"] = amounts.sum(axis=1) + donor["auto_loan_interest"] = np.sum(amounts * rates, axis=1) + donor = donor.loc[ + donor[_DONOR_WEIGHT_COLUMN] > 0, + [ + *SCF_WEALTH_PREDICTORS, + "auto_loan_balance", + "auto_loan_interest", + _DONOR_WEIGHT_COLUMN, + ], + ] + return donor.reset_index(drop=True) + + +def _scf_reference_person_mask(person: pd.DataFrame) -> np.ndarray: + """Select one recipient per household using the retired SCF rules.""" + + required = { + _HOUSEHOLD_ID_COLUMN, + "age", + "is_female", + "A_MARITL", + "A_LINENO", + } + missing = sorted(required - set(person.columns)) + if missing: + raise ValueError( + f"SCF auto-loan reference-person selection missing column(s): {missing}." + ) + + working = pd.DataFrame( + { + "household_id": person[_HOUSEHOLD_ID_COLUMN].to_numpy(), + "age": pd.to_numeric(person["age"], errors="coerce").fillna(0.0).to_numpy(), + "is_female": person["is_female"].astype(bool).to_numpy(), + "is_married": pd.to_numeric(person["A_MARITL"], errors="coerce") + .isin([1, 2]) + .to_numpy(), + "position": np.arange(len(person)), + } + ) + mask = np.zeros(len(person), dtype=bool) + for _, group in working.groupby("household_id", sort=False): + adults = group.loc[group["age"] >= 18] + if adults.empty: + chosen = int(group.loc[group["age"].idxmax(), "position"]) + elif len(adults) == 1: + chosen = int(adults.iloc[0]["position"]) + elif len(adults) == 2 and bool(adults["is_married"].all()): + if adults["is_female"].nunique() == 1: + chosen = int(adults.loc[adults["age"].idxmax(), "position"]) + else: + chosen = int(adults.loc[~adults["is_female"]].iloc[0]["position"]) + else: + chosen = int(adults.loc[adults["age"].idxmax(), "position"]) + mask[chosen] = True + return mask + + +def _recipient_household_predictor_table(frame: Frame) -> pd.DataFrame: + person = frame.table("person") + household = frame.table("household") + missing = sorted(set(_REQUIRED_RECIPIENT_PERSON_COLUMNS) - set(person.columns)) + if missing: + raise ValueError( + f"US SCF auto-loan imputation requires recipient person column(s): " + f"{missing}." + ) + + mask = _scf_reference_person_mask(person) + per_person = pd.DataFrame(index=person.index) + per_person["household_id"] = person[_HOUSEHOLD_ID_COLUMN].to_numpy() + per_person["age"] = pd.to_numeric(person["age"], errors="coerce").fillna(0.0) + per_person["is_female"] = person["is_female"].astype(bool).astype(float) + per_person["cps_race"] = _recipient_cps_race(person) + per_person["is_married"] = ( + pd.to_numeric(person["A_MARITL"], errors="coerce").isin([1, 2]).astype(float) + ) + per_person["own_children_in_household"] = _own_children_in_household(person) + + income = pd.DataFrame( + { + "household_id": person[_HOUSEHOLD_ID_COLUMN].to_numpy(), + "employment_income": np.maximum( + pd.to_numeric(person["employment_income_before_lsr"], errors="coerce") + .fillna(0.0) + .to_numpy(dtype=np.float64), + 0.0, + ), + "interest_dividend_income": _sum_present( + person, _RECIPIENT_INTEREST_DIVIDEND_COLUMNS + ), + "social_security_pension_income": _sum_present( + person, _RECIPIENT_SS_PENSION_COLUMNS + ), + } + ) + household_income = income.groupby("household_id", sort=False).sum() + reference = per_person.loc[mask].set_index("household_id") + if reference.index.duplicated().any(): + raise ValueError("SCF auto-loan receiver selected multiple reference persons.") + reference = reference.join(household_income, how="left") + + household_ids = household["household_id"].to_numpy() + missing_ids = sorted(set(household_ids) - set(reference.index)) + extra_ids = sorted(set(reference.index) - set(household_ids)) + if missing_ids or extra_ids: + raise ValueError( + "SCF auto-loan receiver household alignment failed: " + f"missing={missing_ids[:5]}, extra={extra_ids[:5]}." + ) + return reference.reindex(household_ids).loc[:, list(SCF_WEALTH_PREDICTORS)] + + +def impute_us_scf_auto_loans( + frame: Frame, + donor: pd.DataFrame, + *, + seed: int, + n_estimators: int = _DEFAULT_N_ESTIMATORS, +) -> pd.DataFrame: + """Draw legacy auto balance/interest jointly, once per household.""" + + from populace.fit import QRF + + targets = ("auto_loan_balance", "auto_loan_interest") + required = (*SCF_WEALTH_PREDICTORS, *targets, _DONOR_WEIGHT_COLUMN) + missing = [column for column in required if column not in donor.columns] + if missing: + raise ValueError(f"SCF auto-loan donor table missing column(s): {missing}.") + fit_frame = donor.loc[:, list(required)].copy() + for column in required: + fit_frame[column] = pd.to_numeric(fit_frame[column], errors="coerce").fillna( + 0.0 + ) + fitted = QRF(n_estimators=int(n_estimators), seed=int(seed)).fit( + fit_frame, + predictors=list(SCF_WEALTH_PREDICTORS), + targets=list(targets), + weights=_DONOR_WEIGHT_COLUMN, + ) + drawn = fitted.predict(_recipient_household_predictor_table(frame)) + result = pd.DataFrame(index=frame.table("household").index) + for column in targets: + result[column] = np.maximum(np.asarray(drawn[column], dtype=np.float64), 0.0) + return result + + +def qualified_auto_loan_interest_proxy( + auto_loan_interest: np.ndarray, + household_weights: np.ndarray, + *, + annual_issuance_target: float = QUALIFIED_AUTO_LOAN_ANNUAL_ISSUANCE_TARGET, +) -> tuple[np.ndarray, float]: + """Return expected qualifying interest and the applied incidence share.""" + + interest = np.maximum(np.asarray(auto_loan_interest, dtype=np.float64), 0.0) + weights = np.asarray(household_weights, dtype=np.float64) + if interest.shape != weights.shape: + raise ValueError("Auto-loan interest and household weights must align.") + positive_mass = float(weights[interest > 0].sum()) + share = ( + min(float(annual_issuance_target) / positive_mass, 1.0) + if positive_mass > 0 + else 0.0 + ) + return interest * share, share + + +def _surface_has_signal(household: pd.DataFrame) -> bool: + if not all(column in household for column in US_SCF_AUTO_LOAN_OUTPUT_COLUMNS): + return False + for column in US_SCF_AUTO_LOAN_OUTPUT_COLUMNS: + values = pd.to_numeric(household[column], errors="coerce").dropna() + if values.nunique() < 2 or not (values > 0).any(): + return False + qualified = pd.to_numeric( + household["qualified_passenger_vehicle_loan_interest"], errors="coerce" + ).fillna(0.0) + interest = pd.to_numeric(household["auto_loan_interest"], errors="coerce").fillna( + 0.0 + ) + return bool((qualified <= interest).all()) + + +def with_us_scf_auto_loan_inputs( + frame: Frame, + *, + seed: int, + time_period: int, + scf_auto_loan_donor: pd.DataFrame, + n_estimators: int = _DEFAULT_N_ESTIMATORS, +) -> Frame: + """Restore both legacy and OBBBA auto-loan inputs on ``household``.""" + + if frame.schema != US_SCHEMA: + raise ValueError("US SCF auto-loan inputs require the US schema.") + us_scf_auto_loans_stage_spec() + household = frame.table("household") + if _surface_has_signal(household): + return frame + + imputed = impute_us_scf_auto_loans( + frame, + scf_auto_loan_donor, + seed=int(seed), + n_estimators=int(n_estimators), + ) + qualified, _ = qualified_auto_loan_interest_proxy( + imputed["auto_loan_interest"].to_numpy(dtype=np.float64), + frame.weights_for("household").values, + ) + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + tables["household"]["auto_loan_balance"] = imputed["auto_loan_balance"].to_numpy( + dtype=np.float64 + ) + tables["household"]["auto_loan_interest"] = imputed["auto_loan_interest"].to_numpy( + dtype=np.float64 + ) + tables["household"]["qualified_passenger_vehicle_loan_interest"] = qualified + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + ) + + +def us_scf_auto_loans_summary(frame: Frame) -> dict[str, object]: + """Weighted incidence/totals and relation diagnostics for the family.""" + + household = frame.table("household") + weights = np.asarray(frame.weights_for("household").values, dtype=np.float64) + total_weight = float(weights.sum()) + values = { + column: pd.to_numeric(household[column], errors="coerce") + .fillna(0.0) + .to_numpy(dtype=np.float64) + for column in US_SCF_AUTO_LOAN_OUTPUT_COLUMNS + if column in household + } + interest = values.get("auto_loan_interest", np.zeros(len(household))) + qualified = values.get( + "qualified_passenger_vehicle_loan_interest", np.zeros(len(household)) + ) + positive_mass = float(weights[interest > 0].sum()) + qualified_total = float(np.dot(qualified, weights)) + interest_total = float(np.dot(interest, weights)) + return { + "auto_loan_balance_weighted_total": float( + np.dot(values.get("auto_loan_balance", np.zeros(len(household))), weights) + ), + "auto_loan_interest_weighted_total": interest_total, + "qualified_interest_weighted_total": qualified_total, + "auto_loan_interest_nonzero_share": ( + positive_mass / total_weight if total_weight > 0 else 0.0 + ), + "qualified_interest_share": ( + qualified_total / interest_total if interest_total > 0 else 0.0 + ), + "qualified_exceeds_interest_count": int((qualified > interest).sum()), + "negative_counts": { + column: int((column_values < 0).sum()) + for column, column_values in values.items() + }, + "unique_counts": { + column: int(pd.Series(column_values).nunique()) + for column, column_values in values.items() + }, + "auto_loan_nonzero_share_band": list(_AUTO_LOAN_NONZERO_SHARE_BAND), + } + + +def us_scf_auto_loans_signal_gate(frame: Frame) -> GateResult: + """Require nonnegative, nonconstant legacy and OBBBA auto-loan signal.""" + + household = frame.table("household") + missing = [ + column + for column in US_SCF_AUTO_LOAN_OUTPUT_COLUMNS + if column not in household.columns + ] + if missing: + return GateResult( + name="scf_auto_loans_signal", + passed=False, + failures=(f"household columns missing: {missing}.",), + details={"missing": missing}, + ) + + summary = us_scf_auto_loans_summary(frame) + failures: list[str] = [] + for column, count in summary["unique_counts"].items(): + if count < 2: + failures.append(f"{column}: constant column carries no signal.") + for column, count in summary["negative_counts"].items(): + if count: + failures.append(f"{column}: {count} negative value(s).") + if summary["qualified_exceeds_interest_count"]: + failures.append( + "qualified_passenger_vehicle_loan_interest exceeds auto_loan_interest " + f"on {summary['qualified_exceeds_interest_count']} household(s)." + ) + share = float(summary["auto_loan_interest_nonzero_share"]) + low, high = _AUTO_LOAN_NONZERO_SHARE_BAND + if not (low <= share <= high): + failures.append( + f"auto-loan nonzero share {share:.3f} outside plausibility band " + f"[{low}, {high}]." + ) + if float(summary["qualified_interest_weighted_total"]) <= 0: + failures.append("qualified auto-loan interest weighted total is zero.") + return GateResult( + name="scf_auto_loans_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) diff --git a/packages/populace-build/src/populace/build/us_runtime/scf_wealth.py b/packages/populace-build/src/populace/build/us_runtime/scf_wealth.py index 9136b13a..f1ddcc88 100644 --- a/packages/populace-build/src/populace/build/us_runtime/scf_wealth.py +++ b/packages/populace-build/src/populace/build/us_runtime/scf_wealth.py @@ -1,4 +1,4 @@ -"""SCF-imputed household financial-asset inputs (SSI countable resources). +"""SCF-imputed financial-asset inputs and household net worth. Without this stage the published dataset stores none of the liquid-asset input columns SSI's resource test reads, so each silently takes the @@ -10,7 +10,7 @@ populace issues #356/#368 (the dropped-column counterpart of #278's zeroed input bases). -The three columns are the SSI countable-resource leaves in PolicyEngine-US +The three person columns are the SSI countable-resource leaves in PolicyEngine-US (``policyengine_us/parameters/gov/ssa/ssi/eligibility/resources/ countable.yaml``: ``bank_account_assets`` / ``stock_assets`` / ``bond_assets``). They are imputed from the Federal Reserve Survey of @@ -80,16 +80,21 @@ The nonzero *incidence* already matches the dense-native reference (40.3% vs 42.5%); the gap is amount and calibration, not the grain. -The household-total SCF wealth components the manifest also declares -(net_worth, primary-residence value, auto-loan interest, etc.) remain -follow-up work (populace #356); this stage ships only the three SSI -liquid-asset leaves, which is what makes the SSI resource-limit reform class -score again. - -Healing behavior: a frame already carrying all three columns *with signal* -passes through untouched (idempotent). A frame carrying a constant -``bank_account_assets`` — indistinguishable from the engine's 0 default — is -re-imputed rather than trusted. +The stage also restores the household ``net_worth`` input from the SCF +``networth`` summary-extract field. At archived commit ``42ed5d45``, +``utils/asset_imputation.py`` lines 15-19 and 182-184 make that field the direct +``scf_net_worth`` QRF anchor; ``calibration/source_impute.py`` lines 1146-1164 +fits it, and lines 1324-1338 reconcile construction-only balance-sheet +components to the anchor before exporting their signed sum as ``net_worth``. +Persisting the source-backed anchor is therefore the policy-facing result of +the retired reconciliation without manufacturing or exporting its internal +``scf_*`` component columns. ``net_worth`` is signed: indebted households may +have negative values, so it must never be clipped at zero. + +Healing behavior: a frame already carrying all three person columns and the +household net-worth column *with signal* passes through untouched (idempotent). +A constant ``bank_account_assets`` or ``net_worth`` column — indistinguishable +from an engine broadcast default — is re-imputed rather than trusted. """ from __future__ import annotations @@ -112,12 +117,16 @@ "SCF_2022_SUMMARY_EXTRACT_URL", "SCF_2022_SUMMARY_EXTRACT_ZIP_SHA256", "SCF_FINANCIAL_ASSET_TARGET_COMPONENTS", + "SCF_NET_WORTH_TARGET_COMPONENTS", "SCF_WEALTH_PREDICTORS", "US_SCF_FINANCIAL_ASSET_OUTPUT_COLUMNS", + "US_SCF_NET_WORTH_OUTPUT_COLUMNS", + "US_SCF_WEALTH_NONCONSTANT_HOUSEHOLD_COLUMNS", "US_SCF_WEALTH_NONCONSTANT_PERSON_COLUMNS", "US_SCF_WEALTH_STAGE_NAME", "fetch_scf_2022_summary_extract", "impute_us_scf_financial_assets", + "impute_us_scf_net_worth", "load_scf_2022_financial_asset_donor", "us_scf_wealth_signal_gate", "us_scf_wealth_stage_spec", @@ -163,10 +172,16 @@ "bond_assets", ) +#: The signed household-grain PolicyEngine input restored from SCF networth. +US_SCF_NET_WORTH_OUTPUT_COLUMNS: tuple[str, ...] = ("net_worth",) + #: Release gates require these person columns to carry signal (≥2 values). US_SCF_WEALTH_NONCONSTANT_PERSON_COLUMNS: tuple[str, ...] = ( US_SCF_FINANCIAL_ASSET_OUTPUT_COLUMNS ) +US_SCF_WEALTH_NONCONSTANT_HOUSEHOLD_COLUMNS: tuple[str, ...] = ( + US_SCF_NET_WORTH_OUTPUT_COLUMNS +) #: Each output column and the SCF summary-extract components it sums #: (retired pipeline @ 42ed5d45, ``utils/asset_imputation.py``). @@ -175,6 +190,9 @@ "stock_assets": ("stocks", "nmmf"), "bond_assets": ("bond",), } +SCF_NET_WORTH_TARGET_COMPONENTS: dict[str, tuple[str, ...]] = { + "net_worth": ("networth",), +} #: The predictor columns the imputation conditions on, in the manifest's #: declared order. @@ -203,6 +221,23 @@ #: feature). _DONOR_WEIGHT_COLUMN = "scf_weight" +#: Raw SCF summary-extract columns used to construct the shared eight-feature +#: donor surface. The auto-loan family reads the full SCF for its targets but +#: deliberately reuses this summary-extract predictor construction so its QRF +#: conditions on the same concepts as the retired pipeline and the SSI-assets +#: stage. +_SCF_SUMMARY_PREDICTOR_SOURCE_COLUMNS: tuple[str, ...] = ( + "wgt", + "age", + "hhsex", + "racecl5", + "married", + "kids", + "wageinc", + "intdivinc", + "ssretinc", +) + _PERSON_WEIGHT_COLUMN = "person_weight" _HOUSEHOLD_ID_COLUMN = "person_household_id" @@ -213,12 +248,18 @@ _BANK_NONZERO_SHARE_BAND = (0.25, 0.85) _STOCK_NONZERO_SHARE_BAND = (0.03, 0.40) _BOND_NONZERO_SHARE_BAND = (0.001, 0.12) +#: Household net worth is nearly always nonzero and legitimately signed. The +#: pinned eCPS nonzero share is 0.997241; the SHA-pinned SCF donor is 0.996678 +#: nonzero, 0.921249 positive, and 0.075429 negative by survey weight. +_NET_WORTH_NONZERO_SHARE_BAND = (0.90, 1.0) +_NET_WORTH_POSITIVE_SHARE_BAND = (0.75, 1.0) +_NET_WORTH_NEGATIVE_SHARE_BAND = (0.001, 0.25) def us_scf_wealth_stage_spec() -> SourceStageSpec: """Load the packaged ``scf_wealth`` source-stage manifest entry. - The three produced columns must be a subset of the stage's declared + The policy-facing produced columns must be a subset of the stage's declared outputs, so the manifest stays the single source declaration (survey, citation, sentinel policy) this runtime implements. """ @@ -233,7 +274,11 @@ def us_scf_wealth_stage_spec() -> SourceStageSpec: ) spec = stage_map[US_SCF_WEALTH_STAGE_NAME] declared = set(spec.outputs) - missing = [c for c in US_SCF_FINANCIAL_ASSET_OUTPUT_COLUMNS if c not in declared] + required_outputs = ( + *US_SCF_FINANCIAL_ASSET_OUTPUT_COLUMNS, + *US_SCF_NET_WORTH_OUTPUT_COLUMNS, + ) + missing = [column for column in required_outputs if column not in declared] if missing: raise ValueError( f"{US_SCF_WEALTH_STAGE_NAME!r} manifest stage does not declare " @@ -249,6 +294,52 @@ def _replace_sentinels(values: pd.Series) -> pd.Series: return numeric.mask(numeric.isin(_SCF_SENTINELS), 0.0).fillna(0.0) +def _scf_summary_predictor_table(raw: pd.DataFrame) -> pd.DataFrame: + """Build the common SCF predictor/weight table from summary-extract rows. + + The function keeps row order and row count unchanged. Callers that join a + second SCF artifact (the auto-loan port) can therefore align the returned + predictors through the summary extract's ``(y1, yy1)`` keys before applying + the positive-weight filter. + """ + + missing = sorted(set(_SCF_SUMMARY_PREDICTOR_SOURCE_COLUMNS) - set(raw.columns)) + if missing: + raise ValueError( + f"SCF 2022 summary extract missing required column(s): {missing}." + ) + + donor = pd.DataFrame(index=raw.index) + donor["age"] = _replace_sentinels(raw["age"]).to_numpy(dtype=np.float64) + donor["is_female"] = (pd.to_numeric(raw["hhsex"], errors="coerce") == 2).to_numpy( + dtype=np.float64 + ) + donor["cps_race"] = ( + pd.to_numeric(raw["racecl5"], errors="coerce") + .map(_SCF_RACECL5_TO_CPS_RACE) + .fillna(7) + .to_numpy(dtype=np.float64) + ) + donor["is_married"] = ( + pd.to_numeric(raw["married"], errors="coerce") == 1 + ).to_numpy(dtype=np.float64) + donor["own_children_in_household"] = np.maximum( + _replace_sentinels(raw["kids"]).to_numpy(dtype=np.float64), 0.0 + ) + donor["employment_income"] = _replace_sentinels(raw["wageinc"]).to_numpy( + dtype=np.float64 + ) + donor["interest_dividend_income"] = _replace_sentinels(raw["intdivinc"]).to_numpy( + dtype=np.float64 + ) + donor["social_security_pension_income"] = _replace_sentinels( + raw["ssretinc"] + ).to_numpy(dtype=np.float64) + weight = pd.to_numeric(raw["wgt"], errors="coerce").fillna(0.0) + donor[_DONOR_WEIGHT_COLUMN] = np.maximum(weight.to_numpy(dtype=np.float64), 0.0) + return donor.loc[:, [*SCF_WEALTH_PREDICTORS, _DONOR_WEIGHT_COLUMN]] + + def _sha256_hexdigest(payload: bytes) -> str: """SHA-256 hex digest of a byte payload.""" @@ -343,10 +434,9 @@ def load_scf_2022_financial_asset_donor(path: str | Path) -> pd.DataFrame: ``https://www.federalreserve.gov/econres/files/scfp2022s.zip``). Returns: - A household-grain donor DataFrame carrying the three target columns - (``bank_account_assets`` / ``stock_assets`` / ``bond_assets``), the - eight predictor columns, and the donor weight column - (``scf_weight``). + A household-grain donor DataFrame carrying the three liquid-asset + targets, signed ``net_worth``, the eight predictor columns, and the + donor weight column (``scf_weight``). Raises: ValueError: If the extract is missing a required source column. @@ -358,15 +448,8 @@ def load_scf_2022_financial_asset_donor(path: str | Path) -> pd.DataFrame: "stocks", "nmmf", "bond", - "wgt", - "age", - "hhsex", - "racecl5", - "married", - "kids", - "wageinc", - "intdivinc", - "ssretinc", + "networth", + *_SCF_SUMMARY_PREDICTOR_SOURCE_COLUMNS, } missing = sorted(required - set(raw.columns)) if missing: @@ -383,34 +466,13 @@ def load_scf_2022_financial_asset_donor(path: str | Path) -> pd.DataFrame: dtype=np.float64 ) donor[output] = np.maximum(total, 0.0) - # Predictors (head/household values from the summary extract). - donor["age"] = _replace_sentinels(raw["age"]).to_numpy(dtype=np.float64) - donor["is_female"] = ( - pd.to_numeric(raw["hhsex"], errors="coerce") == 2 - ).to_numpy(dtype=np.float64) - donor["cps_race"] = ( - pd.to_numeric(raw["racecl5"], errors="coerce") - .map(_SCF_RACECL5_TO_CPS_RACE) - .fillna(7) - .to_numpy(dtype=np.float64) - ) - donor["is_married"] = ( - pd.to_numeric(raw["married"], errors="coerce") == 1 - ).to_numpy(dtype=np.float64) - donor["own_children_in_household"] = np.maximum( - _replace_sentinels(raw["kids"]).to_numpy(dtype=np.float64), 0.0 - ) - donor["employment_income"] = _replace_sentinels(raw["wageinc"]).to_numpy( - dtype=np.float64 - ) - donor["interest_dividend_income"] = _replace_sentinels(raw["intdivinc"]).to_numpy( - dtype=np.float64 - ) - donor["social_security_pension_income"] = _replace_sentinels( - raw["ssretinc"] - ).to_numpy(dtype=np.float64) - weight = pd.to_numeric(raw["wgt"], errors="coerce").fillna(0.0) - donor[_DONOR_WEIGHT_COLUMN] = np.maximum(weight.to_numpy(dtype=np.float64), 0.0) + # SCF networth is already the complete signed balance-sheet aggregate. + # Preserve negative values: they are indebted-household source signal, not + # missing-value sentinels. + donor["net_worth"] = pd.to_numeric(raw["networth"], errors="coerce").fillna(0.0) + predictors = _scf_summary_predictor_table(raw) + for column in (*SCF_WEALTH_PREDICTORS, _DONOR_WEIGHT_COLUMN): + donor[column] = predictors[column].to_numpy(dtype=np.float64) return donor.loc[donor[_DONOR_WEIGHT_COLUMN] > 0].reset_index(drop=True) @@ -482,9 +544,7 @@ def _household_head_mask(person: pd.DataFrame) -> np.ndarray: person of each household). """ - household = pd.to_numeric( - person[_HOUSEHOLD_ID_COLUMN], errors="coerce" - ).to_numpy() + household = pd.to_numeric(person[_HOUSEHOLD_ID_COLUMN], errors="coerce").to_numpy() lineno = pd.to_numeric(person["A_LINENO"], errors="coerce").fillna(9_999).to_numpy() position = np.arange(len(person)) # Sort by household, then line number, then original position; the first @@ -549,6 +609,46 @@ def _household_head_predictor_table(person: pd.DataFrame) -> pd.DataFrame: return features.loc[:, list(SCF_WEALTH_PREDICTORS)] +def _draw_us_scf_targets( + person: pd.DataFrame, + donor: pd.DataFrame, + *, + targets: tuple[str, ...], + seed: int, + n_estimators: int, +) -> tuple[pd.DataFrame, np.ndarray]: + """Draw SCF household targets for one reference person per household.""" + + from populace.fit import QRF + + donor_missing = [ + column + for column in (*SCF_WEALTH_PREDICTORS, *targets, _DONOR_WEIGHT_COLUMN) + if column not in donor.columns + ] + if donor_missing: + raise ValueError(f"SCF donor table missing column(s): {donor_missing}.") + + fit_frame = donor.loc[ + :, + [*SCF_WEALTH_PREDICTORS, *targets, _DONOR_WEIGHT_COLUMN], + ].copy() + for column in (*SCF_WEALTH_PREDICTORS, *targets): + fit_frame[column] = pd.to_numeric(fit_frame[column], errors="coerce") + if not np.isfinite(fit_frame[column].to_numpy(dtype=np.float64)).all(): + raise ValueError(f"SCF donor column {column!r} contains nonfinite values.") + + fitted = QRF(n_estimators=int(n_estimators), seed=int(seed)).fit( + fit_frame, + predictors=list(SCF_WEALTH_PREDICTORS), + targets=list(targets), + weights=_DONOR_WEIGHT_COLUMN, + ) + head_mask = _household_head_mask(person) + head_features = _household_head_predictor_table(person).loc[head_mask] + return fitted.predict(head_features), head_mask + + def impute_us_scf_financial_assets( person: pd.DataFrame, donor: pd.DataFrame, @@ -577,42 +677,13 @@ def impute_us_scf_financial_assets( each non-negative float64 and zero off the household heads. """ - from populace.fit import QRF - - donor_missing = [ - column - for column in ( - *SCF_WEALTH_PREDICTORS, - *US_SCF_FINANCIAL_ASSET_OUTPUT_COLUMNS, - _DONOR_WEIGHT_COLUMN, - ) - if column not in donor.columns - ] - if donor_missing: - raise ValueError(f"SCF donor table missing column(s): {donor_missing}.") - - fit_frame = donor.loc[ - :, - [ - *SCF_WEALTH_PREDICTORS, - *US_SCF_FINANCIAL_ASSET_OUTPUT_COLUMNS, - _DONOR_WEIGHT_COLUMN, - ], - ].copy() - for column in (*SCF_WEALTH_PREDICTORS, *US_SCF_FINANCIAL_ASSET_OUTPUT_COLUMNS): - fit_frame[column] = pd.to_numeric(fit_frame[column], errors="coerce").fillna( - 0.0 - ) - - fitted = QRF(n_estimators=int(n_estimators), seed=int(seed)).fit( - fit_frame, - predictors=list(SCF_WEALTH_PREDICTORS), - targets=list(US_SCF_FINANCIAL_ASSET_OUTPUT_COLUMNS), - weights=_DONOR_WEIGHT_COLUMN, + drawn, head_mask = _draw_us_scf_targets( + person, + donor, + targets=US_SCF_FINANCIAL_ASSET_OUTPUT_COLUMNS, + seed=seed, + n_estimators=n_estimators, ) - head_mask = _household_head_mask(person) - head_features = _household_head_predictor_table(person).loc[head_mask] - drawn = fitted.predict(head_features) result = pd.DataFrame(index=person.index) head_positions = np.flatnonzero(head_mask) for column in US_SCF_FINANCIAL_ASSET_OUTPUT_COLUMNS: @@ -624,6 +695,57 @@ def impute_us_scf_financial_assets( return result +def impute_us_scf_net_worth( + person: pd.DataFrame, + household: pd.DataFrame, + donor: pd.DataFrame, + *, + seed: int, + n_estimators: int = _DEFAULT_N_ESTIMATORS, +) -> pd.Series: + """Impute signed SCF net worth and align it to household-table order. + + The retired source imputer treats SCF ``networth`` as a direct QRF anchor, + predicts at person grain, and keeps one reference-person value for each + household. Its internal component reconciliation is constructed to equal + this anchor. This port persists that policy-facing household result without + exposing the construction-only ``scf_*`` components. + """ + + if "household_id" not in household: + raise ValueError("US SCF net-worth imputation requires household_id.") + drawn, head_mask = _draw_us_scf_targets( + person, + donor, + targets=US_SCF_NET_WORTH_OUTPUT_COLUMNS, + seed=seed, + n_estimators=n_estimators, + ) + reference_household_ids = person.loc[head_mask, _HOUSEHOLD_ID_COLUMN].to_numpy() + if pd.Series(reference_household_ids).duplicated().any(): + raise ValueError( + "US SCF net-worth imputation produced duplicate reference households." + ) + values = pd.to_numeric(drawn["net_worth"], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not np.isfinite(values).all(): + raise ValueError("US SCF net-worth predictions contain nonfinite values.") + by_household = pd.Series(values, index=reference_household_ids) + household_ids = household["household_id"].to_numpy() + aligned = by_household.reindex(household_ids) + if aligned.isna().any(): + missing = household_ids[aligned.isna().to_numpy()][:5].tolist() + raise ValueError( + f"US SCF net-worth imputation does not cover household id(s) {missing}." + ) + return pd.Series( + aligned.to_numpy(dtype=np.float64), + index=household.index, + name="net_worth", + ) + + def _bank_assets_carry_signal(person: pd.DataFrame) -> bool: """Whether the persisted bank-assets column is trustworthy as-is. @@ -635,6 +757,22 @@ def _bank_assets_carry_signal(person: pd.DataFrame) -> bool: return values.nunique() > 1 +def _net_worth_carries_signal(household: pd.DataFrame) -> bool: + """Whether persisted household net worth is finite and nonconstant.""" + + if "net_worth" not in household: + return False + values = pd.to_numeric(household["net_worth"], errors="coerce").to_numpy( + dtype=np.float64 + ) + return ( + len(values) == len(household) + and bool(np.isfinite(values).all()) + and np.unique(values).size > 1 + and bool((values != 0).any()) + ) + + def with_us_scf_wealth_inputs( frame: Frame, *, @@ -642,12 +780,12 @@ def with_us_scf_wealth_inputs( time_period: int, scf_donor: pd.DataFrame, ) -> Frame: - """Impute the three SSI financial-asset columns onto a US frame. + """Impute SSI financial assets and signed household net worth. - A frame already carrying all three output columns with a non-constant - bank-assets distribution passes through untouched (idempotent). Any - other surface — columns missing, or bank assets constant at the engine - default — is re-imputed from the SCF donor. + A frame already carrying all three person outputs and household net worth + with nonconstant signal passes through untouched (idempotent). Missing or + default-valued surfaces are independently healed from the SCF donor, so an + existing liquid-asset draw is not replaced merely because net worth is new. Args: frame: A US-schema frame whose person table carries the recipient @@ -659,7 +797,8 @@ def with_us_scf_wealth_inputs( :func:`load_scf_2022_financial_asset_donor`). Returns: - A new frame whose person table carries all three asset columns. + A new frame whose person table carries all three asset columns and + whose household table carries signed ``net_worth``. Raises: ValueError: If the frame is not US-schema. @@ -671,18 +810,29 @@ def with_us_scf_wealth_inputs( # imputation is skipped, so a manifest/runtime drift fails loudly. us_scf_wealth_stage_spec() person = frame.table("person") - have_all = all( + household = frame.table("household") + have_all_assets = all( column in person.columns for column in US_SCF_FINANCIAL_ASSET_OUTPUT_COLUMNS ) - if have_all and _bank_assets_carry_signal(person): + assets_carry_signal = have_all_assets and _bank_assets_carry_signal(person) + net_worth_carries_signal = _net_worth_carries_signal(household) + if assets_carry_signal and net_worth_carries_signal: return frame - imputed = impute_us_scf_financial_assets( - person, scf_donor, seed=int(seed) - ) tables = {entity: frame.table(entity).copy() for entity in frame.entities} - for column in US_SCF_FINANCIAL_ASSET_OUTPUT_COLUMNS: - tables["person"][column] = imputed[column].to_numpy(dtype=np.float64) + if not assets_carry_signal: + imputed_assets = impute_us_scf_financial_assets( + person, scf_donor, seed=int(seed) + ) + for column in US_SCF_FINANCIAL_ASSET_OUTPUT_COLUMNS: + tables["person"][column] = imputed_assets[column].to_numpy(dtype=np.float64) + if not net_worth_carries_signal: + tables["household"]["net_worth"] = impute_us_scf_net_worth( + person, + household, + scf_donor, + seed=int(seed), + ).to_numpy(dtype=np.float64) return Frame( tables, frame.schema, @@ -693,11 +843,16 @@ def with_us_scf_wealth_inputs( def us_scf_wealth_summary(frame: Frame) -> dict[str, object]: - """Weighted financial-asset summary for gates and release manifests.""" + """Weighted financial-asset and net-worth diagnostics.""" person = frame.table("person") + household = frame.table("household") weights = np.asarray(frame.resolve_weights("person").values, dtype=np.float64) total_weight = float(weights.sum()) + household_weights = np.asarray( + frame.resolve_weights("household").values, dtype=np.float64 + ) + total_household_weight = float(household_weights.sum()) def _nonzero_share(column: str) -> float: values = pd.to_numeric(person[column], errors="coerce").fillna(0.0).to_numpy() @@ -717,12 +872,26 @@ def _mean_positive(column: str) -> float: ) unique_counts = { - column: int( - pd.to_numeric(person[column], errors="coerce").dropna().nunique() - ) + column: int(pd.to_numeric(person[column], errors="coerce").dropna().nunique()) for column in US_SCF_FINANCIAL_ASSET_OUTPUT_COLUMNS if column in person.columns } + net_worth = ( + pd.to_numeric(household["net_worth"], errors="coerce").to_numpy( + dtype=np.float64 + ) + if "net_worth" in household + else np.asarray([], dtype=np.float64) + ) + finite_net_worth = np.isfinite(net_worth) + + def _net_worth_share(mask: np.ndarray) -> float: + return ( + float(household_weights[mask].sum()) / total_household_weight + if total_household_weight > 0 and len(mask) == len(household_weights) + else 0.0 + ) + return { "bank_account_assets_nonzero_share": _nonzero_share("bank_account_assets") if "bank_account_assets" in person.columns @@ -740,11 +909,34 @@ def _mean_positive(column: str) -> float: "stock_nonzero_share_band": list(_STOCK_NONZERO_SHARE_BAND), "bond_nonzero_share_band": list(_BOND_NONZERO_SHARE_BAND), "unique_counts": unique_counts, + "net_worth_unique_count": int( + pd.to_numeric(household["net_worth"], errors="coerce").dropna().nunique() + ) + if "net_worth" in household + else 0, + "net_worth_nonfinite": int(np.count_nonzero(~np.isfinite(net_worth))), + "net_worth_nonzero_share": _net_worth_share( + finite_net_worth & (net_worth != 0) + ), + "net_worth_positive_share": _net_worth_share( + finite_net_worth & (net_worth > 0) + ), + "net_worth_negative_share": _net_worth_share( + finite_net_worth & (net_worth < 0) + ), + "net_worth_weighted_total": float( + np.sum(np.where(finite_net_worth, net_worth, 0.0) * household_weights) + ) + if len(net_worth) == len(household_weights) + else 0.0, + "net_worth_nonzero_share_band": list(_NET_WORTH_NONZERO_SHARE_BAND), + "net_worth_positive_share_band": list(_NET_WORTH_POSITIVE_SHARE_BAND), + "net_worth_negative_share_band": list(_NET_WORTH_NEGATIVE_SHARE_BAND), } def us_scf_wealth_signal_gate(frame: Frame) -> GateResult: - """Require the financial-asset surface to carry a plausible distribution. + """Require financial assets and net worth to carry plausible distributions. Fails when a column is missing or constant, or when a weighted nonzero-share leaves its plausibility band — each of which reproduces @@ -758,12 +950,23 @@ def us_scf_wealth_signal_gate(frame: Frame) -> GateResult: for column in US_SCF_FINANCIAL_ASSET_OUTPUT_COLUMNS if column not in person.columns ] - if missing: + missing_household = [ + column + for column in US_SCF_NET_WORTH_OUTPUT_COLUMNS + if column not in frame.table("household") + ] + if missing or missing_household: return GateResult( name="scf_wealth_signal", passed=False, - failures=(f"person columns missing: {missing}.",), - details={"missing": missing}, + failures=( + f"person columns missing: {missing}; household columns missing: " + f"{missing_household}.", + ), + details={ + "missing_person": missing, + "missing_household": missing_household, + }, ) summary = us_scf_wealth_summary(frame) @@ -786,6 +989,36 @@ def us_scf_wealth_signal_gate(frame: Frame) -> GateResult: f"{label} nonzero share {share:.3f} outside plausibility band " f"[{low}, {high}]." ) + if int(summary["net_worth_nonfinite"]): + failures.append( + f"net_worth: {summary['net_worth_nonfinite']} nonfinite value(s)." + ) + if int(summary["net_worth_unique_count"]) < 2: + failures.append("net_worth: constant column carries no signal.") + for share_key, band, label in ( + ( + "net_worth_nonzero_share", + _NET_WORTH_NONZERO_SHARE_BAND, + "nonzero", + ), + ( + "net_worth_positive_share", + _NET_WORTH_POSITIVE_SHARE_BAND, + "positive", + ), + ( + "net_worth_negative_share", + _NET_WORTH_NEGATIVE_SHARE_BAND, + "negative", + ), + ): + share = float(summary[share_key]) + low, high = band + if not low <= share <= high: + failures.append( + f"net_worth {label} share {share:.3f} outside plausibility " + f"band [{low}, {high}]." + ) return GateResult( name="scf_wealth_signal", passed=not failures, diff --git a/packages/populace-build/src/populace/build/us_runtime/sipp_head_start.py b/packages/populace-build/src/populace/build/us_runtime/sipp_head_start.py new file mode 100644 index 00000000..337cdcdf --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/sipp_head_start.py @@ -0,0 +1,848 @@ +"""Measured SIPP Head Start proxy for eligible preschool-age children. + +The 2023 SIPP asks whether a nursery-school or preschool enrollee attends a +federally sponsored program, naming Head Start, Even Start, and Fair Start as +examples. This stage uses that measured December response as the available +Head Start proxy to restore ``takes_up_head_start_if_eligible``; it does not +misdescribe the instrument as a program-only identifier. It also does not use +the retired NIEER scalar or stand in for Early Head Start, whose infant/toddler +population is outside the source question and PolicyEngine Head Start domain. + +Labels fail closed. Direct ``EEDHEADST`` answers are retained only when +``AEDHEADST == 1`` (as reported). A not-in-universe row is a negative only +when an upstream reported screen proves either no school enrollment or a +reported grade other than nursery/preschool. Hot-decked Head Start answers, +imputed upstream screens, and unknown screen states never enter training. +That yields 45 positive and 740 strict negative labels; the latter deliberately +differs from a looser status-only 743-negative mask because three rows lack a +reported upstream screen and therefore cannot support a measured negative. + +A weighted QRF is fit to the immutable, SHA-pinned full 2023 SIPP donor. It +predicts once for each stable ``person_source_id`` in the PolicyEngine age +3--5 domain and fans that decision to all support clones. Existing output is +used only for an exact idempotence check; a stale or default surface is healed. +""" + +from __future__ import annotations + +from importlib.resources import files +from pathlib import Path +from typing import Any, BinaryIO + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import SourceStageSpec, load_source_manifest +from populace.build.us_runtime.voluntary_filing import ( + SIPP_2023_VOLUNTARY_FILING_DONOR_REVISION, + SIPP_2023_VOLUNTARY_FILING_DONOR_SHA256, + SIPP_2023_VOLUNTARY_FILING_DONOR_SIZE_BYTES, + SIPP_2023_VOLUNTARY_FILING_DONOR_URL, + fetch_sipp_2023_voluntary_filing_donor, +) +from populace.frame import Frame +from populace.frame.units import US_SCHEMA + +__all__ = [ + "HEAD_START_SIPP_DICTIONARY_URL", + "SIPP_2023_HEAD_START_DONOR_REVISION", + "SIPP_2023_HEAD_START_DONOR_SHA256", + "SIPP_2023_HEAD_START_DONOR_SIZE_BYTES", + "SIPP_2023_HEAD_START_DONOR_URL", + "SIPP_HEAD_START_FIT_PARAMETERS", + "SIPP_HEAD_START_MODEL_PREDICTORS", + "SIPP_HEAD_START_READ_PARAMETERS", + "SIPP_HEAD_START_SOURCE_COLUMNS", + "US_SIPP_HEAD_START_NONCONSTANT_PERSON_COLUMNS", + "US_SIPP_HEAD_START_OUTPUT_COLUMNS", + "US_SIPP_HEAD_START_REQUIRED_SOURCE_COLUMNS", + "US_SIPP_HEAD_START_STAGE_NAME", + "fetch_sipp_2023_head_start_donor", + "impute_us_sipp_head_start", + "load_sipp_2023_head_start_donor", + "us_sipp_head_start_signal_gate", + "us_sipp_head_start_stage_spec", + "us_sipp_head_start_summary", + "with_us_sipp_head_start_input", +] + +QRF: Any | None = None + +HEAD_START_SIPP_DICTIONARY_URL = ( + "https://www2.census.gov/programs-surveys/sipp/tech-documentation/" + "data-dictionaries/2023/2023_SIPP_Data_Dictionary.pdf" +) + +# This family deliberately aliases the already reviewed full-file coordinate +# instead of introducing another mutable donor locator. +SIPP_2023_HEAD_START_DONOR_REVISION = SIPP_2023_VOLUNTARY_FILING_DONOR_REVISION +SIPP_2023_HEAD_START_DONOR_SHA256 = SIPP_2023_VOLUNTARY_FILING_DONOR_SHA256 +SIPP_2023_HEAD_START_DONOR_SIZE_BYTES = SIPP_2023_VOLUNTARY_FILING_DONOR_SIZE_BYTES +SIPP_2023_HEAD_START_DONOR_URL = SIPP_2023_VOLUNTARY_FILING_DONOR_URL + +US_SIPP_HEAD_START_STAGE_NAME = "sipp_head_start" +US_SIPP_HEAD_START_OUTPUT_COLUMNS: tuple[str, ...] = ( + "takes_up_head_start_if_eligible", +) +US_SIPP_HEAD_START_NONCONSTANT_PERSON_COLUMNS = US_SIPP_HEAD_START_OUTPUT_COLUMNS +US_SIPP_HEAD_START_REQUIRED_SOURCE_COLUMNS: tuple[str, ...] = ( + "person_source_id", + "person_household_id", + "age", + "is_female", + "employment_income_before_lsr", +) + +_OUTPUT = US_SIPP_HEAD_START_OUTPUT_COLUMNS[0] +_DONOR_WEIGHT_COLUMN = "sipp_weight" +_PERSON_SOURCE_ID_COLUMN = "person_source_id" +_PERSON_SUPPORT_CHANNEL_COLUMN = "person_support_channel" +_ASEC_CHANNEL = "asec" +_PUF_CHANNEL = "puf_tax_detail" +_DEFAULT_N_ESTIMATORS = 100 +_ELIGIBLE_MIN_AGE = 3 +_ELIGIBLE_MAX_AGE = 5 +_ELIGIBLE_TAKE_UP_SHARE_BAND = (0.005, 0.25) + +_SIPP_JOB_EARNINGS_COLUMNS = tuple(f"TJB{i}_MSUM" for i in range(1, 8)) +SIPP_HEAD_START_SOURCE_COLUMNS: tuple[str, ...] = ( + "SSUID", + "PNUM", + "MONTHCODE", + "WPFINWGT", + "TAGE", + "ESEX", + "EED_SCRNR", + "AED_SCRNR", + "EEDGRADE", + "AEDGRADE", + "EEDEMONTH", + "AEDMONTH", + "EEDHEADST", + "AEDHEADST", + *_SIPP_JOB_EARNINGS_COLUMNS, +) +SIPP_HEAD_START_MODEL_PREDICTORS: tuple[str, ...] = ( + "age", + "is_female", + "household_size", + "count_under_18", + "count_under_6", + "household_employment_income", +) + +SIPP_HEAD_START_READ_PARAMETERS: dict[str, object] = { + "table": "sipp_person", + "delimiter": "|", + "month_column": "MONTHCODE", + "month": 12, + "source_columns": list(SIPP_HEAD_START_SOURCE_COLUMNS), +} +SIPP_HEAD_START_FIT_PARAMETERS: dict[str, object] = { + "predictors": list(SIPP_HEAD_START_MODEL_PREDICTORS), + "target": _OUTPUT, + "weight": _DONOR_WEIGHT_COLUMN, + "age_domain": [_ELIGIBLE_MIN_AGE, _ELIGIBLE_MAX_AGE], + "direct_response_filter": "AEDHEADST == 1 and EEDHEADST in [1, 2]", + "structural_no_filters": [ + "AEDHEADST == 0 and AED_SCRNR == 1 and EED_SCRNR == 2", + ( + "AEDHEADST == 0 and AED_SCRNR == 1 and EED_SCRNR == 1 " + "and AEDGRADE == 1 and EEDGRADE != 21" + ), + ], + "excluded_statuses": "AEDHEADST == 2 or AED_SCRNR == 4", + "n_estimators": _DEFAULT_N_ESTIMATORS, + "seed_from_build_config": True, + "assignment_unit": _PERSON_SOURCE_ID_COLUMN, + "fan_to_support_clones": True, +} + +# Exact audit of the immutable full 2023 artifact under the transform above. +_PINNED_RAW_ROWS = 476_744 +_PINNED_DECEMBER_ROWS = 39_513 +_PINNED_AGE_DOMAIN_ROWS = 1_177 +_PINNED_TRAINING_ROWS = 785 +_PINNED_POSITIVE_ROWS = 45 +_PINNED_NEGATIVE_ROWS = 740 +_PINNED_DIRECT_RESPONSE_ROWS = 215 +_PINNED_REPORTED_NO_ENROLLMENT_ROWS = 440 +_PINNED_REPORTED_OTHER_GRADE_ROWS = 130 +_PINNED_WEIGHT_SUM = 7_978_494.5412483 +_PINNED_POSITIVE_WEIGHT_SUM = 491_970.1041311 +_PINNED_WEIGHTED_TRUE_SHARE = 0.06166202177461505 + + +def us_sipp_head_start_stage_spec() -> SourceStageSpec: + """Load and validate the packaged Head Start stage declaration.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + stage_map = manifest.stage_map() + if US_SIPP_HEAD_START_STAGE_NAME not in stage_map: + raise ValueError( + f"US source manifest declares no {US_SIPP_HEAD_START_STAGE_NAME!r} stage." + ) + spec = stage_map[US_SIPP_HEAD_START_STAGE_NAME] + if spec.grain != "person": + raise ValueError("US SIPP Head Start stage must have person grain.") + if tuple(spec.outputs) != US_SIPP_HEAD_START_OUTPUT_COLUMNS: + raise ValueError( + "US SIPP Head Start manifest outputs do not match the runtime-owned family." + ) + if [operation.kind for operation in spec.operations] != [ + "read_table", + "fit_weighted_qrf", + ]: + raise ValueError( + "US SIPP Head Start stage must contain read_table then fit_weighted_qrf." + ) + if dict(spec.operations[0].parameters) != SIPP_HEAD_START_READ_PARAMETERS: + raise ValueError( + "US SIPP Head Start read contract drifted from the pinned transform." + ) + if dict(spec.operations[1].parameters) != SIPP_HEAD_START_FIT_PARAMETERS: + raise ValueError( + "US SIPP Head Start fit contract drifted from the reviewed labels." + ) + pinned = [ + artifact + for artifact in spec.artifacts + if artifact.get("sha256") == SIPP_2023_HEAD_START_DONOR_SHA256 + and artifact.get("size_bytes") == SIPP_2023_HEAD_START_DONOR_SIZE_BYTES + ] + if not pinned: + raise ValueError( + "US SIPP Head Start stage does not pin the full donor SHA-256 and " + "byte length." + ) + return spec + + +def _sha256_stream(stream: BinaryIO, *, chunk_size: int = 8 * 1024 * 1024) -> str: + import hashlib + + digest = hashlib.sha256() + for chunk in iter(lambda: stream.read(chunk_size), b""): + digest.update(chunk) + return digest.hexdigest() + + +def _sha256_file(path: Path) -> str: + with path.open("rb") as stream: + return _sha256_stream(stream) + + +def fetch_sipp_2023_head_start_donor( + cache_dir: str | Path | None = None, + *, + expected_sha256: str | None = SIPP_2023_HEAD_START_DONOR_SHA256, + expected_size_bytes: int | None = SIPP_2023_HEAD_START_DONOR_SIZE_BYTES, + chunk_size: int = 8 * 1024 * 1024, +) -> Path: + """Fetch the already-pinned full SIPP donor through its shared cache.""" + + return fetch_sipp_2023_voluntary_filing_donor( + cache_dir, + expected_sha256=expected_sha256, + expected_size_bytes=expected_size_bytes, + chunk_size=chunk_size, + ) + + +def _numeric(source: pd.DataFrame, column: str) -> np.ndarray: + return pd.to_numeric(source[column], errors="coerce").to_numpy(dtype=np.float64) + + +def _require_finite( + values: np.ndarray, + *, + label: str, + minimum: float | None = None, + maximum: float | None = None, +) -> np.ndarray: + valid = np.isfinite(values) + if minimum is not None: + valid &= values >= minimum + if maximum is not None: + valid &= values <= maximum + if not valid.all(): + rows = np.flatnonzero(~valid)[:5].tolist() + raise ValueError(f"SIPP Head Start {label} is invalid at row(s): {rows}.") + return values + + +def _assert_pinned_audit(audit: dict[str, int | float]) -> None: + expected_counts = { + "raw_rows": _PINNED_RAW_ROWS, + "december_rows": _PINNED_DECEMBER_ROWS, + "age_domain_rows": _PINNED_AGE_DOMAIN_ROWS, + "training_rows": _PINNED_TRAINING_ROWS, + "positive_rows": _PINNED_POSITIVE_ROWS, + "negative_rows": _PINNED_NEGATIVE_ROWS, + "direct_response_rows": _PINNED_DIRECT_RESPONSE_ROWS, + "reported_no_enrollment_rows": _PINNED_REPORTED_NO_ENROLLMENT_ROWS, + "reported_other_grade_rows": _PINNED_REPORTED_OTHER_GRADE_ROWS, + } + for key, expected in expected_counts.items(): + if int(audit[key]) != expected: + raise ValueError( + f"Pinned SIPP Head Start audit drifted for {key}: " + f"expected {expected}, got {audit[key]}." + ) + expected_floats = { + "weight_sum": _PINNED_WEIGHT_SUM, + "positive_weight_sum": _PINNED_POSITIVE_WEIGHT_SUM, + "weighted_true_share": _PINNED_WEIGHTED_TRUE_SHARE, + } + for key, expected in expected_floats.items(): + if not np.isclose(float(audit[key]), expected, rtol=0.0, atol=1e-6): + raise ValueError( + f"Pinned SIPP Head Start audit drifted for {key}: " + f"expected {expected}, got {audit[key]}." + ) + + +def load_sipp_2023_head_start_donor( + path: str | Path, + *, + expected_sha256: str | None = SIPP_2023_HEAD_START_DONOR_SHA256, + expected_size_bytes: int | None = SIPP_2023_HEAD_START_DONOR_SIZE_BYTES, + chunksize: int = 100_000, +) -> pd.DataFrame: + """Load strict December age-3--5 labels from the pinned full SIPP file.""" + + source_path = Path(path).expanduser() + if not source_path.is_file(): + raise FileNotFoundError(source_path) + if chunksize < 1: + raise ValueError("chunksize must be positive") + if ( + expected_size_bytes is not None + and source_path.stat().st_size != expected_size_bytes + ): + raise ValueError( + "SIPP Head Start donor byte length does not match the pinned artifact." + ) + if expected_sha256 is not None: + digest = _sha256_file(source_path) + if digest != expected_sha256: + raise ValueError( + "SIPP Head Start donor SHA-256 does not match the pinned artifact." + ) + + header = pd.read_csv(source_path, sep="|", nrows=0) + missing = sorted(set(SIPP_HEAD_START_SOURCE_COLUMNS) - set(header.columns)) + if missing: + raise ValueError(f"SIPP Head Start donor missing source column(s): {missing}.") + + december_parts: list[pd.DataFrame] = [] + raw_rows = 0 + for chunk in pd.read_csv( + source_path, + sep="|", + usecols=list(SIPP_HEAD_START_SOURCE_COLUMNS), + chunksize=chunksize, + low_memory=False, + ): + raw_rows += len(chunk) + month = pd.to_numeric(chunk["MONTHCODE"], errors="coerce") + december_parts.append(chunk.loc[month.eq(12)].copy()) + december = pd.concat(december_parts, ignore_index=True) + if december.empty: + raise ValueError("SIPP Head Start donor contains no December rows.") + if december[["SSUID", "PNUM"]].isna().any(axis=None): + raise ValueError("SIPP Head Start December identity contains missing values.") + if december.duplicated(["SSUID", "PNUM"]).any(): + raise ValueError("SIPP Head Start December identity is not person-unique.") + + age = _require_finite( + _numeric(december, "TAGE"), label="age", minimum=0.0, maximum=120.0 + ) + sex = _require_finite( + _numeric(december, "ESEX"), label="sex code", minimum=1.0, maximum=2.0 + ) + if not np.isin(sex, [1.0, 2.0]).all(): + raise ValueError("SIPP Head Start ESEX must contain only 1 or 2.") + weight = _require_finite( + _numeric(december, "WPFINWGT"), label="weight", minimum=0.0 + ) + + jobs = december.loc[:, list(_SIPP_JOB_EARNINGS_COLUMNS)].apply( + pd.to_numeric, errors="coerce" + ) + # Blank job slots are structural zeroes. Negative earnings are not a + # valid value for the SIPP monthly job-earnings recodes. + jobs = jobs.fillna(0.0) + if (jobs.to_numpy(dtype=np.float64) < 0.0).any(): + raise ValueError("SIPP Head Start job earnings must be nonnegative.") + monthly_person_earnings = jobs.sum(axis=1).to_numpy(dtype=np.float64) + + features = pd.DataFrame(index=december.index) + features["age"] = age + features["is_female"] = sex == 2.0 + household = december["SSUID"] + features["household_size"] = household.groupby(household, sort=False).transform( + "size" + ) + features["count_under_18"] = ( + pd.Series(age < 18.0, index=december.index) + .groupby(household, sort=False) + .transform("sum") + ) + features["count_under_6"] = ( + pd.Series(age < 6.0, index=december.index) + .groupby(household, sort=False) + .transform("sum") + ) + features["household_employment_income"] = ( + pd.Series(monthly_person_earnings * 12.0, index=december.index) + .groupby(household, sort=False) + .transform("sum") + ) + + in_age_domain = (age >= _ELIGIBLE_MIN_AGE) & (age <= _ELIGIBLE_MAX_AGE) + direct_status = _numeric(december, "AEDHEADST") + direct_answer = _numeric(december, "EEDHEADST") + screen_status = _numeric(december, "AED_SCRNR") + screen = _numeric(december, "EED_SCRNR") + grade_status = _numeric(december, "AEDGRADE") + grade = _numeric(december, "EEDGRADE") + + reported_direct = ( + in_age_domain & (direct_status == 1.0) & np.isin(direct_answer, [1.0, 2.0]) + ) + reported_no_enrollment = ( + in_age_domain + & (direct_status == 0.0) + & (screen_status == 1.0) + & (screen == 2.0) + ) + reported_other_grade = ( + in_age_domain + & (direct_status == 0.0) + & (screen_status == 1.0) + & (screen == 1.0) + & (grade_status == 1.0) + & np.isfinite(grade) + & (grade != 21.0) + ) + training_mask = reported_direct | reported_no_enrollment | reported_other_grade + positive = reported_direct & (direct_answer == 1.0) + + donor = features.loc[training_mask, list(SIPP_HEAD_START_MODEL_PREDICTORS)].copy() + donor[_OUTPUT] = positive[training_mask] + donor[_DONOR_WEIGHT_COLUMN] = weight[training_mask] + donor = donor.reset_index(drop=True) + if (donor[_DONOR_WEIGHT_COLUMN] <= 0.0).any(): + raise ValueError("SIPP Head Start training weights must be positive.") + if donor.empty or donor[_OUTPUT].nunique() != 2: + raise ValueError("SIPP Head Start donor must contain both observed classes.") + + positive_weight = float(donor.loc[donor[_OUTPUT], _DONOR_WEIGHT_COLUMN].sum()) + weight_sum = float(donor[_DONOR_WEIGHT_COLUMN].sum()) + audit: dict[str, int | float] = { + "raw_rows": raw_rows, + "december_rows": len(december), + "age_domain_rows": int(np.count_nonzero(in_age_domain)), + "training_rows": len(donor), + "positive_rows": int(donor[_OUTPUT].sum()), + "negative_rows": int((~donor[_OUTPUT]).sum()), + "direct_response_rows": int(np.count_nonzero(reported_direct)), + "reported_no_enrollment_rows": int(np.count_nonzero(reported_no_enrollment)), + "reported_other_grade_rows": int(np.count_nonzero(reported_other_grade)), + "weight_sum": weight_sum, + "positive_weight_sum": positive_weight, + "weighted_true_share": positive_weight / weight_sum, + } + pinned_transform = bool( + expected_sha256 == SIPP_2023_HEAD_START_DONOR_SHA256 + and expected_size_bytes == SIPP_2023_HEAD_START_DONOR_SIZE_BYTES + ) + if pinned_transform: + _assert_pinned_audit(audit) + audit["pinned_transform"] = pinned_transform + donor.attrs["source_audit"] = audit + return donor + + +def _decoded_strings(series: pd.Series) -> pd.Series: + return series.map( + lambda value: value.decode() if isinstance(value, (bytes, bytearray)) else value + ).astype(str) + + +def _recipient_predictors(frame: Frame) -> tuple[pd.DataFrame, pd.Series, np.ndarray]: + person = frame.table("person") + missing = [ + column + for column in US_SIPP_HEAD_START_REQUIRED_SOURCE_COLUMNS + if column not in person + ] + if missing: + raise ValueError( + f"US SIPP Head Start receiver missing person column(s): {missing}." + ) + if person[_PERSON_SOURCE_ID_COLUMN].isna().any(): + raise ValueError("US SIPP Head Start requires complete person_source_id.") + + age = pd.to_numeric(person["age"], errors="coerce").to_numpy(dtype=np.float64) + female_raw = pd.to_numeric(person["is_female"], errors="coerce").to_numpy( + dtype=np.float64 + ) + earnings = pd.to_numeric( + person["employment_income_before_lsr"], errors="coerce" + ).to_numpy(dtype=np.float64) + valid = ( + np.isfinite(age) + & (age >= 0.0) + & (age <= 120.0) + & np.isfinite(female_raw) + & np.isin(female_raw, [0.0, 1.0]) + & np.isfinite(earnings) + & (earnings >= 0.0) + ) + if not valid.all(): + rows = np.flatnonzero(~valid)[:5].tolist() + raise ValueError( + "US SIPP Head Start receiver age, sex, and earnings must be valid; " + f"invalid row(s): {rows}." + ) + if person["person_household_id"].isna().any(): + raise ValueError("US SIPP Head Start requires complete household linkage.") + + household = person["person_household_id"] + features = pd.DataFrame(index=person.index) + features["age"] = age + features["is_female"] = female_raw + features["household_size"] = household.groupby(household, sort=False).transform( + "size" + ) + features["count_under_18"] = ( + pd.Series(age < 18.0, index=person.index) + .groupby(household, sort=False) + .transform("sum") + ) + features["count_under_6"] = ( + pd.Series(age < 6.0, index=person.index) + .groupby(household, sort=False) + .transform("sum") + ) + features["household_employment_income"] = ( + pd.Series(earnings, index=person.index) + .groupby(household, sort=False) + .transform("sum") + ) + + source_id = _decoded_strings(person[_PERSON_SOURCE_ID_COLUMN]) + age_unique = ( + pd.Series(age, index=person.index).groupby(source_id, sort=False).nunique() + ) + inconsistent = age_unique.index[age_unique > 1].tolist() + if inconsistent: + raise ValueError( + "US SIPP Head Start source clones disagree on age for " + f"person_source_id(s): {inconsistent[:5]}." + ) + + order = pd.DataFrame(index=person.index) + order["source_id"] = source_id + if _PERSON_SUPPORT_CHANNEL_COLUMN in person: + if person[_PERSON_SUPPORT_CHANNEL_COLUMN].isna().any(): + raise ValueError("US SIPP Head Start support channel is missing.") + channels = _decoded_strings(person[_PERSON_SUPPORT_CHANNEL_COLUMN]) + unexpected = sorted(set(channels) - {_ASEC_CHANNEL, _PUF_CHANNEL}) + if unexpected: + raise ValueError( + "US SIPP Head Start found unsupported support channel(s): " + f"{unexpected}." + ) + order["channel_priority"] = channels.map({_ASEC_CHANNEL: 0, _PUF_CHANNEL: 1}) + else: + order["channel_priority"] = 0 + order["person_key"] = person["person_id"].astype(str) + canonical_index = ( + order.sort_values( + ["source_id", "channel_priority", "person_key"], kind="mergesort" + ) + .drop_duplicates("source_id", keep="first") + .index + ) + canonical = features.loc[canonical_index].copy() + canonical.insert(0, "source_id", source_id.loc[canonical_index].to_numpy()) + canonical = canonical.sort_values("source_id", kind="mergesort").reset_index( + drop=True + ) + values = canonical.loc[:, list(SIPP_HEAD_START_MODEL_PREDICTORS)].to_numpy( + dtype=np.float64 + ) + if not np.isfinite(values).all(): + raise ValueError("US SIPP Head Start receiver predictors must be finite.") + eligible = ( + canonical["age"] + .between(_ELIGIBLE_MIN_AGE, _ELIGIBLE_MAX_AGE, inclusive="both") + .to_numpy() + ) + return canonical, source_id, eligible + + +def _coerce_boolean_prediction(values: pd.Series | np.ndarray) -> np.ndarray: + series = pd.Series(values) + if series.dtype == bool: + return series.to_numpy(dtype=bool) + numeric = pd.to_numeric(series, errors="coerce").to_numpy(dtype=np.float64) + if not (np.isfinite(numeric) & (numeric >= 0.0) & (numeric <= 1.0)).all(): + raise ValueError("SIPP Head Start QRF produced invalid boolean predictions.") + return numeric >= 0.5 + + +def impute_us_sipp_head_start( + frame: Frame, + donor: pd.DataFrame, + *, + seed: int, + n_estimators: int = _DEFAULT_N_ESTIMATORS, +) -> pd.Series: + """Fit the measured SIPP model and predict once per stable source person.""" + + if frame.schema != US_SCHEMA: + raise ValueError("US SIPP Head Start imputation requires the US schema.") + if n_estimators < 1: + raise ValueError("n_estimators must be positive") + required = [ + *SIPP_HEAD_START_MODEL_PREDICTORS, + _OUTPUT, + _DONOR_WEIGHT_COLUMN, + ] + missing = [column for column in required if column not in donor] + if missing: + raise ValueError(f"SIPP Head Start donor missing column(s): {missing}.") + training = donor.loc[:, required].copy() + for column in [*SIPP_HEAD_START_MODEL_PREDICTORS, _DONOR_WEIGHT_COLUMN]: + training[column] = pd.to_numeric(training[column], errors="coerce") + numeric = training.loc[ + :, [*SIPP_HEAD_START_MODEL_PREDICTORS, _DONOR_WEIGHT_COLUMN] + ].to_numpy(dtype=np.float64) + if not np.isfinite(numeric).all() or np.any( + training[_DONOR_WEIGHT_COLUMN].to_numpy(dtype=np.float64) <= 0.0 + ): + raise ValueError( + "SIPP Head Start donor predictors and weights must be finite, with " + "positive weights." + ) + target = pd.to_numeric(training[_OUTPUT], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not (np.isfinite(target) & np.isin(target, [0.0, 1.0])).all(): + raise ValueError("SIPP Head Start donor target must be boolean.") + if np.unique(target).size != 2: + raise ValueError("SIPP Head Start donor target must contain both classes.") + training[_OUTPUT] = target.astype(bool) + + canonical, source_id, eligible = _recipient_predictors(frame) + global QRF + if QRF is None: + from importlib import import_module + + QRF = import_module("populace.fit").QRF + fitted = QRF(n_estimators=int(n_estimators), seed=int(seed)).fit( + training, + predictors=list(SIPP_HEAD_START_MODEL_PREDICTORS), + targets=[_OUTPUT], + weights=_DONOR_WEIGHT_COLUMN, + ) + decisions = np.zeros(len(canonical), dtype=bool) + if eligible.any(): + predicted = fitted.predict( + canonical.loc[eligible, list(SIPP_HEAD_START_MODEL_PREDICTORS)] + ) + if _OUTPUT not in predicted: + raise ValueError(f"SIPP Head Start QRF prediction missing {_OUTPUT!r}.") + decisions[eligible] = _coerce_boolean_prediction(predicted[_OUTPUT]) + decision_by_source = pd.Series(decisions, index=canonical["source_id"]) + result = source_id.map(decision_by_source) + if result.isna().any(): + raise ValueError( + "SIPP Head Start prediction did not cover every source person." + ) + return pd.Series( + result.to_numpy(dtype=bool), index=frame.table("person").index, name=_OUTPUT + ) + + +def with_us_sipp_head_start_input( + frame: Frame, + *, + seed: int, + time_period: int, + sipp_donor: pd.DataFrame, + n_estimators: int = _DEFAULT_N_ESTIMATORS, +) -> Frame: + """Recompute and materialize measured Head Start take-up on a US frame.""" + + if frame.schema != US_SCHEMA: + raise ValueError("US SIPP Head Start input requires the US schema.") + del time_period # The immutable donor fixes the observation vintage. + us_sipp_head_start_stage_spec() + predicted = impute_us_sipp_head_start( + frame, + sipp_donor, + seed=int(seed), + n_estimators=int(n_estimators), + ) + person = frame.table("person") + if _OUTPUT in person: + current = pd.to_numeric(person[_OUTPUT], errors="coerce").to_numpy( + dtype=np.float64 + ) + if ( + np.isfinite(current).all() + and np.isin(current, [0.0, 1.0]).all() + and np.array_equal(current.astype(bool), predicted.to_numpy()) + ): + return frame + + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + tables["person"][_OUTPUT] = predicted.to_numpy(dtype=bool) + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + ) + + +def us_sipp_head_start_summary(frame: Frame) -> dict[str, object]: + """Return age-domain incidence, provenance, and clone diagnostics.""" + + person = frame.table("person") + if _OUTPUT not in person: + raise ValueError(f"US SIPP Head Start summary requires {_OUTPUT!r}.") + raw = pd.to_numeric(person[_OUTPUT], errors="coerce").to_numpy(dtype=np.float64) + finite = np.isfinite(raw) + boolean = finite & np.isin(raw, [0.0, 1.0]) + values = boolean & (raw == 1.0) + age = pd.to_numeric(person.get("age"), errors="coerce").to_numpy(dtype=np.float64) + weights = np.asarray(frame.resolve_weights("person").values, dtype=np.float64) + if not ( + np.isfinite(age).all() + and np.isfinite(weights).all() + and (weights >= 0.0).all() + and weights.sum() > 0.0 + ): + raise ValueError("US SIPP Head Start summary requires valid age and weights.") + eligible = (age >= _ELIGIBLE_MIN_AGE) & (age <= _ELIGIBLE_MAX_AGE) + eligible_weight = float(weights[eligible].sum()) + + provenance_missing = bool( + _PERSON_SOURCE_ID_COLUMN not in person + or person.get(_PERSON_SOURCE_ID_COLUMN, pd.Series(dtype=object)).isna().any() + ) + channel_invalid = False + if _PERSON_SUPPORT_CHANNEL_COLUMN in person: + channel_source = person[_PERSON_SUPPORT_CHANNEL_COLUMN] + channel_invalid = bool( + channel_source.isna().any() + or not set(_decoded_strings(channel_source)).issubset( + {_ASEC_CHANNEL, _PUF_CHANNEL} + ) + ) + clone_groups = clone_mismatches = 0 + if not provenance_missing: + keys = _decoded_strings(person[_PERSON_SOURCE_ID_COLUMN]) + work = pd.DataFrame({"key": keys, "value": values}) + sizes = work.groupby("key", sort=False).size() + clone_groups = int((sizes > 1).sum()) + clone_mismatches = int( + (work.groupby("key", sort=False)["value"].nunique() > 1).sum() + ) + + summary: dict[str, object] = { + "missing_count": int(np.count_nonzero(~finite)), + "invalid_count": int(np.count_nonzero(finite & ~boolean)), + "unique_count": int(np.unique(raw[boolean]).size), + "positive_count": int(np.count_nonzero(values)), + "eligible_count": int(np.count_nonzero(eligible)), + "eligible_weight": eligible_weight, + "eligible_weighted_take_up_share": ( + float(weights[eligible & values].sum()) / eligible_weight + if eligible_weight > 0.0 + else 0.0 + ), + "eligible_weighted_take_up_share_band": list(_ELIGIBLE_TAKE_UP_SHARE_BAND), + "out_of_domain_positive_count": int(np.count_nonzero(~eligible & values)), + "support_provenance_missing": provenance_missing, + "support_channel_invalid": channel_invalid, + "clone_group_count": clone_groups, + "clone_mismatch_count": clone_mismatches, + } + if _PERSON_SUPPORT_CHANNEL_COLUMN in person: + channel = _decoded_strings(person[_PERSON_SUPPORT_CHANNEL_COLUMN]) + channel_shares: dict[str, float] = {} + for name in sorted(channel.unique()): + mask = channel.eq(name).to_numpy() & eligible + denominator = float(weights[mask].sum()) + channel_shares[name] = ( + float(weights[mask & values].sum()) / denominator + if denominator > 0.0 + else 0.0 + ) + summary["channel_eligible_weighted_take_up_shares"] = channel_shares + return summary + + +def us_sipp_head_start_signal_gate(frame: Frame) -> GateResult: + """Require a nondefault, in-domain, source-clone-consistent take-up flag.""" + + person = frame.table("person") + if _OUTPUT not in person: + return GateResult( + name="sipp_head_start_signal", + passed=False, + failures=(f"person.{_OUTPUT}: missing",), + details={"missing": [_OUTPUT]}, + ) + try: + summary = us_sipp_head_start_summary(frame) + except (TypeError, ValueError) as exc: + return GateResult( + name="sipp_head_start_signal", + passed=False, + failures=(str(exc),), + details={}, + ) + failures: list[str] = [] + if int(summary["missing_count"]): + failures.append(f"{_OUTPUT}: missing values") + if int(summary["invalid_count"]): + failures.append(f"{_OUTPUT}: non-boolean values") + if int(summary["unique_count"]) < 2: + failures.append(f"{_OUTPUT}: constant column") + if int(summary["eligible_count"]) == 0 or float(summary["eligible_weight"]) <= 0: + failures.append(f"{_OUTPUT}: no weighted age-3--5 domain") + share = float(summary["eligible_weighted_take_up_share"]) + low, high = _ELIGIBLE_TAKE_UP_SHARE_BAND + if not low <= share <= high: + failures.append( + f"{_OUTPUT}: eligible weighted take-up share {share:.6f} outside " + f"[{low:.3f}, {high:.3f}]" + ) + if int(summary["out_of_domain_positive_count"]): + failures.append(f"{_OUTPUT}: positive values outside the age-3--5 domain") + if bool(summary["support_provenance_missing"]): + failures.append(f"{_OUTPUT}: person_source_id provenance is missing") + if bool(summary["support_channel_invalid"]): + failures.append(f"{_OUTPUT}: support-channel provenance is invalid") + if int(summary["clone_mismatch_count"]): + failures.append( + f"{_OUTPUT}: {summary['clone_mismatch_count']} support-clone mismatch(es)" + ) + return GateResult( + name="sipp_head_start_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) diff --git a/packages/populace-build/src/populace/build/us_runtime/sipp_tips.py b/packages/populace-build/src/populace/build/us_runtime/sipp_tips.py new file mode 100644 index 00000000..c78ead54 --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/sipp_tips.py @@ -0,0 +1,553 @@ +"""SIPP-imputed tip income and CPS tipped-occupation inputs. + +The retired eCPS pipeline did not carry tips from ASEC: ASEC has no tip- +amount field. It trained a weighted QRF on SIPP person records, with annual +tip income equal to the December monthly total across up to seven jobs times +12. The predictors were employment income, age, counts of children under 18 +and under 6 in the household, and a tipped-occupation indicator. The latter +was derived on both donor and recipient records by mapping detailed Census +occupation codes to the Treasury tipped-occupation list. + +This is a direct port of that method, pinned to the retired implementation at +``9a823603e6b5fb916d65ec45d74c9c7eb0043db1`` +(``datasets/sipp/sipp.py`` and ``calibration/source_impute.py``). The final +pipeline switched its live downloader to SIPP 2024 without pinning the Census +zip. Populace instead uses the same pipeline's fixed SIPP 2023 slim extract, +whose immutable mirror revision and content SHA-256 are pinned below. That is +the donor used to build the reference eCPS and avoids accepting a silently +reissued public-use file. + +Two PolicyEngine-US person inputs are produced: + +* ``treasury_tipped_occupation_code`` is a direct CPS carry-through derived + from raw ASEC ``PEIOOCC``. +* ``tip_income`` is a non-negative SIPP QRF draw. Treasury-listed occupation + is a predictor, not a domain mask: the pinned reference eCPS carries positive + tips for some people outside the list, and the OBBBA formula applies its own + occupation-qualification test downstream. + +Healing behavior matches the other source stages: an existing pair of columns +with signal is passed through untouched; missing or constant-default columns +are rebuilt from the raw ASEC occupation and the pinned donor. +""" + +from __future__ import annotations + +from importlib.resources import files +from pathlib import Path + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import SourceStageSpec, load_source_manifest +from populace.frame import Frame +from populace.frame.units import US_SCHEMA + +__all__ = [ + "CENSUS_OCCUPATION_CODE_TO_TTOC", + "SIPP_2023_TIP_DONOR_REVISION", + "SIPP_2023_TIP_DONOR_SHA256", + "SIPP_2023_TIP_DONOR_URL", + "SIPP_TIP_OUTPUT_COLUMNS", + "SIPP_TIP_PREDICTORS", + "US_SIPP_TIPS_NONCONSTANT_PERSON_COLUMNS", + "US_SIPP_TIPS_OUTPUT_COLUMNS", + "US_SIPP_TIPS_REQUIRED_SOURCE_COLUMNS", + "US_SIPP_TIPS_STAGE_NAME", + "derive_treasury_tipped_occupation_code", + "fetch_sipp_2023_tip_donor", + "impute_us_sipp_tips", + "load_sipp_2023_tip_donor", + "us_sipp_tips_signal_gate", + "us_sipp_tips_stage_spec", + "us_sipp_tips_summary", + "with_us_sipp_tip_inputs", +] + +US_SIPP_TIPS_STAGE_NAME = "sipp_tips" + +# The immutable mirror revision containing the retired pipeline's SIPP 2023 +# slim donor. The repository name is assembled so the live-tree guard does not +# mistake this historical input coordinate for a runtime package dependency. +SIPP_2023_TIP_DONOR_REVISION = "21280dca5995e978d706740a8a4b9b7860cfd7b6" +_RETIRED_DATA_REPOSITORY = "policyengine-" + "us-data" +SIPP_2023_TIP_DONOR_URL = ( + "https://huggingface.co/policyengine/" + f"{_RETIRED_DATA_REPOSITORY}/resolve/" + f"{SIPP_2023_TIP_DONOR_REVISION}/pu2023_slim.csv" +) +SIPP_2023_TIP_DONOR_SHA256 = ( + "1f0bcb8e045ef1118e8eba4b4a2997bdaaf947bd0dd09d41fa7c7d5657a3d7d5" +) +_SIPP_2023_TIP_DONOR_FILENAME = "pu2023_slim.csv" + +US_SIPP_TIPS_OUTPUT_COLUMNS: tuple[str, ...] = ( + "tip_income", + "treasury_tipped_occupation_code", +) +# Short compatibility alias used by stage-local callers; the ``US_`` name is +# retained for consistency with the release-export registries. +SIPP_TIP_OUTPUT_COLUMNS = US_SIPP_TIPS_OUTPUT_COLUMNS +US_SIPP_TIPS_NONCONSTANT_PERSON_COLUMNS = US_SIPP_TIPS_OUTPUT_COLUMNS + +SIPP_TIP_PREDICTORS: tuple[str, ...] = ( + "employment_income", + "age", + "count_under_18", + "count_under_6", + "is_tipped_occupation", +) + +US_SIPP_TIPS_REQUIRED_SOURCE_COLUMNS: tuple[str, ...] = ( + "person_household_id", + "employment_income_before_lsr", + "age", + "PEIOOCC", +) + +_SIPP_JOB_OCCUPATION_COLUMNS = tuple(f"TJB{i}_OCC" for i in range(1, 8)) +_SIPP_TIP_AMOUNT_COLUMNS = tuple(f"TJB{i}_TXAMT" for i in range(1, 8)) +_SIPP_TIP_ALLOCATION_COLUMNS = tuple(f"AJB{i}_TXAMT" for i in range(1, 8)) +_SIPP_OBSERVED_STATUS_VALUES = frozenset((0, 1, 9)) +_DONOR_WEIGHT_COLUMN = "sipp_weight" +_MAX_TRAIN_SAMPLES = 10_000 +_DEFAULT_N_ESTIMATORS = 100 +# Stable seed produced by the retired pipeline's ``seeded_rng`` for +# ``calibration_sipp_tip_training_sample:tip_income``. Keeping the named-source +# seed matters on this sparse target: seed 0 materially under-samples positive +# tip rows in the 10,000-row cap. +_TIP_TRAINING_SAMPLE_SEED = 5_559_651_045_748_063_828 + +# Weighted all-person plausibility bands. The pinned reference eCPS has 7.09% +# in a listed occupation and 0.79% with positive tip income. These deliberately +# broad bands catch a zero/constant/wrong-column surface without point-pinning +# one stochastic QRF draw. +_TIPPED_OCCUPATION_SHARE_BAND = (0.02, 0.15) +_TIP_INCOME_NONZERO_SHARE_BAND = (0.001, 0.03) + +# Retired ``datasets/cps/tipped_occupation.py`` mapping, itself derived from the +# IRS/Treasury 2025-42 tipped-occupation list joined to the Census 2018 +# occupation-to-SOC crosswalk. A few SOC collisions use one representative +# TTOC because PolicyEngine only needs listed (>0) versus unlisted (=0). +CENSUS_OCCUPATION_CODE_TO_TTOC: dict[int, int] = { + 725: 502, + 2350: 507, + 2633: 502, + 2752: 206, + 2755: 207, + 2770: 208, + 2910: 503, + 3602: 501, + 3630: 602, + 4000: 105, + 4010: 106, + 4030: 106, + 4040: 101, + 4055: 107, + 4110: 102, + 4120: 103, + 4130: 104, + 4140: 108, + 4150: 109, + 4160: 106, + 4230: 304, + 4251: 402, + 4350: 506, + 4420: 210, + 4500: 603, + 4510: 603, + 4521: 605, + 4522: 601, + 4600: 508, + 4621: 607, + 4655: 501, + 5130: 203, + 5300: 303, + 6355: 403, + 6442: 404, + 7120: 401, + 7200: 409, + 7315: 405, + 7320: 406, + 7340: 401, + 7540: 408, + 7610: 401, + 7800: 110, + 8510: 401, + 9122: 806, + 9141: 803, + 9142: 802, + 9350: 801, + 9610: 805, + 9620: 809, +} + + +def us_sipp_tips_stage_spec() -> SourceStageSpec: + """Load and validate the packaged ``sipp_tips`` stage declaration.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + stage_map = manifest.stage_map() + if US_SIPP_TIPS_STAGE_NAME not in stage_map: + raise ValueError( + f"US source manifest declares no {US_SIPP_TIPS_STAGE_NAME!r} stage." + ) + spec = stage_map[US_SIPP_TIPS_STAGE_NAME] + missing = sorted(set(US_SIPP_TIPS_OUTPUT_COLUMNS) - set(spec.outputs)) + if missing: + raise ValueError( + f"{US_SIPP_TIPS_STAGE_NAME!r} manifest stage does not declare " + f"output(s) {missing}; the runtime and manifest have drifted." + ) + return spec + + +def _sha256_hexdigest(payload: bytes) -> str: + import hashlib + + return hashlib.sha256(payload).hexdigest() + + +def fetch_sipp_2023_tip_donor( + cache_dir: str | Path | None = None, + *, + expected_sha256: str | None = SIPP_2023_TIP_DONOR_SHA256, +) -> Path: + """Download, SHA-verify, and cache the fixed SIPP 2023 slim donor.""" + + import urllib.request + + root = ( + Path(cache_dir) + if cache_dir is not None + else Path.home() / ".cache" / "populace" / "sipp" + ) + root.mkdir(parents=True, exist_ok=True) + target = root / _SIPP_2023_TIP_DONOR_FILENAME + if target.exists() and target.stat().st_size > 0: + if expected_sha256 is None: + return target + if _sha256_hexdigest(target.read_bytes()) == expected_sha256: + return target + + with urllib.request.urlopen(SIPP_2023_TIP_DONOR_URL) as response: # noqa: S310 + payload = response.read() + if expected_sha256 is not None: + digest = _sha256_hexdigest(payload) + if digest != expected_sha256: + raise ValueError( + "SIPP 2023 tip donor failed sha-256 verification: " + f"expected {expected_sha256}, got {digest}." + ) + target.write_bytes(payload) + return target + + +def derive_treasury_tipped_occupation_code( + census_occupation_codes: pd.Series | np.ndarray, +) -> np.ndarray: + """Map detailed Census occupation codes to Treasury tipped codes.""" + + values = pd.to_numeric( + pd.Series(census_occupation_codes, copy=False), errors="coerce" + ).fillna(-1) + return ( + values.astype(int) + .map(CENSUS_OCCUPATION_CODE_TO_TTOC) + .fillna(0) + .astype(np.int16) + .to_numpy() + ) + + +def _derive_any_tipped_occupation_code(frame: pd.DataFrame) -> np.ndarray: + mapped = [ + derive_treasury_tipped_occupation_code(frame[column]) + for column in _SIPP_JOB_OCCUPATION_COLUMNS + ] + return np.column_stack(mapped).max(axis=1).astype(np.int16) + + +def _sha256_file(path: Path) -> str: + import hashlib + + digest = hashlib.sha256() + with path.open("rb") as stream: + for chunk in iter(lambda: stream.read(1024 * 1024), b""): + digest.update(chunk) + return digest.hexdigest() + + +def load_sipp_2023_tip_donor( + path: str | Path, + *, + expected_sha256: str | None = None, +) -> pd.DataFrame: + """Load the retired SIPP tip donor transformation from its pinned CSV.""" + + path = Path(path) + if expected_sha256 is not None: + digest = _sha256_file(path) + if digest != expected_sha256: + raise ValueError( + "SIPP 2023 tip donor failed sha-256 verification: " + f"expected {expected_sha256}, got {digest}." + ) + + required = { + "SSUID", + "MONTHCODE", + "WPFINWGT", + "TAGE", + "TPTOTINC", + *_SIPP_JOB_OCCUPATION_COLUMNS, + *_SIPP_TIP_AMOUNT_COLUMNS, + *_SIPP_TIP_ALLOCATION_COLUMNS, + } + raw = pd.read_csv(path, usecols=lambda column: column in required) + missing = sorted(required - set(raw.columns)) + if missing: + raise ValueError(f"SIPP 2023 tip donor missing required column(s): {missing}.") + + raw = raw.loc[pd.to_numeric(raw["MONTHCODE"], errors="coerce") == 12].copy() + if raw.empty: + raise ValueError("SIPP 2023 tip donor has no December person records.") + + observed = pd.Series(True, index=raw.index) + for column in _SIPP_TIP_ALLOCATION_COLUMNS: + flag = pd.to_numeric(raw[column], errors="coerce").fillna(0).astype(int) + observed &= flag.isin(_SIPP_OBSERVED_STATUS_VALUES) + if not observed.any(): + raise ValueError("SIPP 2023 tip donor has no observed tip-amount records.") + + # Derive household composition on every December member before applying + # the target-specific observed-source mask. An allocated child is not a tip + # training target, but still counts in a retained adult's household. + donor = pd.DataFrame(index=raw.index) + tip_amounts = raw.loc[:, list(_SIPP_TIP_AMOUNT_COLUMNS)].apply( + pd.to_numeric, errors="coerce" + ) + donor["tip_income"] = np.maximum( + tip_amounts.fillna(0.0).sum(axis=1).to_numpy(dtype=np.float64) * 12.0, + 0.0, + ) + donor["employment_income"] = ( + pd.to_numeric(raw["TPTOTINC"], errors="coerce").fillna(0.0).to_numpy() * 12.0 + ) + donor["age"] = pd.to_numeric(raw["TAGE"], errors="coerce").fillna(0.0) + donor["treasury_tipped_occupation_code"] = _derive_any_tipped_occupation_code(raw) + donor["is_tipped_occupation"] = ( + donor["treasury_tipped_occupation_code"] > 0 + ).astype(np.float64) + household = pd.DataFrame( + { + "SSUID": raw["SSUID"], + "under_18": donor["age"] < 18, + "under_6": donor["age"] < 6, + }, + index=raw.index, + ) + donor["count_under_18"] = household.groupby("SSUID")["under_18"].transform("sum") + donor["count_under_6"] = household.groupby("SSUID")["under_6"].transform("sum") + donor[_DONOR_WEIGHT_COLUMN] = pd.to_numeric( + raw["WPFINWGT"], errors="coerce" + ).fillna(0.0) + finite = np.isfinite(donor.loc[:, [*SIPP_TIP_PREDICTORS, "tip_income"]]).all(axis=1) + positive_weight = donor[_DONOR_WEIGHT_COLUMN] > 0 + return donor.loc[observed & finite & positive_weight].reset_index(drop=True) + + +def _recipient_features(person: pd.DataFrame) -> tuple[pd.DataFrame, np.ndarray]: + missing = [ + column + for column in US_SIPP_TIPS_REQUIRED_SOURCE_COLUMNS + if column not in person.columns + ] + if missing: + raise ValueError( + f"US SIPP-tip imputation requires recipient person column(s): {missing}." + ) + + occupation_code = derive_treasury_tipped_occupation_code(person["PEIOOCC"]) + age = pd.to_numeric(person["age"], errors="coerce").fillna(0.0) + household = pd.DataFrame( + { + "household_id": person["person_household_id"].to_numpy(), + "under_18": age.to_numpy() < 18, + "under_6": age.to_numpy() < 6, + }, + index=person.index, + ) + features = pd.DataFrame(index=person.index) + features["employment_income"] = pd.to_numeric( + person["employment_income_before_lsr"], errors="coerce" + ).fillna(0.0) + features["age"] = age + features["count_under_18"] = household.groupby("household_id")[ + "under_18" + ].transform("sum") + features["count_under_6"] = household.groupby("household_id")["under_6"].transform( + "sum" + ) + features["is_tipped_occupation"] = (occupation_code > 0).astype(np.float64) + return features.loc[:, list(SIPP_TIP_PREDICTORS)], occupation_code + + +def impute_us_sipp_tips( + person: pd.DataFrame, + donor: pd.DataFrame, + *, + seed: int, + n_estimators: int = _DEFAULT_N_ESTIMATORS, +) -> pd.DataFrame: + """QRF-impute tip income and carry the CPS-derived Treasury code.""" + + from populace.fit import QRF + + required = [*SIPP_TIP_PREDICTORS, "tip_income", _DONOR_WEIGHT_COLUMN] + missing = [column for column in required if column not in donor.columns] + if missing: + raise ValueError(f"SIPP tip donor table missing column(s): {missing}.") + fit_frame = donor.loc[:, required].copy() + for column in required: + fit_frame[column] = pd.to_numeric(fit_frame[column], errors="coerce").fillna( + 0.0 + ) + if len(fit_frame) > _MAX_TRAIN_SAMPLES: + selected = np.random.default_rng(_TIP_TRAINING_SAMPLE_SEED).choice( + len(fit_frame), size=_MAX_TRAIN_SAMPLES, replace=False + ) + fit_frame = fit_frame.iloc[np.sort(selected)].reset_index(drop=True) + + fitted = QRF(n_estimators=int(n_estimators), seed=int(seed)).fit( + fit_frame, + predictors=list(SIPP_TIP_PREDICTORS), + targets=["tip_income"], + weights=_DONOR_WEIGHT_COLUMN, + ) + features, occupation_code = _recipient_features(person) + predicted = np.maximum( + np.asarray(fitted.predict(features)["tip_income"], dtype=np.float64), 0.0 + ) + return pd.DataFrame( + { + "tip_income": predicted, + "treasury_tipped_occupation_code": occupation_code, + }, + index=person.index, + ) + + +def _tip_surface_carries_signal(person: pd.DataFrame) -> bool: + return all( + person[column].dropna().nunique() > 1 for column in US_SIPP_TIPS_OUTPUT_COLUMNS + ) + + +def with_us_sipp_tip_inputs( + frame: Frame, + *, + seed: int, + time_period: int, + sipp_donor: pd.DataFrame, +) -> Frame: + """Apply the SIPP tip stage to a US frame before either release arm.""" + + if frame.schema != US_SCHEMA: + raise ValueError("US SIPP-tip inputs require the US schema.") + us_sipp_tips_stage_spec() + person = frame.table("person") + if all(column in person for column in US_SIPP_TIPS_OUTPUT_COLUMNS) and ( + _tip_surface_carries_signal(person) + ): + return frame + + imputed = impute_us_sipp_tips(person, sipp_donor, seed=int(seed)) + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + tables["person"]["tip_income"] = imputed["tip_income"].to_numpy(dtype=np.float64) + tables["person"]["treasury_tipped_occupation_code"] = imputed[ + "treasury_tipped_occupation_code" + ].to_numpy(dtype=np.int16) + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + ) + + +def us_sipp_tips_summary(frame: Frame) -> dict[str, object]: + """Weighted tip-income and tipped-occupation incidence diagnostics.""" + + person = frame.table("person") + weights = np.asarray(frame.resolve_weights("person").values, dtype=np.float64) + total_weight = float(weights.sum()) + + def share(mask: np.ndarray) -> float: + return float(weights[mask].sum()) / total_weight if total_weight > 0 else 0.0 + + tip = pd.to_numeric(person["tip_income"], errors="coerce").fillna(0.0) + occupation = pd.to_numeric( + person["treasury_tipped_occupation_code"], errors="coerce" + ).fillna(0) + return { + "tip_income_nonzero_share": share(tip.to_numpy() > 0), + "tipped_occupation_share": share(occupation.to_numpy() > 0), + "tip_income_nonzero_share_band": list(_TIP_INCOME_NONZERO_SHARE_BAND), + "tipped_occupation_share_band": list(_TIPPED_OCCUPATION_SHARE_BAND), + "unique_counts": { + "tip_income": int(tip.nunique()), + "treasury_tipped_occupation_code": int(occupation.nunique()), + }, + } + + +def us_sipp_tips_signal_gate(frame: Frame) -> GateResult: + """Require both tip-family columns to carry plausible non-default signal.""" + + person = frame.table("person") + missing = [column for column in US_SIPP_TIPS_OUTPUT_COLUMNS if column not in person] + if missing: + return GateResult( + name="sipp_tips_signal", + passed=False, + failures=(f"person columns missing: {missing}.",), + details={"missing": missing}, + ) + + summary = us_sipp_tips_summary(frame) + failures: list[str] = [] + for column, count in summary["unique_counts"].items(): + if count < 2: + failures.append( + f"{column}: constant column (one observed value) — the tip " + "surface carries no signal." + ) + for key, band, label in ( + ( + "tip_income_nonzero_share", + _TIP_INCOME_NONZERO_SHARE_BAND, + "tip-income", + ), + ( + "tipped_occupation_share", + _TIPPED_OCCUPATION_SHARE_BAND, + "tipped-occupation", + ), + ): + value = float(summary[key]) + low, high = band + if not (low <= value <= high): + failures.append( + f"{label} share {value:.3f} outside plausibility band [{low}, {high}]." + ) + return GateResult( + name="sipp_tips_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) diff --git a/packages/populace-build/src/populace/build/us_runtime/sipp_vehicles.py b/packages/populace-build/src/populace/build/us_runtime/sipp_vehicles.py new file mode 100644 index 00000000..13fd792c --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/sipp_vehicles.py @@ -0,0 +1,1035 @@ +"""Household vehicle count and value restored from the full 2023 SIPP. + +The retired enhanced-CPS pipeline trained a household-grain, weighted model +from December SIPP records. Vehicle count was the household maximum of +``TVEH_NUM`` and vehicle value was the first household ``THVAL_VEH``. The +model conditioned on household-summed income, the oldest SIPP member's +demographics, household composition, and homeownership. On the CPS receiver, +the same income concepts were summed to households while demographics came +from the CPS household head. + +This module ports that source layer without folding vehicle value into +``net_worth``. The retired pipeline did that only later, after a stable +SIPP/SCF source draw and a complete SCF balance-sheet reconciliation. Adding +vehicle value to Populace's currently excluded/default net-worth surface would +therefore create a false partial measure. + +The final archived implementation moved to a mutable Census SIPP 2024 URL. +For a hermetic build, this port uses the archived pipeline's full SIPP 2023 +artifact at an immutable Hugging Face revision and verifies both byte length +and SHA-256. The 65.6 MB ``pu2023_slim.csv`` used by the tip stage cannot be +reused: it omits the vehicle targets, allocation flags, asset-income fields, +and home value required here. +""" + +from __future__ import annotations + +from dataclasses import dataclass +from importlib.resources import files +from pathlib import Path +from typing import BinaryIO + +import numpy as np +import pandas as pd +from sklearn.ensemble import RandomForestClassifier + +from populace.build.gates import GateResult +from populace.build.source_manifest import SourceStageSpec, load_source_manifest +from populace.frame import Frame +from populace.frame.units import US_SCHEMA + +__all__ = [ + "ARCHIVED_SIPP_VEHICLE_IMPUTE_URL", + "ARCHIVED_SIPP_VEHICLE_RECEIVER_URL", + "ARCHIVED_SIPP_VEHICLE_SOURCE_URL", + "ARCHIVED_SIPP_VEHICLE_TRANSFORM_URL", + "SIPP_2023_VEHICLE_DONOR_REVISION", + "SIPP_2023_VEHICLE_DONOR_SHA256", + "SIPP_2023_VEHICLE_DONOR_SIZE_BYTES", + "SIPP_2023_VEHICLE_DONOR_URL", + "SIPP_VEHICLE_MODEL_PREDICTORS", + "SIPP_VEHICLE_SOURCE_COLUMNS", + "US_SIPP_VEHICLE_NONCONSTANT_HOUSEHOLD_COLUMNS", + "US_SIPP_VEHICLE_OUTPUT_COLUMNS", + "fetch_sipp_2023_vehicle_donor", + "impute_us_sipp_vehicles", + "load_sipp_2023_vehicle_donor", + "us_sipp_vehicles_signal_gate", + "us_sipp_vehicles_stage_spec", + "us_sipp_vehicles_summary", + "with_us_sipp_vehicle_inputs", +] + +_RETIRED_DATA_REPOSITORY = "policyengine-" + "us-data" +_ARCHIVED_REPOSITORY = ( + "https://github.com/PolicyEngine/" + _RETIRED_DATA_REPOSITORY +) +_ARCHIVED_COMMIT = "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe" +_RETIRED_PACKAGE_PATH = "policyengine_" + "us_data" +ARCHIVED_SIPP_VEHICLE_SOURCE_URL = ( + f"{_ARCHIVED_REPOSITORY}/blob/{_ARCHIVED_COMMIT}/" + f"{_RETIRED_PACKAGE_PATH}/datasets/sipp/sipp.py#L430-L456" +) +ARCHIVED_SIPP_VEHICLE_TRANSFORM_URL = ( + f"{_ARCHIVED_REPOSITORY}/blob/{_ARCHIVED_COMMIT}/" + f"{_RETIRED_PACKAGE_PATH}/datasets/sipp/sipp.py#L953-L1059" +) +ARCHIVED_SIPP_VEHICLE_IMPUTE_URL = ( + f"{_ARCHIVED_REPOSITORY}/blob/{_ARCHIVED_COMMIT}/" + f"{_RETIRED_PACKAGE_PATH}/calibration/source_impute.py#L992-L1084" +) +ARCHIVED_SIPP_VEHICLE_RECEIVER_URL = ( + f"{_ARCHIVED_REPOSITORY}/blob/{_ARCHIVED_COMMIT}/" + f"{_RETIRED_PACKAGE_PATH}/utils/asset_imputation.py#L503-L603" +) + +SIPP_2023_VEHICLE_DONOR_REVISION = "21280dca5995e978d706740a8a4b9b7860cfd7b6" +SIPP_2023_VEHICLE_DONOR_SHA256 = ( + "5c30439e365fc26483318ef61d1d8f4bb2f0e9d6bb47c22c06756a7698733ee2" +) +SIPP_2023_VEHICLE_DONOR_SIZE_BYTES = 3_726_010_471 +SIPP_2023_VEHICLE_DONOR_URL = ( + "https://huggingface.co/policyengine/" + f"{_RETIRED_DATA_REPOSITORY}/resolve/" + f"{SIPP_2023_VEHICLE_DONOR_REVISION}/pu2023.csv" +) + +SIPP_VEHICLE_SOURCE_COLUMNS: tuple[str, ...] = ( + "SSUID", + "PNUM", + "MONTHCODE", + "WPFINWGT", + "TAGE", + "ESEX", + "EMS", + "TPTOTINC", + "TINC_BANK", + "TINC_STMF", + "TINC_BOND", + "TINC_RENT", + "TVEH_NUM", + "THVAL_VEH", + "THVAL_HOME", + "AVEH_NUM", + "AHVAL_VEH", + "AVEH1VAL", + "AVEH2VAL", + "AVEH3VAL", +) + +SIPP_VEHICLE_MODEL_PREDICTORS: tuple[str, ...] = ( + "household_employment_income", + "household_interest_income", + "household_dividend_income", + "household_rental_income", + "reference_age", + "reference_is_female", + "reference_is_married", + "count_under_18", + "household_size", + "is_homeowner", +) + +US_SIPP_VEHICLE_OUTPUT_COLUMNS: tuple[str, ...] = ( + "household_vehicles_owned", + "household_vehicles_value", +) +US_SIPP_VEHICLE_NONCONSTANT_HOUSEHOLD_COLUMNS = US_SIPP_VEHICLE_OUTPUT_COLUMNS + +_STAGE_NAME = "vehicle_assets" +_DONOR_FILENAME = "pu2023.csv" +_DONOR_WEIGHT_COLUMN = "household_weight" +_OWNED_OBSERVED_COLUMN = "household_vehicles_owned_is_observed" +_VALUE_OBSERVED_COLUMN = "household_vehicles_value_is_observed" +_DEFAULT_N_ESTIMATORS = 100 +_ARCHIVED_MODEL_SEED = 42 +_MAX_TRAIN_SAMPLES = 20_000 +_TRAINING_SAMPLE_SEED_NAME = "calibration_sipp_vehicle_training_sample" +_OBSERVED_SIPP_STATUSES = frozenset((0, 1, 9)) +_OWNED_NONZERO_SHARE_BAND = (0.40, 0.99) +_VALUE_NONZERO_SHARE_BAND = (0.25, 0.99) + +_REQUIRED_RECIPIENT_PERSON_COLUMNS: tuple[str, ...] = ( + "person_household_id", + "age", + "is_female", + "A_LINENO", + "employment_income_before_lsr", + "rental_income", +) + +_RECIPIENT_INTEREST_COLUMNS = ( + "taxable_interest_income", + "tax_exempt_interest_income", +) +_RECIPIENT_DIVIDEND_COLUMNS = ( + "qualified_dividend_income", + "non_qualified_dividend_income", +) + + +def us_sipp_vehicles_stage_spec() -> SourceStageSpec: + """Load and validate the packaged ``vehicle_assets`` stage declaration.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + stage_map = manifest.stage_map() + if _STAGE_NAME not in stage_map: + raise ValueError(f"US source manifest declares no {_STAGE_NAME!r} stage.") + spec = stage_map[_STAGE_NAME] + missing = sorted(set(US_SIPP_VEHICLE_OUTPUT_COLUMNS) - set(spec.outputs)) + if missing: + raise ValueError( + f"{_STAGE_NAME!r} manifest stage does not declare output(s) " + f"{missing}; the runtime and manifest have drifted." + ) + return spec + + +def _sha256_stream(stream: BinaryIO, *, chunk_size: int = 8 * 1024 * 1024) -> str: + import hashlib + + digest = hashlib.sha256() + for chunk in iter(lambda: stream.read(chunk_size), b""): + digest.update(chunk) + return digest.hexdigest() + + +def _sha256_file(path: Path) -> str: + with path.open("rb") as stream: + return _sha256_stream(stream) + + +def _file_matches( + path: Path, + *, + expected_sha256: str | None, + expected_size_bytes: int | None, +) -> bool: + if not path.exists() or not path.is_file(): + return False + if expected_size_bytes is not None and path.stat().st_size != expected_size_bytes: + return False + return expected_sha256 is None or _sha256_file(path) == expected_sha256 + + +def fetch_sipp_2023_vehicle_donor( + cache_dir: str | Path | None = None, + *, + expected_sha256: str | None = SIPP_2023_VEHICLE_DONOR_SHA256, + expected_size_bytes: int | None = SIPP_2023_VEHICLE_DONOR_SIZE_BYTES, + chunk_size: int = 8 * 1024 * 1024, +) -> Path: + """Stream, verify, and atomically cache the pinned full SIPP donor. + + Streaming is mandatory for this 3.73 GB artifact: reading the response or + file into one bytes object would needlessly double peak memory. + """ + + if chunk_size < 1: + raise ValueError("chunk_size must be a positive integer") + + import hashlib + import urllib.request + + root = ( + Path(cache_dir).expanduser() + if cache_dir is not None + else Path.home() / ".cache" / "populace" / "sipp" + ) + if cache_dir is None: + snapshot = ( + Path.home() + / ".cache" + / "huggingface" + / "hub" + / f"models--policyengine--{_RETIRED_DATA_REPOSITORY}" + / "snapshots" + / SIPP_2023_VEHICLE_DONOR_REVISION + / _DONOR_FILENAME + ) + if _file_matches( + snapshot, + expected_sha256=expected_sha256, + expected_size_bytes=expected_size_bytes, + ): + return snapshot + root.mkdir(parents=True, exist_ok=True) + target = root / _DONOR_FILENAME + if _file_matches( + target, + expected_sha256=expected_sha256, + expected_size_bytes=expected_size_bytes, + ): + return target + + partial = target.with_name(f"{target.name}.part") + digest = hashlib.sha256() + written = 0 + try: + with ( + urllib.request.urlopen(SIPP_2023_VEHICLE_DONOR_URL) as response, # noqa: S310 + partial.open("wb") as output, + ): + while True: + chunk = response.read(chunk_size) + if not chunk: + break + output.write(chunk) + digest.update(chunk) + written += len(chunk) + + if expected_size_bytes is not None and written != expected_size_bytes: + raise ValueError( + "SIPP 2023 vehicle donor failed byte-length verification: " + f"expected {expected_size_bytes}, got {written}." + ) + actual_sha256 = digest.hexdigest() + if expected_sha256 is not None and actual_sha256 != expected_sha256: + raise ValueError( + "SIPP 2023 vehicle donor failed sha-256 verification: " + f"expected {expected_sha256}, got {actual_sha256}." + ) + partial.replace(target) + except Exception: + partial.unlink(missing_ok=True) + raise + return target + + +def _stable_string_seed(value: str) -> int: + """Match the archived pipeline's uint64 stable-string seed.""" + + mask = 2**64 - 1 + hashed = 0 + for byte in value.encode("utf-8"): + hashed = (hashed * 31 + byte) & mask + hashed ^= hashed >> 33 + hashed = (hashed * 0xFF51AFD7ED558CCD) & mask + hashed ^= hashed >> 33 + return hashed % (2**63) + + +def _sample_rng(seed_name: str, *, salt: str | None = None) -> np.random.Generator: + key = seed_name if salt is None else f"{seed_name}:{salt}" + return np.random.default_rng(_stable_string_seed(key)) + + +def _cap_vehicle_training_sample(donor: pd.DataFrame) -> pd.DataFrame: + """Apply the archived target-balanced 20,000-row positional cap.""" + + filters = { + US_SIPP_VEHICLE_OUTPUT_COLUMNS[0]: donor[_OWNED_OBSERVED_COLUMN].astype(bool), + US_SIPP_VEHICLE_OUTPUT_COLUMNS[1]: donor[_VALUE_OBSERVED_COLUMN].astype(bool), + } + union = filters[US_SIPP_VEHICLE_OUTPUT_COLUMNS[0]] | filters[ + US_SIPP_VEHICLE_OUTPUT_COLUMNS[1] + ] + union_positions = np.flatnonzero(union.to_numpy()) + if len(union_positions) <= _MAX_TRAIN_SAMPLES: + positions = union_positions + else: + selected: list[int] = [] + selected_set: set[int] = set() + per_target_cap = _MAX_TRAIN_SAMPLES // len(filters) + for target, target_filter in filters.items(): + target_positions = np.flatnonzero(target_filter.to_numpy()) + sampled = _sample_rng( + _TRAINING_SAMPLE_SEED_NAME, salt=target + ).choice( + target_positions, + size=min(per_target_cap, len(target_positions)), + replace=False, + ) + for position in sampled: + if int(position) not in selected_set: + selected.append(int(position)) + selected_set.add(int(position)) + + remaining_n = _MAX_TRAIN_SAMPLES - len(selected) + remaining = np.asarray( + [ + position + for position in union_positions + if int(position) not in selected_set + ], + dtype=int, + ) + if remaining_n > 0 and len(remaining): + fill = _sample_rng(_TRAINING_SAMPLE_SEED_NAME, salt="fill").choice( + remaining, + size=min(remaining_n, len(remaining)), + replace=False, + ) + selected.extend(int(position) for position in fill) + positions = np.asarray(selected, dtype=int) + + return donor.iloc[positions].copy().reset_index(drop=True) + + +def load_sipp_2023_vehicle_donor( + path: str | Path, + *, + expected_sha256: str | None = None, + expected_size_bytes: int | None = SIPP_2023_VEHICLE_DONOR_SIZE_BYTES, + chunksize: int = 100_000, +) -> pd.DataFrame: + """Load and transform the pinned person-month file to household donors.""" + + path = Path(path) + if expected_size_bytes is not None and path.stat().st_size != expected_size_bytes: + raise ValueError( + "SIPP 2023 vehicle donor failed byte-length verification: " + f"expected {expected_size_bytes}, got {path.stat().st_size}." + ) + if expected_sha256 is not None: + actual_sha256 = _sha256_file(path) + if actual_sha256 != expected_sha256: + raise ValueError( + "SIPP 2023 vehicle donor failed sha-256 verification: " + f"expected {expected_sha256}, got {actual_sha256}." + ) + if chunksize < 1: + raise ValueError("chunksize must be a positive integer") + + header = pd.read_csv(path, delimiter="|", nrows=0) + missing = sorted(set(SIPP_VEHICLE_SOURCE_COLUMNS) - set(header.columns)) + if missing: + raise ValueError(f"SIPP 2023 vehicle donor missing column(s): {missing}.") + + december_parts: list[pd.DataFrame] = [] + reader = pd.read_csv( + path, + delimiter="|", + usecols=list(SIPP_VEHICLE_SOURCE_COLUMNS), + chunksize=int(chunksize), + low_memory=False, + ) + for chunk in reader: + month = pd.to_numeric(chunk["MONTHCODE"], errors="coerce") + december = chunk.loc[month.eq(12)].copy() + if not december.empty: + december_parts.append(december) + if not december_parts: + raise ValueError("SIPP 2023 vehicle donor has no December person records.") + person = pd.concat(december_parts, ignore_index=True) + + for column in SIPP_VEHICLE_SOURCE_COLUMNS: + if column != "SSUID": + person[column] = pd.to_numeric(person[column], errors="coerce") + + person["employment_income"] = person["TPTOTINC"].fillna(0.0) * 12.0 + person["interest_income"] = ( + person["TINC_BANK"].fillna(0.0) + person["TINC_BOND"].fillna(0.0) + ) * 12.0 + person["dividend_income"] = person["TINC_STMF"].fillna(0.0) * 12.0 + person["rental_income"] = person["TINC_RENT"].fillna(0.0) * 12.0 + person["is_under_18"] = person["TAGE"].fillna(0.0) < 18 + + grouped = person.groupby("SSUID", sort=True) + reference_index = grouped["TAGE"].idxmax() + if reference_index.isna().any(): + raise ValueError("SIPP vehicle donor has a household with no finite age.") + reference = ( + person.loc[reference_index.astype(int), ["SSUID", "TAGE", "ESEX", "EMS"]] + .rename( + columns={ + "TAGE": "reference_age", + "ESEX": "reference_sex", + "EMS": "reference_marital_status", + } + ) + .set_index("SSUID") + ) + + donor = pd.DataFrame( + { + "household_id": grouped["SSUID"].first(), + _DONOR_WEIGHT_COLUMN: grouped["WPFINWGT"].first().fillna(0.0), + "household_employment_income": grouped["employment_income"].sum(), + "household_interest_income": grouped["interest_income"].sum(), + "household_dividend_income": grouped["dividend_income"].sum(), + "household_rental_income": grouped["rental_income"].sum(), + "count_under_18": grouped["is_under_18"].sum(), + "household_size": grouped.size(), + "household_vehicles_owned": grouped["TVEH_NUM"].max().fillna(0.0), + "household_vehicles_value": grouped["THVAL_VEH"].first().fillna(0.0), + "AVEH_NUM": grouped["AVEH_NUM"].max().fillna(0.0), + "AHVAL_VEH": grouped["AHVAL_VEH"].first().fillna(0.0), + "AVEH1VAL": grouped["AVEH1VAL"].max().fillna(0.0), + "AVEH2VAL": grouped["AVEH2VAL"].max().fillna(0.0), + "AVEH3VAL": grouped["AVEH3VAL"].max().fillna(0.0), + "is_homeowner": ( + grouped["THVAL_HOME"].first().fillna(0.0) > 0 + ).astype(np.float32), + } + ).reset_index(drop=True) + donor = donor.merge( + reference, + left_on="household_id", + right_index=True, + how="left", + ) + donor["reference_is_female"] = ( + donor["reference_sex"].fillna(1.0).eq(2) + ).astype(np.float32) + donor["reference_is_married"] = ( + donor["reference_marital_status"].fillna(0.0).eq(1) + ).astype(np.float32) + donor = donor.drop( + columns=["reference_sex", "reference_marital_status"], + errors="ignore", + ).fillna(0.0) + + donor[_OWNED_OBSERVED_COLUMN] = pd.to_numeric( + donor["AVEH_NUM"], errors="coerce" + ).fillna(0).isin(_OBSERVED_SIPP_STATUSES) + value_observed = pd.Series(True, index=donor.index) + for column in ("AVEH1VAL", "AVEH2VAL", "AVEH3VAL"): + value_observed &= pd.to_numeric(donor[column], errors="coerce").fillna( + 0 + ).isin(_OBSERVED_SIPP_STATUSES) + donor[_VALUE_OBSERVED_COLUMN] = value_observed + + weights = pd.to_numeric(donor[_DONOR_WEIGHT_COLUMN], errors="coerce") + valid_weight = np.isfinite(weights) & weights.gt(0) + donor = donor.loc[valid_weight].copy().reset_index(drop=True) + if donor.empty: + raise ValueError("SIPP vehicle donor has no positive finite household weights.") + if not donor[_OWNED_OBSERVED_COLUMN].any(): + raise ValueError("SIPP vehicle donor has no observed vehicle-count rows.") + if not donor[_VALUE_OBSERVED_COLUMN].any(): + raise ValueError("SIPP vehicle donor has no observed vehicle-value rows.") + return _cap_vehicle_training_sample(donor) + + +def _is_numeric_categorical(values: pd.Series) -> bool: + """Replicate MicroImpute 2.1's low-cardinality numeric test.""" + + if not pd.api.types.is_numeric_dtype(values) or values.nunique() >= 10: + return False + unique = np.sort(pd.to_numeric(values, errors="coerce").dropna().unique()) + if len(unique) < 2: + return True + differences = np.diff(unique) + return bool(np.allclose(differences, differences[0], rtol=1e-9)) + + +@dataclass(frozen=True) +class _PredictorEncoding: + numeric_columns: tuple[str, ...] + categorical_levels: dict[str, tuple[float, ...]] + + @property + def columns(self) -> tuple[str, ...]: + encoded: list[str] = list(self.numeric_columns) + for column, levels in self.categorical_levels.items(): + encoded.extend(_dummy_name(column, level) for level in levels[1:]) + return tuple(encoded) + + def transform(self, frame: pd.DataFrame) -> pd.DataFrame: + result = pd.DataFrame(index=frame.index) + for column in self.numeric_columns: + result[column] = pd.to_numeric(frame[column], errors="coerce").fillna(0.0) + for column, levels in self.categorical_levels.items(): + values = pd.to_numeric(frame[column], errors="coerce").fillna(levels[0]) + for level in levels[1:]: + result[_dummy_name(column, level)] = values.eq(level).astype(np.float64) + return result.loc[:, list(self.columns)] + + +def _dummy_name(column: str, level: float) -> str: + return f"{column}__{float(level):g}" + + +def _predictor_encoding(donor: pd.DataFrame) -> _PredictorEncoding: + numeric: list[str] = [] + categorical: dict[str, tuple[float, ...]] = {} + for column in SIPP_VEHICLE_MODEL_PREDICTORS: + values = pd.to_numeric(donor[column], errors="coerce").fillna(0.0) + if _is_numeric_categorical(values): + categorical[column] = tuple(float(value) for value in np.sort(values.unique())) + else: + numeric.append(column) + return _PredictorEncoding(tuple(numeric), categorical) + + +def _append_owned_dummies( + frame: pd.DataFrame, + owned: pd.Series | np.ndarray, + *, + levels: tuple[float, ...], +) -> tuple[pd.DataFrame, tuple[str, ...]]: + result = frame.copy() + values = np.asarray(owned, dtype=np.float64) + columns: list[str] = [] + for level in levels[1:]: + name = _dummy_name("household_vehicles_owned", level) + result[name] = (values == level).astype(np.float64) + columns.append(name) + return result, tuple(columns) + + +def _household_tenure_status( + frame: Frame, + *, + reference_people: pd.DataFrame, + household_ids: np.ndarray, +) -> np.ndarray: + """Align raw CPS ``SPM_TENMORTSTATUS`` to household order.""" + + household = frame.table("household") + person = frame.table("person") + if "SPM_TENMORTSTATUS" in household.columns: + return pd.to_numeric( + household["SPM_TENMORTSTATUS"], errors="coerce" + ).fillna(3).to_numpy() + if "SPM_TENMORTSTATUS" in person.columns: + values = pd.Series( + pd.to_numeric( + reference_people["SPM_TENMORTSTATUS"], errors="coerce" + ).fillna(3).to_numpy(), + index=reference_people["household_id"].to_numpy(), + ) + return values.reindex(household_ids).fillna(3).to_numpy() + if "spm_unit" in frame.entities: + spm_unit = frame.table("spm_unit") + if ( + "SPM_TENMORTSTATUS" in spm_unit.columns + and "spm_unit_id" in spm_unit.columns + and "person_spm_unit_id" in reference_people.columns + ): + by_spm_unit = pd.Series( + pd.to_numeric( + spm_unit["SPM_TENMORTSTATUS"], errors="coerce" + ).fillna(3).to_numpy(), + index=spm_unit["spm_unit_id"].to_numpy(), + ) + by_household = pd.Series( + reference_people["person_spm_unit_id"].map(by_spm_unit).to_numpy(), + index=reference_people["household_id"].to_numpy(), + ) + return by_household.reindex(household_ids).fillna(3).to_numpy() + raise ValueError( + "US SIPP vehicle imputation requires raw CPS SPM_TENMORTSTATUS on " + "household, person, or spm_unit." + ) + + +def _recipient_income( + person: pd.DataFrame, + *, + aggregate_column: str, + component_columns: tuple[str, ...], +) -> pd.Series: + """Use the archived aggregate concept or its current-frame components.""" + + if aggregate_column in person.columns: + return pd.to_numeric(person[aggregate_column], errors="coerce").fillna(0.0) + present = [column for column in component_columns if column in person.columns] + if not present: + raise ValueError( + "US SIPP vehicle imputation requires either person." + f"{aggregate_column} or component column(s) {list(component_columns)}." + ) + return ( + person.loc[:, present] + .apply(pd.to_numeric, errors="coerce") + .fillna(0.0) + .sum(axis=1) + ) + + +def _recipient_is_married(person: pd.DataFrame) -> np.ndarray: + """Match the archive: existing flag, marital-unit pairing, then CPS code.""" + + if "is_married" in person.columns: + return person["is_married"].astype(bool).to_numpy(dtype=np.float64) + if "person_marital_unit_id" in person.columns: + marital_unit_id = person["person_marital_unit_id"] + counts = marital_unit_id.map(marital_unit_id.value_counts()) + return counts.gt(1).to_numpy(dtype=np.float64) + if "A_MARITL" in person.columns: + return ( + pd.to_numeric(person["A_MARITL"], errors="coerce") + .isin([1, 2]) + .to_numpy(dtype=np.float64) + ) + return np.zeros(len(person), dtype=np.float64) + + +def _recipient_household_predictor_table(frame: Frame) -> pd.DataFrame: + person = frame.table("person") + household = frame.table("household") + missing = sorted(set(_REQUIRED_RECIPIENT_PERSON_COLUMNS) - set(person.columns)) + if missing: + raise ValueError( + f"US SIPP vehicle imputation requires recipient person column(s): {missing}." + ) + if "household_id" not in household.columns: + raise ValueError("US SIPP vehicle imputation requires household.household_id.") + + work = pd.DataFrame(index=person.index) + work["household_id"] = person["person_household_id"].to_numpy() + work["age"] = pd.to_numeric(person["age"], errors="coerce").fillna(0.0) + work["is_female"] = person["is_female"].astype(bool).astype(np.float64) + work["is_married"] = _recipient_is_married(person) + work["A_LINENO"] = pd.to_numeric( + person["A_LINENO"], errors="coerce" + ).fillna(np.inf) + if "person_spm_unit_id" in person.columns: + work["person_spm_unit_id"] = person["person_spm_unit_id"].to_numpy() + if "SPM_TENMORTSTATUS" in person.columns: + work["SPM_TENMORTSTATUS"] = person["SPM_TENMORTSTATUS"].to_numpy() + + income = pd.DataFrame( + { + "household_id": work["household_id"], + "household_employment_income": pd.to_numeric( + person["employment_income_before_lsr"], errors="coerce" + ).fillna(0.0), + "household_interest_income": _recipient_income( + person, + aggregate_column="interest_income", + component_columns=_RECIPIENT_INTEREST_COLUMNS, + ), + "household_dividend_income": _recipient_income( + person, + aggregate_column="dividend_income", + component_columns=_RECIPIENT_DIVIDEND_COLUMNS, + ), + "household_rental_income": pd.to_numeric( + person["rental_income"], errors="coerce" + ).fillna(0.0), + "is_under_18": work["age"].lt(18), + }, + index=person.index, + ) + aggregated = ( + income.groupby("household_id", sort=False) + .agg( + household_employment_income=("household_employment_income", "sum"), + household_interest_income=("household_interest_income", "sum"), + household_dividend_income=("household_dividend_income", "sum"), + household_rental_income=("household_rental_income", "sum"), + count_under_18=("is_under_18", "sum"), + household_size=("household_id", "size"), + ) + .reset_index() + ) + + # CPS ASEC line 1 is the household reference person. Selecting the minimum + # line number retains that rule while remaining robust to a missing line 1. + reference = ( + work.sort_values(["household_id", "A_LINENO"], kind="stable") + .drop_duplicates("household_id") + .copy() + ) + reference = reference.rename( + columns={ + "age": "reference_age", + "is_female": "reference_is_female", + "is_married": "reference_is_married", + } + ) + household_ids = household["household_id"].to_numpy() + tenure_status = _household_tenure_status( + frame, + reference_people=reference, + household_ids=household_ids, + ) + reference = reference.loc[ + :, + [ + "household_id", + "reference_age", + "reference_is_female", + "reference_is_married", + ], + ] + receiver = aggregated.merge(reference, on="household_id", how="left") + receiver = receiver.set_index("household_id").reindex(household_ids) + if receiver.isna().any(axis=None): + missing_ids = receiver.index[receiver.isna().any(axis=1)].tolist() + raise ValueError( + "SIPP vehicle receiver household alignment failed for household(s) " + f"{missing_ids[:5]}." + ) + receiver["is_homeowner"] = np.isin(tenure_status, [1, 2]).astype(np.float64) + return receiver.loc[:, list(SIPP_VEHICLE_MODEL_PREDICTORS)].astype(np.float64) + + +def impute_us_sipp_vehicles( + frame: Frame, + donor: pd.DataFrame, + *, + seed: int = _ARCHIVED_MODEL_SEED, + n_estimators: int = _DEFAULT_N_ESTIMATORS, +) -> pd.DataFrame: + """Fit count first, then draw value conditional on predicted count.""" + + from populace.fit import QRF + + required = { + *SIPP_VEHICLE_MODEL_PREDICTORS, + *US_SIPP_VEHICLE_OUTPUT_COLUMNS, + _DONOR_WEIGHT_COLUMN, + _OWNED_OBSERVED_COLUMN, + _VALUE_OBSERVED_COLUMN, + } + missing = sorted(required - set(donor.columns)) + if missing: + raise ValueError(f"SIPP vehicle donor table missing column(s): {missing}.") + + donor = donor.copy().reset_index(drop=True) + for column in ( + *SIPP_VEHICLE_MODEL_PREDICTORS, + *US_SIPP_VEHICLE_OUTPUT_COLUMNS, + _DONOR_WEIGHT_COLUMN, + ): + donor[column] = pd.to_numeric(donor[column], errors="coerce").fillna(0.0) + encoding = _predictor_encoding(donor) + donor_features = encoding.transform(donor) + receiver = _recipient_household_predictor_table(frame) + receiver_features = encoding.transform(receiver) + + owned_mask = donor[_OWNED_OBSERVED_COLUMN].astype(bool).to_numpy() + value_mask = donor[_VALUE_OBSERVED_COLUMN].astype(bool).to_numpy() + if not owned_mask.any() or not value_mask.any(): + raise ValueError("SIPP vehicle donor lacks observed rows for one or both targets.") + + owned_target = donor.loc[owned_mask, "household_vehicles_owned"] + owned_levels = tuple(float(value) for value in np.sort(owned_target.unique())) + count_model = RandomForestClassifier( + n_estimators=int(n_estimators), + max_depth=None, + min_samples_split=2, + min_samples_leaf=1, + max_features="sqrt", + random_state=int(seed), + ) + count_model.fit( + donor_features.loc[owned_mask], + owned_target, + sample_weight=donor.loc[owned_mask, _DONOR_WEIGHT_COLUMN].to_numpy( + dtype=np.float64 + ), + ) + predicted_owned = np.asarray( + count_model.predict(receiver_features), dtype=np.float64 + ) + + donor_value_features, owned_dummy_columns = _append_owned_dummies( + donor_features, + donor["household_vehicles_owned"], + levels=owned_levels, + ) + receiver_value_features, _ = _append_owned_dummies( + receiver_features, + predicted_owned, + levels=owned_levels, + ) + value_predictors = [*encoding.columns, *owned_dummy_columns] + value_fit = donor_value_features.loc[value_mask, value_predictors].copy() + value_fit["household_vehicles_value"] = donor.loc[ + value_mask, "household_vehicles_value" + ].to_numpy(dtype=np.float64) + value_fit[_DONOR_WEIGHT_COLUMN] = donor.loc[ + value_mask, _DONOR_WEIGHT_COLUMN + ].to_numpy(dtype=np.float64) + fitted_value = QRF(n_estimators=int(n_estimators), seed=int(seed)).fit( + value_fit, + predictors=value_predictors, + targets=["household_vehicles_value"], + weights=_DONOR_WEIGHT_COLUMN, + ) + predicted_value = np.asarray( + fitted_value.predict(receiver_value_features.loc[:, value_predictors])[ + "household_vehicles_value" + ], + dtype=np.float64, + ) + + return pd.DataFrame( + { + "household_vehicles_owned": np.clip( + np.rint(predicted_owned), 0, None + ).astype(np.int32), + "household_vehicles_value": np.clip( + predicted_value, 0, None + ).astype(np.float32), + }, + index=frame.table("household").index, + ) + + +def _surface_has_signal(household: pd.DataFrame) -> bool: + if not all(column in household.columns for column in US_SIPP_VEHICLE_OUTPUT_COLUMNS): + return False + owned = pd.to_numeric( + household["household_vehicles_owned"], errors="coerce" + ).dropna() + value = pd.to_numeric( + household["household_vehicles_value"], errors="coerce" + ).dropna() + return bool( + owned.nunique() > 1 + and value.nunique() > 1 + and (owned > 0).any() + and (value > 0).any() + and (owned >= 0).all() + and (value >= 0).all() + and np.allclose(owned, np.rint(owned)) + ) + + +def with_us_sipp_vehicle_inputs( + frame: Frame, + *, + seed: int, + time_period: int, + sipp_donor: pd.DataFrame, + n_estimators: int = _DEFAULT_N_ESTIMATORS, +) -> Frame: + """Restore both vehicle inputs on the household table.""" + + if frame.schema != US_SCHEMA: + raise ValueError("US SIPP vehicle inputs require the US schema.") + del time_period # Source transformation is fixed by the pinned donor vintage. + us_sipp_vehicles_stage_spec() + household = frame.table("household") + if _surface_has_signal(household): + return frame + + imputed = impute_us_sipp_vehicles( + frame, + sipp_donor, + seed=int(seed), + n_estimators=int(n_estimators), + ) + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + for column in US_SIPP_VEHICLE_OUTPUT_COLUMNS: + tables["household"][column] = imputed[column].to_numpy() + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + ) + + +def us_sipp_vehicles_summary(frame: Frame) -> dict[str, object]: + """Return weighted incidence, amount, type, and relation diagnostics.""" + + household = frame.table("household") + weights = np.asarray(frame.weights_for("household").values, dtype=np.float64) + total_weight = float(weights.sum()) + values: dict[str, np.ndarray] = {} + for column in US_SIPP_VEHICLE_OUTPUT_COLUMNS: + if column in household.columns: + values[column] = pd.to_numeric( + household[column], errors="coerce" + ).to_numpy(dtype=np.float64) + owned = values.get("household_vehicles_owned", np.zeros(len(household))) + value = values.get("household_vehicles_value", np.zeros(len(household))) + finite_owned = np.nan_to_num(owned, nan=0.0, posinf=0.0, neginf=0.0) + finite_value = np.nan_to_num(value, nan=0.0, posinf=0.0, neginf=0.0) + nonzero_shares = { + "household_vehicles_owned": ( + float(weights[finite_owned > 0].sum()) / total_weight + if total_weight > 0 + else 0.0 + ), + "household_vehicles_value": ( + float(weights[finite_value > 0].sum()) / total_weight + if total_weight > 0 + else 0.0 + ), + } + return { + "nonzero_shares": nonzero_shares, + "weighted_totals": { + "household_vehicles_owned": float(np.dot(finite_owned, weights)), + "household_vehicles_value": float(np.dot(finite_value, weights)), + }, + "unique_counts": { + column: int(pd.Series(column_values).dropna().nunique()) + for column, column_values in values.items() + }, + "negative_counts": { + column: int((column_values < 0).sum()) + for column, column_values in values.items() + }, + "nonfinite_counts": { + column: int((~np.isfinite(column_values)).sum()) + for column, column_values in values.items() + }, + "owned_noninteger_count": int( + (~np.isclose(finite_owned, np.rint(finite_owned))).sum() + ), + "positive_value_with_positive_owned_weighted_share": ( + float(weights[(finite_owned > 0) & (finite_value > 0)].sum()) / total_weight + if total_weight > 0 + else 0.0 + ), + "nonzero_share_bands": { + "household_vehicles_owned": list(_OWNED_NONZERO_SHARE_BAND), + "household_vehicles_value": list(_VALUE_NONZERO_SHARE_BAND), + }, + } + + +def us_sipp_vehicles_signal_gate(frame: Frame) -> GateResult: + """Require finite, nonnegative, nonconstant household vehicle signal.""" + + household = frame.table("household") + missing = [ + column + for column in US_SIPP_VEHICLE_OUTPUT_COLUMNS + if column not in household.columns + ] + if missing: + return GateResult( + name="sipp_vehicles_signal", + passed=False, + failures=(f"household columns missing: {missing}.",), + details={"missing": missing}, + ) + + summary = us_sipp_vehicles_summary(frame) + failures: list[str] = [] + for column in US_SIPP_VEHICLE_OUTPUT_COLUMNS: + if summary["unique_counts"][column] < 2: + failures.append(f"{column}: constant column carries no signal.") + if summary["negative_counts"][column]: + failures.append( + f"{column}: {summary['negative_counts'][column]} negative value(s)." + ) + if summary["nonfinite_counts"][column]: + failures.append( + f"{column}: {summary['nonfinite_counts'][column]} non-finite value(s)." + ) + share = float(summary["nonzero_shares"][column]) + low, high = summary["nonzero_share_bands"][column] + if not (low <= share <= high): + failures.append( + f"{column}: nonzero share {share:.3f} outside plausibility band " + f"[{low}, {high}]." + ) + if float(summary["weighted_totals"][column]) <= 0: + failures.append(f"{column}: weighted total is not positive.") + if summary["owned_noninteger_count"]: + failures.append( + "household_vehicles_owned: " + f"{summary['owned_noninteger_count']} non-integer value(s)." + ) + if float(summary["positive_value_with_positive_owned_weighted_share"]) <= 0: + failures.append("No weighted households carry both positive count and value.") + return GateResult( + name="sipp_vehicles_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) diff --git a/packages/populace-build/src/populace/build/us_runtime/source_runtime.py b/packages/populace-build/src/populace/build/us_runtime/source_runtime.py index 22e5c614..6345bba2 100644 --- a/packages/populace-build/src/populace/build/us_runtime/source_runtime.py +++ b/packages/populace-build/src/populace/build/us_runtime/source_runtime.py @@ -17,29 +17,79 @@ from populace.build.us_runtime.capital_gain_distributions import ( split_us_component_by_share_from_manifest, ) +from populace.build.us_runtime.child_support import ( + derive_us_child_support_from_manifest, + impute_us_child_support_to_puf_support_from_manifest, +) +from populace.build.us_runtime.childcare import ( + derive_us_childcare_from_manifest, + impute_us_childcare_to_puf_support_from_manifest, +) +from populace.build.us_runtime.disability_benefits import ( + derive_us_disability_benefits_from_manifest, + impute_us_disability_benefits_to_puf_support_from_manifest, +) +from populace.build.us_runtime.education_inputs import ( + derive_us_education_inputs_from_manifest, +) from populace.build.us_runtime.eligibility_inputs import ( derive_us_eligibility_inputs_from_manifest, ) +from populace.build.us_runtime.energy_subsidy import ( + derive_us_energy_subsidy_from_manifest, + impute_us_energy_subsidy_to_puf_support_from_manifest, +) from populace.build.us_runtime.hours_worked import ( derive_us_hours_worked_from_manifest, ) from populace.build.us_runtime.immigration import ( derive_us_immigration_status_from_manifest, ) +from populace.build.us_runtime.medicare_take_up import ( + derive_us_medicare_take_up_from_manifest, +) +from populace.build.us_runtime.other_health_insurance import ( + derive_us_other_health_insurance_from_manifest, + impute_us_other_health_insurance_to_puf_support_from_manifest, +) from populace.build.us_runtime.pregnancy import ( derive_us_pregnancy_from_manifest, ) +from populace.build.us_runtime.prior_year_income import ( + derive_us_prior_year_income_from_manifest, + impute_us_prior_year_income_to_puf_support_from_manifest, +) from populace.build.us_runtime.puf_aggregate_records import ( derive_puf_policyengine_variables, disaggregate_puf_aggregate_records, load_default_puf_aggregate_disaggregation_spec, ) +from populace.build.us_runtime.relationship_inputs import ( + derive_us_relationship_inputs_from_manifest, +) +from populace.build.us_runtime.retirement_contributions import ( + derive_us_retirement_contributions_from_manifest, + impute_us_retirement_contributions_to_puf_support_from_manifest, +) +from populace.build.us_runtime.retirement_distributions import ( + derive_us_retirement_distributions_from_manifest, + impute_us_retirement_distributions_to_puf_support_from_manifest, +) from populace.build.us_runtime.snap_discretionary_exemption import ( derive_us_snap_discretionary_exemption_from_manifest, ) from populace.build.us_runtime.snap_take_up import ( derive_us_snap_take_up_from_manifest, ) +from populace.build.us_runtime.weeks_unemployed import ( + derive_us_weeks_unemployed_from_manifest, + impute_us_weeks_unemployed_to_puf_support_from_manifest, +) +from populace.build.us_runtime.wic_claim import derive_us_wic_claim_from_manifest +from populace.build.us_runtime.workers_compensation import ( + derive_us_workers_compensation_from_manifest, + impute_us_workers_compensation_to_puf_support_from_manifest, +) __all__ = [ "aggregate_us_person_to_tax_unit_from_manifest", @@ -47,8 +97,31 @@ "calibrate_us_binary_assignment_from_manifest", "calibrate_us_binary_assignment_joint_targets_from_manifest", "compute_us_ratio_from_manifest", + "derive_us_childcare_from_manifest", + "derive_us_child_support_from_manifest", + "derive_us_disability_benefits_from_manifest", + "derive_us_energy_subsidy_from_manifest", + "derive_us_medicare_take_up_from_manifest", + "derive_us_other_health_insurance_from_manifest", + "derive_us_prior_year_income_from_manifest", + "derive_us_relationship_inputs_from_manifest", + "derive_us_retirement_distributions_from_manifest", + "derive_us_workers_compensation_from_manifest", + "derive_us_weeks_unemployed_from_manifest", + "derive_us_wic_claim_from_manifest", + "impute_us_retirement_distributions_to_puf_support_from_manifest", "derive_us_puf_policyengine_variables_from_manifest", + "derive_us_retirement_contributions_from_manifest", "disaggregate_us_puf_aggregate_records_from_manifest", + "impute_us_childcare_to_puf_support_from_manifest", + "impute_us_child_support_to_puf_support_from_manifest", + "impute_us_disability_benefits_to_puf_support_from_manifest", + "impute_us_energy_subsidy_to_puf_support_from_manifest", + "impute_us_other_health_insurance_to_puf_support_from_manifest", + "impute_us_prior_year_income_to_puf_support_from_manifest", + "impute_us_retirement_contributions_to_puf_support_from_manifest", + "impute_us_workers_compensation_to_puf_support_from_manifest", + "impute_us_weeks_unemployed_to_puf_support_from_manifest", "support_clip_us_source_output_from_manifest", "us_source_operation_handlers", ] @@ -70,6 +143,33 @@ "qualified_dividend_source", "qualified_dividend_output", "non_qualified_dividend_output", + "qualified_tuition_primary_source", + "qualified_tuition_optional_source", + "qualified_tuition_output", + "alimony_income_source", + "alimony_income_output", + "alimony_expense_source", + "alimony_expense_output", + "casualty_loss_source", + "casualty_loss_output", + "domestic_production_ald_source", + "domestic_production_ald_output", + "educator_expense_source", + "educator_expense_output", + "unreimbursed_business_employee_expenses_source", + "unreimbursed_business_employee_expenses_output", + "farm_operations_income_source", + "farm_operations_income_output", + "farm_rent_income_source", + "farm_rent_income_output", + "investment_income_elected_form_4952_source", + "investment_income_elected_form_4952_output", + "salt_refund_income_source", + "salt_refund_income_output", + "collectibles_capital_gain_source", + "collectibles_capital_gain_output", + "unrecaptured_section_1250_gain_source", + "unrecaptured_section_1250_gain_output", } ) @@ -194,9 +294,59 @@ def us_source_operation_handlers() -> Mapping[str, SourceOperationHandler]: calibrate_us_binary_assignment_joint_targets_from_manifest ), "compute_ratio": compute_us_ratio_from_manifest, + "derive_childcare_inputs": derive_us_childcare_from_manifest, + "derive_child_support_inputs": derive_us_child_support_from_manifest, + "derive_disability_benefits": derive_us_disability_benefits_from_manifest, + "derive_energy_subsidy": derive_us_energy_subsidy_from_manifest, + "derive_other_health_insurance_premiums": ( + derive_us_other_health_insurance_from_manifest + ), "derive_eligibility_inputs": derive_us_eligibility_inputs_from_manifest, + "derive_education_inputs": derive_us_education_inputs_from_manifest, "derive_hours_worked": derive_us_hours_worked_from_manifest, + "derive_medicare_take_up": derive_us_medicare_take_up_from_manifest, "derive_pregnancy": derive_us_pregnancy_from_manifest, + "derive_prior_year_income": derive_us_prior_year_income_from_manifest, + "derive_relationship_inputs": derive_us_relationship_inputs_from_manifest, + "derive_retirement_distributions": ( + derive_us_retirement_distributions_from_manifest + ), + "derive_retirement_contributions": ( + derive_us_retirement_contributions_from_manifest + ), + "derive_workers_compensation": derive_us_workers_compensation_from_manifest, + "derive_weeks_unemployed": derive_us_weeks_unemployed_from_manifest, + "derive_wic_claim": derive_us_wic_claim_from_manifest, + "impute_childcare_to_puf_support": ( + impute_us_childcare_to_puf_support_from_manifest + ), + "impute_child_support_to_puf_support": ( + impute_us_child_support_to_puf_support_from_manifest + ), + "impute_disability_benefits_to_puf_support": ( + impute_us_disability_benefits_to_puf_support_from_manifest + ), + "impute_energy_subsidy_to_puf_support": ( + impute_us_energy_subsidy_to_puf_support_from_manifest + ), + "impute_other_health_insurance_premiums_to_puf_support": ( + impute_us_other_health_insurance_to_puf_support_from_manifest + ), + "impute_prior_year_income_to_puf_support": ( + impute_us_prior_year_income_to_puf_support_from_manifest + ), + "impute_retirement_distributions_to_puf_support": ( + impute_us_retirement_distributions_to_puf_support_from_manifest + ), + "impute_retirement_contributions_to_puf_support": ( + impute_us_retirement_contributions_to_puf_support_from_manifest + ), + "impute_workers_compensation_to_puf_support": ( + impute_us_workers_compensation_to_puf_support_from_manifest + ), + "impute_weeks_unemployed_to_puf_support": ( + impute_us_weeks_unemployed_to_puf_support_from_manifest + ), "derive_snap_abawd_discretionary_exemption": derive_us_snap_discretionary_exemption_from_manifest, "derive_immigration_status": derive_us_immigration_status_from_manifest, "derive_snap_take_up": derive_us_snap_take_up_from_manifest, @@ -742,6 +892,156 @@ def derive_us_puf_policyengine_variables_from_manifest( default="non_qualified_dividend_income", label="PUF PolicyEngine-variable derivation", ), + qualified_tuition_primary_source=_optional_string_param( + params, + "qualified_tuition_primary_source", + label="PUF PolicyEngine-variable derivation", + ), + qualified_tuition_optional_source=_optional_string_param( + params, + "qualified_tuition_optional_source", + label="PUF PolicyEngine-variable derivation", + ), + qualified_tuition_output=_string_param_with_default( + params, + "qualified_tuition_output", + default="qualified_tuition_expenses", + label="PUF PolicyEngine-variable derivation", + ), + alimony_income_source=_optional_string_param( + params, + "alimony_income_source", + label="PUF PolicyEngine-variable derivation", + ), + alimony_income_output=_string_param_with_default( + params, + "alimony_income_output", + default="alimony_income", + label="PUF PolicyEngine-variable derivation", + ), + alimony_expense_source=_optional_string_param( + params, + "alimony_expense_source", + label="PUF PolicyEngine-variable derivation", + ), + alimony_expense_output=_string_param_with_default( + params, + "alimony_expense_output", + default="alimony_expense", + label="PUF PolicyEngine-variable derivation", + ), + casualty_loss_source=_optional_string_param( + params, + "casualty_loss_source", + label="PUF PolicyEngine-variable derivation", + ), + casualty_loss_output=_string_param_with_default( + params, + "casualty_loss_output", + default="casualty_loss", + label="PUF PolicyEngine-variable derivation", + ), + domestic_production_ald_source=_optional_string_param( + params, + "domestic_production_ald_source", + label="PUF PolicyEngine-variable derivation", + ), + domestic_production_ald_output=_string_param_with_default( + params, + "domestic_production_ald_output", + default="domestic_production_ald", + label="PUF PolicyEngine-variable derivation", + ), + educator_expense_source=_optional_string_param( + params, + "educator_expense_source", + label="PUF PolicyEngine-variable derivation", + ), + educator_expense_output=_string_param_with_default( + params, + "educator_expense_output", + default="educator_expense", + label="PUF PolicyEngine-variable derivation", + ), + unreimbursed_business_employee_expenses_source=_optional_string_param( + params, + "unreimbursed_business_employee_expenses_source", + label="PUF PolicyEngine-variable derivation", + ), + unreimbursed_business_employee_expenses_output=( + _string_param_with_default( + params, + "unreimbursed_business_employee_expenses_output", + default="unreimbursed_business_employee_expenses", + label="PUF PolicyEngine-variable derivation", + ) + ), + farm_operations_income_source=_optional_string_param( + params, + "farm_operations_income_source", + label="PUF PolicyEngine-variable derivation", + ), + farm_operations_income_output=_string_param_with_default( + params, + "farm_operations_income_output", + default="farm_operations_income", + label="PUF PolicyEngine-variable derivation", + ), + farm_rent_income_source=_optional_string_param( + params, + "farm_rent_income_source", + label="PUF PolicyEngine-variable derivation", + ), + farm_rent_income_output=_string_param_with_default( + params, + "farm_rent_income_output", + default="farm_rent_income", + label="PUF PolicyEngine-variable derivation", + ), + investment_income_elected_form_4952_source=_optional_string_param( + params, + "investment_income_elected_form_4952_source", + label="PUF PolicyEngine-variable derivation", + ), + investment_income_elected_form_4952_output=_string_param_with_default( + params, + "investment_income_elected_form_4952_output", + default="investment_income_elected_form_4952", + label="PUF PolicyEngine-variable derivation", + ), + salt_refund_income_source=_optional_string_param( + params, + "salt_refund_income_source", + label="PUF PolicyEngine-variable derivation", + ), + salt_refund_income_output=_string_param_with_default( + params, + "salt_refund_income_output", + default="salt_refund_income", + label="PUF PolicyEngine-variable derivation", + ), + collectibles_capital_gain_source=_optional_string_param( + params, + "collectibles_capital_gain_source", + label="PUF PolicyEngine-variable derivation", + ), + collectibles_capital_gain_output=_string_param_with_default( + params, + "collectibles_capital_gain_output", + default="long_term_capital_gains_on_collectibles", + label="PUF PolicyEngine-variable derivation", + ), + unrecaptured_section_1250_gain_source=_optional_string_param( + params, + "unrecaptured_section_1250_gain_source", + label="PUF PolicyEngine-variable derivation", + ), + unrecaptured_section_1250_gain_output=_string_param_with_default( + params, + "unrecaptured_section_1250_gain_output", + default="unrecaptured_section_1250_gain", + label="PUF PolicyEngine-variable derivation", + ), ) except ValueError as exc: raise SourceRuntimeError(str(exc)) from exc diff --git a/packages/populace-build/src/populace/build/us_runtime/ssi_disability_criteria.py b/packages/populace-build/src/populace/build/us_runtime/ssi_disability_criteria.py new file mode 100644 index 00000000..673dd645 --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/ssi_disability_criteria.py @@ -0,0 +1,1244 @@ +"""SIPP-imputed latent SSI disability-criteria input. + +The retired enhanced-CPS pipeline trained a person-grain boolean QRF on the +December SIPP. Its label was deliberately narrower than a general disability +flag: an under-65 SIPP SSI recipient whose *reported* benefit reason was +disabled/blind was positive. Observed nonrecipients were usable negatives only +when they passed a source-time approximation of the SSI resource, countable- +income, and substantial-gainful-activity screens. The exact nineteen archived +predictors are retained below. + +The source transform also retains two distinctions that are easy to lose: + +* the receipt and benefit-reason allocation flags decide whether a label is + observed; and +* the final QRF result is accepted only for a person carrying at least one of + the six measured disability difficulties, Social Security disability, or + other disability income. + +The extended-CPS pipeline predicted its ASEC and PUF-support people separately, +because the latter carried separately imputed income and asset predictors. We +do the same. Direct under-65 ASEC ``SSI_VAL`` reporters are then preserved as +positive anchors; that anchor is not copied onto the PUF channel. An arbitrary +pre-existing criterion column is never trusted: every run recomputes the full +source-backed surface and uses an equality check only for idempotent return. + +The full 2023 SIPP public-use file is the same immutable 3.73 GB artifact +already pinned by the vehicle and voluntary-filing stages. It contains 39,513 +December person rows. Under the archived 2024 financial screen it yields +9,346 training candidates (577 positive and 8,769 negative). +""" + +from __future__ import annotations + +from copy import deepcopy +from importlib.resources import files +from pathlib import Path +from typing import Any + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import SourceStageSpec, load_source_manifest +from populace.build.us_runtime.voluntary_filing import ( + SIPP_2023_VOLUNTARY_FILING_DONOR_REVISION, + SIPP_2023_VOLUNTARY_FILING_DONOR_SHA256, + SIPP_2023_VOLUNTARY_FILING_DONOR_SIZE_BYTES, + SIPP_2023_VOLUNTARY_FILING_DONOR_URL, + fetch_sipp_2023_voluntary_filing_donor, +) +from populace.frame import Frame +from populace.frame.units import US_SCHEMA + +__all__ = [ + "SIPP_2023_SSI_DISABILITY_DONOR_REVISION", + "SIPP_2023_SSI_DISABILITY_DONOR_SHA256", + "SIPP_2023_SSI_DISABILITY_DONOR_SIZE_BYTES", + "SIPP_2023_SSI_DISABILITY_DONOR_URL", + "SIPP_SSI_DISABILITY_DIFFICULTY_PREDICTORS", + "SIPP_SSI_DISABILITY_FIT_PARAMETERS", + "SIPP_SSI_DISABILITY_MODEL_PREDICTORS", + "SIPP_SSI_DISABILITY_READ_PARAMETERS", + "SIPP_SSI_DISABILITY_SOURCE_COLUMNS", + "SSI_DISABILITY_ARCHIVED_CPS_URL", + "SSI_DISABILITY_ARCHIVED_EXTENDED_CPS_URL", + "SSI_DISABILITY_ARCHIVED_SIPP_URL", + "SSI_DISABILITY_ARCHIVED_SOURCE_IMPUTE_URL", + "SSI_DISABILITY_SIPP_DICTIONARY_URL", + "US_SSI_DISABILITY_CRITERIA_NONCONSTANT_PERSON_COLUMNS", + "US_SSI_DISABILITY_CRITERIA_OUTPUT_COLUMNS", + "US_SSI_DISABILITY_CRITERIA_STAGE_NAME", + "fetch_sipp_2023_ssi_disability_donor", + "impute_us_ssi_disability_criteria", + "load_sipp_2023_ssi_disability_donor", + "us_ssi_disability_criteria_signal_gate", + "us_ssi_disability_criteria_stage_spec", + "us_ssi_disability_criteria_summary", + "with_us_ssi_disability_criteria", +] + +QRF: Any | None = None + +_ARCHIVED_COMMIT = "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe" +_RETIRED_REPOSITORY = "policyengine-" + "us-data" +_RETIRED_PACKAGE = "policyengine_" + "us_data" +_ARCHIVED_ROOT = ( + f"https://github.com/PolicyEngine/{_RETIRED_REPOSITORY}/blob/" + f"{_ARCHIVED_COMMIT}/{_RETIRED_PACKAGE}/" +) +SSI_DISABILITY_ARCHIVED_SIPP_URL = _ARCHIVED_ROOT + "datasets/sipp/sipp.py#L63-L105" +SSI_DISABILITY_ARCHIVED_CPS_URL = _ARCHIVED_ROOT + "datasets/cps/cps.py#L2853-L2886" +SSI_DISABILITY_ARCHIVED_SOURCE_IMPUTE_URL = ( + _ARCHIVED_ROOT + "calibration/source_impute.py#L869-L990" +) +SSI_DISABILITY_ARCHIVED_EXTENDED_CPS_URL = ( + _ARCHIVED_ROOT + "datasets/cps/extended_cps.py#L392-L424" +) +SSI_DISABILITY_SIPP_DICTIONARY_URL = ( + "https://www2.census.gov/programs-surveys/sipp/tech-documentation/" + "data-dictionaries/2023/2023_SIPP_Data_Dictionary.pdf" +) + +# Reuse the exact full-file coordinate and fetch/cache implementation already +# owned by the voluntary-filing stage. Aliases keep this family's provenance +# legible to manifests and release telemetry. +SIPP_2023_SSI_DISABILITY_DONOR_REVISION = SIPP_2023_VOLUNTARY_FILING_DONOR_REVISION +SIPP_2023_SSI_DISABILITY_DONOR_SHA256 = SIPP_2023_VOLUNTARY_FILING_DONOR_SHA256 +SIPP_2023_SSI_DISABILITY_DONOR_SIZE_BYTES = SIPP_2023_VOLUNTARY_FILING_DONOR_SIZE_BYTES +SIPP_2023_SSI_DISABILITY_DONOR_URL = SIPP_2023_VOLUNTARY_FILING_DONOR_URL + +US_SSI_DISABILITY_CRITERIA_STAGE_NAME = "ssi_disability_criteria" +US_SSI_DISABILITY_CRITERIA_OUTPUT_COLUMNS: tuple[str, ...] = ( + "meets_ssi_disability_criteria", +) +US_SSI_DISABILITY_CRITERIA_NONCONSTANT_PERSON_COLUMNS = ( + US_SSI_DISABILITY_CRITERIA_OUTPUT_COLUMNS +) + +SIPP_SSI_DISABILITY_DIFFICULTY_PREDICTORS: tuple[str, ...] = ( + "difficulty_dressing_or_bathing", + "difficulty_hearing", + "difficulty_seeing", + "difficulty_doing_errands", + "difficulty_walking_or_climbing_stairs", + "difficulty_remembering_or_making_decisions", +) +_SIPP_DIFFICULTY_SOURCE_COLUMNS: dict[str, str] = { + "difficulty_dressing_or_bathing": "ESELFCARE", + "difficulty_hearing": "EHEARING", + "difficulty_seeing": "ESEEING", + "difficulty_doing_errands": "EERRANDS", + "difficulty_walking_or_climbing_stairs": "EAMBULAT", + "difficulty_remembering_or_making_decisions": "ECOGNIT", +} +_ASEC_DIFFICULTY_SOURCE_COLUMNS: dict[str, str] = { + "difficulty_dressing_or_bathing": "PEDISDRS", + "difficulty_hearing": "PEDISEAR", + "difficulty_seeing": "PEDISEYE", + "difficulty_doing_errands": "PEDISOUT", + "difficulty_walking_or_climbing_stairs": "PEDISPHY", + "difficulty_remembering_or_making_decisions": "PEDISREM", +} +SIPP_SSI_DISABILITY_MODEL_PREDICTORS: tuple[str, ...] = ( + "age", + "is_female", + "is_married", + "employment_income", + "interest_income", + "dividend_income", + "rental_income", + "bank_account_assets", + "stock_assets", + "bond_assets", + "count_under_18", + *SIPP_SSI_DISABILITY_DIFFICULTY_PREDICTORS, + "social_security_disability", + "has_disability_income", +) + +_SIPP_JOB_EARNINGS_COLUMNS = tuple(f"TJB{i}_MSUM" for i in range(1, 8)) +_SIPP_DISABILITY_INCOME_AMOUNT_COLUMNS = tuple(f"TDIS{i}AMT" for i in range(1, 11)) +_SIPP_ASSET_ALLOCATION_COLUMNS: tuple[str, ...] = ( + "AJSSAVVAL", + "AJOSAVVAL", + "AOSAVVAL", + "AJSMMVAL", + "AJOMMVAL", + "AOMMVAL", + "AJSCDVAL", + "AJOCDVAL", + "AOCDVAL", + "AJSCHKVAL", + "AJOCHKVAL", + "AOCHKVAL", + "AJSSTVAL", + "AJOSTVAL", + "AOSTVAL", + "AJSMFVAL", + "AJOMFVAL", + "AOMFVAL", + "AJSGOVSVAL", + "AJOGOVSVAL", + "AOGOVSVAL", + "AJSMCBDVAL", + "AJOMCBDVAL", + "AOMCBDVAL", +) +_SIPP_BASE_ASSET_COLUMNS: tuple[str, ...] = ( + "SSUID", + "PNUM", + "MONTHCODE", + "SPANEL", + "SWAVE", + "WPFINWGT", + "TAGE", + "ESEX", + "EMS", + "TSSSAMT", + "TRETINCAMT", + "TVAL_BANK", + "TVAL_STMF", + "TVAL_BOND", + "TINC_BANK", + "TINC_STMF", + "TINC_BOND", + "TINC_RENT", + *_SIPP_JOB_EARNINGS_COLUMNS, + *_SIPP_ASSET_ALLOCATION_COLUMNS, +) +SIPP_SSI_DISABILITY_SOURCE_COLUMNS: tuple[str, ...] = tuple( + sorted( + { + *_SIPP_BASE_ASSET_COLUMNS, + "TPTOTINC", + "RSSI_YRYN", + "EDISABL", + "EHLTHCOND", + "RDIS", + "RDIS_ALT", + "EDISANY", + "ENJ_NOWRK3", + "ESSRSN2YN", + "ESSI_BRSN", + *_SIPP_DIFFICULTY_SOURCE_COLUMNS.values(), + *_SIPP_DISABILITY_INCOME_AMOUNT_COLUMNS, + "ASSI_YRYN", + "ASSI_BRSN", + } + ) +) + +_OUTPUT = US_SSI_DISABILITY_CRITERIA_OUTPUT_COLUMNS[0] +_DONOR_WEIGHT_COLUMN = "household_weight" +_TRAINING_CANDIDATE_COLUMN = "ssi_disability_training_candidate" +_DEFAULT_N_ESTIMATORS = 100 +_MAX_TRAIN_SAMPLES = 20_000 +_TRAINING_SAMPLE_SEED_NAME = "sipp_ssi_disability_model_training_sample" +_TRAINING_SAMPLE_SEED = 8_386_123_572_872_638_692 +_ARCHIVED_MODEL_SEED = 42 +_PERSON_SUPPORT_CHANNEL_COLUMN = "person_support_channel" +_PERSON_SOURCE_ID_COLUMN = "person_source_id" +_BASE_ASEC_SUPPORT_CHANNEL = "asec" +_PUF_TAX_DETAIL_SUPPORT_CHANNEL = "puf_tax_detail" +_RECIPIENT_INTEREST_COMPONENT_COLUMNS = ( + "taxable_interest_income", + "tax_exempt_interest_income", +) +_RECIPIENT_DIVIDEND_COMPONENT_COLUMNS = ( + "qualified_dividend_income", + "non_qualified_dividend_income", +) +_TRUE_SHARE_BAND = (0.001, 0.25) +_OBSERVED_ALLOCATION_VALUES = frozenset((0, 1, 9)) + +_PINNED_DECEMBER_ROWS = 39_513 +_PINNED_TRAINING_ROWS = 9_346 +_PINNED_POSITIVE_ROWS = 577 +_PINNED_NEGATIVE_ROWS = 8_769 +_PINNED_WEIGHT_SUM = 88_690_359.47893329 +_PINNED_POSITIVE_WEIGHT_SUM = 4_937_167.914119501 +_PINNED_WEIGHTED_TRUE_SHARE = 0.05566746987074994 +_PINNED_RESAMPLE_ROWS = 9_346 +_PINNED_RESAMPLE_UNIQUE_SOURCE_ROWS = 5_314 +_PINNED_RESAMPLE_POSITIVE_ROWS = 524 +_PINNED_RESAMPLE_TRUE_SHARE = 0.05606676653113631 +_PINNED_FLOAT_ATOL = 1e-9 + +SIPP_SSI_DISABILITY_READ_PARAMETERS: dict[str, object] = { + "table": "sipp_person", + "delimiter": "|", + "month_column": "MONTHCODE", + "month": 12, + "source_columns": list(SIPP_SSI_DISABILITY_SOURCE_COLUMNS), +} +SIPP_SSI_DISABILITY_FIT_PARAMETERS: dict[str, object] = { + "predictors": list(SIPP_SSI_DISABILITY_MODEL_PREDICTORS), + "target": _OUTPUT, + "weight": _DONOR_WEIGHT_COLUMN, + "training_candidate": _TRAINING_CANDIDATE_COLUMN, + "label_source_columns": ["RSSI_YRYN", "ESSI_BRSN"], + "label_allocation_columns": ["ASSI_YRYN", "ASSI_BRSN"], + "max_train_samples": _MAX_TRAIN_SAMPLES, + "sample_with_replacement": True, + "training_sample_seed_name": _TRAINING_SAMPLE_SEED_NAME, + "training_sample_seed": _TRAINING_SAMPLE_SEED, + "n_estimators": _DEFAULT_N_ESTIMATORS, + "model_seed": _ARCHIVED_MODEL_SEED, + "seed_from_build_config": False, + "postprediction_signal_predictors": [ + *SIPP_SSI_DISABILITY_DIFFICULTY_PREDICTORS, + "social_security_disability", + "has_disability_income", + ], + "preserve_under_65_asec_ssi_reporters": True, +} + + +def us_ssi_disability_criteria_stage_spec() -> SourceStageSpec: + """Load and strictly validate the packaged source-stage declaration.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + stage_map = manifest.stage_map() + if US_SSI_DISABILITY_CRITERIA_STAGE_NAME not in stage_map: + raise ValueError( + "US source manifest declares no " + f"{US_SSI_DISABILITY_CRITERIA_STAGE_NAME!r} stage." + ) + spec = stage_map[US_SSI_DISABILITY_CRITERIA_STAGE_NAME] + if spec.grain != "person": + raise ValueError("US SSI disability-criteria stage must have person grain.") + if tuple(spec.outputs) != US_SSI_DISABILITY_CRITERIA_OUTPUT_COLUMNS: + raise ValueError( + "US SSI disability-criteria manifest outputs drifted from the " + "runtime-owned family." + ) + if [operation.kind for operation in spec.operations] != [ + "read_table", + "fit_weighted_qrf", + ]: + raise ValueError( + "US SSI disability-criteria stage must contain read_table then " + "fit_weighted_qrf." + ) + if dict(spec.operations[0].parameters) != SIPP_SSI_DISABILITY_READ_PARAMETERS: + raise ValueError( + "US SSI disability-criteria read_table contract drifted from the " + "pinned SIPP transform." + ) + if dict(spec.operations[1].parameters) != SIPP_SSI_DISABILITY_FIT_PARAMETERS: + raise ValueError( + "US SSI disability-criteria QRF contract drifted from the archived method." + ) + pinned = [ + artifact + for artifact in spec.artifacts + if artifact.get("sha256") == SIPP_2023_SSI_DISABILITY_DONOR_SHA256 + and artifact.get("size_bytes") == SIPP_2023_SSI_DISABILITY_DONOR_SIZE_BYTES + ] + if not pinned: + raise ValueError( + "US SSI disability-criteria stage does not pin the full SIPP " + "SHA-256 and byte length." + ) + return spec + + +def _sha256_file(path: Path) -> str: + import hashlib + + digest = hashlib.sha256() + with path.open("rb") as stream: + for chunk in iter(lambda: stream.read(8 * 1024 * 1024), b""): + digest.update(chunk) + return digest.hexdigest() + + +def fetch_sipp_2023_ssi_disability_donor( + cache_dir: str | Path | None = None, + *, + expected_sha256: str | None = SIPP_2023_SSI_DISABILITY_DONOR_SHA256, + expected_size_bytes: int | None = SIPP_2023_SSI_DISABILITY_DONOR_SIZE_BYTES, + chunk_size: int = 8 * 1024 * 1024, +) -> Path: + """Fetch the shared pinned full SIPP artifact through its canonical cache.""" + + return fetch_sipp_2023_voluntary_filing_donor( + cache_dir, + expected_sha256=expected_sha256, + expected_size_bytes=expected_size_bytes, + chunk_size=chunk_size, + ) + + +def _numeric(series: pd.Series) -> pd.Series: + return pd.to_numeric(series, errors="coerce").astype("float64") + + +def _yes(frame: pd.DataFrame, column: str) -> pd.Series: + return _numeric(frame[column]).fillna(0.0).eq(1.0) + + +def _monthly_earned_income(frame: pd.DataFrame) -> pd.Series: + return frame.loc[:, list(_SIPP_JOB_EARNINGS_COLUMNS)].fillna(0.0).sum(axis=1) + + +def _ssi_policy_screen_values(time_period: int) -> dict[str, float]: + """Return the exact 2024-style values read by the archived source screen.""" + + try: + from policyengine_us import CountryTaxBenefitSystem + + parameters = CountryTaxBenefitSystem().parameters(f"{time_period}-01-01") + ssi = parameters.gov.ssa.ssi + exclusions = ssi.income.exclusions + return { + "individual_resource_limit": float( + ssi.eligibility.resources.limit.individual + ), + "couple_resource_limit": float(ssi.eligibility.resources.limit.couple), + "individual_fbr": float(ssi.amount.individual), + "couple_fbr": float(ssi.amount.couple), + "general_exclusion": float(exclusions.general), + "earned_exclusion": float(exclusions.earned), + "earned_share_excluded": float(exclusions.earned_share), + "non_blind_sga": float(parameters.gov.ssa.sga.non_blind), + } + except Exception: + # These are the archived function's explicit 2024 fallback constants. + return { + "individual_resource_limit": 2_000.0, + "couple_resource_limit": 3_000.0, + "individual_fbr": 943.0, + "couple_fbr": 1_415.0, + "general_exclusion": 20.0, + "earned_exclusion": 65.0, + "earned_share_excluded": 0.5, + "non_blind_sga": 1_550.0, + } + + +def _financial_candidate_mask( + frame: pd.DataFrame, + *, + time_period: int, +) -> pd.Series: + values = _ssi_policy_screen_values(time_period) + married = frame["is_married"].astype(bool) + resource_limit = np.where( + married, + values["couple_resource_limit"], + values["individual_resource_limit"], + ) + income_limit = np.where( + married, + values["couple_fbr"], + values["individual_fbr"], + ) + liquid_resources = ( + frame["bank_account_assets"].fillna(0.0) + + frame["stock_assets"].fillna(0.0) + + frame["bond_assets"].fillna(0.0) + ) + earned = _monthly_earned_income(frame) + unearned = (frame["TPTOTINC"].fillna(0.0) - earned).clip(lower=0.0) + applied_general = np.minimum(values["general_exclusion"], unearned) + countable_unearned = unearned - applied_general + leftover_general = values["general_exclusion"] - applied_general + earned_after_flat = (earned - values["earned_exclusion"] - leftover_general).clip( + lower=0.0 + ) + countable_earned = earned_after_flat * (1.0 - values["earned_share_excluded"]) + countable_income = countable_unearned + countable_earned + is_blind = frame["difficulty_seeing"].fillna(False).astype(bool) + passes_sga = is_blind | earned.le(values["non_blind_sga"]) + return ( + liquid_resources.le(resource_limit) + & countable_income.le(income_limit) + & passes_sga + ) + + +def _observed_label_mask( + frame: pd.DataFrame, + received_ssi: pd.Series, +) -> pd.Series: + receipt_flag = _numeric(frame["ASSI_YRYN"]).fillna(0.0) + reason_flag = _numeric(frame["ASSI_BRSN"]).fillna(0.0) + receipt_observed = ( + frame["RSSI_YRYN"].notna() + & receipt_flag.isin(_OBSERVED_ALLOCATION_VALUES) + & _numeric(frame["RSSI_YRYN"]).isin([1.0, 2.0]) + ) + reason_observed = ( + frame["ESSI_BRSN"].notna() + & reason_flag.isin(_OBSERVED_ALLOCATION_VALUES) + & _numeric(frame["ESSI_BRSN"]).isin([1.0, 2.0]) + ) + return receipt_observed & (~received_ssi | reason_observed) + + +def _assert_pinned_source_audit(audit: dict[str, object]) -> None: + expected_counts = { + "december_rows": _PINNED_DECEMBER_ROWS, + "training_rows": _PINNED_TRAINING_ROWS, + "positive_rows": _PINNED_POSITIVE_ROWS, + "negative_rows": _PINNED_NEGATIVE_ROWS, + "resample_rows": _PINNED_RESAMPLE_ROWS, + "resample_unique_source_rows": _PINNED_RESAMPLE_UNIQUE_SOURCE_ROWS, + "resample_positive_rows": _PINNED_RESAMPLE_POSITIVE_ROWS, + } + mismatches = { + key: (expected, audit.get(key)) + for key, expected in expected_counts.items() + if audit.get(key) != expected + } + expected_floats = { + "weight_sum": _PINNED_WEIGHT_SUM, + "positive_weight_sum": _PINNED_POSITIVE_WEIGHT_SUM, + "weighted_true_share": _PINNED_WEIGHTED_TRUE_SHARE, + "resample_true_share": _PINNED_RESAMPLE_TRUE_SHARE, + } + for key, expected in expected_floats.items(): + actual = float(audit[key]) + if not np.isclose(actual, expected, rtol=0.0, atol=_PINNED_FLOAT_ATOL): + mismatches[key] = (expected, actual) + if mismatches: + raise ValueError( + "Pinned SIPP SSI disability source audit drifted from the reviewed " + f"2024 transform: {mismatches}." + ) + + +def load_sipp_2023_ssi_disability_donor( + path: str | Path, + *, + expected_sha256: str | None = None, + expected_size_bytes: int | None = SIPP_2023_SSI_DISABILITY_DONOR_SIZE_BYTES, + chunksize: int = 100_000, + time_period: int = 2024, +) -> pd.DataFrame: + """Build the exact observed/candidate SIPP SSI disability training frame.""" + + source_path = Path(path) + actual_size = source_path.stat().st_size + if expected_size_bytes is not None and actual_size != expected_size_bytes: + raise ValueError( + "SIPP 2023 SSI disability donor failed byte-length verification: " + f"expected {expected_size_bytes}, got {actual_size}." + ) + if expected_sha256 is not None: + actual_sha256 = _sha256_file(source_path) + if actual_sha256 != expected_sha256: + raise ValueError( + "SIPP 2023 SSI disability donor failed sha-256 verification: " + f"expected {expected_sha256}, got {actual_sha256}." + ) + if chunksize < 1: + raise ValueError("chunksize must be a positive integer") + + header = pd.read_csv(source_path, delimiter="|", nrows=0) + missing = sorted(set(SIPP_SSI_DISABILITY_SOURCE_COLUMNS) - set(header.columns)) + if missing: + raise ValueError(f"SIPP SSI disability donor missing column(s): {missing}.") + + parts: list[pd.DataFrame] = [] + for chunk in pd.read_csv( + source_path, + delimiter="|", + usecols=list(SIPP_SSI_DISABILITY_SOURCE_COLUMNS), + chunksize=int(chunksize), + low_memory=False, + ): + month = _numeric(chunk["MONTHCODE"]) + december = chunk.loc[month.eq(12.0)].copy() + if not december.empty: + parts.append(december) + if not parts: + raise ValueError("SIPP SSI disability donor has no December rows.") + frame = pd.concat(parts, ignore_index=True) + + for column in SIPP_SSI_DISABILITY_SOURCE_COLUMNS: + if column == "SSUID": + continue + original = frame[column] + converted = _numeric(original) + invalid = original.notna() & converted.isna() + if invalid.any(): + rows = np.flatnonzero(invalid.to_numpy())[:5].tolist() + raise ValueError( + f"SIPP SSI disability source {column!r} contains nonnumeric " + f"value(s) at row(s) {rows}." + ) + frame[column] = converted + if frame[["SSUID", "PNUM"]].isna().any(axis=None): + raise ValueError("SIPP SSI disability December source has missing IDs.") + + frame["bank_account_assets"] = frame["TVAL_BANK"].fillna(0.0) + frame["stock_assets"] = frame["TVAL_STMF"].fillna(0.0) + frame["bond_assets"] = frame["TVAL_BOND"].fillna(0.0) + frame["age"] = frame["TAGE"] + frame["is_female"] = frame["ESEX"].eq(2.0) + frame["is_married"] = frame["EMS"].eq(1.0) + earned = _monthly_earned_income(frame) + frame["employment_income"] = earned * 12.0 + frame["interest_income"] = ( + frame["TINC_BANK"].fillna(0.0) + frame["TINC_BOND"].fillna(0.0) + ) * 12.0 + frame["dividend_income"] = frame["TINC_STMF"].fillna(0.0) * 12.0 + frame["rental_income"] = frame["TINC_RENT"].fillna(0.0) * 12.0 + frame[_DONOR_WEIGHT_COLUMN] = frame["WPFINWGT"].fillna(0.0) + is_under_18 = frame["TAGE"].lt(18.0) + frame["count_under_18"] = is_under_18.groupby(frame["SSUID"]).transform("sum") + for predictor, source_column in _SIPP_DIFFICULTY_SOURCE_COLUMNS.items(): + frame[predictor] = _yes(frame, source_column) + + disability_income_amount = pd.Series(0.0, index=frame.index) + for column in _SIPP_DISABILITY_INCOME_AMOUNT_COLUMNS: + disability_income_amount += frame[column].fillna(0.0) + frame["social_security_disability"] = np.where( + _yes(frame, "ESSRSN2YN"), + frame["TSSSAMT"].fillna(0.0) * 12.0, + 0.0, + ) + frame["has_disability_income"] = _yes( + frame, "EDISANY" + ) | disability_income_amount.gt(0.0) + + received_ssi = _yes(frame, "RSSI_YRYN") + under_65 = frame["age"].lt(65.0) + reason = frame["ESSI_BRSN"].fillna(-9.0).astype(float) + disabled_or_blind_reason = reason.eq(1.0) + aged_reason = reason.eq(2.0) + frame[_OUTPUT] = received_ssi & under_65 & (disabled_or_blind_reason | ~aged_reason) + financial_candidate = _financial_candidate_mask( + frame, + time_period=int(time_period), + ) + frame[_TRAINING_CANDIDATE_COLUMN] = (financial_candidate & under_65) | frame[ + _OUTPUT + ] + observed = _observed_label_mask(frame, received_ssi) + + columns = [ + *SIPP_SSI_DISABILITY_MODEL_PREDICTORS, + _OUTPUT, + _TRAINING_CANDIDATE_COLUMN, + _DONOR_WEIGHT_COLUMN, + ] + observed_frame = frame.loc[observed, columns].dropna().copy() + donor = observed_frame.loc[observed_frame[_TRAINING_CANDIDATE_COLUMN]].copy() + donor = donor.drop(columns=[_TRAINING_CANDIDATE_COLUMN]).reset_index(drop=True) + + numeric_columns = [*SIPP_SSI_DISABILITY_MODEL_PREDICTORS, _DONOR_WEIGHT_COLUMN] + finite = np.isfinite(donor.loc[:, numeric_columns].to_numpy(dtype=np.float64)).all( + axis=1 + ) + if not finite.all(): + rows = np.flatnonzero(~finite)[:5].tolist() + raise ValueError( + "SIPP SSI disability donor contains nonfinite training values at " + f"row(s) {rows}." + ) + weights = donor[_DONOR_WEIGHT_COLUMN].to_numpy(dtype=np.float64) + if not (weights > 0.0).all(): + rows = np.flatnonzero(weights <= 0.0)[:5].tolist() + raise ValueError( + "SIPP SSI disability donor weights must be positive; invalid " + f"row(s) {rows}." + ) + target = donor[_OUTPUT].astype(bool).to_numpy() + if np.unique(target).size != 2: + raise ValueError("SIPP SSI disability donor target must contain both classes.") + + probability = weights / weights.sum() + selected = np.random.default_rng(_TRAINING_SAMPLE_SEED).choice( + len(donor), + size=min(_MAX_TRAIN_SAMPLES, len(donor)), + replace=True, + p=probability, + ) + resample_target = target[selected] + audit: dict[str, object] = { + "december_rows": int(len(frame)), + "observed_label_rows_before_predictor_dropna": int(observed.sum()), + "observed_complete_rows": int(len(observed_frame)), + "training_rows": int(len(donor)), + "positive_rows": int(target.sum()), + "negative_rows": int((~target).sum()), + "weight_sum": float(weights.sum()), + "positive_weight_sum": float(weights[target].sum()), + "weighted_true_share": float(weights[target].sum() / weights.sum()), + "resample_rows": int(len(selected)), + "resample_unique_source_rows": int(np.unique(selected).size), + "resample_positive_rows": int(resample_target.sum()), + "resample_true_share": float(resample_target.mean()), + "time_period": int(time_period), + } + pinned_transform = ( + actual_size == SIPP_2023_SSI_DISABILITY_DONOR_SIZE_BYTES + and int(time_period) == 2024 + ) + audit["pinned_transform"] = pinned_transform + if pinned_transform: + _assert_pinned_source_audit(audit) + donor.attrs["source_audit"] = audit + return donor + + +def _decoded_strings(values: pd.Series) -> pd.Series: + return values.map( + lambda value: ( + value.decode() if isinstance(value, (bytes, np.bytes_)) else str(value) + ) + ) + + +def _strict_person_numeric( + person: pd.DataFrame, + columns: tuple[str, ...], + *, + label: str, +) -> np.ndarray: + for column in columns: + if column in person: + values = pd.to_numeric(person[column], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not np.isfinite(values).all(): + raise ValueError( + f"US SSI disability receiver {label} source {column!r} " + "contains nonfinite values." + ) + return values + raise ValueError( + f"US SSI disability receiver requires one of {list(columns)} for {label}." + ) + + +def _strict_person_aggregate( + person: pd.DataFrame, + *, + aggregate_column: str, + component_columns: tuple[str, ...], +) -> np.ndarray: + if aggregate_column in person: + return _strict_person_numeric( + person, + (aggregate_column,), + label=aggregate_column, + ) + missing = [column for column in component_columns if column not in person] + if missing: + raise ValueError( + f"US SSI disability receiver requires person.{aggregate_column} or " + f"all measured component columns {list(component_columns)}; missing " + f"{missing}." + ) + return np.sum( + np.column_stack( + [ + _strict_person_numeric(person, (column,), label=column) + for column in component_columns + ] + ), + axis=1, + ) + + +def _person_ssi_disability_predictors(frame: Frame) -> pd.DataFrame: + """Build the exact nineteen predictors on every recipient support row.""" + + person = frame.table("person") + required = { + "person_household_id", + "bank_account_assets", + "stock_assets", + "bond_assets", + "rental_income", + "social_security_disability", + "disability_benefits", + *_ASEC_DIFFICULTY_SOURCE_COLUMNS.values(), + } + missing = sorted(required - set(person.columns)) + if missing: + raise ValueError( + f"US SSI disability receiver missing measured person column(s): {missing}." + ) + + receiver = pd.DataFrame(index=person.index) + receiver["age"] = _strict_person_numeric(person, ("age", "A_AGE"), label="age") + if "is_female" in person: + female = _strict_person_numeric(person, ("is_female",), label="sex") + if not np.isin(female, [0.0, 1.0]).all(): + raise ValueError("US SSI disability receiver is_female must be boolean.") + receiver["is_female"] = female + elif "is_male" in person: + male = _strict_person_numeric(person, ("is_male",), label="sex") + if not np.isin(male, [0.0, 1.0]).all(): + raise ValueError("US SSI disability receiver is_male must be boolean.") + receiver["is_female"] = 1.0 - male + elif "A_SEX" in person: + sex = _strict_person_numeric(person, ("A_SEX",), label="sex") + if not np.isin(sex, [1.0, 2.0]).all(): + raise ValueError("US SSI disability receiver A_SEX must be coded 1 or 2.") + receiver["is_female"] = (sex == 2.0).astype(np.float64) + else: + raise ValueError( + "US SSI disability receiver requires is_female, is_male, or A_SEX." + ) + + if "is_married" in person: + married = _strict_person_numeric(person, ("is_married",), label="marriage") + if not np.isin(married, [0.0, 1.0]).all(): + raise ValueError("US SSI disability receiver is_married must be boolean.") + receiver["is_married"] = married + elif "A_MARITL" in person: + marital = _strict_person_numeric(person, ("A_MARITL",), label="marriage") + receiver["is_married"] = np.isin(marital, [1.0, 2.0]).astype(np.float64) + else: + raise ValueError( + "US SSI disability receiver requires measured is_married or A_MARITL." + ) + + receiver["employment_income"] = _strict_person_numeric( + person, + ("employment_income_before_lsr", "employment_income", "WSAL_VAL"), + label="employment income", + ) + receiver["interest_income"] = _strict_person_aggregate( + person, + aggregate_column="interest_income", + component_columns=_RECIPIENT_INTEREST_COMPONENT_COLUMNS, + ) + receiver["dividend_income"] = _strict_person_aggregate( + person, + aggregate_column="dividend_income", + component_columns=_RECIPIENT_DIVIDEND_COMPONENT_COLUMNS, + ) + for column in ( + "rental_income", + "bank_account_assets", + "stock_assets", + "bond_assets", + "social_security_disability", + ): + receiver[column] = _strict_person_numeric(person, (column,), label=column) + + age = receiver["age"] + receiver["count_under_18"] = ( + age.lt(18.0) + .groupby(person["person_household_id"], sort=False) + .transform("sum") + .to_numpy(dtype=np.float64) + ) + for predictor, source_column in _ASEC_DIFFICULTY_SOURCE_COLUMNS.items(): + source = _strict_person_numeric( + person, + (source_column,), + label=predictor, + ) + receiver[predictor] = (source == 1.0).astype(np.float64) + disability_income = _strict_person_numeric( + person, + ("disability_benefits",), + label="disability income", + ) + receiver["has_disability_income"] = (disability_income > 0.0).astype(np.float64) + + values = receiver.loc[:, list(SIPP_SSI_DISABILITY_MODEL_PREDICTORS)].to_numpy( + dtype=np.float64 + ) + if not np.isfinite(values).all(): + raise ValueError("US SSI disability receiver predictors must be finite.") + return receiver.loc[:, list(SIPP_SSI_DISABILITY_MODEL_PREDICTORS)] + + +def _validate_support_provenance(person: pd.DataFrame) -> None: + if _PERSON_SUPPORT_CHANNEL_COLUMN not in person: + return + if person[_PERSON_SUPPORT_CHANNEL_COLUMN].isna().any(): + raise ValueError( + "US SSI disability support rows contain missing support-channel values." + ) + channels = _decoded_strings(person[_PERSON_SUPPORT_CHANNEL_COLUMN]) + known = {_BASE_ASEC_SUPPORT_CHANNEL, _PUF_TAX_DETAIL_SUPPORT_CHANNEL} + observed = set(channels.unique()) + unexpected = sorted(observed - known) + missing = sorted(known - observed) + if unexpected or missing: + raise ValueError( + "US SSI disability receiver requires exact ASEC/PUF support " + f"channels; missing {missing}, unsupported support channel(s) " + f"{unexpected}." + ) + if ( + _PERSON_SOURCE_ID_COLUMN not in person + or person[_PERSON_SOURCE_ID_COLUMN].isna().any() + ): + raise ValueError( + "US SSI disability support rows require complete person_source_id " + "provenance; PUF-only IDs are allowed but missing IDs are not." + ) + + +def _coerce_boolean_predictions(values: pd.Series | np.ndarray) -> np.ndarray: + series = pd.Series(values) + if series.dtype == bool: + return series.to_numpy(dtype=bool) + if np.issubdtype(series.dtype, np.number): + numeric = pd.to_numeric(series, errors="coerce").to_numpy(dtype=np.float64) + if not np.isfinite(numeric).all(): + raise ValueError("SIPP SSI disability QRF produced nonfinite values.") + if ((numeric < 0.0) | (numeric > 1.0)).any(): + raise ValueError("SIPP SSI disability QRF produced values outside [0, 1].") + return numeric >= 0.5 + normalized = series.fillna("").astype(str).str.strip().str.lower() + valid = normalized.isin(["true", "false", "1", "0", "yes", "no"]) + if not valid.all(): + raise ValueError("SIPP SSI disability QRF produced invalid class labels.") + return normalized.isin(["true", "1", "yes"]).to_numpy(dtype=bool) + + +def _weighted_replacement_sample(donor: pd.DataFrame) -> pd.DataFrame: + weights = pd.to_numeric(donor[_DONOR_WEIGHT_COLUMN], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not (np.isfinite(weights) & (weights > 0.0)).all(): + raise ValueError( + "SIPP SSI disability donor weights must be finite and positive." + ) + probability = weights / weights.sum() + rng = np.random.default_rng(_TRAINING_SAMPLE_SEED) + selected = rng.choice( + len(donor), + size=min(_MAX_TRAIN_SAMPLES, len(donor)), + replace=True, + p=probability, + ) + return donor.iloc[selected].reset_index(drop=True) + + +def impute_us_ssi_disability_criteria( + frame: Frame, + donor: pd.DataFrame, + *, + seed: int, + n_estimators: int = _DEFAULT_N_ESTIMATORS, +) -> pd.Series: + """Fit the archived SIPP model and predict every ASEC/PUF support row.""" + + required = { + *SIPP_SSI_DISABILITY_MODEL_PREDICTORS, + _OUTPUT, + _DONOR_WEIGHT_COLUMN, + } + missing = sorted(required - set(donor.columns)) + if missing: + raise ValueError(f"SIPP SSI disability donor missing column(s): {missing}.") + del seed # The archived model fixed both its source draw and forest seed. + if n_estimators < 1: + raise ValueError("n_estimators must be positive") + + training = donor.loc[ + :, [*SIPP_SSI_DISABILITY_MODEL_PREDICTORS, _OUTPUT, _DONOR_WEIGHT_COLUMN] + ].copy() + for column in [*SIPP_SSI_DISABILITY_MODEL_PREDICTORS, _DONOR_WEIGHT_COLUMN]: + training[column] = pd.to_numeric(training[column], errors="coerce") + values = training.loc[ + :, [*SIPP_SSI_DISABILITY_MODEL_PREDICTORS, _DONOR_WEIGHT_COLUMN] + ].to_numpy(dtype=np.float64) + if not np.isfinite(values).all(): + raise ValueError("SIPP SSI disability donor predictors/weights must be finite.") + target = pd.to_numeric(training[_OUTPUT], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not (np.isfinite(target) & np.isin(target, [0.0, 1.0])).all(): + raise ValueError("SIPP SSI disability donor target must be boolean.") + if np.unique(target).size != 2: + raise ValueError("SIPP SSI disability donor target must contain both classes.") + training[_OUTPUT] = target + training = _weighted_replacement_sample(training) + + person = frame.table("person") + _validate_support_provenance(person) + receiver = _person_ssi_disability_predictors(frame) + global QRF + if QRF is None: + from importlib import import_module + + QRF = import_module("populace.fit").QRF + fitted = QRF(n_estimators=int(n_estimators), seed=_ARCHIVED_MODEL_SEED).fit( + training, + predictors=list(SIPP_SSI_DISABILITY_MODEL_PREDICTORS), + targets=[_OUTPUT], + weights="none", + ) + if _PERSON_SUPPORT_CHANNEL_COLUMN not in person: + prediction = fitted.predict(receiver) + if _OUTPUT not in prediction: + raise ValueError(f"SIPP SSI disability QRF prediction missing {_OUTPUT!r}.") + predicted = _coerce_boolean_predictions(prediction[_OUTPUT]) + else: + # The retired direct-CPS and extended-CPS paths each loaded a fresh + # copy of the same fitted model before drawing their channel. Populace + # QRF predictions advance a seeded draw RNG, so one combined call would + # make PUF outcomes depend on the number/order of preceding ASEC rows. + # Deep-copying the pristine fit per channel reproduces the two fresh + # prediction streams without refitting the reviewed forest. + channels = _decoded_strings(person[_PERSON_SUPPORT_CHANNEL_COLUMN]) + predicted = np.zeros(len(person), dtype=bool) + for channel in ( + _BASE_ASEC_SUPPORT_CHANNEL, + _PUF_TAX_DETAIL_SUPPORT_CHANNEL, + ): + mask = channels.eq(channel).to_numpy() + channel_model = deepcopy(fitted) + prediction = channel_model.predict(receiver.loc[mask]) + if _OUTPUT not in prediction: + raise ValueError( + f"SIPP SSI disability QRF prediction missing {_OUTPUT!r} " + f"for support channel {channel!r}." + ) + predicted[mask] = _coerce_boolean_predictions(prediction[_OUTPUT]) + del channel_model + + difficulty_signal = ( + receiver.loc[:, list(SIPP_SSI_DISABILITY_DIFFICULTY_PREDICTORS)] + .astype(bool) + .any(axis=1) + .to_numpy() + ) + disability_signal = ( + difficulty_signal + | (receiver["social_security_disability"].to_numpy(dtype=np.float64) > 0.0) + | receiver["has_disability_income"].to_numpy(dtype=np.float64).astype(bool) + ) + result = predicted & disability_signal + + # The archived direct-CPS pass preserves measured SSI reporters. Its PUF + # clone override does not, even though raw ASEC columns were duplicated. + if "SSI_VAL" not in person: + raise ValueError( + "US SSI disability receiver requires measured ASEC SSI_VAL for the " + "under-65 reporter anchor." + ) + reported_ssi = ( + _strict_person_numeric( + person, + ("SSI_VAL",), + label="reported SSI", + ) + > 0.0 + ) + under_65 = receiver["age"].to_numpy(dtype=np.float64) < 65.0 + if _PERSON_SUPPORT_CHANNEL_COLUMN in person: + channels = _decoded_strings(person[_PERSON_SUPPORT_CHANNEL_COLUMN]) + asec = channels.eq(_BASE_ASEC_SUPPORT_CHANNEL).to_numpy() + else: + asec = np.ones(len(person), dtype=bool) + result |= asec & under_65 & reported_ssi + return pd.Series(result, index=person.index, name=_OUTPUT, dtype=bool) + + +def with_us_ssi_disability_criteria( + frame: Frame, + *, + seed: int, + time_period: int, + sipp_donor: pd.DataFrame, + n_estimators: int = _DEFAULT_N_ESTIMATORS, +) -> Frame: + """Recompute and materialize the source-backed person input.""" + + if frame.schema != US_SCHEMA: + raise ValueError("US SSI disability criteria require the US schema.") + del time_period # Donor construction owns the reviewed parameter vintage. + us_ssi_disability_criteria_stage_spec() + predicted = impute_us_ssi_disability_criteria( + frame, + sipp_donor, + seed=int(seed), + n_estimators=int(n_estimators), + ) + person = frame.table("person") + if _OUTPUT in person: + current = pd.to_numeric(person[_OUTPUT], errors="coerce").to_numpy( + dtype=np.float64 + ) + if ( + np.isfinite(current).all() + and np.isin(current, [0.0, 1.0]).all() + and np.array_equal(current.astype(bool), predicted.to_numpy()) + ): + return frame + + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + tables["person"][_OUTPUT] = predicted.to_numpy(dtype=bool) + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + ) + + +def us_ssi_disability_criteria_summary(frame: Frame) -> dict[str, object]: + """Return weighted incidence, channel signal, and source-anchor diagnostics.""" + + person = frame.table("person") + values = pd.to_numeric(person[_OUTPUT], errors="coerce").to_numpy(dtype=np.float64) + weights = np.asarray(frame.resolve_weights("person").values, dtype=np.float64) + if not (np.isfinite(weights) & (weights >= 0.0)).all() or weights.sum() <= 0.0: + raise ValueError( + "US SSI disability gate requires finite nonnegative person weights " + "with positive total." + ) + finite = np.isfinite(values) + boolean = finite & np.isin(values, [0.0, 1.0]) + positive = boolean & (values == 1.0) + total_weight = float(weights.sum()) + + channels: dict[str, dict[str, float | int]] = {} + provenance_missing = bool( + _PERSON_SUPPORT_CHANNEL_COLUMN not in person + or person[_PERSON_SUPPORT_CHANNEL_COLUMN].isna().any() + or _PERSON_SOURCE_ID_COLUMN not in person + or person[_PERSON_SOURCE_ID_COLUMN].isna().any() + ) + clone_divergence_source_people = 0 + if _PERSON_SUPPORT_CHANNEL_COLUMN in person: + channel_values = _decoded_strings(person[_PERSON_SUPPORT_CHANNEL_COLUMN]) + for channel in sorted(channel_values.unique()): + mask = channel_values.eq(channel).to_numpy() + channel_weight = float(weights[mask].sum()) + channels[channel] = { + "rows": int(mask.sum()), + "unique_count": int(pd.Series(values[mask & finite]).nunique()), + "weighted_true_share": ( + float(weights[mask & positive].sum()) / channel_weight + if channel_weight > 0.0 + else 0.0 + ), + } + if not provenance_missing: + clone_table = pd.DataFrame( + { + "source_id": _decoded_strings(person[_PERSON_SOURCE_ID_COLUMN]), + "value": values, + } + ) + unique = clone_table.groupby("source_id", sort=False)["value"].nunique( + dropna=False + ) + clone_divergence_source_people = int((unique > 1).sum()) + + reporter_mismatches = 0 + if "SSI_VAL" in person: + reported = pd.to_numeric(person["SSI_VAL"], errors="coerce").fillna(0.0) > 0.0 + age_column = "age" if "age" in person else "A_AGE" + age = pd.to_numeric(person[age_column], errors="coerce") + asec = pd.Series(True, index=person.index) + if _PERSON_SUPPORT_CHANNEL_COLUMN in person: + asec = _decoded_strings(person[_PERSON_SUPPORT_CHANNEL_COLUMN]).eq( + _BASE_ASEC_SUPPORT_CHANNEL + ) + anchor = (reported & age.lt(65.0) & asec).to_numpy() + reporter_mismatches = int(np.count_nonzero(anchor & ~positive)) + + return { + "weighted_true_share": float(weights[positive].sum()) / total_weight, + "weighted_true_total": float(weights[positive].sum()), + "weighted_false_total": float(weights[boolean & ~positive].sum()), + "true_share_band": list(_TRUE_SHARE_BAND), + "unique_count": int(pd.Series(values[finite]).nunique()), + "missing_or_nonfinite_count": int((~finite).sum()), + "invalid_boolean_count": int((finite & ~boolean).sum()), + "reporter_anchor_mismatches": reporter_mismatches, + "support_provenance_missing": provenance_missing, + "clone_divergence_source_people": clone_divergence_source_people, + "channels": channels, + } + + +def us_ssi_disability_criteria_signal_gate(frame: Frame) -> GateResult: + """Require a valid, nonconstant signal in ASEC and PUF support channels.""" + + person = frame.table("person") + if _OUTPUT not in person: + return GateResult( + name="ssi_disability_criteria_signal", + passed=False, + failures=(f"person column missing: {_OUTPUT!r}.",), + details={"missing": [_OUTPUT]}, + ) + summary = us_ssi_disability_criteria_summary(frame) + failures: list[str] = [] + if summary["missing_or_nonfinite_count"]: + failures.append( + f"{_OUTPUT}: {summary['missing_or_nonfinite_count']} missing or " + "nonfinite value(s)." + ) + if summary["invalid_boolean_count"]: + failures.append( + f"{_OUTPUT}: {summary['invalid_boolean_count']} non-boolean value(s)." + ) + if summary["unique_count"] < 2: + failures.append(f"{_OUTPUT}: constant column carries no signal.") + low, high = summary["true_share_band"] + share = float(summary["weighted_true_share"]) + if not (low <= share <= high): + failures.append( + f"{_OUTPUT}: weighted true share {share:.4f} outside plausibility " + f"band [{low}, {high}]." + ) + if float(summary["weighted_true_total"]) <= 0.0: + failures.append(f"{_OUTPUT}: weighted true total is not positive.") + if float(summary["weighted_false_total"]) <= 0.0: + failures.append(f"{_OUTPUT}: weighted false total is not positive.") + if summary["reporter_anchor_mismatches"]: + failures.append( + f"{_OUTPUT}: {summary['reporter_anchor_mismatches']} under-65 ASEC " + "SSI reporter anchor(s) were lost." + ) + if summary["support_provenance_missing"]: + failures.append( + "SSI disability support rows lack complete person_support_channel " + "or person_source_id provenance." + ) + channels = summary["channels"] + required_channels = { + _BASE_ASEC_SUPPORT_CHANNEL, + _PUF_TAX_DETAIL_SUPPORT_CHANNEL, + } + unexpected_channels = sorted(set(channels) - required_channels) + if unexpected_channels: + failures.append( + f"{_OUTPUT}: unsupported support channel(s) {unexpected_channels}." + ) + for required_channel in sorted(required_channels): + diagnostics = channels.get(required_channel) + if diagnostics is None: + failures.append( + f"{_OUTPUT}: support channel {required_channel!r} is missing." + ) + elif diagnostics["unique_count"] < 2: + failures.append( + f"{_OUTPUT}: support channel {required_channel!r} is constant." + ) + elif not (low <= float(diagnostics["weighted_true_share"]) <= high): + failures.append( + f"{_OUTPUT}: support channel {required_channel!r} weighted " + "true share " + f"{float(diagnostics['weighted_true_share']):.4f} outside " + f"plausibility band [{low}, {high}]." + ) + return GateResult( + name="ssi_disability_criteria_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) diff --git a/packages/populace-build/src/populace/build/us_runtime/ssi_take_up.py b/packages/populace-build/src/populace/build/us_runtime/ssi_take_up.py new file mode 100644 index 00000000..41ada110 --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/ssi_take_up.py @@ -0,0 +1,950 @@ +"""Reporter-anchored SSI take-up calibrated to SSA recipient counts. + +The retired eCPS exported ``takes_up_ssi_if_eligible`` after preserving every +CPS ASEC ``SSI_VAL > 0`` reporter and filling additional people to a scalar +50 percent rate. That rate cannot be retained: its only citation estimates +participation among adults age 65 or older, while the retired code applied it +to children and working-age disabled people too. + +This stage keeps the source-backed half of that method (the reporter anchor) +and replaces the scope-invalid rate with SSA's December 2024 counts of people +receiving a *federal payment*, split into the three published age bands. It +uses PolicyEngine-US ``uncapped_ssi > 0`` as the current-benefit candidate +domain, makes one stable decision per ``person_source_id``, and fans that +decision to every actual support row for the source person. A target above +modeled capacity saturates at capacity and records the shortfall; eligibility +is never broadened to manufacture recipients. + +The caller supplies December ``uncapped_ssi`` because the fiscal builder owns +PolicyEngine simulations and batching. Assignment is always recomputed from +the supplied frame weights. The builder fixes it on an initial calibrated +weight surface, replays dependent health inputs and fiscal targets, then +refits on the unchanged support. Returned release weights are diagnosed +without rewriting the persisted assignment; only those final diagnostics are +published. +""" + +from __future__ import annotations + +import hashlib +import json +from collections.abc import Collection, Mapping +from dataclasses import dataclass +from importlib.resources import files +from pathlib import Path + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import SourceStageSpec, load_source_manifest +from populace.frame import Frame +from populace.frame.units import US_SCHEMA + +__all__ = [ + "SSI_TAKE_UP_ARCHIVED_DERIVATION_URL", + "SSI_TAKE_UP_ARCHIVED_EXPORT_URL", + "SSI_TAKE_UP_ARCHIVED_RANDOMNESS_URL", + "SSI_TAKE_UP_ARCHIVED_REPORTER_URL", + "SSI_TAKE_UP_ARCHIVED_TARGETS_URL", + "SSI_TAKE_UP_SSA_SOURCE_URL", + "SSITakeUpAgeTarget", + "US_SSI_TAKE_UP_AGE_TARGETS", + "US_SSI_TAKE_UP_ANCHOR", + "US_SSI_TAKE_UP_NONCONSTANT_PERSON_COLUMNS", + "US_SSI_TAKE_UP_OUTPUT_COLUMNS", + "US_SSI_TAKE_UP_REQUIRED_SOURCE_COLUMNS", + "US_SSI_TAKE_UP_STAGE_NAME", + "US_SSI_TAKE_UP_TARGET_TABLE_NAME", + "us_ssi_take_up_diagnostics", + "us_ssi_take_up_gate", + "us_ssi_take_up_reporter_source_ids", + "us_ssi_take_up_stage_spec", + "with_us_ssi_take_up", + "write_us_ssi_take_up_diagnostics", +] + +_ARCHIVED_COMMIT = "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe" +_RETIRED_REPOSITORY = "policyengine-" + "us-data" +_RETIRED_PACKAGE = "policyengine_" + "us_data" +_ARCHIVED_ROOT = ( + f"https://github.com/PolicyEngine/{_RETIRED_REPOSITORY}/blob/" + f"{_ARCHIVED_COMMIT}/{_RETIRED_PACKAGE}/" +) +SSI_TAKE_UP_ARCHIVED_DERIVATION_URL = _ARCHIVED_ROOT + "datasets/cps/cps.py#L650-L657" +SSI_TAKE_UP_ARCHIVED_EXPORT_URL = _ARCHIVED_ROOT + "datasets/cps/cps.py#L1497-L1499" +SSI_TAKE_UP_ARCHIVED_REPORTER_URL = _ARCHIVED_ROOT + "datasets/cps/cps.py#L584" +SSI_TAKE_UP_ARCHIVED_RANDOMNESS_URL = _ARCHIVED_ROOT + "datasets/cps/takeup.py#L10-L35" +SSI_TAKE_UP_ARCHIVED_TARGETS_URL = _ARCHIVED_ROOT + "utils/ssi_targets.py#L41-L74" +SSI_TAKE_UP_SSA_SOURCE_URL = ( + "https://www.ssa.gov/policy/docs/statcomps/ssi_monthly/2024-12/table01.html" +) + +US_SSI_TAKE_UP_STAGE_NAME = "ssi_take_up" +US_SSI_TAKE_UP_TARGET_TABLE_NAME = "ssa_ssi_federal_payment_recipients_by_age" +US_SSI_TAKE_UP_ANCHOR = "SSI_VAL" +US_SSI_TAKE_UP_OUTPUT_COLUMNS: tuple[str, ...] = ("takes_up_ssi_if_eligible",) +US_SSI_TAKE_UP_NONCONSTANT_PERSON_COLUMNS = US_SSI_TAKE_UP_OUTPUT_COLUMNS +US_SSI_TAKE_UP_REQUIRED_SOURCE_COLUMNS: tuple[str, ...] = ( + "age", + US_SSI_TAKE_UP_ANCHOR, + "person_source_id", + "person_support_channel", +) + +_OUTPUT = US_SSI_TAKE_UP_OUTPUT_COLUMNS[0] +_SOURCE_ID = "person_source_id" +_SUPPORT_CHANNEL = "person_support_channel" +_ASEC_CHANNEL = "asec" +_PUF_CHANNEL = "puf_tax_detail" +_KNOWN_CHANNELS = frozenset((_ASEC_CHANNEL, _PUF_CHANNEL)) +_TARGET_PERIOD = "2024-12" +_TARGET_MEASURE = "Total with—Federal payment" +_CANDIDATE_DEFINITION = "uncapped_ssi > 0 at 2024-12" +_WEIGHTS_BASIS = "current_frame_resolved_person_weights" +_DIAGNOSTICS_SCHEMA_VERSION = 1 + + +@dataclass(frozen=True) +class SSITakeUpAgeTarget: + """One disjoint SSA federal-payment recipient age stratum.""" + + key: str + minimum_age: int | None + maximum_age: int | None + person_count: float + label: str + + def contains(self, age: np.ndarray) -> np.ndarray: + result = np.ones(len(age), dtype=bool) + if self.minimum_age is not None: + result &= age >= self.minimum_age + if self.maximum_age is not None: + result &= age <= self.maximum_age + return result + + +US_SSI_TAKE_UP_AGE_TARGETS: tuple[SSITakeUpAgeTarget, ...] = ( + SSITakeUpAgeTarget("under_18", None, 17, 1_001_922.0, "Under 18"), + SSITakeUpAgeTarget("18_64", 18, 64, 3_905_779.0, "Ages 18–64"), + SSITakeUpAgeTarget("65_plus", 65, None, 2_382_142.0, "Ages 65+"), +) +_TARGETS_BY_KEY = {target.key: target for target in US_SSI_TAKE_UP_AGE_TARGETS} +_TARGET_TOTAL = float(sum(target.person_count for target in US_SSI_TAKE_UP_AGE_TARGETS)) + +_READ_PARAMETERS: dict[str, object] = { + "table": "person", + "weight": "person_weight", +} +_ASSIGN_PARAMETERS: dict[str, object] = { + "output": _OUTPUT, + "draw": "stable_source_person_draw", + "rate_key": "ssi_age_band_count_prior", + "rate_column": "ssi_take_up_assignment_prior", + "reported_true_anchor": "SSI_VAL > 0", + "assignment_unit": _SOURCE_ID, + "fan_to_support_clones": True, +} +_CALIBRATE_PARAMETERS: dict[str, object] = { + "variable": _OUTPUT, + "targets": [US_SSI_TAKE_UP_TARGET_TABLE_NAME], + "preserve_true_anchors": True, + "preserve_true_anchor": "SSI_VAL > 0", + "domain": "uncapped_ssi > 0", + "weight": "person_weight", + "draw": "stable_source_person_draw", + "calibration_unit": _SOURCE_ID, + "age_bands": { + "under_18": "age < 18", + "18_64": "18 <= age < 65", + "65_plus": "age >= 65", + }, + "target_source": SSI_TAKE_UP_SSA_SOURCE_URL, + "target_period": _TARGET_PERIOD, + "target_measure": _TARGET_MEASURE, + "target_values": { + target.key: int(target.person_count) for target in US_SSI_TAKE_UP_AGE_TARGETS + }, + "aggregate_target": int(_TARGET_TOTAL), +} + + +def us_ssi_take_up_stage_spec() -> SourceStageSpec: + """Load and strictly validate the packaged source-stage declaration.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + stage_map = manifest.stage_map() + if US_SSI_TAKE_UP_STAGE_NAME not in stage_map: + raise ValueError( + f"US source manifest declares no {US_SSI_TAKE_UP_STAGE_NAME!r} stage." + ) + spec = stage_map[US_SSI_TAKE_UP_STAGE_NAME] + if spec.grain != "person": + raise ValueError("US SSI take-up stage must have person grain.") + if tuple(spec.outputs) != US_SSI_TAKE_UP_OUTPUT_COLUMNS: + raise ValueError("US SSI take-up manifest outputs drifted from runtime.") + expected_kinds = [ + "read_table", + "assign_binary_from_rate", + "calibrate_binary_assignment", + ] + if [operation.kind for operation in spec.operations] != expected_kinds: + raise ValueError( + "US SSI take-up stage must declare read, assignment, then age-count " + "calibration." + ) + expected_parameters = ( + _READ_PARAMETERS, + _ASSIGN_PARAMETERS, + _CALIBRATE_PARAMETERS, + ) + for operation, expected in zip(spec.operations, expected_parameters, strict=True): + if dict(operation.parameters) != expected: + raise ValueError( + f"US SSI take-up {operation.kind} contract drifted from runtime." + ) + target_artifacts = [ + artifact + for artifact in spec.artifacts + if artifact.get("source") == SSI_TAKE_UP_SSA_SOURCE_URL + ] + if len(target_artifacts) != 1: + raise ValueError("US SSI take-up stage must pin exactly one SSA target table.") + values = target_artifacts[0].get("target_values") + expected_values = { + **{ + target.key: int(target.person_count) + for target in US_SSI_TAKE_UP_AGE_TARGETS + }, + "total": int(_TARGET_TOTAL), + } + if values != expected_values: + raise ValueError("US SSI take-up SSA target cells drifted from runtime.") + return spec + + +def _decoded_strings(values: pd.Series) -> pd.Series: + return values.map( + lambda value: ( + value.decode() if isinstance(value, (bytes, np.bytes_)) else str(value) + ) + ) + + +def _age_band_values(age: np.ndarray) -> np.ndarray: + bands = np.full(len(age), "", dtype=object) + for target in US_SSI_TAKE_UP_AGE_TARGETS: + bands[target.contains(age)] = target.key + if (bands == "").any(): # pragma: no cover - disjoint bands cover finite ages + raise ValueError("US SSI take-up could not classify every age.") + return bands + + +def _normalize_targets(targets: Mapping[str, float] | None) -> dict[str, float]: + expected_keys = tuple(target.key for target in US_SSI_TAKE_UP_AGE_TARGETS) + if targets is None: + return { + target.key: float(target.person_count) + for target in US_SSI_TAKE_UP_AGE_TARGETS + } + actual = {str(key): float(value) for key, value in targets.items()} + if set(actual) != set(expected_keys): + raise ValueError( + "US SSI take-up targets require exactly age bands " + f"{list(expected_keys)}; got {sorted(actual)}." + ) + invalid = { + key: value + for key, value in actual.items() + if not np.isfinite(value) or value <= 0 + } + if invalid: + raise ValueError( + f"US SSI take-up targets must be finite and positive: {invalid}." + ) + return {key: actual[key] for key in expected_keys} + + +def _stable_source_draw(source_id: str, *, seed: int) -> float: + value = int.from_bytes( + hashlib.blake2b( + f"{seed}:{_OUTPUT}:{source_id}".encode(), + digest_size=8, + ).digest(), + byteorder="big", + signed=False, + ) + return value / float(2**64) + + +def us_ssi_take_up_reporter_source_ids(frame: Frame) -> frozenset[str]: + """Capture direct ASEC SSI reporter lineage before support pruning.""" + + if frame.schema != US_SCHEMA: + raise ValueError("US SSI take-up requires the US schema.") + person = frame.table("person") + required = {_SOURCE_ID, _SUPPORT_CHANNEL, US_SSI_TAKE_UP_ANCHOR} + missing = sorted(required - set(person.columns)) + if missing: + raise ValueError( + f"US SSI take-up reporter lineage missing person column(s): {missing}." + ) + if person[_SOURCE_ID].isna().any() or person[_SUPPORT_CHANNEL].isna().any(): + raise ValueError("US SSI take-up reporter lineage requires provenance.") + source_ids = _decoded_strings(person[_SOURCE_ID]) + channels = _decoded_strings(person[_SUPPORT_CHANNEL]) + reported = pd.to_numeric(person[US_SSI_TAKE_UP_ANCHOR], errors="coerce").to_numpy( + dtype=np.float64 + ) + if source_ids.str.strip().eq("").any() or not np.isfinite(reported).all(): + raise ValueError( + "US SSI take-up reporter lineage requires nonblank identities and " + "finite SSI_VAL values." + ) + reporter_ids = frozenset( + source_ids[(channels == _ASEC_CHANNEL).to_numpy() & (reported > 0.0)] + ) + if not reporter_ids: + raise ValueError("US SSI take-up found no direct ASEC SSI reporters.") + return reporter_ids + + +def _source_table( + frame: Frame, + *, + uncapped_ssi: np.ndarray, + seed: int, + reporter_source_ids: Collection[str] | None = None, +) -> tuple[pd.DataFrame, pd.DataFrame]: + """Return row- and source-person-grain tables for assignment.""" + + if frame.schema != US_SCHEMA: + raise ValueError("US SSI take-up requires the US schema.") + person = frame.table("person") + missing = sorted(set(US_SSI_TAKE_UP_REQUIRED_SOURCE_COLUMNS) - set(person.columns)) + if missing: + raise ValueError(f"US SSI take-up missing person source column(s): {missing}.") + if len(uncapped_ssi) != len(person): + raise ValueError( + "US SSI take-up uncapped_ssi must align with person rows: " + f"{len(person)} rows, {len(uncapped_ssi)} values." + ) + + age = pd.to_numeric(person["age"], errors="coerce").to_numpy(dtype=np.float64) + reported = pd.to_numeric(person[US_SSI_TAKE_UP_ANCHOR], errors="coerce").to_numpy( + dtype=np.float64 + ) + potential = np.asarray(uncapped_ssi, dtype=np.float64) + weights = np.asarray(frame.resolve_weights("person").values, dtype=np.float64) + if not np.isfinite(age).all() or (age < 0).any(): + raise ValueError("US SSI take-up ages must be finite and nonnegative.") + if not np.isfinite(reported).all(): + raise ValueError("US SSI take-up SSI_VAL anchors must be finite.") + if not np.isfinite(potential).all(): + raise ValueError("US SSI take-up uncapped_ssi values must be finite.") + if not (np.isfinite(weights) & (weights >= 0)).all() or weights.sum() <= 0: + raise ValueError( + "US SSI take-up requires finite nonnegative person weights with " + "positive total." + ) + if person[_SOURCE_ID].isna().any() or person[_SUPPORT_CHANNEL].isna().any(): + raise ValueError("US SSI take-up requires complete support provenance.") + + source_ids = _decoded_strings(person[_SOURCE_ID]) + channels = _decoded_strings(person[_SUPPORT_CHANNEL]) + if source_ids.str.strip().eq("").any(): + raise ValueError("US SSI take-up source identities must be nonblank.") + observed_channels = set(channels.unique()) + if observed_channels != set(_KNOWN_CHANNELS): + raise ValueError( + "US SSI take-up requires exact ASEC/PUF support channels; " + f"missing {sorted(_KNOWN_CHANNELS - observed_channels)}, " + f"unsupported {sorted(observed_channels - _KNOWN_CHANNELS)}." + ) + + direct_anchor = (reported > 0.0) & channels.eq(_ASEC_CHANNEL).to_numpy() + if reporter_source_ids is None: + anchored_source_ids = frozenset(source_ids[direct_anchor]) + else: + anchored_source_ids = frozenset(str(value) for value in reporter_source_ids) + if not anchored_source_ids: + raise ValueError("US SSI take-up reporter lineage cannot be empty.") + if any(not value.strip() for value in anchored_source_ids): + raise ValueError("US SSI take-up reporter source identities are nonblank.") + omitted_direct = sorted(set(source_ids[direct_anchor]) - anchored_source_ids) + if omitted_direct: + raise ValueError( + "US SSI take-up reporter lineage omitted direct ASEC anchors; " + f"examples {omitted_direct[:5]}." + ) + + rows = pd.DataFrame( + { + "source_id": source_ids.to_numpy(), + "channel": channels.to_numpy(), + "age": age, + "age_band": _age_band_values(age), + "weight": weights, + "candidate": potential > 0.0, + # Capture lineage on the full support before L0. When pruning keeps + # only a PUF clone, the explicit source-ID set still preserves the + # underlying direct ASEC measurement without promoting PUF-only + # SSI_VAL copies into independent anchors. + "anchor": source_ids.isin(anchored_source_ids).to_numpy(), + }, + index=person.index, + ) + duplicated_channel = rows.duplicated(["source_id", "channel"], keep=False) + if duplicated_channel.any(): + examples = ( + rows.loc[duplicated_channel, ["source_id", "channel"]] + .drop_duplicates() + .head(5) + .to_dict("records") + ) + raise ValueError( + "US SSI take-up requires at most one row per source identity and " + f"support channel; duplicate examples {examples}." + ) + rows["candidate_weight"] = rows["weight"].where(rows["candidate"], 0.0) + + source = ( + rows.groupby("source_id", sort=True) + .agg( + age_band=("age_band", "first"), + age_band_count=("age_band", "nunique"), + age_count=("age", "nunique"), + candidate_weight=("candidate_weight", "sum"), + total_weight=("weight", "sum"), + anchor=("anchor", "any"), + row_count=("source_id", "size"), + ) + .copy() + ) + inconsistent = source.index[source["age_band_count"] != 1].tolist() + if inconsistent: + raise ValueError( + "US SSI take-up source identities cross SSA age bands; examples " + f"{inconsistent[:5]}." + ) + age_inconsistent = source.index[source["age_count"] != 1].tolist() + if age_inconsistent: + raise ValueError( + "US SSI take-up support rows disagree on source-person age; " + f"examples {age_inconsistent[:5]}." + ) + source["draw"] = [ + _stable_source_draw(str(source_id), seed=int(seed)) + for source_id in source.index + ] + return rows, source + + +def _age_band_diagnostics( + source: pd.DataFrame, + selected: pd.Series, + *, + targets: Mapping[str, float], +) -> list[dict[str, object]]: + """Summarize one existing source-grain assignment on current weights.""" + + bands: list[dict[str, object]] = [] + for target_definition in US_SSI_TAKE_UP_AGE_TARGETS: + key = target_definition.key + target = float(targets[key]) + in_band = source["age_band"].eq(key) + candidate = in_band & source["candidate_weight"].gt(0.0) + anchored = in_band & source["anchor"].astype(bool) + capacity = float(source.loc[candidate, "candidate_weight"].sum()) + reporter_floor = float( + source.loc[candidate & anchored, "candidate_weight"].sum() + ) + reachable_goal = min(max(target, reporter_floor), capacity) + selected_recipient_weight = float( + source.loc[candidate & selected, "candidate_weight"].sum() + ) + max_source_weight = ( + float(source.loc[candidate, "candidate_weight"].max()) + if candidate.any() + else 0.0 + ) + bands.append( + { + "age_band": key, + "label": target_definition.label, + "target": target, + "source_identity_count": int(in_band.sum()), + "candidate_source_identity_count": int(candidate.sum()), + "reporter_source_identity_count": int(anchored.sum()), + "candidate_capacity": capacity, + "reporter_candidate_floor": reporter_floor, + "reachable_goal": reachable_goal, + "selected_recipient_weight": selected_recipient_weight, + "signed_target_error": selected_recipient_weight - target, + "target_shortfall": max(target - selected_recipient_weight, 0.0), + "anchor_excess": max(reporter_floor - target, 0.0), + "saturated": bool(capacity < target), + "assignment_prior": ( + min(target / capacity, 1.0) if capacity > 0 else 0.0 + ), + "max_source_candidate_weight": max_source_weight, + } + ) + return bands + + +def _assign_sources( + source: pd.DataFrame, + *, + targets: Mapping[str, float], +) -> tuple[pd.Series, list[dict[str, object]]]: + """Assign one flag per source identity and return age-band diagnostics.""" + + selected = pd.Series(False, index=source.index, dtype=bool) + for target_definition in US_SSI_TAKE_UP_AGE_TARGETS: + key = target_definition.key + target = float(targets[key]) + in_band = source["age_band"].eq(key) + candidate = in_band & source["candidate_weight"].gt(0.0) + anchored = in_band & source["anchor"].astype(bool) + capacity = float(source.loc[candidate, "candidate_weight"].sum()) + reporter_floor = float( + source.loc[candidate & anchored, "candidate_weight"].sum() + ) + reachable_goal = min(max(target, reporter_floor), capacity) + prior = min(target / capacity, 1.0) if capacity > 0 else 0.0 + + # Keep reform propensities off today's candidate domain. Candidate + # decisions are then greedily count-calibrated below. + selected.loc[in_band] = source.loc[in_band, "draw"].to_numpy() < prior + selected.loc[anchored] = True + selectable = candidate & ~anchored + selected.loc[selectable] = False + + current = reporter_floor + if capacity < target: + # Select the saturated domain explicitly instead of depending on + # floating-point accumulation reaching the same capacity sum. + selected.loc[candidate] = True + else: + ordered = sorted( + source.index[selectable], + key=lambda source_id: ( + float(source.at[source_id, "draw"]), + str(source_id), + ), + ) + for source_id in ordered: + if current >= reachable_goal: + break + selected.at[source_id] = True + current += float(source.at[source_id, "candidate_weight"]) + + return selected, _age_band_diagnostics(source, selected, targets=targets) + + +def _diagnostics( + frame: Frame, + *, + rows: pd.DataFrame, + source: pd.DataFrame, + assigned: np.ndarray, + bands: list[dict[str, object]], + targets: Mapping[str, float], +) -> dict[str, object]: + weights = np.asarray(frame.resolve_weights("person").values, dtype=np.float64) + reporter_lost = int(np.count_nonzero(rows["anchor"].to_numpy() & ~assigned)) + by_source = pd.DataFrame( + {"source_id": rows["source_id"].to_numpy(), "assigned": assigned} + ) + mismatches = int( + (by_source.groupby("source_id", sort=False)["assigned"].nunique() > 1).sum() + ) + channel_diagnostics: dict[str, dict[str, object]] = {} + for channel in sorted(_KNOWN_CHANNELS): + mask = rows["channel"].eq(channel).to_numpy() + channel_weight = float(weights[mask].sum()) + channel_diagnostics[channel] = { + "rows": int(mask.sum()), + "unique_count": int(np.unique(assigned[mask]).size), + "weighted_true_share": ( + float(weights[mask & assigned].sum()) / channel_weight + if channel_weight > 0 + else 0.0 + ), + } + total_weight = float(weights.sum()) + selected_total = float( + sum(float(band["selected_recipient_weight"]) for band in bands) + ) + target_total = float(sum(targets.values())) + return { + "schema_version": _DIAGNOSTICS_SCHEMA_VERSION, + "classification": "release_diagnostics", + "issues": ["PolicyEngine/populace#312"], + "variable": _OUTPUT, + "anchor": US_SSI_TAKE_UP_ANCHOR, + "anchor_channel": _ASEC_CHANNEL, + "candidate_definition": _CANDIDATE_DEFINITION, + "target_table": US_SSI_TAKE_UP_TARGET_TABLE_NAME, + "target_source": SSI_TAKE_UP_SSA_SOURCE_URL, + "target_period": _TARGET_PERIOD, + "target_measure": _TARGET_MEASURE, + "weights_basis": _WEIGHTS_BASIS, + "target_total": target_total, + "selected_recipient_weight_total": selected_total, + "target_shortfall_total": float( + sum(float(band["target_shortfall"]) for band in bands) + ), + "source_identity_count": int(len(source)), + "weighted_flag_true_count": float(weights[assigned].sum()), + "weighted_flag_universe": total_weight, + "weighted_flag_true_share": ( + float(weights[assigned].sum()) / total_weight if total_weight > 0 else 0.0 + ), + "unique_count": int(np.unique(assigned).size), + "missing_or_invalid_count": 0, + "reporter_anchor_lost_count": reporter_lost, + "source_identity_mismatch_count": mismatches, + "channel_diagnostics": channel_diagnostics, + "age_bands": bands, + } + + +def with_us_ssi_take_up( + frame: Frame, + *, + uncapped_ssi: np.ndarray, + seed: int, + targets: Mapping[str, float] | None = None, + reporter_source_ids: Collection[str] | None = None, +) -> tuple[Frame, dict[str, object]]: + """Recompute SSI take-up and return the frame plus count diagnostics.""" + + us_ssi_take_up_stage_spec() + normalized_targets = _normalize_targets(targets) + rows, source = _source_table( + frame, + uncapped_ssi=np.asarray(uncapped_ssi, dtype=np.float64), + seed=int(seed), + reporter_source_ids=reporter_source_ids, + ) + selected, bands = _assign_sources(source, targets=normalized_targets) + assigned = rows["source_id"].map(selected).to_numpy(dtype=bool) + diagnostics = _diagnostics( + frame, + rows=rows, + source=source, + assigned=assigned, + bands=bands, + targets=normalized_targets, + ) + + person = frame.table("person") + if _OUTPUT in person: + current = pd.to_numeric(person[_OUTPUT], errors="coerce").to_numpy( + dtype=np.float64 + ) + if ( + pd.api.types.is_bool_dtype(person[_OUTPUT].dtype) + and not person[_OUTPUT].isna().any() + and np.isfinite(current).all() + and np.isin(current, [0.0, 1.0]).all() + and np.array_equal(current.astype(bool), assigned) + ): + return frame, diagnostics + + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + tables["person"][_OUTPUT] = assigned + result = Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + ) + return result, diagnostics + + +def us_ssi_take_up_diagnostics( + frame: Frame, + *, + uncapped_ssi: np.ndarray, + seed: int, + targets: Mapping[str, float] | None = None, + reporter_source_ids: Collection[str] | None = None, +) -> dict[str, object]: + """Diagnose a persisted assignment without changing its decisions. + + The fiscal builder uses this after its dependency-aware reconciliation + calibration: the assignment must remain count-faithful on returned release + weights, while the exact flags used to materialize SSI and Medicaid target + vectors must not be mutated after optimization. + """ + + us_ssi_take_up_stage_spec() + normalized_targets = _normalize_targets(targets) + person = frame.table("person") + if _OUTPUT not in person: + raise ValueError(f"US SSI take-up diagnostics require person.{_OUTPUT}.") + output = person[_OUTPUT] + if not pd.api.types.is_bool_dtype(output.dtype) or output.isna().any(): + raise ValueError("US SSI take-up diagnostics require a complete boolean flag.") + assigned = output.to_numpy(dtype=bool) + rows, source = _source_table( + frame, + uncapped_ssi=np.asarray(uncapped_ssi, dtype=np.float64), + seed=int(seed), + reporter_source_ids=reporter_source_ids, + ) + source_assignment = ( + pd.DataFrame({"source_id": rows["source_id"].to_numpy(), "assigned": assigned}) + .groupby("source_id", sort=True)["assigned"] + .first() + .reindex(source.index) + .astype(bool) + ) + bands = _age_band_diagnostics( + source, + source_assignment, + targets=normalized_targets, + ) + return _diagnostics( + frame, + rows=rows, + source=source, + assigned=assigned, + bands=bands, + targets=normalized_targets, + ) + + +def us_ssi_take_up_gate( + diagnostics: Mapping[str, object], + *, + targets: Mapping[str, float] | None = None, +) -> GateResult: + """Require source-faithful anchors and count calibration by age.""" + + expected_targets = _normalize_targets(targets) + failures: list[str] = [] + if diagnostics.get("schema_version") != _DIAGNOSTICS_SCHEMA_VERSION: + failures.append("SSI take-up diagnostics schema version is invalid.") + if diagnostics.get("classification") != "release_diagnostics": + failures.append("SSI take-up diagnostics classification is invalid.") + if diagnostics.get("variable") != _OUTPUT: + failures.append("SSI take-up diagnostics name the wrong output variable.") + if diagnostics.get("anchor") != US_SSI_TAKE_UP_ANCHOR: + failures.append("SSI take-up diagnostics name the wrong reporter anchor.") + if diagnostics.get("anchor_channel") != _ASEC_CHANNEL: + failures.append("SSI take-up diagnostics name the wrong anchor channel.") + if diagnostics.get("candidate_definition") != _CANDIDATE_DEFINITION: + failures.append("SSI take-up diagnostics carry the wrong candidate definition.") + if diagnostics.get("target_table") != US_SSI_TAKE_UP_TARGET_TABLE_NAME: + failures.append("SSI take-up diagnostics carry the wrong SSA target table.") + if diagnostics.get("target_source") != SSI_TAKE_UP_SSA_SOURCE_URL: + failures.append("SSI take-up diagnostics carry the wrong SSA target source.") + if diagnostics.get("target_period") != _TARGET_PERIOD: + failures.append("SSI take-up diagnostics carry the wrong target period.") + if diagnostics.get("target_measure") != _TARGET_MEASURE: + failures.append("SSI take-up diagnostics carry the wrong target measure.") + if diagnostics.get("weights_basis") != _WEIGHTS_BASIS: + failures.append("SSI take-up diagnostics carry the wrong weights basis.") + if int(diagnostics.get("missing_or_invalid_count", -1)) != 0: + failures.append("SSI take-up output contains missing or invalid values.") + if int(diagnostics.get("unique_count", 0)) < 2: + failures.append("SSI take-up output is constant and carries no signal.") + if int(diagnostics.get("reporter_anchor_lost_count", -1)) != 0: + failures.append("SSI take-up lost one or more direct ASEC reporter anchors.") + if int(diagnostics.get("source_identity_mismatch_count", -1)) != 0: + failures.append("SSI take-up support rows disagree within a source identity.") + + channels = diagnostics.get("channel_diagnostics") + if not isinstance(channels, Mapping): + failures.append("SSI take-up channel diagnostics are missing.") + channels = {} + if set(channels) != set(_KNOWN_CHANNELS): + failures.append( + "SSI take-up requires exact ASEC/PUF channel diagnostics; got " + f"{sorted(channels)}." + ) + for channel in sorted(_KNOWN_CHANNELS): + values = channels.get(channel) + if not isinstance(values, Mapping): + continue + if int(values.get("rows", 0)) <= 0: + failures.append(f"SSI take-up channel {channel!r} has no rows.") + if int(values.get("unique_count", 0)) < 2: + failures.append(f"SSI take-up channel {channel!r} is constant.") + share = float(values.get("weighted_true_share", np.nan)) + if not np.isfinite(share) or not 0 < share < 1: + failures.append( + f"SSI take-up channel {channel!r} has an invalid weighted share." + ) + + age_bands = diagnostics.get("age_bands") + if not isinstance(age_bands, list): + failures.append("SSI take-up age-band diagnostics are missing.") + age_bands = [] + if len(age_bands) != len(expected_targets) or not all( + isinstance(row, Mapping) for row in age_bands + ): + failures.append( + "SSI take-up requires exactly one diagnostic row per SSA age band." + ) + by_key = { + str(row.get("age_band")): row for row in age_bands if isinstance(row, Mapping) + } + if set(by_key) != set(expected_targets): + failures.append( + "SSI take-up age bands drifted from the three disjoint SSA strata." + ) + for key, target in expected_targets.items(): + row = by_key.get(key) + if row is None: + continue + recorded_target = float(row.get("target", np.nan)) + if not np.isclose(recorded_target, target, rtol=0.0, atol=1e-6): + failures.append( + f"SSI take-up age band {key!r} target {recorded_target} does " + f"not match {target}." + ) + continue + capacity = float(row.get("candidate_capacity", np.nan)) + floor = float(row.get("reporter_candidate_floor", np.nan)) + reachable = float(row.get("reachable_goal", np.nan)) + selected = float(row.get("selected_recipient_weight", np.nan)) + signed_error = float(row.get("signed_target_error", np.nan)) + shortfall = float(row.get("target_shortfall", np.nan)) + anchor_excess = float(row.get("anchor_excess", np.nan)) + prior = float(row.get("assignment_prior", np.nan)) + max_weight = float(row.get("max_source_candidate_weight", np.nan)) + numeric_values = ( + capacity, + floor, + reachable, + selected, + signed_error, + shortfall, + anchor_excess, + prior, + max_weight, + ) + if not all(np.isfinite(value) for value in numeric_values): + failures.append(f"SSI take-up age band {key!r} has nonfinite diagnostics.") + continue + if capacity <= 0: + failures.append( + f"SSI take-up age band {key!r} has zero candidate capacity." + ) + continue + if floor < 0 or floor > capacity + 1e-6: + failures.append( + f"SSI take-up age band {key!r} has an invalid anchor floor." + ) + continue + goal = min(max(target, floor), capacity) + epsilon = max(1e-6, np.finfo(np.float64).eps * max(goal, 1.0) * 16.0) + allowance = max(max_weight, epsilon) + if abs(reachable - goal) > epsilon: + failures.append( + f"SSI take-up age band {key!r} carries the wrong reachable goal." + ) + if selected < -epsilon or selected > capacity + epsilon: + failures.append( + f"SSI take-up age band {key!r} selected count is outside capacity." + ) + if abs(selected - goal) > allowance + epsilon: + failures.append( + f"SSI take-up age band {key!r} selected weight {selected:.3f} " + f"misses reachable goal {goal:.3f} by more than one source-" + f"identity weight ({allowance:.3f})." + ) + if abs(signed_error - (selected - target)) > epsilon: + failures.append( + f"SSI take-up age band {key!r} carries the wrong signed error." + ) + if abs(shortfall - max(target - selected, 0.0)) > epsilon: + failures.append( + f"SSI take-up age band {key!r} carries the wrong target shortfall." + ) + if abs(anchor_excess - max(floor - target, 0.0)) > epsilon: + failures.append( + f"SSI take-up age band {key!r} carries the wrong anchor excess." + ) + expected_prior = min(target / capacity, 1.0) + if abs(prior - expected_prior) > epsilon: + failures.append( + f"SSI take-up age band {key!r} carries the wrong assignment prior." + ) + saturated = bool(row.get("saturated")) + if saturated != (capacity < target): + failures.append( + f"SSI take-up age band {key!r} saturation status is inconsistent." + ) + if saturated: + if abs(selected - capacity) > epsilon: + failures.append( + f"SSI take-up saturated age band {key!r} did not select " + "every candidate." + ) + if shortfall <= 0: + failures.append( + f"SSI take-up saturated age band {key!r} hides its shortfall." + ) + + expected_total = float(sum(expected_targets.values())) + recorded_total = float(diagnostics.get("target_total", np.nan)) + if not np.isclose(recorded_total, expected_total, rtol=0.0, atol=1e-6): + failures.append( + "SSI take-up aggregate target is not the sum of its disjoint age bands." + ) + selected_total = float( + sum( + float(row.get("selected_recipient_weight", np.nan)) + for row in by_key.values() + ) + ) + recorded_selected_total = float( + diagnostics.get("selected_recipient_weight_total", np.nan) + ) + if not np.isclose( + recorded_selected_total, + selected_total, + rtol=0.0, + atol=1e-6, + ): + failures.append("SSI take-up aggregate selected count drifted from age bands.") + shortfall_total = float( + sum(float(row.get("target_shortfall", np.nan)) for row in by_key.values()) + ) + recorded_shortfall_total = float(diagnostics.get("target_shortfall_total", np.nan)) + if not np.isclose( + recorded_shortfall_total, + shortfall_total, + rtol=0.0, + atol=1e-6, + ): + failures.append("SSI take-up aggregate shortfall drifted from age bands.") + return GateResult( + name="ssi_take_up", + passed=not failures, + failures=tuple(failures), + details=dict(diagnostics), + ) + + +def write_us_ssi_take_up_diagnostics( + diagnostics: Mapping[str, object], path: str | Path +) -> Path: + """Write final-weight SSI take-up diagnostics as strict JSON.""" + + output = Path(path) + output.parent.mkdir(parents=True, exist_ok=True) + output.write_text( + json.dumps(dict(diagnostics), indent=2, sort_keys=True, allow_nan=False) + "\n", + encoding="utf-8", + ) + return output diff --git a/packages/populace-build/src/populace/build/us_runtime/validation_input_coverage.py b/packages/populace-build/src/populace/build/us_runtime/validation_input_coverage.py index 991681a1..5738b367 100644 --- a/packages/populace-build/src/populace/build/us_runtime/validation_input_coverage.py +++ b/packages/populace-build/src/populace/build/us_runtime/validation_input_coverage.py @@ -10,13 +10,13 @@ validation row reads as a silent zero until an external benchmark exposes it. Two shipped instances: -- education credits (``soi_baseline_levels`` row ``soi_education_credits``, - ``tax_expenditure_reforms`` has no education row) depend on - ``qualified_tuition_expenses``, populated on 641 of 160,858 person records — - education credits validate ~40% low (populace#253); and +- education credits (``soi_baseline_levels`` row ``soi_education_credits``) + depend on ``qualified_tuition_expenses``; the PUF education stage now + produces that leaf and its five AOTC factual inputs (populace#253); and - the OBBBA auto-loan-interest deduction depends on - ``qualified_passenger_vehicle_loan_interest``, never imputed, so the - provision is structurally $0 in reform validation (populace#252). + ``qualified_passenger_vehicle_loan_interest``; the SCF auto-loan stage now + derives a documented qualifying-share proxy so the provision binds + (populace#252). Both were invisible because nothing checked that a validation row's inputs are actually produced. :func:`us_validation_input_coverage_gate` is that check. @@ -104,32 +104,70 @@ def __post_init__(self) -> None: #: leaf. Each entry is proven against the live engine graph by #: :func:`assert_validation_leaf_registry_current`. US_VALIDATION_PROVISION_INPUT_LEAVES: tuple[ValidationInputLeaf, ...] = ( + ValidationInputLeaf( + leaf="tip_income", + provision_variables=("tip_income_deduction",), + validation_rows=("obbba_no_tax_on_tips",), + ), + ValidationInputLeaf( + leaf="treasury_tipped_occupation_code", + provision_variables=("tip_income_deduction",), + validation_rows=("obbba_no_tax_on_tips",), + ), + ValidationInputLeaf( + leaf="fsla_overtime_premium", + provision_variables=("overtime_income_deduction",), + validation_rows=("obbba_no_tax_on_overtime",), + ), ValidationInputLeaf( leaf="qualified_tuition_expenses", provision_variables=("education_tax_credits",), validation_rows=("soi_education_credits",), - # #253: populated on 641/160,858 person records; education credits - # validate ~40% below the IRS SOI actual. A reviewed exclusion until - # the leaf is imputed from enrollment + published tuition - # distributions (NCES/IPEDS) or education claims are calibrated. - reason=( - "Qualifying tuition input essentially unimputed — 641/160,858 person " - "records (PolicyEngine/populace#253); education credits validate ~40% " - "low until it is imputed or education claims are calibrated." - ), + ), + ValidationInputLeaf( + leaf="traditional_401k_contributions_desired", + provision_variables=("savers_credit",), + validation_rows=("soi_savers_credit",), + ), + ValidationInputLeaf( + leaf="roth_401k_contributions_desired", + provision_variables=("savers_credit",), + validation_rows=("soi_savers_credit",), + ), + ValidationInputLeaf( + leaf="traditional_ira_contributions_desired", + provision_variables=("savers_credit",), + validation_rows=("soi_savers_credit",), + ), + ValidationInputLeaf( + leaf="roth_ira_contributions_desired", + provision_variables=("savers_credit",), + validation_rows=("soi_savers_credit",), + ), + ValidationInputLeaf( + leaf="self_employed_pension_contributions_desired", + provision_variables=("savers_credit",), + validation_rows=("soi_savers_credit",), + ), + ValidationInputLeaf( + leaf="casualty_loss", + provision_variables=("casualty_loss_deduction",), + validation_rows=("obbba_casualty_loss_limit",), + ), + ValidationInputLeaf( + leaf="unreimbursed_business_employee_expenses", + provision_variables=("total_misc_deductions",), + validation_rows=("obbba_misc_itemized_deductions",), + ), + ValidationInputLeaf( + leaf="spm_unit_pre_subsidy_childcare_expenses", + provision_variables=("cdcc",), + validation_rows=("te_cdcc", "obbba_cdcc", "soi_cdcc"), ), ValidationInputLeaf( leaf="qualified_passenger_vehicle_loan_interest", provision_variables=("auto_loan_interest_deduction",), validation_rows=("obbba_auto_loan_interest",), - # #252: never imputed, so the OBBBA auto-loan deduction is structurally - # $0. A reviewed exclusion until the qualifying share of the existing - # auto_loan_interest is imputed or the deduction is calibrated. - reason=( - "OBBBA qualifying auto-loan interest never imputed — the auto-loan " - "deduction is structurally $0 (PolicyEngine/populace#252); reviewed " - "until the qualifying share of auto_loan_interest is imputed." - ), ), ) diff --git a/packages/populace-build/src/populace/build/us_runtime/voluntary_filing.py b/packages/populace-build/src/populace/build/us_runtime/voluntary_filing.py new file mode 100644 index 00000000..2e5854c3 --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/voluntary_filing.py @@ -0,0 +1,1055 @@ +"""Measured SIPP filing propensity for the voluntary-filer input. + +The retired enhanced-CPS pipeline seeded +``would_file_taxes_voluntarily`` from a synthetic demographic rate table. Its +three dimensions were tax-unit wages, children, and head age. The full 2023 +SIPP donor already pinned for the vehicle stage contains directly reported +federal filing and expected-filing answers, so this stage replaces the +synthetic rate draw with measured survey signal while preserving those three +dimensions. + +The source transform is deliberately strict. It keeps December responses +only when ``AFILING == 1`` (the filing answer is reported) and, for people who +have not filed, ``AWILLFILE == 1`` (the expected-filing answer is reported). +The target is ``EFILING == 1 or EWILLFILE == 1``. Respondents reported as +claimed dependents (``EDEPCLM == 1``) are not standalone tax units. Reciprocal +``EPNSPOUSE`` pairs are collapsed once, and disagreeing spouse targets fail +closed. Unit earnings are the annualized sum of the seven SIPP monthly job- +earnings recodes; age and sex come from the minimum-PNUM reference member; +marriage is the reciprocal-pair fact; and under-18 count is household context +computed before response filtering. No parent pointer is used to invent a +federal dependent attachment. + +A weighted QRF is fit at source-tax-unit grain. On an expanded support frame, +the ASEC member of each ``tax_unit_source_id`` supplies predictors, the model +draws once, and the result fans out to every clone. Re-running with the same +donor and build seed recomputes the same surface rather than trusting an +arbitrary pre-existing nonconstant column. +""" + +from __future__ import annotations + +from pathlib import Path +from typing import Any, BinaryIO + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import SourceStageSpec, load_source_manifest +from populace.frame import Frame +from populace.frame.units import US_SCHEMA + +__all__ = [ + "SIPP_2023_VOLUNTARY_FILING_DONOR_REVISION", + "SIPP_2023_VOLUNTARY_FILING_DONOR_SHA256", + "SIPP_2023_VOLUNTARY_FILING_DONOR_SIZE_BYTES", + "SIPP_2023_VOLUNTARY_FILING_DONOR_URL", + "SIPP_VOLUNTARY_FILING_MODEL_PREDICTORS", + "SIPP_VOLUNTARY_FILING_SOURCE_COLUMNS", + "US_VOLUNTARY_FILING_NONCONSTANT_TAX_UNIT_COLUMNS", + "US_VOLUNTARY_FILING_OUTPUT_COLUMNS", + "US_VOLUNTARY_FILING_STAGE_NAME", + "VOLUNTARY_FILING_ARCHIVED_DERIVATION_URL", + "VOLUNTARY_FILING_ARCHIVED_PARAMETERS_URL", + "VOLUNTARY_FILING_SIPP_DICTIONARY_URL", + "fetch_sipp_2023_voluntary_filing_donor", + "impute_us_voluntary_filing", + "load_sipp_2023_voluntary_filing_donor", + "us_voluntary_filing_signal_gate", + "us_voluntary_filing_stage_spec", + "us_voluntary_filing_summary", + "with_us_voluntary_filing_input", +] + +QRF: Any | None = None + +_ARCHIVED_COMMIT = "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe" +_RETIRED_REPOSITORY = "policyengine-" + "us-data" +_RETIRED_PACKAGE = "policyengine_" + "us_data" +_ARCHIVED_ROOT = ( + f"https://github.com/PolicyEngine/{_RETIRED_REPOSITORY}/blob/" + f"{_ARCHIVED_COMMIT}/{_RETIRED_PACKAGE}/" +) +VOLUNTARY_FILING_ARCHIVED_DERIVATION_URL = ( + _ARCHIVED_ROOT + "datasets/cps/cps.py#L726-L747" +) +VOLUNTARY_FILING_ARCHIVED_PARAMETERS_URL = ( + _ARCHIVED_ROOT + "parameters/take_up/voluntary_filing.yaml#L1-L43" +) +VOLUNTARY_FILING_SIPP_DICTIONARY_URL = ( + "https://www2.census.gov/programs-surveys/sipp/tech-documentation/" + "data-dictionaries/2023/2023_SIPP_Data_Dictionary.pdf" +) + +SIPP_2023_VOLUNTARY_FILING_DONOR_REVISION = "21280dca5995e978d706740a8a4b9b7860cfd7b6" +SIPP_2023_VOLUNTARY_FILING_DONOR_SHA256 = ( + "5c30439e365fc26483318ef61d1d8f4bb2f0e9d6bb47c22c06756a7698733ee2" +) +SIPP_2023_VOLUNTARY_FILING_DONOR_SIZE_BYTES = 3_726_010_471 +SIPP_2023_VOLUNTARY_FILING_DONOR_URL = ( + f"https://huggingface.co/policyengine/{_RETIRED_REPOSITORY}/resolve/" + f"{SIPP_2023_VOLUNTARY_FILING_DONOR_REVISION}/pu2023.csv" +) + +US_VOLUNTARY_FILING_STAGE_NAME = "voluntary_filing_input" +US_VOLUNTARY_FILING_OUTPUT_COLUMNS: tuple[str, ...] = ("would_file_taxes_voluntarily",) +US_VOLUNTARY_FILING_NONCONSTANT_TAX_UNIT_COLUMNS = US_VOLUNTARY_FILING_OUTPUT_COLUMNS + +_JOB_MONTHLY_EARNINGS_COLUMNS = tuple(f"TJB{i}_MSUM" for i in range(1, 8)) +SIPP_VOLUNTARY_FILING_SOURCE_COLUMNS: tuple[str, ...] = ( + "SSUID", + "PNUM", + "MONTHCODE", + "WPFINWGT", + "TAGE", + "ESEX", + "EPNSPOUSE", + "AFILING", + "EFILING", + "AWILLFILE", + "EWILLFILE", + "EDEPCLM", + *_JOB_MONTHLY_EARNINGS_COLUMNS, +) +SIPP_VOLUNTARY_FILING_MODEL_PREDICTORS: tuple[str, ...] = ( + "employment_income", + "reference_age", + "reference_is_female", + "reference_is_married", + "count_under_18", +) + +_OUTPUT = US_VOLUNTARY_FILING_OUTPUT_COLUMNS[0] +_DONOR_FILENAME = "pu2023.csv" +_DONOR_WEIGHT_COLUMN = "tax_unit_weight" +_DONOR_SOURCE_KEY_COLUMN = "source_tax_unit_key" +_DEFAULT_N_ESTIMATORS = 100 +_BASE_ASEC_SUPPORT_CHANNEL = "asec" +_TAX_UNIT_SUPPORT_CHANNEL_COLUMN = "tax_unit_support_channel" +_TAX_UNIT_SOURCE_ID_COLUMN = "tax_unit_source_id" +_TRUE_SHARE_BAND = (0.45, 0.95) +_EXPECTED_READ_PARAMETERS: dict[str, object] = { + "table": "sipp_person", + "delimiter": "|", + "month_column": "MONTHCODE", + "month": 12, + "source_columns": list(SIPP_VOLUNTARY_FILING_SOURCE_COLUMNS), +} +_EXPECTED_FIT_PARAMETERS: dict[str, object] = { + "predictors": list(SIPP_VOLUNTARY_FILING_MODEL_PREDICTORS), + "target": _OUTPUT, + "weight": _DONOR_WEIGHT_COLUMN, + "response_filter": ( + "AFILING == 1 and (EFILING == 1 or (EFILING == 2 and AWILLFILE == 1))" + ), + "dependent_exclusion": "EDEPCLM == 1", + "canonical_unit": ( + "SSUID plus the sorted PNUM/EPNSPOUSE pair for reciprocal spouses; " + "otherwise PNUM" + ), + "n_estimators": _DEFAULT_N_ESTIMATORS, + "seed_from_build_config": True, +} + + +def us_voluntary_filing_stage_spec() -> SourceStageSpec: + """Load and validate the packaged voluntary-filing stage declaration.""" + + from importlib.resources import files + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + stage_map = manifest.stage_map() + if US_VOLUNTARY_FILING_STAGE_NAME not in stage_map: + raise ValueError( + f"US source manifest declares no {US_VOLUNTARY_FILING_STAGE_NAME!r} stage." + ) + spec = stage_map[US_VOLUNTARY_FILING_STAGE_NAME] + if spec.grain != "tax_unit": + raise ValueError("US voluntary-filing stage must have tax_unit grain.") + if tuple(spec.outputs) != US_VOLUNTARY_FILING_OUTPUT_COLUMNS: + raise ValueError( + "US voluntary-filing manifest outputs do not match the runtime-owned " + "family." + ) + if [operation.kind for operation in spec.operations] != [ + "read_table", + "fit_weighted_qrf", + ]: + raise ValueError( + "US voluntary-filing stage must contain read_table then fit_weighted_qrf." + ) + if dict(spec.operations[0].parameters) != _EXPECTED_READ_PARAMETERS: + raise ValueError( + "US voluntary-filing read_table contract drifted from the pinned " + "SIPP transform." + ) + if dict(spec.operations[1].parameters) != _EXPECTED_FIT_PARAMETERS: + raise ValueError( + "US voluntary-filing QRF contract drifted from the reviewed source " + "semantics." + ) + matching_artifacts = [ + artifact + for artifact in spec.artifacts + if artifact.get("sha256") == SIPP_2023_VOLUNTARY_FILING_DONOR_SHA256 + and artifact.get("size_bytes") == SIPP_2023_VOLUNTARY_FILING_DONOR_SIZE_BYTES + ] + if not matching_artifacts: + raise ValueError( + "US voluntary-filing stage does not pin the full 2023 SIPP donor " + "SHA-256 and byte length." + ) + return spec + + +def _sha256_stream(stream: BinaryIO, *, chunk_size: int = 8 * 1024 * 1024) -> str: + import hashlib + + digest = hashlib.sha256() + for chunk in iter(lambda: stream.read(chunk_size), b""): + digest.update(chunk) + return digest.hexdigest() + + +def _sha256_file(path: Path) -> str: + with path.open("rb") as stream: + return _sha256_stream(stream) + + +def _file_matches( + path: Path, + *, + expected_sha256: str | None, + expected_size_bytes: int | None, +) -> bool: + if not path.is_file(): + return False + if expected_size_bytes is not None and path.stat().st_size != expected_size_bytes: + return False + return expected_sha256 is None or _sha256_file(path) == expected_sha256 + + +def fetch_sipp_2023_voluntary_filing_donor( + cache_dir: str | Path | None = None, + *, + expected_sha256: str | None = SIPP_2023_VOLUNTARY_FILING_DONOR_SHA256, + expected_size_bytes: int | None = SIPP_2023_VOLUNTARY_FILING_DONOR_SIZE_BYTES, + chunk_size: int = 8 * 1024 * 1024, +) -> Path: + """Stream, verify, and atomically cache the pinned 3.73 GB SIPP file.""" + + if chunk_size < 1: + raise ValueError("chunk_size must be a positive integer") + + import hashlib + import urllib.request + + root = ( + Path(cache_dir).expanduser() + if cache_dir is not None + else Path.home() / ".cache" / "populace" / "sipp" + ) + if cache_dir is None: + snapshot = ( + Path.home() + / ".cache" + / "huggingface" + / "hub" + / f"models--policyengine--{_RETIRED_REPOSITORY}" + / "snapshots" + / SIPP_2023_VOLUNTARY_FILING_DONOR_REVISION + / _DONOR_FILENAME + ) + if _file_matches( + snapshot, + expected_sha256=expected_sha256, + expected_size_bytes=expected_size_bytes, + ): + return snapshot + + root.mkdir(parents=True, exist_ok=True) + target = root / _DONOR_FILENAME + if _file_matches( + target, + expected_sha256=expected_sha256, + expected_size_bytes=expected_size_bytes, + ): + return target + + partial = target.with_name(f"{target.name}.part") + digest = hashlib.sha256() + written = 0 + try: + with ( + urllib.request.urlopen(SIPP_2023_VOLUNTARY_FILING_DONOR_URL) as response, # noqa: S310 + partial.open("wb") as output, + ): + while True: + chunk = response.read(chunk_size) + if not chunk: + break + output.write(chunk) + digest.update(chunk) + written += len(chunk) + + if expected_size_bytes is not None and written != expected_size_bytes: + raise ValueError( + "SIPP 2023 voluntary-filing donor failed byte-length " + f"verification: expected {expected_size_bytes}, got {written}." + ) + actual_sha256 = digest.hexdigest() + if expected_sha256 is not None and actual_sha256 != expected_sha256: + raise ValueError( + "SIPP 2023 voluntary-filing donor failed sha-256 verification: " + f"expected {expected_sha256}, got {actual_sha256}." + ) + partial.replace(target) + except Exception: + partial.unlink(missing_ok=True) + raise + return target + + +def _numeric(series: pd.Series) -> pd.Series: + return pd.to_numeric(series, errors="coerce").astype("float64") + + +def _source_unit_keys(december: pd.DataFrame) -> pd.Series: + """Return reciprocal-spouse pair keys, leaving every other person singleton.""" + + if december.duplicated(["SSUID", "PNUM"]).any(): + raise ValueError("SIPP December donor contains duplicate SSUID/PNUM rows.") + spouse_by_person = pd.Series( + december["EPNSPOUSE"].to_numpy(dtype=np.float64), + index=pd.MultiIndex.from_frame(december[["SSUID", "PNUM"]]), + ) + keys: list[str] = [] + for ssuid, person_number, spouse_number in december[ + ["SSUID", "PNUM", "EPNSPOUSE"] + ].itertuples(index=False, name=None): + reciprocal = False + if np.isfinite(spouse_number): + partner_key = (ssuid, spouse_number) + if partner_key in spouse_by_person.index: + reciprocal = bool(spouse_by_person.loc[partner_key] == person_number) + if reciprocal: + low = min(float(person_number), float(spouse_number)) + high = max(float(person_number), float(spouse_number)) + else: + low = high = float(person_number) + keys.append(f"{ssuid}:{low:g}:{high:g}") + return pd.Series(keys, index=december.index, dtype="string") + + +def load_sipp_2023_voluntary_filing_donor( + path: str | Path, + *, + expected_sha256: str | None = None, + expected_size_bytes: int | None = SIPP_2023_VOLUNTARY_FILING_DONOR_SIZE_BYTES, + chunksize: int = 100_000, +) -> pd.DataFrame: + """Transform the pinned SIPP person file to measured filing tax units.""" + + path = Path(path) + if expected_size_bytes is not None and path.stat().st_size != expected_size_bytes: + raise ValueError( + "SIPP 2023 voluntary-filing donor failed byte-length verification: " + f"expected {expected_size_bytes}, got {path.stat().st_size}." + ) + if expected_sha256 is not None: + actual_sha256 = _sha256_file(path) + if actual_sha256 != expected_sha256: + raise ValueError( + "SIPP 2023 voluntary-filing donor failed sha-256 verification: " + f"expected {expected_sha256}, got {actual_sha256}." + ) + if chunksize < 1: + raise ValueError("chunksize must be a positive integer") + + header = pd.read_csv(path, delimiter="|", nrows=0) + missing = sorted(set(SIPP_VOLUNTARY_FILING_SOURCE_COLUMNS) - set(header.columns)) + if missing: + raise ValueError( + f"SIPP 2023 voluntary-filing donor missing column(s): {missing}." + ) + + parts: list[pd.DataFrame] = [] + reader = pd.read_csv( + path, + delimiter="|", + usecols=list(SIPP_VOLUNTARY_FILING_SOURCE_COLUMNS), + chunksize=int(chunksize), + low_memory=False, + ) + for chunk in reader: + month = _numeric(chunk["MONTHCODE"]) + december = chunk.loc[month.eq(12)].copy() + if not december.empty: + parts.append(december) + if not parts: + raise ValueError("SIPP 2023 voluntary-filing donor has no December rows.") + december = pd.concat(parts, ignore_index=True) + + numeric_columns = [ + column for column in SIPP_VOLUNTARY_FILING_SOURCE_COLUMNS if column != "SSUID" + ] + for column in numeric_columns: + original = december[column] + converted = _numeric(original) + invalid = original.notna() & converted.isna() + if invalid.any(): + rows = np.flatnonzero(invalid.to_numpy())[:5].tolist() + raise ValueError( + f"SIPP voluntary-filing source {column!r} contains nonnumeric " + f"value(s) at row(s) {rows}." + ) + december[column] = converted + + if december[["SSUID", "PNUM"]].isna().any(axis=None): + raise ValueError("SIPP voluntary-filing December source has missing IDs.") + december[_DONOR_SOURCE_KEY_COLUMN] = _source_unit_keys(december) + december["_is_under_18"] = december["TAGE"] < 18 + under_18_by_household = december.groupby("SSUID", sort=True)["_is_under_18"].sum() + monthly_earnings = ( + december.loc[:, list(_JOB_MONTHLY_EARNINGS_COLUMNS)].fillna(0.0).sum(axis=1) + ) + invalid_earnings = ~np.isfinite(monthly_earnings) | monthly_earnings.lt(0.0) + if invalid_earnings.any(): + rows = np.flatnonzero(invalid_earnings.to_numpy())[:5].tolist() + raise ValueError( + "SIPP voluntary-filing monthly job earnings must be finite and " + f"nonnegative; invalid row(s) {rows}." + ) + december["_annual_employment_income"] = monthly_earnings * 12.0 + + filing = december["EFILING"] + reported = december["AFILING"].eq(1) & ( + filing.eq(1) | (filing.eq(2) & december["AWILLFILE"].eq(1)) + ) + reported_target = december["EFILING"].eq(1) | december["EWILLFILE"].eq(1) + dependent_reported = reported & december["EDEPCLM"].eq(1) + nondependent = ~december["EDEPCLM"].eq(1) + response = december.loc[reported & nondependent].copy() + if response.empty: + raise ValueError( + "SIPP voluntary-filing donor has no directly reported, " + "nondependent filing responses." + ) + response[_OUTPUT] = reported_target.loc[response.index] + + disagreement = response.groupby(_DONOR_SOURCE_KEY_COLUMN, sort=True)[ + _OUTPUT + ].nunique() + disagreeing_keys = disagreement.index[disagreement > 1].tolist() + if disagreeing_keys: + raise ValueError( + "SIPP voluntary-filing reciprocal spouses disagree on the filing " + f"target for source unit(s) {disagreeing_keys[:5]}." + ) + + ordered = december.sort_values([_DONOR_SOURCE_KEY_COLUMN, "PNUM"], kind="stable") + source_units = ordered.groupby(_DONOR_SOURCE_KEY_COLUMN, sort=True) + # Reference attributes and weight come from the minimum-PNUM member of the + # complete canonical unit, before response filtering. The full reciprocal + # pair likewise supplies unit earnings; only the filing target is limited + # to directly reported, nondependent responses. + reference = ordered.drop_duplicates(_DONOR_SOURCE_KEY_COLUMN).set_index( + _DONOR_SOURCE_KEY_COLUMN + ) + employment = source_units["_annual_employment_income"].sum() + target = response.groupby(_DONOR_SOURCE_KEY_COLUMN, sort=True)[_OUTPUT].first() + + donor = pd.DataFrame(index=target.index) + donor[_DONOR_SOURCE_KEY_COLUMN] = donor.index.astype(str) + donor["employment_income"] = employment.reindex(donor.index).to_numpy( + dtype=np.float64 + ) + donor["reference_age"] = reference.reindex(donor.index)["TAGE"].to_numpy( + dtype=np.float64 + ) + donor["reference_is_female"] = ( + reference.reindex(donor.index)["ESEX"].eq(2).to_numpy(dtype=np.float64) + ) + donor["reference_is_married"] = ( + source_units.size().reindex(donor.index).gt(1).to_numpy(dtype=np.float64) + ) + donor["count_under_18"] = ( + reference.reindex(donor.index)["SSUID"] + .map(under_18_by_household) + .to_numpy(dtype=np.float64) + ) + donor[_OUTPUT] = target.to_numpy(dtype=bool) + donor[_DONOR_WEIGHT_COLUMN] = reference.reindex(donor.index)["WPFINWGT"].to_numpy( + dtype=np.float64 + ) + donor = donor.reset_index(drop=True) + + age = donor["reference_age"].to_numpy(dtype=np.float64) + if not (np.isfinite(age) & (age >= 15.0) & (age <= 120.0)).all(): + raise ValueError( + "SIPP voluntary-filing reference ages must be finite and in [15, 120]." + ) + reference_sex = reference.reindex(target.index)["ESEX"].to_numpy(dtype=np.float64) + if not np.isin(reference_sex, [1.0, 2.0]).all(): + raise ValueError("SIPP voluntary-filing reference sex must be coded 1 or 2.") + for column in SIPP_VOLUNTARY_FILING_MODEL_PREDICTORS: + values = donor[column].to_numpy(dtype=np.float64) + if not np.isfinite(values).all() or (values < 0.0).any(): + raise ValueError( + f"SIPP voluntary-filing predictor {column!r} must be finite " + "and nonnegative." + ) + + preweight_units = len(donor) + preweight_true = int(donor[_OUTPUT].sum()) + weights = donor[_DONOR_WEIGHT_COLUMN].to_numpy(dtype=np.float64) + valid_weight = np.isfinite(weights) & (weights > 0.0) + donor = donor.loc[valid_weight].copy().reset_index(drop=True) + if donor.empty: + raise ValueError( + "SIPP voluntary-filing donor has no positive finite tax-unit weights." + ) + if donor[_OUTPUT].nunique() < 2: + raise ValueError( + "SIPP voluntary-filing donor target is constant after source filtering." + ) + final_weights = donor[_DONOR_WEIGHT_COLUMN].to_numpy(dtype=np.float64) + final_target = donor[_OUTPUT].to_numpy(dtype=bool) + donor.attrs["source_audit"] = { + "december_rows": int(len(december)), + "observed_response_rows": int(reported.sum()), + "observed_response_true_rows": int((reported & reported_target).sum()), + "claimed_dependent_observed_rows": int(dependent_reported.sum()), + "claimed_dependent_observed_true_rows": int( + (dependent_reported & reported_target).sum() + ), + "spouse_target_disagreement_units": int(len(disagreeing_keys)), + "canonical_preweight_units": int(preweight_units), + "canonical_preweight_true_units": int(preweight_true), + "positive_finite_weight_units": int(len(donor)), + "positive_finite_weight_true_units": int(final_target.sum()), + "positive_finite_weight_sum": float(final_weights.sum()), + "weighted_true_share": float( + final_weights[final_target].sum() / final_weights.sum() + ), + } + return donor + + +def _decoded_strings(values: pd.Series) -> pd.Series: + return values.map( + lambda value: ( + value.decode() if isinstance(value, (bytes, np.bytes_)) else str(value) + ) + ) + + +def _strict_recipient_numeric( + person: pd.DataFrame, + columns: tuple[str, ...], + *, + label: str, +) -> np.ndarray: + for column in columns: + if column in person.columns: + values = pd.to_numeric(person[column], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not np.isfinite(values).all(): + raise ValueError( + f"US voluntary-filing receiver {label} source {column!r} " + "contains nonfinite values." + ) + return values + raise ValueError( + f"US voluntary-filing receiver requires one of {list(columns)} for {label}." + ) + + +def _recipient_tax_unit_predictor_table(frame: Frame) -> pd.DataFrame: + """Build the approved five predictors in receiver tax-unit order.""" + + person = frame.table("person") + tax_unit = frame.table("tax_unit") + required = { + "person_id", + "person_tax_unit_id", + "person_household_id", + "age", + "is_female", + } + missing = sorted(required - set(person.columns)) + if missing: + raise ValueError( + f"US voluntary-filing receiver missing person column(s): {missing}." + ) + if "tax_unit_id" not in tax_unit: + raise ValueError("US voluntary-filing receiver requires tax_unit_id.") + + age = _strict_recipient_numeric(person, ("age", "A_AGE"), label="age") + if not ((age >= 0.0) & (age <= 120.0)).all(): + raise ValueError("US voluntary-filing receiver age must be in [0, 120].") + income = _strict_recipient_numeric( + person, + ("employment_income_before_lsr", "employment_income", "WSAL_VAL"), + label="employment income", + ) + if (income < 0.0).any(): + raise ValueError("US voluntary-filing receiver wages must be nonnegative.") + female_numeric = pd.to_numeric(person["is_female"], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not (np.isfinite(female_numeric) & np.isin(female_numeric, [0.0, 1.0])).all(): + raise ValueError("US voluntary-filing receiver is_female must be boolean.") + + work = pd.DataFrame( + { + "person_id": person["person_id"].to_numpy(), + "tax_unit_id": person["person_tax_unit_id"].to_numpy(), + "household_id": person["person_household_id"].to_numpy(), + "age": age, + "is_female": female_numeric, + "employment_income": income, + }, + index=person.index, + ) + if "tax_unit_role_input" not in person: + raise ValueError( + "US voluntary-filing receiver requires tax_unit_role_input to " + "identify exactly one tax-unit head." + ) + roles = _decoded_strings(person["tax_unit_role_input"]) + head = roles.eq("HEAD") + head_counts = head.groupby(person["person_tax_unit_id"], sort=False).sum() + if not head_counts.eq(1).all(): + bad = head_counts.index[~head_counts.eq(1)].tolist() + raise ValueError( + "US voluntary-filing receiver requires exactly one HEAD per tax " + f"unit; invalid unit(s) {bad[:5]}." + ) + work["_head_rank"] = (~head).astype(np.int8) + if "A_LINENO" in person: + work["_line"] = pd.to_numeric(person["A_LINENO"], errors="coerce").fillna( + np.inf + ) + else: + work["_line"] = work["person_id"] + + unit_household_counts = work.groupby("tax_unit_id", sort=False)[ + "household_id" + ].nunique() + if (unit_household_counts != 1).any(): + bad = unit_household_counts.index[unit_household_counts != 1].tolist() + raise ValueError( + f"US voluntary-filing receiver tax units span households: {bad[:5]}." + ) + household_under_18 = ( + pd.DataFrame( + { + "household_id": work["household_id"], + "under_18": work["age"].lt(18), + } + ) + .groupby("household_id", sort=False)["under_18"] + .sum() + ) + grouped = work.groupby("tax_unit_id", sort=False) + unit_income = grouped["employment_income"].sum() + unit_household = grouped["household_id"].first() + reference = ( + work.sort_values( + ["tax_unit_id", "_head_rank", "_line", "person_id"], kind="stable" + ) + .drop_duplicates("tax_unit_id") + .set_index("tax_unit_id") + ) + + filing_status_column = next( + ( + column + for column in ("filing_status_input", "filing_status") + if column in tax_unit + ), + None, + ) + if filing_status_column is None: + raise ValueError( + "US voluntary-filing receiver requires tax-unit filing_status_input." + ) + filing_status = _decoded_strings(tax_unit[filing_status_column]) + married = pd.Series( + filing_status.eq("JOINT").to_numpy(), + index=tax_unit["tax_unit_id"].to_numpy(), + ) + + unit_ids = tax_unit["tax_unit_id"].to_numpy() + receiver = pd.DataFrame(index=unit_ids) + receiver["employment_income"] = unit_income.reindex(unit_ids).to_numpy( + dtype=np.float64 + ) + receiver["reference_age"] = reference.reindex(unit_ids)["age"].to_numpy( + dtype=np.float64 + ) + receiver["reference_is_female"] = reference.reindex(unit_ids)["is_female"].to_numpy( + dtype=np.float64 + ) + receiver["reference_is_married"] = ( + married.reindex(unit_ids).fillna(False).to_numpy(dtype=np.float64) + ) + receiver["count_under_18"] = ( + unit_household.reindex(unit_ids) + .map(household_under_18) + .to_numpy(dtype=np.float64) + ) + if receiver.isna().any(axis=None): + bad = receiver.index[receiver.isna().any(axis=1)].tolist() + raise ValueError( + f"US voluntary-filing receiver does not cover tax unit(s) {bad[:5]}." + ) + return receiver.loc[:, list(SIPP_VOLUNTARY_FILING_MODEL_PREDICTORS)] + + +def _source_receiver_rows( + frame: Frame, + receiver: pd.DataFrame, +) -> tuple[pd.DataFrame, pd.Series]: + """Select one ASEC predictor row per source unit and return fan-out keys.""" + + tax_unit = frame.table("tax_unit") + if _TAX_UNIT_SOURCE_ID_COLUMN in tax_unit: + source_ids = tax_unit[_TAX_UNIT_SOURCE_ID_COLUMN] + if source_ids.isna().any(): + raise ValueError( + "US voluntary-filing receiver has missing tax_unit_source_id." + ) + else: + if _TAX_UNIT_SUPPORT_CHANNEL_COLUMN in tax_unit: + raise ValueError( + "US voluntary-filing support clones require tax_unit_source_id." + ) + source_ids = tax_unit["tax_unit_id"] + + rows = receiver.copy() + rows["_source_id"] = _decoded_strings(source_ids).to_numpy() + rows["_tax_unit_id"] = tax_unit["tax_unit_id"].to_numpy() + if _TAX_UNIT_SUPPORT_CHANNEL_COLUMN in tax_unit: + rows["_support_channel"] = _decoded_strings( + tax_unit[_TAX_UNIT_SUPPORT_CHANNEL_COLUMN] + ).to_numpy() + asec_counts = ( + rows["_support_channel"] + .eq(_BASE_ASEC_SUPPORT_CHANNEL) + .groupby(rows["_source_id"]) + .sum() + ) + if not asec_counts.eq(1).all(): + bad = asec_counts.index[~asec_counts.eq(1)].tolist() + raise ValueError( + "US voluntary-filing support source units require exactly one " + f"ASEC row; invalid source unit(s) {bad[:5]}." + ) + source_rows = rows.loc[ + rows["_support_channel"].eq(_BASE_ASEC_SUPPORT_CHANNEL) + ].copy() + else: + if rows["_source_id"].duplicated().any(): + duplicates = rows.loc[ + rows["_source_id"].duplicated(), "_source_id" + ].tolist() + raise ValueError( + "US voluntary-filing unexpanded receiver has duplicate source " + f"unit(s) {duplicates[:5]}." + ) + source_rows = rows + + source_rows = source_rows.sort_values("_source_id", kind="stable") + prediction_rows = source_rows.loc[:, list(SIPP_VOLUNTARY_FILING_MODEL_PREDICTORS)] + prediction_rows.index = source_rows["_source_id"].to_numpy() + return prediction_rows, rows["_source_id"] + + +def impute_us_voluntary_filing( + frame: Frame, + donor: pd.DataFrame, + *, + seed: int, + n_estimators: int = _DEFAULT_N_ESTIMATORS, +) -> pd.Series: + """Fit the weighted SIPP QRF and draw once per source tax unit.""" + + required = { + *SIPP_VOLUNTARY_FILING_MODEL_PREDICTORS, + _OUTPUT, + _DONOR_WEIGHT_COLUMN, + } + missing = sorted(required - set(donor.columns)) + if missing: + raise ValueError( + f"SIPP voluntary-filing donor table missing column(s): {missing}." + ) + if n_estimators < 1: + raise ValueError("n_estimators must be positive") + + training = ( + donor.loc[:, [*SIPP_VOLUNTARY_FILING_MODEL_PREDICTORS, _OUTPUT]] + .copy() + .reset_index(drop=True) + ) + for column in SIPP_VOLUNTARY_FILING_MODEL_PREDICTORS: + training[column] = pd.to_numeric(training[column], errors="coerce") + values = training[column].to_numpy(dtype=np.float64) + if not np.isfinite(values).all() or (values < 0.0).any(): + raise ValueError( + f"SIPP voluntary-filing donor predictor {column!r} must be " + "finite and nonnegative." + ) + target = pd.to_numeric(training[_OUTPUT], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not (np.isfinite(target) & np.isin(target, [0.0, 1.0])).all(): + raise ValueError("SIPP voluntary-filing donor target must be boolean.") + if np.unique(target).size < 2: + raise ValueError("SIPP voluntary-filing donor target must be nonconstant.") + training[_OUTPUT] = target + weights = pd.to_numeric(donor[_DONOR_WEIGHT_COLUMN], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not (np.isfinite(weights) & (weights > 0.0)).all(): + raise ValueError( + "SIPP voluntary-filing donor weights must be finite and positive." + ) + + # Sort on source identity when available, otherwise on the complete fit + # row, so reruns do not depend on incidental input ordering. + if _DONOR_SOURCE_KEY_COLUMN in donor: + order = np.argsort( + donor[_DONOR_SOURCE_KEY_COLUMN].astype(str).to_numpy(), kind="stable" + ) + else: + order = ( + training.assign(_weight=weights) + .sort_values( + [ + *SIPP_VOLUNTARY_FILING_MODEL_PREDICTORS, + _OUTPUT, + "_weight", + ], + kind="stable", + ) + .index.to_numpy(dtype=np.int64) + ) + training = training.iloc[order].reset_index(drop=True) + weights = weights[order] + + receiver = _recipient_tax_unit_predictor_table(frame) + source_receiver, fanout_keys = _source_receiver_rows(frame, receiver) + global QRF + if QRF is None: + from importlib import import_module + + QRF = import_module("populace.fit").QRF + fitted = QRF(n_estimators=int(n_estimators), seed=int(seed)).fit( + training, + predictors=list(SIPP_VOLUNTARY_FILING_MODEL_PREDICTORS), + targets=[_OUTPUT], + weights=weights, + ) + prediction = fitted.predict(source_receiver) + if _OUTPUT not in prediction: + raise ValueError( + f"SIPP voluntary-filing QRF prediction is missing {_OUTPUT!r}." + ) + values = pd.to_numeric(prediction[_OUTPUT], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not np.isfinite(values).all(): + raise ValueError("SIPP voluntary-filing QRF produced nonfinite values.") + if ((values < 0.0) | (values > 1.0)).any(): + raise ValueError("SIPP voluntary-filing QRF produced values outside [0, 1].") + by_source = pd.Series(values >= 0.5, index=source_receiver.index) + fanout = fanout_keys.map(by_source) + if fanout.isna().any(): + raise ValueError("SIPP voluntary-filing source-unit fan-out is incomplete.") + return pd.Series( + fanout.to_numpy(dtype=bool), + index=frame.table("tax_unit").index, + name=_OUTPUT, + ) + + +def with_us_voluntary_filing_input( + frame: Frame, + *, + seed: int, + time_period: int, + sipp_donor: pd.DataFrame, + n_estimators: int = _DEFAULT_N_ESTIMATORS, +) -> Frame: + """Materialize the deterministic, source-backed tax-unit input.""" + + if frame.schema != US_SCHEMA: + raise ValueError("US voluntary-filing input requires the US schema.") + del time_period # The measured donor vintage is pinned independently. + us_voluntary_filing_stage_spec() + predicted = impute_us_voluntary_filing( + frame, + sipp_donor, + seed=int(seed), + n_estimators=int(n_estimators), + ) + tax_unit = frame.table("tax_unit") + if _OUTPUT in tax_unit: + current = pd.to_numeric(tax_unit[_OUTPUT], errors="coerce").to_numpy( + dtype=np.float64 + ) + if ( + np.isfinite(current).all() + and np.isin(current, [0.0, 1.0]).all() + and np.array_equal(current.astype(bool), predicted.to_numpy()) + ): + return frame + + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + tables["tax_unit"][_OUTPUT] = predicted.to_numpy(dtype=bool) + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + ) + + +def us_voluntary_filing_summary(frame: Frame) -> dict[str, object]: + """Return weighted incidence, boolean validity, and clone diagnostics.""" + + tax_unit = frame.table("tax_unit") + values = pd.to_numeric(tax_unit[_OUTPUT], errors="coerce").to_numpy( + dtype=np.float64 + ) + weights = np.asarray(frame.resolve_weights("tax_unit").values, dtype=np.float64) + valid_weights = np.isfinite(weights) & (weights >= 0.0) + if not valid_weights.all() or float(weights.sum()) <= 0.0: + rows = np.flatnonzero(~valid_weights)[:5].tolist() + raise ValueError( + "US voluntary-filing gate requires finite nonnegative tax-unit " + f"weights with positive total; invalid row(s) {rows}." + ) + finite = np.isfinite(values) + boolean = finite & np.isin(values, [0.0, 1.0]) + true = boolean & (values == 1.0) + total_weight = float(weights.sum()) + + clone_mismatch_source_units = 0 + clone_source_units = 0 + clone_metadata_missing = False + channel_diagnostics: dict[str, dict[str, float | int]] = {} + if _TAX_UNIT_SUPPORT_CHANNEL_COLUMN in tax_unit: + channels = _decoded_strings(tax_unit[_TAX_UNIT_SUPPORT_CHANNEL_COLUMN]) + for channel in sorted(channels.unique()): + mask = channels.eq(channel).to_numpy() + channel_weight = float(weights[mask].sum()) + channel_true = true & mask + channel_diagnostics[channel] = { + "unique_count": int(pd.Series(values[mask & finite]).nunique()), + "weighted_true_share": ( + float(weights[channel_true].sum()) / channel_weight + if channel_weight > 0.0 + else 0.0 + ), + } + if _TAX_UNIT_SOURCE_ID_COLUMN not in tax_unit: + clone_metadata_missing = True + else: + source_ids = tax_unit[_TAX_UNIT_SOURCE_ID_COLUMN] + if source_ids.isna().any(): + clone_metadata_missing = True + else: + clone_table = pd.DataFrame( + {"source_id": source_ids.astype(str), "value": values} + ) + sizes = clone_table.groupby("source_id", sort=False).size() + clone_source_units = int((sizes > 1).sum()) + unique = clone_table.groupby("source_id", sort=False)["value"].nunique( + dropna=False + ) + clone_mismatch_source_units = int((unique > 1).sum()) + + return { + "weighted_true_share": float(weights[true].sum()) / total_weight, + "weighted_true_total": float(weights[true].sum()), + "weighted_false_total": float(weights[boolean & ~true].sum()), + "true_share_band": list(_TRUE_SHARE_BAND), + "unique_count": int(pd.Series(values[finite]).nunique()), + "missing_or_nonfinite_count": int((~finite).sum()), + "invalid_boolean_count": int((finite & ~boolean).sum()), + "clone_source_units": clone_source_units, + "clone_mismatch_source_units": clone_mismatch_source_units, + "clone_metadata_missing": clone_metadata_missing, + "channel_diagnostics": channel_diagnostics, + } + + +def us_voluntary_filing_signal_gate(frame: Frame) -> GateResult: + """Require nonconstant boolean signal and identical support clones.""" + + tax_unit = frame.table("tax_unit") + if _OUTPUT not in tax_unit: + return GateResult( + name="voluntary_filing_signal", + passed=False, + failures=(f"tax_unit column missing: {_OUTPUT!r}.",), + details={"missing": [_OUTPUT]}, + ) + summary = us_voluntary_filing_summary(frame) + failures: list[str] = [] + if summary["missing_or_nonfinite_count"]: + failures.append( + f"{_OUTPUT}: {summary['missing_or_nonfinite_count']} missing or " + "nonfinite value(s)." + ) + if summary["invalid_boolean_count"]: + failures.append( + f"{_OUTPUT}: {summary['invalid_boolean_count']} non-boolean value(s)." + ) + if summary["unique_count"] < 2: + failures.append(f"{_OUTPUT}: constant column carries no signal.") + share = float(summary["weighted_true_share"]) + low, high = summary["true_share_band"] + if not (low <= share <= high): + failures.append( + f"{_OUTPUT}: weighted true share {share:.4f} outside plausibility " + f"band [{low}, {high}]." + ) + if float(summary["weighted_true_total"]) <= 0.0: + failures.append(f"{_OUTPUT}: weighted true total is not positive.") + if float(summary["weighted_false_total"]) <= 0.0: + failures.append(f"{_OUTPUT}: weighted false total is not positive.") + if summary["clone_metadata_missing"]: + failures.append( + "Voluntary-filing support clones lack complete tax_unit_source_id " + "provenance." + ) + if summary["clone_mismatch_source_units"]: + failures.append( + "Voluntary-filing support clones disagree for " + f"{summary['clone_mismatch_source_units']} source unit(s)." + ) + for channel, diagnostics in summary["channel_diagnostics"].items(): + if diagnostics["unique_count"] < 2: + failures.append(f"{_OUTPUT}: support channel {channel!r} is constant.") + channel_share = float(diagnostics["weighted_true_share"]) + if not (low <= channel_share <= high): + failures.append( + f"{_OUTPUT}: support channel {channel!r} weighted true share " + f"{channel_share:.4f} outside plausibility band [{low}, {high}]." + ) + return GateResult( + name="voluntary_filing_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) diff --git a/packages/populace-build/src/populace/build/us_runtime/weeks_unemployed.py b/packages/populace-build/src/populace/build/us_runtime/weeks_unemployed.py new file mode 100644 index 00000000..4627cce7 --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/weeks_unemployed.py @@ -0,0 +1,1377 @@ +"""Measured CPS ASEC weeks looking for work and the PUF-half QRF. + +The retired eCPS build carried person-level ``LKWEEKS`` directly, replacing +only its historical ``-1`` not-in-universe sentinel with zero. The locked +2022-income-year ASEC HDF5 omitted that raw column, so this stage repairs only +that cohort from the official, immutable 2023 ASEC public-use archive. The +repair is an exact ``PERIDNUM`` join with redundant Census identity checks; it +never predicts an ASEC source value. + +After support cloning, the retired PUF-specific routine replaces only the PUF +half with a QRF using age, sex, joint filing, the three tax-unit roles, and +unemployment compensation when available. Populace adds typed person weights +and a build seed. QRF interpolation can be fractional, so predictions are +clipped and rounded to the integer-valued 0--52 input domain before the +archived unemployment-compensation zero rule is applied. +""" + +from __future__ import annotations + +import hashlib +import os +import urllib.request +import zipfile +from importlib.resources import files +from pathlib import Path +from typing import Any, BinaryIO + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import ( + SourceOperationSpec, + SourceStageSpec, + load_source_manifest, +) +from populace.build.source_runtime import ( + SourceRuntimeConfig, + SourceRuntimeContext, + SourceRuntimeError, + run_source_stage, +) +from populace.frame import Frame +from populace.frame.units import US_SCHEMA + +__all__ = [ + "ASEC_2023_WEEKS_UNEMPLOYED_MEMBER", + "ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_CRC32", + "ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_SHA256", + "ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_SIZE_BYTES", + "ASEC_2023_WEEKS_UNEMPLOYED_POSITIVE_ROWS", + "ASEC_2023_WEEKS_UNEMPLOYED_RAW_ROWS", + "ASEC_2023_WEEKS_UNEMPLOYED_SOURCE_COLUMNS", + "ASEC_2023_WEEKS_UNEMPLOYED_SOURCE_SHA256", + "ASEC_2023_WEEKS_UNEMPLOYED_SOURCE_YEAR", + "ASEC_2023_WEEKS_UNEMPLOYED_UNIQUE_KEYS", + "ASEC_2023_WEEKS_UNEMPLOYED_WEIGHTED_SOURCE_SHARE", + "ASEC_2023_WEEKS_UNEMPLOYED_WEIGHTED_WEEKS", + "ASEC_2023_WEEKS_UNEMPLOYED_ZIP_SHA256", + "ASEC_2023_WEEKS_UNEMPLOYED_ZIP_SIZE_BYTES", + "ASEC_2023_WEEKS_UNEMPLOYED_ZIP_URL", + "WEEKS_UNEMPLOYED_ARCHIVED_DERIVATION_URL", + "WEEKS_UNEMPLOYED_ARCHIVED_PUF_IMPUTATION_URL", + "WEEKS_UNEMPLOYED_ARCHIVED_SOURCE_URL", + "WEEKS_UNEMPLOYED_DERIVE_PARAMETERS", + "WEEKS_UNEMPLOYED_PUF_IMPUTATION_PARAMETERS", + "WEEKS_UNEMPLOYED_PUF_PREDICTORS", + "WEEKS_UNEMPLOYED_READ_PARAMETERS", + "US_WEEKS_UNEMPLOYED_NONCONSTANT_PERSON_COLUMNS", + "US_WEEKS_UNEMPLOYED_OUTPUT_COLUMNS", + "US_WEEKS_UNEMPLOYED_REQUIRED_SOURCE_COLUMNS", + "US_WEEKS_UNEMPLOYED_STAGE_NAME", + "derive_us_weeks_unemployed_from_manifest", + "fetch_asec_2023_weeks_unemployed_source", + "fill_asec_2022_weeks_unemployed_source", + "impute_us_weeks_unemployed_to_puf_support_from_manifest", + "load_asec_2023_weeks_unemployed_source", + "us_weeks_unemployed_signal_gate", + "us_weeks_unemployed_stage_spec", + "us_weeks_unemployed_summary", + "with_us_weeks_unemployed", +] + +QRF: Any | None = None + +_ARCHIVED_COMMIT = "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe" +_ARCHIVED_REPOSITORY = "policyengine-" + "us-data" +_ARCHIVED_PACKAGE = "policyengine_" + "us_data" +_ARCHIVED_ROOT = ( + f"https://github.com/PolicyEngine/{_ARCHIVED_REPOSITORY}/blob/" + f"{_ARCHIVED_COMMIT}/{_ARCHIVED_PACKAGE}" +) +WEEKS_UNEMPLOYED_ARCHIVED_DERIVATION_URL = ( + _ARCHIVED_ROOT + "/datasets/cps/cps.py#L1438-L1442" +) +WEEKS_UNEMPLOYED_ARCHIVED_SOURCE_URL = ( + _ARCHIVED_ROOT + "/datasets/cps/census_cps.py#L39-L58" +) +WEEKS_UNEMPLOYED_ARCHIVED_PUF_IMPUTATION_URL = ( + _ARCHIVED_ROOT + "/calibration/puf_impute.py#L703-L785" +) + +ASEC_2023_WEEKS_UNEMPLOYED_ZIP_URL = ( + "https://www2.census.gov/programs-surveys/cps/datasets/2023/march/asecpub23csv.zip" +) +ASEC_2023_WEEKS_UNEMPLOYED_ZIP_SIZE_BYTES = 150_165_063 +ASEC_2023_WEEKS_UNEMPLOYED_ZIP_SHA256 = ( + "d2e000250782adfbdd7f29c82b66d866591a30f0d330496698ec19f9c784ce11" +) +# Stable upstream identity used by release checkpoints and reform-vector +# caches. Fetching may return the extracted member, but both representations +# are accepted only after verification against this archive and its member. +ASEC_2023_WEEKS_UNEMPLOYED_SOURCE_SHA256 = ASEC_2023_WEEKS_UNEMPLOYED_ZIP_SHA256 +ASEC_2023_WEEKS_UNEMPLOYED_MEMBER = "pppub23.csv" +ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_SIZE_BYTES = 281_065_733 +ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_CRC32 = "49c09e5f" +ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_SHA256 = ( + "19b56537e50e7663f954361ef2bb5ce9cef8d9d45f156fe1a69a99b654198ffe" +) +ASEC_2023_WEEKS_UNEMPLOYED_SOURCE_YEAR = 2022 +ASEC_2023_WEEKS_UNEMPLOYED_RAW_ROWS = 146_133 +ASEC_2023_WEEKS_UNEMPLOYED_UNIQUE_KEYS = 146_133 +ASEC_2023_WEEKS_UNEMPLOYED_POSITIVE_ROWS = 4_200 +ASEC_2023_WEEKS_UNEMPLOYED_WEIGHTED_SOURCE_SHARE = 0.03145609427589182 +ASEC_2023_WEEKS_UNEMPLOYED_WEIGHTED_WEEKS = 172_563_622.04 + +ASEC_2023_WEEKS_UNEMPLOYED_SOURCE_COLUMNS: tuple[str, ...] = ( + "PH_SEQ", + "P_SEQ", + "A_LINENO", + "PERIDNUM", + "LKWEEKS", +) +_AUDIT_WEIGHT_COLUMN = "A_FNLWGT" + +US_WEEKS_UNEMPLOYED_STAGE_NAME = "weeks_unemployed_input" +US_WEEKS_UNEMPLOYED_OUTPUT_COLUMNS: tuple[str, ...] = ("weeks_unemployed",) +US_WEEKS_UNEMPLOYED_NONCONSTANT_PERSON_COLUMNS = US_WEEKS_UNEMPLOYED_OUTPUT_COLUMNS +US_WEEKS_UNEMPLOYED_REQUIRED_SOURCE_COLUMNS: tuple[str, ...] = ( + "source_year", + "PERIDNUM", + "LKWEEKS", +) + +_OUTPUT = US_WEEKS_UNEMPLOYED_OUTPUT_COLUMNS[0] +_SOURCE = "LKWEEKS" +_PERSON_WEIGHT_COLUMN = "person_weight" +_PERSON_SUPPORT_CHANNEL_COLUMN = "person_support_channel" +_ASEC_CHANNEL = "asec" +_PUF_CHANNEL = "puf_tax_detail" +_PREDICTOR_PREFIX = "weeks_unemployed_predictor_" +_REQUIRED_PUF_PREDICTORS: tuple[str, ...] = ( + "age", + "is_male", + "tax_unit_is_joint", + "is_tax_unit_head", + "is_tax_unit_spouse", + "is_tax_unit_dependent", +) +_OPTIONAL_UC_PREDICTOR = "unemployment_compensation" +WEEKS_UNEMPLOYED_PUF_PREDICTORS: tuple[str, ...] = ( + *_REQUIRED_PUF_PREDICTORS, + _OPTIONAL_UC_PREDICTOR, +) + +WEEKS_UNEMPLOYED_READ_PARAMETERS: dict[str, object] = { + "table": "person", + "weight": _PERSON_WEIGHT_COLUMN, +} +WEEKS_UNEMPLOYED_DERIVE_PARAMETERS: dict[str, object] = { + "source": _SOURCE, + "niu_value": -1, + "niu_replacement": 0, + "output": _OUTPUT, +} +WEEKS_UNEMPLOYED_PUF_IMPUTATION_PARAMETERS: dict[str, object] = { + "predictors": list(WEEKS_UNEMPLOYED_PUF_PREDICTORS), + "max_train_samples": 5_000, + "n_estimators": 100, + "seed_from_build_config": True, + "weight": _PERSON_WEIGHT_COLUMN, + "clip": [0, 52], + "round": "nearest_integer", + "zero_unless": "unemployment_compensation > 0", +} + +# Deliberately broad source-plausibility bands. The full official 2023 ASEC +# source has a 3.1456% weighted positive share and 0.5234 weighted weeks per +# person; the persisted sparse production smoke has 5.2182% / 0.9036 on ASEC +# and 0.5594% / 0.0913 on PUF support. These floors reject a one-row or +# effectively collapsed QRF surface while allowing support selection and +# adjacent ASEC vintages ample room to move. +_CHANNEL_POSITIVE_SHARE_BANDS: dict[str, tuple[float, float]] = { + _ASEC_CHANNEL: (0.015, 0.10), + _PUF_CHANNEL: (0.001, 0.03), +} +_CHANNEL_WEIGHTED_MEAN_WEEKS_BANDS: dict[str, tuple[float, float]] = { + _ASEC_CHANNEL: (0.20, 2.0), + _PUF_CHANNEL: (0.01, 0.50), +} + + +def us_weeks_unemployed_stage_spec() -> SourceStageSpec: + """Load and strictly validate the packaged weeks-unemployed stage.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + spec = manifest.stage_map().get(US_WEEKS_UNEMPLOYED_STAGE_NAME) + if spec is None: + raise ValueError( + f"US source manifest declares no {US_WEEKS_UNEMPLOYED_STAGE_NAME!r} stage." + ) + if spec.grain != "person": + raise ValueError("US weeks-unemployed stage must have person grain.") + if tuple(spec.outputs) != US_WEEKS_UNEMPLOYED_OUTPUT_COLUMNS: + raise ValueError( + "US weeks-unemployed manifest outputs do not match the runtime family." + ) + expected_operations = ( + ("read_table", WEEKS_UNEMPLOYED_READ_PARAMETERS), + ("derive_weeks_unemployed", WEEKS_UNEMPLOYED_DERIVE_PARAMETERS), + ( + "impute_weeks_unemployed_to_puf_support", + WEEKS_UNEMPLOYED_PUF_IMPUTATION_PARAMETERS, + ), + ) + actual_operations = tuple( + (operation.kind, dict(operation.parameters)) for operation in spec.operations + ) + if actual_operations != expected_operations: + raise ValueError( + "US weeks-unemployed operations drifted from the reviewed direct " + "carry and PUF-imputation contract." + ) + pinned = [ + artifact + for artifact in spec.artifacts + if artifact.get("locator") == ASEC_2023_WEEKS_UNEMPLOYED_ZIP_URL + and artifact.get("size_bytes") == ASEC_2023_WEEKS_UNEMPLOYED_ZIP_SIZE_BYTES + and artifact.get("sha256") == ASEC_2023_WEEKS_UNEMPLOYED_ZIP_SHA256 + and artifact.get("member") == ASEC_2023_WEEKS_UNEMPLOYED_MEMBER + and artifact.get("member_size_bytes") + == ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_SIZE_BYTES + and str(artifact.get("member_crc32", "")).lower() + == ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_CRC32 + and artifact.get("member_sha256") == ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_SHA256 + ] + if not pinned: + raise ValueError( + "US weeks-unemployed stage does not pin the official archive and " + "person member identities." + ) + return spec + + +def _sha256_stream(stream: BinaryIO, *, chunk_size: int) -> tuple[str, int]: + digest = hashlib.sha256() + size = 0 + for chunk in iter(lambda: stream.read(chunk_size), b""): + digest.update(chunk) + size += len(chunk) + return digest.hexdigest(), size + + +def _verify_file( + path: Path, + *, + label: str, + expected_size_bytes: int | None, + expected_sha256: str | None, + chunk_size: int, +) -> None: + if not path.is_file(): + raise FileNotFoundError(path) + size = path.stat().st_size + if expected_size_bytes is not None and size != expected_size_bytes: + raise ValueError( + f"{label} byte length mismatch: expected {expected_size_bytes}, got {size}." + ) + if expected_sha256 is not None: + with path.open("rb") as stream: + digest, _ = _sha256_stream(stream, chunk_size=chunk_size) + if digest != expected_sha256: + raise ValueError( + f"{label} SHA-256 mismatch: expected {expected_sha256}, got {digest}." + ) + + +def _verified_member_info( + archive: zipfile.ZipFile, + *, + expected_size_bytes: int | None, + expected_crc32: str | None, +) -> zipfile.ZipInfo: + members = [ + info + for info in archive.infolist() + if info.filename == ASEC_2023_WEEKS_UNEMPLOYED_MEMBER + ] + if len(members) != 1: + raise ValueError( + "ASEC 2023 archive must contain exactly one " + f"{ASEC_2023_WEEKS_UNEMPLOYED_MEMBER!r} member; found {len(members)}." + ) + info = members[0] + if expected_size_bytes is not None and info.file_size != expected_size_bytes: + raise ValueError( + "ASEC 2023 person member byte length mismatch: expected " + f"{expected_size_bytes}, got {info.file_size}." + ) + actual_crc = f"{info.CRC:08x}" + if expected_crc32 is not None and actual_crc != expected_crc32.lower(): + raise ValueError( + "ASEC 2023 person member CRC32 mismatch: expected " + f"{expected_crc32.lower()}, got {actual_crc}." + ) + return info + + +def fetch_asec_2023_weeks_unemployed_source( + cache_dir: str | Path | None = None, + *, + expected_zip_size_bytes: int | None = (ASEC_2023_WEEKS_UNEMPLOYED_ZIP_SIZE_BYTES), + expected_zip_sha256: str | None = ASEC_2023_WEEKS_UNEMPLOYED_ZIP_SHA256, + expected_member_size_bytes: int | None = ( + ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_SIZE_BYTES + ), + expected_member_crc32: str | None = (ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_CRC32), + expected_member_sha256: str | None = (ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_SHA256), + chunk_size: int = 8 * 1024 * 1024, +) -> Path: + """Download, verify, extract, and cache the official ASEC person member.""" + + if chunk_size < 1: + raise ValueError("chunk_size must be positive") + root = ( + Path(cache_dir).expanduser() + if cache_dir is not None + else Path.home() / ".cache" / "populace" / "cps" / "asec_2023" + ) + root.mkdir(parents=True, exist_ok=True) + archive_path = root / "asecpub23csv.zip" + member_path = root / ASEC_2023_WEEKS_UNEMPLOYED_MEMBER + + if not archive_path.exists(): + temporary = archive_path.with_suffix(".zip.part") + request = urllib.request.Request( + ASEC_2023_WEEKS_UNEMPLOYED_ZIP_URL, + headers={"User-Agent": "populace-build/asec-weeks-unemployed"}, + ) + try: + with urllib.request.urlopen(request, timeout=180) as response: # noqa: S310 + with temporary.open("wb") as destination: + digest = hashlib.sha256() + size = 0 + for chunk in iter(lambda: response.read(chunk_size), b""): + destination.write(chunk) + digest.update(chunk) + size += len(chunk) + if expected_zip_size_bytes is not None and size != expected_zip_size_bytes: + raise ValueError( + "ASEC 2023 archive download byte length mismatch: expected " + f"{expected_zip_size_bytes}, got {size}." + ) + actual_digest = digest.hexdigest() + if expected_zip_sha256 is not None and actual_digest != expected_zip_sha256: + raise ValueError( + "ASEC 2023 archive download SHA-256 mismatch: expected " + f"{expected_zip_sha256}, got {actual_digest}." + ) + os.replace(temporary, archive_path) + finally: + temporary.unlink(missing_ok=True) + + _verify_file( + archive_path, + label="ASEC 2023 archive", + expected_size_bytes=expected_zip_size_bytes, + expected_sha256=expected_zip_sha256, + chunk_size=chunk_size, + ) + with zipfile.ZipFile(archive_path) as archive: + info = _verified_member_info( + archive, + expected_size_bytes=expected_member_size_bytes, + expected_crc32=expected_member_crc32, + ) + if member_path.exists(): + _verify_file( + member_path, + label="ASEC 2023 person member", + expected_size_bytes=expected_member_size_bytes, + expected_sha256=expected_member_sha256, + chunk_size=chunk_size, + ) + return member_path + + temporary = member_path.with_suffix(".csv.part") + try: + with archive.open(info) as source, temporary.open("wb") as destination: + digest = hashlib.sha256() + size = 0 + for chunk in iter(lambda: source.read(chunk_size), b""): + destination.write(chunk) + digest.update(chunk) + size += len(chunk) + if ( + expected_member_size_bytes is not None + and size != expected_member_size_bytes + ): + raise ValueError( + "ASEC 2023 extracted person member byte length mismatch: " + f"expected {expected_member_size_bytes}, got {size}." + ) + actual_digest = digest.hexdigest() + if ( + expected_member_sha256 is not None + and actual_digest != expected_member_sha256 + ): + raise ValueError( + "ASEC 2023 extracted person member SHA-256 mismatch: " + f"expected {expected_member_sha256}, got {actual_digest}." + ) + os.replace(temporary, member_path) + finally: + temporary.unlink(missing_ok=True) + return member_path + + +def _fixed_width_peridnum(values: pd.Series, *, label: str) -> pd.Series: + if values.isna().any(): + rows = values.index[values.isna()].tolist()[:5] + raise ValueError(f"{label} PERIDNUM is missing at row(s): {rows}.") + decoded = values.map( + lambda value: ( + value.decode() + if isinstance(value, (bytes, bytearray, np.bytes_)) + else value + ) + ).astype(str) + valid = decoded.str.fullmatch(r"[0-9]{22}", na=False) + if not valid.all(): + rows = decoded.index[~valid].tolist()[:5] + raise ValueError( + f"{label} PERIDNUM must be an exact 22-digit string at row(s): {rows}." + ) + return decoded + + +def _strict_integer( + values: pd.Series, + *, + label: str, + minimum: int | None = None, + maximum: int | None = None, +) -> np.ndarray: + numeric = pd.to_numeric(values, errors="coerce").to_numpy(dtype=np.float64) + valid = np.isfinite(numeric) & (numeric == np.floor(numeric)) + if minimum is not None: + valid &= numeric >= minimum + if maximum is not None: + valid &= numeric <= maximum + if not valid.all(): + rows = np.flatnonzero(~valid)[:5].tolist() + raise ValueError(f"{label} must be finite integers at row(s): {rows}.") + return numeric.astype(np.int64) + + +def _load_csv_source( + source: Path, + *, + expected_zip_size_bytes: int | None, + expected_zip_sha256: str | None, + expected_member_size_bytes: int | None, + expected_member_crc32: str | None, + expected_member_sha256: str | None, + chunk_size: int, +) -> pd.DataFrame: + usecols = [*ASEC_2023_WEEKS_UNEMPLOYED_SOURCE_COLUMNS, _AUDIT_WEIGHT_COLUMN] + if zipfile.is_zipfile(source): + _verify_file( + source, + label="ASEC 2023 archive", + expected_size_bytes=expected_zip_size_bytes, + expected_sha256=expected_zip_sha256, + chunk_size=chunk_size, + ) + with zipfile.ZipFile(source) as archive: + info = _verified_member_info( + archive, + expected_size_bytes=expected_member_size_bytes, + expected_crc32=expected_member_crc32, + ) + with archive.open(info) as member: + digest, size = _sha256_stream(member, chunk_size=chunk_size) + if ( + expected_member_size_bytes is not None + and size != expected_member_size_bytes + ): + raise ValueError("ASEC 2023 person member byte length mismatch.") + if expected_member_sha256 is not None and digest != expected_member_sha256: + raise ValueError("ASEC 2023 person member SHA-256 mismatch.") + with archive.open(info) as member: + return pd.read_csv( + member, + usecols=usecols, + dtype={"PERIDNUM": "string"}, + low_memory=False, + ) + + _verify_file( + source, + label="ASEC 2023 person member", + expected_size_bytes=expected_member_size_bytes, + expected_sha256=expected_member_sha256, + chunk_size=chunk_size, + ) + return pd.read_csv( + source, + usecols=usecols, + dtype={"PERIDNUM": "string"}, + low_memory=False, + ) + + +def _assert_pinned_source_audit(audit: dict[str, int | float]) -> None: + expected_counts = { + "raw_rows": ASEC_2023_WEEKS_UNEMPLOYED_RAW_ROWS, + "unique_keys": ASEC_2023_WEEKS_UNEMPLOYED_UNIQUE_KEYS, + "positive_rows": ASEC_2023_WEEKS_UNEMPLOYED_POSITIVE_ROWS, + } + for key, expected in expected_counts.items(): + if int(audit[key]) != expected: + raise ValueError( + f"Pinned ASEC weeks-unemployed audit drifted for {key}: " + f"expected {expected}, got {audit[key]}." + ) + expected_floats = { + "weighted_source_share": ( + ASEC_2023_WEEKS_UNEMPLOYED_WEIGHTED_SOURCE_SHARE, + 1e-12, + ), + # A 172.6m person-week dot product accumulates a harmless ~2e-8 + # binary-float residue on the pinned file. One millionth of a + # person-week remains far tighter than the source's 0.01 precision. + "weighted_weeks": (ASEC_2023_WEEKS_UNEMPLOYED_WEIGHTED_WEEKS, 1e-6), + } + for key, (expected, tolerance) in expected_floats.items(): + if not np.isclose(float(audit[key]), expected, rtol=0.0, atol=tolerance): + raise ValueError( + f"Pinned ASEC weeks-unemployed audit drifted for {key}: " + f"expected {expected}, got {audit[key]}." + ) + + +def load_asec_2023_weeks_unemployed_source( + path: str | Path, + *, + expected_zip_size_bytes: int | None = (ASEC_2023_WEEKS_UNEMPLOYED_ZIP_SIZE_BYTES), + expected_zip_sha256: str | None = ASEC_2023_WEEKS_UNEMPLOYED_ZIP_SHA256, + expected_member_size_bytes: int | None = ( + ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_SIZE_BYTES + ), + expected_member_crc32: str | None = (ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_CRC32), + expected_member_sha256: str | None = (ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_SHA256), + expected_rows: int | None = ASEC_2023_WEEKS_UNEMPLOYED_RAW_ROWS, + expected_unique_keys: int | None = ASEC_2023_WEEKS_UNEMPLOYED_UNIQUE_KEYS, + chunk_size: int = 8 * 1024 * 1024, +) -> pd.DataFrame: + """Load the exact identity/LKWEEKS sidecar and attach its source audit.""" + + source = Path(path).expanduser() + if chunk_size < 1: + raise ValueError("chunk_size must be positive") + raw = _load_csv_source( + source, + expected_zip_size_bytes=expected_zip_size_bytes, + expected_zip_sha256=expected_zip_sha256, + expected_member_size_bytes=expected_member_size_bytes, + expected_member_crc32=expected_member_crc32, + expected_member_sha256=expected_member_sha256, + chunk_size=chunk_size, + ) + missing = sorted( + { + *ASEC_2023_WEEKS_UNEMPLOYED_SOURCE_COLUMNS, + _AUDIT_WEIGHT_COLUMN, + } + - set(raw.columns) + ) + if missing: + raise ValueError(f"ASEC weeks-unemployed source missing column(s): {missing}.") + if expected_rows is not None and len(raw) != expected_rows: + raise ValueError( + "ASEC weeks-unemployed source row count mismatch: expected " + f"{expected_rows}, got {len(raw)}." + ) + + raw["PERIDNUM"] = _fixed_width_peridnum(raw["PERIDNUM"], label="ASEC source") + unique_keys = int(raw["PERIDNUM"].nunique(dropna=False)) + if raw["PERIDNUM"].duplicated(keep=False).any(): + examples = ( + raw.loc[raw["PERIDNUM"].duplicated(keep=False), "PERIDNUM"] + .drop_duplicates() + .tolist()[:5] + ) + raise ValueError( + "ASEC weeks-unemployed PERIDNUM must be unique; duplicate key(s): " + f"{examples}." + ) + if expected_unique_keys is not None and unique_keys != expected_unique_keys: + raise ValueError( + "ASEC weeks-unemployed unique-key count mismatch: expected " + f"{expected_unique_keys}, got {unique_keys}." + ) + + raw["PH_SEQ"] = _strict_integer( + raw["PH_SEQ"], label="ASEC PH_SEQ", minimum=1, maximum=99_999 + ) + raw["P_SEQ"] = _strict_integer( + raw["P_SEQ"], label="ASEC P_SEQ", minimum=0, maximum=16 + ) + raw["A_LINENO"] = _strict_integer( + raw["A_LINENO"], label="ASEC A_LINENO", minimum=1, maximum=16 + ) + weeks = _strict_integer(raw[_SOURCE], label="ASEC LKWEEKS") + valid_weeks = (weeks == -1) | ((weeks >= 0) & (weeks <= 52)) + if not valid_weeks.all(): + rows = np.flatnonzero(~valid_weeks)[:5].tolist() + raise ValueError(f"ASEC LKWEEKS is outside [-1, 52] at row(s): {rows}.") + raw[_SOURCE] = weeks.astype(np.int16) + weights = pd.to_numeric(raw[_AUDIT_WEIGHT_COLUMN], errors="coerce").to_numpy( + dtype=np.float64 + ) + if ( + not np.isfinite(weights).all() + or (weights < 0.0).any() + or float(weights.sum()) <= 0.0 + ): + raise ValueError( + "ASEC weeks-unemployed A_FNLWGT must be finite and nonnegative " + "with positive total mass." + ) + normalized = np.where(weeks == -1, 0, weeks).astype(np.int16) + positive = normalized > 0 + scaled_weights = weights / 100.0 + audit: dict[str, int | float] = { + "raw_rows": int(len(raw)), + "unique_keys": unique_keys, + "positive_rows": int(np.count_nonzero(positive)), + "niu_rows": int(np.count_nonzero(weeks == -1)), + "minimum": int(weeks.min()), + "maximum": int(weeks.max()), + "weighted_source_share": float( + scaled_weights[positive].sum() / scaled_weights.sum() + ), + "weighted_weeks": float(np.dot(scaled_weights, normalized)), + } + pinned_transform = bool( + expected_member_size_bytes == ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_SIZE_BYTES + and expected_member_sha256 == ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_SHA256 + and expected_rows == ASEC_2023_WEEKS_UNEMPLOYED_RAW_ROWS + and expected_unique_keys == ASEC_2023_WEEKS_UNEMPLOYED_UNIQUE_KEYS + ) + if pinned_transform: + _assert_pinned_source_audit(audit) + audit["pinned_transform"] = int(pinned_transform) + result = raw.loc[:, list(ASEC_2023_WEEKS_UNEMPLOYED_SOURCE_COLUMNS)].copy() + result.attrs["source_audit"] = audit + return result + + +def fill_asec_2022_weeks_unemployed_source( + person: pd.DataFrame, + source: pd.DataFrame, +) -> pd.DataFrame: + """Fill only missing 2022 ``LKWEEKS`` via exact Census person identity.""" + + required_person = ("source_year", "PERIDNUM") + missing_person = [column for column in required_person if column not in person] + if missing_person: + raise ValueError( + f"ASEC weeks-unemployed repair requires person column(s): {missing_person}." + ) + missing_source = [ + column + for column in ASEC_2023_WEEKS_UNEMPLOYED_SOURCE_COLUMNS + if column not in source + ] + if missing_source: + raise ValueError( + f"ASEC weeks-unemployed sidecar missing column(s): {missing_source}." + ) + result = person.copy(deep=True) + year = pd.to_numeric(result["source_year"], errors="coerce").to_numpy( + dtype=np.float64 + ) + valid_year = np.isfinite(year) & (year == np.floor(year)) + if not valid_year.all(): + rows = np.flatnonzero(~valid_year)[:5].tolist() + raise ValueError( + f"ASEC weeks-unemployed source_year is invalid at row(s): {rows}." + ) + repair_mask = year.astype(np.int64) == ASEC_2023_WEEKS_UNEMPLOYED_SOURCE_YEAR + if not repair_mask.any(): + return result + + donor = source.copy(deep=True) + donor["PERIDNUM"] = _fixed_width_peridnum(donor["PERIDNUM"], label="ASEC sidecar") + if donor["PERIDNUM"].duplicated(keep=False).any(): + raise ValueError("ASEC weeks-unemployed sidecar PERIDNUM is not unique.") + target_keys = _fixed_width_peridnum( + result.loc[repair_mask, "PERIDNUM"], label="ASEC frame" + ) + donor = donor.set_index("PERIDNUM") + missing_keys = target_keys[~target_keys.isin(donor.index)].drop_duplicates() + if not missing_keys.empty: + raise ValueError( + "ASEC weeks-unemployed sidecar does not cover frame PERIDNUM key(s): " + f"{missing_keys.tolist()[:5]}." + ) + aligned = donor.reindex(target_keys.to_numpy()) + aligned.index = result.index[repair_mask] + + identity_pairs = [ + ( + "PH_SEQ", + "source_household_id" if "source_household_id" in result else "PH_SEQ", + ), + ("P_SEQ", "P_SEQ"), + ("A_LINENO", "A_LINENO"), + ] + for donor_column, frame_column in identity_pairs: + if frame_column not in result: + continue + observed = pd.to_numeric( + result.loc[repair_mask, frame_column], errors="coerce" + ).to_numpy(dtype=np.float64) + expected = pd.to_numeric(aligned[donor_column], errors="coerce").to_numpy( + dtype=np.float64 + ) + mismatch = ( + ~np.isfinite(observed) | ~np.isfinite(expected) | (observed != expected) + ) + if mismatch.any(): + rows = result.index[repair_mask].to_numpy()[mismatch][:5].tolist() + raise ValueError( + "ASEC weeks-unemployed redundant identity mismatch for " + f"{frame_column} against sidecar {donor_column} at row(s): {rows}." + ) + + donor_weeks = _strict_integer(aligned[_SOURCE], label="ASEC sidecar LKWEEKS") + valid_donor_weeks = (donor_weeks == -1) | ((donor_weeks >= 0) & (donor_weeks <= 52)) + if not valid_donor_weeks.all(): + rows = result.index[repair_mask].to_numpy()[~valid_donor_weeks][:5].tolist() + raise ValueError( + f"ASEC sidecar LKWEEKS is outside [-1, 52] at frame row(s): {rows}." + ) + if _SOURCE not in result: + result[_SOURCE] = np.nan + existing = pd.to_numeric( + result.loc[repair_mask, _SOURCE], errors="coerce" + ).to_numpy(dtype=np.float64) + raw_existing = result.loc[repair_mask, _SOURCE] + present = raw_existing.notna().to_numpy() + invalid_existing = present & ( + ~np.isfinite(existing) | (existing != np.floor(existing)) + ) + if invalid_existing.any(): + rows = result.index[repair_mask].to_numpy()[invalid_existing][:5].tolist() + raise ValueError(f"Existing ASEC LKWEEKS is invalid at frame row(s): {rows}.") + mismatch = present & (existing != donor_weeks) + if mismatch.any(): + rows = result.index[repair_mask].to_numpy()[mismatch][:5].tolist() + raise ValueError( + "Existing ASEC LKWEEKS disagrees with the pinned sidecar at frame " + f"row(s): {rows}." + ) + missing = ~present + repair_indices = result.index[repair_mask].to_numpy() + if missing.any(): + result.loc[repair_indices[missing], _SOURCE] = donor_weeks[missing] + result.attrs["weeks_unemployed_source_audit"] = source.attrs.get("source_audit", {}) + return result + + +def _direct_weeks(values: pd.Series, *, label: str) -> np.ndarray: + numeric = pd.to_numeric(values, errors="coerce").to_numpy(dtype=np.float64) + valid = np.isfinite(numeric) & (numeric == np.floor(numeric)) + valid &= (numeric == -1.0) | ((numeric >= 0.0) & (numeric <= 52.0)) + if not valid.all(): + rows = np.flatnonzero(~valid)[:5].tolist() + raise SourceRuntimeError( + f"{label} must contain only finite integer -1 or 0--52 at row(s) {rows}." + ) + return np.where(numeric == -1.0, 0.0, numeric).astype(np.int16) + + +def derive_us_weeks_unemployed_from_manifest( + frame: pd.DataFrame | None, + operation: SourceOperationSpec, + _context: SourceRuntimeContext | None, +) -> pd.DataFrame: + """Apply the retired direct ``LKWEEKS`` carry.""" + + if operation.kind != "derive_weeks_unemployed": + raise SourceRuntimeError( + "US weeks-unemployed derivation received unexpected operation " + f"{operation.kind!r}." + ) + if frame is None: + raise SourceRuntimeError( + "US weeks-unemployed derivation requires the person table." + ) + if dict(operation.parameters) != WEEKS_UNEMPLOYED_DERIVE_PARAMETERS: + raise SourceRuntimeError( + "US weeks-unemployed derivation parameters drifted from the " + "archived LKWEEKS carry." + ) + if _SOURCE not in frame: + raise SourceRuntimeError( + f"US weeks-unemployed derivation requires source column {_SOURCE!r}." + ) + result = frame.copy(deep=True) + result[_OUTPUT] = _direct_weeks(frame[_SOURCE], label="ASEC LKWEEKS") + return result + + +def impute_us_weeks_unemployed_to_puf_support_from_manifest( + frame: pd.DataFrame | None, + operation: SourceOperationSpec, + context: SourceRuntimeContext | None, +) -> pd.DataFrame: + """QRF-replace weeks on the PUF support half only.""" + + if operation.kind != "impute_weeks_unemployed_to_puf_support": + raise SourceRuntimeError( + "US weeks-unemployed PUF imputation received unexpected operation " + f"{operation.kind!r}." + ) + if frame is None: + raise SourceRuntimeError( + "US weeks-unemployed PUF imputation requires the person table." + ) + if dict(operation.parameters) != WEEKS_UNEMPLOYED_PUF_IMPUTATION_PARAMETERS: + raise SourceRuntimeError( + "US weeks-unemployed PUF imputation parameters drifted from the " + "reviewed QRF contract." + ) + if _PERSON_SUPPORT_CHANNEL_COLUMN not in frame: + return frame.copy(deep=True) + + channel = frame[_PERSON_SUPPORT_CHANNEL_COLUMN].astype(str) + unexpected_channels = sorted(set(channel) - {_ASEC_CHANNEL, _PUF_CHANNEL}) + if unexpected_channels: + raise SourceRuntimeError( + "US weeks-unemployed found unsupported support channel(s): " + f"{unexpected_channels}." + ) + asec_mask = channel.eq(_ASEC_CHANNEL) + puf_mask = channel.eq(_PUF_CHANNEL) + if not asec_mask.any() or not puf_mask.any(): + raise SourceRuntimeError( + "US weeks-unemployed PUF imputation requires nonempty ASEC and " + "PUF-tax-detail channels." + ) + + required_columns = [ + _PERSON_WEIGHT_COLUMN, + _OUTPUT, + *(_PREDICTOR_PREFIX + value for value in _REQUIRED_PUF_PREDICTORS), + ] + missing = [column for column in required_columns if column not in frame] + if missing: + raise SourceRuntimeError( + f"US weeks-unemployed PUF imputation is missing column(s): {missing}." + ) + optional_uc_column = _PREDICTOR_PREFIX + _OPTIONAL_UC_PREDICTOR + predictors = list(_REQUIRED_PUF_PREDICTORS) + if optional_uc_column in frame: + predictors.append(_OPTIONAL_UC_PREDICTOR) + predictor_columns = [_PREDICTOR_PREFIX + value for value in predictors] + + training = frame.loc[asec_mask, [*predictor_columns, _OUTPUT]].copy() + training.columns = [*predictors, _OUTPUT] + test = frame.loc[puf_mask, predictor_columns].copy() + test.columns = predictors + weights = pd.to_numeric( + frame.loc[asec_mask, _PERSON_WEIGHT_COLUMN], errors="coerce" + ) + weight_values = weights.to_numpy(dtype=np.float64) + if ( + not np.isfinite(weight_values).all() + or (weight_values < 0.0).any() + or float(weight_values.sum()) <= 0.0 + ): + raise SourceRuntimeError( + "US weeks-unemployed QRF requires finite, nonnegative typed person " + "weights with positive total mass." + ) + + max_train_samples = int(operation.parameters["max_train_samples"]) + n_estimators = int(operation.parameters["n_estimators"]) + seed = context.config.seed if context is not None else 0 + if len(training) > max_train_samples: + sampled_index = training.sample( + n=max_train_samples, + random_state=seed, + ).index + training = training.loc[sampled_index] + weights = weights.loc[sampled_index] + + for column in (*predictors, _OUTPUT): + training[column] = pd.to_numeric(training[column], errors="coerce") + values = training[column].to_numpy(dtype=np.float64) + if not np.isfinite(values).all(): + raise SourceRuntimeError( + f"US weeks-unemployed QRF training column {column!r} contains " + "nonfinite values." + ) + target = training[_OUTPUT].to_numpy(dtype=np.float64) + if not ((target == np.floor(target)) & (target >= 0.0) & (target <= 52.0)).all(): + raise SourceRuntimeError( + "US weeks-unemployed QRF target must be integer-valued in 0--52." + ) + for column in predictors: + test[column] = pd.to_numeric(test[column], errors="coerce") + values = test[column].to_numpy(dtype=np.float64) + if not np.isfinite(values).all(): + raise SourceRuntimeError( + f"US weeks-unemployed QRF prediction column {column!r} contains " + "nonfinite values." + ) + if column == _OPTIONAL_UC_PREDICTOR and (values < 0.0).any(): + raise SourceRuntimeError( + "US weeks-unemployed unemployment-compensation predictor must " + "be nonnegative." + ) + + global QRF + if QRF is None: + from importlib import import_module + + QRF = import_module("populace.fit").QRF + fitted = QRF(n_estimators=n_estimators, seed=seed).fit( + training, + predictors, + [_OUTPUT], + weights=weights.to_numpy(dtype=np.float64), + ) + predictions = fitted.predict(test) + if not isinstance(predictions, pd.DataFrame) or _OUTPUT not in predictions: + raise SourceRuntimeError( + f"US weeks-unemployed QRF prediction is missing {_OUTPUT!r}." + ) + if len(predictions) != int(puf_mask.sum()): + raise SourceRuntimeError( + "US weeks-unemployed QRF returned " + f"{len(predictions)} rows; expected {int(puf_mask.sum())}." + ) + predicted = pd.to_numeric(predictions[_OUTPUT], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not np.isfinite(predicted).all(): + raise SourceRuntimeError( + "US weeks-unemployed QRF produced nonfinite predictions." + ) + lower, upper = operation.parameters["clip"] + predicted = np.rint(np.clip(predicted, float(lower), float(upper))) + if _OPTIONAL_UC_PREDICTOR in predictors: + uc = test[_OPTIONAL_UC_PREDICTOR].to_numpy(dtype=np.float64) + predicted = np.where(uc > 0.0, predicted, 0.0) + if not ( + np.isfinite(predicted).all() + and (predicted == np.floor(predicted)).all() + and (predicted >= 0.0).all() + and (predicted <= 52.0).all() + ): + raise SourceRuntimeError( + "US weeks-unemployed postprocessed predictions left the integer " + "0--52 domain." + ) + result = frame.copy(deep=True) + result[_OUTPUT] = pd.to_numeric(result[_OUTPUT], errors="coerce").astype("float64") + result.loc[puf_mask, _OUTPUT] = predicted + return result + + +def _decoded_strings(values: pd.Series) -> pd.Series: + return values.map( + lambda value: ( + value.decode() + if isinstance(value, (bytes, bytearray, np.bytes_)) + else value + ) + ).astype(str) + + +def _boolean_predictor(values: pd.Series, *, label: str) -> np.ndarray: + if values.isna().any(): + rows = values.index[values.isna()].tolist()[:5] + raise SourceRuntimeError(f"{label} is missing at row(s): {rows}.") + if pd.api.types.is_bool_dtype(values.dtype): + return values.to_numpy(dtype=bool) + numeric = pd.to_numeric(values, errors="coerce").to_numpy(dtype=np.float64) + valid = np.isfinite(numeric) & np.isin(numeric, [0.0, 1.0]) + if not valid.all(): + rows = np.flatnonzero(~valid)[:5].tolist() + raise SourceRuntimeError( + f"{label} must be complete boolean values at row(s): {rows}." + ) + return numeric.astype(bool) + + +def _person_weeks_unemployed_predictors(frame: Frame) -> pd.DataFrame: + person = frame.table("person") + tax_unit = frame.table("tax_unit") + predictors = pd.DataFrame(index=person.index) + + age_column = "age" if "age" in person else "A_AGE" + if age_column not in person: + raise SourceRuntimeError("US weeks-unemployed QRF requires person age.") + age = pd.to_numeric(person[age_column], errors="coerce").to_numpy(dtype=np.float64) + if not (np.isfinite(age) & (age >= 0.0) & (age <= 120.0)).all(): + raise SourceRuntimeError("US weeks-unemployed QRF age must be in 0--120.") + predictors["age"] = age + + if "is_male" in person: + predictors["is_male"] = _boolean_predictor(person["is_male"], label="is_male") + elif "is_female" in person: + predictors["is_male"] = ~_boolean_predictor( + person["is_female"], label="is_female" + ) + elif "A_SEX" in person: + sex = pd.to_numeric(person["A_SEX"], errors="coerce").to_numpy(dtype=np.float64) + if not (np.isfinite(sex) & np.isin(sex, [1.0, 2.0])).all(): + raise SourceRuntimeError("US weeks-unemployed A_SEX must be 1 or 2.") + predictors["is_male"] = sex == 1.0 + else: + raise SourceRuntimeError( + "US weeks-unemployed QRF requires is_male, is_female, or A_SEX." + ) + + if "tax_unit_is_joint" in person: + predictors["tax_unit_is_joint"] = _boolean_predictor( + person["tax_unit_is_joint"], label="tax_unit_is_joint" + ) + else: + if "person_tax_unit_id" not in person: + raise SourceRuntimeError( + "US weeks-unemployed QRF requires person tax-unit linkage." + ) + filing_column = next( + ( + column + for column in ("filing_status_input", "filing_status") + if column in tax_unit + ), + None, + ) + if filing_column is None: + raise SourceRuntimeError( + "US weeks-unemployed QRF requires tax-unit filing status." + ) + filing = _decoded_strings(tax_unit[filing_column]) + joint_by_id = pd.Series( + filing.eq("JOINT").to_numpy(), + index=tax_unit["tax_unit_id"].to_numpy(), + ) + joint = person["person_tax_unit_id"].map(joint_by_id) + if joint.isna().any(): + raise SourceRuntimeError( + "US weeks-unemployed QRF found missing tax-unit filing linkage." + ) + predictors["tax_unit_is_joint"] = joint.to_numpy(dtype=bool) + + role_outputs = { + "is_tax_unit_head": "HEAD", + "is_tax_unit_spouse": "SPOUSE", + "is_tax_unit_dependent": "DEPENDENT", + } + if all(column in person for column in role_outputs): + for column in role_outputs: + predictors[column] = _boolean_predictor(person[column], label=column) + else: + if "tax_unit_role_input" not in person: + raise SourceRuntimeError( + "US weeks-unemployed QRF requires explicit tax-unit roles or " + "tax_unit_role_input." + ) + if person["tax_unit_role_input"].isna().any(): + raise SourceRuntimeError( + "US weeks-unemployed tax_unit_role_input must be complete." + ) + role = _decoded_strings(person["tax_unit_role_input"]) + unexpected = sorted(set(role) - set(role_outputs.values())) + if unexpected: + raise SourceRuntimeError( + f"US weeks-unemployed found unsupported tax-unit role(s): {unexpected}." + ) + for column, value in role_outputs.items(): + predictors[column] = role.eq(value).to_numpy(dtype=bool) + + uc_column: str | None = None + if _OPTIONAL_UC_PREDICTOR in person: + uc_column = _OPTIONAL_UC_PREDICTOR + elif _PERSON_SUPPORT_CHANNEL_COLUMN not in person and "UC_VAL" in person: + uc_column = "UC_VAL" + if uc_column is not None: + uc = pd.to_numeric(person[uc_column], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not (np.isfinite(uc) & (uc >= 0.0)).all(): + raise SourceRuntimeError( + "US weeks-unemployed unemployment compensation must be finite " + "and nonnegative." + ) + predictors[_OPTIONAL_UC_PREDICTOR] = uc + return predictors + + +def _replace_person_table(frame: Frame, person: pd.DataFrame) -> Frame: + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + tables["person"] = person + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + ) + + +def with_us_weeks_unemployed( + frame: Frame, + *, + seed: int, + time_period: int, + asec_2023_source: pd.DataFrame | None = None, +) -> Frame: + """Repair the 2022 source, carry ASEC weeks, and replace the PUF half.""" + + if frame.schema != US_SCHEMA: + raise ValueError("US weeks unemployed requires the US schema.") + person = frame.table("person") + stage_person = person.copy(deep=True) + has_2022 = False + if "source_year" in stage_person: + source_year = pd.to_numeric(stage_person["source_year"], errors="coerce") + has_2022 = bool(source_year.eq(ASEC_2023_WEEKS_UNEMPLOYED_SOURCE_YEAR).any()) + needs_2022_source = bool( + has_2022 + and ( + _SOURCE not in stage_person + or stage_person.loc[ + pd.to_numeric(stage_person["source_year"], errors="coerce").eq( + ASEC_2023_WEEKS_UNEMPLOYED_SOURCE_YEAR + ), + _SOURCE, + ] + .isna() + .any() + ) + ) + if asec_2023_source is not None and has_2022: + stage_person = fill_asec_2022_weeks_unemployed_source( + stage_person, asec_2023_source + ) + elif needs_2022_source: + raise ValueError( + "US weeks-unemployed stage requires the pinned ASEC 2023 sidecar " + "to repair missing source-year 2022 LKWEEKS." + ) + + stage_person[_PERSON_WEIGHT_COLUMN] = frame.resolve_weights("person").values + if _PERSON_SUPPORT_CHANNEL_COLUMN in stage_person: + # Predictor construction only needs modeled person/tax-unit columns; + # keep the temporary reserved person_weight column outside Frame. + predictors = _person_weeks_unemployed_predictors(frame) + for column in predictors: + stage_person[_PREDICTOR_PREFIX + column] = predictors[column].to_numpy() + output = run_source_stage( + us_weeks_unemployed_stage_spec(), + tables={"person": stage_person}, + operation_handlers={ + "derive_weeks_unemployed": derive_us_weeks_unemployed_from_manifest, + "impute_weeks_unemployed_to_puf_support": ( + impute_us_weeks_unemployed_to_puf_support_from_manifest + ), + }, + config=SourceRuntimeConfig(seed=int(seed), target_year=int(time_period)), + ).drop( + columns=[ + _PERSON_WEIGHT_COLUMN, + *(_PREDICTOR_PREFIX + value for value in WEEKS_UNEMPLOYED_PUF_PREDICTORS), + ], + errors="ignore", + ) + return _replace_person_table(frame, output) + + +def us_weeks_unemployed_summary(frame: Frame) -> dict[str, object]: + """Return integer-domain, source-reconciliation, and channel diagnostics.""" + + person = frame.table("person") + if _OUTPUT not in person: + raise ValueError(f"US weeks-unemployed summary requires {_OUTPUT!r}.") + values = pd.to_numeric(person[_OUTPUT], errors="coerce").to_numpy(dtype=np.float64) + weights = np.asarray(frame.resolve_weights("person").values, dtype=np.float64) + if ( + not np.isfinite(weights).all() + or (weights < 0.0).any() + or float(weights.sum()) <= 0.0 + ): + raise ValueError( + "US weeks-unemployed summary requires finite nonnegative person " + "weights with positive total mass." + ) + finite = np.isfinite(values) + integer = finite & (values == np.rint(values)) + in_range = integer & (values >= 0.0) & (values <= 52.0) + positive = in_range & (values > 0.0) + + if _PERSON_SUPPORT_CHANNEL_COLUMN in person: + channel = _decoded_strings(person[_PERSON_SUPPORT_CHANNEL_COLUMN]).to_numpy() + else: + channel = np.full(len(person), _ASEC_CHANNEL, dtype=object) + channels: dict[str, dict[str, float | int]] = {} + for name in (_ASEC_CHANNEL, _PUF_CHANNEL): + mask = channel == name + channel_weight = float(weights[mask].sum()) + channels[name] = { + "rows": int(np.count_nonzero(mask)), + "positive_rows": int(np.count_nonzero(mask & positive)), + "positive_share": ( + float(weights[mask & positive].sum()) / channel_weight + if channel_weight > 0.0 + else 0.0 + ), + "weighted_weeks": float(np.dot(weights[mask], np.nan_to_num(values[mask]))), + "weighted_mean_weeks": ( + float(np.dot(weights[mask], np.nan_to_num(values[mask]))) + / channel_weight + if channel_weight > 0.0 + else 0.0 + ), + "positive_share_band": list(_CHANNEL_POSITIVE_SHARE_BANDS[name]), + "weighted_mean_weeks_band": list( + _CHANNEL_WEIGHTED_MEAN_WEEKS_BANDS[name] + ), + } + + source_missing = _SOURCE not in person + source_invalid = source_mismatch = 0 + if not source_missing: + source_raw = pd.to_numeric(person[_SOURCE], errors="coerce").to_numpy( + dtype=np.float64 + ) + asec_mask = channel == _ASEC_CHANNEL + source_valid = np.isfinite(source_raw) & (source_raw == np.floor(source_raw)) + source_valid &= (source_raw == -1.0) | ( + (source_raw >= 0.0) & (source_raw <= 52.0) + ) + source_invalid = int(np.count_nonzero(asec_mask & ~source_valid)) + expected = np.where(source_raw == -1.0, 0.0, source_raw) + source_mismatch = int( + np.count_nonzero(asec_mask & source_valid & finite & (values != expected)) + ) + puf_uc_zero_mismatch = 0 + if _OPTIONAL_UC_PREDICTOR in person: + uc = pd.to_numeric(person[_OPTIONAL_UC_PREDICTOR], errors="coerce").to_numpy( + dtype=np.float64 + ) + puf_mask = channel == _PUF_CHANNEL + puf_uc_zero_mismatch = int( + np.count_nonzero( + puf_mask & np.isfinite(uc) & (uc <= 0.0) & finite & (values != 0.0) + ) + ) + + return { + "rows": int(len(person)), + "missing_or_nonfinite": int(np.count_nonzero(~finite)), + "noninteger": int(np.count_nonzero(finite & ~integer)), + "out_of_range": int(np.count_nonzero(integer & ~in_range)), + "positive_rows": int(np.count_nonzero(positive)), + "positive_share": float(weights[positive].sum() / weights.sum()), + "weighted_weeks": float(np.dot(weights, np.nan_to_num(values))), + "source_missing": source_missing, + "source_invalid": source_invalid, + "source_mismatch_count": source_mismatch, + "puf_uc_zero_mismatch_count": puf_uc_zero_mismatch, + "channels": channels, + } + + +def us_weeks_unemployed_signal_gate(frame: Frame) -> GateResult: + """Require exact ASEC carry and integer, nondefault signal on both halves.""" + + person = frame.table("person") + if _OUTPUT not in person: + return GateResult( + name="weeks_unemployed_signal", + passed=False, + failures=(f"person column missing: {_OUTPUT}.",), + details={"missing": [_OUTPUT]}, + ) + try: + summary = us_weeks_unemployed_summary(frame) + except (TypeError, ValueError) as error: + return GateResult( + name="weeks_unemployed_signal", + passed=False, + failures=(str(error),), + details={}, + ) + failures: list[str] = [] + for key, label in ( + ("missing_or_nonfinite", "missing or nonfinite value(s)"), + ("noninteger", "noninteger value(s)"), + ("out_of_range", "value(s) outside 0--52"), + ): + if int(summary[key]): + failures.append(f"{_OUTPUT}: {summary[key]} {label}.") + if bool(summary["source_missing"]): + failures.append("ASEC LKWEEKS source column is missing.") + if int(summary["source_invalid"]): + failures.append( + f"ASEC LKWEEKS has {summary['source_invalid']} invalid source value(s)." + ) + if int(summary["source_mismatch_count"]): + failures.append( + f"{_OUTPUT} has {summary['source_mismatch_count']} ASEC source " + "reconciliation mismatch(es)." + ) + if int(summary["puf_uc_zero_mismatch_count"]): + failures.append( + f"{_OUTPUT} has {summary['puf_uc_zero_mismatch_count']} PUF row(s) " + "positive without unemployment compensation." + ) + has_support_channels = _PERSON_SUPPORT_CHANNEL_COLUMN in person + required_channels = ( + (_ASEC_CHANNEL, _PUF_CHANNEL) if has_support_channels else (_ASEC_CHANNEL,) + ) + channels = summary["channels"] + for name in required_channels: + channel = channels[name] + if int(channel["rows"]) == 0: + failures.append(f"{name} weeks-unemployed support is empty.") + elif ( + int(channel["positive_rows"]) == 0 or float(channel["weighted_weeks"]) <= 0 + ): + failures.append(f"{name} weeks_unemployed is default-only.") + else: + positive_share = float(channel["positive_share"]) + positive_low, positive_high = channel["positive_share_band"] + if not (positive_low <= positive_share <= positive_high): + failures.append( + f"{name} weeks_unemployed positive share " + f"{positive_share:.6f} outside plausibility band " + f"[{positive_low}, {positive_high}]." + ) + weighted_mean = float(channel["weighted_mean_weeks"]) + mean_low, mean_high = channel["weighted_mean_weeks_band"] + if not (mean_low <= weighted_mean <= mean_high): + failures.append( + f"{name} weeks_unemployed weighted mean weeks " + f"{weighted_mean:.6f} outside plausibility band " + f"[{mean_low}, {mean_high}]." + ) + return GateResult( + name="weeks_unemployed_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) diff --git a/packages/populace-build/src/populace/build/us_runtime/wic_claim.py b/packages/populace-build/src/populace/build/us_runtime/wic_claim.py new file mode 100644 index 00000000..10d9f6ba --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/wic_claim.py @@ -0,0 +1,688 @@ +"""Source-backed WIC claim propensity from official FNS category rates. + +The retired eCPS pipeline calculated PolicyEngine-US ``wic_category_str`` and +then drew ``would_claim_wic`` from the USDA FNS participation rate for that +category. This stage restores that exported input without inventing a survey +receipt anchor: the hermetic ASEC spine has the age, sex, parent, family, and +newly seeded pregnancy inputs needed to reproduce the demographic categories, +while the official CY2022 FNS estimates supply the rates. + +PolicyEngine-US 1.764.6 evaluates categories in this order: pregnant, +breastfeeding mother of an infant, postpartum mother, infant, child, none. +There is no breastfeeding assessment in the hermetic sources. Rather than +invent one, this stage assigns every female parent whose family includes an +infant to FNS's all-postpartum category. Pregnancy remains first, so this +stage deliberately runs after the pregnancy stage and corrects the retired +file's accidental ordering (it calculated WIC categories before seeding +pregnancy). + +Draws are seeded blake2b hashes keyed by stable source-person identity. PUF +support clones therefore retain identical claim flags, and reruns with the +same build seed are bit-reproducible. The frame wrapper passes an existing +surface through only when it exactly equals a fresh deterministic derivation; +this heals the retired pre-pregnancy category ordering while remaining +idempotent for the same build seed. +""" + +from __future__ import annotations + +import hashlib +from collections.abc import Mapping +from importlib.resources import files + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import ( + SourceOperationSpec, + SourceStageSpec, + load_source_manifest, +) +from populace.build.source_runtime import ( + SourceRuntimeConfig, + SourceRuntimeContext, + SourceRuntimeError, + run_source_stage, +) +from populace.frame import Frame +from populace.frame.units import US_SCHEMA + +__all__ = [ + "US_WIC_CLAIM_NONCONSTANT_PERSON_COLUMNS", + "US_WIC_CLAIM_OUTPUT_COLUMNS", + "US_WIC_CLAIM_REQUIRED_SOURCE_COLUMNS", + "US_WIC_CLAIM_STAGE_NAME", + "WIC_CLAIM_ARCHIVED_DERIVATION_URL", + "WIC_CLAIM_ARCHIVED_PARAMETERS_URL", + "WIC_CLAIM_ARCHIVED_RANDOMNESS_URL", + "WIC_CLAIM_FNS_SOURCE_URL", + "derive_us_wic_claim_from_manifest", + "us_wic_claim_signal_gate", + "us_wic_claim_stage_spec", + "us_wic_claim_summary", + "with_us_wic_claim_input", +] + +_ARCHIVED_DATA_REPOSITORY = "policyengine-" + "us-data" +_ARCHIVED_ROOT = ( + "https://github.com/PolicyEngine/" + f"{_ARCHIVED_DATA_REPOSITORY}/blob/" + "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe/" + "policyengine_" + "us_data/" +) +WIC_CLAIM_ARCHIVED_DERIVATION_URL = _ARCHIVED_ROOT + "datasets/cps/cps.py#L684-L691" +WIC_CLAIM_ARCHIVED_PARAMETERS_URL = ( + _ARCHIVED_ROOT + "parameters/take_up/wic_takeup.yaml#L1-L33" +) +WIC_CLAIM_ARCHIVED_RANDOMNESS_URL = _ARCHIVED_ROOT + "utils/randomness.py#L5-L28" +WIC_CLAIM_FNS_SOURCE_URL = ( + "https://fns-prod.azureedge.us/sites/default/files/resource-files/" + "wic-eer-2022-summary.pdf" +) + +US_WIC_CLAIM_STAGE_NAME = "wic_claim_input" +US_WIC_CLAIM_OUTPUT_COLUMNS: tuple[str, ...] = ("would_claim_wic",) +US_WIC_CLAIM_NONCONSTANT_PERSON_COLUMNS = US_WIC_CLAIM_OUTPUT_COLUMNS + +# These are persisted PolicyEngine-facing columns already carried or derived +# on the ASEC spine. ``own_children_in_household`` supplies ``is_parent``'s +# measured fact; family minimum age reproduces the engine's category formula. +US_WIC_CLAIM_REQUIRED_SOURCE_COLUMNS: tuple[str, ...] = ( + "age", + "is_female", + "is_pregnant", + "own_children_in_household", + "person_family_id", +) + +_OUTPUT = US_WIC_CLAIM_OUTPUT_COLUMNS[0] +_PERSON_WEIGHT_COLUMN = "person_weight" +_PERSON_SUPPORT_CHANNEL_COLUMN = "person_support_channel" +_PERSON_SUPPORT_SOURCE_ID_COLUMN = "person_support_source_id" +_SOURCE_IDENTITY_COLUMNS = ( + "source_year", + "source_household_id", + "source_person_id", +) + +_CATEGORY_RATES: dict[str, float] = { + "pregnant": 0.456, + "postpartum": 0.689, + "breastfeeding": 0.663, + "infant": 0.784, + "child": 0.460, + "none": 0.0, +} +_ACTIVE_CATEGORIES = ("pregnant", "postpartum", "infant", "child", "none") +_CATEGORY_ASSIGNMENT_ORDER = _ACTIVE_CATEGORIES +_EXPECTED_OPERATION_PARAMETERS = { + "seed_from_build_config": True, + "category_rates": { + "source": WIC_CLAIM_FNS_SOURCE_URL, + "vintage": "CY2022", + "values": _CATEGORY_RATES, + }, +} + +# The locked 2022-2024 ASEC pools produce an all-person claim share around +# 3.4%. Per-category bands bracket the official probability while remaining +# wide enough for weighted stochastic sampling. +_WEIGHTED_CLAIM_SHARE_BAND = (0.015, 0.060) +_CATEGORY_WEIGHTED_CLAIM_SHARE_BANDS: dict[str, tuple[float, float]] = { + "pregnant": (0.25, 0.67), + "postpartum": (0.45, 0.90), + "infant": (0.55, 0.96), + "child": (0.30, 0.63), + "none": (0.0, 0.0), +} + + +def _validate_operation_parameters(operation: SourceOperationSpec) -> dict[str, float]: + """Validate the exact FNS vintage, source, categories, and rates.""" + + parameters = dict(operation.parameters) + if set(parameters) != {"seed_from_build_config", "category_rates"}: + raise SourceRuntimeError( + "US WIC claim derivation requires exactly seed_from_build_config " + "and category_rates manifest parameters." + ) + if parameters["seed_from_build_config"] is not True: + raise SourceRuntimeError( + "US WIC claim derivation requires seed_from_build_config=true." + ) + declared = parameters["category_rates"] + if not isinstance(declared, Mapping): + raise SourceRuntimeError( + "US WIC claim category_rates must be a sourced manifest mapping." + ) + if set(declared) != {"source", "vintage", "values"}: + raise SourceRuntimeError( + "US WIC claim category_rates requires exactly source, vintage, and values." + ) + if declared["source"] != WIC_CLAIM_FNS_SOURCE_URL: + raise SourceRuntimeError( + "US WIC claim rates must cite the locked official FNS CY2022 PDF." + ) + if declared["vintage"] != "CY2022": + raise SourceRuntimeError( + "US WIC claim rates must use the locked CY2022 vintage." + ) + values = declared["values"] + if not isinstance(values, Mapping) or set(values) != set(_CATEGORY_RATES): + raise SourceRuntimeError( + "US WIC claim rates must declare exactly pregnant, postpartum, " + "breastfeeding, infant, child, and none." + ) + validated: dict[str, float] = {} + for category, expected in _CATEGORY_RATES.items(): + value = values[category] + if isinstance(value, bool) or not isinstance(value, (int, float)): + raise SourceRuntimeError( + f"US WIC claim rate for {category!r} must be numeric." + ) + rate = float(value) + if rate != expected: + raise SourceRuntimeError( + f"US WIC claim rate for {category!r} must equal the official " + f"CY2022 value {expected}, got {rate}." + ) + validated[category] = rate + return validated + + +def us_wic_claim_stage_spec() -> SourceStageSpec: + """Load and validate the packaged WIC-claim source-stage declaration.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + stage_map = manifest.stage_map() + if US_WIC_CLAIM_STAGE_NAME not in stage_map: + raise ValueError( + f"US source manifest declares no {US_WIC_CLAIM_STAGE_NAME!r} stage." + ) + spec = stage_map[US_WIC_CLAIM_STAGE_NAME] + if tuple(spec.outputs) != US_WIC_CLAIM_OUTPUT_COLUMNS: + raise ValueError( + f"{US_WIC_CLAIM_STAGE_NAME!r} manifest outputs do not match the " + "runtime-owned WIC claim family." + ) + if [operation.kind for operation in spec.operations] != [ + "read_table", + "derive_wic_claim", + ]: + raise ValueError( + f"{US_WIC_CLAIM_STAGE_NAME!r} must contain only read_table then " + "derive_wic_claim." + ) + if dict(spec.operations[0].parameters) != { + "table": "person", + "weight": _PERSON_WEIGHT_COLUMN, + }: + raise ValueError( + f"{US_WIC_CLAIM_STAGE_NAME!r} must read the weighted person table." + ) + try: + _validate_operation_parameters(spec.operations[1]) + except SourceRuntimeError as exc: + raise ValueError( + f"{US_WIC_CLAIM_STAGE_NAME!r} manifest parameter drift: {exc}" + ) from exc + artifact_sources = { + str(artifact.get("source")) + for artifact in spec.artifacts + if artifact.get("source") is not None + } + if WIC_CLAIM_FNS_SOURCE_URL not in artifact_sources: + raise ValueError( + f"{US_WIC_CLAIM_STAGE_NAME!r} does not pin the official FNS PDF artifact." + ) + return spec + + +def _strict_numeric( + person: pd.DataFrame, + column: str, + *, + minimum: float, + maximum: float | None = None, + integer: bool = False, +) -> np.ndarray: + numeric = pd.to_numeric(person[column], errors="coerce").to_numpy(dtype=np.float64) + valid = np.isfinite(numeric) & (numeric >= minimum) + if maximum is not None: + valid &= numeric <= maximum + if integer: + valid &= numeric == np.floor(numeric) + if not valid.all(): + rows = np.flatnonzero(~valid)[:5].tolist() + raise SourceRuntimeError( + f"US WIC claim derivation requires valid {column!r}; invalid row(s): " + f"{rows}." + ) + return numeric + + +def _strict_bool(person: pd.DataFrame, column: str) -> np.ndarray: + values = pd.to_numeric(person[column], errors="coerce").to_numpy(dtype=np.float64) + valid = np.isfinite(values) & np.isin(values, np.array([0.0, 1.0])) + if not valid.all(): + rows = np.flatnonzero(~valid)[:5].tolist() + raise SourceRuntimeError( + f"US WIC claim derivation requires boolean {column!r}; invalid " + f"row(s): {rows}." + ) + return values.astype(bool) + + +def _wic_categories(person: pd.DataFrame) -> np.ndarray: + """Reproduce the PE-US category precedence from persisted inputs.""" + + missing = [ + column + for column in US_WIC_CLAIM_REQUIRED_SOURCE_COLUMNS + if column not in person.columns + ] + if missing: + raise SourceRuntimeError( + f"US WIC claim derivation requires person source column(s): {missing}." + ) + if person["person_family_id"].isna().any(): + rows = np.flatnonzero(person["person_family_id"].isna().to_numpy())[:5].tolist() + raise SourceRuntimeError( + "US WIC claim derivation requires nonmissing person_family_id; " + f"invalid row(s): {rows}." + ) + + # ASEC age is integer years. Enforcing that source contract makes the + # collapsed ``< 1`` infant-family test equivalent to PE-US's non- + # breastfeeding postpartum ``< 0.5`` branch on this hermetic surface. + age = _strict_numeric( + person, + "age", + minimum=0.0, + maximum=120.0, + integer=True, + ) + female = _strict_bool(person, "is_female") + pregnant = _strict_bool(person, "is_pregnant") + own_children = _strict_numeric( + person, + "own_children_in_household", + minimum=0.0, + integer=True, + ) + if np.any(pregnant & ~female): + rows = np.flatnonzero(pregnant & ~female)[:5].tolist() + raise SourceRuntimeError( + "US WIC claim derivation found is_pregnant=true for a nonfemale " + f"person; invalid row(s): {rows}." + ) + + min_family_age = ( + pd.Series(age, index=person.index) + .groupby(person["person_family_id"], sort=False) + .transform("min") + .to_numpy(dtype=np.float64) + ) + postpartum = female & (own_children > 0.0) & (min_family_age < 1.0) + return np.select( + [pregnant, postpartum, age < 1.0, age < 5.0], + ["pregnant", "postpartum", "infant", "child"], + default="none", + ) + + +def _stable_person_keys(person: pd.DataFrame) -> pd.Series: + present = [column in person.columns for column in _SOURCE_IDENTITY_COLUMNS] + if any(present) and not all(present): + missing = [ + column + for column, is_present in zip( + _SOURCE_IDENTITY_COLUMNS, present, strict=True + ) + if not is_present + ] + raise SourceRuntimeError( + "US WIC claim stable source identity is partial; missing column(s): " + f"{missing}." + ) + if all(present): + identity = person.loc[:, list(_SOURCE_IDENTITY_COLUMNS)] + if identity.isna().any(axis=None): + rows = np.flatnonzero(identity.isna().any(axis=1).to_numpy())[:5].tolist() + raise SourceRuntimeError( + "US WIC claim stable source identity contains missing values at " + f"row(s): {rows}." + ) + return ( + identity["source_year"].astype(str) + + ":" + + identity["source_household_id"].astype(str) + + ":" + + identity["source_person_id"].astype(str) + ) + if _PERSON_SUPPORT_SOURCE_ID_COLUMN in person.columns: + source_id = person[_PERSON_SUPPORT_SOURCE_ID_COLUMN] + if source_id.isna().any(): + rows = np.flatnonzero(source_id.isna().to_numpy())[:5].tolist() + raise SourceRuntimeError( + "US WIC claim support source identity contains missing values at " + f"row(s): {rows}." + ) + return "support:" + source_id.astype(str) + if "person_id" not in person.columns or person["person_id"].isna().any(): + raise SourceRuntimeError( + "US WIC claim derivation requires person_id when stable source " + "identity is unavailable." + ) + return "person:" + person["person_id"].astype(str) + + +def _stable_person_draws(person: pd.DataFrame, *, seed: int) -> np.ndarray: + keys = _stable_person_keys(person) + denominator = float(2**64) + return np.asarray( + [ + int.from_bytes( + hashlib.blake2b( + f"{seed}:{_OUTPUT}:{key}".encode(), + digest_size=8, + ).digest(), + byteorder="big", + signed=False, + ) + / denominator + for key in keys + ], + dtype=np.float64, + ) + + +def derive_us_wic_claim_from_manifest( + frame: pd.DataFrame | None, + operation: SourceOperationSpec, + context: SourceRuntimeContext, +) -> pd.DataFrame: + """Assign ``would_claim_wic`` from source categories and FNS rates.""" + + if operation.kind != "derive_wic_claim": + raise SourceRuntimeError( + f"US WIC claim derivation received unexpected operation {operation.kind!r}." + ) + if frame is None: + raise SourceRuntimeError( + "US WIC claim derivation requires the person table to be read first." + ) + rates = _validate_operation_parameters(operation) + categories = _wic_categories(frame) + draws = _stable_person_draws(frame, seed=int(context.config.seed)) + thresholds = np.fromiter( + (rates[str(category)] for category in categories), + dtype=np.float64, + count=len(categories), + ) + result = frame.copy(deep=True) + result[_OUTPUT] = draws < thresholds + return result + + +def _output_values(person: pd.DataFrame) -> tuple[np.ndarray, int]: + series = person[_OUTPUT] + missing = int(series.isna().sum()) + converted = pd.to_numeric(series, errors="coerce") + coerced_invalid = converted.isna() & ~series.isna() + if coerced_invalid.any(): + rows = np.flatnonzero(coerced_invalid.to_numpy())[:5].tolist() + raise SourceRuntimeError( + f"US WIC claim output {_OUTPUT!r} must be boolean; nonnumeric " + f"value(s) at row(s): {rows}." + ) + numeric = converted.to_numpy(dtype=np.float64) + valid = np.isnan(numeric) | np.isin(numeric, np.array([0.0, 1.0])) + if not valid.all(): + rows = np.flatnonzero(~valid)[:5].tolist() + raise SourceRuntimeError( + f"US WIC claim output {_OUTPUT!r} must be boolean; invalid row(s): {rows}." + ) + return np.nan_to_num(numeric, nan=0.0).astype(bool), missing + + +def with_us_wic_claim_input( + frame: Frame, + *, + seed: int, + time_period: int, +) -> Frame: + """Materialize the source-backed WIC claim input on a US frame.""" + + if frame.schema != US_SCHEMA: + raise ValueError("US WIC claim input requires the US schema.") + + person = frame.table("person") + stage_person = person.copy(deep=True) + stage_person[_PERSON_WEIGHT_COLUMN] = frame.resolve_weights("person").values + output = run_source_stage( + us_wic_claim_stage_spec(), + tables={"person": stage_person}, + operation_handlers={ + "derive_wic_claim": derive_us_wic_claim_from_manifest, + }, + config=SourceRuntimeConfig(seed=int(seed), target_year=int(time_period)), + ) + if output["person_id"].duplicated().any(): + raise ValueError("US WIC claim stage produced duplicate person_id rows.") + aligned = output.set_index("person_id").reindex(person["person_id"]) + if aligned[_OUTPUT].isna().any(): + raise ValueError("US WIC claim stage output does not cover every person.") + expected = aligned[_OUTPUT].to_numpy(dtype=bool) + if _OUTPUT in person: + try: + current, missing = _output_values(person) + except SourceRuntimeError: + current, missing = np.zeros(len(person), dtype=bool), len(person) + if not missing and np.array_equal(current, expected): + return frame + + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + tables["person"][_OUTPUT] = expected + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + ) + + +def _person_weights(frame: Frame) -> np.ndarray: + weights = np.asarray(frame.resolve_weights("person").values, dtype=np.float64) + valid = np.isfinite(weights) & (weights >= 0.0) + if not valid.all() or float(weights.sum()) <= 0.0: + rows = np.flatnonzero(~valid)[:5].tolist() + raise SourceRuntimeError( + "US WIC claim gate requires finite nonnegative person weights with " + f"positive total; invalid row(s): {rows}." + ) + return weights + + +def _clone_diagnostics( + person: pd.DataFrame, + *, + categories: np.ndarray, + claims: np.ndarray, +) -> dict[str, int]: + keys = _stable_person_keys(person) + work = pd.DataFrame( + { + "key": keys.to_numpy(), + "category": categories, + "claim": claims, + } + ) + sizes = work.groupby("key", sort=False).size() + repeated = sizes.index[sizes > 1] + claim_unique = work.groupby("key", sort=False)["claim"].nunique() + category_unique = work.groupby("key", sort=False)["category"].nunique() + result = { + "clone_group_count": int(len(repeated)), + "clone_claim_mismatch_count": int((claim_unique.reindex(repeated) > 1).sum()), + "clone_category_mismatch_count": int( + (category_unique.reindex(repeated) > 1).sum() + ), + } + return result + + +def us_wic_claim_summary(frame: Frame) -> dict[str, object]: + """Return category, channel, and clone diagnostics for the release gate.""" + + person = frame.table("person") + if _OUTPUT not in person: + raise SourceRuntimeError(f"US WIC claim summary requires {_OUTPUT!r}.") + categories = _wic_categories(person) + claims, missing_count = _output_values(person) + weights = _person_weights(frame) + total_weight = float(weights.sum()) + + category_counts: dict[str, int] = {} + category_weights: dict[str, float] = {} + category_claim_shares: dict[str, float] = {} + for category in _ACTIVE_CATEGORIES: + mask = categories == category + category_weight = float(weights[mask].sum()) + category_counts[category] = int(np.count_nonzero(mask)) + category_weights[category] = category_weight + category_claim_shares[category] = ( + float(weights[mask & claims].sum()) / category_weight + if category_weight > 0.0 + else 0.0 + ) + + summary: dict[str, object] = { + "weighted_claim_share": float(weights[claims].sum()) / total_weight, + "weighted_claim_share_band": list(_WEIGHTED_CLAIM_SHARE_BAND), + "positive_count": int(np.count_nonzero(claims)), + "unique_count": int(np.unique(claims).size) if not missing_count else 0, + "missing_count": missing_count, + "category_assignment_order": list(_CATEGORY_ASSIGNMENT_ORDER), + "category_rates": dict(_CATEGORY_RATES), + "category_counts": category_counts, + "category_weights": category_weights, + "category_weighted_claim_shares": category_claim_shares, + "category_weighted_claim_share_bands": { + key: list(value) + for key, value in _CATEGORY_WEIGHTED_CLAIM_SHARE_BANDS.items() + }, + "breastfeeding_source_available": False, + "breastfeeding_rate_validated_but_unassigned": _CATEGORY_RATES["breastfeeding"], + **_clone_diagnostics(person, categories=categories, claims=claims), + } + + if _PERSON_SUPPORT_CHANNEL_COLUMN in person.columns: + channels = ( + person[_PERSON_SUPPORT_CHANNEL_COLUMN] + .fillna("") + .astype(str) + .to_numpy() + ) + channel_shares: dict[str, float] = {} + channel_unique_counts: dict[str, int] = {} + for channel in sorted(set(channels.tolist())): + mask = channels == channel + channel_weight = float(weights[mask].sum()) + channel_shares[channel] = ( + float(weights[mask & claims].sum()) / channel_weight + if channel_weight > 0.0 + else 0.0 + ) + channel_unique_counts[channel] = int(np.unique(claims[mask]).size) + summary["channel_weighted_claim_shares"] = channel_shares + summary["channel_unique_counts"] = channel_unique_counts + summary["channel_weighted_claim_share_band"] = list(_WEIGHTED_CLAIM_SHARE_BAND) + return summary + + +def us_wic_claim_signal_gate(frame: Frame) -> GateResult: + """Require nondefault, category-plausible, clone-consistent WIC claims.""" + + person = frame.table("person") + if _OUTPUT not in person: + return GateResult( + name="wic_claim_input_signal", + passed=False, + failures=(f"person.{_OUTPUT}: missing",), + details={"missing": [_OUTPUT]}, + ) + try: + summary = us_wic_claim_summary(frame) + except SourceRuntimeError as exc: + return GateResult( + name="wic_claim_input_signal", + passed=False, + failures=(str(exc),), + details={}, + ) + + failures: list[str] = [] + if int(summary["missing_count"]): + failures.append(f"{_OUTPUT}: missing values") + if int(summary["unique_count"]) < 2: + failures.append( + f"{_OUTPUT}: constant column — the default-True WIC claim surface " + "carries no source-backed signal" + ) + claim_share = float(summary["weighted_claim_share"]) + low, high = _WEIGHTED_CLAIM_SHARE_BAND + if not low <= claim_share <= high: + failures.append( + f"{_OUTPUT}: weighted claim share {claim_share:.6f} outside " + f"[{low:.3f}, {high:.3f}]" + ) + + category_counts = dict(summary["category_counts"]) + category_weights = dict(summary["category_weights"]) + category_shares = dict(summary["category_weighted_claim_shares"]) + for category in _ACTIVE_CATEGORIES: + if ( + int(category_counts[category]) == 0 + or float(category_weights[category]) <= 0 + ): + failures.append(f"{_OUTPUT}: WIC category {category!r} is absent") + continue + observed = float(category_shares[category]) + category_low, category_high = _CATEGORY_WEIGHTED_CLAIM_SHARE_BANDS[category] + if not category_low <= observed <= category_high: + failures.append( + f"{_OUTPUT}: {category} weighted claim share {observed:.6f} " + f"outside [{category_low:.3f}, {category_high:.3f}]" + ) + + for field, label in ( + ("clone_claim_mismatch_count", "claim"), + ("clone_category_mismatch_count", "category"), + ): + count = int(summary[field]) + if count: + failures.append(f"{_OUTPUT}: {count} support-clone {label} mismatch(es)") + + channel_shares = dict(summary.get("channel_weighted_claim_shares", {})) + channel_unique = dict(summary.get("channel_unique_counts", {})) + for channel, share in channel_shares.items(): + if int(channel_unique[channel]) < 2: + failures.append(f"{_OUTPUT}: {channel} support channel is constant") + if not low <= float(share) <= high: + failures.append( + f"{_OUTPUT}: {channel} weighted claim share {float(share):.6f} " + f"outside [{low:.3f}, {high:.3f}]" + ) + + return GateResult( + name="wic_claim_input_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) diff --git a/packages/populace-build/src/populace/build/us_runtime/workers_compensation.py b/packages/populace-build/src/populace/build/us_runtime/workers_compensation.py new file mode 100644 index 00000000..4b9049f4 --- /dev/null +++ b/packages/populace-build/src/populace/build/us_runtime/workers_compensation.py @@ -0,0 +1,648 @@ +"""CPS ASEC workers' compensation carried from measured annual ``WC_VAL``. + +The retired eCPS pipeline mapped ``WC_VAL`` directly at person grain and kept +workers' compensation out of its separate two-slot disability-benefits sum. +It then used the common eight-predictor, at-most-5,000-person QRF to replace +this CPS-only leaf on the PUF support half. This port preserves that source +separation and uses typed person weights under Populace's weighting contract. +""" + +from __future__ import annotations + +from importlib.resources import files +from typing import Any + +import numpy as np +import pandas as pd + +from populace.build.gates import GateResult +from populace.build.source_manifest import ( + SourceOperationSpec, + SourceStageSpec, + load_source_manifest, +) +from populace.build.source_runtime import ( + SourceRuntimeConfig, + SourceRuntimeContext, + SourceRuntimeError, + run_source_stage, +) +from populace.frame import Frame +from populace.frame.units import US_SCHEMA + +__all__ = [ + "WORKERS_COMPENSATION_ARCHIVED_DERIVATION_URL", + "WORKERS_COMPENSATION_ARCHIVED_PUF_IMPUTATION_URL", + "WORKERS_COMPENSATION_ARCHIVED_PUF_OUTPUTS_URL", + "WORKERS_COMPENSATION_ARCHIVED_SOURCE_COLUMNS_URL", + "US_WORKERS_COMPENSATION_NONCONSTANT_PERSON_COLUMNS", + "US_WORKERS_COMPENSATION_OUTPUT_COLUMNS", + "US_WORKERS_COMPENSATION_REQUIRED_SOURCE_COLUMNS", + "US_WORKERS_COMPENSATION_STAGE_NAME", + "derive_us_workers_compensation_from_asec", + "derive_us_workers_compensation_from_manifest", + "impute_us_workers_compensation_to_puf_support_from_manifest", + "us_workers_compensation_signal_gate", + "us_workers_compensation_stage_spec", + "us_workers_compensation_summary", + "with_us_workers_compensation", +] + +QRF: Any | None = None + +_ARCHIVED_DATA_REPOSITORY = "policyengine-" + "us-data" +_ARCHIVED_ROOT = ( + "https://github.com/PolicyEngine/" + f"{_ARCHIVED_DATA_REPOSITORY}/blob/" + "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe/" + "policyengine_" + "us_data/" +) +WORKERS_COMPENSATION_ARCHIVED_DERIVATION_URL = ( + _ARCHIVED_ROOT + "datasets/cps/cps.py#L1559-L1571" +) +WORKERS_COMPENSATION_ARCHIVED_SOURCE_COLUMNS_URL = ( + _ARCHIVED_ROOT + "datasets/cps/census_cps.py#L306-L381" +) +WORKERS_COMPENSATION_ARCHIVED_PUF_OUTPUTS_URL = ( + _ARCHIVED_ROOT + "datasets/cps/extended_cps.py#L135-L194" +) +WORKERS_COMPENSATION_ARCHIVED_PUF_IMPUTATION_URL = ( + _ARCHIVED_ROOT + "datasets/cps/extended_cps.py#L639-L745" +) + +US_WORKERS_COMPENSATION_STAGE_NAME = "workers_compensation_input" +US_WORKERS_COMPENSATION_OUTPUT_COLUMNS: tuple[str, ...] = ("workers_compensation",) +US_WORKERS_COMPENSATION_NONCONSTANT_PERSON_COLUMNS = ( + US_WORKERS_COMPENSATION_OUTPUT_COLUMNS +) +US_WORKERS_COMPENSATION_REQUIRED_SOURCE_COLUMNS: tuple[str, ...] = ("WC_VAL",) + +_OUTPUT = US_WORKERS_COMPENSATION_OUTPUT_COLUMNS[0] +_PERSON_WEIGHT_COLUMN = "person_weight" +_PERSON_SUPPORT_CHANNEL_COLUMN = "person_support_channel" +_BASE_ASEC_SUPPORT_CHANNEL = "asec" +_PUF_TAX_DETAIL_SUPPORT_CHANNEL = "puf_tax_detail" +_PREDICTORS: tuple[str, ...] = ( + "age", + "is_male", + "has_esi", + "tax_unit_is_joint", + "tax_unit_count_dependents", + "employment_income", + "self_employment_income", + "social_security", +) +_PREDICTOR_PREFIX = "workers_compensation_predictor_" +_EXPECTED_DIRECT_PARAMETERS = { + "source": "WC_VAL", + "output": _OUTPUT, +} +_PUF_IMPUTATION_PARAMETER_KEYS = frozenset( + { + "predictors", + "max_train_samples", + "n_estimators", + "seed_from_build_config", + "weight", + } +) +# Broad source-plausibility bounds: reject default/near-universal surfaces, +# while leaving the exact raw-source and support-channel tests to pin semantics. +_NONZERO_SHARE_BAND = (0.0005, 0.05) +_CHANNEL_NONZERO_SHARE_BANDS = { + _BASE_ASEC_SUPPORT_CHANNEL: (0.001, 0.01), + _PUF_TAX_DETAIL_SUPPORT_CHANNEL: (0.0002, 0.10), +} + + +def us_workers_compensation_stage_spec() -> SourceStageSpec: + """Load and validate the packaged workers-compensation stage.""" + + manifest = load_source_manifest( + files("populace.build.us").joinpath("source_stages.json") + ) + stage_map = manifest.stage_map() + if US_WORKERS_COMPENSATION_STAGE_NAME not in stage_map: + raise ValueError( + "US source manifest declares no " + f"{US_WORKERS_COMPENSATION_STAGE_NAME!r} stage." + ) + spec = stage_map[US_WORKERS_COMPENSATION_STAGE_NAME] + missing = sorted(set(US_WORKERS_COMPENSATION_OUTPUT_COLUMNS) - set(spec.outputs)) + if missing: + raise ValueError( + f"{US_WORKERS_COMPENSATION_STAGE_NAME!r} manifest stage does not " + f"declare output(s) {missing}; the runtime and manifest have drifted." + ) + return spec + + +def _strict_numeric_source( + person: pd.DataFrame, + column: str, + *, + nonnegative: bool, +) -> np.ndarray: + if column not in person.columns: + raise SourceRuntimeError( + f"US workers-compensation derivation requires ASEC source column {column!r}." + ) + values = pd.to_numeric(person[column], errors="coerce").to_numpy(dtype=np.float64) + nonfinite = int(np.count_nonzero(~np.isfinite(values))) + if nonfinite: + raise SourceRuntimeError( + f"US workers-compensation source {column!r} contains {nonfinite} " + "nonnumeric or nonfinite value(s)." + ) + if nonnegative: + negative = int(np.count_nonzero(values < 0.0)) + if negative: + raise SourceRuntimeError( + f"US workers-compensation source {column!r} contains {negative} " + "negative value(s)." + ) + return values + + +def derive_us_workers_compensation_from_asec( + person: pd.DataFrame, + *, + source: str = "WC_VAL", + output_column: str = _OUTPUT, +) -> pd.DataFrame: + """Apply the retired direct annual ``WC_VAL`` mapping.""" + + result = person.copy(deep=True) + result[output_column] = _strict_numeric_source( + person, + source, + nonnegative=True, + ) + return result + + +def derive_us_workers_compensation_from_manifest( + frame: pd.DataFrame | None, + operation: SourceOperationSpec, + _context: SourceRuntimeContext | None, +) -> pd.DataFrame: + """Interpret the manifest's exact ASEC ``WC_VAL`` derivation.""" + + if operation.kind != "derive_workers_compensation": + raise SourceRuntimeError( + "US workers-compensation derivation received unexpected operation " + f"{operation.kind!r}." + ) + if frame is None: + raise SourceRuntimeError( + "US workers-compensation derivation requires the person table first." + ) + parameters = dict(operation.parameters) + if parameters != _EXPECTED_DIRECT_PARAMETERS: + raise SourceRuntimeError( + "US workers-compensation derivation drifted from the archived method: " + f"expected {_EXPECTED_DIRECT_PARAMETERS}, got {parameters}." + ) + return derive_us_workers_compensation_from_asec( + frame, + source=parameters["source"], + output_column=parameters["output"], + ) + + +def impute_us_workers_compensation_to_puf_support_from_manifest( + frame: pd.DataFrame | None, + operation: SourceOperationSpec, + context: SourceRuntimeContext | None, +) -> pd.DataFrame: + """QRF-impute the CPS-only leaf onto PUF-clone people.""" + + if operation.kind != "impute_workers_compensation_to_puf_support": + raise SourceRuntimeError( + "US workers-compensation PUF imputation received unexpected operation " + f"{operation.kind!r}." + ) + if frame is None: + raise SourceRuntimeError( + "US workers-compensation PUF imputation requires the person table first." + ) + unexpected = sorted(set(operation.parameters) - _PUF_IMPUTATION_PARAMETER_KEYS) + missing_parameters = sorted( + _PUF_IMPUTATION_PARAMETER_KEYS - set(operation.parameters) + ) + if unexpected or missing_parameters: + raise SourceRuntimeError( + "US workers-compensation PUF imputation parameters must match the " + f"archived method; missing={missing_parameters}, " + f"unexpected={unexpected}." + ) + if _PERSON_SUPPORT_CHANNEL_COLUMN not in frame.columns: + return frame.copy(deep=True) + + predictors = tuple(str(value) for value in operation.parameters["predictors"]) + if predictors != _PREDICTORS: + raise SourceRuntimeError( + "US workers-compensation PUF predictors drifted from the archived " + f"method: expected {list(_PREDICTORS)}, got {list(predictors)}." + ) + if operation.parameters["weight"] != _PERSON_WEIGHT_COLUMN: + raise SourceRuntimeError( + "US workers-compensation PUF imputation must use typed person weights." + ) + if operation.parameters["seed_from_build_config"] is not True: + raise SourceRuntimeError( + "US workers-compensation PUF imputation seed must come from the build " + "config." + ) + max_train_samples = int(operation.parameters["max_train_samples"]) + n_estimators = int(operation.parameters["n_estimators"]) + if max_train_samples <= 0 or n_estimators <= 0: + raise SourceRuntimeError( + "US workers-compensation PUF max_train_samples and n_estimators must " + "be positive." + ) + + predictor_columns = [_PREDICTOR_PREFIX + name for name in predictors] + required = [_PERSON_WEIGHT_COLUMN, *predictor_columns, _OUTPUT] + missing = [column for column in required if column not in frame.columns] + if missing: + raise SourceRuntimeError( + f"US workers-compensation PUF imputation is missing column(s): {missing}." + ) + + channel = frame[_PERSON_SUPPORT_CHANNEL_COLUMN].astype(str) + asec_mask = channel == _BASE_ASEC_SUPPORT_CHANNEL + puf_mask = channel == _PUF_TAX_DETAIL_SUPPORT_CHANNEL + if not asec_mask.any() or not puf_mask.any(): + raise SourceRuntimeError( + "US workers-compensation PUF imputation requires nonempty ASEC and " + "PUF-tax-detail support channels." + ) + + training = frame.loc[asec_mask, [*predictor_columns, _OUTPUT]].copy() + training.columns = [*predictors, _OUTPUT] + test = frame.loc[puf_mask, predictor_columns].copy() + test.columns = list(predictors) + weights = pd.to_numeric( + frame.loc[asec_mask, _PERSON_WEIGHT_COLUMN], errors="coerce" + ) + numeric_weights = weights.to_numpy(dtype=np.float64) + if not np.isfinite(numeric_weights).all() or bool((numeric_weights < 0.0).any()): + raise SourceRuntimeError( + "US workers-compensation QRF person weights must be finite and nonnegative." + ) + if float(numeric_weights.sum()) <= 0.0: + raise SourceRuntimeError( + "US workers-compensation QRF person weights sum to zero." + ) + + if len(training) > max_train_samples: + sample = training.sample( + n=max_train_samples, + random_state=(context.config.seed if context is not None else 0), + ).index + training = training.loc[sample] + weights = weights.loc[sample] + + for column in (*predictors, _OUTPUT): + training[column] = pd.to_numeric(training[column], errors="coerce") + if not np.isfinite(training[column].to_numpy(dtype=np.float64)).all(): + raise SourceRuntimeError( + f"US workers-compensation QRF training column {column!r} " + "contains nonfinite values." + ) + for column in predictors: + test[column] = pd.to_numeric(test[column], errors="coerce") + if not np.isfinite(test[column].to_numpy(dtype=np.float64)).all(): + raise SourceRuntimeError( + f"US workers-compensation QRF prediction column {column!r} " + "contains nonfinite values." + ) + + global QRF + if QRF is None: + from importlib import import_module + + QRF = import_module("populace.fit").QRF + seed = context.config.seed if context is not None else 0 + fitted = QRF(n_estimators=n_estimators, seed=seed).fit( + training, + list(predictors), + [_OUTPUT], + weights=weights.to_numpy(dtype=np.float64), + ) + predictions = fitted.predict(test) + if _OUTPUT not in predictions: + raise SourceRuntimeError( + f"US workers-compensation QRF prediction is missing {_OUTPUT!r}." + ) + predicted = pd.to_numeric(predictions[_OUTPUT], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not np.isfinite(predicted).all(): + raise SourceRuntimeError( + f"US workers-compensation QRF produced nonfinite {_OUTPUT} values." + ) + if bool((predicted < 0.0).any()): + raise SourceRuntimeError( + f"US workers-compensation QRF produced negative {_OUTPUT} values." + ) + result = frame.copy(deep=True) + result.loc[puf_mask, _OUTPUT] = predicted + return result + + +def _person_workers_compensation_predictors(frame: Frame) -> pd.DataFrame: + """Build the retired eight second-stage QRF predictors on people.""" + + person = frame.table("person") + tax_unit = frame.table("tax_unit") + predictors = pd.DataFrame(index=person.index) + + def _numeric(*columns: str) -> np.ndarray: + for column in columns: + if column in person: + values = pd.to_numeric(person[column], errors="coerce").to_numpy( + dtype=np.float64 + ) + if not np.isfinite(values).all(): + raise SourceRuntimeError( + "US workers-compensation QRF predictor source " + f"{column!r} contains nonfinite values." + ) + return values + raise SourceRuntimeError( + "US workers-compensation PUF imputation cannot construct a predictor " + f"from any of {list(columns)}." + ) + + predictors["age"] = _numeric("age", "A_AGE") + if "is_male" in person: + predictors["is_male"] = person["is_male"].fillna(False).astype(bool) + elif "is_female" in person: + predictors["is_male"] = ~person["is_female"].fillna(False).astype(bool) + elif "A_SEX" in person: + predictors["is_male"] = _numeric("A_SEX") == 1 + else: + raise SourceRuntimeError( + "US workers-compensation PUF imputation requires is_male, is_female, " + "or measured A_SEX." + ) + if "has_esi" not in person: + raise SourceRuntimeError( + "US workers-compensation PUF imputation requires has_esi." + ) + predictors["has_esi"] = person["has_esi"].fillna(False).astype(bool) + + filing_status_column = next( + ( + column + for column in ("filing_status_input", "filing_status") + if column in tax_unit + ), + None, + ) + if filing_status_column is None: + raise SourceRuntimeError( + "US workers-compensation PUF imputation requires tax-unit filing status." + ) + filing_status = tax_unit[filing_status_column].map( + lambda value: ( + value.decode() if isinstance(value, (bytes, np.bytes_)) else str(value) + ) + ) + joint_by_id = pd.Series( + filing_status.eq("JOINT").to_numpy(), + index=tax_unit["tax_unit_id"].to_numpy(), + ) + predictors["tax_unit_is_joint"] = ( + person["person_tax_unit_id"].map(joint_by_id).fillna(False).astype(bool) + ) + + if "tax_unit_role_input" not in person: + raise SourceRuntimeError( + "US workers-compensation PUF imputation requires tax_unit_role_input " + "to count dependents." + ) + dependent = ( + person["tax_unit_role_input"] + .map( + lambda value: ( + value.decode() if isinstance(value, (bytes, np.bytes_)) else str(value) + ) + ) + .eq("DEPENDENT") + ) + predictors["tax_unit_count_dependents"] = ( + dependent.groupby(person["person_tax_unit_id"]) + .transform("sum") + .to_numpy(dtype=np.float64) + ) + predictors["employment_income"] = _numeric( + "employment_income_before_lsr", "WSAL_VAL" + ) + predictors["self_employment_income"] = _numeric( + "self_employment_income_before_lsr", "SEMP_VAL" + ) + social_security_columns = ( + "social_security_retirement", + "social_security_disability", + "social_security_dependents", + "social_security_survivors", + ) + if all(column in person for column in social_security_columns): + predictors["social_security"] = np.sum( + np.column_stack([_numeric(column) for column in social_security_columns]), + axis=1, + ) + else: + predictors["social_security"] = _numeric("SS_VAL") + return predictors + + +def with_us_workers_compensation( + frame: Frame, + *, + seed: int, + time_period: int, + allow_existing_without_source: bool = False, +) -> Frame: + """Materialize direct ASEC and post-PUF-clone workers' compensation.""" + + if frame.schema != US_SCHEMA: + raise ValueError("US workers' compensation require the US schema.") + person = frame.table("person") + source_available = all( + column in person for column in US_WORKERS_COMPENSATION_REQUIRED_SOURCE_COLUMNS + ) + if not source_available: + if allow_existing_without_source and _surface_carries_signal(frame): + return frame + missing = [ + column + for column in US_WORKERS_COMPENSATION_REQUIRED_SOURCE_COLUMNS + if column not in person + ] + raise ValueError( + "US workers-compensation stage cannot heal a default surface without " + f"measured ASEC source column(s): {missing}." + ) + + stage_person = person.copy(deep=True) + stage_person[_PERSON_WEIGHT_COLUMN] = frame.resolve_weights("person").values + if _PERSON_SUPPORT_CHANNEL_COLUMN in person: + predictors = _person_workers_compensation_predictors(frame) + for column in _PREDICTORS: + stage_person[_PREDICTOR_PREFIX + column] = predictors[column].to_numpy() + output = run_source_stage( + us_workers_compensation_stage_spec(), + tables={"person": stage_person}, + operation_handlers={ + "derive_workers_compensation": ( + derive_us_workers_compensation_from_manifest + ), + "impute_workers_compensation_to_puf_support": ( + impute_us_workers_compensation_to_puf_support_from_manifest + ), + }, + config=SourceRuntimeConfig(seed=int(seed), target_year=int(time_period)), + ) + aligned = output.set_index("person_id").reindex(person["person_id"]) + if aligned[_OUTPUT].isna().any(): + raise ValueError( + f"US workers-compensation stage output {_OUTPUT!r} does not cover " + "every person." + ) + + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + tables["person"][_OUTPUT] = aligned[_OUTPUT].to_numpy(dtype=np.float64) + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + ) + + +def us_workers_compensation_summary(frame: Frame) -> dict[str, object]: + """Return weighted signal and validity diagnostics.""" + + person = frame.table("person") + weights = np.asarray(frame.resolve_weights("person").values, dtype=np.float64) + values = pd.to_numeric(person[_OUTPUT], errors="coerce").to_numpy(dtype=np.float64) + finite = np.isfinite(values) + positive = finite & (values > 0.0) + total_weight = float(weights.sum()) + summary: dict[str, object] = { + "positive_share": ( + float(weights[positive].sum()) / total_weight if total_weight > 0.0 else 0.0 + ), + "positive_share_band": list(_NONZERO_SHARE_BAND), + "weighted_total": float((np.nan_to_num(values) * weights).sum()), + "nonfinite": int(np.count_nonzero(~finite)), + "negative": int(np.count_nonzero(finite & (values < 0.0))), + } + if _PERSON_SUPPORT_CHANNEL_COLUMN in person: + channel = person[_PERSON_SUPPORT_CHANNEL_COLUMN].astype(str).to_numpy() + channels: dict[str, dict[str, float | int]] = {} + for name in (_BASE_ASEC_SUPPORT_CHANNEL, _PUF_TAX_DETAIL_SUPPORT_CHANNEL): + mask = channel == name + channel_weight = float(weights[mask].sum()) + channels[name] = { + "positive_rows": int(np.count_nonzero(mask & positive)), + "positive_share": ( + float(weights[mask & positive].sum()) / channel_weight + if channel_weight > 0.0 + else 0.0 + ), + "weighted_total": float( + (np.nan_to_num(values[mask]) * weights[mask]).sum() + ), + } + summary["channels"] = channels + if "WC_VAL" in person: + source = pd.to_numeric(person["WC_VAL"], errors="coerce").to_numpy( + dtype=np.float64 + ) + source_mask = np.ones(len(person), dtype=bool) + if _PERSON_SUPPORT_CHANNEL_COLUMN in person: + source_mask = ( + person[_PERSON_SUPPORT_CHANNEL_COLUMN].astype(str).to_numpy() + == _BASE_ASEC_SUPPORT_CHANNEL + ) + source_valid = np.isfinite(source) & (source >= 0.0) + summary["source_invalid"] = int(np.count_nonzero(source_mask & ~source_valid)) + summary["source_mismatch_count"] = int( + np.count_nonzero( + source_mask + & source_valid + & finite + & ~np.isclose(values, source, rtol=0.0, atol=0.0) + ) + ) + return summary + + +def us_workers_compensation_signal_gate(frame: Frame) -> GateResult: + """Require finite, nonnegative, nondefault signal on both support halves.""" + + person = frame.table("person") + if _OUTPUT not in person: + return GateResult( + name="workers_compensation_signal", + passed=False, + failures=(f"person column missing: {_OUTPUT}.",), + details={"missing": [_OUTPUT]}, + ) + + summary = us_workers_compensation_summary(frame) + failures: list[str] = [] + if summary["nonfinite"]: + failures.append(f"{_OUTPUT}: {int(summary['nonfinite'])} nonfinite values.") + if summary["negative"]: + failures.append(f"{_OUTPUT}: {int(summary['negative'])} negative values.") + if int(summary.get("source_invalid", 0)): + failures.append( + f"WC_VAL: {int(summary['source_invalid'])} invalid ASEC source values." + ) + if int(summary.get("source_mismatch_count", 0)): + failures.append( + f"{_OUTPUT}: {int(summary['source_mismatch_count'])} ASEC WC_VAL " + "reconciliation mismatch(es)." + ) + share = float(summary["positive_share"]) + low, high = summary["positive_share_band"] + if not (low <= share <= high): + failures.append( + f"{_OUTPUT}: nonzero share {share:.4f} outside plausibility band " + f"[{low}, {high}]." + ) + channels = summary.get("channels") + if isinstance(channels, dict): + for name in (_BASE_ASEC_SUPPORT_CHANNEL, _PUF_TAX_DETAIL_SUPPORT_CHANNEL): + detail = channels.get(name) + if not isinstance(detail, dict): + failures.append(f"{_OUTPUT}: missing {name} channel diagnostics.") + continue + channel_low, channel_high = _CHANNEL_NONZERO_SHARE_BANDS[name] + channel_share = float(detail["positive_share"]) + if not (channel_low <= channel_share <= channel_high): + failures.append( + f"{_OUTPUT}: {name} nonzero share {channel_share:.4f} " + "outside plausibility band " + f"[{channel_low}, {channel_high}]." + ) + return GateResult( + name="workers_compensation_signal", + passed=not failures, + failures=tuple(failures), + details=summary, + ) + + +def _surface_carries_signal(frame: Frame) -> bool: + return ( + _OUTPUT in frame.table("person") + and us_workers_compensation_signal_gate(frame).passed + ) diff --git a/packages/populace-build/tests/test_release_input_coverage.py b/packages/populace-build/tests/test_release_input_coverage.py index 5c9d30e6..1b95f020 100644 --- a/packages/populace-build/tests/test_release_input_coverage.py +++ b/packages/populace-build/tests/test_release_input_coverage.py @@ -32,6 +32,7 @@ import populace.build.us_runtime.reform_coverage_smoke as smoke_module from populace.build.us_runtime import ( SSI_COUNTABLE_RESOURCE_ASSETS, + US_QBI_OUTPUT_COLUMNS, US_RELEASE_INPUT_COVERAGE_RESOURCE, ReformCoverageProbe, ReleaseInputColumn, @@ -68,6 +69,36 @@ def _person_frame(columns: dict[str, np.ndarray]) -> Frame: ) +def _household_weight_frame( + typed_values: np.ndarray, + *, + stored_values: np.ndarray | None = None, +) -> Frame: + """A Frame whose authoritative household weights may shadow a stale column.""" + + typed_values = np.asarray(typed_values, dtype=np.float64) + n = len(typed_values) + person = pd.DataFrame( + { + "person_id": np.arange(n, dtype="int64"), + "person_household_id": np.arange(1, n + 1, dtype="int64"), + } + ) + household = pd.DataFrame({"household_id": np.arange(1, n + 1, dtype="int64")}) + if stored_values is not None: + household["household_weight"] = np.asarray(stored_values, dtype=np.float64) + return Frame( + {"person": person, "household": household}, + EntitySchema(group_entities=("household",)), + { + "household": Weights( + values=typed_values, + kind=WeightKind.CALIBRATED, + ) + }, + ) + + class _StubEngine: """Only ``default_values(names)`` — the single engine surface the gate uses.""" @@ -195,6 +226,36 @@ def test_absent_and_degenerate_excluded_column_passes(self) -> None: "alimony_income": "Residual income-source layer not yet sourced; tracked." } + def test_typed_household_weights_count_as_persisted_input_signal(self) -> None: + manifest = _manifest((ReleaseInputColumn("household_weight", "required"),)) + frame = _household_weight_frame(np.asarray([125.0, 275.0])) + assert "household_weight" not in frame.table("household") + + result = us_release_input_coverage_gate( + frame, + _StubEngine({"household_weight": 1.0}), + manifest=manifest, + ) + + assert result.passed + assert result.failures == () + + def test_typed_household_weights_override_stale_table_column(self) -> None: + manifest = _manifest((ReleaseInputColumn("household_weight", "required"),)) + frame = _household_weight_frame( + np.asarray([1.0, 1.0]), + stored_values=np.asarray([125.0, 275.0]), + ) + + result = us_release_input_coverage_gate( + frame, + _StubEngine({"household_weight": 1.0}), + manifest=manifest, + ) + + assert not result.passed + assert result.details["degenerate_required"] == ["household_weight"] + class _Series: def __init__(self, total: float) -> None: @@ -231,6 +292,88 @@ def _probe(min_abs_effect: float = 1_000_000_000.0) -> ReformCoverageProbe: ) +def _tips_probe() -> ReformCoverageProbe: + return ReformCoverageProbe( + id="tips_probe", + name="OBBBA no-tax-on-tips deduction", + parameter_changes={ + "gov.irs.deductions.tip_income.cap": {"2026-01-01.2026-12-31": 0} + }, + budget_measure="income_tax", + binding_inputs=("tip_income", "treasury_tipped_occupation_code"), + min_abs_effect=100_000_000.0, + reason="The cap repeal must bind through qualified tip income.", + issue="PolicyEngine/populace#38", + effect_direction="baseline_minus_reform", + period=2026, + expected_sign="negative", + ) + + +def _overtime_probe() -> ReformCoverageProbe: + return ReformCoverageProbe( + id="obbba_no_tax_on_overtime", + name="OBBBA no-tax-on-overtime deduction", + parameter_changes={ + "gov.irs.deductions.overtime_income.cap.SINGLE": { + "2026-01-01.2026-12-31": 0 + } + }, + budget_measure="income_tax", + binding_inputs=("fsla_overtime_premium",), + min_abs_effect=100_000_000.0, + reason="The cap repeal must bind through the FLSA overtime premium.", + issue="PolicyEngine/populace#242", + effect_direction="baseline_minus_reform", + period=2026, + expected_sign="negative", + ) + + +def _auto_loan_probe() -> ReformCoverageProbe: + return ReformCoverageProbe( + id="obbba_auto_loan_interest", + name="OBBBA no-tax-on-auto-loan-interest deduction", + parameter_changes={ + "gov.irs.deductions.auto_loan_interest.cap": {"2026-01-01.2026-12-31": 0} + }, + budget_measure="income_tax", + binding_inputs=("qualified_passenger_vehicle_loan_interest",), + min_abs_effect=100_000_000.0, + reason="The repeal must bind through qualifying vehicle-loan interest.", + issue="PolicyEngine/populace#252", + effect_direction="baseline_minus_reform", + period=2026, + expected_sign="negative", + ) + + +def test_reform_probe_requires_exactly_one_reform_kind() -> None: + kwargs = { + "id": "form_4952", + "name": "Form 4952 neutralization", + "budget_measure": "income_tax", + "binding_inputs": ("investment_income_elected_form_4952",), + "min_abs_effect": 1_000_000.0, + "reason": "The neutralization binds only through the input.", + "issue": "PolicyEngine/populace#274", + } + with pytest.raises(ValueError, match="exactly one"): + ReformCoverageProbe(parameter_changes={}, **kwargs) + with pytest.raises(ValueError, match="exactly one"): + ReformCoverageProbe( + parameter_changes={"some.parameter": {"2024-01-01": 0}}, + neutralized_variable="investment_income_elected_form_4952", + **kwargs, + ) + with pytest.raises(ValueError, match="binding_inputs"): + ReformCoverageProbe( + parameter_changes={}, + neutralized_variable="different_input", + **kwargs, + ) + + class TestReformCoverageSmokeGate: def test_zero_bound_reform_fails(self, monkeypatch) -> None: # Case 5: with the asset inputs absent, everyone already passes the SSI @@ -263,6 +406,89 @@ def simulate(reform): assert result.passed assert result.details["results"]["ssi_probe"]["effect"] == pytest.approx(1.6e9) + def test_wrong_signed_effect_fails(self, monkeypatch) -> None: + monkeypatch.setattr(smoke_module, "_build_reform", lambda changes: "REFORM") + + def simulate(reform): + return _Sim(3.0e10 if reform == "REFORM" else 4.0e10) + + result = us_reform_coverage_smoke_gate( + simulate=simulate, probes=[_probe()], period=2024 + ) + assert not result.passed + assert result.details["results"]["ssi_probe"]["effect"] == -1.0e10 + assert "expected a positive effect" in result.failures[0] + + def test_negative_tip_effect_uses_probe_period_and_passes( + self, monkeypatch + ) -> None: + monkeypatch.setattr(smoke_module, "_build_reform", lambda changes: "REFORM") + periods: list[int] = [] + + class RecordingSim(_Sim): + def calculate(self, measure: str, period): + periods.append(period) + return super().calculate(measure, period) + + def simulate(reform): + return RecordingSim(10.5e9 if reform == "REFORM" else 10.0e9) + + result = us_reform_coverage_smoke_gate( + simulate=simulate, + probes=[_tips_probe()], + period=2024, + ) + + assert result.passed + assert periods == [2026, 2026] + assert result.details["default_period"] == 2024 + tip_result = result.details["results"]["tips_probe"] + assert tip_result["period"] == 2026 + assert tip_result["effect"] == pytest.approx(-0.5e9) + assert tip_result["expected_sign"] == "negative" + + def test_negative_overtime_effect_passes_and_wrong_sign_fails( + self, monkeypatch + ) -> None: + monkeypatch.setattr(smoke_module, "_build_reform", lambda changes: "REFORM") + + passing = us_reform_coverage_smoke_gate( + simulate=lambda reform: _Sim(10.5e9 if reform else 10.0e9), + probes=[_overtime_probe()], + ) + assert passing.passed + assert passing.details["results"]["obbba_no_tax_on_overtime"][ + "effect" + ] == pytest.approx(-0.5e9) + + wrong_sign = us_reform_coverage_smoke_gate( + simulate=lambda reform: _Sim(9.5e9 if reform else 10.0e9), + probes=[_overtime_probe()], + ) + assert not wrong_sign.passed + assert "expected a negative effect" in wrong_sign.failures[0] + + def test_negative_auto_loan_effect_passes_and_wrong_sign_fails( + self, monkeypatch + ) -> None: + monkeypatch.setattr(smoke_module, "_build_reform", lambda changes: "REFORM") + + passing = us_reform_coverage_smoke_gate( + simulate=lambda reform: _Sim(10.5e9 if reform else 10.0e9), + probes=[_auto_loan_probe()], + ) + assert passing.passed + result = passing.details["results"]["obbba_auto_loan_interest"] + assert result["effect"] == pytest.approx(-0.5e9) + assert result["period"] == 2026 + + wrong_sign = us_reform_coverage_smoke_gate( + simulate=lambda reform: _Sim(9.5e9 if reform else 10.0e9), + probes=[_auto_loan_probe()], + ) + assert not wrong_sign.passed + assert "expected a negative effect" in wrong_sign.failures[0] + def test_probeless_gate_is_refused(self) -> None: # A probe-less smoke gate would pass vacuously — refuse it. with pytest.raises(ValueError, match="at least one probe"): @@ -282,6 +508,279 @@ def test_ssi_assets_are_required_without_exclusion(self) -> None: assert asset in manifest.required_columns assert asset not in manifest.reviewed_exclusions + def test_post_reference_ssi_disability_criterion_has_unique_probe(self) -> None: + manifest = load_release_input_coverage_manifest() + column = "meets_ssi_disability_criteria" + assert column in manifest.required_columns + assert column not in manifest.reviewed_exclusions + + probe = next( + probe + for probe in manifest.probes + if probe.id == "ssi_disability_criteria_neutralization" + ) + assert probe.parameter_changes == {} + assert probe.neutralized_variable == column + assert probe.budget_measure == "ssi" + assert probe.period == 2024 + assert probe.effect_direction == "baseline_minus_reform" + assert probe.expected_sign == "positive" + assert probe.binding_inputs == (column,) + assert probe.min_abs_effect == 100_000_000.0 + + def test_restored_ssi_take_up_has_unique_probe(self) -> None: + manifest = load_release_input_coverage_manifest() + column = "takes_up_ssi_if_eligible" + assert column in manifest.required_columns + assert column not in manifest.reviewed_exclusions + + matches = [ + probe + for probe in manifest.probes + if probe.id == "ssi_take_up_neutralization" + ] + assert len(matches) == 1 + probe = matches[0] + assert probe.parameter_changes == {} + assert probe.neutralized_variable == column + assert probe.budget_measure == "ssi" + assert probe.period == 2024 + assert probe.effect_direction == "baseline_minus_reform" + assert probe.expected_sign == "positive" + assert probe.binding_inputs == (column,) + assert probe.min_abs_effect == 10_000_000_000.0 + + def test_restored_head_start_take_up_has_unique_probe(self) -> None: + manifest = load_release_input_coverage_manifest() + column = "takes_up_head_start_if_eligible" + assert column in manifest.required_columns + assert column not in manifest.reviewed_exclusions + + matches = [ + probe + for probe in manifest.probes + if probe.id == "head_start_take_up_neutralization" + ] + assert len(matches) == 1 + probe = matches[0] + assert probe.parameter_changes == {} + assert probe.neutralized_variable == column + assert probe.budget_measure == "head_start" + assert probe.period == 2024 + assert probe.effect_direction == "baseline_minus_reform" + assert probe.expected_sign == "positive" + assert probe.binding_inputs == (column,) + assert probe.min_abs_effect > 0.0 + + def test_post_reference_obbba_inputs_are_hard_requirements(self) -> None: + manifest = load_release_input_coverage_manifest() + for column in ( + "fsla_overtime_premium", + "qualified_passenger_vehicle_loan_interest", + ): + assert column in manifest.required_columns + assert column not in manifest.reviewed_exclusions + + def test_legacy_auto_loan_columns_are_promoted(self) -> None: + manifest = load_release_input_coverage_manifest() + for column in ("auto_loan_balance", "auto_loan_interest"): + assert column in manifest.required_columns + assert column not in manifest.reviewed_exclusions + + def test_sipp_vehicle_family_is_promoted(self) -> None: + manifest = load_release_input_coverage_manifest() + for column in ("household_vehicles_owned", "household_vehicles_value"): + assert column in manifest.required_columns + assert column not in manifest.reviewed_exclusions + + def test_signed_scf_net_worth_is_promoted(self) -> None: + manifest = load_release_input_coverage_manifest() + assert "net_worth" in manifest.required_columns + assert "net_worth" not in manifest.reviewed_exclusions + + def test_typed_household_weight_is_promoted(self) -> None: + manifest = load_release_input_coverage_manifest() + assert "household_weight" in manifest.required_columns + assert "household_weight" not in manifest.reviewed_exclusions + + def test_education_input_family_is_promoted(self) -> None: + manifest = load_release_input_coverage_manifest() + for column in ( + "qualified_tuition_expenses", + "educational_assistance", + "is_pursuing_credential_for_american_opportunity_credit", + "attends_eligible_educational_institution_for_american_opportunity_credit", + "is_enrolled_at_least_half_time_for_american_opportunity_credit", + "has_american_opportunity_credit_1098_t_or_exception", + "has_american_opportunity_credit_institution_ein", + ): + assert column in manifest.required_columns + assert column not in manifest.reviewed_exclusions + + def test_retirement_contribution_family_is_a_hard_requirement(self) -> None: + manifest = load_release_input_coverage_manifest() + for column in ( + "traditional_401k_contributions_desired", + "roth_401k_contributions_desired", + "traditional_ira_contributions_desired", + "roth_ira_contributions_desired", + "self_employed_pension_contributions_desired", + ): + assert column in manifest.required_columns + assert column not in manifest.reviewed_exclusions + + def test_casualty_loss_is_promoted(self) -> None: + manifest = load_release_input_coverage_manifest() + assert "casualty_loss" in manifest.required_columns + assert "casualty_loss" not in manifest.reviewed_exclusions + + def test_alimony_family_is_promoted(self) -> None: + manifest = load_release_input_coverage_manifest() + for column in ("alimony_income", "alimony_expense"): + assert column in manifest.required_columns + assert column not in manifest.reviewed_exclusions + + def test_misc_itemized_input_is_promoted(self) -> None: + manifest = load_release_input_coverage_manifest() + column = "unreimbursed_business_employee_expenses" + assert column in manifest.required_columns + assert column not in manifest.reviewed_exclusions + + def test_childcare_input_is_a_hard_requirement(self) -> None: + manifest = load_release_input_coverage_manifest() + column = "spm_unit_pre_subsidy_childcare_expenses" + assert column in manifest.required_columns + assert column not in manifest.reviewed_exclusions + + def test_energy_subsidy_is_promoted_with_unique_probe(self) -> None: + manifest = load_release_input_coverage_manifest() + column = "spm_unit_energy_subsidy" + assert column in manifest.required_columns + assert column not in manifest.reviewed_exclusions + + probe = next( + probe + for probe in manifest.probes + if probe.id == "spm_unit_energy_subsidy_neutralization" + ) + assert probe.parameter_changes == {} + assert probe.neutralized_variable == column + assert probe.budget_measure == "spm_unit_benefits" + assert probe.period == 2024 + assert probe.effect_direction == "baseline_minus_reform" + assert probe.expected_sign == "positive" + assert probe.binding_inputs == (column,) + assert probe.min_abs_effect == 100_000_000.0 + + def test_voluntary_filing_is_promoted_with_aca_ptc_probe(self) -> None: + manifest = load_release_input_coverage_manifest() + column = "would_file_taxes_voluntarily" + assert column in manifest.required_columns + assert column not in manifest.reviewed_exclusions + + probe = next( + probe + for probe in manifest.probes + if probe.id == "voluntary_filing_aca_ptc_neutralization" + ) + assert probe.parameter_changes == {} + assert probe.neutralized_variable == column + assert probe.budget_measure == "aca_ptc" + assert probe.period == 2024 + assert probe.effect_direction == "baseline_minus_reform" + assert probe.expected_sign == "positive" + assert probe.binding_inputs == (column,) + assert probe.min_abs_effect == 100_000_000.0 + + def test_child_support_family_is_promoted(self) -> None: + manifest = load_release_input_coverage_manifest() + for column in ("child_support_received", "child_support_expense"): + assert column in manifest.required_columns + assert column not in manifest.reviewed_exclusions + + def test_disability_benefits_is_promoted(self) -> None: + manifest = load_release_input_coverage_manifest() + assert "disability_benefits" in manifest.required_columns + assert "disability_benefits" not in manifest.reviewed_exclusions + + def test_educator_expense_is_promoted(self) -> None: + manifest = load_release_input_coverage_manifest() + assert "educator_expense" in manifest.required_columns + assert "educator_expense" not in manifest.reviewed_exclusions + + def test_other_health_insurance_premiums_is_promoted(self) -> None: + manifest = load_release_input_coverage_manifest() + assert "other_health_insurance_premiums" in manifest.required_columns + assert "other_health_insurance_premiums" not in manifest.reviewed_exclusions + + def test_prior_year_income_family_is_promoted(self) -> None: + manifest = load_release_input_coverage_manifest() + for column in ( + "self_employment_income_last_year", + "previous_year_income_available", + ): + assert column in manifest.required_columns + assert column not in manifest.reviewed_exclusions + + probe = next( + probe + for probe in manifest.probes + if probe.id == "prior_year_self_employment_neutralization" + ) + assert probe.period == 2024 + assert probe.neutralized_variable == "self_employment_income_last_year" + assert probe.parameter_changes == {} + assert probe.binding_inputs == ("self_employment_income_last_year",) + assert probe.budget_measure == "tax_unit_earned_income_last_year" + assert probe.effect_direction == "baseline_minus_reform" + + def test_weeks_unemployed_is_required_without_a_false_policy_probe(self) -> None: + manifest = load_release_input_coverage_manifest() + column = "weeks_unemployed" + + assert column in manifest.required_columns + assert column not in manifest.reviewed_exclusions + # PolicyEngine-US 1.764.6 consumes this leaf only through Pennsylvania + # UC, whose other monetary-eligibility inputs remain default-zero. A + # direct neutralization would test the column against itself, not a + # budget-policy path, so this family deliberately adds no such probe. + assert all( + column not in probe.binding_inputs and probe.neutralized_variable != column + for probe in manifest.probes + ) + + def test_signed_farm_business_income_family_is_promoted(self) -> None: + manifest = load_release_input_coverage_manifest() + for column in ("farm_operations_income", "farm_rent_income"): + assert column in manifest.required_columns + assert column not in manifest.reviewed_exclusions + + def test_qbi_input_family_is_promoted(self) -> None: + manifest = load_release_input_coverage_manifest() + for column in US_QBI_OUTPUT_COLUMNS: + assert column in manifest.required_columns + assert column not in manifest.reviewed_exclusions + + def test_domestic_production_ald_is_promoted_separately_from_qbi(self) -> None: + manifest = load_release_input_coverage_manifest() + assert "domestic_production_ald" in manifest.required_columns + assert "domestic_production_ald" not in manifest.reviewed_exclusions + + def test_form_4952_election_is_promoted_with_unique_probe(self) -> None: + manifest = load_release_input_coverage_manifest() + output = "investment_income_elected_form_4952" + assert output in manifest.required_columns + assert output not in manifest.reviewed_exclusions + + probe = next( + probe + for probe in manifest.probes + if probe.id == "form_4952_election_neutralization" + ) + assert probe.parameter_changes == {} + assert probe.neutralized_variable == output + assert probe.binding_inputs == (output,) + def test_shipped_ssi_probe_binds_through_the_assets(self) -> None: probes = us_release_reform_coverage_probes() assert probes, "the shipped manifest must pin at least one reform probe" @@ -289,6 +788,403 @@ def test_shipped_ssi_probe_binds_through_the_assets(self) -> None: assert set(SSI_COUNTABLE_RESOURCE_ASSETS) <= set(ssi.binding_inputs) assert ssi.budget_measure == "ssi" assert ssi.min_abs_effect > 0 + assert ssi.expected_sign == "positive" + + def test_shipped_tip_probe_has_2026_period_sign_and_inputs(self) -> None: + tip = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "obbba_no_tax_on_tips" + ) + assert tip.period == 2026 + assert tip.expected_sign == "negative" + assert tip.effect_direction == "baseline_minus_reform" + assert tip.budget_measure == "income_tax" + assert set(tip.binding_inputs) == { + "tip_income", + "treasury_tipped_occupation_code", + } + + def test_shipped_aotc_probe_binds_through_education_inputs(self) -> None: + aotc = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "aotc_abolition" + ) + assert aotc.period == 2024 + assert aotc.expected_sign == "positive" + assert aotc.effect_direction == "baseline_minus_reform" + assert aotc.budget_measure == "american_opportunity_credit" + assert set(aotc.binding_inputs) == { + "qualified_tuition_expenses", + "is_pursuing_credential_for_american_opportunity_credit", + "attends_eligible_educational_institution_for_american_opportunity_credit", + "is_enrolled_at_least_half_time_for_american_opportunity_credit", + "has_american_opportunity_credit_1098_t_or_exception", + "has_american_opportunity_credit_institution_ein", + } + assert aotc.min_abs_effect > 0 + assert set(aotc.parameter_changes) == { + "gov.irs.credits.education.american_opportunity_credit.abolition" + } + + def test_shipped_savers_credit_probe_binds_through_contributions(self) -> None: + probe = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "savers_credit_abolition" + ) + assert probe.period == 2024 + assert probe.expected_sign == "positive" + assert probe.effect_direction == "baseline_minus_reform" + assert probe.budget_measure == "savers_credit" + assert set(probe.binding_inputs) == { + "traditional_401k_contributions_desired", + "roth_401k_contributions_desired", + "traditional_ira_contributions_desired", + "roth_ira_contributions_desired", + "self_employed_pension_contributions_desired", + } + assert probe.min_abs_effect == 100_000_000.0 + assert set(probe.parameter_changes) == { + "gov.irs.credits.retirement_saving.contributions_cap" + } + + def test_shipped_casualty_probe_has_2026_period_sign_and_input(self) -> None: + probe = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "obbba_casualty_loss_limit" + ) + assert probe.period == 2026 + assert probe.expected_sign == "positive" + assert probe.effect_direction == "baseline_minus_reform" + assert probe.budget_measure == "income_tax" + assert probe.binding_inputs == ("casualty_loss",) + assert probe.min_abs_effect == 1_000_000.0 + assert set(probe.parameter_changes) == { + "gov.irs.deductions.itemized.casualty.active" + } + + def test_shipped_alimony_probe_has_sign_period_and_expense_input(self) -> None: + probe = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "alimony_expense_ald_abolition" + ) + assert probe.period == 2024 + assert probe.expected_sign == "negative" + assert probe.effect_direction == "baseline_minus_reform" + assert probe.budget_measure == "income_tax" + assert probe.binding_inputs == ("alimony_expense",) + assert probe.min_abs_effect == 1_000_000.0 + assert set(probe.parameter_changes) == { + "gov.irs.ald.alimony_expense.divorce_year_threshold[0].amount" + } + + def test_shipped_misc_itemized_probe_has_2026_period_sign_and_input(self) -> None: + probe = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "obbba_misc_itemized_deductions" + ) + assert probe.period == 2026 + assert probe.expected_sign == "positive" + assert probe.effect_direction == "baseline_minus_reform" + assert probe.budget_measure == "income_tax" + assert probe.binding_inputs == ("unreimbursed_business_employee_expenses",) + assert probe.min_abs_effect == 100_000_000.0 + assert set(probe.parameter_changes) == { + "gov.irs.deductions.itemized.misc.applies" + } + + def test_shipped_cdcc_probe_has_2026_period_sign_and_input(self) -> None: + probe = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "obbba_cdcc" + ) + assert probe.period == 2026 + assert probe.expected_sign == "negative" + assert probe.effect_direction == "baseline_minus_reform" + assert probe.budget_measure == "income_tax" + assert probe.binding_inputs == ("spm_unit_pre_subsidy_childcare_expenses",) + assert probe.min_abs_effect == 1_000_000.0 + assert set(probe.parameter_changes) == { + "gov.irs.credits.cdcc.phase_out.max", + "gov.irs.credits.cdcc.phase_out.min", + "gov.irs.credits.cdcc.phase_out.amended_structure.applies", + } + + def test_shipped_child_support_received_probe_removes_only_snap_source( + self, + ) -> None: + probe = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "child_support_received_snap_exclusion" + ) + assert probe.period == 2024 + assert probe.expected_sign == "positive" + assert probe.effect_direction == "reform_minus_baseline" + assert probe.budget_measure == "snap" + assert probe.binding_inputs == ("child_support_received",) + assert probe.min_abs_effect == 1_000_000.0 + assert probe.parameter_changes == { + "gov.usda.snap.income.sources.unearned": { + "2024-01-01.2024-12-31": [ + "ssi", + "tanf", + "general_assistance", + "pension_income", + "veterans_benefits", + "unemployment_compensation", + "disability_benefits", + "workers_compensation", + "social_security", + "retirement_distributions", + "rental_income", + "alimony_income", + "financial_assistance", + "survivor_benefits", + "dividend_income", + "interest_income", + "miscellaneous_income", + ] + } + } + + def test_shipped_child_support_expense_probe_removes_only_snap_deduction( + self, + ) -> None: + probe = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "child_support_expense_snap_deduction_abolition" + ) + assert probe.period == 2024 + assert probe.expected_sign == "positive" + assert probe.effect_direction == "baseline_minus_reform" + assert probe.budget_measure == "snap" + assert probe.binding_inputs == ("child_support_expense",) + assert probe.min_abs_effect == 1_000_000.0 + assert probe.parameter_changes == { + "gov.usda.snap.income.deductions.allowed": { + "2024-01-01.2024-12-31": [ + "snap_standard_deduction", + "snap_earned_income_deduction", + "snap_dependent_care_deduction", + "snap_excess_medical_expense_deduction", + "snap_excess_shelter_expense_deduction", + ] + } + } + + def test_shipped_disability_probe_removes_only_snap_unearned_source( + self, + ) -> None: + probe = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "disability_benefits_snap_exclusion" + ) + assert probe.period == 2024 + assert probe.expected_sign == "positive" + assert probe.effect_direction == "reform_minus_baseline" + assert probe.budget_measure == "snap" + assert probe.binding_inputs == ("disability_benefits",) + assert probe.min_abs_effect == 1_000_000.0 + assert probe.parameter_changes == { + "gov.usda.snap.income.sources.unearned": { + "2024-01-01.2024-12-31": [ + "ssi", + "tanf", + "general_assistance", + "pension_income", + "veterans_benefits", + "unemployment_compensation", + "workers_compensation", + "social_security", + "retirement_distributions", + "rental_income", + "child_support_received", + "alimony_income", + "financial_assistance", + "survivor_benefits", + "dividend_income", + "interest_income", + "miscellaneous_income", + ] + } + } + + def test_shipped_educator_expense_probe_removes_only_its_ald(self) -> None: + probe = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "educator_expense_ald_abolition" + ) + + assert probe.period == 2024 + assert probe.expected_sign == "positive" + assert probe.effect_direction == "reform_minus_baseline" + assert probe.budget_measure == "income_tax" + assert probe.binding_inputs == ("educator_expense",) + assert probe.min_abs_effect == 1_000_000.0 + assert probe.parameter_changes == { + "gov.irs.ald.deductions": { + "2024-01-01.2024-12-31": [ + "loss_ald", + "self_employment_tax_ald", + "student_loan_interest_ald", + "early_withdrawal_penalty", + "alimony_expense_ald", + "health_savings_account_ald", + "self_employed_health_insurance_ald", + "self_employed_pension_contribution_ald", + "traditional_ira_contributions", + "qualified_adoption_assistance_expense", + "us_bonds_for_higher_ed", + "specified_possession_income", + "puerto_rico_income", + ] + } + } + + def test_shipped_qbi_probes_cover_reit_and_wage_property_inputs(self) -> None: + probes = {probe.id: probe for probe in us_release_reform_coverage_probes()} + reit = probes["qbi_reit_ptp_rate_abolition"] + assert reit.period == 2024 + assert reit.expected_sign == "positive" + assert reit.effect_direction == "baseline_minus_reform" + assert reit.budget_measure == "qualified_business_income_deduction" + assert reit.binding_inputs == ("qualified_reit_and_ptp_income",) + assert set(reit.parameter_changes) == { + "gov.irs.deductions.qbi.max.reit_ptp_rate" + } + + guardrails = probes["qbi_wage_property_guardrails_zeroed"] + assert guardrails.period == 2024 + assert guardrails.expected_sign == "positive" + assert guardrails.effect_direction == "baseline_minus_reform" + assert guardrails.budget_measure == "qualified_business_income_deduction" + assert set(guardrails.binding_inputs) == { + "w2_wages_from_qualified_business", + "unadjusted_basis_qualified_property", + } + assert set(guardrails.parameter_changes) == { + "gov.irs.deductions.qbi.max.w2_wages.rate", + "gov.irs.deductions.qbi.max.w2_wages.alt_rate", + "gov.irs.deductions.qbi.max.business_property.rate", + } + + def test_shipped_farm_probes_each_remove_only_the_bound_qbi_leaf(self) -> None: + probes = {probe.id: probe for probe in us_release_reform_coverage_probes()} + current_income_definition = { + "self_employment_income", + "partnership_s_corp_income", + "farm_rent_income", + "farm_operations_income", + "rental_income", + "estate_income", + } + cases = { + "qbi_farm_operations_income_exclusion": ( + "farm_operations_income", + "negative", + ), + "qbi_farm_rent_income_exclusion": ("farm_rent_income", "positive"), + } + + for probe_id, (removed_input, expected_sign) in cases.items(): + probe = probes[probe_id] + assert probe.period == 2026 + assert probe.expected_sign == expected_sign + assert probe.effect_direction == "baseline_minus_reform" + assert probe.budget_measure == "qualified_business_income_deduction" + assert probe.binding_inputs == (removed_input,) + assert probe.min_abs_effect == 1_000_000.0 + assert set(probe.parameter_changes) == { + "gov.irs.deductions.qbi.income_definition" + } + definition = probe.parameter_changes[ + "gov.irs.deductions.qbi.income_definition" + ] + assert set(definition) == {"2026-01-01.2026-12-31"} + assert set(definition["2026-01-01.2026-12-31"]) == ( + current_income_definition - {removed_input} + ) + + def test_shipped_domestic_production_probe_reactivates_only_its_ald(self) -> None: + probe = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "domestic_production_ald_reactivation" + ) + assert probe.period == 2024 + assert probe.expected_sign == "positive" + assert probe.effect_direction == "baseline_minus_reform" + assert probe.budget_measure == "income_tax" + assert probe.binding_inputs == ("domestic_production_ald",) + assert probe.min_abs_effect == 1_000_000.0 + assert set(probe.parameter_changes) == {"gov.irs.ald.deductions"} + deductions = probe.parameter_changes["gov.irs.ald.deductions"] + assert set(deductions) == {"2024-01-01.2024-12-31"} + assert deductions["2024-01-01.2024-12-31"].count("domestic_production_ald") == 1 + + def test_shipped_overtime_probe_has_2026_period_sign_and_input(self) -> None: + overtime = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "obbba_no_tax_on_overtime" + ) + assert overtime.period == 2026 + assert overtime.expected_sign == "negative" + assert overtime.effect_direction == "baseline_minus_reform" + assert overtime.budget_measure == "income_tax" + assert overtime.binding_inputs == ("fsla_overtime_premium",) + assert overtime.min_abs_effect > 0 + assert set(overtime.parameter_changes) == { + "gov.irs.deductions.overtime_income.cap.JOINT", + "gov.irs.deductions.overtime_income.cap.SINGLE", + "gov.irs.deductions.overtime_income.cap.HEAD_OF_HOUSEHOLD", + "gov.irs.deductions.overtime_income.cap.SURVIVING_SPOUSE", + "gov.irs.deductions.overtime_income.cap.SEPARATE", + } + + def test_shipped_auto_loan_probe_has_2026_period_sign_and_input(self) -> None: + auto = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "obbba_auto_loan_interest" + ) + assert auto.period == 2026 + assert auto.expected_sign == "negative" + assert auto.effect_direction == "baseline_minus_reform" + assert auto.budget_measure == "income_tax" + assert auto.binding_inputs == ("qualified_passenger_vehicle_loan_interest",) + assert set(auto.parameter_changes) == { + "gov.irs.deductions.auto_loan_interest.cap" + } + + def test_shipped_vehicle_asset_probe_binds_through_both_inputs(self) -> None: + probe = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "tx_snap_additional_vehicle_exemption_abolition" + ) + assert probe.period == 2026 + assert probe.expected_sign == "positive" + assert probe.effect_direction == "baseline_minus_reform" + assert probe.budget_measure == "snap" + assert set(probe.binding_inputs) == { + "household_vehicles_owned", + "household_vehicles_value", + } + assert probe.min_abs_effect == 1_000_000.0 + assert set(probe.parameter_changes) == { + "gov.hhs.tanf.non_cash.tx_additional_vehicle_exemption" + } def test_demoting_an_ssi_asset_to_exclusion_is_rejected(self) -> None: # The #368 red-gate guarantee cannot be quietly undone: turning an SSI @@ -315,6 +1211,10 @@ def test_demoting_an_ssi_asset_to_exclusion_is_rejected(self) -> None: manifest=tampered, engine=None ) + def test_duplicate_probe_ids_are_rejected(self) -> None: + with pytest.raises(ValueError, match="Duplicate reform coverage probe id"): + _manifest(_CONTRACT.columns, probes=(_probe(), _probe())) + def _load_manifest_generator(): spec = importlib.util.spec_from_file_location( diff --git a/packages/populace-build/tests/test_us_alimony.py b/packages/populace-build/tests/test_us_alimony.py new file mode 100644 index 00000000..37d7db11 --- /dev/null +++ b/packages/populace-build/tests/test_us_alimony.py @@ -0,0 +1,330 @@ +"""Contracts for the retired eCPS alimony input family.""" + +from __future__ import annotations + +import importlib.util + +import numpy as np +import pandas as pd +import pytest + +from populace.build.us_runtime import puf_support as puf_support_module +from populace.build.us_runtime.alimony import ( + ALIMONY_ASEC_ARCHIVED_DERIVATION_URL, + ALIMONY_PUF_ARCHIVED_DERIVATION_URL, + derive_us_alimony_from_asec, + derive_us_alimony_from_puf, + us_alimony_signal_gate, + us_alimony_stage_spec, +) +from populace.build.us_runtime.puf_aggregate_records import ( + _reconcile_puf_alimony_from_sources, +) + +policyengine_us_installed = importlib.util.find_spec("policyengine_us") is not None +requires_us = pytest.mark.skipif( + not policyengine_us_installed, + reason="requires the policyengine-us [us] extra (build environment)", +) + + +class _ResolvedWeights: + def __init__(self, values: np.ndarray) -> None: + self.values = values + + +class _PersonFrame: + def __init__(self, person: pd.DataFrame) -> None: + self._person = person + + def table(self, entity: str) -> pd.DataFrame: + assert entity == "person" + return self._person + + def resolve_weights(self, entity: str) -> _ResolvedWeights: + assert entity == "person" + return _ResolvedWeights(np.ones(len(self._person))) + + +def test_archived_coordinates_are_commit_and_line_pinned() -> None: + for url in ( + ALIMONY_ASEC_ARCHIVED_DERIVATION_URL, + ALIMONY_PUF_ARCHIVED_DERIVATION_URL, + ): + assert "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe" in url + assert "#L" in url + + +def test_asec_mapping_splits_alimony_and_miscellaneous_income_exactly() -> None: + source = pd.DataFrame( + { + "OI_OFF": [20, 12, 19, 0], + "OI_VAL": [5_000.0, 700.0, 900.0, 0.0], + } + ) + + result = derive_us_alimony_from_asec(source) + + assert result["alimony_income"].tolist() == [5_000.0, 0.0, 0.0, 0.0] + assert result["miscellaneous_income"].tolist() == [0.0, 0.0, 900.0, 0.0] + assert "alimony_income" not in source + assert "miscellaneous_income" not in source + + +def test_asec_mapping_preserves_an_already_materialized_pair() -> None: + source = pd.DataFrame( + { + "alimony_income": [10.0], + "miscellaneous_income": [20.0], + } + ) + + result = derive_us_alimony_from_asec(source) + + assert result.equals(source) + assert result is not source + + +def test_asec_mapping_replaces_partial_stale_miscellaneous_carry() -> None: + source = pd.DataFrame( + { + "OI_OFF": [20, 19], + "OI_VAL": [5_000.0, 700.0], + "miscellaneous_income": [5_000.0, 999.0], + } + ) + + result = derive_us_alimony_from_asec(source) + + assert result["alimony_income"].tolist() == [5_000.0, 0.0] + assert result["miscellaneous_income"].tolist() == [0.0, 700.0] + + +@pytest.mark.parametrize( + ("source", "message"), + [ + (pd.DataFrame({"OI_VAL": [1.0]}), "requires both raw source columns"), + ( + pd.DataFrame({"OI_VAL": [1.0], "OI_OFF": [np.inf]}), + "nonnumeric or nonfinite", + ), + ( + pd.DataFrame({"OI_VAL": [-1.0], "OI_OFF": [20]}), + "negative value", + ), + ( + pd.DataFrame({"OI_VAL": [1.0], "OI_OFF": [20.5]}), + "noninteger code", + ), + ], +) +def test_asec_mapping_fails_closed(source: pd.DataFrame, message: str) -> None: + with pytest.raises(ValueError, match=message): + derive_us_alimony_from_asec(source) + + +def test_puf_mapping_is_an_exact_two_field_carry() -> None: + source = pd.DataFrame( + { + "E00800": [0.0, 125.5, 9_000.0], + "E03500": [40.0, 0.0, 700.0], + } + ) + + result = derive_us_alimony_from_puf(source) + + assert result["alimony_income"].tolist() == [0.0, 125.5, 9_000.0] + assert result["alimony_expense"].tolist() == [40.0, 0.0, 700.0] + + +@pytest.mark.parametrize( + ("source", "message"), + [ + (pd.DataFrame({"E00800": [1.0]}), "requires source column 'E03500'"), + ( + pd.DataFrame({"E00800": [np.nan], "E03500": [0.0]}), + "nonnumeric or nonfinite", + ), + ( + pd.DataFrame({"E00800": [0.0], "E03500": [-1.0]}), + "negative value", + ), + ], +) +def test_puf_mapping_fails_closed(source: pd.DataFrame, message: str) -> None: + with pytest.raises(ValueError, match=message): + derive_us_alimony_from_puf(source) + + +def test_post_disaggregation_reconciliation_uses_final_raw_fields() -> None: + source = pd.DataFrame( + { + "E00800": [0.0, 3_000.0], + "E03500": [2_000.0, 0.0], + "alimony_income": [999.0, 999.0], + "alimony_expense": [999.0, 999.0], + } + ) + + result = _reconcile_puf_alimony_from_sources(source) + + assert result["alimony_income"].tolist() == [0.0, 3_000.0] + assert result["alimony_expense"].tolist() == [2_000.0, 0.0] + + +def test_shared_puf_stage_declares_exact_sources_and_outputs() -> None: + stage = us_alimony_stage_spec() + operation = next( + operation + for operation in stage.operations + if operation.kind == "derive_puf_policyengine_variables" + ) + + assert operation.parameters["alimony_income_source"] == "E00800" + assert operation.parameters["alimony_income_output"] == "alimony_income" + assert operation.parameters["alimony_expense_source"] == "E03500" + assert operation.parameters["alimony_expense_output"] == "alimony_expense" + assert set(("alimony_income", "alimony_expense")) <= set(stage.outputs) + assert set(("alimony_income", "alimony_expense")) <= set(stage.nonnegative_outputs) + + +def test_puf_support_marks_both_leaves_sparse_nonnegative_and_preserves_income() -> ( + None +): + outputs = {"alimony_income", "alimony_expense"} + assert outputs <= set(puf_support_module.PUF_TAX_DETAIL_DEFAULT_PERSON_OUTPUTS) + assert outputs <= puf_support_module._PUF_TAX_DETAIL_NONNEGATIVE_OUTPUTS + assert outputs <= puf_support_module._PUF_TAX_DETAIL_SPARSE_PERSON_OUTPUTS + assert "alimony_income" in ( + puf_support_module._PUF_TAX_DETAIL_PRESERVE_BASE_ASEC_OUTPUTS + ) + assert "alimony_expense" not in ( + puf_support_module._PUF_TAX_DETAIL_PRESERVE_BASE_ASEC_OUTPUTS + ) + + +def test_sparse_puf_splice_does_not_prune_reported_asec_alimony_income() -> None: + person = pd.DataFrame( + { + "person_tax_unit_id": [1, 2, 3, 4], + "person_household_id": [1, 2, 3, 4], + "person_support_channel": [ + "asec", + "asec", + "puf_tax_detail", + "puf_tax_detail", + ], + "alimony_income": [1_000.0, 0.0, 500.0, 250.0], + } + ) + tables = { + "person": person, + "household": pd.DataFrame({"household_id": [1, 2, 3, 4]}), + "tax_unit": pd.DataFrame( + { + "tax_unit_id": [1, 2, 3, 4], + "tax_unit_support_channel": [ + "asec", + "asec", + "puf_tax_detail", + "puf_tax_detail", + ], + } + ), + } + + puf_support_module._sparsify_tax_unit_person_output_to_donor_positive_rate( + tables, + column="alimony_income", + donor_positive_rate=0.0, + household_weights=np.ones(4), + person_channel="person_support_channel", + tax_unit_channel="tax_unit_support_channel", + ) + + assert person.loc[:1, "alimony_income"].tolist() == [1_000.0, 0.0] + assert person.loc[2:, "alimony_income"].tolist() == [750.0, 0.0] + + +def test_signal_gate_accepts_reference_like_sparse_values() -> None: + income = np.zeros(1_000) + expense = np.zeros(1_000) + income[10] = 2_000.0 + expense[20] = 3_000.0 + frame = _PersonFrame( + pd.DataFrame({"alimony_income": income, "alimony_expense": expense}) + ) + + result = us_alimony_signal_gate(frame) # type: ignore[arg-type] + + assert result.passed, result.failures + assert result.details["columns"]["alimony_income"][ + "positive_share" + ] == pytest.approx(0.001) + + +def test_signal_gate_rejects_asec_alimony_left_in_miscellaneous_income() -> None: + n = 1_000 + codes = np.zeros(n) + amounts = np.zeros(n) + income = np.zeros(n) + expense = np.zeros(n) + miscellaneous = np.zeros(n) + codes[10] = 20 + amounts[10] = 2_000.0 + income[10] = 2_000.0 + miscellaneous[10] = 2_000.0 + expense[20] = 3_000.0 + frame = _PersonFrame( + pd.DataFrame( + { + "OI_OFF": codes, + "OI_VAL": amounts, + "alimony_income": income, + "alimony_expense": expense, + "miscellaneous_income": miscellaneous, + } + ) + ) + + result = us_alimony_signal_gate(frame) # type: ignore[arg-type] + + assert not result.passed + assert any( + "miscellaneous_income does not exclude alimony/strike" in failure + for failure in result.failures + ) + + +@pytest.mark.parametrize( + "person", + [ + pd.DataFrame({"alimony_income": [1.0]}), + pd.DataFrame({"alimony_income": [0.0], "alimony_expense": [0.0]}), + pd.DataFrame({"alimony_income": [-1.0], "alimony_expense": [1.0]}), + pd.DataFrame({"alimony_income": [1.0], "alimony_expense": [np.nan]}), + ], +) +def test_signal_gate_rejects_missing_default_or_invalid_surface( + person: pd.DataFrame, +) -> None: + result = us_alimony_signal_gate(_PersonFrame(person)) # type: ignore[arg-type] + + assert not result.passed + + +@requires_us +def test_policyengine_us_contract_is_two_person_year_input_leaves() -> None: + from policyengine_us import CountryTaxBenefitSystem + + variables = CountryTaxBenefitSystem().variables + for name in ("alimony_income", "alimony_expense"): + variable = variables[name] + assert variable.is_input_variable() + assert variable.entity.key == "person" + assert str(variable.definition_period).lower() == "year" + assert variable.default_value == 0 + + assert "alimony_expense" in ( + variables["alimony_expense_ald"].formula.__code__.co_consts + ) diff --git a/packages/populace-build/tests/test_us_asec_pool.py b/packages/populace-build/tests/test_us_asec_pool.py index 768dc336..67fdf4df 100644 --- a/packages/populace-build/tests/test_us_asec_pool.py +++ b/packages/populace-build/tests/test_us_asec_pool.py @@ -74,6 +74,7 @@ def _write_asec( "H_SEQ": household_id, "HSUP_WGT": household_weight, "GESTFIPS": state_fips, + "H_TENURE": 1 + (household_id % 2), } ) if n_people >= 2: @@ -197,7 +198,9 @@ def test_build_pooled_asec_unit_frame_runs_unit_assignment_on_pooled_source( assert metadata["weighted_person_population"] == pytest.approx(200.0) assert np.array_equal(frame.table("household")["household_id"], np.array([1, 2])) assert frame.table("household")["state_fips"].tolist() == [6, 36] + assert frame.table("household")["H_TENURE"].tolist() == [2, 2] assert "state_fips" not in frame.table("person") + assert "H_TENURE" not in frame.table("person") def test_build_pooled_asec_unit_frame_aligns_household_weights_by_remapped_id( diff --git a/packages/populace-build/tests/test_us_buildj_base_script.py b/packages/populace-build/tests/test_us_buildj_base_script.py new file mode 100644 index 00000000..111891c7 --- /dev/null +++ b/packages/populace-build/tests/test_us_buildj_base_script.py @@ -0,0 +1,39 @@ +"""Regression contracts for the Build-J reusable-base driver.""" + +from __future__ import annotations + +import ast +import re +from pathlib import Path + +ROOT = Path(__file__).resolve().parents[3] +BUILDJ_BASE_SCRIPT = ROOT / "experiments/build_j_recert/buildj_base.sh" + + +def _cache_required_columns() -> dict[str, tuple[str, ...]]: + source = BUILDJ_BASE_SCRIPT.read_text(encoding="utf-8") + match = re.search( + r"required = (?P\{.*?\n\})\ntry:", + source, + flags=re.DOTALL, + ) + assert match is not None, "Build-J cache required-column map is missing" + required = ast.literal_eval(match.group("required")) + assert isinstance(required, dict) + return required + + +def test_cache_rejects_bases_predating_salt_refund_and_energy_subsidy() -> None: + required = _cache_required_columns() + + assert "salt_refund_income" in required["person"] + assert "spm_unit_energy_subsidy" in required["spm_unit"] + assert "takes_up_housing_assistance_if_eligible" in required["spm_unit"] + + +def test_cache_rejects_bases_predating_measured_medicare_take_up() -> None: + required = _cache_required_columns() + + assert "takes_up_medicare_if_eligible" in required["person"] + assert "workers_compensation" in required["person"] + assert "would_claim_wic" in required["person"] diff --git a/packages/populace-build/tests/test_us_capital_gain_details.py b/packages/populace-build/tests/test_us_capital_gain_details.py new file mode 100644 index 00000000..cd8638ee --- /dev/null +++ b/packages/populace-build/tests/test_us_capital_gain_details.py @@ -0,0 +1,432 @@ +"""Contracts for source-backed PUF capital-gain detail inputs.""" + +from __future__ import annotations + +import importlib.util +import json +from hashlib import sha256 +from importlib.resources import files +from pathlib import Path + +import h5py +import numpy as np +import pandas as pd +import pytest + +from populace.build.us_runtime import puf_support as puf_support_module +from populace.build.us_runtime.capital_gain_details import ( + CAPITAL_GAIN_DETAILS_ARCHIVED_DERIVATION_URL, + CAPITAL_GAIN_DETAILS_ARCHIVED_EXPORT_URL, + CAPITAL_GAIN_DETAILS_ARCHIVED_IMPUTATION_URL, + CAPITAL_GAIN_DETAILS_ARCHIVED_PERSON_ALLOCATION_URL, + CAPITAL_GAIN_DETAILS_ARCHIVED_PUF_ARTIFACT_URL, + US_CAPITAL_GAIN_DETAILS_NONCONSTANT_PERSON_COLUMNS, + US_CAPITAL_GAIN_DETAILS_NONCONSTANT_TAX_UNIT_COLUMNS, + US_CAPITAL_GAIN_DETAILS_OUTPUT_COLUMNS, + derive_us_capital_gain_details_from_puf, + us_capital_gain_details_signal_gate, + us_capital_gain_details_stage_spec, +) +from populace.build.us_runtime.l0_refit_export import ( + US_RELEASE_REQUIRED_PERSON_SOURCE_COLUMNS, + US_RELEASE_REQUIRED_TAX_UNIT_SOURCE_COLUMNS, +) +from populace.build.us_runtime.puf_aggregate_records import ( + _reconcile_puf_capital_gain_details_from_sources, + derive_puf_policyengine_variables, +) +from populace.build.us_runtime.puf_support import puf_tax_unit_donor_from_arrays +from populace.build.us_runtime.reform_coverage_smoke import _build_reform +from populace.build.us_runtime.release_input_coverage import ( + RESTORED_REFERENCE_ECPS_REQUIRED_INPUTS, + load_release_input_coverage_manifest, + us_release_reform_coverage_probes, +) + +policyengine_us_installed = importlib.util.find_spec("policyengine_us") is not None +requires_us = pytest.mark.skipif( + not policyengine_us_installed, + reason="requires the policyengine-us [us] extra (build environment)", +) +ROOT = Path(__file__).resolve().parents[3] + + +class _ResolvedWeights: + def __init__(self, values: np.ndarray) -> None: + self.values = values + + +class _CapitalGainFrame: + def __init__( + self, + *, + person: pd.DataFrame, + tax_unit: pd.DataFrame, + person_weights: np.ndarray | None = None, + tax_unit_weights: np.ndarray | None = None, + ) -> None: + self._tables = {"person": person, "tax_unit": tax_unit} + self._weights = { + "person": ( + np.ones(len(person)) if person_weights is None else person_weights + ), + "tax_unit": ( + np.ones(len(tax_unit)) if tax_unit_weights is None else tax_unit_weights + ), + } + + def table(self, entity: str) -> pd.DataFrame: + return self._tables[entity] + + def resolve_weights(self, entity: str) -> _ResolvedWeights: + return _ResolvedWeights(np.asarray(self._weights[entity], dtype=np.float64)) + + +def test_archived_coordinates_pin_derivation_export_and_imputation() -> None: + commit = "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe" + + assert commit in CAPITAL_GAIN_DETAILS_ARCHIVED_DERIVATION_URL + assert CAPITAL_GAIN_DETAILS_ARCHIVED_DERIVATION_URL.endswith("puf.py#L636-L702") + assert CAPITAL_GAIN_DETAILS_ARCHIVED_EXPORT_URL.endswith("puf.py#L804-L850") + assert CAPITAL_GAIN_DETAILS_ARCHIVED_IMPUTATION_URL.endswith( + "puf_impute.py#L90-L198" + ) + assert CAPITAL_GAIN_DETAILS_ARCHIVED_PERSON_ALLOCATION_URL.endswith( + "puf.py#L1513-L1601" + ) + assert CAPITAL_GAIN_DETAILS_ARCHIVED_PUF_ARTIFACT_URL.endswith("puf.py#L1655-L1660") + + +def test_archived_puf_mappings_are_exact_carries() -> None: + source = pd.DataFrame( + { + "E24518": [0.0, 125.5, 9_000.0], + "E24515": [10.0, 0.0, 7_500.0], + } + ) + + result = derive_us_capital_gain_details_from_puf(source) + + assert set(US_CAPITAL_GAIN_DETAILS_OUTPUT_COLUMNS).isdisjoint(source.columns) + assert result["long_term_capital_gains_on_collectibles"].tolist() == [ + 0.0, + 125.5, + 9_000.0, + ] + assert result["unrecaptured_section_1250_gain"].tolist() == [ + 10.0, + 0.0, + 7_500.0, + ] + + +@pytest.mark.parametrize( + ("source", "message"), + [ + (pd.DataFrame({"E24518": [1.0]}), "requires source columns"), + ( + pd.DataFrame({"E24518": ["bad"], "E24515": [0.0]}), + "nonnumeric or nonfinite", + ), + ( + pd.DataFrame({"E24518": [0.0], "E24515": [np.inf]}), + "nonnumeric or nonfinite", + ), + ( + pd.DataFrame({"E24518": [-1.0], "E24515": [0.0]}), + "negative value", + ), + ], +) +def test_source_derivation_fails_closed( + source: pd.DataFrame, + message: str, +) -> None: + with pytest.raises(ValueError, match=message): + derive_us_capital_gain_details_from_puf(source) + + +def test_shared_puf_derivation_and_post_disaggregation_reconciliation() -> None: + source = pd.DataFrame( + { + "E00600": [0.0, 0.0, 0.0], + "E00650": [0.0, 0.0, 0.0], + "E24518": [0.0, 3_000.0, 7_500.0], + "E24515": [100.0, 0.0, 8_000.0], + } + ) + + derived = derive_puf_policyengine_variables( + source, + collectibles_capital_gain_source="E24518", + unrecaptured_section_1250_gain_source="E24515", + ) + for output in US_CAPITAL_GAIN_DETAILS_OUTPUT_COLUMNS: + derived[output] = 999.0 + result = _reconcile_puf_capital_gain_details_from_sources(derived) + + assert result["long_term_capital_gains_on_collectibles"].tolist() == [ + 0.0, + 3_000.0, + 7_500.0, + ] + assert result["unrecaptured_section_1250_gain"].tolist() == [ + 100.0, + 0.0, + 8_000.0, + ] + + +def test_shared_puf_stage_declares_sources_outputs_and_artifact() -> None: + stage = us_capital_gain_details_stage_spec() + operation = next( + operation + for operation in stage.operations + if operation.kind == "derive_puf_policyengine_variables" + ) + + assert operation.parameters["collectibles_capital_gain_source"] == "E24518" + assert operation.parameters["unrecaptured_section_1250_gain_source"] == "E24515" + assert set(US_CAPITAL_GAIN_DETAILS_OUTPUT_COLUMNS) <= set(stage.outputs) + assert set(US_CAPITAL_GAIN_DETAILS_OUTPUT_COLUMNS) <= set(stage.nonnegative_outputs) + assert any( + all(field in str(artifact.get("locator")) for field in ("E24515", "E24518")) + for artifact in stage.artifacts + ) + assert any( + "irs-soi-puf/1.8.0/puf_2024.h5" in str(artifact.get("locator")) + for artifact in stage.artifacts + ) + + +def test_puf_support_keeps_both_details_nonnegative_and_sparse() -> None: + person_output = US_CAPITAL_GAIN_DETAILS_NONCONSTANT_PERSON_COLUMNS[0] + tax_unit_output = US_CAPITAL_GAIN_DETAILS_NONCONSTANT_TAX_UNIT_COLUMNS[0] + + assert person_output in puf_support_module.PUF_TAX_DETAIL_DEFAULT_PERSON_OUTPUTS + assert tax_unit_output in puf_support_module.PUF_TAX_DETAIL_DEFAULT_TAX_UNIT_OUTPUTS + assert person_output in puf_support_module._PUF_TAX_DETAIL_NONNEGATIVE_OUTPUTS + assert tax_unit_output in puf_support_module._PUF_TAX_DETAIL_NONNEGATIVE_OUTPUTS + assert person_output in puf_support_module._PUF_TAX_DETAIL_SPARSE_PERSON_OUTPUTS + assert tax_unit_output in puf_support_module._PUF_TAX_DETAIL_SPARSE_TAX_UNIT_OUTPUTS + assert person_output not in puf_support_module._PERSON_OUTPUT_DISTRIBUTION_BASIS + + +def test_processed_puf_arrays_preserve_identifiable_tax_unit_totals() -> None: + donor = puf_tax_unit_donor_from_arrays( + { + "tax_unit_id": [10, 20], + "household_weight": [100.0, 200.0], + "filing_status": [b"SINGLE", b"JOINT"], + "person_tax_unit_id": [10, 20, 20], + "long_term_capital_gains_on_collectibles": [0.0, 100.0, 200.0], + "unrecaptured_section_1250_gain": [50.0, 400.0], + }, + person_outputs=("long_term_capital_gains_on_collectibles",), + tax_unit_outputs=("unrecaptured_section_1250_gain",), + ) + + assert donor["long_term_capital_gains_on_collectibles"].tolist() == [0.0, 300.0] + assert donor["unrecaptured_section_1250_gain"].tolist() == [50.0, 400.0] + + +def test_signal_gate_accepts_sparse_source_aligned_values() -> None: + person_values = np.zeros(2_000) + person_values[[1_010, 1_020]] = [1_000.0, 2_000.0] + tax_unit_values = np.zeros(2_000) + tax_unit_values[[1_010, 1_020, 1_030, 1_040]] = [100.0, 200.0, 300.0, 400.0] + channels = np.asarray(["asec"] * 1_000 + ["puf_tax_detail"] * 1_000) + frame = _CapitalGainFrame( + person=pd.DataFrame( + { + "long_term_capital_gains_on_collectibles": person_values, + "person_support_channel": channels, + } + ), + tax_unit=pd.DataFrame( + { + "unrecaptured_section_1250_gain": tax_unit_values, + "tax_unit_support_channel": channels, + } + ), + ) + + result = us_capital_gain_details_signal_gate(frame) # type: ignore[arg-type] + + assert result.passed, result.failures + assert result.details["long_term_capital_gains_on_collectibles"][ + "positive_share" + ] == pytest.approx(0.001) + assert result.details["unrecaptured_section_1250_gain"][ + "positive_share" + ] == pytest.approx(0.002) + + +def test_release_wiring_promotes_only_source_backed_detail_leaves() -> None: + manifest = load_release_input_coverage_manifest() + person_output = US_CAPITAL_GAIN_DETAILS_NONCONSTANT_PERSON_COLUMNS[0] + tax_unit_output = US_CAPITAL_GAIN_DETAILS_NONCONSTANT_TAX_UNIT_COLUMNS[0] + + assert person_output in US_RELEASE_REQUIRED_PERSON_SOURCE_COLUMNS + assert tax_unit_output in US_RELEASE_REQUIRED_TAX_UNIT_SOURCE_COLUMNS + assert set(US_CAPITAL_GAIN_DETAILS_OUTPUT_COLUMNS) <= ( + RESTORED_REFERENCE_ECPS_REQUIRED_INPUTS + ) + assert set(US_CAPITAL_GAIN_DETAILS_OUTPUT_COLUMNS) <= manifest.required_columns + assert "investment_interest_expense" in manifest.reviewed_exclusions + assert manifest.reviewed_exclusions["investment_interest_expense"].startswith( + "SOURCE UNAVAILABILITY WITH EVIDENCE:" + ) + + +def test_investment_interest_exclusion_has_machine_reviewable_source_evidence() -> None: + path = files("populace.build.us").joinpath("ecps_parity_known_gaps.json") + entry = json.loads(path.read_text(encoding="utf-8"))["known_gaps"][ + "investment_interest_expense" + ] + + assert entry["reason"].startswith("SOURCE UNAVAILABILITY WITH EVIDENCE:") + evidence = entry["evidence"] + assert evidence["classification"] == "source_unavailability" + assert evidence["retired_derivation"]["lines"] == "141-155,268-320,327-349" + assert evidence["required_intermediates"] == [ + "interest_deduction", + "deductible_mortgage_interest", + ] + assert evidence["processed_puf"]["positive_values"] == 0 + assert evidence["processed_puf"]["missing_columns"] == [ + "interest_deduction", + "deductible_mortgage_interest", + "E19200", + ] + assert all( + item["missing_columns"] + == ["interest_deduction", "deductible_mortgage_interest", "E19200"] + for item in evidence["hermetic_asec_inputs"] + ) + + +def test_investment_interest_evidence_matches_build_j_artifact_pins() -> None: + entry = json.loads( + files("populace.build.us") + .joinpath("ecps_parity_known_gaps.json") + .read_text(encoding="utf-8") + )["known_gaps"]["investment_interest_expense"] + evidence = entry["evidence"] + summary = json.loads( + (ROOT / "experiments/build_j_recert/base_j.summary.json").read_text() + ) + + assert evidence["processed_puf"]["sha256"] == summary["puf_sha256"] + assert { + item["filename"]: item["sha256"] for item in evidence["hermetic_asec_inputs"] + } == { + Path(item["path"]).name: item["sha256"] + for item in summary["base_source"]["sources"] + } + + +def _sha256(path: Path) -> str: + digest = sha256() + with path.open("rb") as stream: + for chunk in iter(lambda: stream.read(1024 * 1024), b""): + digest.update(chunk) + return digest.hexdigest() + + +def test_mounted_puf_artifact_confirms_zero_and_missing_source() -> None: + entry = json.loads( + files("populace.build.us") + .joinpath("ecps_parity_known_gaps.json") + .read_text(encoding="utf-8") + )["known_gaps"]["investment_interest_expense"] + evidence = entry["evidence"]["processed_puf"] + summary = json.loads( + (ROOT / "experiments/build_j_recert/base_j.summary.json").read_text() + ) + path = Path(summary["puf_h5"]) + if not path.is_file(): + pytest.skip("SHA-locked PUF artifact is not mounted in this environment") + + assert _sha256(path) == evidence["sha256"] + with h5py.File(path) as h5: + assert set(evidence["missing_columns"]).isdisjoint(h5.keys()) + values = h5["investment_interest_expense"][...] + assert len(values) == evidence["rows"] + assert int(np.count_nonzero(values > 0.0)) == evidence["positive_values"] + assert float(values.sum()) == evidence["weighted_total"] + + +def test_shipped_neutralization_probes_bind_each_restored_leaf() -> None: + probes = {probe.id: probe for probe in us_release_reform_coverage_probes()} + collectibles = probes["collectibles_gain_neutralization"] + unrecaptured = probes["unrecaptured_section_1250_gain_neutralization"] + + assert collectibles.neutralized_variable == ( + "long_term_capital_gains_on_collectibles" + ) + assert collectibles.binding_inputs == ("long_term_capital_gains_on_collectibles",) + assert collectibles.budget_measure == "income_tax" + assert collectibles.expected_sign == "positive" + assert unrecaptured.neutralized_variable == "unrecaptured_section_1250_gain" + assert unrecaptured.binding_inputs == ("unrecaptured_section_1250_gain",) + assert unrecaptured.budget_measure == "income_tax" + assert unrecaptured.expected_sign == "positive" + + +@requires_us +def test_policyengine_us_contracts_and_household_bindings() -> None: + from policyengine_us import CountryTaxBenefitSystem, Simulation + + system = CountryTaxBenefitSystem() + collectibles = system.variables["long_term_capital_gains_on_collectibles"] + unrecaptured = system.variables["unrecaptured_section_1250_gain"] + assert collectibles.is_input_variable() + assert collectibles.entity.key == "person" + assert unrecaptured.is_input_variable() + assert unrecaptured.entity.key == "tax_unit" + + adult = { + "age": {2024: 45}, + "employment_income": {2024: 120_000}, + "long_term_capital_gains_before_response": {2024: 100_000}, + "long_term_capital_gains_on_collectibles": {2024: 25_000}, + } + entities = { + "tax_units": { + "unit": { + "members": ["adult"], + "filing_status": {2024: "SINGLE"}, + "unrecaptured_section_1250_gain": {2024: 20_000}, + } + }, + "families": {"family": {"members": ["adult"]}}, + "spm_units": {"spm": {"members": ["adult"]}}, + "households": { + "household": { + "members": ["adult"], + "state_code_str": {2024: "MA"}, + } + }, + "marital_units": {"marital": {"members": ["adult"]}}, + } + baseline = Simulation(situation={"people": {"adult": adult}, **entities}) + + probes = {probe.id: probe for probe in us_release_reform_coverage_probes()} + no_collectibles = Simulation( + situation={"people": {"adult": adult}, **entities}, + reform=_build_reform(probes["collectibles_gain_neutralization"]), + ) + no_unrecaptured = Simulation( + situation={"people": {"adult": adult}, **entities}, + reform=_build_reform(probes["unrecaptured_section_1250_gain_neutralization"]), + ) + + assert ( + baseline.calculate("income_tax", 2024)[0] + > no_collectibles.calculate("income_tax", 2024)[0] + ) + assert ( + baseline.calculate("income_tax", 2024)[0] + > no_unrecaptured.calculate("income_tax", 2024)[0] + ) diff --git a/packages/populace-build/tests/test_us_casualty_losses.py b/packages/populace-build/tests/test_us_casualty_losses.py new file mode 100644 index 00000000..8c090ebc --- /dev/null +++ b/packages/populace-build/tests/test_us_casualty_losses.py @@ -0,0 +1,154 @@ +"""Contracts for the IRS PUF casualty-loss input restoration.""" + +from __future__ import annotations + +import importlib.util + +import numpy as np +import pandas as pd +import pytest + +from populace.build.us_runtime import puf_support as puf_support_module +from populace.build.us_runtime.casualty_losses import ( + derive_us_casualty_loss_from_puf, + us_casualty_loss_signal_gate, + us_casualty_loss_stage_spec, +) +from populace.build.us_runtime.puf_aggregate_records import ( + _reconcile_puf_casualty_loss_from_source, +) + +policyengine_us_installed = importlib.util.find_spec("policyengine_us") is not None +requires_us = pytest.mark.skipif( + not policyengine_us_installed, + reason="requires the policyengine-us [us] extra (build environment)", +) + + +class _ResolvedWeights: + def __init__(self, values: np.ndarray) -> None: + self.values = values + + +class _PersonFrame: + def __init__( + self, + person: pd.DataFrame, + weights: np.ndarray | None = None, + ) -> None: + self._person = person + self._weights = np.ones(len(person)) if weights is None else weights + + def table(self, entity: str) -> pd.DataFrame: + assert entity == "person" + return self._person + + def resolve_weights(self, entity: str) -> _ResolvedWeights: + assert entity == "person" + return _ResolvedWeights(np.asarray(self._weights, dtype=np.float64)) + + +def test_archived_puf_mapping_is_an_exact_carry() -> None: + source = pd.DataFrame({"E20500": [0.0, 125.5, 9_000.0]}) + + result = derive_us_casualty_loss_from_puf(source) + + assert "casualty_loss" not in source.columns + assert result["casualty_loss"].tolist() == [0.0, 125.5, 9_000.0] + + +@pytest.mark.parametrize( + ("source", "message"), + [ + (pd.DataFrame({"other": [1.0]}), "requires source column"), + (pd.DataFrame({"E20500": ["not numeric"]}), "nonnumeric or nonfinite"), + (pd.DataFrame({"E20500": [np.inf]}), "nonnumeric or nonfinite"), + (pd.DataFrame({"E20500": [-1.0]}), "negative value"), + ], +) +def test_source_derivation_fails_closed( + source: pd.DataFrame, + message: str, +) -> None: + with pytest.raises(ValueError, match=message): + derive_us_casualty_loss_from_puf(source) + + +def test_post_disaggregation_reconciliation_uses_final_e20500() -> None: + source = pd.DataFrame( + { + "E20500": [0.0, 3_000.0, 7_500.0], + "casualty_loss": [999.0, 999.0, 999.0], + } + ) + + result = _reconcile_puf_casualty_loss_from_source(source) + + assert result["casualty_loss"].tolist() == [0.0, 3_000.0, 7_500.0] + + +def test_shared_puf_stage_declares_exact_source_and_output() -> None: + stage = us_casualty_loss_stage_spec() + operation = next( + operation + for operation in stage.operations + if operation.kind == "derive_puf_policyengine_variables" + ) + + assert operation.parameters["casualty_loss_source"] == "E20500" + assert operation.parameters["casualty_loss_output"] == "casualty_loss" + assert "casualty_loss" in stage.outputs + assert "casualty_loss" in stage.nonnegative_outputs + assert any("E20500" in str(artifact.get("locator")) for artifact in stage.artifacts) + + +def test_puf_support_keeps_casualty_loss_sparse_and_earnings_distributed() -> None: + assert "casualty_loss" in puf_support_module.PUF_TAX_DETAIL_DEFAULT_PERSON_OUTPUTS + assert "casualty_loss" in puf_support_module._PUF_TAX_DETAIL_NONNEGATIVE_OUTPUTS + assert "casualty_loss" in puf_support_module._PUF_TAX_DETAIL_SPARSE_PERSON_OUTPUTS + assert puf_support_module._PERSON_OUTPUT_DISTRIBUTION_BASIS["casualty_loss"] == ( + "employment_income_before_lsr", + "self_employment_income_before_lsr", + ) + + +def test_signal_gate_accepts_plausibly_sparse_nondefault_values() -> None: + values = np.zeros(1_000) + values[[10, 20, 30]] = [1_000.0, 2_000.0, 3_000.0] + frame = _PersonFrame(pd.DataFrame({"casualty_loss": values})) + + result = us_casualty_loss_signal_gate(frame) # type: ignore[arg-type] + + assert result.passed, result.failures + assert result.details["positive_share"] == pytest.approx(0.003) + + +@pytest.mark.parametrize( + "person", + [ + pd.DataFrame({"other": [0.0, 1.0]}), + pd.DataFrame({"casualty_loss": [0.0, 0.0]}), + pd.DataFrame({"casualty_loss": [0.0, -1.0]}), + pd.DataFrame({"casualty_loss": [0.0, np.nan]}), + ], +) +def test_signal_gate_rejects_missing_default_or_invalid_surface( + person: pd.DataFrame, +) -> None: + result = us_casualty_loss_signal_gate( # type: ignore[arg-type] + _PersonFrame(person) + ) + + assert not result.passed + + +@requires_us +def test_policyengine_us_contract_is_a_person_year_input_leaf() -> None: + from policyengine_us import CountryTaxBenefitSystem + + variable = CountryTaxBenefitSystem().variables["casualty_loss"] + + assert variable.is_input_variable() + assert variable.entity.key == "person" + assert str(variable.definition_period).lower() == "year" + assert variable.default_value == 0 diff --git a/packages/populace-build/tests/test_us_child_support.py b/packages/populace-build/tests/test_us_child_support.py new file mode 100644 index 00000000..dfce2acc --- /dev/null +++ b/packages/populace-build/tests/test_us_child_support.py @@ -0,0 +1,483 @@ +"""ASEC child-support restoration and retired PUF-half joint QRF treatment.""" + +from __future__ import annotations + +import importlib.util + +import numpy as np +import pandas as pd +import pytest + +import populace.build.us_runtime.child_support as module +from populace.build.source_runtime import ( + SourceRuntimeConfig, + SourceRuntimeContext, + SourceRuntimeError, +) +from populace.build.us_runtime.child_support import ( + CHILD_SUPPORT_ARCHIVED_PUF_IMPUTATION_URL, + CHILD_SUPPORT_ARCHIVED_PUF_OUTPUTS_URL, + CHILD_SUPPORT_EXPENSE_ARCHIVED_DERIVATION_URL, + CHILD_SUPPORT_RECEIVED_ARCHIVED_DERIVATION_URL, + US_CHILD_SUPPORT_OUTPUT_COLUMNS, + derive_us_child_support_from_manifest, + impute_us_child_support_to_puf_support_from_manifest, + us_child_support_signal_gate, + us_child_support_stage_spec, + with_us_child_support_inputs, +) +from populace.build.us_runtime.puf_support import clone_us_frame_for_puf_support +from populace.frame import US_SCHEMA, Frame, WeightKind, Weights + +policyengine_us_installed = importlib.util.find_spec("policyengine_us") is not None +requires_us = pytest.mark.skipif( + not policyengine_us_installed, + reason="requires the policyengine-us [us] extra (build environment)", +) + +_RECEIVED, _EXPENSE = US_CHILD_SUPPORT_OUTPUT_COLUMNS + + +def _person_source() -> pd.DataFrame: + count = 10 + return pd.DataFrame( + { + "person_id": np.arange(1, count + 1, dtype="int64"), + "person_household_id": np.arange(1, count + 1, dtype="int64") * 10, + "person_tax_unit_id": np.arange(1, count + 1, dtype="int64") * 100, + "person_spm_unit_id": np.arange(1, count + 1, dtype="int64") * 1_000, + "person_family_id": np.arange(1, count + 1, dtype="int64") * 10_000, + "person_marital_unit_id": ( + np.arange(1, count + 1, dtype="int64") * 100_000 + ), + "CSP_VAL": [3_600.0, *([0.0] * 9)], + "CHSP_VAL": [2_400.0, *([0.0] * 9)], + "WSAL_VAL": np.linspace(10_000.0, 100_000.0, count), + "SEMP_VAL": np.zeros(count), + "employment_income_before_lsr": np.linspace(10_000.0, 100_000.0, count), + "self_employment_income_before_lsr": np.zeros(count), + "age": np.arange(25, 25 + count), + "is_female": np.asarray([False, True] * 5), + "has_esi": np.asarray([True, False] * 5), + "tax_unit_role_input": ["HEAD"] * count, + "social_security_retirement": np.zeros(count), + "social_security_disability": np.zeros(count), + "social_security_dependents": np.zeros(count), + "social_security_survivors": np.zeros(count), + } + ) + + +def _frame() -> Frame: + person = _person_source() + count = len(person) + ids = { + "household": person["person_household_id"].to_numpy(), + "tax_unit": person["person_tax_unit_id"].to_numpy(), + "spm_unit": person["person_spm_unit_id"].to_numpy(), + "family": person["person_family_id"].to_numpy(), + "marital_unit": person["person_marital_unit_id"].to_numpy(), + } + tables = { + entity: pd.DataFrame({f"{entity}_id": values}) for entity, values in ids.items() + } + tables["person"] = person + tables["tax_unit"]["filing_status_input"] = ["SINGLE"] * count + return Frame( + tables, + US_SCHEMA, + { + "household": Weights( + np.arange(1.0, count + 1.0), + WeightKind.DESIGN, + ) + }, + ) + + +def _derive(person: pd.DataFrame) -> pd.DataFrame: + operation = next( + operation + for operation in us_child_support_stage_spec().operations + if operation.kind == "derive_child_support_inputs" + ) + return derive_us_child_support_from_manifest(person, operation, None) + + +def test_archived_sources_are_sha_and_line_pinned() -> None: + commit = "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe" + assert commit in CHILD_SUPPORT_RECEIVED_ARCHIVED_DERIVATION_URL + assert CHILD_SUPPORT_RECEIVED_ARCHIVED_DERIVATION_URL.endswith( + "datasets/cps/cps.py#L1493-L1496" + ) + assert commit in CHILD_SUPPORT_EXPENSE_ARCHIVED_DERIVATION_URL + assert CHILD_SUPPORT_EXPENSE_ARCHIVED_DERIVATION_URL.endswith( + "datasets/cps/cps.py#L1572-L1574" + ) + assert commit in CHILD_SUPPORT_ARCHIVED_PUF_OUTPUTS_URL + assert CHILD_SUPPORT_ARCHIVED_PUF_OUTPUTS_URL.endswith( + "datasets/cps/extended_cps.py#L135-L194" + ) + assert commit in CHILD_SUPPORT_ARCHIVED_PUF_IMPUTATION_URL + assert CHILD_SUPPORT_ARCHIVED_PUF_IMPUTATION_URL.endswith( + "datasets/cps/extended_cps.py#L639-L745" + ) + + +def test_stage_manifest_pins_direct_sources_and_one_joint_qrf() -> None: + spec = us_child_support_stage_spec() + + assert spec.stage == "child_support_inputs" + assert spec.grain == "person" + assert tuple(spec.outputs) == US_CHILD_SUPPORT_OUTPUT_COLUMNS + assert [operation.kind for operation in spec.operations] == [ + "read_table", + "derive_child_support_inputs", + "impute_child_support_to_puf_support", + ] + assert spec.operations[1].parameters == { + "received_source": "CSP_VAL", + "received_output": _RECEIVED, + "expense_source": "CHSP_VAL", + "expense_output": _EXPENSE, + } + assert spec.operations[2].parameters == { + "predictors": [ + "age", + "is_male", + "has_esi", + "tax_unit_is_joint", + "tax_unit_count_dependents", + "employment_income", + "self_employment_income", + "social_security", + ], + "max_train_samples": 5_000, + "n_estimators": 100, + "seed_from_build_config": True, + "weight": "person_weight", + } + + +def test_direct_carry_preserves_positive_annual_values_and_topcodes() -> None: + source = pd.DataFrame( + { + "CSP_VAL": [0.0, 99_999.0, 1_500.0, 800.0], + "CHSP_VAL": [90_000.0, 0.0, 700.0, 800.0], + } + ) + original = source.copy(deep=True) + + result = _derive(source) + + assert result[_RECEIVED].tolist() == source["CSP_VAL"].tolist() + assert result[_EXPENSE].tolist() == source["CHSP_VAL"].tolist() + assert result.loc[3, [_RECEIVED, _EXPENSE]].tolist() == [800.0, 800.0] + pd.testing.assert_frame_equal(source, original) + + +@pytest.mark.parametrize("missing", ["CSP_VAL", "CHSP_VAL"]) +def test_direct_carry_fails_closed_when_source_is_missing(missing: str) -> None: + with pytest.raises(SourceRuntimeError, match=missing): + _derive(_person_source().drop(columns=[missing])) + + +@pytest.mark.parametrize( + ("column", "bad_value", "message"), + [ + ("CSP_VAL", np.nan, "nonfinite"), + ("CHSP_VAL", np.inf, "nonfinite"), + ("CSP_VAL", -1.0, "negative"), + ("CHSP_VAL", -1.0, "negative"), + ], +) +def test_direct_carry_rejects_invalid_source_values( + column: str, + bad_value: float, + message: str, +) -> None: + person = _person_source() + person.loc[0, column] = bad_value + + with pytest.raises(SourceRuntimeError, match=message): + _derive(person) + + +def test_with_inputs_materializes_person_leaves_without_mutating_source() -> None: + frame = _frame() + original = frame.table("person").copy(deep=True) + + result = with_us_child_support_inputs(frame, seed=0, time_period=2024) + + pd.testing.assert_frame_equal(frame.table("person"), original) + assert result.table("person")[_RECEIVED].tolist() == original["CSP_VAL"].tolist() + assert result.table("person")[_EXPENSE].tolist() == original["CHSP_VAL"].tolist() + gate = us_child_support_signal_gate(result) + assert gate.passed, gate.failures + + +def test_release_frame_requires_opt_in_to_preserve_valid_existing_surface() -> None: + materialized = with_us_child_support_inputs(_frame(), seed=0, time_period=2024) + tables = { + entity: materialized.table(entity).copy() for entity in materialized.entities + } + tables["person"] = tables["person"].drop(columns=["CSP_VAL", "CHSP_VAL"]) + release = Frame( + tables, + materialized.schema, + { + entity: materialized.weights_for(entity) + for entity in materialized.weighted_entities + }, + materialized.strata, + mass_log=materialized.mass_log, + ) + + with pytest.raises(ValueError, match="cannot heal.*without measured"): + with_us_child_support_inputs(release, seed=0, time_period=2024) + + assert ( + with_us_child_support_inputs( + release, + seed=0, + time_period=2024, + allow_existing_without_source=True, + ) + is release + ) + + +def test_puf_half_uses_one_joint_qrf_in_archived_target_order( + monkeypatch: pytest.MonkeyPatch, +) -> None: + direct = with_us_child_support_inputs(_frame(), seed=0, time_period=2024) + expanded = clone_us_frame_for_puf_support(direct) + calls: dict[str, object] = {} + + class FakeFitted: + def predict(self, test: pd.DataFrame) -> pd.DataFrame: + calls["test"] = test.copy() + received = np.zeros(len(test), dtype=np.float64) + expense = np.zeros(len(test), dtype=np.float64) + received[0] = 7_200.0 + expense[0] = 4_800.0 + return pd.DataFrame( + {_RECEIVED: received, _EXPENSE: expense}, + index=test.index, + ) + + class FakeQRF: + def __init__(self, **kwargs: object) -> None: + calls["init"] = kwargs + + def fit( + self, + training: pd.DataFrame, + predictors: list[str], + targets: list[str], + *, + weights: np.ndarray, + ) -> FakeFitted: + calls["fit_count"] = int(calls.get("fit_count", 0)) + 1 + calls["training"] = training.copy() + calls["predictors"] = predictors + calls["targets"] = targets + calls["weights"] = weights.copy() + return FakeFitted() + + monkeypatch.setattr(module, "QRF", FakeQRF) + + result = with_us_child_support_inputs(expanded, seed=7, time_period=2024) + + assert calls["fit_count"] == 1 + assert calls["init"] == {"n_estimators": 100, "seed": 7} + assert calls["targets"] == [_RECEIVED, _EXPENSE] + assert calls["predictors"] == list( + us_child_support_stage_spec().operations[2].parameters["predictors"] + ) + training = calls["training"] + assert isinstance(training, pd.DataFrame) + assert len(training) == 10 + weights = calls["weights"] + assert isinstance(weights, np.ndarray) + assert weights.shape == (10,) + + person = result.table("person") + asec = person[person["person_support_channel"] == "asec"] + puf = person[person["person_support_channel"] == "puf_tax_detail"] + assert asec[_RECEIVED].tolist() == _person_source()["CSP_VAL"].tolist() + assert asec[_EXPENSE].tolist() == _person_source()["CHSP_VAL"].tolist() + assert puf[_RECEIVED].tolist() == [7_200.0, *([0.0] * 9)] + assert puf[_EXPENSE].tolist() == [4_800.0, *([0.0] * 9)] + gate = us_child_support_signal_gate(result) + assert gate.passed, gate.failures + + +def test_puf_qrf_caps_training_at_5000_and_keeps_weights_aligned( + monkeypatch: pytest.MonkeyPatch, +) -> None: + operation = us_child_support_stage_spec().operations[2] + predictors = tuple(operation.parameters["predictors"]) + asec_rows = 5_010 + puf_rows = 2 + rows = asec_rows + puf_rows + frame = pd.DataFrame( + { + "person_support_channel": ["asec"] * asec_rows + + ["puf_tax_detail"] * puf_rows, + "person_weight": np.arange(1.0, rows + 1.0), + _RECEIVED: np.tile([0.0, 500.0], (rows + 1) // 2)[:rows], + _EXPENSE: np.tile([300.0, 0.0], (rows + 1) // 2)[:rows], + **{ + f"child_support_predictor_{predictor}": np.arange( + rows, dtype=np.float64 + ) + for predictor in predictors + }, + } + ) + calls: dict[str, object] = {} + + class FakeFitted: + def predict(self, test: pd.DataFrame) -> pd.DataFrame: + calls["test_rows"] = len(test) + return pd.DataFrame( + { + _RECEIVED: np.zeros(len(test)), + _EXPENSE: np.zeros(len(test)), + }, + index=test.index, + ) + + class FakeQRF: + def __init__(self, **kwargs: object) -> None: + calls["init"] = kwargs + + def fit( + self, + training: pd.DataFrame, + predictor_names: list[str], + targets: list[str], + *, + weights: np.ndarray, + ) -> FakeFitted: + calls["training_rows"] = len(training) + calls["training_index"] = training.index.to_numpy() + calls["weights"] = weights.copy() + return FakeFitted() + + monkeypatch.setattr(module, "QRF", FakeQRF) + context = SourceRuntimeContext( + config=SourceRuntimeConfig(seed=19, target_year=2024), + tables={}, + ) + + impute_us_child_support_to_puf_support_from_manifest(frame, operation, context) + + assert calls["init"] == {"n_estimators": 100, "seed": 19} + assert calls["training_rows"] == 5_000 + assert calls["test_rows"] == 2 + np.testing.assert_allclose( + calls["weights"], + frame.loc[calls["training_index"], "person_weight"].to_numpy(), + ) + + +def test_signal_gate_rejects_missing_default_invalid_and_dead_puf_surface( + monkeypatch: pytest.MonkeyPatch, +) -> None: + valid = with_us_child_support_inputs(_frame(), seed=0, time_period=2024) + assert us_child_support_signal_gate(valid).passed + + candidates: list[Frame] = [] + for column, replacement in ( + (_RECEIVED, None), + (_RECEIVED, np.zeros(10)), + (_EXPENSE, np.zeros(10)), + (_RECEIVED, np.asarray([-1.0, *([0.0] * 9)])), + (_EXPENSE, np.asarray([np.nan, *([0.0] * 9)])), + ): + candidate = with_us_child_support_inputs(_frame(), seed=0, time_period=2024) + if replacement is None: + candidate.table("person").drop(columns=[column], inplace=True) + else: + candidate.table("person")[column] = replacement + candidates.append(candidate) + assert all( + not us_child_support_signal_gate(candidate).passed for candidate in candidates + ) + + class ZeroFitted: + def predict(self, test: pd.DataFrame) -> pd.DataFrame: + return pd.DataFrame( + { + _RECEIVED: np.zeros(len(test)), + _EXPENSE: np.zeros(len(test)), + }, + index=test.index, + ) + + class ZeroQRF: + def __init__(self, **kwargs: object) -> None: + pass + + def fit(self, *args: object, **kwargs: object) -> ZeroFitted: + return ZeroFitted() + + monkeypatch.setattr(module, "QRF", ZeroQRF) + expanded = clone_us_frame_for_puf_support(valid) + dead_puf = with_us_child_support_inputs(expanded, seed=0, time_period=2024) + gate = us_child_support_signal_gate(dead_puf) + assert not gate.passed + assert any("puf_tax_detail nonzero share" in failure for failure in gate.failures) + + +@requires_us +def test_policyengine_us_contract_is_two_person_year_input_leaves() -> None: + from policyengine_us import CountryTaxBenefitSystem + + variables = CountryTaxBenefitSystem().variables + for name in US_CHILD_SUPPORT_OUTPUT_COLUMNS: + variable = variables[name] + assert variable.is_input_variable() + assert variable.entity.key == "person" + assert str(variable.definition_period).lower() == "year" + assert variable.default_value == 0 + + +@requires_us +def test_policyengine_us_graph_uses_positive_annual_received_and_expense() -> None: + from policyengine_us import Simulation + + def situation(received: float, expense: float) -> dict[str, object]: + return { + "people": { + "adult": { + "age": {"2024": 35}, + _RECEIVED: {"2024": received}, + _EXPENSE: {"2024": expense}, + } + }, + "spm_units": {"spm_unit": {"members": ["adult"]}}, + "households": { + "household": { + "members": ["adult"], + "state_code": {"2024": "CA"}, + } + }, + } + + baseline = Simulation(situation=situation(0.0, 0.0)) + active = Simulation(situation=situation(3_600.0, 2_400.0)) + + assert ( + active.calculate("spm_unit_benefits", 2024)[0] + - baseline.calculate("spm_unit_benefits", 2024)[0] + ) == pytest.approx(3_600.0) + assert ( + active.calculate("spm_unit_spm_expenses", 2024)[0] + - baseline.calculate("spm_unit_spm_expenses", 2024)[0] + ) == pytest.approx(2_400.0) + assert active.calculate("snap_child_support_deduction", "2024-01")[ + 0 + ] == pytest.approx(200.0) diff --git a/packages/populace-build/tests/test_us_childcare.py b/packages/populace-build/tests/test_us_childcare.py new file mode 100644 index 00000000..9e5b5f67 --- /dev/null +++ b/packages/populace-build/tests/test_us_childcare.py @@ -0,0 +1,387 @@ +"""ASEC childcare restoration and PUF-half QRF treatment.""" + +from __future__ import annotations + +import importlib.util + +import numpy as np +import pandas as pd +import pytest + +import populace.build.us_runtime.childcare as module +from populace.build.source_runtime import ( + SourceRuntimeConfig, + SourceRuntimeContext, + SourceRuntimeError, +) +from populace.build.us_runtime.childcare import ( + US_CHILDCARE_OUTPUT_COLUMNS, + derive_us_childcare_from_manifest, + impute_us_childcare_to_puf_support_from_manifest, + us_childcare_signal_gate, + us_childcare_stage_spec, + with_us_childcare_inputs, +) +from populace.build.us_runtime.puf_support import clone_us_frame_for_puf_support +from populace.frame import US_SCHEMA, Frame, WeightKind, Weights + +policyengine_us_installed = importlib.util.find_spec("policyengine_us") is not None +requires_us = pytest.mark.skipif( + not policyengine_us_installed, + reason="requires the policyengine-us [us] extra (build environment)", +) + +_OUTPUT = US_CHILDCARE_OUTPUT_COLUMNS[0] + + +def _person_source() -> pd.DataFrame: + return pd.DataFrame( + { + "person_id": np.arange(1, 7, dtype="int64"), + "person_household_id": [10, 10, 20, 30, 40, 50], + "person_tax_unit_id": [100, 100, 200, 300, 400, 500], + "person_spm_unit_id": [1_000, 1_000, 2_000, 3_000, 4_000, 5_000], + "person_family_id": [10_000, 10_000, 20_000, 30_000, 40_000, 50_000], + "person_marital_unit_id": [ + 100_000, + 100_000, + 200_000, + 300_000, + 400_000, + 500_000, + ], + "SPM_CHILDCAREXPNS": [800.0, 800.0, 0.0, 1_200.0, 0.0, 0.0], + "WSAL_VAL": [50_000.0, 20_000.0, 0.0, 35_000.0, 10_000.0, 0.0], + "SEMP_VAL": [0.0, 0.0, 20_000.0, 0.0, 5_000.0, 0.0], + "employment_income_before_lsr": [ + 50_000.0, + 20_000.0, + 0.0, + 35_000.0, + 10_000.0, + 0.0, + ], + "self_employment_income_before_lsr": [ + 0.0, + 0.0, + 20_000.0, + 0.0, + 5_000.0, + 0.0, + ], + "age": [35, 33, 45, 29, 55, 70], + "is_female": [False, True, True, False, True, False], + "has_esi": [True, True, False, True, False, False], + "tax_unit_role_input": [ + "HEAD", + "SPOUSE", + "HEAD", + "HEAD", + "HEAD", + "HEAD", + ], + "social_security_retirement": [0.0, 0.0, 0.0, 0.0, 0.0, 15_000.0], + "social_security_disability": [0.0] * 6, + "social_security_dependents": [0.0] * 6, + "social_security_survivors": [0.0] * 6, + } + ) + + +def _frame() -> Frame: + person = _person_source() + ids = { + "household": [10, 20, 30, 40, 50], + "tax_unit": [100, 200, 300, 400, 500], + "spm_unit": [1_000, 2_000, 3_000, 4_000, 5_000], + "family": [10_000, 20_000, 30_000, 40_000, 50_000], + "marital_unit": [100_000, 200_000, 300_000, 400_000, 500_000], + } + tables = { + entity: pd.DataFrame({f"{entity}_id": np.asarray(values, dtype="int64")}) + for entity, values in ids.items() + } + tables["person"] = person + tables["tax_unit"]["filing_status_input"] = [ + "JOINT", + "SINGLE", + "SINGLE", + "SINGLE", + "SINGLE", + ] + return Frame( + tables, + US_SCHEMA, + { + "household": Weights( + np.ones(5, dtype=np.float64), + WeightKind.DESIGN, + ) + }, + ) + + +def _derive(frame: pd.DataFrame) -> pd.DataFrame: + operation = next( + operation + for operation in us_childcare_stage_spec().operations + if operation.kind == "derive_childcare_inputs" + ) + return derive_us_childcare_from_manifest(frame, operation, None) + + +def test_stage_manifest_pins_source_qrf_and_first_person_reduction() -> None: + spec = us_childcare_stage_spec() + + assert spec.stage == "childcare_inputs" + assert spec.grain == "person" + assert tuple(spec.outputs) == US_CHILDCARE_OUTPUT_COLUMNS + assert [operation.kind for operation in spec.operations] == [ + "read_table", + "derive_childcare_inputs", + "impute_childcare_to_puf_support", + ] + impute = spec.operations[2] + assert impute.parameters["max_train_samples"] == 5_000 + assert impute.parameters["n_estimators"] == 100 + assert impute.parameters["weight"] == "person_weight" + assert impute.parameters["reduction"] == "value_from_first_person" + + +def test_direct_source_is_exact_and_replicated() -> None: + result = _derive(_person_source()) + + assert result[_OUTPUT].tolist() == [800.0, 800.0, 0.0, 1_200.0, 0.0, 0.0] + + +def test_direct_source_fails_closed_when_missing() -> None: + with pytest.raises(SourceRuntimeError, match="SPM_CHILDCAREXPNS"): + _derive(_person_source().drop(columns=["SPM_CHILDCAREXPNS"])) + + +@pytest.mark.parametrize( + ("bad_value", "message"), + [(np.nan, "nonfinite"), (-1.0, "negative")], +) +def test_direct_source_rejects_invalid_values( + bad_value: float, + message: str, +) -> None: + person = _person_source() + person.loc[[0, 1], "SPM_CHILDCAREXPNS"] = bad_value + + with pytest.raises(SourceRuntimeError, match=message): + _derive(person) + + +def test_direct_source_rejects_inconsistent_unit_replicas() -> None: + person = _person_source() + person.loc[1, "SPM_CHILDCAREXPNS"] = 700.0 + + with pytest.raises(SourceRuntimeError, match="disagrees within replicated"): + _derive(person) + + +def test_with_inputs_materializes_spm_unit_values() -> None: + result = with_us_childcare_inputs(_frame(), seed=0, time_period=2024) + + assert result.table("spm_unit")[_OUTPUT].tolist() == [ + 800.0, + 0.0, + 1_200.0, + 0.0, + 0.0, + ] + gate = us_childcare_signal_gate(result) + assert gate.passed, gate.failures + + +def test_existing_output_is_recomputed_from_measured_source() -> None: + frame = _frame() + frame.table("spm_unit")[_OUTPUT] = [111.0, 222.0, 333.0, 444.0, 555.0] + + result = with_us_childcare_inputs(frame, seed=0, time_period=2024) + + assert result.table("spm_unit")[_OUTPUT].tolist() == [ + 800.0, + 0.0, + 1_200.0, + 0.0, + 0.0, + ] + + +def test_release_frame_without_raw_source_preserves_valid_output() -> None: + materialized = with_us_childcare_inputs(_frame(), seed=0, time_period=2024) + tables = { + entity: materialized.table(entity).copy() for entity in materialized.entities + } + tables["person"] = tables["person"].drop(columns=["SPM_CHILDCAREXPNS"]) + release = Frame( + tables, + materialized.schema, + { + entity: materialized.weights_for(entity) + for entity in materialized.weighted_entities + }, + materialized.strata, + ) + + with pytest.raises(ValueError, match="cannot heal.*without measured"): + with_us_childcare_inputs(release, seed=0, time_period=2024) + + result = with_us_childcare_inputs( + release, + seed=0, + time_period=2024, + allow_existing_without_source=True, + ) + + assert result is release + + +def test_puf_half_uses_qrf_and_first_person_spm_reduction( + monkeypatch: pytest.MonkeyPatch, +) -> None: + direct = with_us_childcare_inputs(_frame(), seed=0, time_period=2024) + expanded = clone_us_frame_for_puf_support(direct) + calls: dict[str, object] = {} + + class FakeFitted: + def predict(self, test: pd.DataFrame) -> pd.DataFrame: + calls["test"] = test.copy() + return pd.DataFrame( + {_OUTPUT: np.arange(100.0, 100.0 + len(test))}, + index=test.index, + ) + + class FakeQRF: + def __init__(self, **kwargs: object) -> None: + calls["init"] = kwargs + + def fit( + self, + training: pd.DataFrame, + predictors: list[str], + targets: list[str], + *, + weights: np.ndarray, + ) -> FakeFitted: + calls["training"] = training.copy() + calls["predictors"] = predictors + calls["targets"] = targets + calls["weights"] = weights.copy() + return FakeFitted() + + monkeypatch.setattr(module, "QRF", FakeQRF) + + result = with_us_childcare_inputs(expanded, seed=7, time_period=2024) + + assert calls["init"] == {"n_estimators": 100, "seed": 7} + assert len(calls["training"]) == 6 + assert len(calls["test"]) == 6 + spm = result.table("spm_unit") + asec = spm[spm["spm_unit_support_channel"] == "asec"] + puf = spm[spm["spm_unit_support_channel"] == "puf_tax_detail"] + assert asec[_OUTPUT].tolist() == [800.0, 0.0, 1_200.0, 0.0, 0.0] + # The first two PUF people share an SPM unit, so prediction 101 is ignored. + assert puf[_OUTPUT].tolist() == [100.0, 102.0, 103.0, 104.0, 105.0] + + +def test_puf_qrf_caps_training_at_5000_and_keeps_aligned_weights( + monkeypatch: pytest.MonkeyPatch, +) -> None: + operation = next( + operation + for operation in us_childcare_stage_spec().operations + if operation.kind == "impute_childcare_to_puf_support" + ) + predictors = tuple(operation.parameters["predictors"]) + asec_rows = 5_010 + puf_rows = 2 + frame = pd.DataFrame( + { + "person_support_channel": ["asec"] * asec_rows + + ["puf_tax_detail"] * puf_rows, + "person_weight": np.arange(1.0, asec_rows + puf_rows + 1.0), + _OUTPUT: np.tile([0.0, 500.0], (asec_rows + puf_rows + 1) // 2)[ + : asec_rows + puf_rows + ], + **{ + f"childcare_predictor_{predictor}": np.arange( + asec_rows + puf_rows, dtype=np.float64 + ) + for predictor in predictors + }, + } + ) + calls: dict[str, object] = {} + + class FakeFitted: + def predict(self, test: pd.DataFrame) -> pd.DataFrame: + calls["test_rows"] = len(test) + return pd.DataFrame({_OUTPUT: np.zeros(len(test))}, index=test.index) + + class FakeQRF: + def __init__(self, **kwargs: object) -> None: + calls["init"] = kwargs + + def fit( + self, + training: pd.DataFrame, + predictor_names: list[str], + targets: list[str], + *, + weights: np.ndarray, + ) -> FakeFitted: + calls["training_rows"] = len(training) + calls["weights"] = weights.copy() + calls["training_index"] = training.index.to_numpy() + return FakeFitted() + + monkeypatch.setattr(module, "QRF", FakeQRF) + context = SourceRuntimeContext( + config=SourceRuntimeConfig(seed=19, target_year=2024), + tables={}, + ) + + impute_us_childcare_to_puf_support_from_manifest(frame, operation, context) + + assert calls["init"] == {"n_estimators": 100, "seed": 19} + assert calls["training_rows"] == 5_000 + assert calls["test_rows"] == 2 + np.testing.assert_allclose( + calls["weights"], + frame.loc[calls["training_index"], "person_weight"].to_numpy(), + ) + + +def test_signal_gate_rejects_missing_default_and_invalid_surfaces() -> None: + frame = with_us_childcare_inputs(_frame(), seed=0, time_period=2024) + for values in ( + None, + [0.0] * 5, + [800.0, 0.0, -1.0, 0.0, 0.0], + [800.0, 0.0, np.nan, 0.0, 0.0], + ): + candidate = _frame() + if values is not None: + candidate.table("spm_unit")[_OUTPUT] = values + gate = us_childcare_signal_gate(candidate) + assert not gate.passed + + assert us_childcare_signal_gate(frame).passed + + +@requires_us +def test_policyengine_us_contract_is_spm_unit_year_input_leaf() -> None: + from policyengine_us import CountryTaxBenefitSystem + + variables = CountryTaxBenefitSystem().variables + variable = variables[_OUTPUT] + + assert variable.is_input_variable() + assert variable.entity.key == "spm_unit" + assert str(variable.definition_period).lower() == "year" + assert variable.default_value == 0 + assert not variables["cdcc_relevant_expenses"].is_input_variable() diff --git a/packages/populace-build/tests/test_us_dc_ptc_take_up_exclusion.py b/packages/populace-build/tests/test_us_dc_ptc_take_up_exclusion.py new file mode 100644 index 00000000..1d6872e7 --- /dev/null +++ b/packages/populace-build/tests/test_us_dc_ptc_take_up_exclusion.py @@ -0,0 +1,212 @@ +"""Evidence contract for the irreducible DC PTC take-up exclusion.""" + +from __future__ import annotations + +import importlib.util +import json +from hashlib import sha256 +from importlib.metadata import version +from importlib.resources import files +from pathlib import Path + +import h5py +import pytest + +from populace.build.us_runtime.release_input_coverage import ( + load_release_input_coverage_manifest, +) +from populace.build.us_runtime.take_up_contract import ( + load_take_up_contract, + seeded_take_up_programs, +) + +ROOT = Path(__file__).resolve().parents[3] +policyengine_us_installed = importlib.util.find_spec("policyengine_us") is not None +requires_us = pytest.mark.skipif( + not policyengine_us_installed, + reason="requires the policyengine-us [us] extra (build environment)", +) + + +def _entry() -> dict[str, object]: + payload = json.loads( + files("populace.build.us").joinpath("ecps_parity_known_gaps.json").read_text() + ) + return payload["known_gaps"]["takes_up_dc_ptc"] + + +def _sha256(path: Path) -> str: + digest = sha256() + with path.open("rb") as stream: + for chunk in iter(lambda: stream.read(1024 * 1024), b""): + digest.update(chunk) + return digest.hexdigest() + + +def test_exclusion_pins_archived_random_derivation_and_model_relative_rate() -> None: + entry = _entry() + evidence = entry["evidence"] + + assert entry["reason"].startswith("SOURCE UNAVAILABILITY WITH EVIDENCE:") + assert evidence["classification"] == "source_unavailability" + assert evidence["retired_derivation"] == { + "repository_owner": "PolicyEngine", + "repository_name_parts": ["policyengine-", "us-data"], + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": [ + "policyengine_", + "us_data", + "datasets", + "cps", + "cps.py", + ], + "lines": "565-599", + "method": "seeded Bernoulli draw below a scalar dc_ptc_rate", + } + assert evidence["retired_rate_parameter"] == { + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": [ + "policyengine_", + "us_data", + "parameters", + "take_up", + "dc_ptc.yaml", + ], + "lines": "1-11", + "value": 0.32, + "administrative_claim_count": 37133, + "comment_model_estimate": 131791388, + "denominator_description": "PolicyEngine DC PTC value estimate", + "external_eligible_population_denominator": False, + "dimensionally_consistent_participation_rate": False, + "arithmetically_consistent_with_value": False, + } + assert evidence["retired_randomness"]["lines"] == "5-28" + assert evidence["retired_randomness"]["key"] == "takes_up_dc_ptc" + assert evidence["missing_claim_column_patterns"] == [ + "takes_up_dc_ptc", + "dc_ptc", + "schedule_h", + "property_tax_credit", + ] + assert evidence["archived_asec_tax_columns"]["lines"] == "251-262" + assert evidence["archived_puf_credit_mapping"]["lines"] == "704-719" + assert evidence["archived_puf_credit_mapping"]["other_credits_source"] == ("P08000") + assert "synthesize" in evidence["semantic_non_substitutes"]["rejection"] + + +def test_current_take_up_contract_rejects_the_model_relative_rate() -> None: + evidence = _entry()["evidence"] + program = load_take_up_contract().program_map()["takes_up_dc_ptc"] + + assert evidence["current_take_up_contract"]["treatment"] == "rate_unsourced" + assert evidence["current_take_up_contract"]["rate_status"] == "model_relative" + assert program.populace_treatment == "rate_unsourced" + assert program.rate["status"] == "model_relative" + assert "policyengine" in str(program.rate["detail"]).lower() + assert "takes_up_dc_ptc" not in { + seeded.variable for seeded in seeded_take_up_programs() + } + + +def test_exclusion_pins_all_sha_locked_hermetic_inputs() -> None: + evidence = _entry()["evidence"] + build_summary = json.loads( + (ROOT / "experiments/build_j_recert/base_j.summary.json").read_text() + ) + recorded_asec_hashes = { + Path(item["path"]).name: item["sha256"] + for item in build_summary["base_source"]["sources"] + } + evidence_asec_hashes = { + item["filename"]: item["sha256"] for item in evidence["hermetic_asec_inputs"] + } + + assert evidence_asec_hashes == recorded_asec_hashes + assert evidence["processed_puf"]["filename"] == Path(build_summary["puf_h5"]).name + assert evidence["processed_puf"]["sha256"] == build_summary["puf_sha256"] + for item in evidence["hermetic_asec_inputs"]: + assert item["missing_columns"] == ["takes_up_dc_ptc"] + assert item["present_generic_tax_amount_columns"] == [ + "STATETAX_A", + "STATETAX_B", + "SPM_STTAX", + ] + assert evidence["processed_puf"]["missing_arrays"] == [ + "takes_up_dc_ptc", + "state_fips", + "state_code", + "state_code_str", + ] + dictionary = evidence["official_census_variable_dictionary"] + assert dictionary["STATETAX_A"] == "State income tax liability, after credits" + assert dictionary["STATETAX_B"] == "State income tax liability, before credits" + assert dictionary["SPM_STTAX"] == "SPM unit's state tax" + + +@requires_us +def test_sha_locked_artifact_schemas_match_the_recorded_absence() -> None: + from policyengine_us.data import USSingleYearDataset + + evidence = _entry()["evidence"] + build_summary = json.loads( + (ROOT / "experiments/build_j_recert/base_j.summary.json").read_text() + ) + paths = { + Path(item["path"]).name: Path(item["path"]) + for item in build_summary["base_source"]["sources"] + } + puf_path = Path(build_summary["puf_h5"]) + if not all(path.is_file() for path in [*paths.values(), puf_path]): + pytest.skip("SHA-locked ASEC/PUF artifacts are not mounted") + + entity_names = ("person", "tax_unit", "spm_unit", "family", "household") + for item in evidence["hermetic_asec_inputs"]: + path = paths[item["filename"]] + assert _sha256(path) == item["sha256"] + dataset = USSingleYearDataset(file_path=str(path)) + all_columns = { + str(column) + for entity in entity_names + for column in getattr(dataset, entity).columns + } + assert set(item["missing_columns"]).isdisjoint(all_columns) + assert set(item["present_generic_tax_amount_columns"]) <= all_columns + lower_columns = {column.lower() for column in all_columns} + assert not any( + pattern in column + for pattern in evidence["missing_claim_column_patterns"] + for column in lower_columns + ) + + assert _sha256(puf_path) == evidence["processed_puf"]["sha256"] + with h5py.File(puf_path, mode="r") as puf: + assert set(evidence["processed_puf"]["missing_arrays"]).isdisjoint(puf.keys()) + assert "other_credits" in puf + lower_arrays = {str(name).lower() for name in puf.keys()} + assert not any( + pattern in name + for pattern in evidence["missing_claim_column_patterns"] + for name in lower_arrays + ) + + +@requires_us +def test_policyengine_1_764_6_requires_a_tax_unit_year_boolean() -> None: + from policyengine_us import CountryTaxBenefitSystem + + assert version("policyengine-us") == "1.764.6" + variable = CountryTaxBenefitSystem().variables["takes_up_dc_ptc"] + assert variable.is_input_variable() + assert variable.entity.key == "tax_unit" + assert str(variable.definition_period).lower() == "year" + assert variable.value_type is bool + assert bool(variable.default_value) is True + + +def test_generated_release_manifest_preserves_the_evidenced_exclusion() -> None: + reason = load_release_input_coverage_manifest().reviewed_exclusions[ + "takes_up_dc_ptc" + ] + assert reason == _entry()["reason"] + assert reason.startswith("SOURCE UNAVAILABILITY WITH EVIDENCE:") diff --git a/packages/populace-build/tests/test_us_disability_benefits.py b/packages/populace-build/tests/test_us_disability_benefits.py new file mode 100644 index 00000000..2eb7d320 --- /dev/null +++ b/packages/populace-build/tests/test_us_disability_benefits.py @@ -0,0 +1,547 @@ +"""ASEC disability-benefit restoration and PUF-half QRF treatment.""" + +from __future__ import annotations + +import importlib.util +from importlib.metadata import version + +import numpy as np +import pandas as pd +import pytest + +import populace.build.us_runtime.disability_benefits as module +from populace.build.source_runtime import ( + SourceRuntimeConfig, + SourceRuntimeContext, + SourceRuntimeError, +) +from populace.build.us_runtime.disability_benefits import ( + DISABILITY_BENEFITS_ARCHIVED_DERIVATION_URL, + DISABILITY_BENEFITS_ARCHIVED_PUF_IMPUTATION_URL, + DISABILITY_BENEFITS_ARCHIVED_PUF_OUTPUTS_URL, + DISABILITY_BENEFITS_ARCHIVED_SOURCE_COLUMNS_URL, + US_DISABILITY_BENEFITS_OUTPUT_COLUMNS, + US_DISABILITY_BENEFITS_REQUIRED_SOURCE_COLUMNS, + derive_us_disability_benefits_from_manifest, + impute_us_disability_benefits_to_puf_support_from_manifest, + us_disability_benefits_signal_gate, + us_disability_benefits_stage_spec, + with_us_disability_benefits, +) +from populace.build.us_runtime.puf_support import clone_us_frame_for_puf_support +from populace.build.us_runtime.release_input_coverage import ( + us_release_reform_coverage_probes, +) +from populace.build.us_runtime.source_runtime import us_source_operation_handlers +from populace.frame import US_SCHEMA, Frame, WeightKind, Weights + +policyengine_us_installed = importlib.util.find_spec("policyengine_us") is not None +requires_us = pytest.mark.skipif( + not policyengine_us_installed, + reason="requires the policyengine-us [us] extra (build environment)", +) + +_OUTPUT = US_DISABILITY_BENEFITS_OUTPUT_COLUMNS[0] +_PREDICTORS = ( + "age", + "is_male", + "has_esi", + "tax_unit_is_joint", + "tax_unit_count_dependents", + "employment_income", + "self_employment_income", + "social_security", +) + + +def _person_source() -> pd.DataFrame: + count = 100 + first_amount = np.zeros(count) + first_code = np.zeros(count) + second_amount = np.zeros(count) + second_code = np.zeros(count) + first_amount[0] = 3_600.0 + first_code[0] = 2.0 + # A workers'-compensation amount in the other slot must remain excluded. + second_amount[0] = 1_200.0 + second_code[0] = 1.0 + return pd.DataFrame( + { + "person_id": np.arange(1, count + 1, dtype="int64"), + "person_household_id": np.arange(1, count + 1, dtype="int64") * 10, + "person_tax_unit_id": np.arange(1, count + 1, dtype="int64") * 100, + "person_spm_unit_id": np.arange(1, count + 1, dtype="int64") * 1_000, + "person_family_id": np.arange(1, count + 1, dtype="int64") * 10_000, + "person_marital_unit_id": ( + np.arange(1, count + 1, dtype="int64") * 100_000 + ), + "DIS_VAL1": first_amount, + "DIS_SC1": first_code, + "DIS_VAL2": second_amount, + "DIS_SC2": second_code, + "WSAL_VAL": np.linspace(0.0, 99_000.0, count), + "SEMP_VAL": np.zeros(count), + "employment_income_before_lsr": np.linspace(0.0, 99_000.0, count), + "self_employment_income_before_lsr": np.zeros(count), + "age": np.arange(20, 20 + count), + "is_female": np.tile([False, True], count // 2), + "has_esi": np.tile([True, False], count // 2), + "tax_unit_role_input": ["HEAD"] * count, + "social_security_retirement": np.zeros(count), + "social_security_disability": np.zeros(count), + "social_security_dependents": np.zeros(count), + "social_security_survivors": np.zeros(count), + } + ) + + +def _frame() -> Frame: + person = _person_source() + count = len(person) + ids = { + "household": person["person_household_id"].to_numpy(), + "tax_unit": person["person_tax_unit_id"].to_numpy(), + "spm_unit": person["person_spm_unit_id"].to_numpy(), + "family": person["person_family_id"].to_numpy(), + "marital_unit": person["person_marital_unit_id"].to_numpy(), + } + tables = { + entity: pd.DataFrame({f"{entity}_id": values}) for entity, values in ids.items() + } + tables["person"] = person + tables["tax_unit"]["filing_status_input"] = ["SINGLE"] * count + return Frame( + tables, + US_SCHEMA, + { + "household": Weights( + np.ones(count, dtype=np.float64), + WeightKind.DESIGN, + ) + }, + ) + + +def _derive(person: pd.DataFrame) -> pd.DataFrame: + operation = next( + operation + for operation in us_disability_benefits_stage_spec().operations + if operation.kind == "derive_disability_benefits" + ) + return derive_us_disability_benefits_from_manifest(person, operation, None) + + +def test_archived_sources_are_sha_and_line_pinned_and_sources_available() -> None: + commit = "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe" + urls = ( + DISABILITY_BENEFITS_ARCHIVED_DERIVATION_URL, + DISABILITY_BENEFITS_ARCHIVED_SOURCE_COLUMNS_URL, + DISABILITY_BENEFITS_ARCHIVED_PUF_OUTPUTS_URL, + DISABILITY_BENEFITS_ARCHIVED_PUF_IMPUTATION_URL, + ) + assert all(commit in url for url in urls) + assert DISABILITY_BENEFITS_ARCHIVED_DERIVATION_URL.endswith( + "datasets/cps/cps.py#L1561-L1571" + ) + assert DISABILITY_BENEFITS_ARCHIVED_SOURCE_COLUMNS_URL.endswith( + "datasets/cps/census_cps.py#L306-L381" + ) + assert DISABILITY_BENEFITS_ARCHIVED_PUF_OUTPUTS_URL.endswith( + "datasets/cps/extended_cps.py#L135-L194" + ) + assert DISABILITY_BENEFITS_ARCHIVED_PUF_IMPUTATION_URL.endswith( + "datasets/cps/extended_cps.py#L639-L745" + ) + assert US_DISABILITY_BENEFITS_REQUIRED_SOURCE_COLUMNS == ( + "DIS_VAL1", + "DIS_SC1", + "DIS_VAL2", + "DIS_SC2", + ) + + +def test_stage_manifest_pins_two_slot_formula_and_one_output_qrf() -> None: + spec = us_disability_benefits_stage_spec() + + assert spec.stage == "disability_benefits_input" + assert spec.survey == "Census CPS ASEC" + assert spec.grain == "person" + assert tuple(spec.outputs) == (_OUTPUT,) + assert tuple(spec.nonnegative_outputs) == (_OUTPUT,) + assert [operation.kind for operation in spec.operations] == [ + "read_table", + "derive_disability_benefits", + "impute_disability_benefits_to_puf_support", + ] + assert spec.operations[0].parameters == { + "table": "person", + "weight": "person_weight", + } + assert spec.operations[1].parameters == { + "first_amount_source": "DIS_VAL1", + "first_code_source": "DIS_SC1", + "second_amount_source": "DIS_VAL2", + "second_code_source": "DIS_SC2", + "workers_compensation_code": 1, + "output": _OUTPUT, + } + assert spec.operations[2].parameters == { + "predictors": list(_PREDICTORS), + "max_train_samples": 5_000, + "n_estimators": 100, + "seed_from_build_config": True, + "weight": "person_weight", + } + + handlers = us_source_operation_handlers() + assert ( + handlers["derive_disability_benefits"] + is derive_us_disability_benefits_from_manifest + ) + assert ( + handlers["impute_disability_benefits_to_puf_support"] + is impute_us_disability_benefits_to_puf_support_from_manifest + ) + + +def test_direct_formula_excludes_only_code_one_and_preserves_topcodes() -> None: + source = pd.DataFrame( + { + "DIS_VAL1": [100.0, 99_999.0, 500.0, 800.0, 300.0], + "DIS_SC1": [1, 2, 1, 2, 0], + "DIS_VAL2": [200.0, 90_000.0, 700.0, 900.0, 400.0], + "DIS_SC2": [1, 0, 2, 1, 3], + } + ) + original = source.copy(deep=True) + + result = _derive(source) + + assert result[_OUTPUT].tolist() == [0.0, 189_999.0, 700.0, 800.0, 700.0] + pd.testing.assert_frame_equal(source, original) + + +@pytest.mark.parametrize("missing", US_DISABILITY_BENEFITS_REQUIRED_SOURCE_COLUMNS) +def test_direct_formula_fails_closed_when_source_is_missing(missing: str) -> None: + with pytest.raises(SourceRuntimeError, match=missing): + _derive(_person_source().drop(columns=[missing])) + + +@pytest.mark.parametrize( + ("column", "bad_value", "message"), + [ + ("DIS_VAL1", np.nan, "nonfinite"), + ("DIS_SC1", np.inf, "nonfinite"), + ("DIS_VAL2", np.inf, "nonfinite"), + ("DIS_SC2", np.nan, "nonfinite"), + ("DIS_VAL1", -1.0, "negative"), + ("DIS_VAL2", -1.0, "negative"), + ], +) +def test_direct_formula_rejects_invalid_sources( + column: str, + bad_value: float, + message: str, +) -> None: + person = _person_source() + person.loc[0, column] = bad_value + + with pytest.raises(SourceRuntimeError, match=message): + _derive(person) + + +def test_with_inputs_materializes_exact_asec_values_without_mutation() -> None: + frame = _frame() + original = frame.table("person").copy(deep=True) + + result = with_us_disability_benefits(frame, seed=0, time_period=2024) + + pd.testing.assert_frame_equal(frame.table("person"), original) + expected = np.where(original["DIS_SC1"] != 1, original["DIS_VAL1"], 0) + np.where( + original["DIS_SC2"] != 1, + original["DIS_VAL2"], + 0, + ) + np.testing.assert_allclose(result.table("person")[_OUTPUT], expected) + gate = us_disability_benefits_signal_gate(result) + assert gate.passed, gate.failures + + +def test_release_requires_opt_in_to_preserve_valid_existing_surface() -> None: + materialized = with_us_disability_benefits(_frame(), seed=0, time_period=2024) + tables = { + entity: materialized.table(entity).copy() for entity in materialized.entities + } + tables["person"] = tables["person"].drop( + columns=list(US_DISABILITY_BENEFITS_REQUIRED_SOURCE_COLUMNS) + ) + release = Frame( + tables, + materialized.schema, + { + entity: materialized.weights_for(entity) + for entity in materialized.weighted_entities + }, + materialized.strata, + mass_log=materialized.mass_log, + ) + + with pytest.raises(ValueError, match="cannot heal.*without measured"): + with_us_disability_benefits(release, seed=0, time_period=2024) + + assert ( + with_us_disability_benefits( + release, + seed=0, + time_period=2024, + allow_existing_without_source=True, + ) + is release + ) + + +def test_puf_half_uses_one_output_qrf_and_preserves_asec( + monkeypatch: pytest.MonkeyPatch, +) -> None: + direct = with_us_disability_benefits(_frame(), seed=0, time_period=2024) + expanded = clone_us_frame_for_puf_support(direct) + original = expanded.table("person").copy(deep=True) + calls: dict[str, object] = {} + + class FakeFitted: + def predict(self, test: pd.DataFrame) -> pd.DataFrame: + calls["test"] = test.copy() + predicted = np.zeros(len(test), dtype=np.float64) + predicted[0] = 7_200.0 + return pd.DataFrame({_OUTPUT: predicted}, index=test.index) + + class FakeQRF: + def __init__(self, **kwargs: object) -> None: + calls["init"] = kwargs + + def fit( + self, + training: pd.DataFrame, + predictors: list[str], + targets: list[str], + *, + weights: np.ndarray, + ) -> FakeFitted: + calls["fit_count"] = int(calls.get("fit_count", 0)) + 1 + calls["training"] = training.copy() + calls["predictors"] = predictors + calls["targets"] = targets + calls["weights"] = weights.copy() + return FakeFitted() + + monkeypatch.setattr(module, "QRF", FakeQRF) + + result = with_us_disability_benefits(expanded, seed=7, time_period=2024) + + assert calls["fit_count"] == 1 + assert calls["init"] == {"n_estimators": 100, "seed": 7} + assert calls["predictors"] == list(_PREDICTORS) + assert calls["targets"] == [_OUTPUT] + training = calls["training"] + assert isinstance(training, pd.DataFrame) + assert list(training.columns) == [*_PREDICTORS, _OUTPUT] + asec_mask = original["person_support_channel"] == "asec" + expected_weights = expanded.resolve_weights("person").values[asec_mask] + np.testing.assert_allclose(calls["weights"], expected_weights) + + person = result.table("person") + asec = person[person["person_support_channel"] == "asec"] + puf = person[person["person_support_channel"] == "puf_tax_detail"] + assert asec[_OUTPUT].tolist() == [3_600.0, *([0.0] * 99)] + assert puf[_OUTPUT].tolist() == [7_200.0, *([0.0] * 99)] + pd.testing.assert_frame_equal(expanded.table("person"), original) + gate = us_disability_benefits_signal_gate(result) + assert gate.passed, gate.failures + + +def test_puf_qrf_caps_training_at_5000_and_keeps_weights_aligned( + monkeypatch: pytest.MonkeyPatch, +) -> None: + operation = us_disability_benefits_stage_spec().operations[2] + asec_rows = 5_010 + puf_rows = 2 + rows = asec_rows + puf_rows + frame = pd.DataFrame( + { + "person_support_channel": ["asec"] * asec_rows + + ["puf_tax_detail"] * puf_rows, + "person_weight": np.arange(1.0, rows + 1.0), + _OUTPUT: np.tile([0.0, 500.0], (rows + 1) // 2)[:rows], + **{ + f"disability_benefits_predictor_{predictor}": np.arange( + rows, dtype=np.float64 + ) + for predictor in _PREDICTORS + }, + } + ) + calls: dict[str, object] = {} + + class FakeFitted: + def predict(self, test: pd.DataFrame) -> pd.DataFrame: + calls["test_rows"] = len(test) + return pd.DataFrame({_OUTPUT: np.zeros(len(test))}, index=test.index) + + class FakeQRF: + def __init__(self, **kwargs: object) -> None: + calls["init"] = kwargs + + def fit( + self, + training: pd.DataFrame, + predictors: list[str], + targets: list[str], + *, + weights: np.ndarray, + ) -> FakeFitted: + calls["training_rows"] = len(training) + calls["training_index"] = training.index.to_numpy() + calls["weights"] = weights.copy() + return FakeFitted() + + monkeypatch.setattr(module, "QRF", FakeQRF) + context = SourceRuntimeContext( + config=SourceRuntimeConfig(seed=19, target_year=2024), + tables={}, + ) + + impute_us_disability_benefits_to_puf_support_from_manifest( + frame, + operation, + context, + ) + + assert calls["init"] == {"n_estimators": 100, "seed": 19} + assert calls["training_rows"] == 5_000 + assert calls["test_rows"] == 2 + np.testing.assert_allclose( + calls["weights"], + frame.loc[calls["training_index"], "person_weight"].to_numpy(), + ) + + +def test_signal_gate_rejects_missing_default_and_invalid_surfaces() -> None: + valid = with_us_disability_benefits(_frame(), seed=0, time_period=2024) + assert us_disability_benefits_signal_gate(valid).passed + + candidates: list[Frame] = [] + for replacement in ( + None, + np.zeros(100), + np.asarray([-1.0, *([0.0] * 99)]), + np.asarray([np.nan, *([0.0] * 99)]), + ): + candidate = with_us_disability_benefits(_frame(), seed=0, time_period=2024) + if replacement is None: + candidate.table("person").drop(columns=[_OUTPUT], inplace=True) + else: + candidate.table("person")[_OUTPUT] = replacement + candidates.append(candidate) + + assert all( + not us_disability_benefits_signal_gate(candidate).passed + for candidate in candidates + ) + + +@pytest.mark.parametrize("dead_channel", ["asec", "puf_tax_detail"]) +def test_signal_gate_rejects_either_dead_support_channel(dead_channel: str) -> None: + direct = with_us_disability_benefits(_frame(), seed=0, time_period=2024) + expanded = clone_us_frame_for_puf_support(direct) + assert us_disability_benefits_signal_gate(expanded).passed + + channel = expanded.table("person")["person_support_channel"] + expanded.table("person").loc[channel == dead_channel, _OUTPUT] = 0.0 + gate = us_disability_benefits_signal_gate(expanded) + + assert not gate.passed + assert any(dead_channel in failure for failure in gate.failures) + + +@requires_us +def test_policyengine_us_1_764_6_contract_and_positive_annual_behavior() -> None: + from policyengine_us import CountryTaxBenefitSystem, Simulation + + assert version("policyengine-us") == "1.764.6" + variable = CountryTaxBenefitSystem().variables[_OUTPUT] + assert variable.is_input_variable() + assert variable.entity.key == "person" + assert str(variable.definition_period).lower() == "year" + assert variable.default_value == 0 + + situation = { + "people": { + "adult": { + "age": {"2024": 40}, + _OUTPUT: {"2024": 6_000.0}, + } + }, + "tax_units": { + "tax_unit": { + "members": ["adult"], + "filing_status": {"2024": "SINGLE"}, + } + }, + "spm_units": {"spm_unit": {"members": ["adult"]}}, + "households": { + "household": { + "members": ["adult"], + "state_code": {"2024": "CA"}, + } + }, + } + simulation = Simulation(situation=situation) + + assert simulation.calculate(_OUTPUT, 2024)[0] == pytest.approx(6_000.0) + assert simulation.calculate(_OUTPUT, "2024-01")[0] == pytest.approx(500.0) + assert simulation.calculate("snap_unearned_income", "2024-01")[0] == pytest.approx( + 500.0 + ) + + +@requires_us +def test_shipped_snap_exclusion_probe_binds_with_positive_sign() -> None: + from policyengine_core.reforms import Reform + from policyengine_us import CountryTaxBenefitSystem, Simulation + + probe = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "disability_benefits_snap_exclusion" + ) + reform = Reform.from_dict(dict(probe.parameter_changes), country_id="us") + situation = { + "people": { + "adult": { + "age": {"2024": 40}, + "employment_income": {"2024": 12_000.0}, + _OUTPUT: {"2024": 6_000.0}, + } + }, + "tax_units": { + "tax_unit": { + "members": ["adult"], + "filing_status": {"2024": "SINGLE"}, + } + }, + "spm_units": {"spm_unit": {"members": ["adult"]}}, + "households": { + "household": { + "members": ["adult"], + "state_code": {"2024": "CA"}, + } + }, + } + baseline = Simulation(situation=situation) + reformed = Simulation( + tax_benefit_system=CountryTaxBenefitSystem(reform=(reform,)), + situation=situation, + ) + + effect = reformed.calculate("snap", 2024)[0] - baseline.calculate("snap", 2024)[0] + assert effect > 1_000.0 diff --git a/packages/populace-build/tests/test_us_domestic_production.py b/packages/populace-build/tests/test_us_domestic_production.py new file mode 100644 index 00000000..119dc078 --- /dev/null +++ b/packages/populace-build/tests/test_us_domestic_production.py @@ -0,0 +1,291 @@ +"""Contracts for the IRS PUF domestic-production-ALD restoration.""" + +from __future__ import annotations + +import importlib.util + +import numpy as np +import pandas as pd +import pytest + +from populace.build.us_runtime import puf_support as puf_support_module +from populace.build.us_runtime.domestic_production import ( + DOMESTIC_PRODUCTION_ALD_ARCHIVED_DERIVATION_URL, + DOMESTIC_PRODUCTION_ALD_ARCHIVED_EXPORT_URL, + DOMESTIC_PRODUCTION_ALD_ARCHIVED_IMPUTATION_URL, + DOMESTIC_PRODUCTION_ALD_ARCHIVED_PUF_ARTIFACT_URL, + derive_us_domestic_production_ald_from_puf, + us_domestic_production_ald_signal_gate, + us_domestic_production_ald_stage_spec, +) +from populace.build.us_runtime.puf_aggregate_records import ( + _reconcile_puf_domestic_production_ald_from_source, +) +from populace.build.us_runtime.release_input_coverage import ( + us_release_reform_coverage_probes, +) + +policyengine_us_installed = importlib.util.find_spec("policyengine_us") is not None +requires_us = pytest.mark.skipif( + not policyengine_us_installed, + reason="requires the policyengine-us [us] extra (build environment)", +) + + +class _ResolvedWeights: + def __init__(self, values: np.ndarray) -> None: + self.values = values + + +class _TaxUnitFrame: + def __init__( + self, + tax_unit: pd.DataFrame, + weights: np.ndarray | None = None, + ) -> None: + self._tax_unit = tax_unit + self._weights = np.ones(len(tax_unit)) if weights is None else weights + + def table(self, entity: str) -> pd.DataFrame: + assert entity == "tax_unit" + return self._tax_unit + + def resolve_weights(self, entity: str) -> _ResolvedWeights: + assert entity == "tax_unit" + return _ResolvedWeights(np.asarray(self._weights, dtype=np.float64)) + + +def test_archived_coordinates_are_immutable_and_cover_the_full_source_path() -> None: + commit = "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe" + for url in ( + DOMESTIC_PRODUCTION_ALD_ARCHIVED_DERIVATION_URL, + DOMESTIC_PRODUCTION_ALD_ARCHIVED_EXPORT_URL, + DOMESTIC_PRODUCTION_ALD_ARCHIVED_IMPUTATION_URL, + DOMESTIC_PRODUCTION_ALD_ARCHIVED_PUF_ARTIFACT_URL, + ): + assert commit in url + assert DOMESTIC_PRODUCTION_ALD_ARCHIVED_DERIVATION_URL.endswith( + "datasets/puf/puf.py#L646" + ) + assert DOMESTIC_PRODUCTION_ALD_ARCHIVED_EXPORT_URL.endswith( + "datasets/puf/puf.py#L808-L815" + ) + assert DOMESTIC_PRODUCTION_ALD_ARCHIVED_IMPUTATION_URL.endswith( + "calibration/puf_impute.py#L90-L198" + ) + assert DOMESTIC_PRODUCTION_ALD_ARCHIVED_PUF_ARTIFACT_URL.endswith( + "datasets/puf/puf.py#L1655-L1660" + ) + + +def test_archived_e03240_mapping_is_an_exact_carry() -> None: + source = pd.DataFrame({"E03240": [0.0, 125.5, 9_000.0]}) + + result = derive_us_domestic_production_ald_from_puf(source) + + assert "domestic_production_ald" not in source.columns + assert result["domestic_production_ald"].tolist() == [0.0, 125.5, 9_000.0] + + +@pytest.mark.parametrize( + ("source", "message"), + [ + (pd.DataFrame({"other": [1.0]}), "requires source column"), + (pd.DataFrame({"E03240": ["not numeric"]}), "nonnumeric or nonfinite"), + (pd.DataFrame({"E03240": [np.inf]}), "nonnumeric or nonfinite"), + (pd.DataFrame({"E03240": [-1.0]}), "negative value"), + ], +) +def test_source_derivation_fails_closed( + source: pd.DataFrame, + message: str, +) -> None: + with pytest.raises(ValueError, match=message): + derive_us_domestic_production_ald_from_puf(source) + + +def test_post_disaggregation_reconciliation_uses_final_e03240() -> None: + source = pd.DataFrame( + { + "E03240": [0.0, 3_000.0, 7_500.0], + "domestic_production_ald": [999.0, 999.0, 999.0], + } + ) + + result = _reconcile_puf_domestic_production_ald_from_source(source) + + assert result["domestic_production_ald"].tolist() == [0.0, 3_000.0, 7_500.0] + + +def test_shared_puf_stage_declares_exact_source_and_output() -> None: + stage = us_domestic_production_ald_stage_spec() + operation = next( + operation + for operation in stage.operations + if operation.kind == "derive_puf_policyengine_variables" + ) + + assert operation.parameters["domestic_production_ald_source"] == "E03240" + assert ( + operation.parameters["domestic_production_ald_output"] + == "domestic_production_ald" + ) + assert "domestic_production_ald" in stage.outputs + assert "domestic_production_ald" in stage.nonnegative_outputs + assert any("E03240" in str(artifact.get("locator")) for artifact in stage.artifacts) + + +def test_puf_support_keeps_input_at_tax_unit_grain_and_sparse() -> None: + output = "domestic_production_ald" + assert output in puf_support_module.PUF_TAX_DETAIL_DEFAULT_TAX_UNIT_OUTPUTS + assert output not in puf_support_module.PUF_TAX_DETAIL_DEFAULT_PERSON_OUTPUTS + assert output in puf_support_module._PUF_TAX_DETAIL_NONNEGATIVE_OUTPUTS + assert output in puf_support_module._PUF_TAX_DETAIL_SPARSE_TAX_UNIT_OUTPUTS + assert output not in puf_support_module._PUF_TAX_DETAIL_DISCRETE_TAX_UNIT_OUTPUTS + + +def test_tax_unit_sparsifier_preserves_weighted_total_and_source_channel() -> None: + tables = { + "household": pd.DataFrame({"household_id": [1, 2, 3, 4, 5]}), + "tax_unit": pd.DataFrame( + { + "tax_unit_id": [10, 20, 30, 40, 50], + "tax_unit_support_channel": [ + "puf_tax_detail", + "puf_tax_detail", + "puf_tax_detail", + "puf_tax_detail", + "asec", + ], + "domestic_production_ald": [100.0, 200.0, 300.0, 400.0, 0.0], + } + ), + "person": pd.DataFrame( + { + "person_tax_unit_id": [10, 20, 30, 40, 50], + "person_household_id": [1, 2, 3, 4, 5], + } + ), + } + + puf_support_module._sparsify_tax_unit_output_to_donor_positive_rate( + tables, + column="domestic_production_ald", + donor_positive_rate=0.25, + household_weights=np.ones(5), + tax_unit_channel="tax_unit_support_channel", + ) + + values = tables["tax_unit"]["domestic_production_ald"] + assert int((values.iloc[:4] > 0.0).sum()) == 1 + assert values.iloc[:4].sum() == pytest.approx(1_000.0) + assert values.iloc[4] == 0.0 + + +def test_signal_gate_accepts_plausibly_sparse_nondefault_values() -> None: + values = np.zeros(1_000) + values[[10, 20, 30]] = [1_000.0, 2_000.0, 3_000.0] + frame = _TaxUnitFrame(pd.DataFrame({"domestic_production_ald": values})) + + result = us_domestic_production_ald_signal_gate(frame) # type: ignore[arg-type] + + assert result.passed, result.failures + assert result.details["positive_share"] == pytest.approx(0.003) + + +def test_signal_gate_rejects_positive_values_on_the_asec_support_channel() -> None: + values = np.zeros(1_000) + values[[10, 20, 30]] = [1_000.0, 2_000.0, 3_000.0] + channels = np.full(1_000, "puf_tax_detail", dtype=object) + channels[10] = "asec" + frame = _TaxUnitFrame( + pd.DataFrame( + { + "domestic_production_ald": values, + "tax_unit_support_channel": channels, + } + ) + ) + + result = us_domestic_production_ald_signal_gate(frame) # type: ignore[arg-type] + + assert not result.passed + assert any("ASEC support channel" in failure for failure in result.failures) + + +@pytest.mark.parametrize( + "tax_unit", + [ + pd.DataFrame({"other": [0.0, 1.0]}), + pd.DataFrame({"domestic_production_ald": [0.0, 0.0]}), + pd.DataFrame({"domestic_production_ald": [0.0, -1.0]}), + pd.DataFrame({"domestic_production_ald": [0.0, np.nan]}), + ], +) +def test_signal_gate_rejects_missing_default_or_invalid_surface( + tax_unit: pd.DataFrame, +) -> None: + result = us_domestic_production_ald_signal_gate( # type: ignore[arg-type] + _TaxUnitFrame(tax_unit) + ) + + assert not result.passed + + +@requires_us +def test_policyengine_us_contract_is_a_tax_unit_year_input_leaf() -> None: + from policyengine_us import CountryTaxBenefitSystem + + variable = CountryTaxBenefitSystem().variables["domestic_production_ald"] + + assert variable.is_input_variable() + assert variable.entity.key == "tax_unit" + assert str(variable.definition_period).lower() == "year" + assert variable.default_value == 0 + + +@requires_us +def test_2024_reactivation_probe_binds_only_when_the_input_is_populated() -> None: + from policyengine_core.reforms import Reform + from policyengine_us import CountryTaxBenefitSystem, Simulation + + probe = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "domestic_production_ald_reactivation" + ) + reform = Reform.from_dict(dict(probe.parameter_changes), country_id="us") + situation = { + "people": { + "adult": { + "age": {"2024": 40}, + "employment_income": {"2024": 100_000}, + } + }, + "tax_units": { + "tax_unit": { + "members": ["adult"], + "filing_status": {"2024": "SINGLE"}, + "domestic_production_ald": {"2024": 10_000}, + } + }, + "households": { + "household": { + "members": ["adult"], + "state_code": {"2024": "CA"}, + } + }, + } + baseline = Simulation(situation=situation) + reformed = Simulation( + tax_benefit_system=CountryTaxBenefitSystem(reform=(reform,)), + situation=situation, + ) + + assert baseline.calculate("above_the_line_deductions", 2024)[0] == 0.0 + assert reformed.calculate("above_the_line_deductions", 2024)[0] == 10_000.0 + effect = ( + baseline.calculate("income_tax", 2024)[0] + - reformed.calculate("income_tax", 2024)[0] + ) + assert effect > 1_000.0 diff --git a/packages/populace-build/tests/test_us_early_head_start_exclusion.py b/packages/populace-build/tests/test_us_early_head_start_exclusion.py new file mode 100644 index 00000000..03b773e4 --- /dev/null +++ b/packages/populace-build/tests/test_us_early_head_start_exclusion.py @@ -0,0 +1,415 @@ +"""Evidence contract for irreducible Early Head Start source unavailability.""" + +from __future__ import annotations + +import importlib.util +import inspect +import json +import os +from hashlib import sha256 +from importlib.metadata import version +from importlib.resources import files +from pathlib import Path + +import h5py +import pandas as pd +import pytest + +from populace.build.us_runtime.asec_pool import load_asec_h5_tables +from populace.build.us_runtime.release_input_coverage import ( + load_release_input_coverage_manifest, +) +from populace.build.us_runtime.sipp_head_start import ( + SIPP_2023_HEAD_START_DONOR_REVISION, + SIPP_2023_HEAD_START_DONOR_SHA256, + SIPP_2023_HEAD_START_DONOR_SIZE_BYTES, +) +from populace.build.us_runtime.take_up_contract import load_take_up_contract + +ROOT = Path(__file__).resolve().parents[3] +_RUN_LARGE_SOURCE_AUDITS = os.environ.get("POPULACE_RUN_LARGE_SOURCE_AUDITS") == "1" +policyengine_us_installed = importlib.util.find_spec("policyengine_us") is not None +requires_us = pytest.mark.skipif( + not policyengine_us_installed, + reason="requires the policyengine-us [us] extra (build environment)", +) + + +def _entry() -> dict[str, object]: + payload = json.loads( + files("populace.build.us").joinpath("ecps_parity_known_gaps.json").read_text() + ) + return payload["known_gaps"]["takes_up_early_head_start_if_eligible"] + + +def _build_summary() -> dict[str, object]: + return json.loads( + (ROOT / "experiments/build_j_recert/base_j.summary.json").read_text() + ) + + +def _sha256(path: Path) -> str: + digest = sha256() + with path.open("rb") as stream: + for chunk in iter(lambda: stream.read(1024 * 1024), b""): + digest.update(chunk) + return digest.hexdigest() + + +def _resolve_recorded_artifact(recorded_path: str, filename: str) -> Path | None: + """Resolve a Build-J input without silently substituting another artifact.""" + + candidates = ( + Path(recorded_path).expanduser(), + Path.home() + / "PolicyEngine" + / ("policyengine-" + "us-data") + / ("policyengine_" + "us_data") + / "storage" + / filename, + ) + return next((path for path in candidates if path.is_file()), None) + + +def _sipp_snapshot() -> Path: + return ( + Path.home() + / ".cache" + / "huggingface" + / "hub" + / ("models--policyengine--policyengine-" + "us-data") + / "snapshots" + / SIPP_2023_HEAD_START_DONOR_REVISION + / "pu2023.csv" + ) + + +def test_exclusion_pins_exact_archived_scalar_draw_evidence() -> None: + entry = _entry() + evidence = entry["evidence"] + + assert entry["reason"].startswith("SOURCE UNAVAILABILITY WITH EVIDENCE:") + assert evidence["classification"] == "source_unavailability" + assert evidence["retired_derivation"] == { + "repository_owner": "PolicyEngine", + "repository_name_parts": ["policyengine-", "us-data"], + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": [ + "policyengine_", + "us_data", + "datasets", + "cps", + "cps.py", + ], + "lines": "582-583,640-648", + } + assert evidence["retired_rate"] == { + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": [ + "policyengine_", + "us_data", + "parameters", + "take_up", + "early_head_start.yaml", + ], + "lines": "1-9", + "value": 0.09, + "citation_scope": ( + "NIEER research chart; no ACF administrative participation fact" + ), + } + assert evidence["retired_randomness"] == { + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": [ + "policyengine_", + "us_data", + "utils", + "randomness.py", + ], + "lines": "5-28", + "key": "takes_up_early_head_start_if_eligible", + } + assert evidence["required_person_observation"] == ( + "Individual Early Head Start enrollment covering infants, toddlers, " + "and pregnant participants" + ) + assert "synthesize" in entry["reason"] + + +def test_exclusion_pins_exact_locked_artifact_absence_claims() -> None: + evidence = _entry()["evidence"] + missing = [ + "takes_up_early_head_start_if_eligible", + "takes_up_head_start_if_eligible", + "is_enrolled_in_head_start", + ] + assert evidence["hermetic_asec_inputs"] == [ + { + "filename": "census_cps_2022.h5", + "sha256": ( + "7ccca976284bb47815d84460cc4f75a0a65d26d7754ab0a0f417de351b3d474e" + ), + "missing_columns": missing, + }, + { + "filename": "census_cps_2023.h5", + "sha256": ( + "cb57817327799f42b741caed5f9be94d04021c2e6809c1ad7bd0686da5428d88" + ), + "missing_columns": missing, + }, + { + "filename": "census_cps_2024.h5", + "sha256": ( + "ec36604cb735a660b51b0b2f90be27d803b5878f3464fb30d0eacead59c1260d" + ), + "missing_columns": missing, + }, + ] + assert evidence["processed_puf"] == { + "filename": "puf_2024.h5", + "release": "policyengine/irs-soi-puf/1.8.0", + "sha256": ("7669f5b5281f20080e77204f9bd4aabfad0aa101fa283e22caf9ba8d61d4d6df"), + "missing_columns": missing, + } + + summary = _build_summary() + recorded_asec_hashes = { + Path(item["path"]).name: item["sha256"] + for item in summary["base_source"]["sources"] + } + asserted_asec_hashes = { + item["filename"]: item["sha256"] for item in evidence["hermetic_asec_inputs"] + } + assert asserted_asec_hashes == recorded_asec_hashes + assert evidence["processed_puf"]["filename"] == Path(summary["puf_h5"]).name + assert evidence["processed_puf"]["sha256"] == summary["puf_sha256"] + + +@pytest.mark.parametrize( + "filename", + ["census_cps_2022.h5", "census_cps_2023.h5", "census_cps_2024.h5"], +) +def test_mounted_sha_locked_asec_schema_has_no_person_enrollment_signal( + filename: str, +) -> None: + evidence = _entry()["evidence"] + item = next( + record + for record in evidence["hermetic_asec_inputs"] + if record["filename"] == filename + ) + summary_item = next( + record + for record in _build_summary()["base_source"]["sources"] + if Path(record["path"]).name == filename + ) + path = _resolve_recorded_artifact(summary_item["path"], filename) + if path is None: + pytest.skip(f"SHA-locked {filename} artifact is not mounted") + + assert _sha256(path) == item["sha256"] + tables = load_asec_h5_tables(path) + all_columns = {str(column) for table in tables.values() for column in table.columns} + assert set(item["missing_columns"]).isdisjoint(all_columns) + + +def test_mounted_sha_locked_puf_schema_has_no_person_enrollment_signal() -> None: + evidence = _entry()["evidence"] + item = evidence["processed_puf"] + summary = _build_summary() + path = _resolve_recorded_artifact(summary["puf_h5"], item["filename"]) + if path is None: + pytest.skip("SHA-locked processed PUF artifact is not mounted") + + assert _sha256(path) == item["sha256"] + with h5py.File(path, mode="r") as puf: + assert set(item["missing_columns"]).isdisjoint(puf.keys()) + + +def test_pinned_sipp_evidence_has_no_early_head_start_domain_substitute() -> None: + evidence = _entry()["evidence"] + sipp = evidence["pinned_sipp"] + + assert sipp == { + "revision": "21280dca5995e978d706740a8a4b9b7860cfd7b6", + "sha256": ("5c30439e365fc26483318ef61d1d8f4bb2f0e9d6bb47c22c06756a7698733ee2"), + "direct_education_item": ( + "EEDHEADST is nursery/preschool only and has zero reported positives " + "under age 3" + ), + "child_care_scope": ( + "RDAYHS, RHEADST, and RNURHS cover ages 2-7 and do not distinguish " + "Early Head Start" + ), + "pregnancy_signal": "none", + "census_user_note": ( + "https://www.census.gov/programs-surveys/sipp/tech-documentation/" + "user-notes/2023-usernotes/2023-small-inconsist-child-care.html" + ), + } + assert sipp["revision"] == SIPP_2023_HEAD_START_DONOR_REVISION + assert sipp["sha256"] == SIPP_2023_HEAD_START_DONOR_SHA256 + + snapshot = _sipp_snapshot() + if not snapshot.is_file(): + pytest.skip("the independently pinned full SIPP artifact is not mounted") + assert snapshot.stat().st_size == SIPP_2023_HEAD_START_DONOR_SIZE_BYTES + with snapshot.open("r", encoding="utf-8") as stream: + columns = set(stream.readline().rstrip("\n").split("|")) + assert { + "MONTHCODE", + "TAGE", + "EEDHEADST", + "AEDHEADST", + "RDAYHS", + "RHEADST", + "RNURHS", + } <= columns + normalized = {column.upper().replace("_", "") for column in columns} + assert not any("EARLYHEADSTART" in column for column in normalized) + assert not any( + "PREG" in column and ("HEADST" in column or column.endswith("EHS")) + for column in normalized + ) + + +@pytest.mark.skipif( + not _RUN_LARGE_SOURCE_AUDITS, + reason="set POPULACE_RUN_LARGE_SOURCE_AUDITS=1 for the 3.7 GB SIPP scan", +) +def test_opt_in_full_sipp_scan_confirms_early_head_start_domain_absence() -> None: + snapshot = _sipp_snapshot() + if not snapshot.is_file(): + pytest.skip("the independently pinned full SIPP artifact is not mounted") + assert _sha256(snapshot) == SIPP_2023_HEAD_START_DONOR_SHA256 + + direct_positive_under_3 = 0 + observed_child_care_rows: dict[str, list[pd.DataFrame]] = { + "RDAYHS": [], + "RHEADST": [], + "RNURHS": [], + } + allocation_status = { + "RDAYHS": "ARDAYHS", + "RHEADST": "ARHEADST", + "RNURHS": "ARNURHS", + } + for chunk in pd.read_csv( + snapshot, + sep="|", + usecols=[ + "SSUID", + "PNUM", + "MONTHCODE", + "TAGE", + "EEDHEADST", + "AEDHEADST", + *observed_child_care_rows, + *allocation_status.values(), + ], + chunksize=100_000, + low_memory=False, + ): + age = pd.to_numeric(chunk["TAGE"], errors="coerce") + direct_positive_under_3 += int( + ( + age.lt(3) + & pd.to_numeric(chunk["AEDHEADST"], errors="coerce").eq(1) + & pd.to_numeric(chunk["EEDHEADST"], errors="coerce").eq(1) + ).sum() + ) + for column, status_column in allocation_status.items(): + # Annual child-care recodes repeat on person-month rows. Status 1 + # is as reported; reduce to the first represented month per source + # person before checking the documented age-2--7 respondent scope, + # so a within-year birthday does not introduce a spurious age 8. + observed = pd.to_numeric(chunk[status_column], errors="coerce").eq( + 1 + ) & pd.to_numeric(chunk[column], errors="coerce").isin([1, 2]) + observed_child_care_rows[column].append( + chunk.loc[ + observed, + ["SSUID", "PNUM", "MONTHCODE", "TAGE"], + ].copy() + ) + + assert direct_positive_under_3 == 0 + for column, parts in observed_child_care_rows.items(): + observed = pd.concat(parts, ignore_index=True) + first_person_month = observed.sort_values( + "MONTHCODE", kind="mergesort" + ).drop_duplicates(["SSUID", "PNUM"], keep="first") + ages = pd.to_numeric(first_person_month["TAGE"], errors="coerce").dropna() + assert len(ages), f"{column} has no directly observed source people" + assert int(ages.min()) >= 2 + assert int(ages.max()) <= 7 + + +def test_administrative_counts_are_semantic_non_substitutes() -> None: + substitutes = _entry()["evidence"]["administrative_non_substitutes"] + assert substitutes == { + "cumulative_enrollment": ( + "Counts every participant served at any point in the program year, " + "including turnover replacements; it is not point-in-time take-up " + "and supplies no person identity." + ), + "funded_enrollment": ( + "Capacity or federally supported slots, not observed occupancy or " + "person identity." + ), + "source": ( + "https://headstart.gov/sites/default/files/pdf/" + "service-snapshot-ehs-2023-2024.pdf" + ), + } + + +def test_take_up_contract_retains_source_unavailable_treatment() -> None: + program = load_take_up_contract().program_map()[ + "takes_up_early_head_start_if_eligible" + ] + assert program.populace_treatment == "rate_unsourced" + assert program.rate["status"] == "source_unavailable" + assert "locked individual-level source" in program.raw["followup"] + assert "synthesize" in program.raw["notes"] + + +@requires_us +def test_policyengine_1_764_6_requires_under_3_or_pregnant_person_input() -> None: + from policyengine_us import CountryTaxBenefitSystem + + assert version("policyengine-us") == "1.764.6" + system = CountryTaxBenefitSystem() + evidence = _entry()["evidence"]["policyengine_variable"] + assert evidence == { + "version": "1.764.6", + "entity": "person", + "definition_period": "year", + "value_type": "bool", + "default": True, + "eligibility_domain": ( + "age < 3 or is_pregnant, plus Head Start income or categorical eligibility" + ), + } + variable = system.variables["takes_up_early_head_start_if_eligible"] + assert variable.is_input_variable() + assert variable.entity.key == "person" + assert str(variable.definition_period).lower() == "year" + assert variable.value_type is bool + assert variable.default_value is True + + formula = system.variables["is_early_head_start_eligible"].get_formula("2024") + assert formula is not None + source = inspect.getsource(formula) + assert "age < p.early_head_start.age_limit" in source + assert "is_pregnant" in source + assert system.parameters("2024").gov.hhs.head_start.early_head_start.age_limit == 3 + + +def test_generated_release_manifest_preserves_evidenced_exclusion() -> None: + reason = load_release_input_coverage_manifest().reviewed_exclusions[ + "takes_up_early_head_start_if_eligible" + ] + assert reason == _entry()["reason"] + assert reason.startswith("SOURCE UNAVAILABILITY WITH EVIDENCE:") diff --git a/packages/populace-build/tests/test_us_education_inputs.py b/packages/populace-build/tests/test_us_education_inputs.py new file mode 100644 index 00000000..fba1aa4c --- /dev/null +++ b/packages/populace-build/tests/test_us_education_inputs.py @@ -0,0 +1,399 @@ +"""US education-assistance and education-credit input stage tests. + +The retired eCPS pipeline sourced educational assistance directly from CPS +ASEC ``ED_VAL`` and imputed ``qualified_tuition_expenses`` from the PUF. The +PUF transformation has already combined the applicable source fields before +this stage runs; this stage preserves that tuition amount and turns its +positive-support mask into the five factual AOTC eligibility inputs exported +by the reference eCPS. +""" + +from __future__ import annotations + +import numpy as np +import pandas as pd +import pytest + +from populace.build.source_manifest import SourceStageSpec +from populace.build.source_runtime import SourceRuntimeError +from populace.build.us_runtime import ( + BASE_ASEC_SUPPORT_CHANNEL, + PUF_TAX_DETAIL_SUPPORT_CHANNEL, + US_AOTC_ELIGIBILITY_OUTPUT_COLUMNS, + US_EDUCATION_INPUTS_NONCONSTANT_PERSON_COLUMNS, + US_EDUCATION_INPUTS_OUTPUT_COLUMNS, + US_EDUCATION_INPUTS_REQUIRED_SOURCE_COLUMNS, + US_EDUCATION_INPUTS_STAGE_NAME, + clone_us_frame_for_puf_support, + derive_us_education_inputs_from_manifest, + impute_us_puf_tax_detail_support, + support_channel_column, + us_education_inputs_signal_gate, + us_education_inputs_stage_spec, + us_education_inputs_summary, + with_us_education_inputs, +) +from populace.build.us_runtime.source_runtime import us_source_operation_handlers +from populace.frame import US_SCHEMA, Frame, WeightKind, Weights + +TIME_PERIOD = 2024 + +_EXPECTED_AOTC_COLUMNS = ( + "is_pursuing_credential_for_american_opportunity_credit", + "attends_eligible_educational_institution_for_american_opportunity_credit", + "is_enrolled_at_least_half_time_for_american_opportunity_credit", + "has_american_opportunity_credit_1098_t_or_exception", + "has_american_opportunity_credit_institution_ein", +) +_EXPECTED_OUTPUT_COLUMNS = ( + "qualified_tuition_expenses", + "educational_assistance", + *_EXPECTED_AOTC_COLUMNS, +) + + +def _person_table(rows: list[dict]) -> pd.DataFrame: + """Return a person table with both education source columns present.""" + + records: list[dict] = [] + for index, row in enumerate(rows): + record = { + "ED_VAL": 0.0, + "qualified_tuition_expenses": 0.0, + } + record.update(row) + record.setdefault("person_id", index + 1) + record.setdefault("person_household_id", index + 1) + records.append(record) + return pd.DataFrame(records) + + +def _us_frame( + person_rows: list[dict], + *, + household_weights: list[float] | None = None, +) -> Frame: + person = _person_table(person_rows) + n = len(person) + household_ids = person["person_household_id"].to_numpy(dtype=np.int64) + unique_households = np.unique(household_ids) + person["person_tax_unit_id"] = household_ids + 1_000 + person["person_spm_unit_id"] = household_ids + 2_000 + person["person_family_id"] = household_ids + 3_000 + person["person_marital_unit_id"] = np.arange(n, dtype=np.int64) + 4_000 + tables = { + "person": person, + "household": pd.DataFrame({"household_id": unique_households}), + "tax_unit": pd.DataFrame({"tax_unit_id": unique_households + 1_000}), + "spm_unit": pd.DataFrame({"spm_unit_id": unique_households + 2_000}), + "family": pd.DataFrame({"family_id": unique_households + 3_000}), + "marital_unit": pd.DataFrame( + {"marital_unit_id": np.arange(n, dtype=np.int64) + 4_000} + ), + } + weights = ( + [1.0] * len(unique_households) + if household_weights is None + else household_weights + ) + return Frame( + tables, + US_SCHEMA, + { + "household": Weights( + values=np.asarray(weights, dtype=np.float64), + kind=WeightKind.DESIGN, + ) + }, + ) + + +def _plausible_rows() -> list[dict]: + """100 people: 2% with tuition and 4% with assistance.""" + + rows = [ + {"qualified_tuition_expenses": 1_000.0}, + {"qualified_tuition_expenses": 4_000.0}, + {"ED_VAL": 2_500.0}, + {"ED_VAL": 5_000.0}, + {"ED_VAL": 7_500.0}, + {"ED_VAL": 10_000.0}, + ] + rows.extend({} for _ in range(94)) + return rows + + +def _operation(): + spec = us_education_inputs_stage_spec() + return next(op for op in spec.operations if op.kind == "derive_education_inputs") + + +class TestManifestDeclaration: + def test_stage_declares_the_complete_seven_column_family(self) -> None: + spec = us_education_inputs_stage_spec() + + assert spec.stage == US_EDUCATION_INPUTS_STAGE_NAME == "education_inputs" + assert US_AOTC_ELIGIBILITY_OUTPUT_COLUMNS == _EXPECTED_AOTC_COLUMNS + assert US_EDUCATION_INPUTS_OUTPUT_COLUMNS == _EXPECTED_OUTPUT_COLUMNS + assert ( + US_EDUCATION_INPUTS_NONCONSTANT_PERSON_COLUMNS + == _EXPECTED_OUTPUT_COLUMNS + ) + assert tuple(spec.outputs) == _EXPECTED_OUTPUT_COLUMNS + assert US_EDUCATION_INPUTS_REQUIRED_SOURCE_COLUMNS == ( + "ED_VAL", + "qualified_tuition_expenses", + ) + assert set(spec.nonnegative_outputs) == { + "qualified_tuition_expenses", + "educational_assistance", + } + + def test_stage_reads_person_then_derives_education_inputs(self) -> None: + kinds = [ + operation.kind for operation in us_education_inputs_stage_spec().operations + ] + assert kinds == ["read_table", "derive_education_inputs"] + + def test_handler_is_registered(self) -> None: + handlers = us_source_operation_handlers() + assert ( + handlers["derive_education_inputs"] + is derive_us_education_inputs_from_manifest + ) + + +class TestDerivation: + def _derive(self, table: pd.DataFrame) -> pd.DataFrame: + return derive_us_education_inputs_from_manifest(table, _operation(), None) + + def test_preserves_upstream_tuition_and_flags_exactly_positive_tuition( + self, + ) -> None: + # E03230/E87530 have already been reconciled by the PUF transformation; + # this stage must not reinterpret or rescale their combined result. + table = _person_table( + [ + {"qualified_tuition_expenses": 0.0}, + {"qualified_tuition_expenses": 1_250.0}, + {"qualified_tuition_expenses": 4_000.0}, + ] + ) + + result = self._derive(table) + + assert result["qualified_tuition_expenses"].tolist() == [0.0, 1_250.0, 4_000.0] + expected = [False, True, True] + for column in _EXPECTED_AOTC_COLUMNS: + assert result[column].tolist() == expected + + def test_maps_ed_val_to_educational_assistance(self) -> None: + result = self._derive(_person_table([{"ED_VAL": 750.0}, {"ED_VAL": 0.0}])) + assert result["educational_assistance"].tolist() == [750.0, 0.0] + + def test_amounts_are_numeric_finite_and_clipped_nonnegative(self) -> None: + result = self._derive( + _person_table( + [ + {"ED_VAL": -10.0, "qualified_tuition_expenses": -20.0}, + {"ED_VAL": np.nan, "qualified_tuition_expenses": np.nan}, + {"ED_VAL": "300", "qualified_tuition_expenses": "1200"}, + ] + ) + ) + + assert result["educational_assistance"].tolist() == [0.0, 0.0, 300.0] + assert result["qualified_tuition_expenses"].tolist() == [0.0, 0.0, 1_200.0] + for column in _EXPECTED_AOTC_COLUMNS: + assert result[column].tolist() == [False, False, True] + + @pytest.mark.parametrize("missing", US_EDUCATION_INPUTS_REQUIRED_SOURCE_COLUMNS) + def test_missing_source_column_is_named(self, missing: str) -> None: + table = _person_table([{}]).drop(columns=[missing]) + with pytest.raises(SourceRuntimeError, match=missing): + self._derive(table) + + def test_requires_person_table_first(self) -> None: + with pytest.raises(SourceRuntimeError, match="person table"): + derive_us_education_inputs_from_manifest(None, _operation(), None) + + def test_unexpected_parameters_are_refused(self) -> None: + operation_spec = SourceStageSpec.from_mapping( + { + "stage": US_EDUCATION_INPUTS_STAGE_NAME, + "survey": "test ASEC + PUF", + "source": "https://example.com", + "grain": "person", + "operations": [ + {"kind": "read_table", "table": "person"}, + {"kind": "derive_education_inputs", "surprise": True}, + ], + "outputs": list(US_EDUCATION_INPUTS_OUTPUT_COLUMNS), + } + ) + with pytest.raises(SourceRuntimeError, match="surprise"): + derive_us_education_inputs_from_manifest( + _person_table([{}]), operation_spec.operations[1], None + ) + + +class TestFrameIntegration: + def test_with_inputs_writes_the_complete_family(self) -> None: + frame = with_us_education_inputs( + _us_frame( + [ + {"ED_VAL": 600.0, "qualified_tuition_expenses": 0.0}, + {"ED_VAL": 0.0, "qualified_tuition_expenses": 2_000.0}, + ] + ), + seed=0, + time_period=TIME_PERIOD, + ) + person = frame.table("person") + + assert set(_EXPECTED_OUTPUT_COLUMNS).issubset(person.columns) + assert person["educational_assistance"].tolist() == [600.0, 0.0] + for column in _EXPECTED_AOTC_COLUMNS: + assert person[column].tolist() == [False, True] + + def test_frame_with_coherent_signal_passes_through_untouched(self) -> None: + derived = with_us_education_inputs( + _us_frame(_plausible_rows()), seed=0, time_period=TIME_PERIOD + ) + again = with_us_education_inputs(derived, seed=1, time_period=TIME_PERIOD) + assert again is derived + + def test_incoherent_default_surface_is_healed_from_sources(self) -> None: + constants = { + "educational_assistance": 0.0, + **{column: False for column in _EXPECTED_AOTC_COLUMNS}, + } + frame = with_us_education_inputs( + _us_frame( + [ + { + "ED_VAL": 900.0, + "qualified_tuition_expenses": 1_500.0, + **constants, + }, + {**constants}, + ] + ), + seed=0, + time_period=TIME_PERIOD, + ) + person = frame.table("person") + + assert person["educational_assistance"].tolist() == [900.0, 0.0] + for column in _EXPECTED_AOTC_COLUMNS: + assert person[column].tolist() == [True, False] + + def test_missing_sources_without_a_complete_surface_raise(self) -> None: + frame = _us_frame([{}]) + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + tables["person"] = tables["person"].drop( + columns=list(US_EDUCATION_INPUTS_REQUIRED_SOURCE_COLUMNS) + ) + stripped = Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + ) + + with pytest.raises(SourceRuntimeError, match="ED_VAL"): + with_us_education_inputs(stripped, seed=0, time_period=TIME_PERIOD) + + def test_puf_support_to_education_stage_preserves_real_source_signal( + self, + ) -> None: + rows = [{"ED_VAL": 1_000.0} for _ in range(4)] + rows.extend({} for _ in range(96)) + expanded = clone_us_frame_for_puf_support(_us_frame(rows)) + donor = pd.DataFrame( + { + "puf_predictor_tax_unit_person_count": np.ones(100), + "qualified_tuition_expenses": [1_000.0, 4_000.0, *([0.0] * 98)], + "weight": np.ones(100), + } + ) + + imputed = impute_us_puf_tax_detail_support( + expanded, + donor, + predictors=("puf_predictor_tax_unit_person_count",), + person_outputs=("qualified_tuition_expenses",), + tax_unit_outputs=(), + seed=0, + n_estimators=20, + ) + person = imputed.table("person") + channel = person[support_channel_column("person")] + assert not person.loc[ + channel == BASE_ASEC_SUPPORT_CHANNEL, + "qualified_tuition_expenses", + ].any() + assert person.loc[ + channel == PUF_TAX_DETAIL_SUPPORT_CHANNEL, + "qualified_tuition_expenses", + ].gt(0).any() + + result = with_us_education_inputs( + imputed, + seed=0, + time_period=TIME_PERIOD, + ) + + gate = us_education_inputs_signal_gate(result) + assert gate.passed, gate.failures + result_person = result.table("person") + tuition_positive = result_person["qualified_tuition_expenses"] > 0 + for column in _EXPECTED_AOTC_COLUMNS: + assert result_person[column].equals(tuition_positive) + + +class TestGate: + def test_plausible_surface_passes(self) -> None: + frame = with_us_education_inputs( + _us_frame(_plausible_rows()), seed=0, time_period=TIME_PERIOD + ) + + summary = us_education_inputs_summary(frame) + gate = us_education_inputs_signal_gate(frame) + + assert gate.passed, gate.failures + assert gate.details == summary + + def test_missing_columns_fail(self) -> None: + gate = us_education_inputs_signal_gate(_us_frame([{}])) + assert not gate.passed + assert any("missing" in failure for failure in gate.failures) + + def test_all_zero_surface_fails(self) -> None: + rows = [ + { + "educational_assistance": 0.0, + **{column: False for column in _EXPECTED_AOTC_COLUMNS}, + } + for _ in range(20) + ] + gate = us_education_inputs_signal_gate(_us_frame(rows)) + + assert not gate.passed + assert any( + token in " ".join(gate.failures).lower() + for token in ("constant", "zero", "share") + ) + + def test_flag_disagreeing_with_positive_tuition_fails(self) -> None: + frame = with_us_education_inputs( + _us_frame(_plausible_rows()), seed=0, time_period=TIME_PERIOD + ) + frame.table("person").loc[ + 0, "has_american_opportunity_credit_institution_ein" + ] = False + + gate = us_education_inputs_signal_gate(frame) + + assert not gate.passed + message = " ".join(gate.failures).lower() + assert any(token in message for token in ("coher", "match", "disagree")) diff --git a/packages/populace-build/tests/test_us_educator_expenses.py b/packages/populace-build/tests/test_us_educator_expenses.py new file mode 100644 index 00000000..b9d79856 --- /dev/null +++ b/packages/populace-build/tests/test_us_educator_expenses.py @@ -0,0 +1,443 @@ +"""Contracts for the IRS PUF educator-expense input restoration.""" + +from __future__ import annotations + +import importlib.util +from importlib.metadata import version + +import numpy as np +import pandas as pd +import pytest + +from populace.build.us_runtime import ( + BASE_ASEC_SUPPORT_CHANNEL, + PUF_TAX_DETAIL_SUPPORT_CHANNEL, + clone_us_frame_for_puf_support, + impute_us_puf_tax_detail_support, + puf_tax_unit_donor_from_arrays, + support_channel_column, + us_release_reform_coverage_probes, +) +from populace.build.us_runtime import puf_support as puf_support_module +from populace.build.us_runtime.educator_expenses import ( + EDUCATOR_EXPENSE_ARCHIVED_ALLOCATION_URL, + EDUCATOR_EXPENSE_ARCHIVED_DERIVATION_URL, + EDUCATOR_EXPENSE_ARCHIVED_EXPORT_URL, + EDUCATOR_EXPENSE_ARCHIVED_PUF_IMPUTATION_URL, + derive_us_educator_expense_from_puf, + us_educator_expense_signal_gate, + us_educator_expense_stage_spec, +) +from populace.build.us_runtime.puf_aggregate_records import ( + _reconcile_puf_educator_expense_from_source, +) +from populace.frame import US_SCHEMA, Frame, WeightKind, Weights + +policyengine_us_installed = importlib.util.find_spec("policyengine_us") is not None +requires_us = pytest.mark.skipif( + not policyengine_us_installed, + reason="requires the policyengine-us [us] extra (build environment)", +) + +_ARCHIVED_DATA_REPOSITORY = "policyengine-" + "us-data" +_ARCHIVED_ROOT = ( + "https://github.com/PolicyEngine/" + f"{_ARCHIVED_DATA_REPOSITORY}/blob/" + "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe/" + "policyengine_" + "us_data/" +) + + +class _ResolvedWeights: + def __init__(self, values: np.ndarray) -> None: + self.values = values + + +class _PersonFrame: + def __init__(self, person: pd.DataFrame, weights: np.ndarray | None = None) -> None: + self._person = person + self._weights = np.ones(len(person)) if weights is None else weights + + def table(self, entity: str) -> pd.DataFrame: + assert entity == "person" + return self._person + + def resolve_weights(self, entity: str) -> _ResolvedWeights: + assert entity == "person" + return _ResolvedWeights(np.asarray(self._weights, dtype=np.float64)) + + +def _minimal_us_frame() -> Frame: + person = pd.DataFrame( + { + "person_id": np.asarray([1, 2, 3], dtype="int64"), + "person_household_id": np.asarray([1, 1, 2], dtype="int64"), + "person_tax_unit_id": np.asarray([10, 10, 20], dtype="int64"), + "person_spm_unit_id": np.asarray([100, 100, 200], dtype="int64"), + "person_family_id": np.asarray([1_000, 1_000, 2_000], dtype="int64"), + "person_marital_unit_id": np.asarray( + [10_000, 10_000, 20_000], dtype="int64" + ), + "employment_income_before_lsr": np.asarray([75_000.0, 25_000.0, 50_000.0]), + } + ) + tables = { + "person": person, + "household": pd.DataFrame( + { + "household_id": np.asarray([1, 2], dtype="int64"), + "state_fips": np.asarray([6, 36], dtype="int64"), + } + ), + "tax_unit": pd.DataFrame( + { + "tax_unit_id": np.asarray([10, 20], dtype="int64"), + "filing_status_input": ["JOINT", "SINGLE"], + } + ), + "spm_unit": pd.DataFrame( + {"spm_unit_id": np.asarray([100, 200], dtype="int64")} + ), + "family": pd.DataFrame( + {"family_id": np.asarray([1_000, 2_000], dtype="int64")} + ), + "marital_unit": pd.DataFrame( + {"marital_unit_id": np.asarray([10_000, 20_000], dtype="int64")} + ), + } + return Frame( + tables, + US_SCHEMA, + { + "household": Weights( + values=np.asarray([100.0, 100.0]), + kind=WeightKind.DESIGN, + ) + }, + pd.Series(["asec_2024", "asec_2024", "asec_2024"], name="stratum"), + ) + + +def test_immutable_archive_urls_pin_derivation_allocation_export_and_qrf() -> None: + assert EDUCATOR_EXPENSE_ARCHIVED_DERIVATION_URL == ( + _ARCHIVED_ROOT + "datasets/puf/puf.py#L636-L649" + ) + assert EDUCATOR_EXPENSE_ARCHIVED_ALLOCATION_URL == ( + _ARCHIVED_ROOT + "datasets/puf/puf.py#L617-L620" + ) + assert EDUCATOR_EXPENSE_ARCHIVED_EXPORT_URL == ( + _ARCHIVED_ROOT + "datasets/puf/puf.py#L804-L815" + ) + assert EDUCATOR_EXPENSE_ARCHIVED_PUF_IMPUTATION_URL == ( + _ARCHIVED_ROOT + "calibration/puf_impute.py#L940-L1075" + ) + + +def test_archived_e03220_mapping_is_exact_and_preserves_observed_topcode() -> None: + source = pd.DataFrame( + { + "E03220": [0.0, 250.0, 500.0, 99_999.0], + "other": [1, 2, 3, 4], + } + ) + before = source.copy(deep=True) + + result = derive_us_educator_expense_from_puf(source) + + pd.testing.assert_frame_equal(source, before) + assert "educator_expense" not in source + assert result["educator_expense"].tolist() == [0.0, 250.0, 500.0, 99_999.0] + assert result["other"].tolist() == [1, 2, 3, 4] + + +@pytest.mark.parametrize( + ("source", "message"), + [ + (pd.DataFrame({"other": [1.0]}), "requires source column"), + (pd.DataFrame({"E03220": ["not numeric"]}), "nonnumeric or nonfinite"), + (pd.DataFrame({"E03220": [np.nan]}), "nonnumeric or nonfinite"), + (pd.DataFrame({"E03220": [np.inf]}), "nonnumeric or nonfinite"), + (pd.DataFrame({"E03220": [-1.0]}), "negative value"), + ], +) +def test_source_derivation_fails_closed( + source: pd.DataFrame, + message: str, +) -> None: + with pytest.raises(ValueError, match=message): + derive_us_educator_expense_from_puf(source) + + +def test_shared_puf_stage_pins_source_artifact_mapping_and_qrf_contract() -> None: + stage = us_educator_expense_stage_spec() + operations = {operation.kind: operation for operation in stage.operations} + derive = operations["derive_puf_policyengine_variables"] + qrf = operations["fit_weighted_qrf"] + + assert stage.stage == "puf_tax_detail" + assert stage.survey == "IRS PUF 2015 (uprated)" + assert stage.source == ( + "https://www.irs.gov/statistics/" + "soi-tax-stats-individual-public-use-microdata-files" + ) + assert stage.grain == "tax_unit" + assert derive.parameters["educator_expense_source"] == "E03220" + assert derive.parameters["educator_expense_output"] == "educator_expense" + assert qrf.parameters["predictors"] == [ + "employment_income", + "self_employment_income", + "taxable_interest_income", + "qualified_dividend_income", + "non_qualified_dividend_income", + "capital_gains", + "filing_status", + ] + assert "educator_expense" in stage.outputs + assert "educator_expense" in stage.nonnegative_outputs + locators = [str(artifact.get("locator")) for artifact in stage.artifacts] + assert any("E03220" in locator for locator in locators) + assert any( + "release://policyengine/irs-soi-puf/1.8.0/puf_2024.h5" in locator + for locator in locators + ) + assert any( + "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe" in locator for locator in locators + ) + + +def test_post_disaggregation_reconciliation_uses_final_e03220() -> None: + source = pd.DataFrame( + { + "E03220": [0.0, 300.0, 600.0], + "educator_expense": [999.0, 999.0, 999.0], + } + ) + before = source.copy(deep=True) + + result = _reconcile_puf_educator_expense_from_source(source) + + pd.testing.assert_frame_equal(source, before) + assert result["educator_expense"].tolist() == [0.0, 300.0, 600.0] + + +@requires_us +def test_processed_puf_person_array_is_available_to_tax_unit_donor() -> None: + donor = puf_tax_unit_donor_from_arrays( + { + "tax_unit_id": [10, 20], + "household_weight": [100.0, 200.0], + "filing_status": [b"SINGLE", b"JOINT"], + "person_tax_unit_id": [10, 20, 20], + "educator_expense": [300.0, 150.0, 150.0], + }, + person_outputs=("educator_expense",), + tax_unit_outputs=(), + ) + + assert donor["educator_expense"].tolist() == [300.0, 300.0] + assert donor["weight"].tolist() == [100.0, 200.0] + assert donor["tax_unit_person_count"].tolist() == [1, 2] + + +def test_puf_support_declares_sparse_nonnegative_employment_allocation() -> None: + assert ( + "educator_expense" in puf_support_module.PUF_TAX_DETAIL_DEFAULT_PERSON_OUTPUTS + ) + assert ( + "educator_expense" in puf_support_module._PUF_TAX_DETAIL_SPARSE_PERSON_OUTPUTS + ) + assert "educator_expense" in puf_support_module._PUF_TAX_DETAIL_NONNEGATIVE_OUTPUTS + assert puf_support_module._PERSON_OUTPUT_DISTRIBUTION_BASIS["educator_expense"] == ( + "employment_income_before_lsr", + ) + + +@requires_us +def test_weighted_qrf_writes_only_puf_channel_and_allocates_by_employment( + monkeypatch: pytest.MonkeyPatch, +) -> None: + class FakeQRF: + def __init__(self, *, n_estimators: int, seed: int) -> None: + assert n_estimators == 4 + assert seed == 9 + + def fit( + self, + frame, + predictors, + outputs, + *, + weights, + ) -> FakeQRF: + assert predictors == [ + "puf_predictor_filing_status_code", + "puf_predictor_tax_unit_person_count", + ] + assert outputs == ["educator_expense"] + assert weights == "design" + return self + + def predict(self, features: pd.DataFrame) -> pd.DataFrame: + assert len(features) == 2 + # Sparse snapping maps these to observed donor values 600 and 0. + return pd.DataFrame( + {"educator_expense": [500.0, 100.0]}, + index=features.index, + ) + + monkeypatch.setattr(puf_support_module, "QRF", FakeQRF) + expanded = clone_us_frame_for_puf_support(_minimal_us_frame()) + donor = pd.DataFrame( + { + "puf_predictor_filing_status_code": [1.0, 2.0], + "puf_predictor_tax_unit_person_count": [1.0, 2.0], + "educator_expense": [0.0, 600.0], + "weight": [1.0, 1.0], + } + ) + + imputed = impute_us_puf_tax_detail_support( + expanded, + donor, + predictors=( + "puf_predictor_filing_status_code", + "puf_predictor_tax_unit_person_count", + ), + person_outputs=("educator_expense",), + tax_unit_outputs=(), + n_estimators=4, + seed=9, + ) + + person = imputed.table("person") + channel = support_channel_column("person") + asec = person[person[channel] == BASE_ASEC_SUPPORT_CHANNEL] + puf = person[person[channel] == PUF_TAX_DETAIL_SUPPORT_CHANNEL] + assert asec["educator_expense"].tolist() == [0.0, 0.0, 0.0] + assert puf["educator_expense"].tolist() == [450.0, 150.0, 0.0] + assert puf.groupby("person_tax_unit_id")["educator_expense"].sum().tolist() == [ + 600.0, + 0.0, + ] + assert "educator_expense" not in expanded.table("person") + + +def test_signal_gate_requires_exact_zero_asec_and_sparse_nonzero_puf() -> None: + n_per_channel = 1_000 + values = np.zeros(2 * n_per_channel) + values[n_per_channel : n_per_channel + 20] = 300.0 + channels = np.asarray( + [BASE_ASEC_SUPPORT_CHANNEL] * n_per_channel + + [PUF_TAX_DETAIL_SUPPORT_CHANNEL] * n_per_channel + ) + frame = _PersonFrame( + pd.DataFrame( + { + "educator_expense": values, + support_channel_column("person"): channels, + } + ) + ) + + result = us_educator_expense_signal_gate(frame) # type: ignore[arg-type] + + assert result.passed, result.failures + assert result.details["positive_share"] == pytest.approx(0.01) + assert result.details["channels"][BASE_ASEC_SUPPORT_CHANNEL][ + "positive_share" + ] == pytest.approx(0.0) + assert result.details["channels"][PUF_TAX_DETAIL_SUPPORT_CHANNEL][ + "positive_share" + ] == pytest.approx(0.02) + + values[0] = 300.0 + contaminated = _PersonFrame( + pd.DataFrame( + { + "educator_expense": values, + support_channel_column("person"): channels, + } + ) + ) + failed = us_educator_expense_signal_gate(contaminated) # type: ignore[arg-type] + assert not failed.passed + assert any("ASEC support channel" in failure for failure in failed.failures) + + +@requires_us +def test_policyengine_us_17646_contract_is_person_year_input_in_ald_graph() -> None: + from policyengine_us import CountryTaxBenefitSystem + + assert version("policyengine-us") == "1.764.6" + system = CountryTaxBenefitSystem() + variable = system.variables["educator_expense"] + + assert variable.is_input_variable() + assert variable.entity.key == "person" + assert str(variable.definition_period).lower() == "year" + assert variable.default_value == 0 + assert ( + variable.uprating + == "calibration.gov.cbo.income_by_source.adjusted_gross_income" + ) + assert system.variables["above_the_line_deductions"].adds == ( + "gov.irs.ald.deductions" + ) + assert "educator_expense" in system.parameters.gov.irs.ald.deductions("2024-01-01") + + +@requires_us +def test_shipped_abolition_probe_has_positive_sign_and_binds_live() -> None: + from policyengine_core.reforms import Reform + from policyengine_us import CountryTaxBenefitSystem, Simulation + + probe = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "educator_expense_ald_abolition" + ) + assert probe.period == 2024 + assert probe.expected_sign == "positive" + assert probe.effect_direction == "reform_minus_baseline" + assert probe.budget_measure == "income_tax" + assert probe.binding_inputs == ("educator_expense",) + assert probe.min_abs_effect == 1_000_000.0 + deductions = probe.parameter_changes["gov.irs.ald.deductions"] + assert set(deductions) == {"2024-01-01.2024-12-31"} + assert "educator_expense" not in deductions["2024-01-01.2024-12-31"] + + situation = { + "people": { + "adult": { + "age": {"2024": 40}, + "employment_income": {"2024": 100_000}, + "educator_expense": {"2024": 300}, + } + }, + "tax_units": { + "tax_unit": { + "members": ["adult"], + "filing_status": {"2024": "SINGLE"}, + } + }, + "households": { + "household": { + "members": ["adult"], + "state_code": {"2024": "CA"}, + } + }, + } + reform = Reform.from_dict(dict(probe.parameter_changes), country_id="us") + baseline = Simulation(situation=situation) + reformed = Simulation( + tax_benefit_system=CountryTaxBenefitSystem(reform=(reform,)), + situation=situation, + ) + + assert baseline.calculate("above_the_line_deductions", 2024)[0] == 300.0 + assert reformed.calculate("above_the_line_deductions", 2024)[0] == 0.0 + assert ( + reformed.calculate("income_tax", 2024)[0] + - baseline.calculate("income_tax", 2024)[0] + > 0.0 + ) diff --git a/packages/populace-build/tests/test_us_energy_subsidy.py b/packages/populace-build/tests/test_us_energy_subsidy.py new file mode 100644 index 00000000..f52f38e4 --- /dev/null +++ b/packages/populace-build/tests/test_us_energy_subsidy.py @@ -0,0 +1,550 @@ +"""ASEC energy-subsidy restoration and PUF-half QRF treatment.""" + +from __future__ import annotations + +import importlib.util +import json +from hashlib import sha256 +from importlib.metadata import version +from pathlib import Path + +import numpy as np +import pandas as pd +import pytest + +import populace.build.us_runtime.energy_subsidy as module +from populace.build.source_runtime import ( + SourceRuntimeConfig, + SourceRuntimeContext, + SourceRuntimeError, +) +from populace.build.us_runtime.asec_pool import load_asec_h5_tables +from populace.build.us_runtime.energy_subsidy import ( + ENERGY_SUBSIDY_ARCHIVED_CPS_DERIVATION_URL, + ENERGY_SUBSIDY_ARCHIVED_PUF_IMPUTATION_URL, + US_ENERGY_SUBSIDY_OUTPUT_COLUMNS, + US_ENERGY_SUBSIDY_REQUIRED_SOURCE_COLUMNS, + US_ENERGY_SUBSIDY_STAGE_NAME, + derive_us_energy_subsidy_from_manifest, + impute_us_energy_subsidy_to_puf_support_from_manifest, + us_energy_subsidy_signal_gate, + us_energy_subsidy_stage_spec, + us_energy_subsidy_summary, + with_us_energy_subsidy_input, +) +from populace.build.us_runtime.puf_support import clone_us_frame_for_puf_support +from populace.frame import US_SCHEMA, Frame, WeightKind, Weights + +ROOT = Path(__file__).resolve().parents[3] +policyengine_us_installed = importlib.util.find_spec("policyengine_us") is not None +requires_us = pytest.mark.skipif( + not policyengine_us_installed, + reason="requires the policyengine-us [us] extra (build environment)", +) + +_OUTPUT = US_ENERGY_SUBSIDY_OUTPUT_COLUMNS[0] +_ARCHIVED_COMMIT = "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe" + + +def _person_source() -> pd.DataFrame: + return pd.DataFrame( + { + "person_id": np.arange(1, 7, dtype="int64"), + "person_household_id": [10, 10, 20, 30, 40, 50], + "person_tax_unit_id": [100, 100, 200, 300, 400, 500], + "person_spm_unit_id": [1_000, 1_000, 2_000, 3_000, 4_000, 5_000], + "person_family_id": [10_000, 10_000, 20_000, 30_000, 40_000, 50_000], + "person_marital_unit_id": [ + 100_000, + 100_000, + 200_000, + 300_000, + 400_000, + 500_000, + ], + "SPM_ENGVAL": [600.0, 600.0, 0.0, 0.0, 0.0, 0.0], + "WSAL_VAL": [50_000.0, 20_000.0, 0.0, 35_000.0, 10_000.0, 0.0], + "SEMP_VAL": [0.0, 0.0, 20_000.0, 0.0, 5_000.0, 0.0], + "employment_income_before_lsr": [ + 50_000.0, + 20_000.0, + 0.0, + 35_000.0, + 10_000.0, + 0.0, + ], + "self_employment_income_before_lsr": [ + 0.0, + 0.0, + 20_000.0, + 0.0, + 5_000.0, + 0.0, + ], + "age": [35, 33, 45, 29, 55, 70], + "is_female": [False, True, True, False, True, False], + "has_esi": [True, True, False, True, False, False], + "tax_unit_role_input": [ + "HEAD", + "SPOUSE", + "HEAD", + "HEAD", + "HEAD", + "HEAD", + ], + "social_security_retirement": [0.0, 0.0, 0.0, 0.0, 0.0, 15_000.0], + "social_security_disability": [0.0] * 6, + "social_security_dependents": [0.0] * 6, + "social_security_survivors": [0.0] * 6, + } + ) + + +def _frame() -> Frame: + person = _person_source() + ids = { + "household": [10, 20, 30, 40, 50], + "tax_unit": [100, 200, 300, 400, 500], + "spm_unit": [1_000, 2_000, 3_000, 4_000, 5_000], + "family": [10_000, 20_000, 30_000, 40_000, 50_000], + "marital_unit": [100_000, 200_000, 300_000, 400_000, 500_000], + } + tables = { + entity: pd.DataFrame({f"{entity}_id": np.asarray(values, dtype="int64")}) + for entity, values in ids.items() + } + tables["person"] = person + tables["tax_unit"]["filing_status_input"] = [ + "JOINT", + "SINGLE", + "SINGLE", + "SINGLE", + "SINGLE", + ] + return Frame( + tables, + US_SCHEMA, + { + "household": Weights( + np.ones(5, dtype=np.float64), + WeightKind.DESIGN, + ) + }, + ) + + +def _derive(frame: pd.DataFrame) -> pd.DataFrame: + operation = next( + operation + for operation in us_energy_subsidy_stage_spec().operations + if operation.kind == "derive_energy_subsidy" + ) + return derive_us_energy_subsidy_from_manifest(frame, operation, None) + + +def _sha256(path: Path) -> str: + digest = sha256() + with path.open("rb") as stream: + for chunk in iter(lambda: stream.read(1024 * 1024), b""): + digest.update(chunk) + return digest.hexdigest() + + +def test_archived_urls_are_immutable_and_pin_both_retired_steps() -> None: + assert f"/blob/{_ARCHIVED_COMMIT}/" in (ENERGY_SUBSIDY_ARCHIVED_CPS_DERIVATION_URL) + assert f"/blob/{_ARCHIVED_COMMIT}/" in (ENERGY_SUBSIDY_ARCHIVED_PUF_IMPUTATION_URL) + assert ENERGY_SUBSIDY_ARCHIVED_CPS_DERIVATION_URL.endswith( + "datasets/cps/cps.py#L1612-L1622" + ) + assert ENERGY_SUBSIDY_ARCHIVED_PUF_IMPUTATION_URL.endswith( + "datasets/cps/extended_cps.py#L639-L739" + ) + + +def test_stage_manifest_pins_exact_source_qrf_and_reduction_contract() -> None: + spec = us_energy_subsidy_stage_spec() + + assert US_ENERGY_SUBSIDY_STAGE_NAME == "energy_subsidy" + assert spec.stage == US_ENERGY_SUBSIDY_STAGE_NAME + assert spec.grain == "person" + assert tuple(spec.outputs) == US_ENERGY_SUBSIDY_OUTPUT_COLUMNS + assert tuple(spec.nonnegative_outputs) == US_ENERGY_SUBSIDY_OUTPUT_COLUMNS + assert [operation.kind for operation in spec.operations] == [ + "read_table", + "derive_energy_subsidy", + "impute_energy_subsidy_to_puf_support", + ] + assert spec.operations[0].parameters == { + "table": "person", + "weight": "person_weight", + } + impute = spec.operations[2] + assert tuple(impute.parameters["predictors"]) == ( + "age", + "is_male", + "has_esi", + "tax_unit_is_joint", + "tax_unit_count_dependents", + "employment_income", + "self_employment_income", + "social_security", + ) + assert impute.parameters["max_train_samples"] == 5_000 + assert impute.parameters["n_estimators"] == 100 + assert impute.parameters["seed_from_build_config"] is True + assert impute.parameters["weight"] == "person_weight" + assert impute.parameters["reduction"] == "value_from_first_person" + + +def test_direct_source_is_exact_and_replicated() -> None: + result = _derive(_person_source()) + + assert result[_OUTPUT].tolist() == [600.0, 600.0, 0.0, 0.0, 0.0, 0.0] + + +@pytest.mark.parametrize("missing", US_ENERGY_SUBSIDY_REQUIRED_SOURCE_COLUMNS) +def test_direct_source_fails_closed_when_required_source_is_missing( + missing: str, +) -> None: + with pytest.raises(SourceRuntimeError, match=missing): + _derive(_person_source().drop(columns=[missing])) + + +@pytest.mark.parametrize( + ("bad_value", "message"), + [(np.nan, "nonfinite"), (np.inf, "nonfinite"), (-1.0, "negative")], +) +def test_direct_source_rejects_invalid_values( + bad_value: float, + message: str, +) -> None: + person = _person_source() + person.loc[[0, 1], "SPM_ENGVAL"] = bad_value + + with pytest.raises(SourceRuntimeError, match=message): + _derive(person) + + +def test_direct_source_rejects_inconsistent_unit_replicas() -> None: + person = _person_source() + person.loc[1, "SPM_ENGVAL"] = 599.0 + + with pytest.raises(SourceRuntimeError, match="disagrees within replicated"): + _derive(person) + + +def test_with_input_materializes_first_person_spm_unit_values() -> None: + result = with_us_energy_subsidy_input(_frame(), seed=0, time_period=2024) + + assert result.table("spm_unit")[_OUTPUT].tolist() == [ + 600.0, + 0.0, + 0.0, + 0.0, + 0.0, + ] + gate = us_energy_subsidy_signal_gate(result) + assert gate.passed, gate.failures + assert us_energy_subsidy_summary(result)["positive_share_band"] == [0.01, 0.2] + + +def test_existing_output_is_recomputed_from_measured_source() -> None: + frame = _frame() + frame.table("spm_unit")[_OUTPUT] = [111.0, 222.0, 333.0, 444.0, 555.0] + + result = with_us_energy_subsidy_input(frame, seed=0, time_period=2024) + + assert result.table("spm_unit")[_OUTPUT].tolist() == [ + 600.0, + 0.0, + 0.0, + 0.0, + 0.0, + ] + + +def test_production_release_can_preserve_only_a_valid_existing_surface() -> None: + materialized = with_us_energy_subsidy_input(_frame(), seed=0, time_period=2024) + tables = { + entity: materialized.table(entity).copy() for entity in materialized.entities + } + tables["person"] = tables["person"].drop(columns=["SPM_ENGVAL"]) + release = Frame( + tables, + materialized.schema, + { + entity: materialized.weights_for(entity) + for entity in materialized.weighted_entities + }, + materialized.strata, + ) + + with pytest.raises(ValueError, match="cannot heal.*without measured"): + with_us_energy_subsidy_input(release, seed=0, time_period=2024) + + assert ( + with_us_energy_subsidy_input( + release, + seed=0, + time_period=2024, + allow_existing_without_source=True, + ) + is release + ) + + release.table("spm_unit")[_OUTPUT] = 0.0 + with pytest.raises(ValueError, match="cannot heal.*without measured"): + with_us_energy_subsidy_input( + release, + seed=0, + time_period=2024, + allow_existing_without_source=True, + ) + + +def test_puf_half_uses_weighted_qrf_and_first_person_spm_reduction( + monkeypatch: pytest.MonkeyPatch, +) -> None: + direct = with_us_energy_subsidy_input(_frame(), seed=0, time_period=2024) + expanded = clone_us_frame_for_puf_support(direct) + calls: dict[str, object] = {} + + class FakeFitted: + def predict(self, test: pd.DataFrame) -> pd.DataFrame: + calls["test"] = test.copy() + return pd.DataFrame( + {_OUTPUT: np.arange(100.0, 100.0 + len(test))}, + index=test.index, + ) + + class FakeQRF: + def __init__(self, **kwargs: object) -> None: + calls["init"] = kwargs + + def fit( + self, + training: pd.DataFrame, + predictors: list[str], + targets: list[str], + *, + weights: np.ndarray, + ) -> FakeFitted: + calls["training"] = training.copy() + calls["predictors"] = predictors + calls["targets"] = targets + calls["weights"] = weights.copy() + return FakeFitted() + + monkeypatch.setattr(module, "QRF", FakeQRF) + + result = with_us_energy_subsidy_input(expanded, seed=7, time_period=2024) + + assert calls["init"] == {"n_estimators": 100, "seed": 7} + assert len(calls["training"]) == 6 + assert len(calls["test"]) == 6 + assert calls["targets"] == [_OUTPUT] + spm = result.table("spm_unit") + asec = spm[spm["spm_unit_support_channel"] == "asec"] + puf = spm[spm["spm_unit_support_channel"] == "puf_tax_detail"] + assert asec[_OUTPUT].tolist() == [600.0, 0.0, 0.0, 0.0, 0.0] + # The first two PUF people share an SPM unit, so prediction 101 is ignored. + assert puf[_OUTPUT].tolist() == [100.0, 102.0, 103.0, 104.0, 105.0] + + +def test_puf_qrf_caps_training_at_5000_and_keeps_aligned_weights( + monkeypatch: pytest.MonkeyPatch, +) -> None: + operation = next( + operation + for operation in us_energy_subsidy_stage_spec().operations + if operation.kind == "impute_energy_subsidy_to_puf_support" + ) + predictors = tuple(operation.parameters["predictors"]) + asec_rows = 5_010 + puf_rows = 2 + frame = pd.DataFrame( + { + "person_support_channel": ["asec"] * asec_rows + + ["puf_tax_detail"] * puf_rows, + "person_weight": np.arange(1.0, asec_rows + puf_rows + 1.0), + _OUTPUT: np.tile([0.0, 500.0], (asec_rows + puf_rows + 1) // 2)[ + : asec_rows + puf_rows + ], + **{ + f"energy_subsidy_predictor_{predictor}": np.arange( + asec_rows + puf_rows, dtype=np.float64 + ) + for predictor in predictors + }, + } + ) + calls: dict[str, object] = {} + + class FakeFitted: + def predict(self, test: pd.DataFrame) -> pd.DataFrame: + calls["test_rows"] = len(test) + return pd.DataFrame({_OUTPUT: np.zeros(len(test))}, index=test.index) + + class FakeQRF: + def __init__(self, **kwargs: object) -> None: + calls["init"] = kwargs + + def fit( + self, + training: pd.DataFrame, + predictor_names: list[str], + targets: list[str], + *, + weights: np.ndarray, + ) -> FakeFitted: + calls["training_rows"] = len(training) + calls["weights"] = weights.copy() + calls["training_index"] = training.index.to_numpy() + return FakeFitted() + + monkeypatch.setattr(module, "QRF", FakeQRF) + context = SourceRuntimeContext( + config=SourceRuntimeConfig(seed=19, target_year=2024), + tables={}, + ) + + impute_us_energy_subsidy_to_puf_support_from_manifest(frame, operation, context) + + assert calls["init"] == {"n_estimators": 100, "seed": 19} + assert calls["training_rows"] == 5_000 + assert calls["test_rows"] == 2 + np.testing.assert_allclose( + calls["weights"], + frame.loc[calls["training_index"], "person_weight"].to_numpy(), + ) + + +def test_signal_gate_rejects_missing_default_invalid_and_channel_drift() -> None: + valid = with_us_energy_subsidy_input(_frame(), seed=0, time_period=2024) + for values in ( + None, + [0.0] * 5, + [600.0, 0.0, -1.0, 0.0, 0.0], + [600.0, 0.0, np.nan, 0.0, 0.0], + [600.0] * 5, + ): + candidate = _frame() + if values is not None: + candidate.table("spm_unit")[_OUTPUT] = values + assert not us_energy_subsidy_signal_gate(candidate).passed + + assert us_energy_subsidy_signal_gate(valid).passed + + valid.table("spm_unit")["spm_unit_support_channel"] = "asec" + channel_gate = us_energy_subsidy_signal_gate(valid) + assert not channel_gate.passed + assert any("asec positive share" in failure for failure in channel_gate.failures) + assert any( + "puf_tax_detail positive share" in failure for failure in channel_gate.failures + ) + + +def test_all_sha_locked_asec_artifacts_carry_exact_source_signal() -> None: + expected = { + "census_cps_2022.h5": ( + "7ccca976284bb47815d84460cc4f75a0a65d26d7754ab0a0f417de351b3d474e", + 4_740, + 3_439_573.0, + 2_083, + 1_442_882.0, + ), + "census_cps_2023.h5": ( + "cb57817327799f42b741caed5f9be94d04021c2e6809c1ad7bd0686da5428d88", + 4_562, + 3_078_344.0, + 1_995, + 1_270_213.0, + ), + "census_cps_2024.h5": ( + "ec36604cb735a660b51b0b2f90be27d803b5878f3464fb30d0eacead59c1260d", + 4_297, + 2_866_931.0, + 1_955, + 1_246_009.0, + ), + } + summary = json.loads( + (ROOT / "experiments/build_j_recert/base_j.summary.json").read_text() + ) + paths = { + Path(item["path"]).name: Path(item["path"]) + for item in summary["base_source"]["sources"] + } + if not all(paths[name].is_file() for name in expected): + pytest.skip("SHA-locked ASEC artifacts are not mounted") + + for name, expected_values in expected.items(): + digest, positive_rows, person_total, positive_units, unit_total = ( + expected_values + ) + path = paths[name] + assert _sha256(path) == digest + raw_person = load_asec_h5_tables(path)["person"] + assert {"SPM_ID", "SPM_ENGVAL"} <= set(raw_person.columns) + source = raw_person[["SPM_ID", "SPM_ENGVAL"]].rename( + columns={"SPM_ID": "person_spm_unit_id"} + ) + derived = _derive(source) + values = derived[_OUTPUT] + assert int((values > 0.0).sum()) == positive_rows + assert float(values.sum()) == person_total + by_unit = derived.groupby("person_spm_unit_id", sort=False)[_OUTPUT] + assert int((by_unit.first() > 0.0).sum()) == positive_units + assert float(by_unit.first().sum()) == unit_total + np.testing.assert_array_equal(by_unit.min(), by_unit.max()) + + +@requires_us +def test_policyengine_1_764_6_contract_and_live_neutralization() -> None: + from policyengine_core.reforms import Reform + from policyengine_us import CountryTaxBenefitSystem, Simulation + + assert version("policyengine-us") == "1.764.6" + variable = CountryTaxBenefitSystem().variables[_OUTPUT] + assert variable.is_input_variable() + assert variable.entity.key == "spm_unit" + assert variable.value_type is float + assert variable.default_value == 0 + assert str(variable.definition_period).lower() == "year" + + situation = { + "people": {"adult": {"age": {"2024": 40}}}, + "tax_units": {"tax_unit": {"members": ["adult"]}}, + "families": {"family": {"members": ["adult"]}}, + "spm_units": { + "spm_unit": { + "members": ["adult"], + _OUTPUT: {"2024": 2_400.0}, + } + }, + "households": { + "household": { + "members": ["adult"], + "state_code": {"2024": "CA"}, + } + }, + "marital_units": {"marital_unit": {"members": ["adult"]}}, + } + + class NeutralizeEnergySubsidy(Reform): + def apply(self) -> None: + self.neutralize_variable(_OUTPUT) + + baseline = Simulation(situation=situation) + neutralized = Simulation( + situation=situation, + reform=NeutralizeEnergySubsidy, + ) + assert baseline.calculate(_OUTPUT, 2024)[0] == pytest.approx(2_400.0) + assert neutralized.calculate(_OUTPUT, 2024)[0] == 0.0 + for downstream in ("spm_unit_benefits", "spm_unit_net_income"): + effect = ( + baseline.calculate(downstream, 2024)[0] + - neutralized.calculate(downstream, 2024)[0] + ) + assert effect == pytest.approx(2_400.0) diff --git a/packages/populace-build/tests/test_us_farm_business_income.py b/packages/populace-build/tests/test_us_farm_business_income.py new file mode 100644 index 00000000..cb260948 --- /dev/null +++ b/packages/populace-build/tests/test_us_farm_business_income.py @@ -0,0 +1,385 @@ +"""Signed ASEC/PUF farm-business input restoration contracts.""" + +from __future__ import annotations + +import importlib.util +from importlib.metadata import version + +import numpy as np +import pandas as pd +import pytest + +import populace.build.us_runtime.puf_support as puf_support_module +from populace.build.us_runtime.cps_carried import derive_us_cps_carried_inputs +from populace.build.us_runtime.farm_business_income import ( + FARM_BUSINESS_INCOME_ARCHIVED_CPS_FARM_INCOME_URL, + FARM_BUSINESS_INCOME_ARCHIVED_DERIVATION_URL, + FARM_BUSINESS_INCOME_ARCHIVED_EXPORT_URL, + FARM_BUSINESS_INCOME_ARCHIVED_IMPUTATION_URL, + FARM_BUSINESS_INCOME_ARCHIVED_OVERRIDE_URL, + FARM_BUSINESS_INCOME_ARCHIVED_PUF_ARTIFACT_URL, + US_FARM_BUSINESS_INCOME_OUTPUT_COLUMNS, + derive_us_farm_business_income_from_puf, + us_farm_business_income_signal_gate, + us_farm_business_income_stage_spec, +) +from populace.build.us_runtime.puf_aggregate_records import ( + _reconcile_puf_farm_business_income_from_sources, +) +from populace.build.us_runtime.puf_support import ( + PUF_TAX_DETAIL_DEFAULT_PERSON_OUTPUTS, + clone_us_frame_for_puf_support, + impute_us_puf_tax_detail_support, + puf_tax_unit_donor_from_arrays, +) +from populace.frame import US_SCHEMA, Frame, WeightKind, Weights + +policyengine_us_installed = importlib.util.find_spec("policyengine_us") is not None +requires_us = pytest.mark.skipif( + not policyengine_us_installed, + reason="requires the policyengine-us [us] extra (build environment)", +) + +_OPERATIONS, _RENT = US_FARM_BUSINESS_INCOME_OUTPUT_COLUMNS + + +def _asec_frame() -> Frame: + count = 20 + ids = np.arange(1, count + 1, dtype=np.int64) + farm_operations = np.zeros(count, dtype=np.float64) + farm_operations[:4] = [500.0, -200.0, 800.0, -300.0] + person = pd.DataFrame( + { + "person_id": ids, + "person_household_id": ids + 100, + "person_tax_unit_id": ids + 200, + "person_spm_unit_id": ids + 300, + "person_family_id": ids + 400, + "person_marital_unit_id": ids + 500, + "A_AGE": np.arange(30, 30 + count), + "A_SEX": np.tile([1, 2], count // 2), + "WSAL_VAL": np.arange(count) * 1_000.0, + "SEMP_VAL": np.arange(count) * 100.0, + "FRSE_VAL": farm_operations, + "OI_VAL": np.zeros(count), + "OI_OFF": np.zeros(count), + } + ) + tables = { + "person": person, + "household": pd.DataFrame( + {"household_id": ids + 100, "state_fips": np.full(count, 6)} + ), + "tax_unit": pd.DataFrame( + { + "tax_unit_id": ids + 200, + "filing_status_input": ["SINGLE"] * count, + } + ), + "spm_unit": pd.DataFrame({"spm_unit_id": ids + 300}), + "family": pd.DataFrame({"family_id": ids + 400}), + "marital_unit": pd.DataFrame({"marital_unit_id": ids + 500}), + } + return Frame( + tables, + US_SCHEMA, + { + "household": Weights( + np.ones(count, dtype=np.float64), + WeightKind.DESIGN, + ) + }, + ) + + +def _imputed_frame(monkeypatch: pytest.MonkeyPatch) -> Frame: + asec = derive_us_cps_carried_inputs(_asec_frame()) + expanded = clone_us_frame_for_puf_support(asec) + predictions = { + _OPERATIONS: np.asarray([700.0, -400.0, 900.0, -500.0, *([0.0] * 16)]), + _RENT: np.asarray([300.0, -100.0, 600.0, -200.0, *([0.0] * 16)]), + } + + class FakeFitted: + def predict(self, test: pd.DataFrame) -> pd.DataFrame: + return pd.DataFrame(predictions, index=test.index) + + class FakeQRF: + def __init__(self, *, n_estimators: int, seed: int) -> None: + assert n_estimators == 100 + assert seed == 11 + + def fit( + self, + frame: Frame, + predictors: list[str], + outcomes: list[str], + *, + weights: str, + ) -> FakeFitted: + assert predictors == [] + assert outcomes == list(US_FARM_BUSINESS_INCOME_OUTPUT_COLUMNS) + assert weights == "design" + return FakeFitted() + + monkeypatch.setattr(puf_support_module, "QRF", FakeQRF) + donor = pd.DataFrame( + { + _OPERATIONS: [10.0, -20.0, 0.0, 30.0], + _RENT: [-5.0, 15.0, 0.0, 25.0], + "weight": [1.0, 1.0, 1.0, 1.0], + } + ) + return impute_us_puf_tax_detail_support( + expanded, + donor, + predictors=(), + person_outputs=US_FARM_BUSINESS_INCOME_OUTPUT_COLUMNS, + tax_unit_outputs=(), + seed=11, + ) + + +def test_archive_urls_pin_asec_puf_export_qrf_override_and_artifact() -> None: + commit = "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe" + urls = ( + FARM_BUSINESS_INCOME_ARCHIVED_CPS_FARM_INCOME_URL, + FARM_BUSINESS_INCOME_ARCHIVED_DERIVATION_URL, + FARM_BUSINESS_INCOME_ARCHIVED_EXPORT_URL, + FARM_BUSINESS_INCOME_ARCHIVED_IMPUTATION_URL, + FARM_BUSINESS_INCOME_ARCHIVED_OVERRIDE_URL, + FARM_BUSINESS_INCOME_ARCHIVED_PUF_ARTIFACT_URL, + ) + assert all(commit in url for url in urls) + assert FARM_BUSINESS_INCOME_ARCHIVED_CPS_FARM_INCOME_URL.endswith( + "datasets/cps/cps.py#L1363-L1382" + ) + assert FARM_BUSINESS_INCOME_ARCHIVED_DERIVATION_URL.endswith( + "datasets/puf/puf.py#L636-L704" + ) + assert FARM_BUSINESS_INCOME_ARCHIVED_EXPORT_URL.endswith( + "datasets/puf/puf.py#L804-L875" + ) + assert FARM_BUSINESS_INCOME_ARCHIVED_IMPUTATION_URL.endswith( + "calibration/puf_impute.py#L80-L198" + ) + assert FARM_BUSINESS_INCOME_ARCHIVED_OVERRIDE_URL.endswith( + "calibration/puf_impute.py#L513-L672" + ) + assert FARM_BUSINESS_INCOME_ARCHIVED_PUF_ARTIFACT_URL.endswith( + "datasets/puf/puf.py#L1655-L1660" + ) + + +def test_shared_puf_stage_pins_exact_signed_sources_and_outputs() -> None: + spec = us_farm_business_income_stage_spec() + operation = next( + operation + for operation in spec.operations + if operation.kind == "derive_puf_policyengine_variables" + ) + + assert operation.parameters["farm_operations_income_source"] == "E02100" + assert operation.parameters["farm_operations_income_output"] == _OPERATIONS + assert operation.parameters["farm_rent_income_source"] == "E27200" + assert operation.parameters["farm_rent_income_output"] == _RENT + assert set(US_FARM_BUSINESS_INCOME_OUTPUT_COLUMNS) <= set(spec.outputs) + assert not set(US_FARM_BUSINESS_INCOME_OUTPUT_COLUMNS).intersection( + spec.nonnegative_outputs + ) + assert "preserves measured ASEC operations income" in spec.notes + + +def test_direct_puf_mapping_preserves_both_signs_and_does_not_mutate() -> None: + source = pd.DataFrame( + { + "E02100": [10_000.0, -8_000.0, 0.0, 99_999.0], + "E27200": [-3_000.0, 7_000.0, 0.0, 88_888.0], + "T27800": [1.0, 2.0, 3.0, 4.0], + } + ) + original = source.copy(deep=True) + + result = derive_us_farm_business_income_from_puf(source) + + assert result[_OPERATIONS].tolist() == [10_000.0, -8_000.0, 0.0, 99_999.0] + assert result[_RENT].tolist() == [-3_000.0, 7_000.0, 0.0, 88_888.0] + assert result["T27800"].tolist() == [1.0, 2.0, 3.0, 4.0] + pd.testing.assert_frame_equal(source, original) + + +@pytest.mark.parametrize( + ("column", "bad_value", "message"), + [ + ("E02100", np.nan, "nonnumeric or nonfinite"), + ("E02100", np.inf, "nonnumeric or nonfinite"), + ("E27200", -np.inf, "nonnumeric or nonfinite"), + ], +) +def test_direct_puf_mapping_fails_closed_on_invalid_not_negative_values( + column: str, + bad_value: float, + message: str, +) -> None: + source = pd.DataFrame({"E02100": [-1.0], "E27200": [-2.0]}) + source.loc[0, column] = bad_value + + with pytest.raises(ValueError, match=message): + derive_us_farm_business_income_from_puf(source) + + +def test_cps_frse_maps_to_operations_not_schedule_j_farm_income() -> None: + source = _asec_frame() + original = source.table("person").copy(deep=True) + + result = derive_us_cps_carried_inputs(source) + + pd.testing.assert_frame_equal(source.table("person"), original) + np.testing.assert_allclose( + result.table("person")[_OPERATIONS], + original["FRSE_VAL"], + ) + assert "farm_income" not in result.table("person") + assert _RENT not in result.table("person") + + +def test_processed_puf_donor_aggregates_signed_person_values() -> None: + donor = puf_tax_unit_donor_from_arrays( + { + "tax_unit_id": [10, 20], + "household_weight": [100.0, 200.0], + "filing_status": [b"SINGLE", b"JOINT"], + "person_tax_unit_id": [10, 10, 20], + _OPERATIONS: [2_000.0, -3_000.0, -4_000.0], + _RENT: [-500.0, 1_500.0, -2_000.0], + }, + person_outputs=US_FARM_BUSINESS_INCOME_OUTPUT_COLUMNS, + tax_unit_outputs=(), + ) + + assert donor[_OPERATIONS].tolist() == [-1_000.0, -4_000.0] + assert donor[_RENT].tolist() == [1_000.0, -2_000.0] + + +def test_post_disaggregation_reconciliation_uses_final_signed_sources() -> None: + source = pd.DataFrame( + { + "E02100": [-7_000.0, 9_000.0], + "E27200": [3_000.0, -5_000.0], + _OPERATIONS: [999.0, 999.0], + _RENT: [999.0, 999.0], + } + ) + + result = _reconcile_puf_farm_business_income_from_sources(source) + + assert result[_OPERATIONS].tolist() == [-7_000.0, 9_000.0] + assert result[_RENT].tolist() == [3_000.0, -5_000.0] + + +def test_puf_runtime_keeps_signed_outputs_out_of_clipping_and_sparse_sets() -> None: + for output in US_FARM_BUSINESS_INCOME_OUTPUT_COLUMNS: + assert output in PUF_TAX_DETAIL_DEFAULT_PERSON_OUTPUTS + assert output not in puf_support_module._PUF_TAX_DETAIL_NONNEGATIVE_OUTPUTS + assert output not in puf_support_module._PUF_TAX_DETAIL_SPARSE_PERSON_OUTPUTS + + +def test_weighted_qrf_preserves_asec_and_writes_signed_puf_support( + monkeypatch: pytest.MonkeyPatch, +) -> None: + result = _imputed_frame(monkeypatch) + person = result.table("person") + asec = person[person["person_support_channel"] == "asec"] + puf = person[person["person_support_channel"] == "puf_tax_detail"] + + assert asec[_OPERATIONS].tolist()[:4] == [500.0, -200.0, 800.0, -300.0] + assert not asec[_RENT].any() + assert puf[_OPERATIONS].tolist()[:4] == [700.0, -400.0, 900.0, -500.0] + assert puf[_RENT].tolist()[:4] == [300.0, -100.0, 600.0, -200.0] + gate = us_farm_business_income_signal_gate(result) + assert gate.passed, gate.failures + + +def test_signal_gate_rejects_clipped_losses_and_asec_rent( + monkeypatch: pytest.MonkeyPatch, +) -> None: + clipped = _imputed_frame(monkeypatch) + person = clipped.table("person") + person[_OPERATIONS] = person[_OPERATIONS].clip(lower=0.0) + person[_RENT] = person[_RENT].clip(lower=0.0) + gate = us_farm_business_income_signal_gate(clipped) + assert not gate.passed + assert any("farm losses" in failure for failure in gate.failures) + + misplaced = _imputed_frame(monkeypatch) + asec_mask = misplaced.table("person")["person_support_channel"] == "asec" + misplaced.table("person").loc[asec_mask, _RENT] = 1.0 + gate = us_farm_business_income_signal_gate(misplaced) + assert not gate.passed + assert any("ASEC support carries nonzero" in failure for failure in gate.failures) + + +@requires_us +def test_policyengine_us_1_764_6_qbi_graph_reads_each_farm_leaf() -> None: + from policyengine_core.reforms import Reform + from policyengine_us import CountryTaxBenefitSystem, Simulation + + assert version("policyengine-us") == "1.764.6" + system = CountryTaxBenefitSystem() + for output in US_FARM_BUSINESS_INCOME_OUTPUT_COLUMNS: + variable = system.variables[output] + assert variable.is_input_variable() + assert variable.entity.key == "person" + assert str(variable.definition_period).lower() == "year" + assert variable.default_value == 0 + + reform = Reform.from_dict( + { + "gov.irs.deductions.qbi.income_definition": { + "2026-01-01.2026-12-31": [ + "self_employment_income", + "partnership_s_corp_income", + "rental_income", + "estate_income", + ] + } + }, + country_id="us", + ) + + def situation(output: str) -> dict[str, object]: + return { + "people": { + "adult": { + "age": {"2026": 40}, + output: {"2026": 10_000.0}, + } + }, + "tax_units": { + "tax_unit": { + "members": ["adult"], + "filing_status": {"2026": "SINGLE"}, + } + }, + "spm_units": {"spm_unit": {"members": ["adult"]}}, + "households": { + "household": { + "members": ["adult"], + "state_code": {"2026": "CA"}, + } + }, + } + + reformed_system = CountryTaxBenefitSystem(reform=(reform,)) + for output in US_FARM_BUSINESS_INCOME_OUTPUT_COLUMNS: + baseline = Simulation(tax_benefit_system=system, situation=situation(output)) + reformed = Simulation( + tax_benefit_system=reformed_system, + situation=situation(output), + ) + assert baseline.calculate("qualified_business_income", 2026)[0] > 9_000.0 + assert baseline.calculate("qualified_business_income_deduction", 2026)[ + 0 + ] == pytest.approx(400.0) + assert reformed.calculate("qualified_business_income", 2026)[0] == 0.0 + assert reformed.calculate("qualified_business_income_deduction", 2026)[0] == 0.0 diff --git a/packages/populace-build/tests/test_us_financial_assistance_exclusion.py b/packages/populace-build/tests/test_us_financial_assistance_exclusion.py new file mode 100644 index 00000000..1f455ec8 --- /dev/null +++ b/packages/populace-build/tests/test_us_financial_assistance_exclusion.py @@ -0,0 +1,185 @@ +"""Evidence contract for the irreducible financial-assistance exclusion.""" + +from __future__ import annotations + +import importlib.util +import json +from hashlib import sha256 +from importlib.metadata import version +from importlib.resources import files +from pathlib import Path + +import numpy as np +import pytest + +from populace.build.us_runtime.release_input_coverage import ( + load_release_input_coverage_manifest, +) + +ROOT = Path(__file__).resolve().parents[3] +policyengine_us_installed = importlib.util.find_spec("policyengine_us") is not None +requires_us = pytest.mark.skipif( + not policyengine_us_installed, + reason="requires the policyengine-us [us] extra (build environment)", +) + + +def _entry() -> dict[str, object]: + payload = json.loads( + files("populace.build.us").joinpath("ecps_parity_known_gaps.json").read_text() + ) + return payload["known_gaps"]["financial_assistance"] + + +def _sha256(path: Path) -> str: + digest = sha256() + with path.open("rb") as stream: + for chunk in iter(lambda: stream.read(1024 * 1024), b""): + digest.update(chunk) + return digest.hexdigest() + + +def test_exclusion_pins_exact_archived_person_source_and_qrf_dependency() -> None: + entry = _entry() + evidence = entry["evidence"] + + assert entry["reason"].startswith("SOURCE UNAVAILABILITY WITH EVIDENCE:") + assert evidence["classification"] == "source_unavailability" + assert evidence["retired_derivation"] == { + "repository_owner": "PolicyEngine", + "repository_name_parts": ["policyengine-", "us-data"], + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": [ + "policyengine_", + "us_data", + "datasets", + "cps", + "cps.py", + ], + "lines": "1493-1496", + } + assert evidence["required_person_source_declaration"]["lines"] == ("39-58,306-359") + assert evidence["puf_clone_qrf"]["lines"] == "140-194,639-745" + assert evidence["puf_clone_qrf"]["target_line"] == 166 + assert evidence["no_independent_puf_source"] == { + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": [ + "policyengine_", + "us_data", + "calibration", + "puf_impute.py", + ], + "tax_detail_target_lines": "90-149,158-198", + "financial_assistance_occurrences": 0, + } + assert evidence["required_person_columns"] == ["FIN_VAL"] + assert evidence["missing_person_context_columns"] == ["FIN_YN", "I_FINVAL"] + + +def test_exclusion_pins_all_sha_locked_hermetic_inputs_and_grain() -> None: + evidence = _entry()["evidence"] + evidence_hashes = { + item["filename"]: item["sha256"] for item in evidence["hermetic_inputs"] + } + build_summary = json.loads( + (ROOT / "experiments/build_j_recert/base_j.summary.json").read_text() + ) + recorded_hashes = { + Path(item["path"]).name: item["sha256"] + for item in build_summary["base_source"]["sources"] + } + + assert evidence_hashes == recorded_hashes + assert set(evidence_hashes) == { + "census_cps_2022.h5", + "census_cps_2023.h5", + "census_cps_2024.h5", + } + for item in evidence["hermetic_inputs"]: + assert item["missing_person_columns"] == ["FIN_VAL", "FIN_YN", "I_FINVAL"] + assert item["present_household_columns"] == ["HFINVAL", "HFIN_YN"] + assert item["present_family_columns"] == ["FFINVAL", "FINC_FIN"] + assert item["positive_multi_person_families"] > 0 + + dictionary = evidence["official_variable_dictionary"] + assert dictionary["url"] == ( + "https://api.census.gov/data/2024/cps/asec/mar/variables.html" + ) + assert dictionary["FIN_VAL"]["entity"] == "person" + assert dictionary["HFINVAL"]["entity"] == "household" + assert dictionary["FFINVAL"]["entity"] == "family" + substitutes = evidence["semantic_non_substitutes"] + assert "synthesize" in substitutes["rejection"] + assert "recipient person" in substitutes["HFINVAL"] + assert "recipient person" in substitutes["FFINVAL"] + + build_script = (ROOT / "experiments/build_j_recert/buildj_base.sh").read_text() + for year in (2022, 2023, 2024): + assert f'--asec-h5 {year}="$USD/census_cps_{year}.h5"' in build_script + assert "buildj_base.sh lines 65-69" in evidence["hermetic_build_contract"] + assert "base_j.summary.json lines 55-75" in evidence["hermetic_build_contract"] + asec_pool = ( + ROOT / "packages/populace-build/src/populace/build/us_runtime/asec_pool.py" + ).read_text() + assert 'pd.HDFStore(path, mode="r")' in asec_pool + + +@requires_us +def test_sha_locked_artifact_schemas_match_the_recorded_absence() -> None: + from policyengine_us.data import USSingleYearDataset + + evidence = _entry()["evidence"] + build_summary = json.loads( + (ROOT / "experiments/build_j_recert/base_j.summary.json").read_text() + ) + paths = { + Path(item["path"]).name: Path(item["path"]) + for item in build_summary["base_source"]["sources"] + } + if not all(path.is_file() for path in paths.values()): + pytest.skip("SHA-locked ASEC artifacts are not mounted in this environment") + + for item in evidence["hermetic_inputs"]: + path = paths[item["filename"]] + assert _sha256(path) == item["sha256"] + dataset = USSingleYearDataset(file_path=str(path)) + assert set(item["missing_person_columns"]).isdisjoint(dataset.person.columns) + assert set(item["present_household_columns"]) <= set(dataset.household.columns) + assert set(item["present_family_columns"]) <= set(dataset.family.columns) + positive_multi_person_families = int( + ((dataset.family["FFINVAL"] > 0) & (dataset.family["FPERSONS"] > 1)).sum() + ) + assert positive_multi_person_families == item["positive_multi_person_families"] + family_amount_by_household = dataset.family.groupby("FH_SEQ", sort=False)[ + "FFINVAL" + ].sum() + household_amount = dataset.household.set_index("H_SEQ")["HFINVAL"] + np.testing.assert_array_equal( + family_amount_by_household.reindex( + household_amount.index, + fill_value=0, + ).to_numpy(), + household_amount.to_numpy(), + ) + + +@requires_us +def test_policyengine_1_764_6_requires_a_person_year_input() -> None: + from policyengine_us import CountryTaxBenefitSystem + + assert version("policyengine-us") == "1.764.6" + variable = CountryTaxBenefitSystem().variables["financial_assistance"] + assert variable.is_input_variable() + assert variable.entity.key == "person" + assert str(variable.definition_period).lower() == "year" + assert variable.documentation == ( + "Cash financial assistance from outside the household." + ) + + +def test_generated_release_manifest_preserves_the_evidenced_exclusion() -> None: + reason = load_release_input_coverage_manifest().reviewed_exclusions[ + "financial_assistance" + ] + assert reason == _entry()["reason"] + assert reason.startswith("SOURCE UNAVAILABILITY WITH EVIDENCE:") diff --git a/packages/populace-build/tests/test_us_fiscal_refresh_builder.py b/packages/populace-build/tests/test_us_fiscal_refresh_builder.py index 15bd51ca..0bb64847 100644 --- a/packages/populace-build/tests/test_us_fiscal_refresh_builder.py +++ b/packages/populace-build/tests/test_us_fiscal_refresh_builder.py @@ -113,8 +113,11 @@ def test__given_target_frame_checkpoint__then_builder_round_trips_frame( seed=0, target_period=builder.PERIOD, target_registry_version="registry-sha", + weeks_unemployed_source_sha256="weeks-source-sha", congressional_district_vintage_crosswalk_sha256="crosswalk-sha", ) + assert identity["materializer_version"] == 6 + assert identity["weeks_unemployed_source_sha256"] == "weeks-source-sha" path = tmp_path / "target_frame_checkpoint.h5" payload = builder._write_target_frame_checkpoint( @@ -161,6 +164,381 @@ def test__given_target_frame_checkpoint__then_builder_round_trips_frame( ) +def test_ssi_candidate_amount_uses_december_person_values() -> None: + builder = _load_builder_module() + + class FakeSimulation: + def calculate(self, variable, *, period, map_to): + assert variable == "uncapped_ssi" + assert period == "2024-12" + assert map_to == "person" + return np.asarray([0.0, 125.0, -2.0]) + + values = builder._ssi_person_uncapped_amount( + SimpleNamespace(), + simulation=FakeSimulation(), + ) + + np.testing.assert_array_equal(values, np.asarray([0.0, 125.0, -2.0])) + + +def test_ssi_reconciliation_replays_dependencies_before_refit_and_only_diagnoses_final( + monkeypatch, + small_frame, +) -> None: + builder = _load_builder_module() + calls: list[str] = [] + reporter_ids = frozenset({"asec-reporter"}) + initial_result = SimpleNamespace( + weights=np.asarray([1_100.0, 1_900.0]), + initial_weights=np.asarray([1_000.0, 2_000.0]), + ) + first_reconciled_result = SimpleNamespace( + weights=np.asarray([1_200.0, 1_800.0]), + initial_weights=np.asarray([1_000.0, 2_000.0]), + final_loss=0.5, + ) + reconciled_result = SimpleNamespace( + weights=np.asarray([1_250.0, 1_750.0]), + initial_weights=np.asarray([1_000.0, 2_000.0]), + final_loss=0.25, + ) + + monkeypatch.setattr(builder, "_assert_no_formula_owned_columns", lambda frame: None) + monkeypatch.setattr( + builder, + "us_ssi_take_up_reporter_source_ids", + lambda frame: reporter_ids, + ) + + uncapped_calls = 0 + + def fake_uncapped(frame, *, maximum_microsim_batch_size): + nonlocal uncapped_calls + uncapped_calls += 1 + calls.append(f"uncapped_{uncapped_calls}") + return np.zeros(frame.n("person"), dtype=np.float64) + + monkeypatch.setattr(builder, "_ssi_person_uncapped_amount", fake_uncapped) + + assignment_calls = 0 + + def fake_assign(frame, *, uncapped_ssi, seed, reporter_source_ids): + nonlocal assignment_calls + assignment_calls += 1 + calls.append("ssi_assign") + assert reporter_source_ids == reporter_ids + return frame, {"phase": "assigned"} + + monkeypatch.setattr(builder, "with_us_ssi_take_up", fake_assign) + + def fake_aca(frame, target_specs, *, seed, maximum_microsim_batch_size): + calls.append("aca") + return frame + + monkeypatch.setattr(builder, "_with_aca_marketplace_source_outputs", fake_aca) + monkeypatch.setattr( + builder, + "_health_input_signal_gate", + lambda frame: builder.GateResult(name="health", passed=True), + ) + + def fake_medicaid( + frame, + target_specs, + *, + seed, + substitutions, + maximum_microsim_batch_size, + ): + calls.append("medicaid") + return frame, {"phase": "medicaid"} + + monkeypatch.setattr(builder, "_with_medicaid_take_up_outputs", fake_medicaid) + + def fake_final_medicaid(*args, **kwargs): + calls.append("medicaid_diagnose") + return {"phase": f"final_medicaid_{calibrate_calls}"} + + monkeypatch.setattr( + builder, + "_medicaid_diagnostics_for_existing_output", + fake_final_medicaid, + ) + + def fake_medicaid_gate(diagnostics): + passed = diagnostics.get("phase") != "final_medicaid_1" + return builder.GateResult( + name="medicaid", + passed=passed, + failures=() if passed else ("first returned weights drifted",), + ) + + monkeypatch.setattr( + builder, + "us_medicaid_take_up_gate", + fake_medicaid_gate, + ) + + def fake_other_health( + frame, + *, + seed, + time_period, + maximum_microsim_batch_size, + ): + calls.append("other_health") + return frame + + monkeypatch.setattr( + builder, + "with_us_other_health_insurance_inputs", + fake_other_health, + ) + monkeypatch.setattr( + builder, + "us_other_health_insurance_signal_gate", + lambda frame: builder.GateResult(name="other_health", passed=True), + ) + + class FakeRegistry: + specs = (SimpleNamespace(),) + + def to_target_set(self): + return "targets" + + registry = FakeRegistry() + + def fake_materialize(frame, target_specs, **kwargs): + calls.append("materialize") + np.testing.assert_array_equal( + frame.weights_for("household").values, + initial_result.initial_weights, + ) + return frame, registry, {"declared_targets": 1} + + monkeypatch.setattr(builder, "_materialize_target_frame", fake_materialize) + monkeypatch.setattr( + builder, + "_fiscal_target_loss_weights", + lambda active_registry: np.ones(1), + ) + + calibrate_calls = 0 + + def fake_calibrate(frame, targets, **kwargs): + nonlocal calibrate_calls + calibrate_calls += 1 + calls.append("calibrate") + np.testing.assert_array_equal( + kwargs["warm_start_weights"], + ( + initial_result.weights + if calibrate_calls == 1 + else first_reconciled_result.weights + ), + ) + return first_reconciled_result if calibrate_calls == 1 else reconciled_result + + monkeypatch.setattr(builder, "calibrate", fake_calibrate) + + def fake_final_diagnostics( + frame, + *, + uncapped_ssi, + seed, + reporter_source_ids, + ): + calls.append("ssi_diagnose") + assert reporter_source_ids == reporter_ids + return {"phase": f"final_{calibrate_calls}"} + + monkeypatch.setattr( + builder, + "us_ssi_take_up_diagnostics", + fake_final_diagnostics, + ) + + def fake_ssi_gate(diagnostics): + return builder.GateResult(name="ssi", passed=True, details=diagnostics) + + monkeypatch.setattr(builder, "us_ssi_take_up_gate", fake_ssi_gate) + + reconciliation = builder._reconcile_ssi_take_up_and_refit( + small_frame, + initial_result, + (), + dense_default_dataset=True, + seed=3, + epochs=5, + learning_rate=0.01, + max_weight_ratio=10.0, + l2_lambda=0.0, + target_loss_cap=1.0, + ) + + assert assignment_calls == 2 + assert calls == [ + "uncapped_1", + "ssi_assign", + "aca", + "medicaid", + "other_health", + "materialize", + "calibrate", + "uncapped_2", + "ssi_diagnose", + "medicaid_diagnose", + "uncapped_3", + "ssi_assign", + "aca", + "medicaid", + "other_health", + "materialize", + "calibrate", + "uncapped_4", + "ssi_diagnose", + "medicaid_diagnose", + ] + np.testing.assert_array_equal( + reconciliation.export_frame.weights_for("household").values, + reconciled_result.weights, + ) + assert reconciliation.ssi_diagnostics == {"phase": "final_2"} + assert reconciliation.compilation["ssi_take_up_reconciliation"]["passes"] == 2 + + +def test_ssi_reconciliation_fails_closed_after_bounded_weight_drift( + monkeypatch, + small_frame, +) -> None: + builder = _load_builder_module() + counts = {"assign": 0, "calibrate": 0, "diagnose": 0} + initial_result = SimpleNamespace( + weights=np.asarray([1_100.0, 1_900.0]), + initial_weights=np.asarray([1_000.0, 2_000.0]), + ) + returned_result = SimpleNamespace( + weights=np.asarray([1_100.0, 1_900.0]), + initial_weights=np.asarray([1_000.0, 2_000.0]), + final_loss=1.0, + ) + + monkeypatch.setattr(builder, "_assert_no_formula_owned_columns", lambda frame: None) + monkeypatch.setattr( + builder, + "us_ssi_take_up_reporter_source_ids", + lambda frame: frozenset({"reporter"}), + ) + monkeypatch.setattr( + builder, + "_ssi_person_uncapped_amount", + lambda frame, **kwargs: np.zeros(frame.n("person")), + ) + + def fake_assign(frame, **kwargs): + counts["assign"] += 1 + return frame, {"stage": True} + + monkeypatch.setattr(builder, "with_us_ssi_take_up", fake_assign) + monkeypatch.setattr( + builder, + "_with_aca_marketplace_source_outputs", + lambda frame, *args, **kwargs: frame, + ) + + def passing_gate(name): + return builder.GateResult(name=name, passed=True) + + monkeypatch.setattr( + builder, + "_health_input_signal_gate", + lambda frame: passing_gate("health"), + ) + monkeypatch.setattr( + builder, + "_with_medicaid_take_up_outputs", + lambda frame, *args, **kwargs: (frame, {"stage": True}), + ) + monkeypatch.setattr( + builder, + "_medicaid_diagnostics_for_existing_output", + lambda *args, **kwargs: {"final": True}, + ) + monkeypatch.setattr( + builder, + "us_medicaid_take_up_gate", + lambda diagnostics: passing_gate("medicaid"), + ) + monkeypatch.setattr( + builder, + "with_us_other_health_insurance_inputs", + lambda frame, **kwargs: frame, + ) + monkeypatch.setattr( + builder, + "us_other_health_insurance_signal_gate", + lambda frame: passing_gate("other_health"), + ) + + class FakeRegistry: + specs = (SimpleNamespace(),) + + def to_target_set(self): + return "targets" + + monkeypatch.setattr( + builder, + "_materialize_target_frame", + lambda frame, *args, **kwargs: (frame, FakeRegistry(), {}), + ) + monkeypatch.setattr( + builder, + "_fiscal_target_loss_weights", + lambda registry: np.ones(1), + ) + + def fake_calibrate(*args, **kwargs): + counts["calibrate"] += 1 + return returned_result + + monkeypatch.setattr(builder, "calibrate", fake_calibrate) + + def fake_diagnostics(*args, **kwargs): + counts["diagnose"] += 1 + return {"final": True} + + monkeypatch.setattr(builder, "us_ssi_take_up_diagnostics", fake_diagnostics) + + def fake_ssi_gate(diagnostics): + if diagnostics.get("stage"): + return passing_gate("ssi") + return builder.GateResult( + name="ssi", + passed=False, + failures=("returned weights drifted",), + ) + + monkeypatch.setattr(builder, "us_ssi_take_up_gate", fake_ssi_gate) + + with pytest.raises(RuntimeError, match="after 2 pass"): + builder._reconcile_ssi_take_up_and_refit( + small_frame, + initial_result, + (), + dense_default_dataset=True, + seed=0, + epochs=1, + learning_rate=0.01, + max_weight_ratio=10.0, + l2_lambda=0.0, + target_loss_cap=1.0, + max_passes=2, + ) + + assert counts == {"assign": 2, "calibrate": 2, "diagnose": 2} + + def test__given_stale_target_frame_checkpoint__then_builder_ignores_it( tmp_path, small_frame, @@ -172,11 +550,12 @@ def test__given_stale_target_frame_checkpoint__then_builder_ignores_it( seed=0, target_period=builder.PERIOD, target_registry_version="registry-sha", + weeks_unemployed_source_sha256="weeks-source-sha", congressional_district_vintage_crosswalk_sha256="crosswalk-sha", ) stale_identity = { **fresh_identity, - "target_registry_version": "old-registry-sha", + "weeks_unemployed_source_sha256": "old-weeks-source-sha", } path = tmp_path / "target_frame_checkpoint.h5" builder._write_target_frame_checkpoint( @@ -215,6 +594,7 @@ def test__given_matching_target_frame_checkpoint__then_builder_skips_materializa seed=0, target_period=builder.PERIOD, target_registry_version=registry.version, + weeks_unemployed_source_sha256="weeks-source-sha", congressional_district_vintage_crosswalk_sha256=None, ) @@ -526,6 +906,127 @@ def test_export_target_audit_is_opt_in(monkeypatch) -> None: assert args.audit_export_targets +def test_sipp_tip_donor_override_parses(monkeypatch) -> None: + builder = _load_builder_module() + monkeypatch.setattr( + sys, + "argv", + [ + "build_us_fiscal_refresh_release.py", + "--ledger-facts", + "facts.jsonl", + "--out", + "release", + "--sipp-tip-donor", + "pu2023_slim.csv", + ], + ) + + args = builder._parse_args() + + assert args.sipp_tip_donor == Path("pu2023_slim.csv") + + +def test_weeks_unemployed_source_override_parses(monkeypatch) -> None: + builder = _load_builder_module() + monkeypatch.setattr( + sys, + "argv", + [ + "build_us_fiscal_refresh_release.py", + "--ledger-facts", + "facts.jsonl", + "--out", + "release", + "--asec-2023-weeks-unemployed-source", + "asecpub23csv.zip", + ], + ) + + args = builder._parse_args() + + assert args.asec_2023_weeks_unemployed_source == Path("asecpub23csv.zip") + + +def test_frozen_support_selection_is_followed_by_weeks_unemployed_regate() -> None: + builder = _load_builder_module() + source = Path(builder.__file__).read_text(encoding="utf-8") + + selection = source.index("base_frame, selection_report = select_frozen_support(") + regate = source.index( + "post_selection_weeks_unemployed_gate = us_weeks_unemployed_signal_gate(" + ) + mass_repair = source.index( + "base_frame, base_population_repair = _with_base_population_mass_repair(" + ) + + assert selection < regate < mass_repair + assert "Post-selection weeks-unemployed input signal failed" in source + + +def test_sipp_vehicle_donor_override_parses(monkeypatch) -> None: + builder = _load_builder_module() + monkeypatch.setattr( + sys, + "argv", + [ + "build_us_fiscal_refresh_release.py", + "--ledger-facts", + "facts.jsonl", + "--out", + "release", + "--sipp-vehicle-donor", + "pu2023.csv", + ], + ) + + args = builder._parse_args() + + assert args.sipp_vehicle_donor == Path("pu2023.csv") + + +def test_scf_full_extract_override_parses(monkeypatch) -> None: + builder = _load_builder_module() + monkeypatch.setattr( + sys, + "argv", + [ + "build_us_fiscal_refresh_release.py", + "--ledger-facts", + "facts.jsonl", + "--out", + "release", + "--scf-full-extract", + "p22i6.dta", + ], + ) + + args = builder._parse_args() + + assert args.scf_full_extract == Path("p22i6.dta") + + +def test_org_wages_donor_override_parses(monkeypatch) -> None: + builder = _load_builder_module() + monkeypatch.setattr( + sys, + "argv", + [ + "build_us_fiscal_refresh_release.py", + "--ledger-facts", + "facts.jsonl", + "--out", + "release", + "--org-wages-donor", + "census_cps_org_2024_wages.csv.gz", + ], + ) + + args = builder._parse_args() + + assert args.org_wages_donor == Path("census_cps_org_2024_wages.csv.gz") + + def test_cd_targets_require_vintage_crosswalk(monkeypatch) -> None: builder = _load_builder_module() @@ -1698,6 +2199,7 @@ def test_main_writes_diagnostics_before_post_calibration_gate_failure( builder = _load_builder_module() release_id = "populace-us-2024-gate-failure-test" base_h5 = tmp_path / "base.h5" + weeks_source = tmp_path / "asecpub23csv.zip" facts = tmp_path / "facts.jsonl" out = tmp_path / "out" base_h5.write_bytes(b"h5") @@ -1722,7 +2224,10 @@ def test_main_writes_diagnostics_before_post_calibration_gate_failure( weight_entity="household", selection=SimpleNamespace(n_nonzero=2, final_loss=1.5), ) - captured: dict[str, object] = {} + captured: dict[str, object] = { + "health_stage_events": [], + "source_stage_events": [], + } class FakeFrame: def n(self, entity): @@ -1742,15 +2247,27 @@ def n(self, entity): str(out), "--release-id", release_id, + "--asec-2023-weeks-unemployed-source", + str(weeks_source), "--no-target-frame-checkpoint", "--no-staging", ], ) monkeypatch.setattr(builder, "_git_dirty", lambda: False) - monkeypatch.setattr(builder, "_sha256", lambda path: "base-sha") + monkeypatch.setattr( + builder, + "_sha256", + lambda path: "weeks-source-sha" if Path(path) == weeks_source else "base-sha", + ) monkeypatch.setattr(builder, "_git_output", lambda *args: "commit") # The consistency/contract preflights hit the installed policyengine-us # (absent in CI); this test pins diagnostics ordering, not engine metadata. + monkeypatch.setattr( + builder, "assert_validation_leaf_registry_current", lambda: None + ) + monkeypatch.setattr( + builder, "assert_release_input_coverage_manifest_current", lambda: None + ) monkeypatch.setattr(builder, "assert_take_up_contract_current", lambda: None) monkeypatch.setattr(builder, "assert_take_up_treatments_consistent", lambda: None) monkeypatch.setattr( @@ -1781,16 +2298,80 @@ def n(self, entity): details={"checked": True}, ), ) - monkeypatch.setattr(builder, "_load_frame", lambda path: FakeFrame()) + + def fake_load_frame(path): + captured["source_stage_events"].append("load_frame") + return FakeFrame() + + monkeypatch.setattr(builder, "_load_frame", fake_load_frame) + + def fake_load_weeks_unemployed_source(path, **kwargs): + captured["source_stage_events"].append("load_weeks_source") + captured["weeks_unemployed_source_path"] = Path(path) + captured["weeks_unemployed_source_load_kwargs"] = kwargs + return pd.DataFrame({"LKWEEKS": [0, 12]}) + + def fake_with_weeks_unemployed( + frame, + *, + seed, + time_period, + asec_2023_source, + ): + captured["source_stage_events"].append("weeks_stage") + captured["weeks_unemployed_stage_seed"] = seed + captured["weeks_unemployed_stage_period"] = time_period + captured["weeks_unemployed_stage_source"] = asec_2023_source + return frame + + def fake_weeks_unemployed_signal_gate(frame): + captured["source_stage_events"].append("weeks_gate") + captured["weeks_unemployed_gate_called"] = True + return builder.GateResult( + name="weeks_unemployed_signal", + passed=True, + details={"checked": True}, + ) + + monkeypatch.setattr( + builder, + "load_asec_2023_weeks_unemployed_source", + fake_load_weeks_unemployed_source, + ) + monkeypatch.setattr( + builder, + "with_us_weeks_unemployed", + fake_with_weeks_unemployed, + ) + monkeypatch.setattr( + builder, + "us_weeks_unemployed_signal_gate", + fake_weeks_unemployed_signal_gate, + ) + + def fake_ssi_reporter_source_ids(frame): + captured["source_stage_events"].append("ssi_reporters") + return frozenset({"asec-reporter"}) + + monkeypatch.setattr( + builder, + "us_ssi_take_up_reporter_source_ids", + fake_ssi_reporter_source_ids, + ) repair_payload = { "method": "rescale_household_weights_to_census_person_population", "applied": True, "factor": 2.0, } + + def fake_base_population_mass_repair(frame): + captured["source_stage_events"].append("population_repair") + return frame, repair_payload + monkeypatch.setattr( builder, "_with_base_population_mass_repair", - lambda frame: (frame, repair_payload), + fake_base_population_mass_repair, ) ss_repair_payload = { "method": "rescale_social_security_component_leaves_to_ssa_targets", @@ -1810,6 +2391,211 @@ def n(self, entity): details={"checked": True, "mass_repair": mass_repair}, ), ) + monkeypatch.setattr( + builder, + "with_us_qbi_input_reconciliation", + lambda frame: frame, + ) + monkeypatch.setattr( + builder, + "us_qbi_inputs_signal_gate", + lambda frame: builder.GateResult( + name="qbi_inputs_signal", + passed=True, + details={"checked": True}, + ), + ) + + def fake_farm_business_income_signal_gate(frame): + captured["farm_business_income_gate_called"] = True + return builder.GateResult( + name="farm_business_income_signal", + passed=True, + details={"checked": True}, + ) + + monkeypatch.setattr( + builder, + "us_farm_business_income_signal_gate", + fake_farm_business_income_signal_gate, + ) + monkeypatch.setattr( + builder, + "us_domestic_production_ald_signal_gate", + lambda frame: builder.GateResult( + name="domestic_production_ald_signal", + passed=True, + details={"checked": True}, + ), + ) + + def fake_child_support_signal_gate(frame): + captured["child_support_gate_called"] = True + return builder.GateResult( + name="child_support_inputs_signal", + passed=True, + details={"checked": True}, + ) + + monkeypatch.setattr( + builder, + "us_child_support_signal_gate", + fake_child_support_signal_gate, + ) + + def fake_disability_benefits_signal_gate(frame): + captured["disability_benefits_gate_called"] = True + return builder.GateResult( + name="disability_benefits_signal", + passed=True, + details={"checked": True}, + ) + + monkeypatch.setattr( + builder, + "us_disability_benefits_signal_gate", + fake_disability_benefits_signal_gate, + ) + monkeypatch.setattr( + builder, + "us_workers_compensation_signal_gate", + lambda frame: builder.GateResult( + name="workers_compensation_signal", + passed=True, + details={"checked": True}, + ), + ) + + def fake_educator_expense_signal_gate(frame): + captured["educator_expense_gate_called"] = True + return builder.GateResult( + name="educator_expense_signal", + passed=True, + details={"checked": True}, + ) + + monkeypatch.setattr( + builder, + "us_educator_expense_signal_gate", + fake_educator_expense_signal_gate, + ) + + def fake_form_4952_election_signal_gate(frame): + captured["form_4952_election_gate_called"] = True + return builder.GateResult( + name="form_4952_election_signal", + passed=True, + details={"checked": True}, + ) + + monkeypatch.setattr( + builder, + "us_form_4952_election_signal_gate", + fake_form_4952_election_signal_gate, + ) + + def fake_salt_refund_income_signal_gate(frame): + captured["salt_refund_income_gate_called"] = True + return builder.GateResult( + name="salt_refund_income_signal", + passed=True, + details={"checked": True}, + ) + + monkeypatch.setattr( + builder, + "us_salt_refund_income_signal_gate", + fake_salt_refund_income_signal_gate, + ) + + def fake_capital_gain_details_signal_gate(frame): + captured["capital_gain_details_gate_called"] = True + return builder.GateResult( + name="capital_gain_details_signal", + passed=True, + details={"checked": True}, + ) + + monkeypatch.setattr( + builder, + "us_capital_gain_details_signal_gate", + fake_capital_gain_details_signal_gate, + ) + monkeypatch.setattr( + builder, + "with_us_childcare_inputs", + lambda frame, *, seed, time_period, allow_existing_without_source: frame, + ) + monkeypatch.setattr( + builder, + "us_childcare_signal_gate", + lambda frame: builder.GateResult( + name="childcare_inputs_signal", + passed=True, + details={"checked": True}, + ), + ) + + monkeypatch.setattr( + builder, + "with_us_energy_subsidy_input", + lambda frame, *, seed, time_period, allow_existing_without_source: frame, + ) + + def fake_energy_subsidy_signal_gate(frame): + captured["energy_subsidy_gate_called"] = True + return builder.GateResult( + name="energy_subsidy_signal", + passed=True, + details={"checked": True}, + ) + + monkeypatch.setattr( + builder, + "us_energy_subsidy_signal_gate", + fake_energy_subsidy_signal_gate, + ) + monkeypatch.setattr( + builder, + "us_alimony_signal_gate", + lambda frame: builder.GateResult( + name="alimony_inputs_signal", + passed=True, + details={"checked": True}, + ), + ) + monkeypatch.setattr( + builder, + "us_casualty_loss_signal_gate", + lambda frame: builder.GateResult( + name="casualty_loss_signal", + passed=True, + details={"checked": True}, + ), + ) + monkeypatch.setattr( + builder, + "us_misc_itemized_signal_gate", + lambda frame: builder.GateResult( + name="misc_itemized_signal", + passed=True, + details={"checked": True}, + ), + ) + monkeypatch.setattr( + builder, + "with_us_retirement_contribution_inputs", + lambda frame, *, seed, time_period: frame, + ) + monkeypatch.setattr( + builder, + "us_retirement_contributions_signal_gate", + lambda frame: builder.GateResult( + name="retirement_contributions_signal", + passed=True, + details={"checked": True}, + ), + ) monkeypatch.setattr( builder, "with_us_immigration_inputs", @@ -1836,17 +2622,42 @@ def n(self, entity): ) monkeypatch.setattr( builder, - "with_us_snap_take_up_inputs", + "with_us_snap_take_up_inputs", + lambda frame, *, seed, time_period: frame, + ) + monkeypatch.setattr( + builder, + "with_us_eligibility_inputs", + lambda frame, *, seed, time_period: frame, + ) + monkeypatch.setattr( + builder, + "with_us_relationship_inputs", + lambda frame, *, seed, time_period: frame, + ) + monkeypatch.setattr( + builder, + "with_us_medicare_take_up_input", + lambda frame, *, seed, time_period: frame, + ) + monkeypatch.setattr( + builder, + "with_us_retirement_distribution_inputs", + lambda frame, *, seed, time_period: frame, + ) + monkeypatch.setattr( + builder, + "with_us_education_inputs", lambda frame, *, seed, time_period: frame, ) monkeypatch.setattr( builder, - "with_us_eligibility_inputs", + "with_us_pregnancy_inputs", lambda frame, *, seed, time_period: frame, ) monkeypatch.setattr( builder, - "with_us_pregnancy_inputs", + "with_us_wic_claim_input", lambda frame, *, seed, time_period: frame, ) monkeypatch.setattr( @@ -1900,6 +2711,60 @@ def n(self, entity): details={"checked": True}, ), ) + monkeypatch.setattr( + builder, + "us_relationship_inputs_signal_gate", + lambda frame: builder.GateResult( + name="relationship_inputs_signal", + passed=True, + details={"checked": True}, + ), + ) + monkeypatch.setattr( + builder, + "us_medicare_take_up_signal_gate", + lambda frame: builder.GateResult( + name="medicare_take_up_input_signal", + passed=True, + details={"checked": True}, + ), + ) + monkeypatch.setattr( + builder, + "us_prior_year_income_signal_gate", + lambda frame: builder.GateResult( + name="prior_year_income_signal", + passed=True, + details={"checked": True}, + ), + ) + monkeypatch.setattr( + builder, + "us_housing_inputs_signal_gate", + lambda frame: builder.GateResult( + name="housing_inputs_signal", + passed=True, + details={"checked": True}, + ), + ) + monkeypatch.setattr( + builder, + "us_retirement_distributions_signal_gate", + lambda frame: builder.GateResult( + name="retirement_distributions_signal", + passed=True, + details={"checked": True}, + ), + ) + monkeypatch.setattr( + builder, + "us_education_inputs_signal_gate", + lambda frame: builder.GateResult( + name="education_inputs_signal", + passed=True, + details={"checked": True}, + ), + ) monkeypatch.setattr( builder, "us_pregnancy_signal_gate", @@ -1909,6 +2774,15 @@ def n(self, entity): details={"checked": True}, ), ) + monkeypatch.setattr( + builder, + "us_wic_claim_signal_gate", + lambda frame: builder.GateResult( + name="wic_claim_signal", + passed=True, + details={"checked": True}, + ), + ) monkeypatch.setattr( builder, "us_snap_discretionary_exemption_signal_gate", @@ -1942,6 +2816,319 @@ def n(self, entity): details={"checked": True}, ), ) + monkeypatch.setattr( + builder, + "fetch_scf_2022_full_extract", + lambda *args, **kwargs: Path("p22i6.dta"), + ) + + def fake_load_scf_auto_loan_donor(summary_path, full_path): + captured["scf_auto_summary_path"] = summary_path + captured["scf_auto_full_path"] = full_path + return pd.DataFrame() + + def fake_with_scf_auto_loan_inputs( + frame, *, seed, time_period, scf_auto_loan_donor + ): + captured["scf_auto_stage_called"] = True + return frame + + monkeypatch.setattr( + builder, + "load_scf_2022_auto_loan_donor", + fake_load_scf_auto_loan_donor, + ) + monkeypatch.setattr( + builder, + "with_us_scf_auto_loan_inputs", + fake_with_scf_auto_loan_inputs, + ) + monkeypatch.setattr( + builder, + "us_scf_auto_loans_signal_gate", + lambda frame: builder.GateResult( + name="scf_auto_loans_signal", + passed=True, + details={"checked": True}, + ), + ) + monkeypatch.setattr( + builder, + "fetch_sipp_2023_vehicle_donor", + lambda *args, **kwargs: Path("pu2023.csv"), + ) + + def fake_load_sipp_vehicle_donor( + path, + *, + expected_sha256=None, + expected_size_bytes=None, + ): + captured["sipp_vehicle_donor_path"] = path + captured["sipp_vehicle_donor_sha256"] = expected_sha256 + captured["sipp_vehicle_donor_size_bytes"] = expected_size_bytes + return pd.DataFrame() + + def fake_with_sipp_vehicle_inputs(frame, *, seed, time_period, sipp_donor): + captured["sipp_vehicle_stage_called"] = True + captured["sipp_vehicle_seed"] = seed + return frame + + def fake_sipp_vehicles_signal_gate(frame): + captured["sipp_vehicle_gate_called"] = True + return builder.GateResult( + name="sipp_vehicles_signal", + passed=True, + details={"checked": True}, + ) + + monkeypatch.setattr( + builder, + "load_sipp_2023_vehicle_donor", + fake_load_sipp_vehicle_donor, + ) + monkeypatch.setattr( + builder, + "with_us_sipp_vehicle_inputs", + fake_with_sipp_vehicle_inputs, + ) + monkeypatch.setattr( + builder, + "us_sipp_vehicles_signal_gate", + fake_sipp_vehicles_signal_gate, + ) + + def fake_load_ssi_disability_donor( + path, + *, + expected_sha256=None, + expected_size_bytes=None, + time_period=2024, + ): + captured["ssi_disability_donor_path"] = path + captured["ssi_disability_donor_sha256"] = expected_sha256 + captured["ssi_disability_donor_size_bytes"] = expected_size_bytes + captured["ssi_disability_donor_period"] = time_period + return pd.DataFrame() + + def fake_with_ssi_disability_criteria(frame, *, seed, time_period, sipp_donor): + captured["ssi_disability_stage_called"] = True + captured["ssi_disability_seed"] = seed + return frame + + def fake_ssi_disability_signal_gate(frame): + captured["ssi_disability_gate_called"] = True + return builder.GateResult( + name="ssi_disability_criteria_signal", + passed=True, + details={"checked": True}, + ) + + monkeypatch.setattr( + builder, + "load_sipp_2023_ssi_disability_donor", + fake_load_ssi_disability_donor, + ) + monkeypatch.setattr( + builder, + "with_us_ssi_disability_criteria", + fake_with_ssi_disability_criteria, + ) + monkeypatch.setattr( + builder, + "us_ssi_disability_criteria_signal_gate", + fake_ssi_disability_signal_gate, + ) + + def fake_load_sipp_head_start_donor( + path, + *, + expected_sha256=None, + expected_size_bytes=None, + ): + captured["sipp_head_start_donor_path"] = path + captured["sipp_head_start_donor_sha256"] = expected_sha256 + captured["sipp_head_start_donor_size_bytes"] = expected_size_bytes + return pd.DataFrame() + + def fake_with_sipp_head_start_input( + frame, + *, + seed, + time_period, + sipp_donor, + ): + captured["sipp_head_start_stage_called"] = True + captured["sipp_head_start_seed"] = seed + captured["sipp_head_start_period"] = time_period + captured["sipp_head_start_donor"] = sipp_donor + return frame + + def fake_sipp_head_start_signal_gate(frame): + captured["sipp_head_start_gate_called"] = True + return builder.GateResult( + name="sipp_head_start_signal", + passed=True, + details={"checked": True}, + ) + + monkeypatch.setattr( + builder, + "load_sipp_2023_head_start_donor", + fake_load_sipp_head_start_donor, + ) + monkeypatch.setattr( + builder, + "with_us_sipp_head_start_input", + fake_with_sipp_head_start_input, + ) + monkeypatch.setattr( + builder, + "us_sipp_head_start_signal_gate", + fake_sipp_head_start_signal_gate, + ) + + def fake_ssi_uncapped_amount( + frame, + *, + simulation=None, + maximum_microsim_batch_size=None, + ): + captured["ssi_uncapped_stage_called"] = True + captured["ssi_uncapped_batch_size"] = maximum_microsim_batch_size + return np.zeros(4, dtype=np.float64) + + def fake_with_ssi_take_up( + frame, + *, + uncapped_ssi, + seed, + reporter_source_ids, + ): + captured["ssi_take_up_stage_called"] = True + captured["ssi_take_up_seed"] = seed + captured["ssi_take_up_uncapped"] = np.asarray(uncapped_ssi) + captured["ssi_reporter_source_ids"] = reporter_source_ids + return frame, {"checked": True} + + def fake_ssi_take_up_gate(diagnostics): + captured["ssi_take_up_gate_called"] = True + captured["ssi_take_up_gate_diagnostics"] = diagnostics + return builder.GateResult( + name="ssi_take_up", + passed=True, + details=diagnostics, + ) + + monkeypatch.setattr( + builder, + "_ssi_person_uncapped_amount", + fake_ssi_uncapped_amount, + ) + monkeypatch.setattr(builder, "with_us_ssi_take_up", fake_with_ssi_take_up) + monkeypatch.setattr(builder, "us_ssi_take_up_gate", fake_ssi_take_up_gate) + + def fake_load_voluntary_filing_donor( + path, + *, + expected_sha256=None, + expected_size_bytes=None, + ): + captured["voluntary_filing_donor_path"] = path + captured["voluntary_filing_donor_sha256"] = expected_sha256 + captured["voluntary_filing_donor_size_bytes"] = expected_size_bytes + return pd.DataFrame() + + def fake_with_voluntary_filing_input(frame, *, seed, time_period, sipp_donor): + captured["voluntary_filing_stage_called"] = True + captured["voluntary_filing_seed"] = seed + return frame + + def fake_voluntary_filing_signal_gate(frame): + captured["voluntary_filing_gate_called"] = True + return builder.GateResult( + name="voluntary_filing_signal", + passed=True, + details={"checked": True}, + ) + + monkeypatch.setattr( + builder, + "load_sipp_2023_voluntary_filing_donor", + fake_load_voluntary_filing_donor, + ) + monkeypatch.setattr( + builder, + "with_us_voluntary_filing_input", + fake_with_voluntary_filing_input, + ) + monkeypatch.setattr( + builder, + "us_voluntary_filing_signal_gate", + fake_voluntary_filing_signal_gate, + ) + monkeypatch.setattr( + builder, + "fetch_sipp_2023_tip_donor", + lambda *args, **kwargs: Path("pu2023_slim.csv"), + ) + + def fake_load_sipp_tip_donor(path, *, expected_sha256=None): + captured["sipp_tip_donor_path"] = path + captured["sipp_tip_donor_sha256"] = expected_sha256 + return pd.DataFrame() + + def fake_with_sipp_tip_inputs(frame, *, seed, time_period, sipp_donor): + captured["sipp_tip_stage_called"] = True + return frame + + def fake_sipp_tips_signal_gate(frame): + captured["sipp_tip_gate_called"] = True + return builder.GateResult( + name="sipp_tips_signal", + passed=True, + details={"checked": True}, + ) + + monkeypatch.setattr( + builder, + "load_sipp_2023_tip_donor", + fake_load_sipp_tip_donor, + ) + monkeypatch.setattr( + builder, + "with_us_sipp_tip_inputs", + fake_with_sipp_tip_inputs, + ) + monkeypatch.setattr( + builder, + "us_sipp_tips_signal_gate", + fake_sipp_tips_signal_gate, + ) + monkeypatch.setattr( + builder, + "fetch_org_2024_donor", + lambda *args, **kwargs: Path("census_cps_org_2024_wages.csv.gz"), + ) + + def fake_load_org_donor(path, *, expected_content_sha256=None): + captured["org_donor_path"] = path + captured["org_donor_sha256"] = expected_content_sha256 + return pd.DataFrame() + + def fake_with_org_inputs(frame, *, seed, time_period, org_donor): + captured["org_stage_called"] = True + return frame + + monkeypatch.setattr(builder, "load_org_2024_donor", fake_load_org_donor) + monkeypatch.setattr(builder, "with_us_org_wages_inputs", fake_with_org_inputs) + monkeypatch.setattr( + builder, + "us_org_wages_signal_gate", + lambda frame: builder.GateResult( + name="org_wages_signal", passed=True, details={"checked": True} + ), + ) monkeypatch.setattr( builder, "_ecps_parity_gate", @@ -1951,10 +3138,21 @@ def n(self, entity): details={"checked": True}, ), ) + + def fake_with_aca_outputs( + frame, + specs, + *, + seed, + maximum_microsim_batch_size=None, + ): + captured["health_stage_events"].append("aca") + return frame + monkeypatch.setattr( builder, "_with_aca_marketplace_source_outputs", - lambda frame, specs, *, seed, maximum_microsim_batch_size=None: frame, + fake_with_aca_outputs, ) monkeypatch.setattr( builder, @@ -1965,22 +3163,69 @@ def n(self, entity): details={"checked": True}, ), ) + + def fake_with_medicaid_outputs( + frame, + specs, + *, + seed, + substitutions=(), + maximum_microsim_batch_size=None, + ): + captured["health_stage_events"].append("medicaid") + return frame, {} + + def fake_medicaid_gate(diagnostics): + captured["health_stage_events"].append("medicaid_gate") + return builder.GateResult( + name="medicaid_take_up", + passed=True, + details={"checked": True}, + ) + monkeypatch.setattr( builder, "_with_medicaid_take_up_outputs", - lambda frame, specs, *, seed, substitutions=(), maximum_microsim_batch_size=None: ( - frame, - {}, - ), + fake_with_medicaid_outputs, ) monkeypatch.setattr( builder, "us_medicaid_take_up_gate", - lambda diagnostics: builder.GateResult( - name="medicaid_take_up", + fake_medicaid_gate, + ) + + def fake_with_other_health_insurance_inputs( + frame, + *, + seed, + time_period, + maximum_microsim_batch_size, + ): + captured["health_stage_events"].append("other_health") + captured["other_health_insurance_stage_called"] = True + captured["other_health_insurance_seed"] = seed + captured["other_health_insurance_period"] = time_period + captured["other_health_insurance_batch_size"] = maximum_microsim_batch_size + return frame + + def fake_other_health_insurance_signal_gate(frame): + captured["health_stage_events"].append("other_health_gate") + captured["other_health_insurance_gate_called"] = True + return builder.GateResult( + name="other_health_insurance_premiums_signal", passed=True, details={"checked": True}, - ), + ) + + monkeypatch.setattr( + builder, + "with_us_other_health_insurance_inputs", + fake_with_other_health_insurance_inputs, + ) + monkeypatch.setattr( + builder, + "us_other_health_insurance_signal_gate", + fake_other_health_insurance_signal_gate, ) monkeypatch.setattr( builder, @@ -2031,6 +3276,37 @@ def fake_write_diagnostics(**kwargs): return release_dir / "calibration_diagnostics.json" monkeypatch.setattr(builder, "calibrate_l0_refit", fake_calibrate_l0_refit) + + def fake_reconcile(frame, initial_result, specs, **kwargs): + captured["ssi_reconciliation_called"] = True + captured["ssi_reconciliation_reporter_source_ids"] = kwargs[ + "reporter_source_ids" + ] + return builder._SSITakeUpReconciliationResult( + export_frame=frame, + calibration_result=initial_result, + registry=registry, + compilation={"dropped_target_names": []}, + ssi_diagnostics={"checked": True}, + medicaid_diagnostics={}, + health_input_gate=builder.GateResult( + name="health_input_signal", + passed=True, + details={"checked": True}, + ), + other_health_insurance_gate=builder.GateResult( + name="other_health_insurance_premiums_signal", + passed=True, + details={"checked": True}, + ), + passes=1, + ) + + monkeypatch.setattr( + builder, + "_reconcile_ssi_take_up_and_refit", + fake_reconcile, + ) monkeypatch.setattr( builder, "_release_gate_failures", @@ -2074,6 +3350,9 @@ def fake_write_diagnostics(**kwargs): "selection_final_loss": 1.5, "refit_initial_loss": 2.0, "refit_final_loss": 1.0, + "pre_ssi_reconciliation_final_loss": 1.0, + "ssi_take_up_reconciliation_passes": 1, + "final_loss": 1.0, } assert captured["l0_kwargs"]["l0_lambda"] == 0.2 assert captured["l0_kwargs"]["l2_lambda"] == 0.0 @@ -2088,6 +3367,120 @@ def fake_write_diagnostics(**kwargs): == out / "artifacts" / "target_materialization_cache" ) assert not captured["materialize_kwargs"]["gate_congressional_district_targets"] + assert captured["sipp_tip_donor_path"] == Path("pu2023_slim.csv") + assert captured["weeks_unemployed_source_path"] == weeks_source + assert captured["weeks_unemployed_stage_seed"] == 0 + assert captured["weeks_unemployed_stage_period"] == builder.PERIOD + assert isinstance(captured["weeks_unemployed_stage_source"], pd.DataFrame) + assert captured["weeks_unemployed_gate_called"] is True + assert captured["source_stage_events"].index("weeks_stage") < captured[ + "source_stage_events" + ].index("ssi_reporters") + assert captured["source_stage_events"].index("weeks_gate") < captured[ + "source_stage_events" + ].index("population_repair") + assert ( + captured["materialize_kwargs"]["target_materialization_cache_context"][ + "weeks_unemployed_source_sha256" + ] + == builder.ASEC_2023_WEEKS_UNEMPLOYED_SOURCE_SHA256 + ) + assert captured["sipp_tip_donor_sha256"] == builder.SIPP_2023_TIP_DONOR_SHA256 + assert captured["sipp_tip_stage_called"] is True + assert captured["sipp_tip_gate_called"] is True + assert captured["sipp_vehicle_donor_path"] == Path("pu2023.csv") + assert ( + captured["sipp_vehicle_donor_sha256"] == builder.SIPP_2023_VEHICLE_DONOR_SHA256 + ) + assert ( + captured["sipp_vehicle_donor_size_bytes"] + == builder.SIPP_2023_VEHICLE_DONOR_SIZE_BYTES + ) + assert captured["sipp_vehicle_stage_called"] is True + assert captured["sipp_vehicle_seed"] == 42 + assert captured["sipp_vehicle_gate_called"] is True + assert captured["ssi_disability_donor_path"] == Path("pu2023.csv") + assert ( + captured["ssi_disability_donor_sha256"] + == builder.SIPP_2023_SSI_DISABILITY_DONOR_SHA256 + ) + assert ( + captured["ssi_disability_donor_size_bytes"] + == builder.SIPP_2023_SSI_DISABILITY_DONOR_SIZE_BYTES + ) + assert captured["ssi_disability_donor_period"] == builder.PERIOD + assert captured["ssi_disability_stage_called"] is True + assert captured["ssi_disability_seed"] == 42 + assert captured["ssi_disability_gate_called"] is True + assert captured["sipp_head_start_donor_path"] == Path("pu2023.csv") + assert ( + captured["sipp_head_start_donor_sha256"] + == builder.SIPP_2023_HEAD_START_DONOR_SHA256 + ) + assert ( + captured["sipp_head_start_donor_size_bytes"] + == builder.SIPP_2023_HEAD_START_DONOR_SIZE_BYTES + ) + assert captured["sipp_head_start_stage_called"] is True + assert captured["sipp_head_start_seed"] == 0 + assert captured["sipp_head_start_period"] == builder.PERIOD + assert isinstance(captured["sipp_head_start_donor"], pd.DataFrame) + assert captured["sipp_head_start_gate_called"] is True + assert captured["ssi_uncapped_stage_called"] is True + assert ( + captured["ssi_uncapped_batch_size"] + == builder.DEFAULT_MAXIMUM_MICROSIM_BATCH_SIZE + ) + assert captured["ssi_take_up_stage_called"] is True + assert captured["ssi_take_up_seed"] == 0 + assert captured["ssi_take_up_uncapped"].shape == (4,) + assert captured["ssi_reporter_source_ids"] == frozenset({"asec-reporter"}) + assert captured["ssi_reconciliation_reporter_source_ids"] == frozenset( + {"asec-reporter"} + ) + assert captured["ssi_take_up_gate_called"] is True + assert captured["ssi_take_up_gate_diagnostics"] == {"checked": True} + assert captured["voluntary_filing_donor_path"] == Path("pu2023.csv") + assert ( + captured["voluntary_filing_donor_sha256"] + == builder.SIPP_2023_VOLUNTARY_FILING_DONOR_SHA256 + ) + assert ( + captured["voluntary_filing_donor_size_bytes"] + == builder.SIPP_2023_VOLUNTARY_FILING_DONOR_SIZE_BYTES + ) + assert captured["voluntary_filing_stage_called"] is True + assert captured["voluntary_filing_seed"] == 0 + assert captured["voluntary_filing_gate_called"] is True + assert captured["org_donor_path"] == Path("census_cps_org_2024_wages.csv.gz") + assert captured["org_donor_sha256"] == builder.ORG_2024_DONOR_CONTENT_SHA256 + assert captured["org_stage_called"] is True + assert captured["scf_auto_summary_path"] == Path("rscfp2022.dta") + assert captured["scf_auto_full_path"] == Path("p22i6.dta") + assert captured["scf_auto_stage_called"] is True + assert captured["child_support_gate_called"] is True + assert captured["disability_benefits_gate_called"] is True + assert captured["educator_expense_gate_called"] is True + assert captured["form_4952_election_gate_called"] is True + assert captured["salt_refund_income_gate_called"] is True + assert captured["capital_gain_details_gate_called"] is True + assert captured["energy_subsidy_gate_called"] is True + assert captured["farm_business_income_gate_called"] is True + assert captured["other_health_insurance_stage_called"] is True + assert captured["other_health_insurance_gate_called"] is True + assert captured["other_health_insurance_seed"] == 0 + assert captured["other_health_insurance_period"] == builder.PERIOD + assert ( + captured["other_health_insurance_batch_size"] + == builder.DEFAULT_MAXIMUM_MICROSIM_BATCH_SIZE + ) + assert captured["health_stage_events"] == [ + "aca", + "medicaid", + "medicaid_gate", + "other_health", + "other_health_gate", + ] def test_release_gate_failures_reject_bad_national_credit_and_ss_fits() -> None: @@ -4052,6 +5445,8 @@ def test_build_manifests_emits_policyengine_certifiable_release_manifest( (artifact_root / builder.CALIBRATION_FILENAME).write_bytes(b"npz") (release_dir / "calibration_diagnostics.json").write_text("{}") (release_dir / "us_source_coverage.json").write_text("{}") + ssi_diagnostics_path = release_dir / "us_ssi_take_up.json" + ssi_diagnostics_path.write_text('{"variable":"takes_up_ssi_if_eligible"}') monkeypatch.setattr( builder, @@ -4218,6 +5613,13 @@ def __len__(self): assert not any( key.startswith(("states/", "districts/")) for key in manifest["artifacts"] ) + assert manifest["artifacts"]["us_ssi_take_up"] == { + "kind": "diagnostics", + "path": "us_ssi_take_up.json", + "repo_id": builder.REPO_ID, + "revision": release_id, + "sha256": builder._sha256(ssi_diagnostics_path), + } for artifact in manifest["artifacts"].values(): assert artifact["repo_id"] == builder.REPO_ID assert artifact["revision"] == release_id @@ -4277,6 +5679,7 @@ def test_build_manifests_records_selection_source_provenance( (artifact_root / builder.CALIBRATION_FILENAME).write_bytes(b"npz") (release_dir / "calibration_diagnostics.json").write_text("{}") (release_dir / "us_source_coverage.json").write_text("{}") + (release_dir / "us_ssi_take_up.json").write_text("{}") monkeypatch.setattr( builder, "_runtime_versions", @@ -4350,6 +5753,7 @@ def test_build_manifests_selection_source_absent_by_default( (artifact_root / builder.CALIBRATION_FILENAME).write_bytes(b"npz") (release_dir / "calibration_diagnostics.json").write_text("{}") (release_dir / "us_source_coverage.json").write_text("{}") + (release_dir / "us_ssi_take_up.json").write_text("{}") monkeypatch.setattr( builder, "_runtime_versions", @@ -4399,6 +5803,7 @@ def test_build_manifests_uses_incumbent_aware_calibration_gate( (artifact_root / builder.CALIBRATION_FILENAME).write_bytes(b"npz") (release_dir / "calibration_diagnostics.json").write_text("{}") (release_dir / "us_source_coverage.json").write_text("{}") + (release_dir / "us_ssi_take_up.json").write_text("{}") monkeypatch.setattr( builder, @@ -4990,6 +6395,7 @@ def counting_reform_household_income_tax(*, reform_spec, **kwargs): def _base_cache_context(builder): return { "base_dataset_sha256": "base-sha-A", + "weeks_unemployed_source_sha256": "weeks-source-sha-A", "build_commit": "commit-A", "policyengine_us_version": "pe-us-A", "seed": 0, @@ -5139,9 +6545,18 @@ def test__given_changed_reform_vector__then_stale_checkpoint_is_not_reused( np.testing.assert_allclose(household["jct_reform_a"], [-115.0, -69.0]) +@pytest.mark.parametrize( + ("identity_key", "new_value"), + [ + ("base_dataset_sha256", "base-sha-B"), + ("weeks_unemployed_source_sha256", "weeks-source-sha-B"), + ], +) def test__given_changed_frame_identity__then_stale_checkpoint_is_not_reused( monkeypatch, tmp_path, + identity_key, + new_value, ) -> None: builder = _load_builder_module() frame = _multi_reform_frame(builder) @@ -5168,10 +6583,10 @@ def test__given_changed_frame_identity__then_stale_checkpoint_is_not_reused( ) assert calls_a == ["credit_a"] - # A DIFFERENT base H5 (incumbent vs candidate) must not share per-household - # vectors even at the same record count (#217 acceptance criterion 3). + # A different base H5 or measured LKWEEKS source must not share + # per-household vectors even at the same record count. context_b = _base_cache_context(builder) - context_b["base_dataset_sha256"] = "base-sha-B" + context_b[identity_key] = new_value calls_b: list[str] = [] _install_multi_reform_fakes( builder, diff --git a/packages/populace-build/tests/test_us_fiscal_targets.py b/packages/populace-build/tests/test_us_fiscal_targets.py index 85b3ba27..77594f2a 100644 --- a/packages/populace-build/tests/test_us_fiscal_targets.py +++ b/packages/populace-build/tests/test_us_fiscal_targets.py @@ -3423,11 +3423,14 @@ def test_macro_realism_bands_cover_issue_40_backstops() -> None: def test_scf_nonnegative_targets_gate_negative_interest() -> None: - assert "auto_loan_interest" in US_NONNEGATIVE_SOURCE_OUTPUTS - result = nonnegative_columns_gate( - {"auto_loan_interest": [120.0, -9.0]}, - US_NONNEGATIVE_SOURCE_OUTPUTS, - ) + assert { + "auto_loan_balance", + "auto_loan_interest", + "qualified_passenger_vehicle_loan_interest", + } <= US_NONNEGATIVE_SOURCE_OUTPUTS + columns = {name: [0.0, 0.0] for name in US_NONNEGATIVE_SOURCE_OUTPUTS} + columns["auto_loan_interest"] = [120.0, -9.0] + result = nonnegative_columns_gate(columns, US_NONNEGATIVE_SOURCE_OUTPUTS) assert not result.passed assert "auto_loan_interest" in result.failures[0] diff --git a/packages/populace-build/tests/test_us_form_4952.py b/packages/populace-build/tests/test_us_form_4952.py new file mode 100644 index 00000000..a59d87b4 --- /dev/null +++ b/packages/populace-build/tests/test_us_form_4952.py @@ -0,0 +1,407 @@ +"""Contracts for the IRS PUF Form 4952 elected-investment-income input.""" + +from __future__ import annotations + +import importlib.util + +import numpy as np +import pandas as pd +import pytest + +from populace.build.us_runtime import ( + BASE_ASEC_SUPPORT_CHANNEL, + PUF_TAX_DETAIL_SUPPORT_CHANNEL, + clone_us_frame_for_puf_support, + impute_us_puf_tax_detail_support, + puf_tax_unit_donor_from_arrays, + support_channel_column, +) +from populace.build.us_runtime import puf_support as puf_support_module +from populace.build.us_runtime.form_4952 import ( + FORM_4952_ARCHIVED_DERIVATION_URL, + FORM_4952_ARCHIVED_EXPORT_URL, + FORM_4952_ARCHIVED_IMPUTATION_URL, + FORM_4952_ARCHIVED_PERSON_ALLOCATION_URL, + FORM_4952_ARCHIVED_PUF_ARTIFACT_URL, + derive_us_form_4952_election_from_puf, + us_form_4952_election_signal_gate, + us_form_4952_election_stage_spec, +) +from populace.build.us_runtime.l0_refit_export import ( + US_RELEASE_REQUIRED_PERSON_SOURCE_COLUMNS, +) +from populace.build.us_runtime.puf_aggregate_records import ( + _reconcile_puf_form_4952_election_from_source, + derive_puf_policyengine_variables, +) +from populace.build.us_runtime.reform_coverage_smoke import _build_reform +from populace.build.us_runtime.release_input_coverage import ( + RESTORED_REFERENCE_ECPS_REQUIRED_INPUTS, + load_release_input_coverage_manifest, + us_release_reform_coverage_probes, +) +from populace.frame import US_SCHEMA, Frame, WeightKind, Weights + +policyengine_us_installed = importlib.util.find_spec("policyengine_us") is not None +requires_us = pytest.mark.skipif( + not policyengine_us_installed, + reason="requires the policyengine-us [us] extra (build environment)", +) + + +class _ResolvedWeights: + def __init__(self, values: np.ndarray) -> None: + self.values = values + + +class _PersonFrame: + def __init__( + self, + person: pd.DataFrame, + weights: np.ndarray | None = None, + ) -> None: + self._person = person + self._weights = np.ones(len(person)) if weights is None else weights + + def table(self, entity: str) -> pd.DataFrame: + assert entity == "person" + return self._person + + def resolve_weights(self, entity: str) -> _ResolvedWeights: + assert entity == "person" + return _ResolvedWeights(np.asarray(self._weights, dtype=np.float64)) + + +def _joint_support_frame() -> Frame: + tables = { + "person": pd.DataFrame( + { + "person_id": np.asarray([1, 2], dtype="int64"), + "person_household_id": np.asarray([10, 10], dtype="int64"), + "person_tax_unit_id": np.asarray([100, 100], dtype="int64"), + "person_spm_unit_id": np.asarray([1_000, 1_000], dtype="int64"), + "person_family_id": np.asarray([10_000, 10_000], dtype="int64"), + "person_marital_unit_id": np.asarray([100_000, 100_000], dtype="int64"), + "employment_income_before_lsr": [50_000.0, 20_000.0], + "self_employment_income_before_lsr": [0.0, 40_000.0], + } + ), + "household": pd.DataFrame({"household_id": np.asarray([10], dtype="int64")}), + "tax_unit": pd.DataFrame( + { + "tax_unit_id": np.asarray([100], dtype="int64"), + "filing_status_input": ["JOINT"], + } + ), + "spm_unit": pd.DataFrame({"spm_unit_id": np.asarray([1_000], dtype="int64")}), + "family": pd.DataFrame({"family_id": np.asarray([10_000], dtype="int64")}), + "marital_unit": pd.DataFrame( + {"marital_unit_id": np.asarray([100_000], dtype="int64")} + ), + } + return Frame( + tables, + US_SCHEMA, + {"household": Weights(np.asarray([100.0]), WeightKind.DESIGN)}, + ) + + +def test_archived_coordinates_pin_derivation_export_and_imputation() -> None: + commit = "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe" + + assert commit in FORM_4952_ARCHIVED_DERIVATION_URL + assert FORM_4952_ARCHIVED_DERIVATION_URL.endswith("puf.py#L708") + assert commit in FORM_4952_ARCHIVED_EXPORT_URL + assert FORM_4952_ARCHIVED_EXPORT_URL.endswith("puf.py#L804-L850") + assert commit in FORM_4952_ARCHIVED_IMPUTATION_URL + assert FORM_4952_ARCHIVED_IMPUTATION_URL.endswith("puf_impute.py#L90-L198") + assert FORM_4952_ARCHIVED_PERSON_ALLOCATION_URL.endswith("puf.py#L477-L546") + assert FORM_4952_ARCHIVED_PUF_ARTIFACT_URL.endswith("puf.py#L1655-L1660") + + +def test_archived_e58990_mapping_is_an_exact_carry() -> None: + source = pd.DataFrame({"E58990": [0.0, 125.5, 9_000.0]}) + + result = derive_us_form_4952_election_from_puf(source) + + assert "investment_income_elected_form_4952" not in source.columns + assert result["investment_income_elected_form_4952"].tolist() == [ + 0.0, + 125.5, + 9_000.0, + ] + + +@pytest.mark.parametrize( + ("source", "message"), + [ + (pd.DataFrame({"other": [1.0]}), "requires source column"), + (pd.DataFrame({"E58990": ["not numeric"]}), "nonnumeric or nonfinite"), + (pd.DataFrame({"E58990": [np.inf]}), "nonnumeric or nonfinite"), + (pd.DataFrame({"E58990": [-1.0]}), "negative value"), + ], +) +def test_source_derivation_fails_closed( + source: pd.DataFrame, + message: str, +) -> None: + with pytest.raises(ValueError, match=message): + derive_us_form_4952_election_from_puf(source) + + +def test_shared_puf_derivation_and_post_disaggregation_reconciliation() -> None: + source = pd.DataFrame( + { + "E00600": [0.0, 0.0, 0.0], + "E00650": [0.0, 0.0, 0.0], + "E58990": [0.0, 3_000.0, 7_500.0], + } + ) + + derived = derive_puf_policyengine_variables( + source, + investment_income_elected_form_4952_source="E58990", + ) + derived["investment_income_elected_form_4952"] = 999.0 + result = _reconcile_puf_form_4952_election_from_source(derived) + + assert result["investment_income_elected_form_4952"].tolist() == [ + 0.0, + 3_000.0, + 7_500.0, + ] + + +def test_shared_puf_stage_declares_exact_source_output_and_artifact() -> None: + stage = us_form_4952_election_stage_spec() + operation = next( + operation + for operation in stage.operations + if operation.kind == "derive_puf_policyengine_variables" + ) + + assert ( + operation.parameters["investment_income_elected_form_4952_source"] == "E58990" + ) + assert ( + operation.parameters["investment_income_elected_form_4952_output"] + == "investment_income_elected_form_4952" + ) + assert "investment_income_elected_form_4952" in stage.outputs + assert "investment_income_elected_form_4952" in stage.nonnegative_outputs + assert any("E58990" in str(artifact.get("locator")) for artifact in stage.artifacts) + assert any( + "irs-soi-puf/1.8.0/puf_2024.h5" in str(artifact.get("locator")) + for artifact in stage.artifacts + ) + + +def test_puf_support_keeps_form_4952_nonnegative_and_sparse() -> None: + output = "investment_income_elected_form_4952" + + assert output in puf_support_module.PUF_TAX_DETAIL_DEFAULT_PERSON_OUTPUTS + assert output in puf_support_module._PUF_TAX_DETAIL_NONNEGATIVE_OUTPUTS + assert output in puf_support_module._PUF_TAX_DETAIL_SPARSE_PERSON_OUTPUTS + assert output not in puf_support_module._PERSON_OUTPUT_DISTRIBUTION_BASIS + + +def test_processed_puf_people_aggregate_to_one_tax_unit_election() -> None: + donor = puf_tax_unit_donor_from_arrays( + { + "tax_unit_id": [10, 20], + "household_weight": [100.0, 200.0], + "filing_status": [b"SINGLE", b"JOINT"], + "person_tax_unit_id": [10, 20, 20], + "investment_income_elected_form_4952": [0.0, 100.0, 200.0], + }, + person_outputs=("investment_income_elected_form_4952",), + tax_unit_outputs=(), + ) + + assert donor["investment_income_elected_form_4952"].tolist() == [0.0, 300.0] + + +def test_puf_imputation_preserves_tax_unit_total_without_inventing_split() -> None: + expanded = clone_us_frame_for_puf_support(_joint_support_frame()) + donor = pd.DataFrame( + { + "filing_status_code": [2.0, 2.0], + "tax_unit_person_count": [2.0, 2.0], + "investment_income_elected_form_4952": [900.0, 900.0], + "weight": [1.0, 1.0], + } + ) + + imputed = impute_us_puf_tax_detail_support( + expanded, + donor, + predictors=( + "puf_predictor_filing_status_code", + "puf_predictor_tax_unit_person_count", + ), + person_outputs=("investment_income_elected_form_4952",), + tax_unit_outputs=(), + n_estimators=4, + seed=0, + ) + person = imputed.table("person") + channel = support_channel_column("person") + asec = person[person[channel] == BASE_ASEC_SUPPORT_CHANNEL] + puf = person[person[channel] == PUF_TAX_DETAIL_SUPPORT_CHANNEL].sort_values( + "person_id" + ) + + assert asec["investment_income_elected_form_4952"].tolist() == [0.0, 0.0] + np.testing.assert_allclose( + puf["investment_income_elected_form_4952"].to_numpy(), + [900.0, 0.0], + ) + assert puf["investment_income_elected_form_4952"].sum() == pytest.approx(900.0) + + +def test_signal_gate_accepts_source_aligned_sparse_nondefault_values() -> None: + values = np.zeros(2_000) + values[[1_010, 1_020, 1_030]] = [1_000.0, 2_000.0, 3_000.0] + channels = np.asarray(["asec"] * 1_000 + ["puf_tax_detail"] * 1_000) + frame = _PersonFrame( + pd.DataFrame( + { + "investment_income_elected_form_4952": values, + "person_support_channel": channels, + } + ) + ) + + result = us_form_4952_election_signal_gate(frame) # type: ignore[arg-type] + + assert result.passed, result.failures + assert result.details["positive_share"] == pytest.approx(0.0015) + assert result.details["channels"]["asec"]["weighted_total"] == 0.0 + assert result.details["channels"]["puf_tax_detail"]["positive_share"] == ( + pytest.approx(0.003) + ) + + +@pytest.mark.parametrize( + "person", + [ + pd.DataFrame({"other": [0.0, 1.0]}), + pd.DataFrame({"investment_income_elected_form_4952": [0.0, 0.0]}), + pd.DataFrame({"investment_income_elected_form_4952": [0.0, -1.0]}), + pd.DataFrame({"investment_income_elected_form_4952": [0.0, np.nan]}), + pd.DataFrame( + { + "investment_income_elected_form_4952": [10.0, 0.0], + "person_support_channel": ["asec", "puf_tax_detail"], + } + ), + ], +) +def test_signal_gate_rejects_missing_default_invalid_or_asec_signal( + person: pd.DataFrame, +) -> None: + result = us_form_4952_election_signal_gate( # type: ignore[arg-type] + _PersonFrame(person) + ) + + assert not result.passed + + +def test_release_wiring_keeps_restored_input_hard_required() -> None: + output = "investment_income_elected_form_4952" + + assert output in US_RELEASE_REQUIRED_PERSON_SOURCE_COLUMNS + assert output in RESTORED_REFERENCE_ECPS_REQUIRED_INPUTS + assert output in load_release_input_coverage_manifest().required_columns + + +def test_shipped_neutralization_probe_binds_only_through_form_4952() -> None: + probe = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "form_4952_election_neutralization" + ) + + assert probe.parameter_changes == {} + assert probe.neutralized_variable == "investment_income_elected_form_4952" + assert probe.binding_inputs == ("investment_income_elected_form_4952",) + assert probe.period == 2024 + assert probe.effect_direction == "baseline_minus_reform" + assert probe.expected_sign == "positive" + assert probe.budget_measure == "income_tax" + assert probe.min_abs_effect == 1_000_000.0 + + +@requires_us +def test_policyengine_us_contract_and_net_capital_gain_binding() -> None: + from policyengine_us import CountryTaxBenefitSystem, Simulation + + variable = CountryTaxBenefitSystem().variables[ + "investment_income_elected_form_4952" + ] + assert variable.is_input_variable() + assert variable.entity.key == "person" + assert str(variable.definition_period).lower() == "year" + assert variable.default_value == 0 + + adult = { + "age": {2024: 45}, + "long_term_capital_gains_before_response": {2024: 100_000}, + "qualified_dividend_income": {2024: 0}, + "short_term_capital_gains": {2024: 0}, + } + entities = { + "tax_units": { + "unit": {"members": ["adult"], "filing_status": {2024: "SINGLE"}} + }, + "families": {"family": {"members": ["adult"]}}, + "spm_units": {"spm": {"members": ["adult"]}}, + "households": { + "household": { + "members": ["adult"], + "state_code_str": {2024: "CA"}, + } + }, + "marital_units": {"marital": {"members": ["adult"]}}, + } + baseline = Simulation(situation={"people": {"adult": adult}, **entities}) + election = Simulation( + situation={ + "people": { + "adult": { + **adult, + "investment_income_elected_form_4952": {2024: 25_000}, + } + }, + **entities, + } + ) + + assert baseline.calculate("net_capital_gain", 2024)[0] == pytest.approx(100_000) + assert election.calculate("net_capital_gain", 2024)[0] == pytest.approx(75_000) + assert ( + election.calculate("income_tax", 2024)[0] + > baseline.calculate("income_tax", 2024)[0] + ) + + probe = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "form_4952_election_neutralization" + ) + neutralized = Simulation( + situation={ + "people": { + "adult": { + **adult, + "investment_income_elected_form_4952": {2024: 25_000}, + } + }, + **entities, + }, + reform=_build_reform(probe), + ) + assert ( + election.calculate("income_tax", 2024)[0] + > neutralized.calculate("income_tax", 2024)[0] + ) diff --git a/packages/populace-build/tests/test_us_housing_inputs.py b/packages/populace-build/tests/test_us_housing_inputs.py new file mode 100644 index 00000000..c3db8be3 --- /dev/null +++ b/packages/populace-build/tests/test_us_housing_inputs.py @@ -0,0 +1,606 @@ +"""Archived CPS/ACS housing-input restoration.""" + +from __future__ import annotations + +import importlib.util +from pathlib import Path + +import h5py +import numpy as np +import pandas as pd +import pytest + +import populace.build.us_runtime.housing_inputs as module +from populace.build.us_runtime.housing_inputs import ( + ACS_2022_RENT_ARTIFACT_SHA256, + HOUSING_INPUTS_ARCHIVED_ACS_DERIVATION_URL, + HOUSING_INPUTS_ARCHIVED_CPS_RENT_URL, + HOUSING_INPUTS_ARCHIVED_CPS_SPM_URL, + HOUSING_INPUTS_ARCHIVED_PUF_IMPUTATION_URL, + HOUSING_TAKE_UP_ARCHIVED_DERIVATION_URL, + HOUSING_TAKE_UP_ARCHIVED_HUD_ETL_URL, + HOUSING_TAKE_UP_ARCHIVED_PARAMETER_URL, + US_HOUSING_INPUTS_OUTPUT_COLUMNS, + US_HOUSING_NONCONSTANT_HOUSEHOLD_COLUMNS, + US_HOUSING_NONCONSTANT_PERSON_COLUMNS, + US_HOUSING_NONCONSTANT_SPM_UNIT_COLUMNS, + derive_us_housing_inputs, + impute_us_housing_assistance_to_puf_support, + load_acs_2022_rent_donor, + us_housing_inputs_signal_gate, + us_housing_inputs_stage_spec, + with_us_housing_inputs, +) +from populace.build.us_runtime.l0_refit_export import ( + US_RELEASE_REQUIRED_HOUSEHOLD_NONCONSTANT_SOURCE_COLUMNS, + US_RELEASE_REQUIRED_PERSON_SOURCE_COLUMNS, + US_RELEASE_REQUIRED_SPM_UNIT_SOURCE_COLUMNS, +) +from populace.build.us_runtime.puf_support import ( + BASE_ASEC_SUPPORT_CHANNEL, + PUF_TAX_DETAIL_SUPPORT_CHANNEL, + clone_us_frame_for_puf_support, + support_channel_column, +) +from populace.build.us_runtime.release_input_coverage import ( + RESTORED_REFERENCE_ECPS_REQUIRED_INPUTS, + load_release_input_coverage_manifest, + us_release_reform_coverage_probes, +) +from populace.build.us_runtime.take_up_contract import load_take_up_contract +from populace.frame import US_SCHEMA, Frame, WeightKind, Weights + +policyengine_us_installed = importlib.util.find_spec("policyengine_us") is not None +requires_us = pytest.mark.skipif( + not policyengine_us_installed, + reason="requires the policyengine-us [us] extra (build environment)", +) + + +def _frame() -> Frame: + n = 20 + household_ids = np.arange(1, n + 1, dtype=np.int64) + person = pd.DataFrame( + { + "person_id": household_ids, + "person_household_id": household_ids, + "person_tax_unit_id": household_ids + 100, + "person_spm_unit_id": household_ids + 200, + "person_family_id": household_ids + 300, + "person_marital_unit_id": household_ids + 400, + "is_household_head": np.ones(n, dtype=bool), + "tax_unit_role_input": ["HEAD"] * n, + "age": np.linspace(25, 75, n), + "is_female": household_ids % 2 == 0, + "has_esi": household_ids % 3 == 0, + "employment_income_before_lsr": household_ids * 2_000.0, + "self_employment_income_before_lsr": household_ids * 100.0, + "social_security_retirement": np.where(household_ids > 15, 12_000.0, 0.0), + "social_security_disability": np.zeros(n), + "social_security_survivors": np.zeros(n), + "social_security_dependents": np.zeros(n), + "taxable_private_pension_income": np.where( + household_ids > 15, 4_000.0, 0.0 + ), + "tax_exempt_private_pension_income": np.zeros(n), + "SPM_CAPHOUSESUB": np.where(household_ids == 1, 5_000.0, 0.0), + "SPM_TENMORTSTATUS": np.resize(np.array([3, 1, 2]), n), + } + ) + h_tenure = np.where(household_ids <= 4, 2, np.where(household_ids == 5, 3, 1)) + tables = { + "person": person, + "household": pd.DataFrame( + { + "household_id": household_ids, + "state_fips": np.where(household_ids <= 10, 6, 36), + "H_TENURE": h_tenure, + } + ), + "tax_unit": pd.DataFrame( + { + "tax_unit_id": household_ids + 100, + "filing_status_input": ["SINGLE"] * n, + } + ), + "spm_unit": pd.DataFrame({"spm_unit_id": household_ids + 200}), + "family": pd.DataFrame({"family_id": household_ids + 300}), + "marital_unit": pd.DataFrame({"marital_unit_id": household_ids + 400}), + } + return Frame( + tables, + US_SCHEMA, + { + "household": Weights( + np.full(n, 1_000_000.0), + WeightKind.DESIGN, + ) + }, + ) + + +def _donor(n: int = 60) -> pd.DataFrame: + rows = np.arange(n, dtype=np.float64) + donor = pd.DataFrame( + { + predictor: rows + position + for position, predictor in enumerate(module.ACS_RENT_PREDICTORS) + } + ) + donor["is_household_head"] = 1.0 + donor["tenure_type"] = np.resize( + np.array(["NONE", "OWNED_WITH_MORTGAGE", "RENTED"]), n + ) + donor["state_code_str"] = np.resize(np.array(["06", "36", "48"]), n) + donor["rent"] = np.where(donor["tenure_type"] == "RENTED", 12_000.0, 0.0) + donor["rent_is_allocated"] = False + donor["real_estate_taxes"] = np.where( + donor["tenure_type"] == "RENTED", 0.0, 4_000.0 + ) + donor["real_estate_taxes_is_allocated"] = False + donor["household_weight"] = np.linspace(1.0, 2.0, n) + return donor + + +class _RentFitted: + def predict(self, test: pd.DataFrame) -> pd.DataFrame: + rented = test["tenure_type__RENTED"] > 0 + return pd.DataFrame( + {"rent": np.where(rented, 12_000.0, 0.0)}, + index=test.index, + ) + + +class _RentQRF: + def __init__(self, **_kwargs: object) -> None: + pass + + def fit(self, *_args: object, **_kwargs: object) -> _RentFitted: + return _RentFitted() + + +def test_stage_manifest_pins_exact_archived_sources_and_two_qrfs() -> None: + spec = us_housing_inputs_stage_spec() + + assert spec.stage == "acs_rent" + assert spec.outputs == US_HOUSING_INPUTS_OUTPUT_COLUMNS + assert [operation.kind for operation in spec.operations] == [ + "derive_housing_tenure_inputs", + "read_acs_rent_donor", + "fit_weighted_acs_rent_qrf", + "impute_housing_assistance_to_puf_support", + ] + rent_fit = spec.operations[2] + assert tuple(rent_fit.parameters["predictors"]) == module.ACS_RENT_PREDICTORS + assert rent_fit.parameters["max_train_samples"] == 10_000 + assert rent_fit.parameters["shared_sample_targets"] == [ + "rent", + "real_estate_taxes", + ] + assert rent_fit.parameters["per_target_initial_cap"] == 5_000 + assert rent_fit.parameters["weight"] == "household_weight" + puf_fit = spec.operations[3] + assert puf_fit.parameters["max_train_samples"] == 5_000 + assert puf_fit.parameters["reduction"] == "value_from_first_person" + assert puf_fit.parameters["take_up_output"] == ( + "takes_up_housing_assistance_if_eligible" + ) + direct = spec.operations[0] + assert direct.parameters["housing_take_up_assignment"] == ( + "equal_to_measured_receipt" + ) + assert all( + url.startswith( + "https://github.com/PolicyEngine/" + "policyengine-" + "us-data/blob/" + "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe/" + ) + for url in ( + HOUSING_INPUTS_ARCHIVED_CPS_RENT_URL, + HOUSING_INPUTS_ARCHIVED_CPS_SPM_URL, + HOUSING_INPUTS_ARCHIVED_ACS_DERIVATION_URL, + HOUSING_INPUTS_ARCHIVED_PUF_IMPUTATION_URL, + HOUSING_TAKE_UP_ARCHIVED_DERIVATION_URL, + HOUSING_TAKE_UP_ARCHIVED_PARAMETER_URL, + HOUSING_TAKE_UP_ARCHIVED_HUD_ETL_URL, + ) + ) + artifact = next(artifact for artifact in spec.artifacts if artifact.get("sha256")) + assert artifact["sha256"] == ACS_2022_RENT_ARTIFACT_SHA256 + assert HOUSING_TAKE_UP_ARCHIVED_DERIVATION_URL.endswith( + "/datasets/cps/cps.py#L664-L682" + ) + assert HOUSING_TAKE_UP_ARCHIVED_PARAMETER_URL.endswith( + "/parameters/take_up/housing_assistance.yaml#L1-L15" + ) + assert HOUSING_TAKE_UP_ARCHIVED_HUD_ETL_URL.endswith( + "/db/etl_housing_assistance.py#L35-L169" + ) + + +def test_direct_asec_mappings_are_exact() -> None: + result = derive_us_housing_inputs(_frame()) + + household = result.table("household") + assert household["tenure_type"].tolist()[:6] == [ + "RENTED", + "RENTED", + "RENTED", + "RENTED", + "NONE", + "OWNED_WITH_MORTGAGE", + ] + spm = result.table("spm_unit") + assert spm["receives_housing_assistance"].tolist().count(True) == 1 + pd.testing.assert_series_equal( + spm["takes_up_housing_assistance_if_eligible"], + spm["receives_housing_assistance"], + check_names=False, + ) + assert spm["spm_unit_tenure_type"].tolist()[:3] == [ + "RENTER", + "OWNER_WITH_MORTGAGE", + "OWNER_WITHOUT_MORTGAGE", + ] + + +def test_direct_mapping_rejects_inconsistent_spm_replicas() -> None: + frame = _frame() + extra = frame.table("person").iloc[[0]].copy() + extra["person_id"] = 999 + extra["SPM_CAPHOUSESUB"] = 0.0 + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + tables["person"] = pd.concat([tables["person"], extra], ignore_index=True) + broken = Frame(tables, frame.schema, {"household": frame.weights_for("household")}) + + with pytest.raises(ValueError, match="constant within its SPM unit"): + derive_us_housing_inputs(broken) + + +def test_materialization_places_rent_only_on_renter_heads( + monkeypatch: pytest.MonkeyPatch, +) -> None: + monkeypatch.setattr(module, "QRF", _RentQRF) + + result = with_us_housing_inputs( + _frame(), + seed=0, + time_period=2024, + acs_rent_donor=_donor(), + ) + + rent = result.table("person")["pre_subsidy_rent"] + assert rent.tolist()[:6] == [12_000.0] * 4 + [0.0, 0.0] + gate = us_housing_inputs_signal_gate(result) + assert gate.passed, gate.failures + + +def test_rent_fit_replays_archived_joint_target_sample( + monkeypatch: pytest.MonkeyPatch, +) -> None: + donor = _donor(12_000) + positions = np.arange(len(donor)) + donor["rent_is_allocated"] = positions >= 6_000 + donor["real_estate_taxes_is_allocated"] = positions < 6_000 + captured: dict[str, object] = {} + + class Fitted: + def predict(self, test: pd.DataFrame) -> pd.DataFrame: + return pd.DataFrame({"rent": np.zeros(len(test))}, index=test.index) + + class QRF: + def __init__(self, **_kwargs: object) -> None: + pass + + def fit( + self, + training: pd.DataFrame, + *, + predictors: list[str], + targets: list[str], + weights: str, + ) -> Fitted: + captured["training"] = training.copy() + captured["predictors"] = predictors + captured["targets"] = targets + captured["weights"] = weights + return Fitted() + + monkeypatch.setattr(module, "QRF", QRF) + module.impute_us_pre_subsidy_rent(derive_us_housing_inputs(_frame()), donor, seed=0) + + # The retired shared cap first draws 5,000 rent rows and 5,000 disjoint + # real-estate-tax rows. Only the rent-filtered half trains this first + # chained target; sampling 10,000 rent rows directly would violate parity. + assert len(captured["training"]) == 5_000 + assert captured["targets"] == ["rent"] + assert captured["weights"] == "household_weight" + + +def test_signal_gate_preserves_archived_qrf_rent_outside_renter_household( + monkeypatch: pytest.MonkeyPatch, +) -> None: + monkeypatch.setattr(module, "QRF", _RentQRF) + result = with_us_housing_inputs( + _frame(), seed=0, time_period=2024, acs_rent_donor=_donor() + ) + result.table("person").loc[5, "pre_subsidy_rent"] = 1.0 + + gate = us_housing_inputs_signal_gate(result) + + assert gate.passed, gate.failures + assert gate.details["positive_rent_nonrenter"] == 1 + + +def test_puf_half_reimputes_only_housing_assistance( + monkeypatch: pytest.MonkeyPatch, +) -> None: + monkeypatch.setattr(module, "QRF", _RentQRF) + direct = with_us_housing_inputs( + _frame(), seed=0, time_period=2024, acs_rent_donor=_donor() + ) + expanded = clone_us_frame_for_puf_support(direct) + calls: dict[str, object] = {} + + class Fitted: + def predict(self, test: pd.DataFrame) -> pd.DataFrame: + values = np.zeros(len(test), dtype=np.float64) + values[1] = 1.0 + return pd.DataFrame( + {"receives_housing_assistance": values}, index=test.index + ) + + class QRF: + def __init__(self, **kwargs: object) -> None: + calls["init"] = kwargs + + def fit( + self, + training: pd.DataFrame, + predictors: list[str], + targets: list[str], + **kwargs: object, + ) -> Fitted: + calls["training"] = training.copy() + calls["predictors"] = predictors + calls["targets"] = targets + calls["fit_kwargs"] = kwargs + return Fitted() + + monkeypatch.setattr(module, "QRF", QRF) + result = impute_us_housing_assistance_to_puf_support(expanded, seed=7) + + spm = result.table("spm_unit") + channel = spm[support_channel_column("spm_unit")] + asec = channel == BASE_ASEC_SUPPORT_CHANNEL + puf = channel == PUF_TAX_DETAIL_SUPPORT_CHANNEL + assert spm.loc[asec, "receives_housing_assistance"].sum() == 1 + assert spm.loc[puf, "receives_housing_assistance"].sum() == 1 + assert not np.array_equal( + spm.loc[asec, "receives_housing_assistance"].to_numpy(), + spm.loc[puf, "receives_housing_assistance"].to_numpy(), + ) + pd.testing.assert_series_equal( + spm["takes_up_housing_assistance_if_eligible"], + spm["receives_housing_assistance"], + check_names=False, + ) + assert calls["targets"] == ["receives_housing_assistance"] + assert len(calls["predictors"]) == 8 + assert "pre_subsidy_rent" not in calls["targets"] + for entity, column in ( + ("person", "pre_subsidy_rent"), + ("spm_unit", "spm_unit_tenure_type"), + ("household", "tenure_type"), + ): + source = expanded.table(entity)[column] + actual = result.table(entity)[column] + pd.testing.assert_series_equal(actual, source) + + result.table("spm_unit").loc[puf, "receives_housing_assistance"] = False + collapsed_gate = us_housing_inputs_signal_gate(result) + assert not collapsed_gate.passed + assert any("PUF-tax-detail" in failure for failure in collapsed_gate.failures) + + +@pytest.mark.parametrize("bad_value", [None, np.nan, 2, -1, "true"]) +def test_signal_gate_rejects_missing_or_invalid_take_up(bad_value: object) -> None: + result = derive_us_housing_inputs(_frame()) + result.table("spm_unit")["takes_up_housing_assistance_if_eligible"] = result.table( + "spm_unit" + )["takes_up_housing_assistance_if_eligible"].astype(object) + result.table("spm_unit").loc[0, "takes_up_housing_assistance_if_eligible"] = ( + bad_value + ) + result.table("person")["pre_subsidy_rent"] = np.where( + result.table("person")["is_household_head"], + 1.0, + 0.0, + ) + + gate = us_housing_inputs_signal_gate(result) + + assert not gate.passed + assert any("takes_up_housing_assistance" in failure for failure in gate.failures) + + +def _write_tiny_acs(path: Path) -> None: + arrays = { + "person_id": np.array([1, 2, 3, 4]), + "person_household_id": np.array([10, 10, 20, 30]), + "is_household_head": np.array([True, False, True, True]), + "age": np.array([40, 38, 50, 60]), + "is_male": np.array([True, False, True, False]), + "employment_income": np.array([50_000, 20_000, 0, 10_000]), + "self_employment_income": np.array([0, 0, 5_000, 0]), + "social_security": np.array([0, 0, 10_000, 12_000]), + "taxable_private_pension_income": np.array([0, 0, 3_000, 4_000]), + "rent": np.array([12_000, 0, 0, 18_000]), + "rent_is_allocated": np.array([False, False, True, False]), + "real_estate_taxes": np.array([0, 0, 8_000, 4_000]), + "real_estate_taxes_is_allocated": np.array([False, False, False, True]), + "household_id": np.array([10, 20, 30]), + "household_weight": np.array([100.0, 200.0, 300.0]), + "state_fips": np.array([6, 36, 48]), + "tenure_type": np.array([b"RENTED", b"OWNED_OUTRIGHT", b"OWNED_WITH_MORTGAGE"]), + } + with h5py.File(path, "w") as h5: + for name, values in arrays.items(): + h5[name] = values + + +def test_acs_loader_aligns_entities_collapses_tenure_and_marks_allocations( + tmp_path: Path, +) -> None: + path = tmp_path / "acs_2022.h5" + _write_tiny_acs(path) + + donor = load_acs_2022_rent_donor(path, expected_sha256=None) + + assert len(donor) == 3 + assert donor["household_size"].tolist() == [2.0, 1.0, 1.0] + assert donor["tenure_type"].tolist() == [ + "RENTED", + "OWNED_WITH_MORTGAGE", + "OWNED_WITH_MORTGAGE", + ] + assert donor["state_code_str"].tolist() == ["06", "36", "48"] + assert donor["rent"].tolist() == [12_000.0, 0.0, 18_000.0] + assert donor["rent_is_allocated"].tolist() == [False, True, False] + assert donor["real_estate_taxes"].tolist() == [0.0, 8_000.0, 4_000.0] + assert donor["real_estate_taxes_is_allocated"].tolist() == [False, False, True] + assert donor["household_weight"].tolist() == [100.0, 200.0, 300.0] + + +def test_acs_loader_rejects_wrong_artifact_hash(tmp_path: Path) -> None: + path = tmp_path / "acs_2022.h5" + _write_tiny_acs(path) + + with pytest.raises(ValueError, match="SHA-256 mismatch"): + load_acs_2022_rent_donor(path, expected_sha256="0" * 64) + + +def test_acs_loader_retains_zero_weight_heads_for_archived_sampling( + tmp_path: Path, +) -> None: + path = tmp_path / "acs_2022.h5" + _write_tiny_acs(path) + with h5py.File(path, "a") as h5: + h5["household_weight"][1] = 0.0 + + donor = load_acs_2022_rent_donor(path, expected_sha256=None) + + assert len(donor) == 3 + assert donor["household_weight"].tolist() == [100.0, 0.0, 300.0] + + +def test_release_wiring_promotes_all_five_inputs_and_probes() -> None: + manifest = load_release_input_coverage_manifest() + for column in US_HOUSING_INPUTS_OUTPUT_COLUMNS: + assert column in RESTORED_REFERENCE_ECPS_REQUIRED_INPUTS + assert column in manifest.required_columns + assert column not in manifest.reviewed_exclusions + assert set(US_HOUSING_NONCONSTANT_PERSON_COLUMNS) <= set( + US_RELEASE_REQUIRED_PERSON_SOURCE_COLUMNS + ) + assert set(US_HOUSING_NONCONSTANT_SPM_UNIT_COLUMNS) <= set( + US_RELEASE_REQUIRED_SPM_UNIT_SOURCE_COLUMNS + ) + assert set(US_HOUSING_NONCONSTANT_HOUSEHOLD_COLUMNS) <= set( + US_RELEASE_REQUIRED_HOUSEHOLD_NONCONSTANT_SOURCE_COLUMNS + ) + probe = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "pre_subsidy_rent_neutralization" + ) + assert probe.neutralized_variable == "pre_subsidy_rent" + assert probe.binding_inputs == ("pre_subsidy_rent",) + assert probe.budget_measure == "snap" + assert probe.effect_direction == "baseline_minus_reform" + assert probe.expected_sign == "positive" + take_up_probe = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "housing_assistance_take_up_neutralization" + ) + assert take_up_probe.neutralized_variable == ( + "takes_up_housing_assistance_if_eligible" + ) + assert take_up_probe.binding_inputs == ("takes_up_housing_assistance_if_eligible",) + assert take_up_probe.budget_measure == "housing_assistance" + assert take_up_probe.min_abs_effect == 100_000_000.0 + contract = load_take_up_contract().program_map()[ + "takes_up_housing_assistance_if_eligible" + ] + assert contract.populace_treatment == "out_of_scope" + assert contract.rate == {"status": "not_used_measured_source"} + + +@requires_us +def test_policyengine_us_input_contracts_and_rent_consumer() -> None: + from policyengine_us import CountryTaxBenefitSystem + + system = CountryTaxBenefitSystem() + expected_entities = { + "pre_subsidy_rent": "person", + "receives_housing_assistance": "spm_unit", + "takes_up_housing_assistance_if_eligible": "spm_unit", + "spm_unit_tenure_type": "spm_unit", + "tenure_type": "household", + } + for name, entity in expected_entities.items(): + variable = system.variables[name] + assert variable.is_input_variable() + assert variable.entity.key == entity + assert "pre_subsidy_rent" in system.variables["hud_gross_rent"].adds + + +@requires_us +def test_policyengine_us_live_housing_take_up_neutralization() -> None: + from importlib.metadata import version + + from policyengine_us import Simulation + + from populace.build.us_runtime.reform_coverage_smoke import _build_reform + + assert version("policyengine-us") == "1.764.6" + situation = { + "people": { + "adult": { + "age": {"2024": 40}, + "employment_income": {"2024": 0}, + "pre_subsidy_rent": {"2024": 12_000}, + } + }, + "tax_units": {"tu": {"members": ["adult"]}}, + "families": {"fam": {"members": ["adult"]}}, + "spm_units": { + "spm": { + "members": ["adult"], + "receives_housing_assistance": {"2024": True}, + "takes_up_housing_assistance_if_eligible": {"2024": True}, + "spm_unit_tenure_type": {"2024": "RENTER"}, + } + }, + "households": { + "hh": { + "members": ["adult"], + "state_code": {"2024": "CA"}, + "county_fips": {"2024": "06037"}, + "bedrooms": {"2024": 1}, + } + }, + "marital_units": {"mu": {"members": ["adult"]}}, + } + probe = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "housing_assistance_take_up_neutralization" + ) + + baseline = Simulation(situation=situation) + neutralized = Simulation(situation=situation, reform=_build_reform(probe)) + + assert baseline.calculate("is_eligible_for_housing_assistance", 2024)[0] + assert baseline.calculate("housing_assistance", 2024)[0] == pytest.approx(11_975) + assert neutralized.calculate("housing_assistance", 2024)[0] == 0 diff --git a/packages/populace-build/tests/test_us_l0_refit_export.py b/packages/populace-build/tests/test_us_l0_refit_export.py index 4382b2cd..5c508221 100644 --- a/packages/populace-build/tests/test_us_l0_refit_export.py +++ b/packages/populace-build/tests/test_us_l0_refit_export.py @@ -33,6 +33,10 @@ def _us_frame(**person_extra: object) -> Frame: "weekly_hours_worked_before_lsr": [40.0, 20.0, 0.0], "hours_worked_last_week": [40.0, 18.0, 0.0], "weeks_worked": [52.0, 26.0, 0.0], + "weeks_unemployed": [0.0, 12.0, 4.0], + "is_household_head": [True, True, False], + "is_separated": [False, False, True], + "is_surviving_spouse": [False, True, False], "is_disabled": [False, True, False], "is_blind": [False, False, True], "is_full_time_college_student": [False, False, True], @@ -43,6 +47,98 @@ def _us_frame(**person_extra: object) -> Frame: "bank_account_assets": [5_000.0, 0.0, 1_200.0], "stock_assets": [0.0, 30_000.0, 0.0], "bond_assets": [0.0, 0.0, 800.0], + "tip_income": [2_400.0, 0.0, 600.0], + "treasury_tipped_occupation_code": [101, 0, 304], + "alimony_income": [5_000.0, 0.0, 1_000.0], + "alimony_expense": [0.0, 2_500.0, 0.0], + "child_support_received": [3_600.0, 0.0, 1_200.0], + "child_support_expense": [0.0, 2_400.0, 600.0], + "disability_benefits": [0.0, 5_000.0, 1_500.0], + "workers_compensation": [0.0, 6_000.0, 1_000.0], + "would_claim_wic": [True, False, True], + "educator_expense": [0.0, 300.0, 150.0], + "other_health_insurance_premiums": [1_200.0, 0.0, 600.0], + # Keep both signs without making the fixture's signed weighted + # total nearly cancel; the mass-parity tests reweight household 20. + "farm_operations_income": [4_000.0, 2_500.0, -600.0], + "farm_rent_income": [0.0, 1_500.0, -600.0], + "casualty_loss": [0.0, 2_500.0, 0.0], + "unreimbursed_business_employee_expenses": [1_200.0, 0.0, 800.0], + "investment_income_elected_form_4952": [0.0, 500.0, 250.0], + "salt_refund_income": [0.0, 1_200.0, 400.0], + "long_term_capital_gains_on_collectibles": [0.0, 2_500.0, 1_000.0], + "qualified_tuition_expenses": [1_000.0, 0.0, 2_500.0], + "educational_assistance": [0.0, 500.0, 0.0], + "traditional_401k_contributions_desired": [1_000.0, 0.0, 500.0], + "roth_401k_contributions_desired": [200.0, 0.0, 100.0], + "traditional_ira_contributions_desired": [300.0, 0.0, 150.0], + "roth_ira_contributions_desired": [400.0, 0.0, 200.0], + "self_employed_pension_contributions_desired": [0.0, 800.0, 0.0], + "taxable_401k_distributions": [1_000.0, 0.0, 500.0], + "taxable_403b_distributions": [0.0, 600.0, 0.0], + "tax_exempt_ira_distributions": [300.0, 0.0, 100.0], + "taxable_ira_distributions": [400.0, 0.0, 200.0], + "keogh_distributions": [0.0, 700.0, 0.0], + "taxable_sep_distributions": [0.0, 0.0, 250.0], + "estate_income_would_be_qualified": [True, False, True], + "farm_operations_income_would_be_qualified": [True, True, False], + "farm_rent_income_would_be_qualified": [True, False, True], + "partnership_s_corp_income_would_be_qualified": [True, False, True], + "rental_income_would_be_qualified": [True, False, True], + "self_employment_income_would_be_qualified": [True, True, False], + "sstb_self_employment_income_would_be_qualified": [False, True, False], + "business_is_sstb": [False, True, False], + "qualified_bdc_income": [0.0, 20.0, 0.0], + "qualified_reit_and_ptp_income": [100.0, 0.0, 50.0], + "sstb_self_employment_income_before_lsr": [0.0, 1_000.0, 0.0], + "sstb_unadjusted_basis_qualified_property": [0.0, 5_000.0, 0.0], + "sstb_w2_wages_from_qualified_business": [0.0, 2_000.0, 0.0], + "unadjusted_basis_qualified_property": [1_000.0, 5_000.0, 0.0], + "w2_wages_from_qualified_business": [500.0, 2_000.0, 0.0], + "is_pursuing_credential_for_american_opportunity_credit": [ + True, + False, + True, + ], + "attends_eligible_educational_institution_for_american_opportunity_credit": [ + True, + False, + True, + ], + "is_enrolled_at_least_half_time_for_american_opportunity_credit": [ + True, + False, + True, + ], + "has_american_opportunity_credit_1098_t_or_exception": [ + True, + False, + True, + ], + "has_american_opportunity_credit_institution_ein": [ + True, + False, + True, + ], + "cps_race": [1, 2, 4], + "is_hispanic": [False, True, False], + "detailed_occupation_recode": [1, 20, 53], + "has_never_worked": [False, False, True], + "is_military": [False, True, False], + "is_computer_scientist": [False, True, False], + "is_executive_administrative_professional": [False, True, False], + "is_farmer_fisher": [False, True, False], + "hourly_wage": [25.0, 18.0, 0.0], + "is_paid_hourly": [False, True, False], + "is_union_member_or_covered": [False, True, False], + "fsla_overtime_premium": [0.0, 1_000.0, 0.0], + "takes_up_medicare_if_eligible": [False, True, False], + "pre_subsidy_rent": [0.0, 12_000.0, 0.0], + "self_employment_income_last_year": [0.0, 8_000.0, -1_000.0], + "previous_year_income_available": [False, True, False], + "meets_ssi_disability_criteria": [False, True, False], + "takes_up_head_start_if_eligible": [False, True, False], + "takes_up_ssi_if_eligible": [False, True, False], **person_extra, } ) @@ -63,6 +159,13 @@ def _us_frame(**person_extra: object) -> Frame: "sldu": ["027", "024"], "sldl": ["075", "051"], "cbsa_code": ["35620", "31080"], + "auto_loan_balance": [0.0, 24_000.0], + "auto_loan_interest": [0.0, 1_200.0], + "qualified_passenger_vehicle_loan_interest": [0.0, 240.0], + "net_worth": [-50_000.0, 350_000.0], + "household_vehicles_owned": [0, 2], + "household_vehicles_value": [0.0, 30_000.0], + "tenure_type": ["OWNED_WITH_MORTGAGE", "RENTED"], } ), "tax_unit": pd.DataFrame( @@ -74,10 +177,25 @@ def _us_frame(**person_extra: object) -> Frame: "selected_marketplace_plan_benchmark_ratio": np.asarray( [1.0, 0.8, 1.2], dtype="float64" ), + "would_file_taxes_voluntarily": np.asarray( + [True, False, True], dtype=bool + ), + "domestic_production_ald": [0.0, 10_000.0, 5_000.0], + "unrecaptured_section_1250_gain": [0.0, 2_000.0, 500.0], } ), "spm_unit": pd.DataFrame( - {"spm_unit_id": np.asarray([1000, 2000], dtype="int64")} + { + "spm_unit_id": np.asarray([1000, 2000], dtype="int64"), + "spm_unit_pre_subsidy_childcare_expenses": [0.0, 1_200.0], + "spm_unit_energy_subsidy": [0.0, 600.0], + "receives_housing_assistance": [False, True], + "takes_up_housing_assistance_if_eligible": [False, True], + "spm_unit_tenure_type": [ + "OWNER_WITH_MORTGAGE", + "RENTER", + ], + } ), "family": pd.DataFrame( {"family_id": np.asarray([10000, 20000], dtype="int64")} @@ -386,6 +504,485 @@ def test_required_us_release_source_columns_rejects_missing_geography_spine() -> assert_required_us_release_source_columns(raw_frame) +@pytest.mark.parametrize( + "column", + [ + "auto_loan_balance", + "auto_loan_interest", + "qualified_passenger_vehicle_loan_interest", + "net_worth", + "household_vehicles_owned", + "household_vehicles_value", + ], +) +def test_required_us_release_source_columns_rejects_missing_auto_input( + column: str, +) -> None: + frame = _us_frame() + raw_households = frame.table("household").drop(columns=[column]) + raw_frame = Frame( + { + **{entity: frame.table(entity).copy() for entity in frame.schema.entities}, + "household": raw_households, + }, + frame.schema, + {"household": frame.weights_for("household")}, + ) + + with pytest.raises(ValueError, match=rf"household.{column}: missing"): + assert_required_us_release_source_columns(raw_frame) + + +def test_required_us_release_source_columns_rejects_constant_auto_input() -> None: + frame = _us_frame() + raw_households = frame.table("household").copy() + raw_households["qualified_passenger_vehicle_loan_interest"] = 0.0 + raw_frame = Frame( + { + **{entity: frame.table(entity).copy() for entity in frame.schema.entities}, + "household": raw_households, + }, + frame.schema, + {"household": frame.weights_for("household")}, + ) + + with pytest.raises( + ValueError, + match="household.qualified_passenger_vehicle_loan_interest: not nonconstant", + ): + assert_required_us_release_source_columns(raw_frame) + + +def test_required_us_release_source_columns_rejects_constant_net_worth() -> None: + frame = _us_frame() + raw_households = frame.table("household").copy() + raw_households["net_worth"] = 0.0 + raw_frame = Frame( + { + **{entity: frame.table(entity).copy() for entity in frame.schema.entities}, + "household": raw_households, + }, + frame.schema, + {"household": frame.weights_for("household")}, + ) + + with pytest.raises( + ValueError, + match="household.net_worth: not nonconstant", + ): + assert_required_us_release_source_columns(raw_frame) + + +@pytest.mark.parametrize( + "column", + ["household_vehicles_owned", "household_vehicles_value"], +) +def test_required_us_release_source_columns_rejects_constant_vehicle_input( + column: str, +) -> None: + frame = _us_frame() + raw_households = frame.table("household").copy() + raw_households[column] = 0 + raw_frame = Frame( + { + **{entity: frame.table(entity).copy() for entity in frame.schema.entities}, + "household": raw_households, + }, + frame.schema, + {"household": frame.weights_for("household")}, + ) + + with pytest.raises( + ValueError, + match=rf"household.{column}: not nonconstant", + ): + assert_required_us_release_source_columns(raw_frame) + + +def test_required_us_release_source_columns_enforces_casualty_loss_signal() -> None: + frame = _us_frame() + raw_people = frame.table("person").copy() + raw_people["casualty_loss"] = 0.0 + raw_frame = Frame( + { + **{entity: frame.table(entity).copy() for entity in frame.schema.entities}, + "person": raw_people, + }, + frame.schema, + {"household": frame.weights_for("household")}, + ) + + with pytest.raises( + ValueError, + match="person.casualty_loss: not nonconstant", + ): + assert_required_us_release_source_columns(raw_frame) + + +def test_required_us_release_source_columns_enforces_form_4952_signal() -> None: + frame = _us_frame() + raw_people = frame.table("person").copy() + raw_people["investment_income_elected_form_4952"] = 0.0 + raw_frame = Frame( + { + **{entity: frame.table(entity).copy() for entity in frame.schema.entities}, + "person": raw_people, + }, + frame.schema, + {"household": frame.weights_for("household")}, + ) + + with pytest.raises( + ValueError, + match="person.investment_income_elected_form_4952: not nonconstant", + ): + assert_required_us_release_source_columns(raw_frame) + + +@pytest.mark.parametrize( + "column", + [ + "taxable_401k_distributions", + "taxable_403b_distributions", + "tax_exempt_ira_distributions", + "taxable_ira_distributions", + "keogh_distributions", + "taxable_sep_distributions", + ], +) +def test_required_us_release_source_columns_enforces_retirement_distribution_signal( + column: str, +) -> None: + frame = _us_frame() + raw_people = frame.table("person").copy() + raw_people[column] = 0.0 + raw_frame = Frame( + { + **{entity: frame.table(entity).copy() for entity in frame.schema.entities}, + "person": raw_people, + }, + frame.schema, + {"household": frame.weights_for("household")}, + ) + + with pytest.raises(ValueError, match=rf"person.{column}: not nonconstant"): + assert_required_us_release_source_columns(raw_frame) + + +def test_required_us_release_source_columns_enforces_domestic_production_signal() -> ( + None +): + frame = _us_frame() + raw_tax_units = frame.table("tax_unit").copy() + raw_tax_units["domestic_production_ald"] = 0.0 + raw_frame = Frame( + { + **{entity: frame.table(entity).copy() for entity in frame.schema.entities}, + "tax_unit": raw_tax_units, + }, + frame.schema, + {"household": frame.weights_for("household")}, + ) + + with pytest.raises( + ValueError, + match="tax_unit.domestic_production_ald: not nonconstant", + ): + assert_required_us_release_source_columns(raw_frame) + + +def test_required_us_release_source_columns_enforces_voluntary_filing_signal() -> None: + frame = _us_frame() + raw_tax_units = frame.table("tax_unit").copy() + raw_tax_units["would_file_taxes_voluntarily"] = False + raw_frame = Frame( + { + **{entity: frame.table(entity).copy() for entity in frame.schema.entities}, + "tax_unit": raw_tax_units, + }, + frame.schema, + {"household": frame.weights_for("household")}, + ) + + with pytest.raises( + ValueError, + match="tax_unit.would_file_taxes_voluntarily: not nonconstant", + ): + assert_required_us_release_source_columns(raw_frame) + + +@pytest.mark.parametrize("column", ["alimony_income", "alimony_expense"]) +def test_required_us_release_source_columns_enforces_alimony_signal( + column: str, +) -> None: + frame = _us_frame() + raw_people = frame.table("person").copy() + raw_people[column] = 0.0 + raw_frame = Frame( + { + **{entity: frame.table(entity).copy() for entity in frame.schema.entities}, + "person": raw_people, + }, + frame.schema, + {"household": frame.weights_for("household")}, + ) + + with pytest.raises(ValueError, match=rf"person.{column}: not nonconstant"): + assert_required_us_release_source_columns(raw_frame) + + +@pytest.mark.parametrize( + "column", + ["child_support_received", "child_support_expense"], +) +def test_required_us_release_source_columns_enforces_child_support_signal( + column: str, +) -> None: + frame = _us_frame() + raw_people = frame.table("person").copy() + raw_people[column] = 0.0 + raw_frame = Frame( + { + **{entity: frame.table(entity).copy() for entity in frame.schema.entities}, + "person": raw_people, + }, + frame.schema, + {"household": frame.weights_for("household")}, + ) + + with pytest.raises(ValueError, match=rf"person.{column}: not nonconstant"): + assert_required_us_release_source_columns(raw_frame) + + +def test_required_us_release_source_columns_enforces_disability_benefits_signal() -> ( + None +): + frame = _us_frame() + raw_people = frame.table("person").copy() + raw_people["disability_benefits"] = 0.0 + raw_frame = Frame( + { + **{entity: frame.table(entity).copy() for entity in frame.schema.entities}, + "person": raw_people, + }, + frame.schema, + {"household": frame.weights_for("household")}, + ) + + with pytest.raises( + ValueError, + match="person.disability_benefits: not nonconstant", + ): + assert_required_us_release_source_columns(raw_frame) + + +def test_required_us_release_source_columns_enforces_ssi_disability_criteria_signal() -> ( + None +): + frame = _us_frame() + raw_people = frame.table("person").copy() + raw_people["meets_ssi_disability_criteria"] = False + raw_frame = Frame( + { + **{entity: frame.table(entity).copy() for entity in frame.schema.entities}, + "person": raw_people, + }, + frame.schema, + {"household": frame.weights_for("household")}, + ) + + with pytest.raises( + ValueError, + match="person.meets_ssi_disability_criteria: not nonconstant", + ): + assert_required_us_release_source_columns(raw_frame) + + +def test_required_us_release_source_columns_enforces_ssi_take_up_signal() -> None: + frame = _us_frame() + raw_people = frame.table("person").copy() + raw_people["takes_up_ssi_if_eligible"] = False + raw_frame = Frame( + { + **{entity: frame.table(entity).copy() for entity in frame.schema.entities}, + "person": raw_people, + }, + frame.schema, + {"household": frame.weights_for("household")}, + ) + + with pytest.raises( + ValueError, + match="person.takes_up_ssi_if_eligible: not nonconstant", + ): + assert_required_us_release_source_columns(raw_frame) + + +def test_required_us_release_source_columns_enforces_head_start_take_up_signal() -> ( + None +): + frame = _us_frame() + raw_people = frame.table("person").copy() + raw_people["takes_up_head_start_if_eligible"] = False + raw_frame = Frame( + { + **{entity: frame.table(entity).copy() for entity in frame.schema.entities}, + "person": raw_people, + }, + frame.schema, + {"household": frame.weights_for("household")}, + ) + + with pytest.raises( + ValueError, + match="person.takes_up_head_start_if_eligible: not nonconstant", + ): + assert_required_us_release_source_columns(raw_frame) + + +def test_required_us_release_source_columns_enforces_weeks_unemployed() -> None: + frame = _us_frame() + raw_people = frame.table("person").copy() + raw_people["weeks_unemployed"] = 0.0 + raw_frame = Frame( + { + **{entity: frame.table(entity).copy() for entity in frame.schema.entities}, + "person": raw_people, + }, + frame.schema, + {"household": frame.weights_for("household")}, + ) + + with pytest.raises( + ValueError, + match="person.weeks_unemployed: not nonconstant", + ): + assert_required_us_release_source_columns(raw_frame) + + +def test_required_us_release_source_columns_enforces_educator_expense_signal() -> None: + frame = _us_frame() + raw_people = frame.table("person").copy() + raw_people["educator_expense"] = 0.0 + raw_frame = Frame( + { + **{entity: frame.table(entity).copy() for entity in frame.schema.entities}, + "person": raw_people, + }, + frame.schema, + {"household": frame.weights_for("household")}, + ) + + with pytest.raises( + ValueError, + match="person.educator_expense: not nonconstant", + ): + assert_required_us_release_source_columns(raw_frame) + + +def test_required_us_release_source_columns_enforces_other_premium_signal() -> None: + frame = _us_frame() + raw_people = frame.table("person").copy() + raw_people["other_health_insurance_premiums"] = 0.0 + raw_frame = Frame( + { + **{entity: frame.table(entity).copy() for entity in frame.schema.entities}, + "person": raw_people, + }, + frame.schema, + {"household": frame.weights_for("household")}, + ) + + with pytest.raises( + ValueError, + match="person.other_health_insurance_premiums: not nonconstant", + ): + assert_required_us_release_source_columns(raw_frame) + + +@pytest.mark.parametrize( + "column", + ["farm_operations_income", "farm_rent_income"], +) +def test_required_us_release_source_columns_enforces_farm_business_signal( + column: str, +) -> None: + frame = _us_frame() + raw_people = frame.table("person").copy() + raw_people[column] = 0.0 + raw_frame = Frame( + { + **{entity: frame.table(entity).copy() for entity in frame.schema.entities}, + "person": raw_people, + }, + frame.schema, + {"household": frame.weights_for("household")}, + ) + + with pytest.raises(ValueError, match=rf"person.{column}: not nonconstant"): + assert_required_us_release_source_columns(raw_frame) + + +def test_required_us_release_source_columns_enforces_misc_itemized_signal() -> None: + frame = _us_frame() + raw_people = frame.table("person").copy() + raw_people["unreimbursed_business_employee_expenses"] = 0.0 + raw_frame = Frame( + { + **{entity: frame.table(entity).copy() for entity in frame.schema.entities}, + "person": raw_people, + }, + frame.schema, + {"household": frame.weights_for("household")}, + ) + + with pytest.raises( + ValueError, + match=("person.unreimbursed_business_employee_expenses: not nonconstant"), + ): + assert_required_us_release_source_columns(raw_frame) + + +def test_required_us_release_source_columns_enforces_childcare_signal() -> None: + frame = _us_frame() + raw_spm_units = frame.table("spm_unit").copy() + raw_spm_units["spm_unit_pre_subsidy_childcare_expenses"] = 0.0 + raw_frame = Frame( + { + **{entity: frame.table(entity).copy() for entity in frame.schema.entities}, + "spm_unit": raw_spm_units, + }, + frame.schema, + {"household": frame.weights_for("household")}, + ) + + with pytest.raises( + ValueError, + match=("spm_unit.spm_unit_pre_subsidy_childcare_expenses: not nonconstant"), + ): + assert_required_us_release_source_columns(raw_frame) + + +def test_required_us_release_source_columns_enforces_energy_subsidy_signal() -> None: + frame = _us_frame() + raw_spm_units = frame.table("spm_unit").copy() + raw_spm_units["spm_unit_energy_subsidy"] = 0.0 + raw_frame = Frame( + { + **{entity: frame.table(entity).copy() for entity in frame.schema.entities}, + "spm_unit": raw_spm_units, + }, + frame.schema, + {"household": frame.weights_for("household")}, + ) + + with pytest.raises( + ValueError, + match="spm_unit.spm_unit_energy_subsidy: not nonconstant", + ): + assert_required_us_release_source_columns(raw_frame) + + def test_export_us_l0_refit_h5_fails_geography_ladder_gate_by_default( monkeypatch: pytest.MonkeyPatch, tmp_path: Path, @@ -448,3 +1045,50 @@ def write_dataset(self, bundle, path, period): assert summary["geography_ladder_gate_enforced"] is False assert summary["geography_ladder_gate"]["passed"] is False assert summary["required_household_source_columns"][0] == "state_fips" + assert "alimony_income" in summary["required_person_source_columns"] + assert "alimony_expense" in summary["required_person_source_columns"] + assert "casualty_loss" in summary["required_person_source_columns"] + assert "child_support_received" in summary["required_person_source_columns"] + assert "child_support_expense" in summary["required_person_source_columns"] + assert "disability_benefits" in summary["required_person_source_columns"] + assert "meets_ssi_disability_criteria" in summary["required_person_source_columns"] + assert ( + "takes_up_head_start_if_eligible" in summary["required_person_source_columns"] + ) + assert "takes_up_ssi_if_eligible" in summary["required_person_source_columns"] + assert "weeks_unemployed" in summary["required_person_source_columns"] + assert "educator_expense" in summary["required_person_source_columns"] + assert ( + "investment_income_elected_form_4952" + in summary["required_person_source_columns"] + ) + assert ( + "other_health_insurance_premiums" in summary["required_person_source_columns"] + ) + assert "farm_operations_income" in summary["required_person_source_columns"] + assert "farm_rent_income" in summary["required_person_source_columns"] + assert "keogh_distributions" in summary["required_person_source_columns"] + assert "business_is_sstb" in summary["required_person_source_columns"] + assert "qualified_reit_and_ptp_income" in summary["required_person_source_columns"] + assert "domestic_production_ald" in summary["required_source_columns"] + assert "would_file_taxes_voluntarily" in summary["required_source_columns"] + assert ( + "unreimbursed_business_employee_expenses" + in summary["required_person_source_columns"] + ) + assert summary["required_spm_unit_source_columns"] == [ + "spm_unit_pre_subsidy_childcare_expenses", + "spm_unit_energy_subsidy", + "receives_housing_assistance", + "takes_up_housing_assistance_if_eligible", + "spm_unit_tenure_type", + ] + assert summary["required_household_nonconstant_source_columns"] == [ + "auto_loan_balance", + "auto_loan_interest", + "qualified_passenger_vehicle_loan_interest", + "net_worth", + "household_vehicles_owned", + "household_vehicles_value", + "tenure_type", + ] diff --git a/packages/populace-build/tests/test_us_medicaid_take_up.py b/packages/populace-build/tests/test_us_medicaid_take_up.py index 0bb93976..b917d7c5 100644 --- a/packages/populace-build/tests/test_us_medicaid_take_up.py +++ b/packages/populace-build/tests/test_us_medicaid_take_up.py @@ -109,7 +109,8 @@ def _scenario(n: int = 400, seed: int = 7): class TestContractTreatment: def test_medicaid_is_count_calibrated_not_seeded(self) -> None: assert [p.variable for p in count_calibrated_take_up_programs()] == [ - "takes_up_medicaid_if_eligible" + "takes_up_medicaid_if_eligible", + "takes_up_ssi_if_eligible", ] assert "takes_up_medicaid_if_eligible" not in { p.variable for p in seeded_take_up_programs() diff --git a/packages/populace-build/tests/test_us_medicare_take_up.py b/packages/populace-build/tests/test_us_medicare_take_up.py new file mode 100644 index 00000000..11441b8d --- /dev/null +++ b/packages/populace-build/tests/test_us_medicare_take_up.py @@ -0,0 +1,436 @@ +"""Measured ASEC Medicare take-up restoration.""" + +from __future__ import annotations + +import importlib.util +import json +from hashlib import sha256 +from importlib.metadata import version +from importlib.resources import files +from pathlib import Path + +import numpy as np +import pandas as pd +import pytest + +from populace.build.source_manifest import SourceOperationSpec +from populace.build.source_runtime import SourceRuntimeError +from populace.build.us_runtime import ( + MEDICARE_TAKE_UP_ARCHIVED_CLONE_URL, + MEDICARE_TAKE_UP_ARCHIVED_DERIVATION_URL, + MEDICARE_TAKE_UP_ARCHIVED_EXPORT_URL, + MEDICARE_TAKE_UP_ARCHIVED_SOURCE_COLUMNS_URL, + PUF_TAX_DETAIL_SUPPORT_CHANNEL, + US_DONORS, + US_MEDICARE_TAKE_UP_NONCONSTANT_PERSON_COLUMNS, + US_MEDICARE_TAKE_UP_OUTPUT_COLUMNS, + US_MEDICARE_TAKE_UP_REQUIRED_SOURCE_COLUMNS, + US_MEDICARE_TAKE_UP_STAGE_NAME, + US_PUF_SUPPORT_STAGE_NAME, + US_STAGE_NAMES, + clone_us_frame_for_puf_support, + derive_us_medicare_take_up_from_manifest, + load_release_input_coverage_manifest, + us_medicare_take_up_signal_gate, + us_medicare_take_up_stage_spec, + us_medicare_take_up_summary, + us_release_reform_coverage_probes, + with_us_medicare_take_up_input, +) +from populace.build.us_runtime.l0_refit_export import ( + US_RELEASE_REQUIRED_PERSON_SOURCE_COLUMNS, +) +from populace.build.us_runtime.release_input_coverage import ( + RESTORED_REFERENCE_ECPS_REQUIRED_INPUTS, +) +from populace.build.us_runtime.source_runtime import us_source_operation_handlers +from populace.build.us_runtime.take_up_contract import load_take_up_contract +from populace.frame import US_SCHEMA, EntitySchema, Frame, WeightKind, Weights + +ROOT = Path(__file__).resolve().parents[3] +policyengine_us_installed = importlib.util.find_spec("policyengine_us") is not None +requires_us = pytest.mark.skipif( + not policyengine_us_installed, + reason="requires the policyengine-us [us] extra (build environment)", +) + +_OUTPUT = "takes_up_medicare_if_eligible" +_SOURCE = "MCARE" + + +def _frame( + source_codes: list[object] | None = None, + *, + output: list[object] | None = None, + weights: list[float] | None = None, +) -> Frame: + codes = source_codes or [1, 2, 2, 2, 0] + count = len(codes) + household_ids = np.arange(1, count + 1, dtype=np.int64) + person = pd.DataFrame( + { + "person_id": household_ids, + "person_household_id": household_ids, + "person_tax_unit_id": household_ids + 100, + "person_spm_unit_id": household_ids + 200, + "person_family_id": household_ids + 300, + "person_marital_unit_id": household_ids + 400, + _SOURCE: codes, + "age": [70, 66, 40, 20, 10][:count], + } + ) + if output is not None: + person[_OUTPUT] = output + tables = { + "person": person, + "household": pd.DataFrame({"household_id": household_ids}), + "tax_unit": pd.DataFrame({"tax_unit_id": household_ids + 100}), + "spm_unit": pd.DataFrame({"spm_unit_id": household_ids + 200}), + "family": pd.DataFrame({"family_id": household_ids + 300}), + "marital_unit": pd.DataFrame({"marital_unit_id": household_ids + 400}), + } + return Frame( + tables, + US_SCHEMA, + { + "household": Weights( + np.asarray(weights or [1.0] * count, dtype=np.float64), + WeightKind.DESIGN, + ) + }, + ) + + +def _operation() -> SourceOperationSpec: + return next( + operation + for operation in us_medicare_take_up_stage_spec().operations + if operation.kind == "derive_medicare_take_up" + ) + + +def _derive(person: pd.DataFrame) -> pd.DataFrame: + return derive_us_medicare_take_up_from_manifest(person, _operation(), None) + + +def _sha256(path: Path) -> str: + digest = sha256() + with path.open("rb") as stream: + for chunk in iter(lambda: stream.read(1024 * 1024), b""): + digest.update(chunk) + return digest.hexdigest() + + +class TestManifestAndPlan: + def test_stage_pins_exact_measured_mapping_and_clone_semantics(self) -> None: + spec = us_medicare_take_up_stage_spec() + + assert spec.stage == US_MEDICARE_TAKE_UP_STAGE_NAME + assert tuple(spec.outputs) == US_MEDICARE_TAKE_UP_OUTPUT_COLUMNS == (_OUTPUT,) + assert US_MEDICARE_TAKE_UP_NONCONSTANT_PERSON_COLUMNS == (_OUTPUT,) + assert US_MEDICARE_TAKE_UP_REQUIRED_SOURCE_COLUMNS == (_SOURCE,) + assert [operation.kind for operation in spec.operations] == [ + "read_table", + "derive_medicare_take_up", + ] + assert _operation().parameters == { + "source": _SOURCE, + "enrolled_code": 1, + "output": _OUTPUT, + } + assert "MCARE == 1" in spec.notes + assert "No take-up rate or stochastic draw" in spec.notes + + def test_archived_coordinates_are_immutable(self) -> None: + assert MEDICARE_TAKE_UP_ARCHIVED_DERIVATION_URL.endswith( + "/datasets/cps/cps.py#L1579-L1585" + ) + assert MEDICARE_TAKE_UP_ARCHIVED_SOURCE_COLUMNS_URL.endswith( + "/datasets/cps/census_cps.py#L39-L58" + ) + assert MEDICARE_TAKE_UP_ARCHIVED_CLONE_URL.endswith( + "/calibration/puf_impute.py#L608-L629" + ) + assert MEDICARE_TAKE_UP_ARCHIVED_EXPORT_URL.endswith( + "/datasets/cps/extended_cps.py#L1747-L1754" + ) + + def test_handler_and_plan_are_wired_before_puf_support(self) -> None: + handlers = us_source_operation_handlers() + assert ( + handlers["derive_medicare_take_up"] + is derive_us_medicare_take_up_from_manifest + ) + assert US_MEDICARE_TAKE_UP_STAGE_NAME in US_DONORS + assert US_STAGE_NAMES.index(US_MEDICARE_TAKE_UP_STAGE_NAME) < ( + US_STAGE_NAMES.index(US_PUF_SUPPORT_STAGE_NAME) + ) + + +class TestDerivation: + def test_maps_only_mcare_code_one_to_true(self) -> None: + person = _frame([0, 1, 2]).table("person") + + result = _derive(person) + + assert result[_OUTPUT].tolist() == [False, True, False] + assert result[_OUTPUT].dtype == bool + + @pytest.mark.parametrize("invalid", [np.nan, 1.5, -1, 3, "unknown"]) + def test_invalid_source_codes_fail_closed(self, invalid: object) -> None: + person = _frame([1, invalid, 2]).table("person") + with pytest.raises(SourceRuntimeError, match="MCARE"): + _derive(person) + + def test_missing_source_wrong_operation_and_parameters_are_rejected(self) -> None: + person = _frame().table("person") + with pytest.raises(SourceRuntimeError, match="MCARE"): + _derive(person.drop(columns=[_SOURCE])) + with pytest.raises(SourceRuntimeError, match="unexpected operation"): + derive_us_medicare_take_up_from_manifest( + person, + SourceOperationSpec(kind="wrong", parameters={}), + None, + ) + with pytest.raises(SourceRuntimeError, match="requires the person table"): + derive_us_medicare_take_up_from_manifest(None, _operation(), None) + with pytest.raises(SourceRuntimeError, match="drifted"): + derive_us_medicare_take_up_from_manifest( + person, + SourceOperationSpec( + kind="derive_medicare_take_up", + parameters={"source": _SOURCE}, + ), + None, + ) + + +class TestFrameStageAndGate: + def test_materializes_exact_source_and_is_idempotent(self) -> None: + frame = _frame() + derived = with_us_medicare_take_up_input(frame, seed=0, time_period=2024) + + assert derived.table("person")[_OUTPUT].tolist() == [ + True, + False, + False, + False, + False, + ] + assert ( + with_us_medicare_take_up_input(derived, seed=99, time_period=2026) + is derived + ) + assert us_medicare_take_up_signal_gate(derived).passed + + def test_stale_nonconstant_output_is_rederived_from_source(self) -> None: + frame = _frame(output=[False, True, False, False, False]) + + derived = with_us_medicare_take_up_input(frame, seed=0, time_period=2024) + + assert derived.table("person")[_OUTPUT].tolist() == [ + True, + False, + False, + False, + False, + ] + assert us_medicare_take_up_summary(derived)["source_mismatch_count"] == 0 + + def test_support_cloning_preserves_measured_values_on_both_channels(self) -> None: + derived = with_us_medicare_take_up_input(_frame(), seed=0, time_period=2024) + cloned = clone_us_frame_for_puf_support(derived) + person = cloned.table("person") + + assert ( + person[_OUTPUT].tolist() + == [ + True, + False, + False, + False, + False, + ] + * 2 + ) + summary = us_medicare_take_up_summary(cloned) + assert summary["source_mismatch_count"] == 0 + assert summary["channel_weighted_enrolled_shares"] == { + "asec": pytest.approx(0.2), + PUF_TAX_DETAIL_SUPPORT_CHANNEL: pytest.approx(0.2), + } + assert us_medicare_take_up_signal_gate(cloned).passed + + def test_gate_rejects_missing_constant_bad_share_and_mismatch(self) -> None: + missing = _frame() + assert not us_medicare_take_up_signal_gate(missing).passed + + constant = _frame(output=[True] * 5) + assert not us_medicare_take_up_signal_gate(constant).passed + + bad_share = _frame( + source_codes=[1, 1, 1, 1, 2], + output=[True, True, True, True, False], + ) + assert not us_medicare_take_up_signal_gate(bad_share).passed + + mismatch = _frame(output=[False, True, False, False, False]) + gate = us_medicare_take_up_signal_gate(mismatch) + assert not gate.passed + assert any("reconciliation mismatch" in failure for failure in gate.failures) + + def test_non_us_schema_is_rejected(self) -> None: + frame = Frame( + { + "person": pd.DataFrame( + {"person_id": [1], "person_household_id": [1], _SOURCE: [1]} + ), + "household": pd.DataFrame({"household_id": [1]}), + }, + EntitySchema(group_entities=("household",)), + {"household": Weights(np.ones(1), WeightKind.DESIGN)}, + ) + with pytest.raises(ValueError, match="US Medicare"): + with_us_medicare_take_up_input(frame, seed=0, time_period=2024) + + +@requires_us +def test_all_sha_locked_asec_sources_have_exact_measured_signal() -> None: + from policyengine_us.data import USSingleYearDataset + + expected = { + "census_cps_2022.h5": { + "sha256": "7ccca976284bb47815d84460cc4f75a0a65d26d7754ab0a0f417de351b3d474e", + "positive": 26495, + "weighted_share": 0.18617172806991097, + }, + "census_cps_2023.h5": { + "sha256": "cb57817327799f42b741caed5f9be94d04021c2e6809c1ad7bd0686da5428d88", + "positive": 26466, + "weighted_share": 0.187909801755554, + }, + "census_cps_2024.h5": { + "sha256": "ec36604cb735a660b51b0b2f90be27d803b5878f3464fb30d0eacead59c1260d", + "positive": 26448, + "weighted_share": 0.19082998028041492, + }, + } + summary = json.loads( + (ROOT / "experiments/build_j_recert/base_j.summary.json").read_text() + ) + paths = { + Path(item["path"]).name: Path(item["path"]) + for item in summary["base_source"]["sources"] + } + if not all(path.is_file() for path in paths.values()): + pytest.skip("SHA-locked ASEC artifacts are not mounted") + + assert set(paths) == set(expected) + for filename, facts in expected.items(): + path = paths[filename] + assert _sha256(path) == facts["sha256"] + person = USSingleYearDataset(file_path=str(path)).person + codes = pd.to_numeric(person[_SOURCE], errors="coerce") + weights = pd.to_numeric(person["A_FNLWGT"], errors="coerce") / 100.0 + enrolled = codes == 1 + assert set(codes.unique()) == {0, 1, 2} + assert int(enrolled.sum()) == facts["positive"] + assert float(weights[enrolled].sum() / weights.sum()) == pytest.approx( + facts["weighted_share"] + ) + derived = _derive(person) + np.testing.assert_array_equal(derived[_OUTPUT].to_numpy(), enrolled.to_numpy()) + + +@requires_us +def test_policyengine_contract_and_live_neutralization() -> None: + from policyengine_us import CountryTaxBenefitSystem, Simulation + + from populace.build.us_runtime.reform_coverage_smoke import _build_reform + + assert version("policyengine-us") == "1.764.6" + system = CountryTaxBenefitSystem() + variable = system.variables[_OUTPUT] + assert variable.is_input_variable() + assert variable.entity.key == "person" + assert variable.value_type is bool + assert bool(variable.default_value) is True + assert str(variable.definition_period).lower() == "year" + assert system.variables["medicare_enrolled"].adds == [_OUTPUT] + + situation = { + "people": { + "adult": { + "age": {"2024": 70}, + _OUTPUT: {"2024": True}, + } + }, + "tax_units": {"tax_unit": {"members": ["adult"]}}, + "families": {"family": {"members": ["adult"]}}, + "spm_units": {"spm_unit": {"members": ["adult"]}}, + "households": { + "household": { + "members": ["adult"], + "state_code": {"2024": "CA"}, + } + }, + "marital_units": {"marital_unit": {"members": ["adult"]}}, + } + + probe = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "medicare_take_up_neutralization" + ) + baseline = Simulation(situation=situation) + neutralized = Simulation(situation=situation, reform=_build_reform(probe)) + assert baseline.calculate("medicare_enrolled", 2024)[0] + assert not neutralized.calculate("medicare_enrolled", 2024)[0] + assert baseline.calculate("medicare_cost", 2024)[0] > 0 + assert neutralized.calculate("medicare_cost", 2024)[0] == 0 + + +def test_release_contract_promotion_probe_and_take_up_inventory() -> None: + manifest = load_release_input_coverage_manifest() + assert _OUTPUT in RESTORED_REFERENCE_ECPS_REQUIRED_INPUTS + assert _OUTPUT in manifest.required_columns + assert _OUTPUT not in manifest.reviewed_exclusions + assert _OUTPUT in US_RELEASE_REQUIRED_PERSON_SOURCE_COLUMNS + + probe = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "medicare_take_up_neutralization" + ) + assert probe.neutralized_variable == _OUTPUT + assert probe.binding_inputs == (_OUTPUT,) + assert probe.budget_measure == "medicare_cost" + assert probe.effect_direction == "baseline_minus_reform" + assert probe.expected_sign == "positive" + assert probe.min_abs_effect == 1_000_000_000.0 + + contract = load_take_up_contract().program_map()[_OUTPUT] + assert contract.populace_treatment == "out_of_scope" + assert contract.rate == {"status": "not_used_measured_source"} + assert "MCARE == 1" in str(contract.raw["notes"]) + + +def test_both_release_builders_run_stage_and_gate() -> None: + support_builder = (ROOT / "tools/build_us_puf_support_base.py").read_text() + fiscal_builder = (ROOT / "tools/build_us_fiscal_refresh_release.py").read_text() + cache_driver = (ROOT / "experiments/build_j_recert/buildj_base.sh").read_text() + + assert support_builder.index("with_us_medicare_take_up_input(") < ( + support_builder.index("clone_us_frame_for_puf_support(base)") + ) + assert support_builder.count("us_medicare_take_up_signal_gate(") == 2 + assert "with_us_medicare_take_up_input(" in fiscal_builder + assert "us_medicare_take_up_signal_gate(" in fiscal_builder + assert f'"{_OUTPUT}"' in cache_driver + + +def test_generated_manifest_drops_the_retired_generic_gap() -> None: + known_gaps = json.loads( + files("populace.build.us").joinpath("ecps_parity_known_gaps.json").read_text() + )["known_gaps"] + assert _OUTPUT not in known_gaps diff --git a/packages/populace-build/tests/test_us_misc_itemized.py b/packages/populace-build/tests/test_us_misc_itemized.py new file mode 100644 index 00000000..51121429 --- /dev/null +++ b/packages/populace-build/tests/test_us_misc_itemized.py @@ -0,0 +1,166 @@ +"""Contracts for the IRS PUF miscellaneous-itemized input restoration.""" + +from __future__ import annotations + +import importlib.util + +import numpy as np +import pandas as pd +import pytest + +from populace.build.us_runtime import puf_support as puf_support_module +from populace.build.us_runtime.misc_itemized import ( + derive_us_misc_itemized_from_puf, + us_misc_itemized_signal_gate, + us_misc_itemized_stage_spec, +) +from populace.build.us_runtime.puf_aggregate_records import ( + _reconcile_puf_misc_itemized_from_source, +) + +policyengine_us_installed = importlib.util.find_spec("policyengine_us") is not None +requires_us = pytest.mark.skipif( + not policyengine_us_installed, + reason="requires the policyengine-us [us] extra (build environment)", +) + + +class _ResolvedWeights: + def __init__(self, values: np.ndarray) -> None: + self.values = values + + +class _PersonFrame: + def __init__( + self, + person: pd.DataFrame, + weights: np.ndarray | None = None, + ) -> None: + self._person = person + self._weights = np.ones(len(person)) if weights is None else weights + + def table(self, entity: str) -> pd.DataFrame: + assert entity == "person" + return self._person + + def resolve_weights(self, entity: str) -> _ResolvedWeights: + assert entity == "person" + return _ResolvedWeights(np.asarray(self._weights, dtype=np.float64)) + + +def test_archived_e20400_proxy_is_an_exact_carry() -> None: + source = pd.DataFrame({"E20400": [0.0, 125.5, 9_000.0]}) + + result = derive_us_misc_itemized_from_puf(source) + + assert "unreimbursed_business_employee_expenses" not in source.columns + assert result["unreimbursed_business_employee_expenses"].tolist() == [ + 0.0, + 125.5, + 9_000.0, + ] + + +@pytest.mark.parametrize( + ("source", "message"), + [ + (pd.DataFrame({"other": [1.0]}), "requires source column"), + (pd.DataFrame({"E20400": ["not numeric"]}), "nonnumeric or nonfinite"), + (pd.DataFrame({"E20400": [np.inf]}), "nonnumeric or nonfinite"), + (pd.DataFrame({"E20400": [-1.0]}), "negative value"), + ], +) +def test_source_derivation_fails_closed( + source: pd.DataFrame, + message: str, +) -> None: + with pytest.raises(ValueError, match=message): + derive_us_misc_itemized_from_puf(source) + + +def test_post_disaggregation_reconciliation_uses_final_e20400() -> None: + output = "unreimbursed_business_employee_expenses" + source = pd.DataFrame( + { + "E20400": [0.0, 3_000.0, 7_500.0], + output: [999.0, 999.0, 999.0], + } + ) + + result = _reconcile_puf_misc_itemized_from_source(source) + + assert result[output].tolist() == [0.0, 3_000.0, 7_500.0] + + +def test_shared_puf_stage_declares_exact_source_and_output() -> None: + stage = us_misc_itemized_stage_spec() + operation = next( + operation + for operation in stage.operations + if operation.kind == "derive_puf_policyengine_variables" + ) + + assert ( + operation.parameters["unreimbursed_business_employee_expenses_source"] + == "E20400" + ) + assert ( + operation.parameters["unreimbursed_business_employee_expenses_output"] + == "unreimbursed_business_employee_expenses" + ) + assert "unreimbursed_business_employee_expenses" in stage.outputs + assert "unreimbursed_business_employee_expenses" in stage.nonnegative_outputs + assert any("E20400" in str(artifact.get("locator")) for artifact in stage.artifacts) + + +def test_puf_support_keeps_misc_itemized_dense_and_nonnegative() -> None: + output = "unreimbursed_business_employee_expenses" + assert output in puf_support_module.PUF_TAX_DETAIL_DEFAULT_PERSON_OUTPUTS + assert output in puf_support_module._PUF_TAX_DETAIL_NONNEGATIVE_OUTPUTS + assert output not in puf_support_module._PUF_TAX_DETAIL_SPARSE_PERSON_OUTPUTS + + +def test_signal_gate_accepts_reference_like_nondefault_values() -> None: + values = np.zeros(100) + values[:30] = np.arange(1, 31, dtype=np.float64) * 100.0 + frame = _PersonFrame( + pd.DataFrame({"unreimbursed_business_employee_expenses": values}) + ) + + result = us_misc_itemized_signal_gate(frame) # type: ignore[arg-type] + + assert result.passed, result.failures + assert result.details["positive_share"] == pytest.approx(0.30) + + +@pytest.mark.parametrize( + "person", + [ + pd.DataFrame({"other": [0.0, 1.0]}), + pd.DataFrame({"unreimbursed_business_employee_expenses": [0.0, 0.0]}), + pd.DataFrame({"unreimbursed_business_employee_expenses": [0.0, -1.0]}), + pd.DataFrame({"unreimbursed_business_employee_expenses": [0.0, np.nan]}), + ], +) +def test_signal_gate_rejects_missing_default_or_invalid_surface( + person: pd.DataFrame, +) -> None: + result = us_misc_itemized_signal_gate( # type: ignore[arg-type] + _PersonFrame(person) + ) + + assert not result.passed + + +@requires_us +def test_policyengine_us_contract_is_a_person_year_input_leaf() -> None: + from policyengine_us import CountryTaxBenefitSystem + + variable = CountryTaxBenefitSystem().variables[ + "unreimbursed_business_employee_expenses" + ] + + assert variable.is_input_variable() + assert variable.entity.key == "person" + assert str(variable.definition_period).lower() == "year" + assert variable.default_value == 0 diff --git a/packages/populace-build/tests/test_us_org_wages.py b/packages/populace-build/tests/test_us_org_wages.py new file mode 100644 index 00000000..154c8ec5 --- /dev/null +++ b/packages/populace-build/tests/test_us_org_wages.py @@ -0,0 +1,351 @@ +"""Exact CPS ORG, occupation, and FLSA overtime stage contracts.""" + +from __future__ import annotations + +import gzip +import hashlib + +import numpy as np +import pandas as pd +import pytest + +from populace.build.source_manifest import SourceStageSpec +from populace.build.us_runtime import ( + ORG_2024_DONOR_CONTENT_SHA256, + ORG_PREDICTORS, + US_ORG_WAGES_OUTPUT_COLUMNS, + derive_flsa_overtime_premium, + derive_us_org_occupation_inputs, + fetch_org_2024_donor, + load_org_2024_donor, + us_org_wages_signal_gate, + us_org_wages_stage_spec, +) +from populace.build.us_runtime import org_wages as module +from populace.frame import US_SCHEMA, Frame, WeightKind, Weights + + +def _person(n: int = 1_000) -> pd.DataFrame: + household = np.arange(1, n + 1, dtype=np.int64) + occupation = np.full(n, 20, dtype=np.int16) + occupation[:180] = 0 + occupation[180:420] = 53 + occupation[420:423] = 52 + occupation[423:443] = 8 + occupation[443:447] = 41 + occupation[447:737] = 1 + employment = np.zeros(n) + employment[400:] = 52_000.0 + return pd.DataFrame( + { + "person_id": np.arange(1, n + 1, dtype=np.int64), + "person_household_id": household, + "person_tax_unit_id": household + 10_000, + "person_spm_unit_id": household + 20_000, + "person_family_id": household + 30_000, + "person_marital_unit_id": household + 40_000, + "age": np.resize(np.arange(18, 78), n), + "is_female": np.arange(n) % 2 == 0, + "PRDTRACE": np.resize(np.asarray([1, 1, 2, 4]), n), + "PRDTHSP": np.arange(n) % 7 == 0, + "POCCU2": occupation, + "employment_income_before_lsr": employment, + "self_employment_income_before_lsr": 0.0, + "weekly_hours_worked_before_lsr": np.where(employment > 0, 40.0, 0.0), + "hours_worked_last_week": np.where( + np.arange(n) >= 950, 50.0, np.where(employment > 0, 40.0, 0.0) + ), + "weeks_worked": np.where(employment > 0, 52.0, 0.0), + } + ) + + +def _frame(person: pd.DataFrame) -> Frame: + household_ids = person["person_household_id"].to_numpy() + n = len(person) + return Frame( + { + "person": person, + "household": pd.DataFrame( + {"household_id": household_ids, "state_fips": np.resize([6, 36], n)} + ), + "tax_unit": pd.DataFrame( + {"tax_unit_id": person["person_tax_unit_id"].to_numpy()} + ), + "spm_unit": pd.DataFrame( + {"spm_unit_id": person["person_spm_unit_id"].to_numpy()} + ), + "family": pd.DataFrame( + {"family_id": person["person_family_id"].to_numpy()} + ), + "marital_unit": pd.DataFrame( + {"marital_unit_id": person["person_marital_unit_id"].to_numpy()} + ), + }, + US_SCHEMA, + {"household": Weights(np.ones(n), WeightKind.DESIGN)}, + ) + + +def _donor(n: int = 200) -> pd.DataFrame: + rng = np.random.default_rng(2) + return pd.DataFrame( + { + "employment_income": rng.uniform(10_000, 120_000, n), + "weekly_hours_worked": rng.uniform(20, 60, n), + "age": rng.integers(18, 80, n), + "is_female": rng.integers(0, 2, n), + "is_hispanic": rng.integers(0, 2, n), + "race_wbho": rng.integers(1, 5, n), + "state_fips": np.resize([6, 36], n), + "hourly_wage": rng.uniform(8, 80, n), + "is_paid_hourly": rng.integers(0, 2, n), + "sample_weight": rng.uniform(1, 5, n), + } + ) + + +def test_stage_spec_declares_complete_family_and_exact_predictors() -> None: + spec = us_org_wages_stage_spec() + assert isinstance(spec, SourceStageSpec) + assert spec.stage == "org_wages" + assert set(US_ORG_WAGES_OUTPUT_COLUMNS) <= set(spec.outputs) + assert ORG_PREDICTORS == ( + "employment_income", + "weekly_hours_worked", + "age", + "is_female", + "is_hispanic", + "race_wbho", + "state_fips", + ) + assert len(ORG_2024_DONOR_CONTENT_SHA256) == 64 + + +def test_donor_loader_verifies_canonical_uncompressed_sha(tmp_path) -> None: + donor = _donor(20) + content = donor.to_csv(index=False).encode() + path = tmp_path / "census_cps_org_2024_wages.csv.gz" + path.write_bytes(gzip.compress(content, mtime=0)) + digest = hashlib.sha256(content).hexdigest() + + loaded = load_org_2024_donor(path, expected_content_sha256=digest) + assert len(loaded) == 20 + with pytest.raises(ValueError, match="sha-256 verification"): + load_org_2024_donor(path, expected_content_sha256="0" * 64) + + +def test_fetch_reuses_only_a_canonical_hash_matching_cache( + tmp_path, monkeypatch +) -> None: + content = _donor(20).to_csv(index=False).encode() + path = tmp_path / "census_cps_org_2024_wages.csv.gz" + path.write_bytes(gzip.compress(content, mtime=0)) + digest = hashlib.sha256(content).hexdigest() + monkeypatch.setattr( + module, + "_load_month_from_network", + lambda month: pytest.fail("valid cache must not fetch a monthly file"), + ) + + assert fetch_org_2024_donor(tmp_path, expected_content_sha256=digest) == path + + +def test_month_transform_matches_retired_raw_cps_fields() -> None: + raw = pd.DataFrame( + { + "HRMIS": [4, 8, 4, 3], + "gestfips": [6, 36, 12, 6], + "prtage": [30, 45, 28, 32], + "pesex": [1, 2, 2, 1], + "ptdtrace": [1, 2, 3, 1], + "pehspnon": [2, 2, 1, 2], + "pworwgt": [100.0, 200.0, 150.0, 0.0], + "pternwa": [100000.0, 80000.0, 120000.0, 90000.0], + "pternhly": [2500.0, -1.0, 3000.0, 2000.0], + "peernhry": [1, 2, 1, 1], + "pehruslt": [40.0, 40.0, 50.0, 40.0], + "prerelg": [1, 1, 1, 1], + "pemlr": [1, 1, 2, 1], + "peio1cow": [1, 4, 2, 1], + } + ) + + transformed = module._transform_month(raw) + + assert len(transformed) == 3 + assert transformed["hourly_wage"].tolist() == [25.0, 20.0, 30.0] + assert transformed["is_paid_hourly"].tolist() == [1.0, 0.0, 1.0] + assert transformed["employment_income"].tolist() == [52_000.0, 41_600.0, 62_400.0] + assert transformed.dtypes.eq(np.dtype("float32")).all() + + +def test_occupation_carries_match_retired_poccu2_codes() -> None: + person = pd.DataFrame( + { + "PRDTRACE": [1, 2, 4, 1, 1, 1, 1], + "PRDTHSP": [0, 1, 0, 0, 0, 0, 0], + "POCCU2": [53, 52, 8, 41, 1, 50, 20], + } + ) + result = derive_us_org_occupation_inputs(person) + + assert result["cps_race"].tolist() == [1, 2, 4, 1, 1, 1, 1] + assert result["is_hispanic"].tolist() == [ + False, + True, + False, + False, + False, + False, + False, + ] + assert result["has_never_worked"].tolist() == [ + True, + False, + False, + False, + False, + False, + False, + ] + assert result["is_military"].tolist() == [ + False, + True, + False, + False, + False, + False, + False, + ] + assert result["is_computer_scientist"].tolist()[2] + assert result["is_farmer_fisher"].tolist()[3] + assert result["is_executive_administrative_professional"].tolist() == [ + False, + False, + False, + False, + True, + True, + False, + ] + + +def test_flsa_proxy_uses_annual_wage_share_and_exemption_screen() -> None: + policy = (107_432.0, 35_568.0, 57_470.4, 40.0, 1.5) + premium = derive_flsa_overtime_premium( + time_period=2024, + employment_income=np.asarray([57_200, 60_000, 60_000, 100_000, 50_000]), + hours_worked_last_week=np.asarray([50, 50, 50, 50, 50]), + weeks_worked=np.asarray([52, 52, 52, 52, 52]), + is_paid_hourly=np.asarray([True, False, False, False, True]), + has_never_worked=np.asarray([False, False, False, False, True]), + is_military=np.zeros(5, dtype=bool), + is_executive_administrative_professional=np.asarray( + [False, True, False, False, False] + ), + is_farmer_fisher=np.zeros(5, dtype=bool), + is_computer_scientist=np.asarray([False, False, True, False, False]), + policy=policy, + ) + np.testing.assert_allclose(premium, [5_200, 0, 0, 100_000 / 11, 0], rtol=1e-6) + + +def test_imputation_zeroes_inactive_and_union_is_deterministic(monkeypatch) -> None: + class Fitted: + def predict(self, features): + return pd.DataFrame( + { + "hourly_wage": np.full(len(features), 25.0), + "is_paid_hourly": np.resize([0.9, 0.1], len(features)), + }, + index=features.index, + ) + + class FakeQRF: + def __init__(self, **kwargs): + pass + + def fit(self, *args, **kwargs): + return Fitted() + + monkeypatch.setattr("populace.fit.QRF", FakeQRF) + frame = _frame(_person(100)) + first_carried, first = module.impute_us_org_wages(frame, _donor(), seed=7) + second_carried, second = module.impute_us_org_wages(frame, _donor(), seed=99) + + inactive = frame.table("person")["employment_income_before_lsr"].to_numpy() <= 0 + assert (first.loc[inactive, "hourly_wage"] == 0).all() + assert not first.loc[inactive, "is_paid_hourly"].any() + assert not first.loc[inactive, "is_union_member_or_covered"].any() + np.testing.assert_array_equal( + first["is_union_member_or_covered"], second["is_union_member_or_covered"] + ) + pd.testing.assert_frame_equal(first_carried, second_carried) + + +def test_union_assignment_matches_unweighted_state_quotas() -> None: + n = 201 + features = pd.DataFrame( + { + "employment_income": np.r_[np.full(200, 50_000.0), 0.0], + "weekly_hours_worked": np.full(n, 40.0), + "age": np.resize(np.arange(20, 70), n), + "is_female": np.arange(n) % 2, + "is_hispanic": np.arange(n) % 7 == 0, + "race_wbho": np.resize([1, 2, 3, 4], n), + "state_fips": np.r_[np.full(100, 6), np.full(100, 37), 6], + } + ) + + assigned = module._assign_union(features) + + assert assigned[:100].sum() == 16 # California 16.3%, rounded. + assert assigned[100:200].sum() == 3 # North Carolina 3.1%, rounded. + assert not assigned[-1] # inactive rows never receive union coverage. + + +def test_live_2024_flsa_policy_values_when_us_extra_is_installed() -> None: + pytest.importorskip("policyengine_us") + assert module._flsa_policy(2024) == pytest.approx( + (107_432.0, 35_568.0, 57_470.4, 40.0, 1.5) + ) + + +def _plausible_surface() -> Frame: + person = _person() + person["cps_race"] = person["PRDTRACE"] + person["is_hispanic"] = person["PRDTHSP"].ne(0) + carried = derive_us_org_occupation_inputs(person) + for column in carried: + person[column] = carried[column] + person["hourly_wage"] = 0.0 + person.loc[400:949, "hourly_wage"] = 25.0 + person["is_paid_hourly"] = False + person.loc[400:649, "is_paid_hourly"] = True + person["is_union_member_or_covered"] = False + person.loc[700:759, "is_union_member_or_covered"] = True + person["fsla_overtime_premium"] = 0.0 + # These 50 are non-exempt code-20 workers with 50 hours and positive wages. + person.loc[950:999, "fsla_overtime_premium"] = 52_000 / 11 + return _frame(person) + + +def test_signal_gate_passes_plausible_coherent_surface() -> None: + gate = us_org_wages_signal_gate(_plausible_surface()) + assert gate.passed, gate.failures + + +def test_signal_gate_rejects_zero_and_structurally_impossible_premium() -> None: + frame = _plausible_surface() + frame.table("person")["fsla_overtime_premium"] = 0.0 + gate = us_org_wages_signal_gate(frame) + assert not gate.passed + assert any("fsla_overtime_premium" in failure for failure in gate.failures) + + impossible = _plausible_surface() + impossible.table("person").loc[500, "fsla_overtime_premium"] = 1_000.0 + impossible.table("person").loc[500, "hours_worked_last_week"] = 40.0 + gate = us_org_wages_signal_gate(impossible) + assert not gate.passed + assert any("positive_without_overtime" in failure for failure in gate.failures) diff --git a/packages/populace-build/tests/test_us_other_health_insurance.py b/packages/populace-build/tests/test_us_other_health_insurance.py new file mode 100644 index 00000000..91838bb4 --- /dev/null +++ b/packages/populace-build/tests/test_us_other_health_insurance.py @@ -0,0 +1,550 @@ +"""ASEC other-health-insurance restoration and ESI source exclusion.""" + +from __future__ import annotations + +import importlib.util +import json +from importlib.metadata import version +from importlib.resources import files +from pathlib import Path + +import numpy as np +import pandas as pd +import pytest + +import populace.build.us_runtime.other_health_insurance as module +from populace.build.source_runtime import SourceRuntimeError +from populace.build.us_runtime import US_SOURCE_MANIFEST +from populace.build.us_runtime.other_health_insurance import ( + OTHER_HEALTH_INSURANCE_ARCHIVED_DERIVATION_URL, + OTHER_HEALTH_INSURANCE_ARCHIVED_PUF_IMPUTATION_URL, + OTHER_HEALTH_INSURANCE_ARCHIVED_PUF_OUTPUTS_URL, + OTHER_HEALTH_INSURANCE_ARCHIVED_PUF_PREDICTORS_URL, + OTHER_HEALTH_INSURANCE_ARCHIVED_PUF_SPLICE_URL, + US_OTHER_HEALTH_INSURANCE_NONCONSTANT_PERSON_COLUMNS, + US_OTHER_HEALTH_INSURANCE_OUTPUT_COLUMNS, + derive_us_other_health_insurance_from_asec, + derive_us_other_health_insurance_from_manifest, + impute_us_other_health_insurance_to_puf_support_from_manifest, + us_other_health_insurance_signal_gate, + us_other_health_insurance_stage_spec, + with_us_other_health_insurance_inputs, +) +from populace.build.us_runtime.puf_support import clone_us_frame_for_puf_support +from populace.build.us_runtime.source_runtime import us_source_operation_handlers +from populace.frame import US_SCHEMA, Frame, WeightKind, Weights + +policyengine_us_installed = importlib.util.find_spec("policyengine_us") is not None +requires_us = pytest.mark.skipif( + not policyengine_us_installed, + reason="requires the policyengine-us [us] extra (build environment)", +) + +_REPORTED, _OTHER = US_OTHER_HEALTH_INSURANCE_OUTPUT_COLUMNS +ROOT = Path(__file__).resolve().parents[3] +_PREDICTORS = ( + "age", + "is_male", + "has_esi", + "tax_unit_is_joint", + "tax_unit_count_dependents", + "employment_income", + "self_employment_income", + "social_security", +) +_REPORTED_VALUES = np.asarray( + [0.0, 1_000.0, 0.0, 2_000.0, 0.0, 3_000.0, 0.0, 4_000.0, 0.0, 0.0] +) + + +def _frame() -> Frame: + count = len(_REPORTED_VALUES) + person = pd.DataFrame( + { + "person_id": np.arange(1, count + 1, dtype="int64"), + "person_household_id": np.arange(101, 101 + count, dtype="int64"), + "person_tax_unit_id": np.arange(201, 201 + count, dtype="int64"), + "person_spm_unit_id": np.arange(301, 301 + count, dtype="int64"), + "person_family_id": np.arange(401, 401 + count, dtype="int64"), + "person_marital_unit_id": np.arange(501, 501 + count, dtype="int64"), + _REPORTED: _REPORTED_VALUES, + "age": np.arange(30, 30 + count), + "is_female": np.tile([True, False], count // 2), + "has_esi": np.tile([False, True], count // 2), + "tax_unit_role_input": ["HEAD"] * count, + "employment_income_before_lsr": np.arange(count) * 5_000.0, + "self_employment_income_before_lsr": np.zeros(count), + "social_security_retirement": np.zeros(count), + "social_security_disability": np.zeros(count), + "social_security_dependents": np.zeros(count), + "social_security_survivors": np.zeros(count), + } + ) + tables = { + "person": person, + "household": pd.DataFrame( + { + "household_id": person["person_household_id"], + "state_fips": np.full(count, 6), + } + ), + "tax_unit": pd.DataFrame( + { + "tax_unit_id": person["person_tax_unit_id"], + "filing_status_input": ["SINGLE"] * count, + } + ), + "spm_unit": pd.DataFrame({"spm_unit_id": person["person_spm_unit_id"]}), + "family": pd.DataFrame({"family_id": person["person_family_id"]}), + "marital_unit": pd.DataFrame( + {"marital_unit_id": person["person_marital_unit_id"]} + ), + } + return Frame( + tables, + US_SCHEMA, + { + "household": Weights( + np.ones(count, dtype=np.float64), + WeightKind.DESIGN, + ) + }, + ) + + +class _ZeroPremiumEngine: + def materialize( + self, + frame: Frame, + variables: list[str], + *, + period: int, + ) -> dict[str, np.ndarray]: + assert period == 2024 + assert variables == list( + module.US_OTHER_HEALTH_INSURANCE_MODELED_PREMIUM_VARIABLES + ) + return { + variable: np.zeros(frame.n("tax_unit"), dtype=np.float64) + for variable in variables + } + + +def test_archive_urls_pin_exact_derivation_outputs_predictors_and_splice() -> None: + commit = "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe" + urls = ( + OTHER_HEALTH_INSURANCE_ARCHIVED_DERIVATION_URL, + OTHER_HEALTH_INSURANCE_ARCHIVED_PUF_OUTPUTS_URL, + OTHER_HEALTH_INSURANCE_ARCHIVED_PUF_PREDICTORS_URL, + OTHER_HEALTH_INSURANCE_ARCHIVED_PUF_IMPUTATION_URL, + OTHER_HEALTH_INSURANCE_ARCHIVED_PUF_SPLICE_URL, + ) + assert all(commit in url for url in urls) + assert OTHER_HEALTH_INSURANCE_ARCHIVED_DERIVATION_URL.endswith( + "datasets/cps/cps.py#L828-L944" + ) + assert OTHER_HEALTH_INSURANCE_ARCHIVED_PUF_OUTPUTS_URL.endswith( + "datasets/cps/extended_cps.py#L135-L194" + ) + assert OTHER_HEALTH_INSURANCE_ARCHIVED_PUF_PREDICTORS_URL.endswith( + "datasets/cps/extended_cps.py#L234-L248" + ) + assert OTHER_HEALTH_INSURANCE_ARCHIVED_PUF_IMPUTATION_URL.endswith( + "datasets/cps/extended_cps.py#L639-L745" + ) + assert OTHER_HEALTH_INSURANCE_ARCHIVED_PUF_SPLICE_URL.endswith( + "datasets/cps/extended_cps.py#L1014-L1076" + ) + + +def test_stage_manifest_pins_residual_and_joint_puf_qrf() -> None: + spec = us_other_health_insurance_stage_spec() + + assert spec.stage == "other_health_insurance_premiums" + assert spec.survey == "Census CPS ASEC" + assert spec.source == "https://www.census.gov/programs-surveys/cps.html" + assert spec.grain == "person" + assert spec.outputs == US_OTHER_HEALTH_INSURANCE_OUTPUT_COLUMNS + assert spec.nonnegative_outputs == spec.outputs + assert US_OTHER_HEALTH_INSURANCE_NONCONSTANT_PERSON_COLUMNS == (_OTHER,) + assert [operation.kind for operation in spec.operations] == [ + "read_table", + "derive_other_health_insurance_premiums", + "impute_other_health_insurance_premiums_to_puf_support", + ] + assert spec.operations[0].parameters == { + "table": "person", + "weight": "person_weight", + } + assert spec.operations[1].parameters == { + "reported_source": _REPORTED, + "chip_premium_source": "chip_premium", + "marketplace_net_premium_source": "marketplace_net_premium", + "medicaid_premium_source": "medicaid_premium", + "output": _OTHER, + } + assert spec.operations[2].parameters == { + "predictors": list(_PREDICTORS), + "max_train_samples": 5_000, + "n_estimators": 100, + "seed_from_build_config": True, + "weight": "person_weight", + } + assert "PUF-only residual-order exceedances remain diagnostics" in spec.notes + + handlers = us_source_operation_handlers() + assert ( + handlers["derive_other_health_insurance_premiums"] + is derive_us_other_health_insurance_from_manifest + ) + assert ( + handlers["impute_other_health_insurance_premiums_to_puf_support"] + is impute_us_other_health_insurance_to_puf_support_from_manifest + ) + + +def test_asec_residual_is_exact_nonnegative_and_does_not_mutate_source() -> None: + source = pd.DataFrame( + { + _REPORTED: [0.0, 100.0, 500.0, 1_000.0], + "chip_premium": [0.0, 80.0, 100.0, 100.0], + "marketplace_net_premium": [0.0, 30.0, 200.0, 200.0], + "medicaid_premium": [0.0, 10.0, 300.0, 300.0], + } + ) + original = source.copy(deep=True) + + result = derive_us_other_health_insurance_from_asec(source) + + assert result[_OTHER].tolist() == [0.0, 0.0, 0.0, 400.0] + pd.testing.assert_frame_equal(source, original) + + +@pytest.mark.parametrize( + ("column", "bad_value", "message"), + [ + (_REPORTED, np.nan, "nonnumeric or nonfinite"), + ("chip_premium", np.inf, "nonnumeric or nonfinite"), + ("marketplace_net_premium", -1.0, "negative"), + ("medicaid_premium", -1.0, "negative"), + ], +) +def test_asec_residual_fails_closed_on_invalid_sources( + column: str, + bad_value: float, + message: str, +) -> None: + source = pd.DataFrame( + { + _REPORTED: [1_000.0], + "chip_premium": [0.0], + "marketplace_net_premium": [0.0], + "medicaid_premium": [0.0], + } + ) + source.loc[0, column] = bad_value + + with pytest.raises(SourceRuntimeError, match=message): + derive_us_other_health_insurance_from_asec(source) + + +def test_with_inputs_preserves_measured_asec_and_carries_signal( + monkeypatch: pytest.MonkeyPatch, +) -> None: + monkeypatch.setattr(module, "PolicyEngineUSEngine", _ZeroPremiumEngine) + frame = _frame() + original = frame.table("person").copy(deep=True) + + result = with_us_other_health_insurance_inputs(frame, seed=0, time_period=2024) + + pd.testing.assert_frame_equal(frame.table("person"), original) + np.testing.assert_allclose(result.table("person")[_REPORTED], _REPORTED_VALUES) + np.testing.assert_allclose(result.table("person")[_OTHER], _REPORTED_VALUES) + gate = us_other_health_insurance_signal_gate(result) + assert gate.passed, gate.failures + + +def test_tax_unit_premiums_are_allocated_wholly_to_first_person() -> None: + frame = _frame() + person = frame.table("person") + person.loc[1, "person_tax_unit_id"] = person.loc[0, "person_tax_unit_id"] + frame.table("tax_unit").drop(index=[1], inplace=True) + + allocated = module._tax_unit_values_on_first_person( + frame, + np.arange(1.0, frame.n("tax_unit") + 1.0) * 100.0, + variable="chip_premium", + ) + + assert allocated[0] == 100.0 + assert allocated[1] == 0.0 + assert np.count_nonzero(allocated) == frame.n("tax_unit") + + +def test_modeled_premiums_batch_by_household_and_realign_tax_units( + monkeypatch: pytest.MonkeyPatch, +) -> None: + frame = _frame() + person_tax_unit_ids = np.asarray( + [210, 201, 209, 202, 208, 203, 207, 204, 206, 205], + dtype=np.int64, + ) + frame.table("person")["person_tax_unit_id"] = person_tax_unit_ids + calls: list[np.ndarray] = [] + + class IdPremiumEngine: + def materialize( + self, + batch: Frame, + variables: list[str], + *, + period: int, + ) -> dict[str, np.ndarray]: + assert period == 2024 + assert variables == list( + module.US_OTHER_HEALTH_INSURANCE_MODELED_PREMIUM_VARIABLES + ) + tax_unit_ids = batch.table("tax_unit")["tax_unit_id"].to_numpy() + calls.append(tax_unit_ids.copy()) + factor = tax_unit_ids.astype(np.float64) - 200.0 + return { + "chip_premium": factor, + "marketplace_net_premium": factor * 2.0, + "medicaid_premium": factor * 3.0, + } + + monkeypatch.setattr(module, "PolicyEngineUSEngine", IdPremiumEngine) + + result = with_us_other_health_insurance_inputs( + frame, + seed=0, + time_period=2024, + maximum_microsim_batch_size=3, + ) + + assert [len(call) for call in calls] == [3, 3, 3, 1] + assert set(np.concatenate(calls)) == set(range(201, 211)) + expected_modeled = (person_tax_unit_ids - 200.0) * 6.0 + np.testing.assert_allclose( + result.table("person")[_OTHER], + np.maximum(_REPORTED_VALUES - expected_modeled, 0.0), + ) + gate = us_other_health_insurance_signal_gate(result) + assert gate.passed, gate.failures + + +def test_joint_qrf_replaces_only_puf_and_does_not_invent_clipping( + monkeypatch: pytest.MonkeyPatch, +) -> None: + monkeypatch.setattr(module, "PolicyEngineUSEngine", _ZeroPremiumEngine) + expanded = clone_us_frame_for_puf_support(_frame()) + original = expanded.table("person").copy(deep=True) + calls: dict[str, object] = {} + + class FakeFitted: + def predict(self, test: pd.DataFrame) -> pd.DataFrame: + calls["test"] = test.copy() + reported = _REPORTED_VALUES.copy() + other = reported.copy() + other[1] = 1_500.0 + return pd.DataFrame( + {_REPORTED: reported, _OTHER: other}, + index=test.index, + ) + + class FakeQRF: + def __init__(self, **kwargs: object) -> None: + calls["init"] = kwargs + + def fit( + self, + training: pd.DataFrame, + predictors: list[str], + targets: list[str], + *, + weights: np.ndarray, + ) -> FakeFitted: + calls["training"] = training.copy() + calls["predictors"] = predictors + calls["targets"] = targets + calls["weights"] = weights.copy() + return FakeFitted() + + monkeypatch.setattr(module, "QRF", FakeQRF) + + result = with_us_other_health_insurance_inputs( + expanded, + seed=17, + time_period=2024, + ) + + assert calls["init"] == {"n_estimators": 100, "seed": 17} + assert calls["predictors"] == list(_PREDICTORS) + assert calls["targets"] == list(US_OTHER_HEALTH_INSURANCE_OUTPUT_COLUMNS) + training = calls["training"] + assert isinstance(training, pd.DataFrame) + assert list(training.columns) == [ + *_PREDICTORS, + *US_OTHER_HEALTH_INSURANCE_OUTPUT_COLUMNS, + ] + asec_mask = original["person_support_channel"] == "asec" + np.testing.assert_allclose( + calls["weights"], + expanded.resolve_weights("person").values[asec_mask], + ) + + person = result.table("person") + puf_mask = person["person_support_channel"] == "puf_tax_detail" + np.testing.assert_allclose(person.loc[~puf_mask, _OTHER], _REPORTED_VALUES) + assert person.loc[puf_mask, _OTHER].iloc[1] == 1_500.0 + assert person.loc[puf_mask, _REPORTED].iloc[1] == 1_000.0 + gate = us_other_health_insurance_signal_gate(result) + assert gate.passed, gate.failures + assert gate.details["other_exceeds_reported_by_channel"] == { + "asec": 0, + "puf_tax_detail": 1, + } + pd.testing.assert_frame_equal(expanded.table("person"), original) + + +def test_signal_gate_rejects_missing_default_and_measured_identity_violation( + monkeypatch: pytest.MonkeyPatch, +) -> None: + monkeypatch.setattr(module, "PolicyEngineUSEngine", _ZeroPremiumEngine) + valid = with_us_other_health_insurance_inputs(_frame(), seed=0, time_period=2024) + assert us_other_health_insurance_signal_gate(valid).passed + + missing = with_us_other_health_insurance_inputs(_frame(), seed=0, time_period=2024) + missing.table("person").drop(columns=[_OTHER], inplace=True) + assert not us_other_health_insurance_signal_gate(missing).passed + + default = with_us_other_health_insurance_inputs(_frame(), seed=0, time_period=2024) + default.table("person")[_OTHER] = 0.0 + assert not us_other_health_insurance_signal_gate(default).passed + + invalid = with_us_other_health_insurance_inputs(_frame(), seed=0, time_period=2024) + invalid.table("person").loc[0, _OTHER] = 1.0 + gate = us_other_health_insurance_signal_gate(invalid) + assert not gate.passed + assert any("measured ASEC" in failure for failure in gate.failures) + + +def test_esi_exclusion_pins_complete_hermetic_source_unavailability_evidence() -> None: + payload = json.loads( + files("populace.build.us").joinpath("ecps_parity_known_gaps.json").read_text() + ) + entry = payload["known_gaps"]["employer_sponsored_insurance_premiums"] + evidence = entry["evidence"] + + assert entry["reason"].startswith("SOURCE UNAVAILABILITY WITH EVIDENCE:") + assert evidence["classification"] == "source_unavailability" + assert evidence["retired_derivation"] == { + "repository_owner": "PolicyEngine", + "repository_name_parts": ["policyengine-", "us-data"], + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": ["policyengine_", "us_data", "datasets", "cps", "cps.py"], + "lines": "197-271,1575-1581", + } + assert evidence["optional_source_fields"]["lines"] == "13-55" + assert evidence["required_columns"] == [ + "NOW_OWNGRP", + "NOW_HIPAID", + "NOW_GRPFTYP", + "PHIP_VAL", + ] + evidence_hashes = { + item["filename"]: item["sha256"] for item in evidence["hermetic_inputs"] + } + build_summary = json.loads( + (ROOT / "experiments/build_j_recert/base_j.summary.json").read_text() + ) + recorded_hashes = { + Path(item["path"]).name: item["sha256"] + for item in build_summary["base_source"]["sources"] + } + assert evidence_hashes == recorded_hashes + assert set(evidence_hashes) == { + "census_cps_2022.h5", + "census_cps_2023.h5", + "census_cps_2024.h5", + } + for item in evidence["hermetic_inputs"]: + assert item["present_columns"] == ["PHIP_VAL"] + assert item["missing_columns"] == [ + "NOW_OWNGRP", + "NOW_HIPAID", + "NOW_GRPFTYP", + ] + assert "buildj_base.sh lines 65-69" in evidence["hermetic_build_contract"] + assert "base_j.summary.json lines 55-75" in evidence["hermetic_build_contract"] + build_script = (ROOT / "experiments/build_j_recert/buildj_base.sh").read_text() + for year in (2022, 2023, 2024): + assert f'--asec-h5 {year}="$USD/census_cps_{year}.h5"' in build_script + + esi_stage = US_SOURCE_MANIFEST.stage_map()["meps_esi_premiums"] + assignment = next( + operation + for operation in esi_stage.operations + if operation.kind == "assign_by_plan_type" + ) + assert assignment.parameters["inputs"] == [ + "NOW_OWNGRP", + "NOW_HIPAID", + "NOW_GRPFTYP", + "PHIP_VAL", + ] + assert assignment.parameters["source_unavailable_when_missing"] == [ + "NOW_OWNGRP", + "NOW_HIPAID", + "NOW_GRPFTYP", + ] + forbidden_proxies = {"has_esi", "is_esi_dependent", "tax_unit_size", "state_fips"} + assert forbidden_proxies.isdisjoint(assignment.parameters["inputs"]) + + +@requires_us +def test_policyengine_us_graph_uses_other_health_insurance_premiums() -> None: + from policyengine_us import CountryTaxBenefitSystem, Simulation + + assert version("policyengine-us") == "1.764.6" + variable = CountryTaxBenefitSystem().variables[_OTHER] + assert variable.is_input_variable() + assert variable.entity.key == "person" + assert str(variable.definition_period).lower() == "year" + assert variable.default_value == 0 + + def situation(premiums: float) -> dict[str, object]: + return { + "people": { + "adult": { + "age": {"2024": 35}, + _OTHER: {"2024": premiums}, + } + }, + "tax_units": { + "tax_unit": { + "members": ["adult"], + "filing_status": {"2024": "SINGLE"}, + } + }, + "spm_units": {"spm_unit": {"members": ["adult"]}}, + "households": { + "household": { + "members": ["adult"], + "state_code": {"2024": "CA"}, + } + }, + } + + baseline = Simulation(situation=situation(0.0)) + active = Simulation(situation=situation(1_200.0)) + expected_deltas = { + "spm_unit_health_insurance_premiums": 1_200.0, + "spm_unit_non_premium_medical_out_of_pocket_expenses": 0.0, + "spm_unit_medical_out_of_pocket_expenses": 1_200.0, + "spm_unit_spm_expenses": 1_200.0, + "spm_unit_net_income": -1_200.0, + } + for name, expected in expected_deltas.items(): + delta = active.calculate(name, 2024)[0] - baseline.calculate(name, 2024)[0] + assert delta == pytest.approx(expected) diff --git a/packages/populace-build/tests/test_us_parity_reference.py b/packages/populace-build/tests/test_us_parity_reference.py index 5bc72598..6529bd12 100644 --- a/packages/populace-build/tests/test_us_parity_reference.py +++ b/packages/populace-build/tests/test_us_parity_reference.py @@ -18,11 +18,19 @@ ) from populace.build.us_runtime.hours_worked import US_HOURS_WORKED_OUTPUT_COLUMNS from populace.build.us_runtime.immigration import US_IMMIGRATION_OUTPUT_COLUMNS +from populace.build.us_runtime.org_wages import US_ORG_WAGES_OUTPUT_COLUMNS from populace.build.us_runtime.parity_reference import ( + ECPS_PARITY_KNOWN_GAPS_RESOURCE, ECPS_PARITY_REFERENCE_RESOURCE, load_ecps_parity_known_gaps, load_ecps_parity_reference, ) +from populace.build.us_runtime.scf_auto_loans import US_SCF_AUTO_LOAN_OUTPUT_COLUMNS +from populace.build.us_runtime.sipp_head_start import ( + US_SIPP_HEAD_START_OUTPUT_COLUMNS, +) +from populace.build.us_runtime.sipp_tips import US_SIPP_TIPS_OUTPUT_COLUMNS +from populace.build.us_runtime.sipp_vehicles import US_SIPP_VEHICLE_OUTPUT_COLUMNS from populace.build.us_runtime.snap_take_up import US_SNAP_TAKE_UP_OUTPUT_COLUMN from populace.build.us_runtime.take_up_contract import load_take_up_contract @@ -158,6 +166,11 @@ def test_register_exempts_no_release_produced_layer(self) -> None: set(US_IMMIGRATION_OUTPUT_COLUMNS) | set(US_HOURS_WORKED_OUTPUT_COLUMNS) | set(US_ELIGIBILITY_INPUTS_OUTPUT_COLUMNS) + | set(US_SIPP_TIPS_OUTPUT_COLUMNS) + | set(US_SIPP_HEAD_START_OUTPUT_COLUMNS) + | set(US_SIPP_VEHICLE_OUTPUT_COLUMNS) + | set(US_ORG_WAGES_OUTPUT_COLUMNS) + | set(US_SCF_AUTO_LOAN_OUTPUT_COLUMNS) | {US_SNAP_TAKE_UP_OUTPUT_COLUMN} | { program.variable @@ -172,6 +185,73 @@ def test_register_exempts_no_release_produced_layer(self) -> None: ) assert stale == [], f"register exempts release-produced layers: {stale}" + def test_legacy_auto_loan_gaps_are_retired(self) -> None: + gaps = {gap.variable for gap in load_ecps_parity_known_gaps()} + assert "auto_loan_balance" not in gaps + assert "auto_loan_interest" not in gaps + + def test_weeks_unemployed_gap_is_retired_with_measured_asec_stage(self) -> None: + gaps = {gap.variable for gap in load_ecps_parity_known_gaps()} + + assert "weeks_unemployed" not in gaps + + def test_ssi_take_up_gap_is_retired_with_count_calibrated_stage(self) -> None: + gaps = {gap.variable for gap in load_ecps_parity_known_gaps()} + program = load_take_up_contract().program_map()["takes_up_ssi_if_eligible"] + + assert program.populace_treatment == "count_calibrated" + assert "takes_up_ssi_if_eligible" not in gaps + + def test_head_start_gap_is_retired_but_early_head_start_is_evidenced(self) -> None: + gaps = {gap.variable: gap for gap in load_ecps_parity_known_gaps()} + programs = load_take_up_contract().program_map() + + standard = "takes_up_head_start_if_eligible" + early = "takes_up_early_head_start_if_eligible" + assert programs[standard].populace_treatment == "out_of_scope" + assert standard not in gaps + + assert early in gaps + assert gaps[early].reason.startswith("SOURCE UNAVAILABILITY WITH EVIDENCE:") + assert "datasets/cps/cps.py lines 582-583 and 640-648" in gaps[early].reason + assert "parameters/take_up/early_head_start.yaml lines 1-9" in ( + gaps[early].reason + ) + + raw = json.loads(_resource(ECPS_PARITY_KNOWN_GAPS_RESOURCE).read_text()) + evidence = raw["known_gaps"][early]["evidence"] + assert evidence["classification"] == "source_unavailability" + assert evidence["retired_derivation"] == { + "repository_owner": "PolicyEngine", + "repository_name_parts": ["policyengine-", "us-data"], + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": [ + "policyengine_", + "us_data", + "datasets", + "cps", + "cps.py", + ], + "lines": "582-583,640-648", + } + assert evidence["retired_rate"]["lines"] == "1-9" + assert evidence["retired_rate"]["value"] == 0.09 + assert evidence["retired_randomness"]["lines"] == "5-28" + assert evidence["policyengine_variable"]["eligibility_domain"].startswith( + "age < 3 or is_pregnant" + ) + assert len(evidence["hermetic_asec_inputs"]) == 3 + assert all( + early in artifact["missing_columns"] + for artifact in evidence["hermetic_asec_inputs"] + ) + assert early in evidence["processed_puf"]["missing_columns"] + assert evidence["pinned_sipp"]["pregnancy_signal"] == "none" + assert ( + "person identity" + in evidence["administrative_non_substitutes"]["cumulative_enrollment"] + ) + def test_register_entries_are_unique(self) -> None: register = load_ecps_parity_known_gaps() names = [gap.variable for gap in register] diff --git a/packages/populace-build/tests/test_us_plan.py b/packages/populace-build/tests/test_us_plan.py index f8fd5f50..d80ba5a2 100644 --- a/packages/populace-build/tests/test_us_plan.py +++ b/packages/populace-build/tests/test_us_plan.py @@ -13,14 +13,45 @@ SupportSpineSpec, ) from populace.build.us_runtime import ( + ASEC_2023_WEEKS_UNEMPLOYED_MEMBER, + ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_CRC32, + ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_SHA256, + ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_SIZE_BYTES, + ASEC_2023_WEEKS_UNEMPLOYED_ZIP_SHA256, + ASEC_2023_WEEKS_UNEMPLOYED_ZIP_SIZE_BYTES, + SIPP_HEAD_START_FIT_PARAMETERS, + SIPP_HEAD_START_READ_PARAMETERS, + SIPP_SSI_DISABILITY_FIT_PARAMETERS, + SIPP_SSI_DISABILITY_READ_PARAMETERS, + SSI_TAKE_UP_SSA_SOURCE_URL, + US_CHILD_SUPPORT_STAGE_NAME, + US_CHILDCARE_STAGE_NAME, + US_DISABILITY_BENEFITS_STAGE_NAME, US_DONORS, + US_EDUCATION_INPUTS_STAGE_NAME, + US_ENERGY_SUBSIDY_STAGE_NAME, US_NONNEGATIVE_SOURCE_OUTPUTS, + US_OTHER_HEALTH_INSURANCE_STAGE_NAME, + US_PRIOR_YEAR_INCOME_STAGE_NAME, US_PUF_SUPPORT_STAGE_NAME, + US_QBI_OUTPUT_COLUMNS, + US_RETIREMENT_CONTRIBUTION_STAGE_NAME, + US_SIPP_HEAD_START_STAGE_NAME, US_SOURCE_MANIFEST, US_SOURCE_STAGE_SPECS, + US_SSI_DISABILITY_CRITERIA_STAGE_NAME, + US_SSI_TAKE_UP_ANCHOR, + US_SSI_TAKE_UP_STAGE_NAME, + US_SSI_TAKE_UP_TARGET_TABLE_NAME, US_STAGE_NAMES, US_SUPPORT_SPINE_MANIFEST, US_SUPPORT_SPINE_SPEC, + US_VOLUNTARY_FILING_STAGE_NAME, + US_WEEKS_UNEMPLOYED_STAGE_NAME, + US_WORKERS_COMPENSATION_STAGE_NAME, + WEEKS_UNEMPLOYED_DERIVE_PARAMETERS, + WEEKS_UNEMPLOYED_PUF_IMPUTATION_PARAMETERS, + WEEKS_UNEMPLOYED_READ_PARAMETERS, BuildConfig, us_plan, ) @@ -186,9 +217,382 @@ def test_source_specs_align_with_declared_plan(self) -> None: assert frame_structural_stages.issubset(US_STAGE_NAMES) def test_puf_support_channel_precedes_puf_detail_donor_stage(self) -> None: + assert US_STAGE_NAMES.index( + US_PRIOR_YEAR_INCOME_STAGE_NAME + ) < US_STAGE_NAMES.index(US_PUF_SUPPORT_STAGE_NAME) + assert US_STAGE_NAMES.index( + US_RETIREMENT_CONTRIBUTION_STAGE_NAME + ) < US_STAGE_NAMES.index(US_PUF_SUPPORT_STAGE_NAME) + assert US_STAGE_NAMES.index(US_CHILDCARE_STAGE_NAME) < US_STAGE_NAMES.index( + US_PUF_SUPPORT_STAGE_NAME + ) + assert US_STAGE_NAMES.index( + US_ENERGY_SUBSIDY_STAGE_NAME + ) < US_STAGE_NAMES.index(US_PUF_SUPPORT_STAGE_NAME) assert US_STAGE_NAMES.index(US_PUF_SUPPORT_STAGE_NAME) < US_STAGE_NAMES.index( "puf_tax_detail" ) + assert US_STAGE_NAMES.index("puf_tax_detail") < US_STAGE_NAMES.index( + US_CHILD_SUPPORT_STAGE_NAME + ) + assert US_STAGE_NAMES.index(US_CHILD_SUPPORT_STAGE_NAME) < US_STAGE_NAMES.index( + US_DISABILITY_BENEFITS_STAGE_NAME + ) + assert US_STAGE_NAMES.index( + US_DISABILITY_BENEFITS_STAGE_NAME + ) < US_STAGE_NAMES.index("education_inputs") + assert US_STAGE_NAMES.index( + US_WORKERS_COMPENSATION_STAGE_NAME + ) < US_STAGE_NAMES.index(US_WEEKS_UNEMPLOYED_STAGE_NAME) + assert US_STAGE_NAMES.index( + US_WEEKS_UNEMPLOYED_STAGE_NAME + ) < US_STAGE_NAMES.index(US_EDUCATION_INPUTS_STAGE_NAME) + assert US_STAGE_NAMES.index("medicaid_take_up") < US_STAGE_NAMES.index( + US_OTHER_HEALTH_INSURANCE_STAGE_NAME + ) + assert US_STAGE_NAMES.index( + US_OTHER_HEALTH_INSURANCE_STAGE_NAME + ) < US_STAGE_NAMES.index("export") + assert US_STAGE_NAMES.index("vehicle_assets") < US_STAGE_NAMES.index( + US_VOLUNTARY_FILING_STAGE_NAME + ) + assert US_STAGE_NAMES.index( + US_VOLUNTARY_FILING_STAGE_NAME + ) < US_STAGE_NAMES.index("entity_placement") + assert US_STAGE_NAMES.index("scf_wealth") < US_STAGE_NAMES.index( + US_SSI_DISABILITY_CRITERIA_STAGE_NAME + ) + assert US_STAGE_NAMES.index( + US_SSI_DISABILITY_CRITERIA_STAGE_NAME + ) < US_STAGE_NAMES.index(US_SIPP_HEAD_START_STAGE_NAME) + assert US_STAGE_NAMES.index( + US_SIPP_HEAD_START_STAGE_NAME + ) < US_STAGE_NAMES.index(US_SSI_TAKE_UP_STAGE_NAME) + assert US_STAGE_NAMES.index(US_SSI_TAKE_UP_STAGE_NAME) < US_STAGE_NAMES.index( + "sipp_tips" + ) + + def test_ssi_disability_criteria_stage_pins_sipp_model_contract(self) -> None: + stage = US_SOURCE_MANIFEST.stage_map()[US_SSI_DISABILITY_CRITERIA_STAGE_NAME] + donor = US_DONORS[US_SSI_DISABILITY_CRITERIA_STAGE_NAME] + + assert stage.survey == donor.survey == "Census SIPP" + assert stage.source == donor.source + assert stage.grain == "person" + assert stage.outputs == ("meets_ssi_disability_criteria",) + assert stage.nonnegative_outputs == () + assert [operation.kind for operation in stage.operations] == [ + "read_table", + "fit_weighted_qrf", + ] + assert ( + dict(stage.operations[0].parameters) == SIPP_SSI_DISABILITY_READ_PARAMETERS + ) + assert ( + dict(stage.operations[1].parameters) == SIPP_SSI_DISABILITY_FIT_PARAMETERS + ) + assert stage.operations[1].parameters["training_sample_seed"] == ( + 8_386_123_572_872_638_692 + ) + assert stage.operations[1].parameters["model_seed"] == 42 + assert stage.operations[1].parameters["seed_from_build_config"] is False + assert "separately predict PUF-support people" in stage.notes + assert "ASEC reporter anchor is never copied" in stage.notes + + def test_head_start_stage_pins_measured_sipp_response_contract(self) -> None: + stage = US_SOURCE_MANIFEST.stage_map()[US_SIPP_HEAD_START_STAGE_NAME] + donor = US_DONORS[US_SIPP_HEAD_START_STAGE_NAME] + + assert stage.survey == donor.survey == "Census SIPP" + assert stage.source == donor.source + assert stage.grain == "person" + assert stage.outputs == ("takes_up_head_start_if_eligible",) + assert stage.nonnegative_outputs == () + assert [operation.kind for operation in stage.operations] == [ + "read_table", + "fit_weighted_qrf", + ] + assert dict(stage.operations[0].parameters) == SIPP_HEAD_START_READ_PARAMETERS + assert dict(stage.operations[1].parameters) == SIPP_HEAD_START_FIT_PARAMETERS + assert stage.operations[1].parameters["target"] == ( + "takes_up_head_start_if_eligible" + ) + assert stage.operations[1].parameters["direct_response_filter"] == ( + "AEDHEADST == 1 and EEDHEADST in [1, 2]" + ) + assert stage.operations[1].parameters["assignment_unit"] == ("person_source_id") + assert stage.operations[1].parameters["fan_to_support_clones"] is True + assert any( + artifact.get("sha256") and artifact.get("size_bytes") + for artifact in stage.artifacts + ) + assert "EEDHEADST" in stage.notes + assert "AEDHEADST" in stage.notes + assert "measured" in stage.notes.lower() + + def test_ssi_take_up_stage_pins_reporter_and_ssa_count_contract(self) -> None: + stage = US_SOURCE_MANIFEST.stage_map()[US_SSI_TAKE_UP_STAGE_NAME] + donor = US_DONORS[US_SSI_TAKE_UP_STAGE_NAME] + + assert ( + stage.survey + == donor.survey + == ("CPS ASEC reported SSI + SSA SSI Monthly Statistics December 2024") + ) + assert stage.source == donor.source == SSI_TAKE_UP_SSA_SOURCE_URL + assert stage.grain == "person" + assert stage.outputs == ("takes_up_ssi_if_eligible",) + assert stage.nonnegative_outputs == () + assert [operation.kind for operation in stage.operations] == [ + "read_table", + "assign_binary_from_rate", + "calibrate_binary_assignment", + ] + assert dict(stage.operations[0].parameters) == { + "table": "person", + "weight": "person_weight", + } + assert dict(stage.operations[1].parameters) == { + "output": "takes_up_ssi_if_eligible", + "draw": "stable_source_person_draw", + "rate_key": "ssi_age_band_count_prior", + "rate_column": "ssi_take_up_assignment_prior", + "reported_true_anchor": f"{US_SSI_TAKE_UP_ANCHOR} > 0", + "assignment_unit": "person_source_id", + "fan_to_support_clones": True, + } + assert dict(stage.operations[2].parameters) == { + "variable": "takes_up_ssi_if_eligible", + "targets": [US_SSI_TAKE_UP_TARGET_TABLE_NAME], + "preserve_true_anchors": True, + "preserve_true_anchor": f"{US_SSI_TAKE_UP_ANCHOR} > 0", + "domain": "uncapped_ssi > 0", + "weight": "person_weight", + "draw": "stable_source_person_draw", + "calibration_unit": "person_source_id", + "age_bands": { + "under_18": "age < 18", + "18_64": "18 <= age < 65", + "65_plus": "age >= 65", + }, + "target_source": SSI_TAKE_UP_SSA_SOURCE_URL, + "target_period": "2024-12", + "target_measure": "Total with—Federal payment", + "target_values": { + "under_18": 1_001_922, + "18_64": 3_905_779, + "65_plus": 2_382_142, + }, + "aggregate_target": 7_289_843, + } + ssa_artifacts = [ + artifact + for artifact in stage.artifacts + if artifact.get("source") == SSI_TAKE_UP_SSA_SOURCE_URL + ] + assert len(ssa_artifacts) == 1 + evidence = next( + artifact + for artifact in stage.artifacts + if artifact.get("kind") == "archived_derivation_evidence" + ) + assert evidence["commit"] == ("42ed5d45c56df80d754fbe24cce21cfeb8d05cbe") + assert evidence["lines"] == "584,650-657,1497-1499" + assert evidence["randomness_path_parts"] == [ + "policyengine_", + "us_data", + "datasets", + "cps", + "takeup.py", + ] + assert evidence["randomness_lines"] == "10-35" + assert evidence["targets_lines"] == "41-74" + + def test_other_health_insurance_donor_matches_manifest(self) -> None: + stage = US_SOURCE_MANIFEST.stage_map()[US_OTHER_HEALTH_INSURANCE_STAGE_NAME] + donor = US_DONORS[US_OTHER_HEALTH_INSURANCE_STAGE_NAME] + + assert donor.survey == stage.survey == "Census CPS ASEC" + assert donor.source == stage.source + assert "PHIP_VAL" in donor.notes + assert "PUF support half" in donor.notes + + def test_prior_year_income_stage_pins_join_fallback_and_joint_qrf(self) -> None: + stage = US_SOURCE_MANIFEST.stage_map()[US_PRIOR_YEAR_INCOME_STAGE_NAME] + + assert stage.outputs == ( + "employment_income_last_year", + "self_employment_income_last_year", + "previous_year_income_available", + ) + assert stage.nonnegative_outputs == ("employment_income_last_year",) + operations = {operation.kind: operation for operation in stage.operations} + assert tuple(operations) == ( + "read_table", + "derive_prior_year_income", + "impute_prior_year_income_to_puf_support", + ) + derive = operations["derive_prior_year_income"].parameters + assert derive["employment_allocation_flag"] == "I_ERNVAL" + assert derive["self_employment_allocation_flag"] == "I_SEVAL" + assert derive["sentinels"] == [-1, -9999] + assert derive["no_prior_artifact"] == "leave_defaults" + qrf = operations["impute_prior_year_income_to_puf_support"].parameters + assert qrf["max_train_samples"] == 5_000 + assert qrf["weight"] == "person_weight" + assert qrf["outputs"] == [ + "employment_income_last_year", + "self_employment_income_last_year", + ] + assert "self_employment_income_last_year" not in US_NONNEGATIVE_SOURCE_OUTPUTS + + donor = US_DONORS[US_PRIOR_YEAR_INCOME_STAGE_NAME] + assert donor.survey == stage.survey + assert donor.source == stage.source + + def test_weeks_unemployed_stage_pins_direct_source_and_puf_qrf(self) -> None: + stage = US_SOURCE_MANIFEST.stage_map()[US_WEEKS_UNEMPLOYED_STAGE_NAME] + donor = US_DONORS[US_WEEKS_UNEMPLOYED_STAGE_NAME] + + assert stage.survey == donor.survey == "Census CPS ASEC" + assert stage.source == donor.source + assert stage.grain == "person" + assert stage.outputs == ("weeks_unemployed",) + assert stage.nonnegative_outputs == stage.outputs + + archive = next( + artifact + for artifact in stage.artifacts + if artifact.get("format") == "zip_csv" + ) + assert archive["vintage"] == "2023 ASEC / 2022 income reference year" + assert archive["sha256"] == ASEC_2023_WEEKS_UNEMPLOYED_ZIP_SHA256 + assert archive["size_bytes"] == ASEC_2023_WEEKS_UNEMPLOYED_ZIP_SIZE_BYTES + assert archive["member"] == ASEC_2023_WEEKS_UNEMPLOYED_MEMBER + assert ( + archive["member_size_bytes"] == ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_SIZE_BYTES + ) + assert archive["member_crc32"] == ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_CRC32 + assert archive["member_sha256"] == ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_SHA256 + + operations = {operation.kind: operation for operation in stage.operations} + assert tuple(operations) == ( + "read_table", + "derive_weeks_unemployed", + "impute_weeks_unemployed_to_puf_support", + ) + assert dict(operations["read_table"].parameters) == ( + WEEKS_UNEMPLOYED_READ_PARAMETERS + ) + assert dict(operations["derive_weeks_unemployed"].parameters) == ( + WEEKS_UNEMPLOYED_DERIVE_PARAMETERS + ) + assert ( + operations["impute_weeks_unemployed_to_puf_support"].parameters + == WEEKS_UNEMPLOYED_PUF_IMPUTATION_PARAMETERS + ) + assert "no 2022 source value is filled statistically" in stage.notes + + def test_disability_benefits_stage_pins_direct_formula_and_puf_imputation( + self, + ) -> None: + stage = US_SOURCE_MANIFEST.stage_map()[US_DISABILITY_BENEFITS_STAGE_NAME] + + assert stage.survey == "Census CPS ASEC" + assert stage.source == "https://www.census.gov/programs-surveys/cps.html" + assert stage.grain == "person" + assert stage.outputs == ("disability_benefits",) + assert stage.nonnegative_outputs == stage.outputs + + operations = {operation.kind: operation for operation in stage.operations} + assert tuple(operations) == ( + "read_table", + "derive_disability_benefits", + "impute_disability_benefits_to_puf_support", + ) + assert operations["read_table"].parameters == { + "table": "person", + "weight": "person_weight", + } + assert operations["derive_disability_benefits"].parameters == { + "first_amount_source": "DIS_VAL1", + "first_code_source": "DIS_SC1", + "second_amount_source": "DIS_VAL2", + "second_code_source": "DIS_SC2", + "workers_compensation_code": 1, + "output": "disability_benefits", + } + assert operations["impute_disability_benefits_to_puf_support"].parameters == { + "predictors": [ + "age", + "is_male", + "has_esi", + "tax_unit_is_joint", + "tax_unit_count_dependents", + "employment_income", + "self_employment_income", + "social_security", + ], + "max_train_samples": 5_000, + "n_estimators": 100, + "seed_from_build_config": True, + "weight": "person_weight", + } + + donor = US_DONORS[US_DISABILITY_BENEFITS_STAGE_NAME] + assert donor.survey == stage.survey + assert donor.source == stage.source + + def test_child_support_stage_declares_direct_carry_and_joint_puf_imputation( + self, + ) -> None: + stage = US_SOURCE_MANIFEST.stage_map()[US_CHILD_SUPPORT_STAGE_NAME] + + assert stage.survey == "Census CPS ASEC" + assert stage.source == "https://www.census.gov/programs-surveys/cps.html" + assert stage.grain == "person" + assert stage.outputs == ( + "child_support_received", + "child_support_expense", + ) + assert stage.nonnegative_outputs == stage.outputs + + operations = {operation.kind: operation for operation in stage.operations} + assert tuple(operations) == ( + "read_table", + "derive_child_support_inputs", + "impute_child_support_to_puf_support", + ) + assert operations["read_table"].parameters == { + "table": "person", + "weight": "person_weight", + } + assert operations["derive_child_support_inputs"].parameters == { + "received_source": "CSP_VAL", + "received_output": "child_support_received", + "expense_source": "CHSP_VAL", + "expense_output": "child_support_expense", + } + assert operations["impute_child_support_to_puf_support"].parameters == { + "predictors": [ + "age", + "is_male", + "has_esi", + "tax_unit_is_joint", + "tax_unit_count_dependents", + "employment_income", + "self_employment_income", + "social_security", + ], + "max_train_samples": 5_000, + "n_estimators": 100, + "seed_from_build_config": True, + "weight": "person_weight", + } + + donor = US_DONORS[US_CHILD_SUPPORT_STAGE_NAME] + assert donor.survey == stage.survey + assert donor.source == stage.source def test_source_specs_are_manifest_only_not_python_loaders(self) -> None: for spec in US_SOURCE_STAGE_SPECS: @@ -304,10 +708,107 @@ def test_legacy_sources_import_path_is_metadata_only(self) -> None: def test_scf_nonnegative_requirement_is_source_declared(self) -> None: specs = US_SOURCE_MANIFEST.stage_map() - assert "auto_loan_interest" in specs["scf_wealth"].nonnegative_outputs - assert "auto_loan_interest" in US_NONNEGATIVE_SOURCE_OUTPUTS + for column in ( + "auto_loan_balance", + "auto_loan_interest", + "qualified_passenger_vehicle_loan_interest", + ): + assert column in specs["scf_wealth"].nonnegative_outputs + assert column in US_NONNEGATIVE_SOURCE_OUTPUTS assert "net_worth" not in US_NONNEGATIVE_SOURCE_OUTPUTS + def test_sipp_vehicle_stage_pins_full_donor_and_no_partial_net_worth(self) -> None: + spec = US_SOURCE_MANIFEST.stage_map()["vehicle_assets"] + artifact = spec.artifacts[0] + assert artifact["vintage"] == "2023" + assert artifact["sha256"] == ( + "5c30439e365fc26483318ef61d1d8f4bb2f0e9d6bb47c22c06756a7698733ee2" + ) + assert artifact["size_bytes"] == 3_726_010_471 + assert "21280dca5995e978d706740a8a4b9b7860cfd7b6" in artifact["locator"] + model = next( + operation + for operation in spec.operations + if operation.kind == "fit_vehicle_model" + ) + assert model.parameters["weight"] == "household_weight" + assert model.parameters["observed_status_values"] == [0, 1, 9] + assert model.parameters["owned_allocation_flag"] == "AVEH_NUM" + assert model.parameters["value_allocation_flags"] == [ + "AVEH1VAL", + "AVEH2VAL", + "AVEH3VAL", + ] + assert model.parameters["seed"] == 42 + assert model.parameters["n_estimators"] == 100 + assert all(operation.kind != "fold_into" for operation in spec.operations) + assert "does not execute the old placeholder fold" in spec.notes + assert set(spec.outputs) == { + "household_vehicles_owned", + "household_vehicles_value", + } + assert set(spec.outputs) <= set(spec.nonnegative_outputs) + + def test_voluntary_filing_stage_pins_measured_sipp_response(self) -> None: + spec = US_SOURCE_MANIFEST.stage_map()[US_VOLUNTARY_FILING_STAGE_NAME] + artifact = spec.artifacts[0] + assert artifact["vintage"] == "2023" + assert artifact["sha256"] == ( + "5c30439e365fc26483318ef61d1d8f4bb2f0e9d6bb47c22c06756a7698733ee2" + ) + assert artifact["size_bytes"] == 3_726_010_471 + assert "21280dca5995e978d706740a8a4b9b7860cfd7b6" in artifact["locator"] + + read, model = spec.operations + assert read.kind == "read_table" + assert read.parameters["month_column"] == "MONTHCODE" + assert read.parameters["month"] == 12 + assert read.parameters["source_columns"] == [ + "SSUID", + "PNUM", + "MONTHCODE", + "WPFINWGT", + "TAGE", + "ESEX", + "EPNSPOUSE", + "AFILING", + "EFILING", + "AWILLFILE", + "EWILLFILE", + "EDEPCLM", + "TJB1_MSUM", + "TJB2_MSUM", + "TJB3_MSUM", + "TJB4_MSUM", + "TJB5_MSUM", + "TJB6_MSUM", + "TJB7_MSUM", + ] + assert model.kind == "fit_weighted_qrf" + assert model.parameters["predictors"] == [ + "employment_income", + "reference_age", + "reference_is_female", + "reference_is_married", + "count_under_18", + ] + assert model.parameters["target"] == "would_file_taxes_voluntarily" + assert model.parameters["weight"] == "tax_unit_weight" + assert model.parameters["response_filter"] == ( + "AFILING == 1 and (EFILING == 1 or (EFILING == 2 and AWILLFILE == 1))" + ) + assert model.parameters["dependent_exclusion"] == "EDEPCLM == 1" + assert "reciprocal spouses" in model.parameters["canonical_unit"] + assert model.parameters["n_estimators"] == 100 + assert model.parameters["seed_from_build_config"] is True + assert spec.outputs == ("would_file_taxes_voluntarily",) + assert spec.nonnegative_outputs == () + assert "datasets/cps/cps.py lines 726-747" in spec.notes + assert "parameters/take_up/voluntary_filing.yaml lines 1-43" in spec.notes + assert "22,313 observed canonical units" in spec.notes + assert "22,296 have positive finite reference weights" in spec.notes + assert "0.7603084563 weighted true share" in spec.notes + def test_puf_stage_declares_partnership_s_corp_leaves_not_aggregate(self) -> None: outputs = set(US_SOURCE_MANIFEST.stage_map()["puf_tax_detail"].outputs) assert "partnership_income" in outputs @@ -327,13 +828,24 @@ def test_puf_stage_declares_partnership_s_corp_leaves_not_aggregate(self) -> Non assert "charitable_deduction" not in outputs assert "interest_deduction" not in outputs - def test_puf_stage_outputs_match_runtime_defaults(self) -> None: + def test_puf_stage_declares_all_materialized_qbi_input_leaves(self) -> None: outputs = set(US_SOURCE_MANIFEST.stage_map()["puf_tax_detail"].outputs) + + assert set(US_QBI_OUTPUT_COLUMNS) <= outputs + assert set(US_QBI_OUTPUT_COLUMNS) <= set(PUF_TAX_DETAIL_DEFAULT_PERSON_OUTPUTS) + assert "qualified_business_income" not in outputs + + def test_puf_stage_outputs_match_runtime_defaults(self) -> None: + stage = US_SOURCE_MANIFEST.stage_map()["puf_tax_detail"] + outputs = set(stage.outputs) runtime_outputs = set(PUF_TAX_DETAIL_DEFAULT_PERSON_OUTPUTS) | set( PUF_TAX_DETAIL_DEFAULT_TAX_UNIT_OUTPUTS ) assert outputs == runtime_outputs + assert "educator_expense" in stage.outputs + assert "educator_expense" in stage.nonnegative_outputs + assert "educator_expense" in PUF_TAX_DETAIL_DEFAULT_PERSON_OUTPUTS def test_puf_stage_disaggregates_aggregate_records_before_uprating(self) -> None: operations = US_SOURCE_MANIFEST.stage_map()["puf_tax_detail"].operations @@ -351,6 +863,37 @@ def test_puf_stage_disaggregates_aggregate_records_before_uprating(self) -> None "qualified_dividend_source": "E00650", "qualified_dividend_output": "qualified_dividend_income", "non_qualified_dividend_output": "non_qualified_dividend_income", + "qualified_tuition_primary_source": "E03230", + "qualified_tuition_optional_source": "E87530", + "qualified_tuition_output": "qualified_tuition_expenses", + "alimony_income_source": "E00800", + "alimony_income_output": "alimony_income", + "alimony_expense_source": "E03500", + "alimony_expense_output": "alimony_expense", + "casualty_loss_source": "E20500", + "casualty_loss_output": "casualty_loss", + "domestic_production_ald_source": "E03240", + "domestic_production_ald_output": "domestic_production_ald", + "educator_expense_source": "E03220", + "educator_expense_output": "educator_expense", + "unreimbursed_business_employee_expenses_source": "E20400", + "unreimbursed_business_employee_expenses_output": "unreimbursed_business_employee_expenses", + "farm_operations_income_source": "E02100", + "farm_operations_income_output": "farm_operations_income", + "farm_rent_income_source": "E27200", + "farm_rent_income_output": "farm_rent_income", + "investment_income_elected_form_4952_source": "E58990", + "investment_income_elected_form_4952_output": ( + "investment_income_elected_form_4952" + ), + "salt_refund_income_source": "E00700", + "salt_refund_income_output": "salt_refund_income", + "collectibles_capital_gain_source": "E24518", + "collectibles_capital_gain_output": ( + "long_term_capital_gains_on_collectibles" + ), + "unrecaptured_section_1250_gain_source": "E24515", + "unrecaptured_section_1250_gain_output": "unrecaptured_section_1250_gain", } operation = operations[kinds.index("disaggregate_aggregate_records")] diff --git a/packages/populace-build/tests/test_us_prior_year_income.py b/packages/populace-build/tests/test_us_prior_year_income.py new file mode 100644 index 00000000..86c9ce6c --- /dev/null +++ b/packages/populace-build/tests/test_us_prior_year_income.py @@ -0,0 +1,491 @@ +"""Archived adjacent-ASEC prior-year-income restoration.""" + +from __future__ import annotations + +import importlib.util +from pathlib import Path + +import numpy as np +import pandas as pd +import pytest + +import populace.build.us_runtime.prior_year_income as module +from populace.build.source_runtime import SourceRuntimeError +from populace.build.us_runtime.l0_refit_export import ( + US_RELEASE_REQUIRED_PERSON_SOURCE_COLUMNS, +) +from populace.build.us_runtime.prior_year_income import ( + PRIOR_YEAR_INCOME_ARCHIVED_DERIVATION_URL, + PRIOR_YEAR_INCOME_ARCHIVED_FINALIZER_URL, + PRIOR_YEAR_INCOME_ARCHIVED_FORMULA_OUTPUT_URL, + PRIOR_YEAR_INCOME_ARCHIVED_PUF_IMPUTATION_URL, + PRIOR_YEAR_INCOME_ARCHIVED_PUF_OUTPUTS_URL, + PRIOR_YEAR_INCOME_ARCHIVED_PUF_SPLICE_URL, + US_PRIOR_YEAR_INCOME_NONCONSTANT_PERSON_COLUMNS, + US_PRIOR_YEAR_INCOME_OUTPUT_COLUMNS, + US_PRIOR_YEAR_INCOME_PERSISTED_OUTPUT_COLUMNS, + derive_us_prior_year_income_from_manifest, + impute_us_prior_year_income_to_puf_support_from_manifest, + us_prior_year_income_signal_gate, + us_prior_year_income_source_reconciliation_gate, + us_prior_year_income_stage_spec, + with_us_prior_year_income_inputs, +) +from populace.build.us_runtime.puf_support import ( + BASE_ASEC_SUPPORT_CHANNEL, + PUF_TAX_DETAIL_SUPPORT_CHANNEL, + clone_us_frame_for_puf_support, +) +from populace.build.us_runtime.release_input_coverage import ( + RESTORED_REFERENCE_ECPS_REQUIRED_INPUTS, + load_release_input_coverage_manifest, + us_release_reform_coverage_probes, +) +from populace.frame import US_SCHEMA, Frame, WeightKind, Weights +from populace.frame.adapters.policyengine_us import PolicyEngineUSEngine + +policyengine_us_installed = importlib.util.find_spec("policyengine_us") is not None +requires_us = pytest.mark.skipif( + not policyengine_us_installed, + reason="requires the policyengine-us [us] extra (build environment)", +) + + +def _frame(person: pd.DataFrame) -> Frame: + person = person.reset_index(drop=True).copy() + n = len(person) + ids = np.arange(1, n + 1, dtype=np.int64) + defaults: dict[str, object] = { + "person_id": ids, + "person_household_id": ids, + "person_tax_unit_id": ids + 100, + "person_spm_unit_id": ids + 200, + "person_family_id": ids + 300, + "person_marital_unit_id": ids + 400, + "tax_unit_role_input": np.full(n, "HEAD", dtype=object), + "age": np.linspace(25, 65, n), + "is_female": ids % 2 == 0, + "has_esi": ids % 3 == 0, + "employment_income_before_lsr": np.zeros(n), + "self_employment_income_before_lsr": np.zeros(n), + "social_security_retirement": np.zeros(n), + "social_security_disability": np.zeros(n), + "social_security_dependents": np.zeros(n), + "social_security_survivors": np.zeros(n), + } + for column, values in defaults.items(): + if column not in person: + person[column] = values + person["employment_income_before_lsr"] = pd.to_numeric( + person.get("WSAL_VAL", person["employment_income_before_lsr"]), + errors="coerce", + ).replace({-1: 0, -9999: 0}) + person["self_employment_income_before_lsr"] = pd.to_numeric( + person.get("SEMP_VAL", person["self_employment_income_before_lsr"]), + errors="coerce", + ).replace({-1: 0, -9999: 0}) + tables = { + "person": person, + "household": pd.DataFrame( + {"household_id": ids, "state_fips": np.resize([6, 36], n)} + ), + "tax_unit": pd.DataFrame( + { + "tax_unit_id": ids + 100, + "filing_status_input": np.resize(["SINGLE", "JOINT"], n), + } + ), + "spm_unit": pd.DataFrame({"spm_unit_id": ids + 200}), + "family": pd.DataFrame({"family_id": ids + 300}), + "marital_unit": pd.DataFrame({"marital_unit_id": ids + 400}), + } + return Frame( + tables, + US_SCHEMA, + {"household": Weights(np.linspace(1.0, 2.0, n), WeightKind.DESIGN)}, + ) + + +def _source_frame() -> Frame: + return _frame( + pd.DataFrame( + { + "source_year": [2022, 2023, 2024, 2023, 2024, 2023, 2024, 2024], + "PERIDNUM": ["A", "A", "A", "B", "B", "C", "C", "D"], + "WSAL_VAL": [100, 200, 300, 10, 400, 300, 500, 600], + "SEMP_VAL": [-20, 30, 40, 20, -5, -9999, 50, 60], + "I_ERNVAL": [0, 0, 0, 1, 0, 0, 0, 0], + "I_SEVAL": [0, 0, 0, 0, 0, 0, 0, 0], + } + ) + ) + + +def test_stage_manifest_pins_archived_join_signedness_and_puf_qrf() -> None: + spec = us_prior_year_income_stage_spec() + + assert spec.stage == "prior_year_income" + assert spec.outputs == US_PRIOR_YEAR_INCOME_OUTPUT_COLUMNS + assert spec.nonnegative_outputs == ("employment_income_last_year",) + operations = {operation.kind: operation for operation in spec.operations} + assert tuple(operations) == ( + "read_table", + "derive_prior_year_income", + "impute_prior_year_income_to_puf_support", + ) + derive = operations["derive_prior_year_income"].parameters + assert derive["person_id"] == "PERIDNUM" + assert derive["prior_year_offset"] == -1 + assert derive["employment_allocation_flag"] == "I_ERNVAL" + assert derive["self_employment_allocation_flag"] == "I_SEVAL" + assert derive["unallocated_flag"] == 0 + assert derive["sentinels"] == [-1, -9999] + assert derive["fallback_to_current"] is True + assert derive["no_prior_artifact"] == "leave_defaults" + qrf = operations["impute_prior_year_income_to_puf_support"].parameters + assert qrf["predictors"] == [ + "age", + "is_male", + "has_esi", + "tax_unit_is_joint", + "tax_unit_count_dependents", + "employment_income", + "self_employment_income", + "social_security", + ] + assert qrf["outputs"] == [ + "employment_income_last_year", + "self_employment_income_last_year", + ] + assert qrf["max_train_samples"] == 5_000 + assert qrf["n_estimators"] == 100 + assert qrf["weight"] == "person_weight" + assert all( + url.startswith( + "https://github.com/PolicyEngine/policyengine-" + "us-data/blob/" + "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe/" + ) + for url in ( + PRIOR_YEAR_INCOME_ARCHIVED_DERIVATION_URL, + PRIOR_YEAR_INCOME_ARCHIVED_PUF_OUTPUTS_URL, + PRIOR_YEAR_INCOME_ARCHIVED_PUF_IMPUTATION_URL, + PRIOR_YEAR_INCOME_ARCHIVED_PUF_SPLICE_URL, + PRIOR_YEAR_INCOME_ARCHIVED_FORMULA_OUTPUT_URL, + PRIOR_YEAR_INCOME_ARCHIVED_FINALIZER_URL, + ) + ) + + +def test_adjacent_join_matches_flags_sentinels_fallback_and_signed_losses() -> None: + result = with_us_prior_year_income_inputs( + _source_frame(), seed=0, time_period=2024 + ).table("person") + + assert result["previous_year_income_available"].tolist() == [ + False, + True, + True, + False, + False, + False, + False, + False, + ] + assert result["employment_income_last_year"].tolist() == [ + 0, + 100, + 200, + 10, + 400, + 300, + 500, + 600, + ] + assert result["self_employment_income_last_year"].tolist() == [ + 0, + -20, + 30, + 20, + -5, + 0, + 50, + 60, + ] + + +def test_valid_zero_prior_values_are_available_not_fallbacks() -> None: + frame = _frame( + pd.DataFrame( + { + "source_year": [2023, 2024], + "PERIDNUM": ["0000000000000000000001"] * 2, + "WSAL_VAL": [0, 50_000], + "SEMP_VAL": [0, 10_000], + "I_ERNVAL": [0, 0], + "I_SEVAL": [0, 0], + } + ) + ) + + person = with_us_prior_year_income_inputs(frame, seed=0, time_period=2024).table( + "person" + ) + + assert person["previous_year_income_available"].tolist() == [False, True] + assert person["employment_income_last_year"].tolist() == [0, 0] + assert person["self_employment_income_last_year"].tolist() == [0, 0] + + +def test_existing_default_outputs_rederive_when_raw_sources_remain() -> None: + frame = _source_frame() + person = frame.table("person").copy() + person["employment_income_last_year"] = 0.0 + person["self_employment_income_last_year"] = 0.0 + person["previous_year_income_available"] = False + stale = module._replace_person_table(frame, person) + + result = with_us_prior_year_income_inputs(stale, seed=0, time_period=2024).table( + "person" + ) + + assert result["previous_year_income_available"].any() + assert result["self_employment_income_last_year"].tolist() != [0.0] * len(result) + + +def test_join_rejects_duplicate_source_year_person_key() -> None: + frame = _source_frame() + person = frame.table("person").copy() + duplicate = person.iloc[[0]].copy() + duplicate["person_id"] = 999 + duplicate["person_household_id"] = 999 + duplicate["person_tax_unit_id"] = 999 + duplicate["person_spm_unit_id"] = 999 + duplicate["person_family_id"] = 999 + duplicate["person_marital_unit_id"] = 999 + operation = us_prior_year_income_stage_spec().operations[1] + + with pytest.raises(SourceRuntimeError, match="duplicate key"): + derive_us_prior_year_income_from_manifest( + pd.concat([person, duplicate], ignore_index=True), operation, None + ) + + +def test_join_refuses_missing_allocation_source() -> None: + person = _source_frame().table("person").drop(columns=["I_SEVAL"]) + operation = us_prior_year_income_stage_spec().operations[1] + + with pytest.raises(SourceRuntimeError, match="I_SEVAL"): + derive_us_prior_year_income_from_manifest(person, operation, None) + + +class _Fitted: + def predict(self, test: pd.DataFrame) -> pd.DataFrame: + n = len(test) + self_employment = np.zeros(n, dtype=np.float64) + self_employment[:2] = [125.0, -25.0] + return pd.DataFrame( + { + "employment_income_last_year": np.arange(n) + 1_000.0, + "self_employment_income_last_year": self_employment, + }, + index=test.index, + ) + + +class _QRF: + calls: list[dict[str, object]] = [] + + def __init__(self, **kwargs: object) -> None: + self.kwargs = kwargs + + def fit( + self, + training: pd.DataFrame, + predictors: list[str], + targets: list[str], + *, + weights: np.ndarray, + ) -> _Fitted: + self.calls.append( + { + "kwargs": self.kwargs, + "training": training.copy(), + "predictors": predictors, + "targets": targets, + "weights": weights.copy(), + } + ) + return _Fitted() + + +def test_puf_support_joint_qrf_is_weighted_signed_and_drops_formula_output( + monkeypatch: pytest.MonkeyPatch, +) -> None: + monkeypatch.setattr(module, "QRF", _QRF) + _QRF.calls.clear() + direct = with_us_prior_year_income_inputs(_source_frame(), seed=7, time_period=2024) + expanded = clone_us_frame_for_puf_support(direct) + + result = with_us_prior_year_income_inputs(expanded, seed=7, time_period=2024) + person = result.table("person") + assert "employment_income_last_year" not in person + channel = person["person_support_channel"].astype(str) + puf_values = person.loc[ + channel == PUF_TAX_DETAIL_SUPPORT_CHANNEL, + "self_employment_income_last_year", + ].to_numpy() + assert puf_values[:2].tolist() == [125.0, -25.0] + asec_values = person.loc[ + channel == BASE_ASEC_SUPPORT_CHANNEL, + "self_employment_income_last_year", + ].to_numpy() + assert asec_values.tolist() == [0, -20, 30, 20, -5, 0, 50, 60] + + assert len(_QRF.calls) == 1 + call = _QRF.calls[0] + assert call["targets"] == [ + "employment_income_last_year", + "self_employment_income_last_year", + ] + assert call["predictors"] == list(module._PUF_PREDICTORS) + assert np.all(np.asarray(call["weights"]) > 0) + assert call["kwargs"] == {"n_estimators": 100, "seed": 7} + + +def test_puf_qrf_rejects_zero_weight_capped_training_sample() -> None: + n_asec = 5_001 + sampled = ( + pd.DataFrame(index=np.arange(n_asec)).sample(n=5_000, random_state=0).index + ) + omitted = int(next(iter(set(range(n_asec)) - set(sampled)))) + weights = np.zeros(n_asec + 1, dtype=np.float64) + weights[omitted] = 1.0 + person = pd.DataFrame( + { + "person_support_channel": [BASE_ASEC_SUPPORT_CHANNEL] * n_asec + + [PUF_TAX_DETAIL_SUPPORT_CHANNEL], + "person_weight": weights, + "employment_income_last_year": np.ones(n_asec + 1), + "self_employment_income_last_year": np.ones(n_asec + 1), + } + ) + for predictor in module._PUF_PREDICTORS: + person[module._PUF_PREDICTOR_PREFIX + predictor] = 1.0 + operation = us_prior_year_income_stage_spec().operations[2] + + with pytest.raises(SourceRuntimeError, match="sampled training weights"): + impute_us_prior_year_income_to_puf_support_from_manifest( + person, operation, None + ) + + +def _signal_frame() -> Frame: + n = 100 + return _frame( + pd.DataFrame( + { + "self_employment_income_last_year": np.where( + np.arange(n) < 4, + np.resize([10_000.0, -2_000.0], n), + 0.0, + ), + "previous_year_income_available": np.arange(n) < 20, + } + ) + ) + + +def test_signal_gate_accepts_signed_source_signal_and_rejects_defaults() -> None: + passing = us_prior_year_income_signal_gate(_signal_frame()) + assert passing.passed, passing.failures + assert passing.details["self_employment_income_last_year_negative_rows"] > 0 + + frame = _signal_frame() + person = frame.table("person").copy() + person["previous_year_income_available"] = False + failing = us_prior_year_income_signal_gate( + module._replace_person_table(frame, person) + ) + assert not failing.passed + assert "availability" in " ".join(failing.failures) + + +def test_source_reconciliation_detects_plausible_but_wrong_asec_carry() -> None: + derived = with_us_prior_year_income_inputs( + _source_frame(), seed=0, time_period=2024 + ) + passing = us_prior_year_income_source_reconciliation_gate(derived) + assert passing.passed, passing.failures + + person = derived.table("person").copy() + person.loc[1, "self_employment_income_last_year"] = 30.0 + corrupted = module._replace_person_table(derived, person) + failing = us_prior_year_income_source_reconciliation_gate(corrupted) + assert not failing.passed + assert failing.details["mismatch_counts"] == { + "employment_income_last_year": 0, + "self_employment_income_last_year": 1, + "previous_year_income_available": 0, + } + + +def test_release_contract_promotes_both_persisted_inputs_without_wage_formula() -> None: + manifest = load_release_input_coverage_manifest() + for column in US_PRIOR_YEAR_INCOME_PERSISTED_OUTPUT_COLUMNS: + assert column in RESTORED_REFERENCE_ECPS_REQUIRED_INPUTS + assert column in US_RELEASE_REQUIRED_PERSON_SOURCE_COLUMNS + assert column in manifest.required_columns + assert column not in manifest.reviewed_exclusions + assert US_PRIOR_YEAR_INCOME_NONCONSTANT_PERSON_COLUMNS == ( + "self_employment_income_last_year", + "previous_year_income_available", + ) + assert "employment_income_last_year" not in manifest.required_columns + + probe = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "prior_year_self_employment_neutralization" + ) + assert probe.neutralized_variable == "self_employment_income_last_year" + assert probe.parameter_changes == {} + assert probe.binding_inputs == ("self_employment_income_last_year",) + assert probe.budget_measure == "tax_unit_earned_income_last_year" + assert probe.effect_direction == "baseline_minus_reform" + assert probe.expected_sign == "positive" + + +@requires_us +def test_policyengine_17646_input_and_dependency_contract() -> None: + engine = PolicyEngineUSEngine() + assert set(US_PRIOR_YEAR_INCOME_PERSISTED_OUTPUT_COLUMNS) <= set(engine.variables()) + assert engine.formula_owned_outputs(US_PRIOR_YEAR_INCOME_OUTPUT_COLUMNS) == { + "employment_income_last_year" + } + system = engine._tax_benefit_system() + assert system.variables["earned_income_last_year"].adds == [ + "employment_income_last_year", + "self_employment_income_last_year", + ] + assert system.variables["tax_unit_earned_income_last_year"].entity.key == "tax_unit" + + +@requires_us +def test_export_ready_support_persists_inputs_and_excludes_formula_owned_wage( + monkeypatch: pytest.MonkeyPatch, + tmp_path: Path, +) -> None: + monkeypatch.setattr(module, "QRF", _QRF) + direct = with_us_prior_year_income_inputs(_source_frame(), seed=0, time_period=2024) + result = with_us_prior_year_income_inputs( + clone_us_frame_for_puf_support(direct), seed=0, time_period=2024 + ) + output = tmp_path / "prior_year_income.h5" + + PolicyEngineUSEngine().write_dataset(result, output, period=2024) + + with pd.HDFStore(output, mode="r") as store: + person = store["person"] + assert "self_employment_income_last_year" in person + assert "previous_year_income_available" in person + assert "employment_income_last_year" not in person diff --git a/packages/populace-build/tests/test_us_puf_support.py b/packages/populace-build/tests/test_us_puf_support.py index cb6d51d1..84418428 100644 --- a/packages/populace-build/tests/test_us_puf_support.py +++ b/packages/populace-build/tests/test_us_puf_support.py @@ -115,6 +115,7 @@ def _raw_asec_predictor_frame() -> Frame: "RNT_VAL": [20.0, 0.0, 0.0], "FRSE_VAL": [5.0, 0.0, 0.0], "UC_VAL": [0.0, 70.0, 0.0], + "OI_OFF": [19, 0, 0], "OI_VAL": [3.0, 0.0, 0.0], "PHIP_VAL": [400.0, 0.0, 50.0], "PEMCPREM": [100.0, 0.0, 25.0], @@ -236,6 +237,7 @@ def test_puf_tax_unit_donor_from_arrays_aggregates_person_values() -> None: "non_qualified_dividend_income": [1.0, 2.0, 3.0], "qualified_dividend_income": [4.0, 5.0, 6.0], "home_mortgage_interest": [10.0, 20.0, 30.0], + "educator_expense": [100.0, 200.0, 300.0], "social_security": [100.0, 200.0, 300.0], "taxable_unemployment_compensation": [13.0, 17.0, 19.0], "state_and_local_sales_or_income_tax": [40.0, 50.0], @@ -245,6 +247,7 @@ def test_puf_tax_unit_donor_from_arrays_aggregates_person_values() -> None: "qualified_dividend_income", "non_qualified_dividend_income", "home_mortgage_interest", + "educator_expense", "social_security_retirement", "social_security_disability", "social_security_dependents", @@ -266,6 +269,7 @@ def test_puf_tax_unit_donor_from_arrays_aggregates_person_values() -> None: assert "social_security" not in donor assert donor["unemployment_compensation"].tolist() == [30.0, 19.0] assert donor["home_mortgage_interest"].tolist() == [30.0, 30.0] + assert donor["educator_expense"].tolist() == [300.0, 300.0] assert "interest_deduction" not in donor assert "state_withheld_income_tax" not in donor assert donor["puf_predictor_employment_income"].tolist() == [12.0, 11.0] @@ -295,8 +299,16 @@ def test_puf_tax_detail_default_person_outputs_are_engine_leaves() -> None: assert "long_term_capital_gains" not in PUF_TAX_DETAIL_DEFAULT_PERSON_OUTPUTS assert "taxable_private_pension_income" in PUF_TAX_DETAIL_DEFAULT_PERSON_OUTPUTS assert "taxable_pension_income" not in PUF_TAX_DETAIL_DEFAULT_PERSON_OUTPUTS + assert "qualified_tuition_expenses" in PUF_TAX_DETAIL_DEFAULT_PERSON_OUTPUTS + assert "educator_expense" in PUF_TAX_DETAIL_DEFAULT_PERSON_OUTPUTS + assert "casualty_loss" in PUF_TAX_DETAIL_DEFAULT_PERSON_OUTPUTS + assert ( + "unreimbursed_business_employee_expenses" + in PUF_TAX_DETAIL_DEFAULT_PERSON_OUTPUTS + ) assert "medical_expense_deduction" not in PUF_TAX_DETAIL_DEFAULT_PERSON_OUTPUTS assert "interest_deduction" not in PUF_TAX_DETAIL_DEFAULT_TAX_UNIT_OUTPUTS + assert "domestic_production_ald" in PUF_TAX_DETAIL_DEFAULT_TAX_UNIT_OUTPUTS assert "first_home_mortgage_interest" in PUF_TAX_DETAIL_DEFAULT_TAX_UNIT_OUTPUTS assert "second_home_mortgage_interest" in PUF_TAX_DETAIL_DEFAULT_TAX_UNIT_OUTPUTS assert "first_home_mortgage_balance" in PUF_TAX_DETAIL_DEFAULT_TAX_UNIT_OUTPUTS @@ -335,6 +347,90 @@ def test_puf_tax_detail_default_person_outputs_are_engine_leaves() -> None: assert "health_savings_account_ald" in PUF_TAX_DETAIL_DEFAULT_TAX_UNIT_OUTPUTS +def test_puf_tax_unit_donor_derives_qualified_tuition_from_raw_fields() -> None: + donor = puf_tax_unit_donor_from_arrays( + { + "tax_unit_id": [10, 20], + "household_weight": [100.0, 200.0], + "filing_status": [b"SINGLE", b"JOINT"], + "person_tax_unit_id": [10, 10, 20], + "E03230": [500.0, 2_000.0, -10.0], + "E87530": [1_000.0, 1_500.0, 4_000.0], + }, + person_outputs=("qualified_tuition_expenses",), + tax_unit_outputs=(), + ) + + assert donor["qualified_tuition_expenses"].tolist() == [3_000.0, 4_000.0] + + +def test_puf_tax_unit_donor_carries_person_educator_expense() -> None: + donor = puf_tax_unit_donor_from_arrays( + { + "tax_unit_id": [10, 20], + "household_weight": [100.0, 200.0], + "filing_status": [b"SINGLE", b"JOINT"], + "person_tax_unit_id": [10, 10, 20], + "educator_expense": [250.0, 50.0, 0.0], + }, + person_outputs=("educator_expense",), + tax_unit_outputs=(), + ) + + assert donor["educator_expense"].tolist() == [300.0, 0.0] + + +def test_puf_tax_unit_donor_carries_person_casualty_losses() -> None: + donor = puf_tax_unit_donor_from_arrays( + { + "tax_unit_id": [10, 20], + "household_weight": [100.0, 200.0], + "filing_status": [b"SINGLE", b"JOINT"], + "person_tax_unit_id": [10, 10, 20], + "casualty_loss": [1_000.0, 2_000.0, 4_000.0], + }, + person_outputs=("casualty_loss",), + tax_unit_outputs=(), + ) + + assert donor["casualty_loss"].tolist() == [3_000.0, 4_000.0] + + +def test_puf_tax_unit_donor_carries_both_person_alimony_leaves() -> None: + donor = puf_tax_unit_donor_from_arrays( + { + "tax_unit_id": [10, 20], + "household_weight": [100.0, 200.0], + "filing_status": [b"SINGLE", b"JOINT"], + "person_tax_unit_id": [10, 10, 20], + "alimony_income": [1_000.0, 2_000.0, 4_000.0], + "alimony_expense": [500.0, 700.0, 3_000.0], + }, + person_outputs=("alimony_income", "alimony_expense"), + tax_unit_outputs=(), + ) + + assert donor["alimony_income"].tolist() == [3_000.0, 4_000.0] + assert donor["alimony_expense"].tolist() == [1_200.0, 3_000.0] + + +def test_puf_tax_unit_donor_carries_person_misc_itemized_expenses() -> None: + output = "unreimbursed_business_employee_expenses" + donor = puf_tax_unit_donor_from_arrays( + { + "tax_unit_id": [10, 20], + "household_weight": [100.0, 200.0], + "filing_status": [b"SINGLE", b"JOINT"], + "person_tax_unit_id": [10, 10, 20], + output: [1_000.0, 2_000.0, 4_000.0], + }, + person_outputs=(output,), + tax_unit_outputs=(), + ) + + assert donor[output].tolist() == [3_000.0, 4_000.0] + + def test_puf_tax_unit_donor_carries_ald_contribution_leaves() -> None: donor = puf_tax_unit_donor_from_arrays( { @@ -597,8 +693,10 @@ def test_cps_carried_derivations_create_leaf_inputs_not_aggregates() -> None: assert person["taxable_private_pension_income"].tolist() == [649.0, 0.0, 0.0] assert person["taxable_ira_distributions"].tolist() == [500.0, 0.0, 0.0] assert person["rental_income"].tolist() == [20.0, 0.0, 0.0] - assert person["farm_income"].tolist() == [5.0, 0.0, 0.0] + assert person["farm_operations_income"].tolist() == [5.0, 0.0, 0.0] + assert "farm_income" not in person assert person["unemployment_compensation"].tolist() == [0.0, 70.0, 0.0] + assert person["alimony_income"].tolist() == [0.0, 0.0, 0.0] assert person["miscellaneous_income"].tolist() == [3.0, 0.0, 0.0] assert person["health_insurance_premiums_without_medicare_part_b"].tolist() == [ 400.0, @@ -801,6 +899,97 @@ def predict(self, features: pd.DataFrame) -> pd.DataFrame: assert puf_people["qualified_dividend_income"].tolist() == [0.25, 0.0, 9.6] +def test_puf_tax_detail_preserves_sparse_educator_rate_and_earnings_allocation( + monkeypatch: pytest.MonkeyPatch, +) -> None: + class EducatorExpenseQRF: + def __init__(self, *, n_estimators: int, seed: int) -> None: + pass + + def fit( + self, + frame, + predictors, + outputs, + *, + weights, + ) -> "EducatorExpenseQRF": + assert outputs == ["educator_expense"] + assert weights == "design" + return self + + def predict(self, features: pd.DataFrame) -> pd.DataFrame: + return pd.DataFrame( + {"educator_expense": [300.0, 200.0]}, + index=features.index, + ) + + monkeypatch.setattr(puf_support_module, "QRF", EducatorExpenseQRF) + expanded = clone_us_frame_for_puf_support(_minimal_us_frame()) + donor = pd.DataFrame( + { + "puf_predictor_filing_status_code": [1.0, 2.0, 4.0], + "puf_predictor_tax_unit_person_count": [1.0, 2.0, 1.0], + "educator_expense": [300.0, 200.0, 0.0], + # Weighted donor positive rate = (0.5 + 0.5) / 4 = 25%. + "weight": [0.5, 0.5, 3.0], + } + ) + + imputed = impute_us_puf_tax_detail_support( + expanded, + donor, + predictors=( + "puf_predictor_filing_status_code", + "puf_predictor_tax_unit_person_count", + ), + person_outputs=("educator_expense",), + tax_unit_outputs=(), + n_estimators=4, + seed=0, + ) + + person = imputed.table("person") + asec_people = person[ + person[support_channel_column("person")] == BASE_ASEC_SUPPORT_CHANNEL + ] + puf_people = person[ + person[support_channel_column("person")] == PUF_TAX_DETAIL_SUPPORT_CHANNEL + ].sort_values(support_source_id_column("person")) + assert asec_people["educator_expense"].tolist() == [0.0, 0.0, 0.0] + + puf_totals = puf_people.groupby("person_tax_unit_id", sort=False)[ + "educator_expense" + ].sum() + np.testing.assert_allclose(puf_totals.to_numpy(), [900.0, 0.0]) + np.testing.assert_allclose( + puf_people.iloc[:2]["educator_expense"].to_numpy(), + [900.0 * 50_000.0 / 70_000.0, 900.0 * 20_000.0 / 70_000.0], + ) + + household = imputed.table("household") + puf_household_mask = ( + household[support_channel_column("household")] == PUF_TAX_DETAIL_SUPPORT_CHANNEL + ) + puf_household_weights = pd.Series( + imputed.weights_for("household").values[puf_household_mask], + index=household.loc[puf_household_mask, "household_id"], + ) + tax_unit_household = puf_people.groupby("person_tax_unit_id", sort=False)[ + "person_household_id" + ].first() + tax_unit_weights = tax_unit_household.map(puf_household_weights) + puf_positive_rate = float( + tax_unit_weights[puf_totals > 0.0].sum() / tax_unit_weights.sum() + ) + donor_positive_rate = float( + donor.loc[donor["educator_expense"] > 0.0, "weight"].sum() + / donor["weight"].sum() + ) + assert donor_positive_rate == pytest.approx(0.25) + assert puf_positive_rate == pytest.approx(donor_positive_rate) + + def test_puf_tax_unit_donor_derives_partnership_and_s_corp_split_from_raw_fields() -> ( None ): @@ -855,12 +1044,16 @@ def test_puf_tax_unit_donor_carries_structural_mortgage_leaves() -> None: "first_home_mortgage_origination_year": [2018, 2016], "second_home_mortgage_origination_year": [0, 2020], "health_savings_account_ald": [1_500.0, 0.0], + "domestic_production_ald": [7_500.0, 0.0], + "unrecaptured_section_1250_gain": [500.0, 0.0], }, person_outputs=(), tax_unit_outputs=PUF_TAX_DETAIL_DEFAULT_TAX_UNIT_OUTPUTS, ) assert donor["health_savings_account_ald"].tolist() == [1_500.0, 0.0] + assert donor["domestic_production_ald"].tolist() == [7_500.0, 0.0] + assert donor["unrecaptured_section_1250_gain"].tolist() == [500.0, 0.0] assert donor["first_home_mortgage_balance"].tolist() == [250_000.0, 500_000.0] assert donor["second_home_mortgage_balance"].tolist() == [0.0, 125_000.0] assert donor["first_home_mortgage_interest"].tolist() == [10_000.0, 20_000.0] diff --git a/packages/populace-build/tests/test_us_puf_support_base_builder.py b/packages/populace-build/tests/test_us_puf_support_base_builder.py index 4f6b9933..e9331acc 100644 --- a/packages/populace-build/tests/test_us_puf_support_base_builder.py +++ b/packages/populace-build/tests/test_us_puf_support_base_builder.py @@ -196,6 +196,26 @@ def test_pooled_asec_mode_rejects_base_h5_at_parse_time() -> None: assert exc.value.code == 2 +def test_weeks_unemployed_source_override_parses() -> None: + builder = _load_support_builder_module() + + args = builder._parse_args( + [ + "--base-h5", + "base.h5", + "--puf-h5", + "puf.h5", + "--asec-2023-weeks-unemployed-source", + "asecpub23csv.zip", + "--out", + "out", + "--without-block-ladder", + ] + ) + + assert args.asec_2023_weeks_unemployed_source == Path("asecpub23csv.zip") + + def test_pooled_asec_mode_loads_sources_with_manifest_metadata( monkeypatch: pytest.MonkeyPatch, tmp_path: Path, @@ -454,3 +474,651 @@ def test_block_ladder_and_opt_out_are_contradictory() -> None: ) assert exc.value.code == 2 + + +@pytest.mark.parametrize( + ("failing_gate", "failure_message"), + [ + ("child_support", "PUF child-support channel is default-only"), + ("disability_benefits", "PUF disability-benefits channel is default-only"), + ("wic_claim", "PUF WIC-claim channel is default-only"), + ("educator_expense", "PUF educator-expense channel is default-only"), + ("form_4952", "PUF Form 4952 channel is default-only"), + ("salt_refund", "PUF SALT-refund channel is default-only"), + ("energy_subsidy", "PUF energy-subsidy channel is default-only"), + ("weeks_unemployed", "PUF weeks-unemployed channel is default-only"), + ( + "retirement_distributions", + "PUF retirement-distribution channel is default-only", + ), + ], +) +def test_main_runs_cps_only_inputs_before_clone_and_after_puf_then_fails_gate( + monkeypatch: pytest.MonkeyPatch, + tmp_path: Path, + failing_gate: str, + failure_message: str, +) -> None: + builder = _load_support_builder_module() + child_support_calls: list[tuple[object, int, int]] = [] + disability_benefits_calls: list[tuple[object, int, int]] = [] + weeks_unemployed_calls: list[tuple[object, int, int, object]] = [] + weeks_unemployed_gate_frames: list[object] = [] + weeks_unemployed_source_loads: list[Path] = [] + educator_expense_gate_frames: list[object] = [] + form_4952_gate_frames: list[object] = [] + salt_refund_gate_frames: list[object] = [] + energy_subsidy_gate_frames: list[object] = [] + retirement_distribution_calls: list[tuple[object, int, int]] = [] + retirement_distribution_gate_frames: list[object] = [] + prior_year_income_calls: list[tuple[object, int, int]] = [] + prior_year_income_gate_frames: list[object] = [] + prior_year_income_reconciliation_frames: list[object] = [] + + monkeypatch.setattr( + builder, + "_parse_args", + lambda: type( + "Args", + (), + { + "out": tmp_path, + "target_year": 2024, + "seed": 7, + "n_estimators": 4, + "puf_h5": tmp_path / "puf.h5", + "acs_h5": tmp_path / "acs.h5", + "asec_2023_weeks_unemployed_source": (tmp_path / "asecpub23csv.zip"), + }, + )(), + ) + monkeypatch.setattr( + builder, + "_load_base_frame_from_args", + lambda args: ("raw", {"kind": "fixture"}), + ) + monkeypatch.setattr(builder, "derive_us_cps_carried_inputs", lambda frame: "cps") + + def fake_prior_year_income(frame, *, seed, time_period): + prior_year_income_calls.append((frame, seed, time_period)) + return "prior-year-direct" if frame == "cps" else "prior-year-puf" + + monkeypatch.setattr( + builder, + "with_us_prior_year_income_inputs", + fake_prior_year_income, + ) + monkeypatch.setattr( + builder, + "with_us_relationship_inputs", + lambda frame, *, seed, time_period: "relationship-inputs", + ) + monkeypatch.setattr( + builder, + "us_relationship_inputs_signal_gate", + lambda frame: type( + "Gate", (), {"passed": True, "failures": (), "details": {}} + )(), + ) + monkeypatch.setattr( + builder, + "with_us_medicare_take_up_input", + lambda frame, *, seed, time_period: frame, + ) + monkeypatch.setattr( + builder, + "us_medicare_take_up_signal_gate", + lambda frame: type( + "Gate", (), {"passed": True, "failures": (), "details": {}} + )(), + ) + housing_gate_frames: list[object] = [] + + def fake_housing_inputs_signal_gate(frame): + housing_gate_frames.append(frame) + return type( + "Gate", + (), + { + "passed": frame != "relationship-inputs", + "failures": ( + ("housing inputs absent",) if frame == "relationship-inputs" else () + ), + "details": {}, + }, + )() + + monkeypatch.setattr( + builder, + "us_housing_inputs_signal_gate", + fake_housing_inputs_signal_gate, + ) + monkeypatch.setattr( + builder, + "load_acs_2022_rent_donor", + lambda path: "acs-rent-donor", + ) + monkeypatch.setattr( + builder, + "with_us_housing_inputs", + lambda frame, *, seed, time_period, acs_rent_donor: "housing-direct", + ) + monkeypatch.setattr( + builder, + "with_us_eligibility_inputs", + lambda frame, *, seed, time_period: frame, + ) + monkeypatch.setattr( + builder, + "with_us_pregnancy_inputs", + lambda frame, *, seed, time_period: frame, + ) + monkeypatch.setattr( + builder, + "with_us_wic_claim_input", + lambda frame, *, seed, time_period: frame, + ) + for gate_name in ( + "us_eligibility_inputs_signal_gate", + "us_pregnancy_signal_gate", + ): + monkeypatch.setattr( + builder, + gate_name, + lambda frame: type( + "Gate", (), {"passed": True, "failures": (), "details": {}} + )(), + ) + monkeypatch.setattr( + builder, + "us_wic_claim_signal_gate", + lambda frame: type( + "Gate", + (), + { + "passed": not ( + failing_gate == "wic_claim" and frame == "qbi-reconciled" + ), + "failures": ( + ("PUF WIC-claim channel is default-only",) + if failing_gate == "wic_claim" and frame == "qbi-reconciled" + else () + ), + "details": {}, + }, + )(), + ) + + def fake_child_support(frame, *, seed, time_period): + child_support_calls.append((frame, seed, time_period)) + return ( + "child-support-direct" if frame == "housing-direct" else "child-support-puf" + ) + + monkeypatch.setattr(builder, "with_us_child_support_inputs", fake_child_support) + + def fake_disability_benefits(frame, *, seed, time_period): + disability_benefits_calls.append((frame, seed, time_period)) + if frame == "child-support-direct": + return "disability-benefits-direct" + return "disability-benefits-puf" + + monkeypatch.setattr( + builder, + "with_us_disability_benefits", + fake_disability_benefits, + ) + + def fake_load_weeks_unemployed_source(path, **_kwargs): + weeks_unemployed_source_loads.append(Path(path)) + return "weeks-source" + + def fake_with_weeks_unemployed( + frame, + *, + seed, + time_period, + asec_2023_source, + ): + weeks_unemployed_calls.append((frame, seed, time_period, asec_2023_source)) + return frame + + monkeypatch.setattr( + builder, + "load_asec_2023_weeks_unemployed_source", + fake_load_weeks_unemployed_source, + ) + monkeypatch.setattr( + builder, + "with_us_weeks_unemployed", + fake_with_weeks_unemployed, + ) + monkeypatch.setattr( + builder, + "with_us_workers_compensation", + lambda frame, *, seed, time_period: frame, + ) + monkeypatch.setattr( + builder, + "with_us_childcare_inputs", + lambda frame, *, seed, time_period: frame, + ) + monkeypatch.setattr( + builder, + "with_us_energy_subsidy_input", + lambda frame, *, seed, time_period: frame, + ) + monkeypatch.setattr( + builder, + "with_us_retirement_contribution_inputs", + lambda frame, *, seed, time_period: frame, + ) + + def fake_retirement_distributions(frame, *, seed, time_period): + retirement_distribution_calls.append((frame, seed, time_period)) + if frame == "disability-benefits-direct": + return "retirement-distributions-direct" + return "retirement-distributions-puf" + + monkeypatch.setattr( + builder, + "with_us_retirement_distribution_inputs", + fake_retirement_distributions, + ) + monkeypatch.setattr( + builder, + "with_us_immigration_inputs", + lambda frame, *, seed, time_period: frame, + ) + monkeypatch.setattr( + builder, + "clone_us_frame_for_puf_support", + lambda frame: "expanded", + ) + monkeypatch.setattr(builder, "_read_h5_arrays", lambda path: {}) + monkeypatch.setattr(builder, "puf_tax_unit_donor_from_arrays", lambda arrays: None) + monkeypatch.setattr( + builder, + "impute_and_audit_us_puf_support", + lambda expanded, donor, **kwargs: ("puf-imputed", {"passed": True}), + ) + monkeypatch.setattr( + builder, + "with_us_qbi_input_reconciliation", + lambda frame: "qbi-reconciled", + ) + monkeypatch.setattr( + builder, + "impute_us_housing_assistance_to_puf_support", + lambda frame, *, seed: "housing-puf", + ) + passing_gate = type( + "Gate", + (), + {"passed": True, "failures": (), "details": {}}, + )() + + def fake_prior_year_income_gate(frame): + prior_year_income_gate_frames.append(frame) + return passing_gate + + monkeypatch.setattr( + builder, + "us_prior_year_income_signal_gate", + fake_prior_year_income_gate, + ) + + def fake_prior_year_income_reconciliation_gate(frame): + prior_year_income_reconciliation_frames.append(frame) + return passing_gate + + monkeypatch.setattr( + builder, + "us_prior_year_income_source_reconciliation_gate", + fake_prior_year_income_reconciliation_gate, + ) + monkeypatch.setattr( + builder, "us_qbi_inputs_signal_gate", lambda frame: passing_gate + ) + monkeypatch.setattr( + builder, + "us_farm_business_income_signal_gate", + lambda frame: passing_gate, + ) + monkeypatch.setattr( + builder, + "us_domestic_production_ald_signal_gate", + lambda frame: passing_gate, + ) + monkeypatch.setattr( + builder, + "us_child_support_signal_gate", + lambda frame: type( + "Gate", + (), + { + "passed": failing_gate != "child_support", + "failures": ( + ("PUF child-support channel is default-only",) + if failing_gate == "child_support" + else () + ), + "details": {}, + }, + )(), + ) + monkeypatch.setattr( + builder, + "us_disability_benefits_signal_gate", + lambda frame: type( + "Gate", + (), + { + "passed": failing_gate != "disability_benefits", + "failures": ( + ("PUF disability-benefits channel is default-only",) + if failing_gate == "disability_benefits" + else () + ), + "details": {}, + }, + )(), + ) + monkeypatch.setattr( + builder, + "us_workers_compensation_signal_gate", + lambda frame: type( + "Gate", (), {"passed": True, "failures": (), "details": {}} + )(), + ) + + def fake_weeks_unemployed_signal_gate(frame): + weeks_unemployed_gate_frames.append(frame) + return type( + "Gate", + (), + { + "passed": failing_gate != "weeks_unemployed", + "failures": ( + ("PUF weeks-unemployed channel is default-only",) + if failing_gate == "weeks_unemployed" + else () + ), + "details": {}, + }, + )() + + monkeypatch.setattr( + builder, + "us_weeks_unemployed_signal_gate", + fake_weeks_unemployed_signal_gate, + ) + + def fake_educator_expense_signal_gate(frame): + educator_expense_gate_frames.append(frame) + return type( + "Gate", + (), + { + "passed": failing_gate != "educator_expense", + "failures": ( + ("PUF educator-expense channel is default-only",) + if failing_gate == "educator_expense" + else () + ), + "details": {}, + }, + )() + + monkeypatch.setattr( + builder, + "us_educator_expense_signal_gate", + fake_educator_expense_signal_gate, + ) + + def fake_form_4952_election_signal_gate(frame): + form_4952_gate_frames.append(frame) + return type( + "Gate", + (), + { + "passed": failing_gate != "form_4952", + "failures": ( + ("PUF Form 4952 channel is default-only",) + if failing_gate == "form_4952" + else () + ), + "details": {}, + }, + )() + + monkeypatch.setattr( + builder, + "us_form_4952_election_signal_gate", + fake_form_4952_election_signal_gate, + ) + + def fake_salt_refund_income_signal_gate(frame): + salt_refund_gate_frames.append(frame) + return type( + "Gate", + (), + { + "passed": failing_gate != "salt_refund", + "failures": ( + ("PUF SALT-refund channel is default-only",) + if failing_gate == "salt_refund" + else () + ), + "details": {}, + }, + )() + + monkeypatch.setattr( + builder, + "us_salt_refund_income_signal_gate", + fake_salt_refund_income_signal_gate, + ) + monkeypatch.setattr( + builder, + "us_capital_gain_details_signal_gate", + lambda frame: passing_gate, + ) + monkeypatch.setattr( + builder, + "us_childcare_signal_gate", + lambda frame: passing_gate, + ) + + def fake_energy_subsidy_signal_gate(frame): + energy_subsidy_gate_frames.append(frame) + return type( + "Gate", + (), + { + "passed": failing_gate != "energy_subsidy", + "failures": ( + ("PUF energy-subsidy channel is default-only",) + if failing_gate == "energy_subsidy" + else () + ), + "details": {}, + }, + )() + + monkeypatch.setattr( + builder, + "us_energy_subsidy_signal_gate", + fake_energy_subsidy_signal_gate, + ) + monkeypatch.setattr(builder, "us_alimony_signal_gate", lambda frame: passing_gate) + monkeypatch.setattr( + builder, + "us_casualty_loss_signal_gate", + lambda frame: passing_gate, + ) + monkeypatch.setattr( + builder, + "us_misc_itemized_signal_gate", + lambda frame: passing_gate, + ) + monkeypatch.setattr( + builder, + "us_retirement_contributions_signal_gate", + lambda frame: passing_gate, + ) + + def fake_retirement_distributions_signal_gate(frame): + retirement_distribution_gate_frames.append(frame) + return type( + "Gate", + (), + { + "passed": failing_gate != "retirement_distributions", + "failures": ( + ("PUF retirement-distribution channel is default-only",) + if failing_gate == "retirement_distributions" + else () + ), + "details": {}, + }, + )() + + monkeypatch.setattr( + builder, + "us_retirement_distributions_signal_gate", + fake_retirement_distributions_signal_gate, + ) + + with pytest.raises(SystemExit, match=failure_message): + builder.main() + + if failing_gate == "wic_claim": + assert child_support_calls == [("housing-direct", 7, 2024)] + assert prior_year_income_calls == [("cps", 7, 2024)] + assert prior_year_income_gate_frames == [] + assert prior_year_income_reconciliation_frames == [] + assert housing_gate_frames == ["relationship-inputs", "housing-direct"] + else: + assert child_support_calls == [ + ("housing-direct", 7, 2024), + ("prior-year-puf", 7, 2024), + ] + assert prior_year_income_calls == [ + ("cps", 7, 2024), + ("housing-puf", 7, 2024), + ] + assert prior_year_income_gate_frames == ["prior-year-puf"] + assert prior_year_income_reconciliation_frames == ["prior-year-puf"] + assert housing_gate_frames == [ + "relationship-inputs", + "housing-direct", + "prior-year-puf", + ] + expected_disability_calls = [("child-support-direct", 7, 2024)] + if failing_gate not in {"child_support", "wic_claim"}: + expected_disability_calls.append(("child-support-puf", 7, 2024)) + assert disability_benefits_calls == expected_disability_calls + expected_weeks_calls = [("disability-benefits-direct", 7, 2024, "weeks-source")] + if failing_gate not in {"wic_claim", "child_support", "disability_benefits"}: + expected_weeks_calls.append( + ("disability-benefits-puf", 7, 2024, "weeks-source") + ) + assert weeks_unemployed_calls == expected_weeks_calls + assert weeks_unemployed_source_loads == [tmp_path / "asecpub23csv.zip"] + assert weeks_unemployed_gate_frames == ( + ["disability-benefits-puf"] + if failing_gate not in {"wic_claim", "child_support", "disability_benefits"} + else [] + ) + assert educator_expense_gate_frames == ( + ["disability-benefits-puf"] + if failing_gate + in { + "educator_expense", + "form_4952", + "salt_refund", + "energy_subsidy", + "retirement_distributions", + } + else [] + ) + assert form_4952_gate_frames == ( + ["disability-benefits-puf"] + if failing_gate + in {"form_4952", "salt_refund", "energy_subsidy", "retirement_distributions"} + else [] + ) + assert salt_refund_gate_frames == ( + ["disability-benefits-puf"] + if failing_gate in {"salt_refund", "energy_subsidy", "retirement_distributions"} + else [] + ) + assert energy_subsidy_gate_frames == ( + ["disability-benefits-puf"] + if failing_gate in {"energy_subsidy", "retirement_distributions"} + else [] + ) + expected_retirement_distribution_calls = [("disability-benefits-direct", 7, 2024)] + if failing_gate == "retirement_distributions": + expected_retirement_distribution_calls.append( + ("disability-benefits-puf", 7, 2024) + ) + assert retirement_distribution_calls == expected_retirement_distribution_calls + assert retirement_distribution_gate_frames == ( + ["retirement-distributions-puf"] + if failing_gate == "retirement_distributions" + else [] + ) + + +def test_main_summary_records_retirement_distribution_gate() -> None: + builder = _load_support_builder_module() + source = Path(builder.__file__).read_text(encoding="utf-8") + + assert '"retirement_distributions_signal": {' in source + assert '"passed": retirement_distributions_gate.passed' in source + assert '"failures": list(retirement_distributions_gate.failures)' in source + assert '"details": dict(retirement_distributions_gate.details)' in source + + +def test_main_summary_records_weeks_unemployed_gate() -> None: + builder = _load_support_builder_module() + source = Path(builder.__file__).read_text(encoding="utf-8") + + assert '"weeks_unemployed_signal": {' in source + assert '"passed": weeks_unemployed_gate.passed' in source + assert '"failures": list(weeks_unemployed_gate.failures)' in source + assert '"details": dict(weeks_unemployed_gate.details)' in source + + +def test_main_summary_records_salt_refund_income_gate() -> None: + builder = _load_support_builder_module() + source = Path(builder.__file__).read_text(encoding="utf-8") + + assert '"salt_refund_income_signal": {' in source + assert '"passed": salt_refund_income_gate.passed' in source + assert '"failures": list(salt_refund_income_gate.failures)' in source + assert '"details": dict(salt_refund_income_gate.details)' in source + + +def test_main_summary_records_energy_subsidy_gate() -> None: + builder = _load_support_builder_module() + source = Path(builder.__file__).read_text(encoding="utf-8") + + assert '"energy_subsidy_signal": {' in source + assert '"passed": energy_subsidy_gate.passed' in source + assert '"failures": list(energy_subsidy_gate.failures)' in source + assert '"details": dict(energy_subsidy_gate.details)' in source + + +def test_main_summary_records_prior_year_income_gate() -> None: + builder = _load_support_builder_module() + source = Path(builder.__file__).read_text(encoding="utf-8") + + assert '"prior_year_income_signal": {' in source + assert '"passed": prior_year_income_gate.passed' in source + assert '"failures": list(prior_year_income_gate.failures)' in source + assert '"details": dict(prior_year_income_gate.details)' in source + assert '"prior_year_income_source_reconciliation": {' in source + assert '"passed": prior_year_income_reconciliation_gate.passed' in source diff --git a/packages/populace-build/tests/test_us_qbi_inputs.py b/packages/populace-build/tests/test_us_qbi_inputs.py new file mode 100644 index 00000000..314c9659 --- /dev/null +++ b/packages/populace-build/tests/test_us_qbi_inputs.py @@ -0,0 +1,329 @@ +"""Contracts for the retired eCPS Section 199A input family.""" + +from __future__ import annotations + +import importlib.util + +import numpy as np +import pandas as pd +import pytest + +from populace.build.us_runtime import puf_support as puf_support_module +from populace.build.us_runtime.qbi_inputs import ( + QBI_ARCHIVED_ASSUMPTIONS_URL, + QBI_ARCHIVED_CLONE_URL, + QBI_ARCHIVED_DERIVATION_URL, + QBI_ARCHIVED_EXPORT_URL, + QBI_ARCHIVED_IMPUTATION_URL, + QBI_ARCHIVED_PUF_ARTIFACT_URL, + QBI_ARCHIVED_SIMULATION_URL, + US_QBI_BOOLEAN_OUTPUT_COLUMNS, + US_QBI_NONNEGATIVE_OUTPUT_COLUMNS, + US_QBI_OUTPUT_COLUMNS, + us_qbi_inputs_signal_gate, + us_qbi_inputs_stage_spec, + us_qbi_inputs_summary, + with_us_qbi_input_reconciliation, +) +from populace.frame import US_SCHEMA, Frame, WeightKind, Weights +from populace.frame.schema import EntitySchema + +policyengine_us_installed = importlib.util.find_spec("policyengine_us") is not None +requires_us = pytest.mark.skipif( + not policyengine_us_installed, + reason="requires the policyengine-us [us] extra (build environment)", +) + + +def _frame(person: pd.DataFrame) -> Frame: + person = person.copy(deep=True).reset_index(drop=True) + n = len(person) + ids = np.arange(1, n + 1, dtype=np.int64) + person.insert(0, "person_id", ids) + for entity in ("household", "tax_unit", "spm_unit", "family", "marital_unit"): + person[f"person_{entity}_id"] = ids + tables = { + "person": person, + "household": pd.DataFrame({"household_id": ids}), + "tax_unit": pd.DataFrame({"tax_unit_id": ids}), + "spm_unit": pd.DataFrame({"spm_unit_id": ids}), + "family": pd.DataFrame({"family_id": ids}), + "marital_unit": pd.DataFrame({"marital_unit_id": ids}), + } + return Frame( + tables, + US_SCHEMA, + {"household": Weights(np.ones(n), WeightKind.DESIGN)}, + ) + + +def _qbi_person(n: int = 200) -> pd.DataFrame: + person = pd.DataFrame( + { + column: np.zeros(n, dtype=bool) + if column in US_QBI_BOOLEAN_OUTPUT_COLUMNS + else np.zeros(n, dtype=np.float64) + for column in US_QBI_OUTPUT_COLUMNS + } + ) + person["person_support_channel"] = np.where( + np.arange(n) < n // 2, "asec", "puf_tax_detail" + ) + person["self_employment_income_before_lsr"] = 0.0 + person["partnership_income"] = 0.0 + person["s_corp_income"] = 0.0 + person["estate_income"] = 0.0 + person["rental_income"] = 0.0 + person["non_qualified_dividend_income"] = 0.0 + + puf = np.arange(n // 2, n) + for index, column in enumerate(US_QBI_BOOLEAN_OUTPUT_COLUMNS): + if column == "business_is_sstb": + continue + person.loc[puf[index % 5 :: 5], column] = False + person.loc[puf, column] |= np.arange(len(puf)) % 5 != index % 5 + + sstb = puf[:10] + person.loc[sstb, "business_is_sstb"] = True + person.loc[sstb, "self_employment_income_would_be_qualified"] = True + person.loc[sstb, "self_employment_income_before_lsr"] = 10_000.0 + person.loc[sstb, "w2_wages_from_qualified_business"] = 2_000.0 + person.loc[sstb, "unadjusted_basis_qualified_property"] = 5_000.0 + + person.loc[puf[:2], "non_qualified_dividend_income"] = 1_000.0 + person.loc[puf[:2], "qualified_bdc_income"] = 20.0 + person.loc[puf[:20], "non_qualified_dividend_income"] = 1_000.0 + person.loc[puf[:20], "qualified_reit_and_ptp_income"] = 40.0 + person.loc[puf[:25], "w2_wages_from_qualified_business"] = 2_000.0 + person.loc[puf[:30], "unadjusted_basis_qualified_property"] = 5_000.0 + return person + + +def test_archived_coordinates_pin_algorithms_export_clone_and_artifact() -> None: + urls = ( + QBI_ARCHIVED_ASSUMPTIONS_URL, + QBI_ARCHIVED_CLONE_URL, + QBI_ARCHIVED_DERIVATION_URL, + QBI_ARCHIVED_EXPORT_URL, + QBI_ARCHIVED_IMPUTATION_URL, + QBI_ARCHIVED_PUF_ARTIFACT_URL, + QBI_ARCHIVED_SIMULATION_URL, + ) + for url in urls: + assert "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe" in url + assert "#L" in url + + +def test_output_contract_is_exact_and_source_declared() -> None: + expected = { + "business_is_sstb", + "estate_income_would_be_qualified", + "farm_operations_income_would_be_qualified", + "farm_rent_income_would_be_qualified", + "partnership_s_corp_income_would_be_qualified", + "rental_income_would_be_qualified", + "self_employment_income_would_be_qualified", + "sstb_self_employment_income_would_be_qualified", + "qualified_bdc_income", + "qualified_reit_and_ptp_income", + "sstb_self_employment_income_before_lsr", + "sstb_unadjusted_basis_qualified_property", + "sstb_w2_wages_from_qualified_business", + "unadjusted_basis_qualified_property", + "w2_wages_from_qualified_business", + } + assert set(US_QBI_OUTPUT_COLUMNS) == expected + assert set(US_QBI_BOOLEAN_OUTPUT_COLUMNS) == { + name for name in expected if name.endswith("_would_be_qualified") + } | {"business_is_sstb"} + assert set(US_QBI_NONNEGATIVE_OUTPUT_COLUMNS) == { + "qualified_bdc_income", + "qualified_reit_and_ptp_income", + "sstb_unadjusted_basis_qualified_property", + "sstb_w2_wages_from_qualified_business", + "unadjusted_basis_qualified_property", + "w2_wages_from_qualified_business", + } + + stage = us_qbi_inputs_stage_spec() + assert expected <= set(stage.outputs) + assert set(US_QBI_NONNEGATIVE_OUTPUT_COLUMNS) <= set(stage.nonnegative_outputs) + assert "release://policyengine/irs-soi-puf/1.8.0/puf_2024.h5" in str( + stage.artifacts + ) + + +def test_processed_puf_sstb_alias_is_carried_without_redrawing() -> None: + source = {"sstb_self_employment_income": np.array([-20.0, 0.0, 300.0])} + + values = puf_support_module._person_source_values( + source, "sstb_self_employment_income_before_lsr" + ) + + assert values.tolist() == [-20.0, 0.0, 300.0] + + +def test_boolean_tax_unit_counts_are_preserved_and_source_aligned() -> None: + person = pd.DataFrame( + { + "person_tax_unit_id": [1, 1, 1, 2, 2], + "self_employment_income_before_lsr": [0.0, 50.0, 100.0, 0.0, 25.0], + "business_is_sstb": 0.0, + } + ) + + puf_support_module._write_person_tax_unit_boolean_counts( + person, + mask=pd.Series(True, index=person.index), + column="business_is_sstb", + totals=pd.Series({1: 2.0, 2: 1.0}), + fallback_basis_columns=("self_employment_income_before_lsr",), + ) + + assert person.groupby("person_tax_unit_id")["business_is_sstb"].sum().to_dict() == { + 1: 2.0, + 2: 1.0, + } + assert person["business_is_sstb"].tolist() == [0.0, 1.0, 1.0, 0.0, 1.0] + + +def test_reconciliation_restores_sstb_routes_total_pools_and_exposure_caps() -> None: + person = _qbi_person(200) + original_total = ( + person["self_employment_income_before_lsr"] + + person["sstb_self_employment_income_before_lsr"] + ).to_numpy(copy=True) + person.loc[100, "qualified_bdc_income"] = 2_000.0 + person.loc[100, "qualified_reit_and_ptp_income"] = 4_000.0 + + result = with_us_qbi_input_reconciliation(_frame(person)) + reconciled = result.table("person") + + np.testing.assert_allclose( + reconciled["self_employment_income_before_lsr"] + + reconciled["sstb_self_employment_income_before_lsr"], + original_total, + ) + sstb = reconciled["business_is_sstb"].to_numpy() + assert np.all(reconciled.loc[sstb, "self_employment_income_before_lsr"] == 0) + assert np.all(reconciled.loc[~sstb, "sstb_self_employment_income_before_lsr"] == 0) + np.testing.assert_allclose( + reconciled["sstb_w2_wages_from_qualified_business"], + np.where(sstb, reconciled["w2_wages_from_qualified_business"], 0.0), + ) + np.testing.assert_allclose( + reconciled["sstb_unadjusted_basis_qualified_property"], + np.where(sstb, reconciled["unadjusted_basis_qualified_property"], 0.0), + ) + # The base W-2 and UBIA leaves remain total pools; they are not zeroed on + # all-SSTB rows. + assert reconciled.loc[sstb, "w2_wages_from_qualified_business"].sum() > 0 + assert reconciled.loc[sstb, "unadjusted_basis_qualified_property"].sum() > 0 + assert reconciled.loc[100, "qualified_bdc_income"] == 1_000.0 + assert reconciled.loc[100, "qualified_reit_and_ptp_income"] == 1_000.0 + + +def test_sstb_requires_a_positive_qualified_mapped_source() -> None: + person = _qbi_person(200) + row = 100 + person.loc[row, "business_is_sstb"] = True + person.loc[row, "self_employment_income_before_lsr"] = 500.0 + person.loc[row, "self_employment_income_would_be_qualified"] = False + person.loc[row, "sstb_self_employment_income_would_be_qualified"] = False + person.loc[row, "partnership_income"] = 0.0 + person.loc[row, "s_corp_income"] = 0.0 + person.loc[row, "estate_income"] = 0.0 + + result = with_us_qbi_input_reconciliation(_frame(person)).table("person") + + assert not result.loc[row, "business_is_sstb"] + assert result.loc[row, "self_employment_income_before_lsr"] == 500.0 + assert result.loc[row, "sstb_self_employment_income_before_lsr"] == 0.0 + + +def test_asec_channel_policy_is_explicit_and_reconciliation_is_idempotent() -> None: + first = with_us_qbi_input_reconciliation(_frame(_qbi_person())) + second = with_us_qbi_input_reconciliation(first) + asec = second.table("person").query("person_support_channel == 'asec'") + + assert not asec["business_is_sstb"].any() + assert not asec["sstb_self_employment_income_would_be_qualified"].any() + for column in US_QBI_BOOLEAN_OUTPUT_COLUMNS: + if column not in { + "business_is_sstb", + "sstb_self_employment_income_would_be_qualified", + }: + assert asec[column].all() + pd.testing.assert_frame_equal(first.table("person"), second.table("person")) + + +def test_signal_gate_passes_reconciled_nondefault_family() -> None: + reconciled = with_us_qbi_input_reconciliation(_frame(_qbi_person())) + + result = us_qbi_inputs_signal_gate(reconciled) + + assert result.passed, result.failures + assert all(value == 0 for value in result.details["invariants"].values()) + + +def test_signal_gate_reports_missing_negative_and_identity_failures() -> None: + person = _qbi_person() + person.loc[100, "qualified_bdc_income"] = -1.0 + person.loc[100, "business_is_sstb"] = True + person.loc[100, "self_employment_income_before_lsr"] = 100.0 + person.loc[100, "sstb_self_employment_income_before_lsr"] = 0.0 + frame = _frame(person) + + result = us_qbi_inputs_signal_gate(frame) + + assert not result.passed + assert any("qualified_bdc_income" in failure for failure in result.failures) + assert any( + "sstb_rows_with_non_sstb_income" in failure for failure in result.failures + ) + + missing = person.drop(columns=["qualified_reit_and_ptp_income"]) + missing_result = us_qbi_inputs_signal_gate(_frame(missing)) + assert not missing_result.passed + assert "qualified_reit_and_ptp_income" in missing_result.failures[0] + + +def test_reconciliation_rejects_wrong_schema_and_missing_inputs() -> None: + wrong = Frame( + { + "person": pd.DataFrame({"person_id": [1], "person_household_id": [1]}), + "household": pd.DataFrame({"household_id": [1]}), + }, + EntitySchema(group_entities=("household",)), + {"household": Weights(np.ones(1), WeightKind.DESIGN)}, + ) + with pytest.raises(ValueError, match="US schema"): + with_us_qbi_input_reconciliation(wrong) + + with pytest.raises(ValueError, match="requires PUF-imputed"): + with_us_qbi_input_reconciliation( + _frame(_qbi_person().drop(columns=["qualified_bdc_income"])) + ) + + +def test_summary_does_not_mutate_source_frame() -> None: + reconciled = with_us_qbi_input_reconciliation(_frame(_qbi_person())) + before = reconciled.table("person").copy(deep=True) + + summary = us_qbi_inputs_summary(reconciled) + + assert set(summary) == {"columns", "invariants"} + pd.testing.assert_frame_equal(before, reconciled.table("person")) + + +@requires_us +def test_all_qbi_outputs_are_live_policyengine_us_input_leaves() -> None: + from policyengine_us import CountryTaxBenefitSystem + + variables = CountryTaxBenefitSystem().variables + for column in US_QBI_OUTPUT_COLUMNS: + variable = variables[column] + assert variable.entity.key == "person" + assert not variable.formulas + assert getattr(variable, "adds", None) is None + assert getattr(variable, "subtracts", None) is None diff --git a/packages/populace-build/tests/test_us_relationship_inputs.py b/packages/populace-build/tests/test_us_relationship_inputs.py new file mode 100644 index 00000000..d2b50bb4 --- /dev/null +++ b/packages/populace-build/tests/test_us_relationship_inputs.py @@ -0,0 +1,438 @@ +"""ASEC relationship-input restoration and partner-source exclusion.""" + +from __future__ import annotations + +import importlib.util +import json +from hashlib import sha256 +from importlib.metadata import version +from importlib.resources import files +from pathlib import Path + +import numpy as np +import pandas as pd +import pytest + +from populace.build.source_manifest import SourceStageSpec +from populace.build.source_runtime import SourceRuntimeError +from populace.build.us_runtime import ( + US_DONORS, + US_PUF_SUPPORT_STAGE_NAME, + US_RELATIONSHIP_INPUTS_NONCONSTANT_PERSON_COLUMNS, + US_RELATIONSHIP_INPUTS_OUTPUT_COLUMNS, + US_RELATIONSHIP_INPUTS_REQUIRED_SOURCE_COLUMNS, + US_RELATIONSHIP_INPUTS_STAGE_NAME, + US_STAGE_NAMES, + derive_us_relationship_inputs_from_manifest, + load_release_input_coverage_manifest, + us_relationship_inputs_signal_gate, + us_relationship_inputs_stage_spec, + us_relationship_inputs_summary, + us_release_reform_coverage_probes, + with_us_relationship_inputs, +) +from populace.build.us_runtime.asec_pool import _with_relationship_recode +from populace.build.us_runtime.l0_refit_export import ( + US_RELEASE_REQUIRED_PERSON_SOURCE_COLUMNS, +) +from populace.build.us_runtime.release_input_coverage import ( + RESTORED_REFERENCE_ECPS_REQUIRED_INPUTS, +) +from populace.build.us_runtime.source_runtime import us_source_operation_handlers +from populace.frame import US_SCHEMA, Frame, WeightKind, Weights + +ROOT = Path(__file__).resolve().parents[3] +policyengine_us_installed = importlib.util.find_spec("policyengine_us") is not None +requires_us = pytest.mark.skipif( + not policyengine_us_installed, + reason="requires the policyengine-us [us] extra (build environment)", +) + +_HEAD = "is_household_head" +_SEPARATED = "is_separated" +_SURVIVING = "is_surviving_spouse" +_UNMARRIED_PARTNER = "is_unmarried_partner_of_household_head" +_OUTPUTS = (_HEAD, _SEPARATED, _SURVIVING) + + +def _person(rows: list[dict[str, object]]) -> pd.DataFrame: + records: list[dict[str, object]] = [] + for index, row in enumerate(rows): + record: dict[str, object] = { + "person_id": index + 1, + "PH_SEQ": 100 + index, + "P_SEQ": 1, + "A_MARITL": 7, + } + record.update(row) + records.append(record) + return pd.DataFrame(records) + + +def _frame(rows: list[dict[str, object]], weights: list[float] | None = None) -> Frame: + person = _person(rows) + household_ids, person_household_ids = np.unique( + person["PH_SEQ"].to_numpy(dtype=np.int64), return_inverse=True + ) + person_household_ids = household_ids[person_household_ids] + person["person_household_id"] = person_household_ids + person["person_tax_unit_id"] = person_household_ids + 1_000 + person["person_spm_unit_id"] = person_household_ids + 2_000 + person["person_family_id"] = person_household_ids + 3_000 + person["person_marital_unit_id"] = np.arange(len(person)) + 4_000 + tables = { + "person": person, + "household": pd.DataFrame({"household_id": household_ids}), + "tax_unit": pd.DataFrame({"tax_unit_id": household_ids + 1_000}), + "spm_unit": pd.DataFrame({"spm_unit_id": household_ids + 2_000}), + "family": pd.DataFrame({"family_id": household_ids + 3_000}), + "marital_unit": pd.DataFrame( + {"marital_unit_id": np.arange(len(person)) + 4_000} + ), + } + return Frame( + tables, + US_SCHEMA, + { + "household": Weights( + np.asarray(weights or [1.0] * len(household_ids), dtype=np.float64), + WeightKind.DESIGN, + ) + }, + ) + + +def _operation(): + return next( + operation + for operation in us_relationship_inputs_stage_spec().operations + if operation.kind == "derive_relationship_inputs" + ) + + +def _known_gap(name: str) -> dict[str, object]: + payload = json.loads( + files("populace.build.us") + .joinpath("ecps_parity_known_gaps.json") + .read_text(encoding="utf-8") + ) + return payload["known_gaps"][name] + + +def _sha256(path: Path) -> str: + digest = sha256() + with path.open("rb") as stream: + for chunk in iter(lambda: stream.read(1024 * 1024), b""): + digest.update(chunk) + return digest.hexdigest() + + +class TestManifestAndPlan: + def test_stage_pins_exact_archived_derivations(self) -> None: + spec = us_relationship_inputs_stage_spec() + + assert spec.stage == US_RELATIONSHIP_INPUTS_STAGE_NAME == "relationship_inputs" + assert US_RELATIONSHIP_INPUTS_OUTPUT_COLUMNS == _OUTPUTS + assert US_RELATIONSHIP_INPUTS_NONCONSTANT_PERSON_COLUMNS == _OUTPUTS + assert US_RELATIONSHIP_INPUTS_REQUIRED_SOURCE_COLUMNS == ( + "PH_SEQ", + "P_SEQ", + "A_MARITL", + ) + assert tuple(spec.outputs) == _OUTPUTS + assert [operation.kind for operation in spec.operations] == [ + "read_table", + "derive_relationship_inputs", + ] + assert "cps.py lines 1069-1075 and 1209-1221" in spec.notes + assert "P_SEQ == 1" in spec.notes + assert "A_MARITL == 6" in spec.notes + assert "A_MARITL == 4" in spec.notes + + def test_handler_and_plan_are_wired_before_puf_support(self) -> None: + handlers = us_source_operation_handlers() + assert ( + handlers["derive_relationship_inputs"] + is derive_us_relationship_inputs_from_manifest + ) + assert US_RELATIONSHIP_INPUTS_STAGE_NAME in US_DONORS + assert US_STAGE_NAMES.index(US_RELATIONSHIP_INPUTS_STAGE_NAME) < ( + US_STAGE_NAMES.index(US_PUF_SUPPORT_STAGE_NAME) + ) + + +class TestDerivation: + def test_maps_exact_asec_codes(self) -> None: + source = _person( + [ + {"PH_SEQ": 1, "P_SEQ": 1, "A_MARITL": 4}, + {"PH_SEQ": 1, "P_SEQ": 2, "A_MARITL": 6}, + {"PH_SEQ": 2, "P_SEQ": 1, "A_MARITL": 7}, + ] + ) + + result = derive_us_relationship_inputs_from_manifest(source, _operation(), None) + + assert result[_HEAD].tolist() == [True, False, True] + assert result[_SEPARATED].tolist() == [False, True, False] + assert result[_SURVIVING].tolist() == [True, False, False] + assert all(result[column].dtype == bool for column in _OUTPUTS) + + @pytest.mark.parametrize("missing", US_RELATIONSHIP_INPUTS_REQUIRED_SOURCE_COLUMNS) + def test_missing_source_is_named(self, missing: str) -> None: + source = _person([{}]).drop(columns=[missing]) + with pytest.raises(SourceRuntimeError, match=missing): + derive_us_relationship_inputs_from_manifest(source, _operation(), None) + + @pytest.mark.parametrize( + ("column", "value"), + [("PH_SEQ", np.nan), ("P_SEQ", 0), ("P_SEQ", 1.5), ("A_MARITL", 0)], + ) + def test_invalid_source_values_fail_closed(self, column: str, value: float) -> None: + source = _person([{}, {}]) + source[column] = source[column].astype(float) + source.loc[0, column] = value + with pytest.raises(SourceRuntimeError, match=column): + derive_us_relationship_inputs_from_manifest(source, _operation(), None) + + def test_requires_exactly_one_head_per_household(self) -> None: + source = _person( + [ + {"PH_SEQ": 1, "P_SEQ": 1}, + {"PH_SEQ": 1, "P_SEQ": 1}, + ] + ) + with pytest.raises(SourceRuntimeError, match="exactly one P_SEQ == 1"): + derive_us_relationship_inputs_from_manifest(source, _operation(), None) + + def test_support_clones_group_by_frame_household_not_raw_ph_seq(self) -> None: + source = _person( + [ + {"PH_SEQ": 1, "person_household_id": 10, "P_SEQ": 1}, + {"PH_SEQ": 1, "person_household_id": 10, "P_SEQ": 2}, + {"PH_SEQ": 1, "person_household_id": 20, "P_SEQ": 1}, + {"PH_SEQ": 1, "person_household_id": 20, "P_SEQ": 2}, + ] + ) + result = derive_us_relationship_inputs_from_manifest(source, _operation(), None) + assert result[_HEAD].tolist() == [True, False, True, False] + + def test_rejects_wrong_operation_and_missing_table(self) -> None: + with pytest.raises(SourceRuntimeError, match="unexpected operation"): + derive_us_relationship_inputs_from_manifest( + _person([{}]), + SourceStageSpec.from_mapping( + { + "stage": "test", + "survey": "test", + "source": "https://example.com", + "grain": "person", + "operations": [{"kind": "derive"}], + "outputs": list(_OUTPUTS), + } + ).operations[0], + None, + ) + with pytest.raises(SourceRuntimeError, match="person table"): + derive_us_relationship_inputs_from_manifest(None, _operation(), None) + + +class TestFrameAndGate: + def test_frame_integration_and_idempotence(self) -> None: + frame = _frame( + [ + {"PH_SEQ": 1, "P_SEQ": 1, "A_MARITL": 4}, + {"PH_SEQ": 1, "P_SEQ": 2, "A_MARITL": 6}, + {"PH_SEQ": 2, "P_SEQ": 1, "A_MARITL": 7}, + ], + weights=[2.0, 1.0], + ) + result = with_us_relationship_inputs(frame, seed=0, time_period=2024) + + assert result.table("person")[_HEAD].tolist() == [True, False, True] + assert with_us_relationship_inputs(result, seed=0, time_period=2024) is result + + def test_summary_and_gate_require_plausible_signal_and_one_head(self) -> None: + rows: list[dict[str, object]] = [] + for household in range(1, 51): + rows.extend( + [ + { + "PH_SEQ": household, + "P_SEQ": 1, + "A_MARITL": 4 if household <= 5 else 7, + }, + { + "PH_SEQ": household, + "P_SEQ": 2, + "A_MARITL": 6 if household <= 2 else 7, + }, + ] + ) + result = with_us_relationship_inputs(_frame(rows), seed=0, time_period=2024) + + summary = us_relationship_inputs_summary(result) + assert summary["household_head_share"] == pytest.approx(0.5) + assert summary["separated_share"] == pytest.approx(0.02) + assert summary["surviving_spouse_share"] == pytest.approx(0.05) + assert summary["households_without_exactly_one_head"] == 0 + assert us_relationship_inputs_signal_gate(result).passed + + result.table("person").loc[1, _HEAD] = True + gate = us_relationship_inputs_signal_gate(result) + assert not gate.passed + assert any("exactly one" in failure for failure in gate.failures) + + def test_gate_rejects_missing_output(self) -> None: + gate = us_relationship_inputs_signal_gate(_frame([{}])) + assert not gate.passed + assert _HEAD in gate.details["missing"] + + +class TestCoverageAndExclusion: + def test_restored_inputs_are_hard_release_requirements(self) -> None: + manifest = load_release_input_coverage_manifest() + assert set(_OUTPUTS) <= RESTORED_REFERENCE_ECPS_REQUIRED_INPUTS + assert set(_OUTPUTS) <= set(US_RELEASE_REQUIRED_PERSON_SOURCE_COLUMNS) + assert set(_OUTPUTS) <= manifest.required_columns + assert set(_OUTPUTS).isdisjoint(manifest.reviewed_exclusions) + + def test_unmarried_partner_exclusion_is_source_unavailability(self) -> None: + entry = _known_gap(_UNMARRIED_PARTNER) + evidence = entry["evidence"] + + assert entry["reason"].startswith("SOURCE UNAVAILABILITY WITH EVIDENCE:") + assert evidence["classification"] == "source_unavailability" + assert evidence["retired_derivation"] == { + "repository_owner": "PolicyEngine", + "repository_name_parts": ["policyengine-", "us-data"], + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": [ + "policyengine_", + "us_data", + "datasets", + "cps", + "cps.py", + ], + "lines": "1214-1221", + } + assert evidence["required_person_columns"] == ["PERRP"] + assert all( + "PERRP" in item["missing_person_columns"] + for item in evidence["hermetic_inputs"] + ) + assert evidence["semantic_non_substitutes"]["rejection"].startswith( + "Using only the 2024" + ) + + manifest = load_release_input_coverage_manifest() + assert manifest.reviewed_exclusions[_UNMARRIED_PARTNER] == entry["reason"] + + @requires_us + def test_locked_artifacts_confirm_source_presence_and_absence(self) -> None: + summary = json.loads( + (ROOT / "experiments/build_j_recert/base_j.summary.json").read_text() + ) + paths = { + Path(item["path"]).name: Path(item["path"]) + for item in summary["base_source"]["sources"] + } + if not all(path.is_file() for path in paths.values()): + pytest.skip("SHA-locked ASEC artifacts are not mounted") + + evidence = _known_gap(_UNMARRIED_PARTNER)["evidence"] + expected_heads = {2022: 56_839, 2023: 56_251, 2024: 55_762} + for item in evidence["hermetic_inputs"]: + path = paths[item["filename"]] + assert _sha256(path) == item["sha256"] + with pd.HDFStore(path, mode="r") as store: + person = store["person"] + assert set(item["missing_person_columns"]).isdisjoint(person.columns) + assert {"PH_SEQ", "P_SEQ", "A_MARITL"} <= set(person.columns) + year = int(item["filename"].split("_")[-1].split(".")[0]) + assert int((person["P_SEQ"] == 1).sum()) == expected_heads[year] + assert ( + person.groupby("PH_SEQ")["P_SEQ"].apply(lambda x: (x == 1).sum()) == 1 + ).all() + if year < 2024: + recoded, source = _with_relationship_recode(person) + assert source == "derived:line_spouse_parent" + assert int((recoded["A_EXPRRP"] == 13).sum()) == 0 + else: + assert ( + int((person["PECOHAB"] > 0).sum()) == item["positive_pecohab_rows"] + ) + assert ( + int((person["A_EXPRRP"] == 13).sum()) + == item["a_exprrp_partner_or_roommate_rows"] + ) + + +@requires_us +def test_policyengine_1_764_6_contract_and_live_bindings() -> None: + from policyengine_us import CountryTaxBenefitSystem, Simulation + + assert version("policyengine-us") == "1.764.6" + variables = CountryTaxBenefitSystem().variables + for name in _OUTPUTS: + variable = variables[name] + assert variable.is_input_variable() + assert variable.entity.key == "person" + assert variable.value_type is bool + assert variable.default_value is False + assert str(variables[_HEAD].definition_period).lower() == "eternity" + assert str(variables[_SEPARATED].definition_period).lower() == "year" + assert str(variables[_SURVIVING].definition_period).lower() == "year" + + common = { + "people": { + "adult": { + "age": {"2024": 40}, + "employment_income": {"2024": 50_000}, + "is_tax_unit_head": {"2024": True}, + _SURVIVING: {"2024": True}, + }, + "child": { + "age": {"2024": 5}, + "is_tax_unit_dependent": {"2024": True}, + }, + }, + "tax_units": {"tax_unit": {"members": ["adult", "child"]}}, + "spm_units": {"spm_unit": {"members": ["adult", "child"]}}, + "households": { + "household": { + "members": ["adult", "child"], + "state_code": {"2024": "CA"}, + } + }, + } + assert ( + Simulation(situation=common).calculate("filing_status", 2024).decode()[0].value + == "Surviving spouse" + ) + + separated = json.loads(json.dumps(common)) + separated["people"]["adult"].pop(_SURVIVING) + separated["people"]["adult"][_SEPARATED] = {"2024": True} + assert ( + Simulation(situation=separated) + .calculate("filing_status", 2024) + .decode()[0] + .value + == "Head of household" + ) + + +def test_shipped_household_head_neutralization_probe() -> None: + probe = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "household_head_childcare_cap_neutralization" + ) + assert probe.neutralized_variable == _HEAD + assert probe.binding_inputs == (_HEAD,) + assert probe.budget_measure == "spm_unit_capped_work_childcare_expenses" + assert probe.period == 2024 + assert probe.effect_direction == "baseline_minus_reform" + assert probe.expected_sign == "negative" + assert probe.min_abs_effect >= 1_000_000.0 diff --git a/packages/populace-build/tests/test_us_retirement_contributions.py b/packages/populace-build/tests/test_us_retirement_contributions.py new file mode 100644 index 00000000..e8c681f4 --- /dev/null +++ b/packages/populace-build/tests/test_us_retirement_contributions.py @@ -0,0 +1,244 @@ +"""ASEC retirement-contribution restoration and PUF-half QRF treatment.""" + +from __future__ import annotations + +import numpy as np +import pandas as pd +import pytest + +import populace.build.us_runtime.retirement_contributions as module +from populace.build.source_runtime import SourceRuntimeError +from populace.build.us_runtime.puf_support import clone_us_frame_for_puf_support +from populace.build.us_runtime.retirement_contributions import ( + US_RETIREMENT_CONTRIBUTION_OUTPUT_COLUMNS, + derive_us_retirement_contributions_from_manifest, + us_retirement_contributions_signal_gate, + us_retirement_contributions_stage_spec, + with_us_retirement_contribution_inputs, +) +from populace.frame import US_SCHEMA, Frame, WeightKind, Weights + + +def _person_source() -> pd.DataFrame: + return pd.DataFrame( + { + "person_id": np.asarray([1, 2, 3, 4], dtype="int64"), + "person_household_id": np.asarray([10, 20, 30, 40], dtype="int64"), + "person_tax_unit_id": np.asarray([100, 200, 300, 400], dtype="int64"), + "person_spm_unit_id": np.asarray([1_000, 2_000, 3_000, 4_000]), + "person_family_id": np.asarray([10_000, 20_000, 30_000, 40_000]), + "person_marital_unit_id": np.asarray([100_000, 200_000, 300_000, 400_000]), + # Wage-only, self-employment-only, both, and neither. + "RETCB_VAL": [1.0, 1.0, 1.0, 1.0], + "WSAL_VAL": [50_000.0, 0.0, 50_000.0, 0.0], + "SEMP_VAL": [0.0, 30_000.0, 30_000.0, 0.0], + "employment_income_before_lsr": [50_000.0, 0.0, 50_000.0, 0.0], + "self_employment_income_before_lsr": [0.0, 30_000.0, 30_000.0, 0.0], + "age": [35, 45, 55, 25], + "is_female": [False, True, False, True], + "has_esi": [True, False, True, False], + "tax_unit_role_input": ["HEAD", "HEAD", "HEAD", "HEAD"], + "social_security_retirement": [0.0, 0.0, 2_000.0, 0.0], + "social_security_disability": [0.0, 0.0, 0.0, 0.0], + "social_security_dependents": [0.0, 0.0, 0.0, 0.0], + "social_security_survivors": [0.0, 0.0, 0.0, 0.0], + } + ) + + +def _frame() -> Frame: + person = _person_source() + ids = { + "household": [10, 20, 30, 40], + "tax_unit": [100, 200, 300, 400], + "spm_unit": [1_000, 2_000, 3_000, 4_000], + "family": [10_000, 20_000, 30_000, 40_000], + "marital_unit": [100_000, 200_000, 300_000, 400_000], + } + tables = { + entity: pd.DataFrame({f"{entity}_id": np.asarray(values, dtype="int64")}) + for entity, values in ids.items() + } + tables["person"] = person + tables["tax_unit"]["filing_status_input"] = [ + "SINGLE", + "SINGLE", + "SINGLE", + "SINGLE", + ] + return Frame( + tables, + US_SCHEMA, + { + "household": Weights( + np.asarray([1.0, 1.0, 1.0, 1.0]), + WeightKind.DESIGN, + ) + }, + ) + + +def _derive(frame: pd.DataFrame) -> pd.DataFrame: + operation = next( + operation + for operation in us_retirement_contributions_stage_spec().operations + if operation.kind == "derive_retirement_contributions" + ) + return derive_us_retirement_contributions_from_manifest(frame, operation, None) + + +def test_stage_manifest_pins_sources_operations_and_five_desired_leaves() -> None: + spec = us_retirement_contributions_stage_spec() + + assert spec.stage == "retirement_contributions" + assert spec.grain == "person" + assert tuple(spec.outputs) == US_RETIREMENT_CONTRIBUTION_OUTPUT_COLUMNS + assert [operation.kind for operation in spec.operations] == [ + "read_table", + "derive_retirement_contributions", + "impute_retirement_contributions_to_puf_support", + ] + derive = spec.operations[1] + assert derive.parameters == { + "se_pension_share": 0.046, + "dc_share_of_remainder": 0.908, + "roth_dc_share": 0.15, + "traditional_ira_share": 0.392, + } + + +def test_direct_split_matches_archived_wage_and_self_employment_cases() -> None: + result = _derive(_person_source()) + + np.testing.assert_allclose( + result[list(US_RETIREMENT_CONTRIBUTION_OUTPUT_COLUMNS)].to_numpy(), + np.asarray( + [ + [0.7718, 0.1362, 0.036064, 0.055936, 0.0], + [0.0, 0.0, 0.373968, 0.580032, 0.046], + [0.7362972, 0.1299348, 0.034405056, 0.053362944, 0.046], + [0.0, 0.0, 0.0, 0.0, 0.0], + ] + ), + rtol=0, + atol=1e-12, + ) + + +def test_direct_split_fails_closed_without_measured_asec_source() -> None: + with pytest.raises(SourceRuntimeError, match="RETCB_VAL"): + _derive(_person_source().drop(columns=["RETCB_VAL"])) + + +@pytest.mark.parametrize( + ("bad_value", "message"), + [(np.nan, "nonfinite"), (-1.0, "negative")], +) +def test_direct_split_rejects_invalid_measured_totals( + bad_value: float, + message: str, +) -> None: + person = _person_source() + person.loc[0, "RETCB_VAL"] = bad_value + + with pytest.raises(SourceRuntimeError, match=message): + _derive(person) + + +def test_with_inputs_materializes_signal_and_preserves_reported_total() -> None: + result = with_us_retirement_contribution_inputs( + _frame(), + seed=0, + time_period=2024, + ) + + gate = us_retirement_contributions_signal_gate(result) + assert gate.passed, gate.failures + person = result.table("person") + allocated = person[list(US_RETIREMENT_CONTRIBUTION_OUTPUT_COLUMNS)].sum(axis=1) + np.testing.assert_allclose(allocated.iloc[:3], person["RETCB_VAL"].iloc[:3]) + assert allocated.iloc[3] == 0.0 + + +def test_puf_half_uses_qrf_predictions_and_applies_income_constraints( + monkeypatch: pytest.MonkeyPatch, +) -> None: + direct = with_us_retirement_contribution_inputs( + _frame(), + seed=0, + time_period=2024, + ) + expanded = clone_us_frame_for_puf_support(direct) + calls: dict[str, object] = {} + + class FakeFitted: + def predict(self, test: pd.DataFrame) -> pd.DataFrame: + calls["test"] = test.copy() + return pd.DataFrame( + { + column: np.full(len(test), index + 1.0) + for index, column in enumerate( + US_RETIREMENT_CONTRIBUTION_OUTPUT_COLUMNS + ) + }, + index=test.index, + ) + + class FakeQRF: + def __init__(self, **kwargs: object) -> None: + calls["init"] = kwargs + + def fit( + self, + training: pd.DataFrame, + predictors: list[str], + targets: list[str], + *, + weights: np.ndarray, + ) -> FakeFitted: + calls["training"] = training.copy() + calls["predictors"] = predictors + calls["targets"] = targets + calls["weights"] = weights.copy() + return FakeFitted() + + monkeypatch.setattr(module, "QRF", FakeQRF) + + result = with_us_retirement_contribution_inputs( + expanded, + seed=7, + time_period=2024, + ) + + assert calls["init"] == {"n_estimators": 100, "seed": 7} + assert len(calls["training"]) == 4 + assert len(calls["test"]) == 4 + person = result.table("person") + puf = person[person["person_support_channel"] == "puf_tax_detail"] + # Fake predictions are 1,2,3,4,5. The no-wage rows lose both 401(k) + # values, and the no-self-employment rows lose the SE-pension value. + np.testing.assert_allclose( + puf[list(US_RETIREMENT_CONTRIBUTION_OUTPUT_COLUMNS)].to_numpy(), + np.asarray( + [ + [1.0, 2.0, 3.0, 4.0, 0.0], + [0.0, 0.0, 3.0, 4.0, 5.0], + [1.0, 2.0, 3.0, 4.0, 5.0], + [0.0, 0.0, 3.0, 4.0, 0.0], + ] + ), + ) + + +def test_signal_gate_rejects_a_default_leaf() -> None: + result = with_us_retirement_contribution_inputs( + _frame(), + seed=0, + time_period=2024, + ) + result.table("person")["roth_ira_contributions_desired"] = 0.0 + + gate = us_retirement_contributions_signal_gate(result) + + assert not gate.passed + assert any("roth_ira_contributions_desired" in failure for failure in gate.failures) diff --git a/packages/populace-build/tests/test_us_retirement_distributions.py b/packages/populace-build/tests/test_us_retirement_distributions.py new file mode 100644 index 00000000..5690569a --- /dev/null +++ b/packages/populace-build/tests/test_us_retirement_distributions.py @@ -0,0 +1,456 @@ +"""ASEC retirement-distribution restoration and PUF-half QRF treatment.""" + +from __future__ import annotations + +import importlib.util +import json +from hashlib import sha256 +from importlib.metadata import version +from pathlib import Path + +import numpy as np +import pandas as pd +import pytest + +import populace.build.us_runtime.retirement_distributions as module +from populace.build.source_manifest import SourceStageSpec +from populace.build.source_runtime import SourceRuntimeError +from populace.build.us_runtime.asec_pool import load_asec_h5_tables +from populace.build.us_runtime.puf_support import clone_us_frame_for_puf_support +from populace.build.us_runtime.release_input_coverage import ( + us_release_reform_coverage_probes, +) +from populace.build.us_runtime.retirement_distributions import ( + RETIREMENT_DISTRIBUTIONS_ARCHIVED_DERIVATION_URL, + RETIREMENT_DISTRIBUTIONS_ARCHIVED_PARAMETERS_URL, + US_RETIREMENT_DISTRIBUTION_OUTPUT_COLUMNS, + US_RETIREMENT_DISTRIBUTION_REQUIRED_SOURCE_COLUMNS, + derive_us_retirement_distributions_from_manifest, + us_retirement_distributions_signal_gate, + us_retirement_distributions_stage_spec, + with_us_retirement_distribution_inputs, +) +from populace.build.us_runtime.source_runtime import us_source_operation_handlers +from populace.frame import US_SCHEMA, Frame, WeightKind, Weights + +ROOT = Path(__file__).resolve().parents[3] +policyengine_us_installed = importlib.util.find_spec("policyengine_us") is not None +requires_us = pytest.mark.skipif( + not policyengine_us_installed, + reason="requires the policyengine-us [us] extra (build environment)", +) + +_OUTPUTS = US_RETIREMENT_DISTRIBUTION_OUTPUT_COLUMNS +_PUF_QRF_OUTPUTS = ( + "taxable_401k_distributions", + "taxable_403b_distributions", + "keogh_distributions", + "taxable_sep_distributions", +) +_SLOT_SUFFIXES = ("1", "2", "1_YNG", "2_YNG") + + +def _person_source() -> pd.DataFrame: + count = 8 + person = pd.DataFrame( + { + "person_id": np.arange(1, count + 1, dtype="int64"), + "person_household_id": np.arange(10, 10 + count, dtype="int64"), + "person_tax_unit_id": np.arange(100, 100 + count, dtype="int64"), + "person_spm_unit_id": np.arange(1_000, 1_000 + count, dtype="int64"), + "person_family_id": np.arange(10_000, 10_000 + count, dtype="int64"), + "person_marital_unit_id": np.arange( + 100_000, 100_000 + count, dtype="int64" + ), + "age": [65, 66, 67, 68, 69, 70, 71, 40], + "is_female": [False, True, False, True, False, True, False, True], + "has_esi": [False] * count, + "tax_unit_role_input": ["HEAD"] * count, + "employment_income_before_lsr": [50_000.0] * count, + "self_employment_income_before_lsr": [0.0] * count, + "social_security_retirement": [10_000.0] * count, + "social_security_disability": [0.0] * count, + "social_security_dependents": [0.0] * count, + "social_security_survivors": [0.0] * count, + } + ) + for suffix in _SLOT_SUFFIXES: + person[f"DST_SC{suffix}"] = 0 + person[f"DST_VAL{suffix}"] = 0.0 + person["DST_SC1"] = np.arange(1, count + 1) % 8 + person["DST_VAL1"] = [100.0, 200.0, 300.0, 400.0, 500.0, 600.0, 700.0, 0.0] + # A second 401(k) slot proves that repeated account codes sum by person. + person.loc[0, ["DST_SC2", "DST_VAL2"]] = [1, 50.0] + return person + + +def _frame() -> Frame: + person = _person_source() + identifiers = { + "household": person["person_household_id"].to_numpy(), + "tax_unit": person["person_tax_unit_id"].to_numpy(), + "spm_unit": person["person_spm_unit_id"].to_numpy(), + "family": person["person_family_id"].to_numpy(), + "marital_unit": person["person_marital_unit_id"].to_numpy(), + } + tables = { + entity: pd.DataFrame({f"{entity}_id": values}) + for entity, values in identifiers.items() + } + tables["person"] = person + tables["tax_unit"]["filing_status_input"] = ["SINGLE"] * len(person) + # Total = 10,000. Rare outputs receive smaller weights so the fixture + # reproduces plausible population shares while retaining every account. + weights = np.asarray([1_000, 100, 100, 1_000, 1, 100, 100, 7_599], dtype=float) + return Frame( + tables, + US_SCHEMA, + {"household": Weights(weights, WeightKind.DESIGN)}, + ) + + +def _operation(kind: str = "derive_retirement_distributions"): + return next( + operation + for operation in us_retirement_distributions_stage_spec().operations + if operation.kind == kind + ) + + +def _derive(frame: pd.DataFrame) -> pd.DataFrame: + return derive_us_retirement_distributions_from_manifest( + frame, + _operation(), + None, + ) + + +def _sha256(path: Path) -> str: + digest = sha256() + with path.open("rb") as stream: + for chunk in iter(lambda: stream.read(1024 * 1024), b""): + digest.update(chunk) + return digest.hexdigest() + + +def test_stage_manifest_pins_direct_mapping_and_retired_qrf() -> None: + spec = us_retirement_distributions_stage_spec() + + assert spec.stage == "retirement_distributions" + assert spec.grain == "person" + assert tuple(spec.outputs) == _OUTPUTS + assert [operation.kind for operation in spec.operations] == [ + "read_table", + "derive_retirement_distributions", + "impute_retirement_distributions_to_puf_support", + ] + assert spec.operations[1].parameters["output_by_account_code"] == { + "1": "taxable_401k_distributions", + "2": "taxable_403b_distributions", + "3": "tax_exempt_ira_distributions", + "4": "taxable_ira_distributions", + "5": "keogh_distributions", + "6": "taxable_sep_distributions", + } + impute = spec.operations[2].parameters + assert impute["predictors"] == [ + "age", + "is_male", + "has_esi", + "tax_unit_is_joint", + "tax_unit_count_dependents", + "employment_income", + "self_employment_income", + "social_security", + ] + assert impute["max_train_samples"] == 5_000 + assert impute["n_estimators"] == 100 + assert RETIREMENT_DISTRIBUTIONS_ARCHIVED_DERIVATION_URL.endswith( + "cps.py#L1448-L1481" + ) + assert RETIREMENT_DISTRIBUTIONS_ARCHIVED_PARAMETERS_URL.endswith( + "imputation_parameters.yaml#L10-L15" + ) + + +def test_handler_is_registered() -> None: + handlers = us_source_operation_handlers() + assert ( + handlers["derive_retirement_distributions"] + is derive_us_retirement_distributions_from_manifest + ) + assert ( + handlers["impute_retirement_distributions_to_puf_support"] + is module.impute_us_retirement_distributions_to_puf_support_from_manifest + ) + + +def test_direct_mapping_sums_all_four_measured_slots() -> None: + person = _person_source() + person.loc[0, ["DST_SC1_YNG", "DST_VAL1_YNG"]] = [1, 25.0] + person.loc[0, ["DST_SC2_YNG", "DST_VAL2_YNG"]] = [1, 75.0] + + result = _derive(person) + + np.testing.assert_array_equal( + result[list(_OUTPUTS)].to_numpy(), + np.asarray( + [ + [250.0, 0.0, 0.0, 0.0, 0.0, 0.0], + [0.0, 200.0, 0.0, 0.0, 0.0, 0.0], + [0.0, 0.0, 300.0, 0.0, 0.0, 0.0], + [0.0, 0.0, 0.0, 400.0, 0.0, 0.0], + [0.0, 0.0, 0.0, 0.0, 500.0, 0.0], + [0.0, 0.0, 0.0, 0.0, 0.0, 600.0], + [0.0, 0.0, 0.0, 0.0, 0.0, 0.0], + [0.0, 0.0, 0.0, 0.0, 0.0, 0.0], + ] + ), + ) + + +@pytest.mark.parametrize("missing", US_RETIREMENT_DISTRIBUTION_REQUIRED_SOURCE_COLUMNS) +def test_direct_mapping_fails_closed_without_each_source_column(missing: str) -> None: + with pytest.raises(SourceRuntimeError, match=missing): + _derive(_person_source().drop(columns=[missing])) + + +@pytest.mark.parametrize( + ("column", "value", "message"), + [ + ("DST_SC1", np.nan, "account codes"), + ("DST_SC1", 8, "account codes"), + ("DST_SC1", 1.5, "account codes"), + ("DST_VAL1", np.nan, "nonnegative amounts"), + ("DST_VAL1", -1, "nonnegative amounts"), + ], +) +def test_direct_mapping_rejects_invalid_source_values( + column: str, + value: float, + message: str, +) -> None: + person = _person_source() + person[column] = person[column].astype(float) + person.loc[0, column] = value + + with pytest.raises(SourceRuntimeError, match=message): + _derive(person) + + +def test_direct_mapping_rejects_amount_on_not_in_universe_slot() -> None: + person = _person_source() + person.loc[7, "DST_VAL1"] = 1.0 + with pytest.raises(SourceRuntimeError, match="NIU code 0"): + _derive(person) + + +def test_wrong_operation_and_missing_table_fail_closed() -> None: + wrong = SourceStageSpec.from_mapping( + { + "stage": "test", + "survey": "test", + "source": "https://example.com", + "grain": "person", + "operations": [{"kind": "derive"}], + "outputs": list(_OUTPUTS), + } + ).operations[0] + with pytest.raises(SourceRuntimeError, match="unexpected operation"): + derive_us_retirement_distributions_from_manifest(_person_source(), wrong, None) + with pytest.raises(SourceRuntimeError, match="person table"): + derive_us_retirement_distributions_from_manifest(None, _operation(), None) + + +def test_frame_integration_gate_and_idempotence() -> None: + result = with_us_retirement_distribution_inputs(_frame(), seed=0, time_period=2024) + gate = us_retirement_distributions_signal_gate(result) + + assert gate.passed, gate.failures + assert all(value == 0 for value in gate.details["source_mismatches"].values()) + assert ( + with_us_retirement_distribution_inputs(result, seed=0, time_period=2024) + is result + ) + + +def test_puf_half_uses_qrf_and_asec_half_remains_exact( + monkeypatch: pytest.MonkeyPatch, +) -> None: + direct = with_us_retirement_distribution_inputs(_frame(), seed=0, time_period=2024) + expanded = clone_us_frame_for_puf_support(direct) + expanded_person = expanded.table("person") + puf_mask = expanded_person["person_support_channel"] == "puf_tax_detail" + # The PUF tax-detail stage owns this leaf; the retirement stage must not + # replace it with either the copied ASEC value or a QRF prediction. + puf_indices = expanded_person.index[puf_mask] + expanded_person.loc[puf_mask, "taxable_ira_distributions"] = 0.0 + expanded_person.loc[puf_indices[3], "taxable_ira_distributions"] = 999.0 + calls: dict[str, object] = {} + + class FakeFitted: + def predict(self, test: pd.DataFrame) -> pd.DataFrame: + calls["test"] = test.copy() + return pd.DataFrame( + 0.0, + index=test.index, + columns=list(_PUF_QRF_OUTPUTS), + ) + + class FakeQRF: + def __init__(self, **kwargs: object) -> None: + calls["init"] = kwargs + + def fit( + self, + training: pd.DataFrame, + predictors: list[str], + targets: list[str], + *, + weights: np.ndarray, + ) -> FakeFitted: + calls["training"] = training.copy() + calls["predictors"] = predictors + calls["targets"] = targets + calls["weights"] = weights.copy() + return FakeFitted() + + monkeypatch.setattr(module, "QRF", FakeQRF) + result = with_us_retirement_distribution_inputs(expanded, seed=7, time_period=2024) + + assert calls["init"] == {"n_estimators": 100, "seed": 7} + assert len(calls["training"]) == len(direct.table("person")) + assert len(calls["test"]) == len(direct.table("person")) + assert calls["targets"] == list(_PUF_QRF_OUTPUTS) + person = result.table("person") + asec = person[person["person_support_channel"] == "asec"] + puf = person[person["person_support_channel"] == "puf_tax_detail"] + np.testing.assert_array_equal( + asec[list(_OUTPUTS)].to_numpy(), + direct.table("person")[list(_OUTPUTS)].to_numpy(), + ) + assert not puf[list(_PUF_QRF_OUTPUTS)].to_numpy().any() + np.testing.assert_array_equal( + puf["tax_exempt_ira_distributions"].to_numpy(), + direct.table("person")["tax_exempt_ira_distributions"].to_numpy(), + ) + np.testing.assert_array_equal( + puf["taxable_ira_distributions"].to_numpy(), + [0.0, 0.0, 0.0, 999.0, 0.0, 0.0, 0.0, 0.0], + ) + gate = us_retirement_distributions_signal_gate(result) + assert gate.passed, gate.failures + + +def test_gate_rejects_a_default_or_source_divergent_leaf() -> None: + result = with_us_retirement_distribution_inputs(_frame(), seed=0, time_period=2024) + result.table("person")["keogh_distributions"] = 0.0 + gate = us_retirement_distributions_signal_gate(result) + assert not gate.passed + assert any("keogh_distributions" in failure for failure in gate.failures) + + +def test_all_sha_locked_asec_artifacts_carry_exact_source_signal() -> None: + expected = { + "census_cps_2022.h5": ( + "7ccca976284bb47815d84460cc4f75a0a65d26d7754ab0a0f417de351b3d474e", + 4, + 44_000.0, + ), + "census_cps_2023.h5": ( + "cb57817327799f42b741caed5f9be94d04021c2e6809c1ad7bd0686da5428d88", + 5, + 130_040.0, + ), + "census_cps_2024.h5": ( + "ec36604cb735a660b51b0b2f90be27d803b5878f3464fb30d0eacead59c1260d", + 4, + 166_600.0, + ), + } + summary = json.loads( + (ROOT / "experiments/build_j_recert/base_j.summary.json").read_text() + ) + paths = { + Path(item["path"]).name: Path(item["path"]) + for item in summary["base_source"]["sources"] + } + if not all(paths[name].is_file() for name in expected): + pytest.skip("SHA-locked ASEC artifacts are not mounted") + + for name, (digest, positives, total) in expected.items(): + path = paths[name] + assert _sha256(path) == digest + person = load_asec_h5_tables(path)["person"] + assert set(US_RETIREMENT_DISTRIBUTION_REQUIRED_SOURCE_COLUMNS) <= set( + person.columns + ) + derived = _derive(person)["keogh_distributions"] + assert int((derived > 0).sum()) == positives + assert float(derived.sum()) == total + + +@requires_us +def test_policyengine_1_764_6_contract_and_live_binding() -> None: + from policyengine_core.reforms import Reform + from policyengine_us import CountryTaxBenefitSystem, Simulation + + assert version("policyengine-us") == "1.764.6" + variables = CountryTaxBenefitSystem().variables + for name in _OUTPUTS: + variable = variables[name] + assert variable.is_input_variable() + assert variable.entity.key == "person" + assert variable.value_type is float + assert variable.default_value == 0 + assert str(variable.definition_period).lower() == "year" + + situation = { + "people": { + "adult": { + "age": {"2024": 65}, + "employment_income": {"2024": 60_000}, + "keogh_distributions": {"2024": 10_000}, + } + }, + "tax_units": {"tax_unit": {"members": ["adult"]}}, + "spm_units": {"spm_unit": {"members": ["adult"]}}, + "households": { + "household": { + "members": ["adult"], + "state_code": {"2024": "CA"}, + } + }, + } + + class NeutralizeKeogh(Reform): + def apply(self) -> None: + self.neutralize_variable("keogh_distributions") + + baseline = Simulation(situation=situation) + neutralized = Simulation(situation=situation, reform=NeutralizeKeogh) + assert baseline.calculate("taxable_retirement_distributions", 2024)[0] == 10_000 + assert neutralized.calculate("taxable_retirement_distributions", 2024)[0] == 0 + for measure in ("irs_gross_income", "adjusted_gross_income"): + assert baseline.calculate(measure, 2024)[0] == pytest.approx( + neutralized.calculate(measure, 2024)[0] + 10_000 + ) + assert ( + baseline.calculate("income_tax", 2024)[0] + > neutralized.calculate("income_tax", 2024)[0] + ) + + +def test_shipped_keogh_neutralization_probe() -> None: + probe = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "keogh_distribution_neutralization" + ) + assert probe.neutralized_variable == "keogh_distributions" + assert probe.binding_inputs == ("keogh_distributions",) + assert probe.budget_measure == "income_tax" + assert probe.period == 2024 + assert probe.effect_direction == "baseline_minus_reform" + assert probe.expected_sign == "positive" + assert probe.min_abs_effect >= 1_000_000.0 diff --git a/packages/populace-build/tests/test_us_salt_refund_income.py b/packages/populace-build/tests/test_us_salt_refund_income.py new file mode 100644 index 00000000..2c686cae --- /dev/null +++ b/packages/populace-build/tests/test_us_salt_refund_income.py @@ -0,0 +1,363 @@ +"""Contracts for the IRS PUF state-and-local-tax-refund income input.""" + +from __future__ import annotations + +import importlib.util + +import numpy as np +import pandas as pd +import pytest + +from populace.build.us_runtime import ( + BASE_ASEC_SUPPORT_CHANNEL, + PUF_TAX_DETAIL_SUPPORT_CHANNEL, + clone_us_frame_for_puf_support, + impute_us_puf_tax_detail_support, + puf_tax_unit_donor_from_arrays, + support_channel_column, +) +from populace.build.us_runtime import puf_support as puf_support_module +from populace.build.us_runtime.l0_refit_export import ( + US_RELEASE_REQUIRED_PERSON_SOURCE_COLUMNS, +) +from populace.build.us_runtime.puf_aggregate_records import ( + _reconcile_puf_salt_refund_income_from_source, + derive_puf_policyengine_variables, +) +from populace.build.us_runtime.reform_coverage_smoke import _build_reform +from populace.build.us_runtime.release_input_coverage import ( + RESTORED_REFERENCE_ECPS_REQUIRED_INPUTS, + load_release_input_coverage_manifest, + us_release_reform_coverage_probes, +) +from populace.build.us_runtime.salt_refund_income import ( + SALT_REFUND_ARCHIVED_DERIVATION_URL, + SALT_REFUND_ARCHIVED_EXPORT_URL, + SALT_REFUND_ARCHIVED_IMPUTATION_URL, + SALT_REFUND_ARCHIVED_PERSON_ALLOCATION_URL, + SALT_REFUND_ARCHIVED_PUF_ARTIFACT_URL, + derive_us_salt_refund_income_from_puf, + us_salt_refund_income_signal_gate, + us_salt_refund_income_stage_spec, +) +from populace.frame import US_SCHEMA, Frame, WeightKind, Weights + +policyengine_us_installed = importlib.util.find_spec("policyengine_us") is not None +requires_us = pytest.mark.skipif( + not policyengine_us_installed, + reason="requires the policyengine-us [us] extra (build environment)", +) + + +class _ResolvedWeights: + def __init__(self, values: np.ndarray) -> None: + self.values = values + + +class _PersonFrame: + def __init__( + self, + person: pd.DataFrame, + weights: np.ndarray | None = None, + ) -> None: + self._person = person + self._weights = np.ones(len(person)) if weights is None else weights + + def table(self, entity: str) -> pd.DataFrame: + assert entity == "person" + return self._person + + def resolve_weights(self, entity: str) -> _ResolvedWeights: + assert entity == "person" + return _ResolvedWeights(np.asarray(self._weights, dtype=np.float64)) + + +def _joint_support_frame() -> Frame: + tables = { + "person": pd.DataFrame( + { + "person_id": np.asarray([1, 2], dtype="int64"), + "person_household_id": np.asarray([10, 10], dtype="int64"), + "person_tax_unit_id": np.asarray([100, 100], dtype="int64"), + "person_spm_unit_id": np.asarray([1_000, 1_000], dtype="int64"), + "person_family_id": np.asarray([10_000, 10_000], dtype="int64"), + "person_marital_unit_id": np.asarray([100_000, 100_000], dtype="int64"), + "employment_income_before_lsr": [50_000.0, 20_000.0], + "self_employment_income_before_lsr": [0.0, 40_000.0], + } + ), + "household": pd.DataFrame({"household_id": np.asarray([10], dtype="int64")}), + "tax_unit": pd.DataFrame( + { + "tax_unit_id": np.asarray([100], dtype="int64"), + "filing_status_input": ["JOINT"], + } + ), + "spm_unit": pd.DataFrame({"spm_unit_id": np.asarray([1_000], dtype="int64")}), + "family": pd.DataFrame({"family_id": np.asarray([10_000], dtype="int64")}), + "marital_unit": pd.DataFrame( + {"marital_unit_id": np.asarray([100_000], dtype="int64")} + ), + } + return Frame( + tables, + US_SCHEMA, + {"household": Weights(np.asarray([100.0]), WeightKind.DESIGN)}, + ) + + +def test_archived_coordinates_pin_derivation_export_and_imputation() -> None: + commit = "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe" + + assert commit in SALT_REFUND_ARCHIVED_DERIVATION_URL + assert SALT_REFUND_ARCHIVED_DERIVATION_URL.endswith("puf.py#L707") + assert commit in SALT_REFUND_ARCHIVED_EXPORT_URL + assert SALT_REFUND_ARCHIVED_EXPORT_URL.endswith("puf.py#L804-L850") + assert commit in SALT_REFUND_ARCHIVED_IMPUTATION_URL + assert SALT_REFUND_ARCHIVED_IMPUTATION_URL.endswith("puf_impute.py#L90-L198") + assert SALT_REFUND_ARCHIVED_PERSON_ALLOCATION_URL.endswith("puf.py#L1513-L1601") + assert SALT_REFUND_ARCHIVED_PUF_ARTIFACT_URL.endswith("puf.py#L1655-L1660") + + +def test_archived_e00700_mapping_is_an_exact_carry() -> None: + source = pd.DataFrame({"E00700": [0.0, 125.5, 9_000.0]}) + + result = derive_us_salt_refund_income_from_puf(source) + + assert "salt_refund_income" not in source.columns + assert result["salt_refund_income"].tolist() == [0.0, 125.5, 9_000.0] + + +@pytest.mark.parametrize( + ("source", "message"), + [ + (pd.DataFrame({"other": [1.0]}), "requires source column"), + (pd.DataFrame({"E00700": ["not numeric"]}), "nonnumeric or nonfinite"), + (pd.DataFrame({"E00700": [np.inf]}), "nonnumeric or nonfinite"), + (pd.DataFrame({"E00700": [-1.0]}), "negative value"), + ], +) +def test_source_derivation_fails_closed( + source: pd.DataFrame, + message: str, +) -> None: + with pytest.raises(ValueError, match=message): + derive_us_salt_refund_income_from_puf(source) + + +def test_shared_puf_derivation_and_post_disaggregation_reconciliation() -> None: + source = pd.DataFrame( + { + "E00600": [0.0, 0.0, 0.0], + "E00650": [0.0, 0.0, 0.0], + "E00700": [0.0, 3_000.0, 7_500.0], + } + ) + + derived = derive_puf_policyengine_variables( + source, + salt_refund_income_source="E00700", + ) + derived["salt_refund_income"] = 999.0 + result = _reconcile_puf_salt_refund_income_from_source(derived) + + assert result["salt_refund_income"].tolist() == [0.0, 3_000.0, 7_500.0] + + +def test_shared_puf_stage_declares_exact_source_output_and_artifact() -> None: + stage = us_salt_refund_income_stage_spec() + operation = next( + operation + for operation in stage.operations + if operation.kind == "derive_puf_policyengine_variables" + ) + + assert operation.parameters["salt_refund_income_source"] == "E00700" + assert operation.parameters["salt_refund_income_output"] == "salt_refund_income" + assert "salt_refund_income" in stage.outputs + assert "salt_refund_income" in stage.nonnegative_outputs + assert any("E00700" in str(artifact.get("locator")) for artifact in stage.artifacts) + assert any( + "irs-soi-puf/1.8.0/puf_2024.h5" in str(artifact.get("locator")) + for artifact in stage.artifacts + ) + + +def test_puf_support_keeps_salt_refunds_nonnegative_sparse_and_unsplit() -> None: + output = "salt_refund_income" + + assert output in puf_support_module.PUF_TAX_DETAIL_DEFAULT_PERSON_OUTPUTS + assert output in puf_support_module._PUF_TAX_DETAIL_NONNEGATIVE_OUTPUTS + assert output in puf_support_module._PUF_TAX_DETAIL_SPARSE_PERSON_OUTPUTS + assert output not in puf_support_module._PERSON_OUTPUT_DISTRIBUTION_BASIS + + +def test_processed_puf_people_aggregate_to_one_tax_unit_refund() -> None: + donor = puf_tax_unit_donor_from_arrays( + { + "tax_unit_id": [10, 20], + "household_weight": [100.0, 200.0], + "filing_status": [b"SINGLE", b"JOINT"], + "person_tax_unit_id": [10, 20, 20], + "salt_refund_income": [0.0, 100.0, 200.0], + }, + person_outputs=("salt_refund_income",), + tax_unit_outputs=(), + ) + + assert donor["salt_refund_income"].tolist() == [0.0, 300.0] + + +def test_puf_imputation_preserves_tax_unit_total_without_inventing_split() -> None: + expanded = clone_us_frame_for_puf_support(_joint_support_frame()) + donor = pd.DataFrame( + { + "filing_status_code": [2.0, 2.0], + "tax_unit_person_count": [2.0, 2.0], + "salt_refund_income": [900.0, 900.0], + "weight": [1.0, 1.0], + } + ) + + imputed = impute_us_puf_tax_detail_support( + expanded, + donor, + predictors=( + "puf_predictor_filing_status_code", + "puf_predictor_tax_unit_person_count", + ), + person_outputs=("salt_refund_income",), + tax_unit_outputs=(), + n_estimators=4, + seed=0, + ) + person = imputed.table("person") + channel = support_channel_column("person") + asec = person[person[channel] == BASE_ASEC_SUPPORT_CHANNEL] + puf = person[person[channel] == PUF_TAX_DETAIL_SUPPORT_CHANNEL].sort_values( + "person_id" + ) + + assert asec["salt_refund_income"].tolist() == [0.0, 0.0] + np.testing.assert_allclose(puf["salt_refund_income"].to_numpy(), [900.0, 0.0]) + assert puf["salt_refund_income"].sum() == pytest.approx(900.0) + + +def test_signal_gate_accepts_source_aligned_nondefault_values() -> None: + values = np.zeros(2_000) + values[np.arange(1_000, 1_100)] = np.linspace(100.0, 10_000.0, 100) + channels = np.asarray(["asec"] * 1_000 + ["puf_tax_detail"] * 1_000) + frame = _PersonFrame( + pd.DataFrame( + { + "salt_refund_income": values, + "person_support_channel": channels, + } + ) + ) + + result = us_salt_refund_income_signal_gate(frame) # type: ignore[arg-type] + + assert result.passed, result.failures + assert result.details["positive_share"] == pytest.approx(0.05) + assert result.details["channels"]["asec"]["weighted_total"] == 0.0 + assert result.details["channels"]["puf_tax_detail"]["positive_share"] == ( + pytest.approx(0.10) + ) + + +@pytest.mark.parametrize( + "person", + [ + pd.DataFrame({"other": [0.0, 1.0]}), + pd.DataFrame({"salt_refund_income": [0.0, 0.0]}), + pd.DataFrame({"salt_refund_income": [0.0, -1.0]}), + pd.DataFrame({"salt_refund_income": [0.0, np.nan]}), + pd.DataFrame( + { + "salt_refund_income": [10.0, 0.0], + "person_support_channel": ["asec", "puf_tax_detail"], + } + ), + ], +) +def test_signal_gate_rejects_missing_default_invalid_or_asec_signal( + person: pd.DataFrame, +) -> None: + result = us_salt_refund_income_signal_gate( # type: ignore[arg-type] + _PersonFrame(person) + ) + + assert not result.passed + + +def test_release_wiring_keeps_restored_input_hard_required() -> None: + output = "salt_refund_income" + + assert output in US_RELEASE_REQUIRED_PERSON_SOURCE_COLUMNS + assert output in RESTORED_REFERENCE_ECPS_REQUIRED_INPUTS + assert output in load_release_input_coverage_manifest().required_columns + + +def test_shipped_neutralization_probe_binds_only_through_salt_refunds() -> None: + probe = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "salt_refund_income_neutralization" + ) + + assert probe.parameter_changes == {} + assert probe.neutralized_variable == "salt_refund_income" + assert probe.binding_inputs == ("salt_refund_income",) + assert probe.period == 2024 + assert probe.effect_direction == "baseline_minus_reform" + assert probe.expected_sign == "negative" + assert probe.budget_measure == "state_income_tax" + assert probe.min_abs_effect == 1_000_000.0 + + +@requires_us +def test_policyengine_us_contract_and_state_tax_binding() -> None: + from policyengine_us import CountryTaxBenefitSystem, Simulation + + variable = CountryTaxBenefitSystem().variables["salt_refund_income"] + assert variable.is_input_variable() + assert variable.entity.key == "person" + assert str(variable.definition_period).lower() == "year" + assert variable.default_value == 0 + + year = 2024 + adult = { + "age": {year: 40}, + "employment_income": {year: 100_000}, + "salt_refund_income": {year: 10_000}, + } + entities = { + "tax_units": {"unit": {"members": ["adult"]}}, + "families": {"family": {"members": ["adult"]}}, + "spm_units": {"spm": {"members": ["adult"]}}, + "households": { + "household": { + "members": ["adult"], + "state_code": {year: "SC"}, + } + }, + "marital_units": {"marital": {"members": ["adult"]}}, + } + baseline = Simulation(situation={"people": {"adult": adult}, **entities}) + probe = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "salt_refund_income_neutralization" + ) + neutralized = Simulation( + situation={"people": {"adult": adult}, **entities}, + reform=_build_reform(probe), + ) + + assert baseline.calculate("sc_subtractions", year)[0] == pytest.approx(10_000) + assert neutralized.calculate("sc_subtractions", year)[0] == 0 + assert ( + baseline.calculate("state_income_tax", year)[0] + < neutralized.calculate("state_income_tax", year)[0] + ) diff --git a/packages/populace-build/tests/test_us_scf_auto_loans.py b/packages/populace-build/tests/test_us_scf_auto_loans.py new file mode 100644 index 00000000..acf22dc5 --- /dev/null +++ b/packages/populace-build/tests/test_us_scf_auto_loans.py @@ -0,0 +1,418 @@ +"""SCF auto-loan and OBBBA qualifying-interest stage tests. + +The retired eCPS imputed household auto-loan balance and annual interest from +the full 2022 SCF, conditioning on an SCF-style household reference person and +household-summed income. The OBBBA deduction uses a later, distinct pure input +(``qualified_passenger_vehicle_loan_interest``), so this port also derives a +documented expected qualifying share rather than leaving the reform structural +zero. +""" + +from __future__ import annotations + +import numpy as np +import pandas as pd +import pytest + +from populace.build.us_runtime import ( + QUALIFIED_AUTO_LOAN_ANNUAL_ISSUANCE_TARGET, + SCF_2022_FULL_EXTRACT_MEMBER, + SCF_2022_FULL_EXTRACT_MEMBER_SHA256, + SCF_2022_FULL_EXTRACT_URL, + SCF_2022_FULL_EXTRACT_ZIP_SHA256, + SCF_AUTO_LOAN_AMOUNT_COLUMNS, + SCF_AUTO_LOAN_RATE_COLUMNS, + US_SCF_AUTO_LOAN_OUTPUT_COLUMNS, + fetch_scf_2022_full_extract, + impute_us_scf_auto_loans, + load_scf_2022_auto_loan_donor, + qualified_auto_loan_interest_proxy, + us_scf_auto_loans_signal_gate, + us_scf_auto_loans_stage_spec, + us_scf_auto_loans_summary, + with_us_scf_auto_loan_inputs, +) +from populace.build.us_runtime.scf_auto_loans import ( + _scf_reference_person_mask, + _sha256_hexdigest, +) +from populace.build.us_runtime.scf_wealth import SCF_WEALTH_PREDICTORS +from populace.frame import US_SCHEMA, Frame, WeightKind, Weights + +TIME_PERIOD = 2024 +_DONOR_WEIGHT_COLUMN = "scf_weight" + + +def _raw_scf_summary(n_families: int = 8) -> pd.DataFrame: + """A summary-extract-shaped table with all five SCF implicates.""" + + records: list[dict[str, float]] = [] + for family in range(1, n_families + 1): + for implicate in range(1, 6): + records.append( + { + "y1": float(family), + "yy1": float(implicate), + "wgt": 1_000.0 + family, + "age": 30.0 + family, + "hhsex": 1.0 if family % 2 else 2.0, + "racecl5": float((family % 5) + 1), + "married": float(family % 2), + "kids": float(family % 3), + "wageinc": 20_000.0 + 1_000.0 * family, + "intdivinc": 200.0 * family, + "ssretinc": 500.0 * family, + } + ) + return pd.DataFrame.from_records(records) + + +def _raw_scf_full(summary: pd.DataFrame) -> pd.DataFrame: + """A full-extract-shaped table deliberately in reverse join order.""" + + full = summary.loc[:, ["y1", "yy1"]].copy() + row = np.arange(1, len(full) + 1, dtype=np.float64) + full["x2209"] = 10_000.0 + row + full["x2309"] = np.where(row % 2 == 0, 2_000.0, -1.0) + full["x2409"] = 0.0 + full["x7158"] = np.where(row % 3 == 0, 3_000.0, -7.0) + full["x2219"] = 500.0 # 5 percent in SCF's rate encoding. + full["x2319"] = 250.0 + full["x2419"] = -1.0 + full["x7170"] = 1_000.0 + return full.iloc[::-1].reset_index(drop=True) + + +def _write_scf_extracts(tmp_path, *, n_families: int = 8): + summary = _raw_scf_summary(n_families) + full = _raw_scf_full(summary) + summary_path = tmp_path / "rscfp2022.dta" + full_path = tmp_path / "p22i6.dta" + summary.to_stata(summary_path, write_index=False) + full.to_stata(full_path, write_index=False) + return summary, full, summary_path, full_path + + +def _ready_donor(n: int = 400) -> pd.DataFrame: + rng = np.random.default_rng(7) + donor = pd.DataFrame({name: rng.normal(size=n) for name in SCF_WEALTH_PREDICTORS}) + donor["age"] = rng.integers(18, 90, n).astype(np.float64) + donor["is_female"] = rng.integers(0, 2, n).astype(np.float64) + donor["cps_race"] = rng.integers(1, 8, n).astype(np.float64) + donor["is_married"] = rng.integers(0, 2, n).astype(np.float64) + donor["own_children_in_household"] = rng.integers(0, 4, n).astype(np.float64) + donor["employment_income"] = rng.gamma(2.0, 20_000.0, n) + donor["interest_dividend_income"] = rng.gamma(1.0, 1_000.0, n) + donor["social_security_pension_income"] = rng.gamma(1.0, 7_000.0, n) + has_loan = rng.random(n) < 0.35 + donor["auto_loan_balance"] = np.where(has_loan, rng.gamma(2.0, 12_000.0, n), 0.0) + donor["auto_loan_interest"] = np.where( + has_loan, donor["auto_loan_balance"] * rng.uniform(0.02, 0.10, n), 0.0 + ) + donor[_DONOR_WEIGHT_COLUMN] = rng.uniform(500.0, 2_000.0, n) + return donor + + +def _person_rows(n_households: int = 120) -> pd.DataFrame: + rng = np.random.default_rng(8) + records: list[dict[str, object]] = [] + person_id = 1 + for household_id in range(1, n_households + 1): + for line, age in ((1, rng.integers(25, 75)), (2, rng.integers(0, 70))): + records.append( + { + "person_id": person_id, + "person_household_id": household_id, + "PH_SEQ": household_id, + "A_LINENO": line, + "age": float(age), + "is_female": bool(rng.integers(0, 2)), + "PRDTRACE": int(rng.integers(1, 7)), + "PRDTHSP": int(rng.integers(0, 2)), + "A_MARITL": int(rng.choice([1, 2, 4, 5, 6, 7])), + "PEPAR1": -1, + "PEPAR2": -1, + "employment_income_before_lsr": float(rng.gamma(2.0, 12_000.0)), + "taxable_interest_income": float(rng.gamma(1.0, 500.0)), + "social_security_retirement": float(rng.gamma(1.0, 3_000.0)), + } + ) + person_id += 1 + return pd.DataFrame.from_records(records) + + +def _us_frame( + person: pd.DataFrame, + *, + weights: np.ndarray | None = None, +) -> Frame: + person = person.copy() + household_ids = np.sort(person["person_household_id"].unique()) + n_person = len(person) + person["person_tax_unit_id"] = person["person_household_id"] + 1_000 + person["person_spm_unit_id"] = person["person_household_id"] + 2_000 + person["person_family_id"] = person["person_household_id"] + 3_000 + person["person_marital_unit_id"] = np.arange(n_person) + 4_000 + tables = { + "person": person, + "household": pd.DataFrame({"household_id": household_ids}), + "tax_unit": pd.DataFrame({"tax_unit_id": household_ids + 1_000}), + "spm_unit": pd.DataFrame({"spm_unit_id": household_ids + 2_000}), + "family": pd.DataFrame({"family_id": household_ids + 3_000}), + "marital_unit": pd.DataFrame({"marital_unit_id": np.arange(n_person) + 4_000}), + } + values = ( + np.asarray(weights, dtype=np.float64) + if weights is not None + else np.full(len(household_ids), 1_000_000.0, dtype=np.float64) + ) + return Frame( + tables, + US_SCHEMA, + {"household": Weights(values=values, kind=WeightKind.DESIGN)}, + ) + + +def test_full_extract_coordinates_and_pin_state_are_explicit() -> None: + assert SCF_2022_FULL_EXTRACT_URL.endswith("/scf2022s.zip") + assert SCF_2022_FULL_EXTRACT_MEMBER == "p22i6.dta" + pins = ( + SCF_2022_FULL_EXTRACT_ZIP_SHA256, + SCF_2022_FULL_EXTRACT_MEMBER_SHA256, + ) + # Never silently mix a pinned and unpinned artifact, and never substitute + # the unrelated summary-extract hashes. This environment cannot make the + # one network fetch needed to fill the full-file pins; once filled, both + # values must be sha-256 hex digests. + assert all(pin is None for pin in pins) or all( + isinstance(pin, str) and len(pin) == 64 and set(pin) <= set("0123456789abcdef") + for pin in pins + ) + + +def test_stage_spec_declares_all_three_auto_loan_outputs_and_artifact() -> None: + spec = us_scf_auto_loans_stage_spec() + assert set(US_SCF_AUTO_LOAN_OUTPUT_COLUMNS) <= set(spec.outputs) + assert set(US_SCF_AUTO_LOAN_OUTPUT_COLUMNS) <= set(spec.nonnegative_outputs) + artifact = next( + artifact + for artifact in spec.artifacts + if artifact.get("locator") == SCF_2022_FULL_EXTRACT_URL + ) + assert artifact.get("member") == SCF_2022_FULL_EXTRACT_MEMBER + assert artifact.get("expected_rows") == 22_975 + if SCF_2022_FULL_EXTRACT_ZIP_SHA256 is None: + assert "pending" in artifact.get("integrity_note", "") + else: + assert artifact.get("sha256") == SCF_2022_FULL_EXTRACT_ZIP_SHA256 + assert artifact.get("member_sha256") == SCF_2022_FULL_EXTRACT_MEMBER_SHA256 + + +def test_source_column_pairs_match_four_retired_vehicle_loans() -> None: + assert SCF_AUTO_LOAN_AMOUNT_COLUMNS == ("x2209", "x2309", "x2409", "x7158") + assert SCF_AUTO_LOAN_RATE_COLUMNS == ("x2219", "x2319", "x2419", "x7170") + + +def test_load_donor_joins_all_five_implicates_and_derives_targets(tmp_path) -> None: + summary, full, summary_path, full_path = _write_scf_extracts(tmp_path, n_families=2) + + donor = load_scf_2022_auto_loan_donor(summary_path, full_path) + + assert len(donor) == 10 + assert list(donor.columns) == [ + *SCF_WEALTH_PREDICTORS, + "auto_loan_balance", + "auto_loan_interest", + _DONOR_WEIGHT_COLUMN, + ] + # The join follows summary order even though the full extract was reversed. + first_key = summary.loc[0, ["y1", "yy1"]].tolist() + full_first = full.loc[ + (full["y1"] == first_key[0]) & (full["yy1"] == first_key[1]) + ].iloc[0] + amounts = np.maximum( + full_first.loc[list(SCF_AUTO_LOAN_AMOUNT_COLUMNS)].to_numpy(float), 0.0 + ) + rates = ( + np.maximum( + full_first.loc[list(SCF_AUTO_LOAN_RATE_COLUMNS)].to_numpy(float), 0.0 + ) + / 10_000.0 + ) + assert donor.loc[0, "auto_loan_balance"] == pytest.approx(amounts.sum()) + assert donor.loc[0, "auto_loan_interest"] == pytest.approx( + float(np.dot(amounts, rates)) + ) + assert (donor[["auto_loan_balance", "auto_loan_interest"]] >= 0).all().all() + + +def test_load_donor_rejects_an_unmatched_implicate(tmp_path) -> None: + _, _, summary_path, full_path = _write_scf_extracts(tmp_path) + full = pd.read_stata(full_path, convert_categoricals=False).iloc[:-1] + full.to_stata(full_path, write_index=False) + + with pytest.raises(ValueError, match="one-to-one|unmatched"): + load_scf_2022_auto_loan_donor(summary_path, full_path) + + +def test_load_donor_rejects_missing_auto_source_column(tmp_path) -> None: + _, _, summary_path, full_path = _write_scf_extracts(tmp_path) + full = pd.read_stata(full_path, convert_categoricals=False).drop(columns=["x7158"]) + full.to_stata(full_path, write_index=False) + + with pytest.raises(ValueError, match="missing required column"): + load_scf_2022_auto_loan_donor(summary_path, full_path) + + +def test_reference_person_mask_matches_retired_scf_rules() -> None: + # hh1 mixed-sex married couple -> male (row 1), even though female is older. + # hh2 same-sex married couple -> older (row 2). + # hh3 non-couple multi-adult -> oldest adult (row 5). + # hh4 child-only -> oldest person (row 6). + person = pd.DataFrame( + { + "person_household_id": [1, 1, 2, 2, 3, 3, 4, 4], + "age": [55, 50, 60, 45, 22, 70, 16, 11], + "is_female": [True, False, True, True, False, True, False, True], + "A_MARITL": [1, 1, 2, 2, 7, 7, 7, 7], + "A_LINENO": [1, 2, 1, 2, 1, 2, 1, 2], + } + ) + + mask = _scf_reference_person_mask(person) + + np.testing.assert_array_equal( + mask, [False, True, True, False, False, True, True, False] + ) + + +def test_imputation_is_household_grain_nonnegative_and_deterministic() -> None: + frame = _us_frame(_person_rows(80)) + donor = _ready_donor() + + a = impute_us_scf_auto_loans(frame, donor, seed=42, n_estimators=20) + b = impute_us_scf_auto_loans(frame, donor, seed=42, n_estimators=20) + + assert list(a.columns) == ["auto_loan_balance", "auto_loan_interest"] + assert len(a) == frame.n("household") + assert (a >= 0).all().all() + assert a["auto_loan_balance"].nunique() > 1 + assert a["auto_loan_interest"].nunique() > 1 + pd.testing.assert_frame_equal(a, b) + + +def test_qualified_interest_proxy_targets_six_million_annual_loans() -> None: + interest = np.asarray([0.0, 1_000.0, 2_000.0]) + weights = np.asarray([2_000_000.0, 4_000_000.0, 6_000_000.0]) + + qualified, share = qualified_auto_loan_interest_proxy(interest, weights) + + # Positive-loan household mass is 10m; the IRS/Treasury annual estimate is + # 6m qualifying loans, so expected qualifying interest uses a 60% share. + assert QUALIFIED_AUTO_LOAN_ANNUAL_ISSUANCE_TARGET == 6_000_000.0 + assert share == pytest.approx(0.6) + np.testing.assert_allclose(qualified, [0.0, 600.0, 1_200.0]) + + +def test_with_inputs_writes_legacy_and_obbba_household_columns() -> None: + frame = _us_frame(_person_rows(120)) + + out = with_us_scf_auto_loan_inputs( + frame, + seed=42, + time_period=TIME_PERIOD, + scf_auto_loan_donor=_ready_donor(), + n_estimators=20, + ) + + household = out.table("household") + for column in US_SCF_AUTO_LOAN_OUTPUT_COLUMNS: + assert column in household.columns + assert household[column].dtype == np.float64 + assert (household[column] >= 0).all() + assert household["auto_loan_interest"].sum() > 0 + assert household["qualified_passenger_vehicle_loan_interest"].sum() > 0 + assert ( + household["qualified_passenger_vehicle_loan_interest"] + <= household["auto_loan_interest"] + ).all() + + +def test_with_inputs_is_idempotent_when_all_columns_have_signal() -> None: + frame = _us_frame(_person_rows(100)) + once = with_us_scf_auto_loan_inputs( + frame, + seed=42, + time_period=TIME_PERIOD, + scf_auto_loan_donor=_ready_donor(), + n_estimators=20, + ) + twice = with_us_scf_auto_loan_inputs( + once, + seed=99, + time_period=TIME_PERIOD, + scf_auto_loan_donor=_ready_donor(), + n_estimators=20, + ) + + for column in US_SCF_AUTO_LOAN_OUTPUT_COLUMNS: + np.testing.assert_array_equal( + once.table("household")[column], twice.table("household")[column] + ) + + +def test_signal_gate_passes_and_summary_reports_proxy_share() -> None: + out = with_us_scf_auto_loan_inputs( + _us_frame(_person_rows(160)), + seed=4, + time_period=TIME_PERIOD, + scf_auto_loan_donor=_ready_donor(), + n_estimators=20, + ) + + gate = us_scf_auto_loans_signal_gate(out) + summary = us_scf_auto_loans_summary(out) + + assert gate.passed, gate.failures + assert 0 < summary["auto_loan_interest_nonzero_share"] < 1 + assert 0 < summary["qualified_interest_share"] <= 1 + assert summary["auto_loan_balance_weighted_total"] > 0 + assert summary["auto_loan_interest_weighted_total"] > 0 + + +def test_signal_gate_fails_on_missing_or_invalid_surface() -> None: + frame = _us_frame(_person_rows(4)) + missing = us_scf_auto_loans_signal_gate(frame) + assert not missing.passed + assert "missing" in missing.failures[0] + + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + tables["household"]["auto_loan_balance"] = [0.0] * frame.n("household") + tables["household"]["auto_loan_interest"] = [0.0] * frame.n("household") + tables["household"]["qualified_passenger_vehicle_loan_interest"] = [1.0] * frame.n( + "household" + ) + invalid = Frame( + tables, + frame.schema, + {"household": frame.weights_for("household")}, + ) + gate = us_scf_auto_loans_signal_gate(invalid) + assert not gate.passed + assert any("constant" in failure for failure in gate.failures) + assert any("exceeds auto_loan_interest" in failure for failure in gate.failures) + + +def test_fetch_returns_sha_matching_cached_full_extract(tmp_path) -> None: + cached = tmp_path / SCF_2022_FULL_EXTRACT_MEMBER + cached.write_bytes(b"pinned-full-scf") + digest = _sha256_hexdigest(cached.read_bytes()) + + result = fetch_scf_2022_full_extract( + cache_dir=tmp_path, + expected_member_sha256=digest, + expected_zip_sha256=None, + ) + + assert result == cached + assert result.read_bytes() == b"pinned-full-scf" diff --git a/packages/populace-build/tests/test_us_scf_wealth.py b/packages/populace-build/tests/test_us_scf_wealth.py index 0908778f..d373ee87 100644 --- a/packages/populace-build/tests/test_us_scf_wealth.py +++ b/packages/populace-build/tests/test_us_scf_wealth.py @@ -1,13 +1,17 @@ -"""US SCF financial-asset (SSI countable-resource) stage tests (populace #356/#368). +"""US SCF wealth and SSI countable-resource tests (#49/#356/#368). The three asset leaves ``bank_account_assets`` / ``stock_assets`` / ``bond_assets`` are what ``ssi_countable_resources`` sums; with them absent the SSI resource-limit reform class scores $0 (the #356 failure). This stage -SCF-imputes them, head-carried onto the household reference person. +SCF-imputes them onto the household reference person and restores signed +household ``net_worth`` from the retired pipeline's direct SCF anchor. """ from __future__ import annotations +import importlib.util +from importlib.metadata import version + import numpy as np import pandas as pd import pytest @@ -15,11 +19,15 @@ from populace.build.source_manifest import SourceStageSpec from populace.build.us_runtime import ( SCF_FINANCIAL_ASSET_TARGET_COMPONENTS, + SCF_NET_WORTH_TARGET_COMPONENTS, SCF_WEALTH_PREDICTORS, US_SCF_FINANCIAL_ASSET_OUTPUT_COLUMNS, + US_SCF_NET_WORTH_OUTPUT_COLUMNS, + US_SCF_WEALTH_NONCONSTANT_HOUSEHOLD_COLUMNS, US_SCF_WEALTH_STAGE_NAME, fetch_scf_2022_summary_extract, impute_us_scf_financial_assets, + impute_us_scf_net_worth, load_scf_2022_financial_asset_donor, us_scf_wealth_signal_gate, us_scf_wealth_stage_spec, @@ -36,6 +44,10 @@ TIME_PERIOD = 2024 _DONOR_WEIGHT_COLUMN = "scf_weight" +requires_us = pytest.mark.skipif( + importlib.util.find_spec("policyengine_us") is None, + reason="policyengine-us extra not installed", +) # --------------------------------------------------------------------------- # @@ -47,12 +59,16 @@ def _raw_scf_summary() -> pd.DataFrame: rng = np.random.default_rng(0) n = 400 liq = rng.gamma(2.0, 3_000.0, n) + net_worth = rng.lognormal(12.0, 1.1, n) + indebted = rng.random(n) < 0.10 + net_worth[indebted] = -rng.gamma(2.0, 20_000.0, indebted.sum()) return pd.DataFrame( { "liq": liq, "stocks": rng.gamma(1.0, 5_000.0, n), "nmmf": rng.gamma(1.0, 4_000.0, n), "bond": np.where(rng.random(n) < 0.05, rng.gamma(1.0, 9_000.0, n), 0.0), + "networth": net_worth, "wgt": rng.uniform(500.0, 2_000.0, n), "age": rng.integers(20, 85, n).astype(float), "hhsex": rng.integers(1, 3, n).astype(float), @@ -87,6 +103,10 @@ def _donor_table() -> pd.DataFrame: frame["bond_assets"] = np.where( rng.random(n) < 0.04, rng.gamma(1.0, 9_000.0, n), 0.0 ) + net_worth = rng.lognormal(12.0, 1.1, n) + indebted = rng.random(n) < 0.12 + net_worth[indebted] = -rng.gamma(2.0, 20_000.0, indebted.sum()) + frame["net_worth"] = net_worth frame[_DONOR_WEIGHT_COLUMN] = rng.uniform(500.0, 2_000.0, n) return frame @@ -140,7 +160,12 @@ def _person_rows(n_households: int = 60) -> pd.DataFrame: return pd.DataFrame(records) -def _us_frame(person: pd.DataFrame, *, weights: list[float] | None = None) -> Frame: +def _us_frame( + person: pd.DataFrame, + *, + weights: list[float] | None = None, + household_extra: dict[str, object] | None = None, +) -> Frame: person = person.copy() n = len(person) household_ids = person["person_household_id"].to_numpy() @@ -149,9 +174,12 @@ def _us_frame(person: pd.DataFrame, *, weights: list[float] | None = None) -> Fr person["person_spm_unit_id"] = person["person_household_id"] + 2_000 person["person_family_id"] = person["person_household_id"] + 3_000 person["person_marital_unit_id"] = np.arange(n, dtype="int64") + 4_000 + household = pd.DataFrame({"household_id": unique_households}) + for column, values in (household_extra or {}).items(): + household[column] = values tables = { "person": person, - "household": pd.DataFrame({"household_id": unique_households}), + "household": household, "tax_unit": pd.DataFrame({"tax_unit_id": unique_households + 1_000}), "spm_unit": pd.DataFrame({"spm_unit_id": unique_households + 2_000}), "family": pd.DataFrame({"family_id": unique_households + 3_000}), @@ -178,7 +206,10 @@ def test_stage_spec_loads_and_declares_outputs() -> None: spec = us_scf_wealth_stage_spec() assert isinstance(spec, SourceStageSpec) assert spec.stage == US_SCF_WEALTH_STAGE_NAME - for column in US_SCF_FINANCIAL_ASSET_OUTPUT_COLUMNS: + for column in ( + *US_SCF_FINANCIAL_ASSET_OUTPUT_COLUMNS, + *US_SCF_NET_WORTH_OUTPUT_COLUMNS, + ): assert column in spec.outputs @@ -196,6 +227,30 @@ def test_stock_target_sums_stocks_and_nmmf() -> None: assert SCF_FINANCIAL_ASSET_TARGET_COMPONENTS["bond_assets"] == ("bond",) +def test_net_worth_target_is_the_direct_signed_scf_anchor() -> None: + assert SCF_NET_WORTH_TARGET_COMPONENTS == {"net_worth": ("networth",)} + assert US_SCF_NET_WORTH_OUTPUT_COLUMNS == ("net_worth",) + assert US_SCF_WEALTH_NONCONSTANT_HOUSEHOLD_COLUMNS == ("net_worth",) + + spec = us_scf_wealth_stage_spec() + derivation = next( + operation + for operation in spec.operations + if operation.kind == "derive" + and "net_worth_anchor" in operation.parameters["outputs"] + ) + assert derivation.parameters["net_worth_source_column"] == "networth" + fit = next( + operation + for operation in spec.operations + if operation.kind == "fit_weighted_qrf" + ) + assert fit.parameters["net_worth_target"] == "networth" + assert fit.parameters["net_worth_signed"] is True + assert "utils/asset_imputation.py lines 15-19" in spec.notes + assert "calibration/source_impute.py lines 1324-1338" in spec.notes + + # --------------------------------------------------------------------------- # # Donor loading # # --------------------------------------------------------------------------- # @@ -206,6 +261,7 @@ def test_load_donor_derives_targets_predictors_and_weight(tmp_path) -> None: donor = load_scf_2022_financial_asset_donor(path) for column in ( *US_SCF_FINANCIAL_ASSET_OUTPUT_COLUMNS, + *US_SCF_NET_WORTH_OUTPUT_COLUMNS, *SCF_WEALTH_PREDICTORS, _DONOR_WEIGHT_COLUMN, ): @@ -219,6 +275,10 @@ def test_load_donor_derives_targets_predictors_and_weight(tmp_path) -> None: assert (donor[_DONOR_WEIGHT_COLUMN] > 0).all() for column in US_SCF_FINANCIAL_ASSET_OUTPUT_COLUMNS: assert (donor[column] >= 0).all() + # Unlike the three SSI leaves, net worth is a signed direct carry. + np.testing.assert_allclose(donor["net_worth"], raw["networth"], rtol=1e-6) + assert (donor["net_worth"] < 0).any() + assert (donor["net_worth"] > 0).any() def test_load_donor_missing_column_raises(tmp_path) -> None: @@ -298,10 +358,29 @@ def test_impute_missing_donor_column_raises() -> None: impute_us_scf_financial_assets(person, donor, seed=0, n_estimators=10) +def test_impute_net_worth_is_signed_finite_and_household_aligned() -> None: + person = _person_rows(200) + household = _us_frame(person).table("household").iloc[::-1].copy() + result = impute_us_scf_net_worth( + person, + household, + _donor_table(), + seed=42, + n_estimators=20, + ) + + assert result.name == "net_worth" + assert result.index.equals(household.index) + assert np.isfinite(result.to_numpy()).all() + assert (result > 0).any() + assert (result < 0).any() + assert result.nunique() > 1 + + # --------------------------------------------------------------------------- # # Frame integration # # --------------------------------------------------------------------------- # -def test_with_inputs_writes_all_three_columns() -> None: +def test_with_inputs_writes_asset_and_net_worth_columns() -> None: frame = _us_frame(_person_rows(60)) donor = _donor_table() out = with_us_scf_wealth_inputs( @@ -312,6 +391,10 @@ def test_with_inputs_writes_all_three_columns() -> None: assert column in person.columns assert person[column].to_numpy().dtype == np.float64 assert person["bank_account_assets"].to_numpy().sum() > 0 + net_worth = out.table("household")["net_worth"] + assert net_worth.to_numpy().dtype == np.float64 + assert net_worth.nunique() > 1 + assert (net_worth < 0).any() def test_with_inputs_is_idempotent_when_signal_present() -> None: @@ -329,6 +412,10 @@ def test_with_inputs_is_idempotent_when_signal_present() -> None: once.table("person")["bank_account_assets"].to_numpy(), twice.table("person")["bank_account_assets"].to_numpy(), ) + np.testing.assert_array_equal( + once.table("household")["net_worth"].to_numpy(), + twice.table("household")["net_worth"].to_numpy(), + ) def test_with_inputs_reimputes_when_bank_assets_constant() -> None: @@ -344,6 +431,35 @@ def test_with_inputs_reimputes_when_bank_assets_constant() -> None: assert out.table("person")["bank_account_assets"].to_numpy().sum() > 0 +def test_with_inputs_heals_nonfinite_net_worth_without_redrawing_assets() -> None: + once = with_us_scf_wealth_inputs( + _us_frame(_person_rows(60)), + seed=42, + time_period=TIME_PERIOD, + scf_donor=_donor_table(), + ) + tables = {entity: once.table(entity).copy() for entity in once.entities} + original_assets = tables["person"]["bank_account_assets"].to_numpy().copy() + tables["household"].loc[tables["household"].index[0], "net_worth"] = np.inf + damaged = Frame( + tables, + once.schema, + {entity: once.weights_for(entity) for entity in once.weighted_entities}, + ) + + healed = with_us_scf_wealth_inputs( + damaged, + seed=99, + time_period=TIME_PERIOD, + scf_donor=_donor_table(), + ) + + np.testing.assert_array_equal( + healed.table("person")["bank_account_assets"].to_numpy(), original_assets + ) + assert np.isfinite(healed.table("household")["net_worth"]).all() + + # --------------------------------------------------------------------------- # # Gate + summary # # --------------------------------------------------------------------------- # @@ -369,12 +485,37 @@ def test_signal_gate_fails_on_constant_zero_surface() -> None: person["bank_account_assets"] = 0.0 person["stock_assets"] = 0.0 person["bond_assets"] = 0.0 - frame = _us_frame(person) + frame = _us_frame( + person, + household_extra={"net_worth": np.zeros(20, dtype=np.float64)}, + ) gate = us_scf_wealth_signal_gate(frame) assert not gate.passed assert any("constant" in f or "nonzero share" in f for f in gate.failures) +def test_signal_gate_requires_signed_net_worth_support() -> None: + out = with_us_scf_wealth_inputs( + _us_frame(_person_rows(200)), + seed=42, + time_period=TIME_PERIOD, + scf_donor=_donor_table(), + ) + tables = {entity: out.table(entity).copy() for entity in out.entities} + tables["household"]["net_worth"] = ( + np.abs(tables["household"]["net_worth"].to_numpy()) + 1.0 + ) + positive_only = Frame( + tables, + out.schema, + {entity: out.weights_for(entity) for entity in out.weighted_entities}, + ) + + gate = us_scf_wealth_signal_gate(positive_only) + assert not gate.passed + assert any("net_worth negative share" in failure for failure in gate.failures) + + def test_summary_reports_shares_and_bands() -> None: frame = _us_frame(_person_rows(200)) out = with_us_scf_wealth_inputs( @@ -384,6 +525,22 @@ def test_summary_reports_shares_and_bands() -> None: assert 0.0 <= summary["bank_account_assets_nonzero_share"] <= 1.0 assert "bank_nonzero_share_band" in summary assert summary["unique_counts"]["bank_account_assets"] >= 2 + assert summary["net_worth_unique_count"] >= 2 + assert 0.0 < summary["net_worth_negative_share"] < 1.0 + assert 0.0 < summary["net_worth_positive_share"] < 1.0 + + +@requires_us +def test_policyengine_1_764_6_net_worth_input_contract() -> None: + from policyengine_us import CountryTaxBenefitSystem + + assert version("policyengine-us") == "1.764.6" + variable = CountryTaxBenefitSystem().variables["net_worth"] + assert variable.is_input_variable() + assert variable.entity.key == "household" + assert variable.value_type is float + assert variable.default_value == 0 + assert str(variable.definition_period).lower() == "year" # --------------------------------------------------------------------------- # diff --git a/packages/populace-build/tests/test_us_sipp_head_start.py b/packages/populace-build/tests/test_us_sipp_head_start.py new file mode 100644 index 00000000..4c6fa625 --- /dev/null +++ b/packages/populace-build/tests/test_us_sipp_head_start.py @@ -0,0 +1,520 @@ +"""Strict measured-SIPP Head Start take-up tests.""" + +from __future__ import annotations + +import importlib.util +from importlib.metadata import version +from pathlib import Path + +import numpy as np +import pandas as pd +import pytest + +import populace.build.us_runtime.sipp_head_start as module +from populace.build.us_runtime.sipp_head_start import ( + HEAD_START_SIPP_DICTIONARY_URL, + SIPP_2023_HEAD_START_DONOR_REVISION, + SIPP_2023_HEAD_START_DONOR_SHA256, + SIPP_2023_HEAD_START_DONOR_SIZE_BYTES, + SIPP_2023_HEAD_START_DONOR_URL, + SIPP_HEAD_START_FIT_PARAMETERS, + SIPP_HEAD_START_MODEL_PREDICTORS, + SIPP_HEAD_START_READ_PARAMETERS, + SIPP_HEAD_START_SOURCE_COLUMNS, + US_SIPP_HEAD_START_NONCONSTANT_PERSON_COLUMNS, + US_SIPP_HEAD_START_OUTPUT_COLUMNS, + US_SIPP_HEAD_START_REQUIRED_SOURCE_COLUMNS, + US_SIPP_HEAD_START_STAGE_NAME, + impute_us_sipp_head_start, + load_sipp_2023_head_start_donor, + us_sipp_head_start_signal_gate, + us_sipp_head_start_summary, + with_us_sipp_head_start_input, +) +from populace.frame import US_SCHEMA, Frame, WeightKind, Weights + +_OUTPUT = US_SIPP_HEAD_START_OUTPUT_COLUMNS[0] +_policyengine_us_installed = importlib.util.find_spec("policyengine_us") is not None +requires_us = pytest.mark.skipif( + not _policyengine_us_installed, + reason="requires the policyengine-us [us] extra", +) + + +def _source_row( + ssuid: str, + pnum: int, + *, + month: int = 12, + age: int = 4, + head_start_status: int = 1, + head_start_answer: float = 2.0, + screen_status: int = 1, + screen: float = 1.0, + grade_status: int = 1, + grade: float = 21.0, + end_month_status: int = 1, + end_month: float = 12.0, + weight: float = 100.0, +) -> dict[str, object]: + row: dict[str, object] = {column: 0.0 for column in SIPP_HEAD_START_SOURCE_COLUMNS} + row.update( + { + "SSUID": ssuid, + "PNUM": pnum, + "MONTHCODE": month, + "WPFINWGT": weight, + "TAGE": age, + "ESEX": 2 if pnum % 2 else 1, + "EED_SCRNR": screen, + "AED_SCRNR": screen_status, + "EEDGRADE": grade, + "AEDGRADE": grade_status, + "EEDEMONTH": end_month, + "AEDMONTH": end_month_status, + "EEDHEADST": head_start_answer, + "AEDHEADST": head_start_status, + } + ) + return row + + +def _write_source(tmp_path: Path, rows: list[dict[str, object]]) -> Path: + path = tmp_path / "pu2023.csv" + pd.DataFrame(rows, columns=SIPP_HEAD_START_SOURCE_COLUMNS).to_csv( + path, sep="|", index=False + ) + return path + + +def _donor() -> pd.DataFrame: + return pd.DataFrame( + { + "age": [3.0, 4.0, 5.0, 4.0], + "is_female": [0.0, 1.0, 0.0, 1.0], + "household_size": [2.0, 3.0, 4.0, 5.0], + "count_under_18": [1.0, 2.0, 3.0, 4.0], + "count_under_6": [1.0, 1.0, 2.0, 2.0], + "household_employment_income": [0.0, 10_000.0, 20_000.0, 30_000.0], + _OUTPUT: [False, True, False, True], + "sipp_weight": [1.0, 2.0, 3.0, 4.0], + } + ) + + +def _frame( + source_ids: list[int], + *, + ages: list[int] | None = None, + female: list[bool] | None = None, + channels: list[str] | None = None, + output: list[object] | None = None, +) -> Frame: + n = len(source_ids) + ids = np.arange(1, n + 1, dtype=np.int64) + if ages is None: + ages = [4] * n + if female is None: + female = [i % 2 == 0 for i in range(n)] + person = pd.DataFrame( + { + "person_id": ids, + "person_household_id": ids, + "person_tax_unit_id": ids + 100, + "person_spm_unit_id": ids + 200, + "person_family_id": ids + 300, + "person_marital_unit_id": ids + 400, + "person_source_id": source_ids, + "age": ages, + "is_female": female, + "employment_income_before_lsr": np.arange(n, dtype=np.float64) * 100.0, + } + ) + if channels is not None: + person["person_support_channel"] = channels + if output is not None: + person[_OUTPUT] = output + tables = { + "person": person, + "household": pd.DataFrame({"household_id": ids}), + "tax_unit": pd.DataFrame({"tax_unit_id": ids + 100}), + "spm_unit": pd.DataFrame({"spm_unit_id": ids + 200}), + "family": pd.DataFrame({"family_id": ids + 300}), + "marital_unit": pd.DataFrame({"marital_unit_id": ids + 400}), + } + return Frame( + tables, + US_SCHEMA, + {"household": Weights(np.ones(n), WeightKind.DESIGN)}, + ) + + +def _replace_person(frame: Frame, person: pd.DataFrame) -> Frame: + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + tables["person"] = person + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + ) + + +class _FakeQRF: + instances: list[_FakeQRF] = [] + + def __init__(self, *, n_estimators: int, seed: int) -> None: + self.n_estimators = n_estimators + self.seed = seed + self.weights: object = None + self.training: pd.DataFrame | None = None + self.receiver: pd.DataFrame | None = None + self.__class__.instances.append(self) + + def fit( + self, + training: pd.DataFrame, + *, + predictors: list[str], + targets: list[str], + weights: object, + ) -> _FakeQRF: + assert predictors == list(SIPP_HEAD_START_MODEL_PREDICTORS) + assert targets == [_OUTPUT] + self.training = training.copy() + self.weights = weights + return self + + def predict(self, receiver: pd.DataFrame) -> pd.DataFrame: + self.receiver = receiver.copy() + return pd.DataFrame( + {_OUTPUT: receiver["is_female"].to_numpy(dtype=bool)}, + index=receiver.index, + ) + + +@pytest.fixture(autouse=True) +def _clear_fake() -> None: + _FakeQRF.instances.clear() + + +def test_source_coordinates_and_operation_contract_are_exact() -> None: + assert US_SIPP_HEAD_START_STAGE_NAME == "sipp_head_start" + assert US_SIPP_HEAD_START_OUTPUT_COLUMNS == ("takes_up_head_start_if_eligible",) + assert US_SIPP_HEAD_START_NONCONSTANT_PERSON_COLUMNS == ( + "takes_up_head_start_if_eligible", + ) + assert US_SIPP_HEAD_START_REQUIRED_SOURCE_COLUMNS == ( + "person_source_id", + "person_household_id", + "age", + "is_female", + "employment_income_before_lsr", + ) + assert SIPP_2023_HEAD_START_DONOR_REVISION == ( + "21280dca5995e978d706740a8a4b9b7860cfd7b6" + ) + assert SIPP_2023_HEAD_START_DONOR_SHA256 == ( + "5c30439e365fc26483318ef61d1d8f4bb2f0e9d6bb47c22c06756a7698733ee2" + ) + assert SIPP_2023_HEAD_START_DONOR_SIZE_BYTES == 3_726_010_471 + assert SIPP_2023_HEAD_START_DONOR_REVISION in SIPP_2023_HEAD_START_DONOR_URL + assert HEAD_START_SIPP_DICTIONARY_URL.endswith("2023/2023_SIPP_Data_Dictionary.pdf") + assert SIPP_HEAD_START_READ_PARAMETERS == { + "table": "sipp_person", + "delimiter": "|", + "month_column": "MONTHCODE", + "month": 12, + "source_columns": list(SIPP_HEAD_START_SOURCE_COLUMNS), + } + assert SIPP_HEAD_START_FIT_PARAMETERS["age_domain"] == [3, 5] + assert SIPP_HEAD_START_FIT_PARAMETERS["assignment_unit"] == "person_source_id" + assert SIPP_HEAD_START_FIT_PARAMETERS["fan_to_support_clones"] is True + assert SIPP_HEAD_START_FIT_PARAMETERS["seed_from_build_config"] is True + assert "rate" not in " ".join(map(str, SIPP_HEAD_START_FIT_PARAMETERS.values())) + + +def test_loader_keeps_only_strict_reported_labels(tmp_path: Path) -> None: + rows = [ + _source_row("yes", 1, head_start_answer=1), + _source_row("no", 1, head_start_answer=2), + _source_row( + "not_enrolled", + 1, + head_start_status=0, + head_start_answer=np.nan, + screen_status=1, + screen=2, + grade_status=0, + grade=np.nan, + ), + _source_row( + "other_grade", + 1, + head_start_status=0, + head_start_answer=np.nan, + screen_status=1, + screen=1, + grade_status=1, + grade=22, + ), + _source_row("hot_deck_head_start", 1, head_start_status=2, head_start_answer=1), + _source_row( + "unknown_screen", + 1, + head_start_status=0, + head_start_answer=np.nan, + screen_status=4, + screen=2, + ), + _source_row( + "hot_deck_grade", + 1, + head_start_status=0, + head_start_answer=np.nan, + screen_status=1, + screen=1, + grade_status=2, + grade=22, + ), + _source_row("age_two", 1, age=2, head_start_answer=1), + _source_row("month_eleven", 1, month=11, head_start_answer=1), + ] + donor = load_sipp_2023_head_start_donor( + _write_source(tmp_path, rows), + expected_sha256=None, + expected_size_bytes=None, + chunksize=3, + ) + + assert donor[_OUTPUT].tolist() == [True, False, False, False] + audit = donor.attrs["source_audit"] + assert audit["raw_rows"] == 9 + assert audit["december_rows"] == 8 + assert audit["age_domain_rows"] == 7 + assert audit["training_rows"] == 4 + assert audit["positive_rows"] == 1 + assert audit["direct_response_rows"] == 2 + assert audit["reported_no_enrollment_rows"] == 1 + assert audit["reported_other_grade_rows"] == 1 + assert audit["pinned_transform"] is False + + +def test_loader_refuses_missing_upstream_status(tmp_path: Path) -> None: + path = _write_source(tmp_path, [_source_row("one", 1)]) + source = pd.read_csv(path, sep="|").drop(columns=["AED_SCRNR"]) + source.to_csv(path, sep="|", index=False) + + with pytest.raises(ValueError, match="AED_SCRNR"): + load_sipp_2023_head_start_donor( + path, + expected_sha256=None, + expected_size_bytes=None, + ) + + +def test_pinned_full_file_audit_constants_are_exact() -> None: + # The 740 negatives are the strict upstream-observed set. A looser + # status-only mask has 743, but the extra three have no reported screening + # fact and therefore cannot be called measured structural negatives. + assert module._PINNED_RAW_ROWS == 476_744 + assert module._PINNED_DECEMBER_ROWS == 39_513 + assert module._PINNED_AGE_DOMAIN_ROWS == 1_177 + assert module._PINNED_TRAINING_ROWS == 785 + assert module._PINNED_POSITIVE_ROWS == 45 + assert module._PINNED_NEGATIVE_ROWS == 740 + assert module._PINNED_DIRECT_RESPONSE_ROWS == 215 + assert module._PINNED_REPORTED_NO_ENROLLMENT_ROWS == 440 + assert module._PINNED_REPORTED_OTHER_GRADE_ROWS == 130 + assert module._PINNED_WEIGHT_SUM == pytest.approx(7_978_494.5412483) + assert module._PINNED_POSITIVE_WEIGHT_SUM == pytest.approx(491_970.1041311) + assert module._PINNED_WEIGHTED_TRUE_SHARE == pytest.approx(0.06166202177461505) + + +def test_imputer_weights_qrf_and_fans_one_asec_decision_to_source_clones( + monkeypatch: pytest.MonkeyPatch, +) -> None: + monkeypatch.setattr(module, "QRF", _FakeQRF) + frame = _frame( + [10, 10, 20, 30, 30, 40, 40], + ages=[4, 4, 5, 3, 3, 40, 40], + # Source 10 deliberately disagrees: canonical ASEC must win. Source 20 + # is PUF-only and remains supported. + female=[True, False, False, True, False, True, False], + channels=[ + "asec", + "puf_tax_detail", + "puf_tax_detail", + "asec", + "puf_tax_detail", + "asec", + "puf_tax_detail", + ], + ) + + first = impute_us_sipp_head_start(frame, _donor(), seed=91) + second = impute_us_sipp_head_start(frame, _donor(), seed=91) + + assert _FakeQRF.instances[0].n_estimators == 100 + assert _FakeQRF.instances[0].seed == 91 + assert _FakeQRF.instances[0].weights == "sipp_weight" + np.testing.assert_array_equal(first, second) + values = pd.DataFrame( + { + "source": frame.table("person")["person_source_id"], + "value": first, + } + ) + assert values.groupby("source")["value"].nunique().max() == 1 + assert values[values["source"] == 10]["value"].all() + assert not values[values["source"] == 20]["value"].any() + assert values[values["source"] == 30]["value"].all() + # An off-domain adult is always false even if the QRF would draw true. + assert not values[values["source"] == 40]["value"].any() + + +def test_imputer_fails_closed_on_provenance_and_clone_age( + monkeypatch: pytest.MonkeyPatch, +) -> None: + monkeypatch.setattr(module, "QRF", _FakeQRF) + frame = _frame( + [1, 1], + ages=[4, 4], + channels=["asec", "puf_tax_detail"], + ) + person = frame.table("person").drop(columns=["person_source_id"]) + missing = _replace_person(frame, person) + with pytest.raises(ValueError, match="person_source_id"): + impute_us_sipp_head_start(missing, _donor(), seed=0) + + person = frame.table("person").copy() + person.loc[1, "person_support_channel"] = "mystery" + unknown = _replace_person(frame, person) + with pytest.raises(ValueError, match="unsupported support channel"): + impute_us_sipp_head_start(unknown, _donor(), seed=0) + + person = frame.table("person").copy() + person.loc[1, "age"] = 5 + inconsistent = _replace_person(frame, person) + with pytest.raises(ValueError, match="disagree on age"): + impute_us_sipp_head_start(inconsistent, _donor(), seed=0) + + +def test_wrapper_heals_stale_output_and_is_exactly_idempotent( + monkeypatch: pytest.MonkeyPatch, +) -> None: + monkeypatch.setattr(module, "QRF", _FakeQRF) + monkeypatch.setattr(module, "us_sipp_head_start_stage_spec", lambda: None) + base = _frame( + [1, 2, 3, 4], + ages=[3, 4, 5, 40], + female=[True, False, True, True], + output=[True, True, True, True], + ) + + healed = with_us_sipp_head_start_input( + base, + seed=5, + time_period=2024, + sipp_donor=_donor(), + ) + twice = with_us_sipp_head_start_input( + healed, + seed=5, + time_period=2024, + sipp_donor=_donor(), + ) + + assert healed.table("person")[_OUTPUT].tolist() == [True, False, True, False] + assert twice is healed + + +def test_summary_and_gate_require_nonconstant_clone_consistent_domain_signal() -> None: + sources = [source for source in range(20) for _ in range(2)] + channels = [channel for _ in range(20) for channel in ("asec", "puf_tax_detail")] + values = [source == 0 for source in range(20) for _ in range(2)] + healthy = _frame(sources, channels=channels, output=values) + + summary = us_sipp_head_start_summary(healthy) + gate = us_sipp_head_start_signal_gate(healthy) + assert summary["eligible_weighted_take_up_share"] == pytest.approx(0.05) + assert summary["clone_mismatch_count"] == 0 + assert gate.passed, gate.failures + + person = healthy.table("person").copy() + person.loc[1, _OUTPUT] = False + mismatch = _replace_person(healthy, person) + mismatch_gate = us_sipp_head_start_signal_gate(mismatch) + assert not mismatch_gate.passed + assert any("clone" in failure for failure in mismatch_gate.failures) + + person = healthy.table("person").copy() + person.loc[2, "age"] = 40 + person.loc[2, _OUTPUT] = True + outside = _replace_person(healthy, person) + outside_gate = us_sipp_head_start_signal_gate(outside) + assert not outside_gate.passed + assert any("outside" in failure for failure in outside_gate.failures) + + person = healthy.table("person").copy() + person[_OUTPUT] = False + constant = _replace_person(healthy, person) + constant_gate = us_sipp_head_start_signal_gate(constant) + assert not constant_gate.passed + assert any("constant" in failure for failure in constant_gate.failures) + + person = healthy.table("person").drop(columns=["person_source_id"]) + no_provenance = _replace_person(healthy, person) + provenance_gate = us_sipp_head_start_signal_gate(no_provenance) + assert not provenance_gate.passed + assert any("provenance" in failure for failure in provenance_gate.failures) + + person = healthy.table("person").copy() + person.loc[0, "person_support_channel"] = "mystery" + bad_channel = _replace_person(healthy, person) + channel_gate = us_sipp_head_start_signal_gate(bad_channel) + assert not channel_gate.passed + assert any("support-channel" in failure for failure in channel_gate.failures) + + +@requires_us +def test_policyengine_us_1_764_6_head_start_is_zero_when_neutralized() -> None: + from policyengine_us import CountryTaxBenefitSystem, Simulation + + assert version("policyengine-us") == "1.764.6" + variable = CountryTaxBenefitSystem().variables[_OUTPUT] + assert variable.is_input_variable() + assert variable.entity.key == "person" + assert variable.value_type is bool + assert variable.default_value is True + + def situation(takes_up: bool) -> dict[str, object]: + return { + "people": { + "child": { + "age": {"2024": 4}, + _OUTPUT: {"2024": takes_up}, + } + }, + "tax_units": { + "unit": { + "members": ["child"], + "filing_status": {"2024": "SINGLE"}, + } + }, + "families": {"family": {"members": ["child"]}}, + "spm_units": {"spm": {"members": ["child"]}}, + "households": { + "household": { + "members": ["child"], + "state_code": {"2024": "CA"}, + } + }, + "marital_units": {"marital": {"members": ["child"]}}, + } + + active = Simulation(situation=situation(True)) + neutralized = Simulation(situation=situation(False)) + assert active.calculate("head_start", "2024")[0] > 0.0 + assert neutralized.calculate("head_start", "2024")[0] == 0.0 diff --git a/packages/populace-build/tests/test_us_sipp_tips.py b/packages/populace-build/tests/test_us_sipp_tips.py new file mode 100644 index 00000000..1d86eb38 --- /dev/null +++ b/packages/populace-build/tests/test_us_sipp_tips.py @@ -0,0 +1,531 @@ +"""US SIPP tip-income and tipped-occupation source-stage tests. + +The retired eCPS pipeline annualized observed-source December SIPP monthly tip +amounts while excluding Census allocation-flag columns from the dollar sum, +and QRF-imputed the resulting person-level income with Treasury-listed +occupation as a predictor rather than a domain mask. +These tests pin that source transformation and the release-stage healing and +signal contracts. +""" + +from __future__ import annotations + +import hashlib + +import numpy as np +import pandas as pd +import pytest + +from populace.build.source_manifest import SourceStageSpec +from populace.build.us_runtime import ( + SIPP_2023_TIP_DONOR_REVISION, + SIPP_2023_TIP_DONOR_SHA256, + SIPP_2023_TIP_DONOR_URL, + SIPP_TIP_OUTPUT_COLUMNS, + SIPP_TIP_PREDICTORS, + derive_treasury_tipped_occupation_code, + fetch_sipp_2023_tip_donor, + impute_us_sipp_tips, + load_sipp_2023_tip_donor, + us_sipp_tips_signal_gate, + us_sipp_tips_stage_spec, + us_sipp_tips_summary, + with_us_sipp_tip_inputs, +) +from populace.build.us_runtime import sipp_tips as module +from populace.frame import US_SCHEMA, Frame, WeightKind, Weights + +TIME_PERIOD = 2024 + +_DONOR_WEIGHT_COLUMN = "sipp_weight" +_EXPECTED_DONOR_SHA256 = ( + "1f0bcb8e045ef1118e8eba4b4a2997bdaaf947bd0dd09d41fa7c7d5657a3d7d5" +) + + +def _raw_sipp_row( + *, + ssuid: int, + month: int, + age: int, + monthly_income: float, + tips: tuple[float, ...] = (), + occupations: tuple[int, ...] = (), + allocation_flags: tuple[int, ...] = (), + weight: float = 1.0, +) -> dict[str, float | int]: + """Return one ``pu2023_slim.csv``-shaped synthetic person-month.""" + + row: dict[str, float | int] = { + "SSUID": ssuid, + "MONTHCODE": month, + "WPFINWGT": weight, + "TAGE": age, + "TPTOTINC": monthly_income, + # A distractor the retired broad ``contains('TXAMT')`` selector would + # have summed. The port must read only the seven TJB amount fields. + "SOME_TXAMT_OTHER": 50_000.0, + } + for job in range(1, 8): + position = job - 1 + row[f"TJB{job}_TXAMT"] = tips[position] if position < len(tips) else 0.0 + row[f"AJB{job}_TXAMT"] = ( + allocation_flags[position] if position < len(allocation_flags) else 0 + ) + row[f"TJB{job}_OCC"] = ( + occupations[position] if position < len(occupations) else 9999 + ) + return row + + +def _raw_sipp() -> pd.DataFrame: + """Small panel with December, non-December, and allocated records.""" + + return pd.DataFrame( + [ + # The same adult's November record must not enter the annual donor. + _raw_sipp_row( + ssuid=1, + month=11, + age=35, + monthly_income=2_000.0, + tips=(9_999.0,), + occupations=(4040,), + ), + # $100 + $50 per month -> exactly $1,800 annually. + _raw_sipp_row( + ssuid=1, + month=12, + age=35, + monthly_income=2_000.0, + tips=(100.0, 50.0), + occupations=(4040, 9999), + weight=2.0, + ), + # A child in the adult's December household pins composition counts. + _raw_sipp_row( + ssuid=1, + month=12, + age=5, + monthly_income=0.0, + ), + # This allocated child is not a training target, but still counts + # in the retained adult's December household composition. + _raw_sipp_row( + ssuid=1, + month=12, + age=3, + monthly_income=0.0, + allocation_flags=(2,), + ), + # SIPP status 2 is an allocated/imputed source amount. It must be + # excluded from training, not added to the dollar amount. Statuses + # 0, 1, and 9 are observed/derivable in the retired source-quality + # contract. + _raw_sipp_row( + ssuid=2, + month=12, + age=40, + monthly_income=4_000.0, + tips=(25_000.0,), + occupations=(4110,), + allocation_flags=(2,), + ), + ] + ) + + +def _donor_table(n: int = 180) -> pd.DataFrame: + """Fast, varied donor table for QRF plumbing tests.""" + + rng = np.random.default_rng(1) + tipped = np.arange(n) % 3 == 0 + donor = pd.DataFrame( + { + "employment_income": rng.gamma(2.0, 18_000.0, n), + "age": rng.integers(18, 80, n).astype(np.float64), + "count_under_18": rng.integers(0, 4, n).astype(np.float64), + "count_under_6": rng.integers(0, 2, n).astype(np.float64), + "is_tipped_occupation": tipped.astype(np.float64), + "tip_income": np.where(tipped, rng.gamma(2.0, 1_200.0, n) + 100.0, 0.0), + "treasury_tipped_occupation_code": np.where(tipped, 101, 0), + _DONOR_WEIGHT_COLUMN: rng.uniform(0.5, 3.0, n), + } + ) + return donor + + +def _person_rows(n: int = 120) -> pd.DataFrame: + """ASEC-shaped recipients with a mix of listed and unlisted occupations.""" + + rng = np.random.default_rng(2) + tipped = np.arange(n) % 10 == 0 + return pd.DataFrame( + { + "person_id": np.arange(1, n + 1, dtype=np.int64), + "person_household_id": np.arange(1, n + 1, dtype=np.int64), + "employment_income_before_lsr": rng.gamma(2.0, 18_000.0, n), + "age": rng.integers(20, 70, n).astype(np.float64), + "PEIOOCC": np.where(tipped, 4040, 9999), + } + ) + + +def _us_frame(person: pd.DataFrame, *, weights: np.ndarray | None = None) -> Frame: + person = person.copy() + n = len(person) + household_ids = person["person_household_id"].to_numpy(dtype=np.int64) + unique_households = np.unique(household_ids) + person["person_tax_unit_id"] = household_ids + 1_000 + person["person_spm_unit_id"] = household_ids + 2_000 + person["person_family_id"] = household_ids + 3_000 + person["person_marital_unit_id"] = np.arange(n, dtype=np.int64) + 4_000 + tables = { + "person": person, + "household": pd.DataFrame({"household_id": unique_households}), + "tax_unit": pd.DataFrame({"tax_unit_id": unique_households + 1_000}), + "spm_unit": pd.DataFrame({"spm_unit_id": unique_households + 2_000}), + "family": pd.DataFrame({"family_id": unique_households + 3_000}), + "marital_unit": pd.DataFrame( + {"marital_unit_id": np.arange(n, dtype=np.int64) + 4_000} + ), + } + household_weights = ( + np.ones(len(unique_households), dtype=np.float64) + if weights is None + else np.asarray(weights, dtype=np.float64) + ) + return Frame( + tables, + US_SCHEMA, + { + "household": Weights( + values=household_weights, + kind=WeightKind.DESIGN, + ) + }, + ) + + +def _frame_with_tip_surface( + *, + n: int = 1_000, + tipped_count: int = 70, + positive_tip_count: int = 8, +) -> Frame: + person = _person_rows(n) + codes = np.zeros(n, dtype=np.int16) + codes[:tipped_count] = 101 + tips = np.zeros(n, dtype=np.float64) + tips[:positive_tip_count] = np.arange(1, positive_tip_count + 1) * 250.0 + person["treasury_tipped_occupation_code"] = codes + person["tip_income"] = tips + return _us_frame(person) + + +# --------------------------------------------------------------------------- # +# Source declaration and pinned artifact +# --------------------------------------------------------------------------- # +def test_stage_spec_loads_and_declares_both_outputs() -> None: + spec = us_sipp_tips_stage_spec() + assert isinstance(spec, SourceStageSpec) + assert spec.stage == "sipp_tips" + assert SIPP_TIP_OUTPUT_COLUMNS == ( + "tip_income", + "treasury_tipped_occupation_code", + ) + assert set(SIPP_TIP_OUTPUT_COLUMNS) <= set(spec.outputs) + assert len(spec.artifacts) == 1 + artifact = spec.artifacts[0] + assert artifact["sha256"] == SIPP_2023_TIP_DONOR_SHA256 + assert SIPP_2023_TIP_DONOR_REVISION in artifact["locator"] + tip_model = next( + operation + for operation in spec.operations + if operation.kind == "fit_tip_income_model" + ) + assert tuple(tip_model.parameters["predictors"]) == SIPP_TIP_PREDICTORS + assert tuple(tip_model.parameters["source_columns"]) == ( + module._SIPP_TIP_AMOUNT_COLUMNS + ) + assert tuple(tip_model.parameters["allocation_flag_columns"]) == ( + module._SIPP_TIP_ALLOCATION_COLUMNS + ) + assert tip_model.parameters["observed_status_values"] == [0, 1, 9] + assert not any(operation.kind == "zero_when_false" for operation in spec.operations) + + +def test_predictors_match_retired_sipp_tip_model() -> None: + assert SIPP_TIP_PREDICTORS == ( + "employment_income", + "age", + "count_under_18", + "count_under_6", + "is_tipped_occupation", + ) + + +def test_sipp_donor_coordinate_and_sha_are_pinned() -> None: + assert SIPP_2023_TIP_DONOR_URL.endswith("/pu2023_slim.csv") + assert "/resolve/" in SIPP_2023_TIP_DONOR_URL + assert SIPP_2023_TIP_DONOR_SHA256 == _EXPECTED_DONOR_SHA256 + + +# --------------------------------------------------------------------------- # +# Donor loading and direct occupation derivation +# --------------------------------------------------------------------------- # +def test_load_donor_uses_december_exact_tip_sum_and_observed_rows(tmp_path) -> None: + path = tmp_path / "pu2023_slim.csv" + _raw_sipp().to_csv(path, index=False) + + donor = load_sipp_2023_tip_donor(path) + + # November and the allocated December record are excluded. + assert sorted(donor["age"].tolist()) == [5.0, 35.0] + adult = donor.loc[donor["age"] == 35].iloc[0] + assert adult["tip_income"] == pytest.approx((100.0 + 50.0) * 12.0) + assert adult["employment_income"] == pytest.approx(2_000.0 * 12.0) + assert adult["treasury_tipped_occupation_code"] == 101 + assert adult["count_under_18"] == 2 + assert adult["count_under_6"] == 2 + assert adult[_DONOR_WEIGHT_COLUMN] == pytest.approx(2.0) + + +def test_load_donor_requires_allocation_flags(tmp_path) -> None: + raw = _raw_sipp().drop(columns=["AJB7_TXAMT"]) + path = tmp_path / "pu2023_slim.csv" + raw.to_csv(path, index=False) + with pytest.raises(ValueError, match="missing required column"): + load_sipp_2023_tip_donor(path) + + +def test_load_donor_verifies_requested_sha(tmp_path) -> None: + path = tmp_path / "pu2023_slim.csv" + _raw_sipp().to_csv(path, index=False) + digest = hashlib.sha256(path.read_bytes()).hexdigest() + + donor = load_sipp_2023_tip_donor(path, expected_sha256=digest) + assert len(donor) == 2 + with pytest.raises(ValueError, match="sha-256 verification"): + load_sipp_2023_tip_donor(path, expected_sha256="0" * 64) + + +def test_load_donor_retains_allowed_statuses_one_and_nine(tmp_path) -> None: + raw = pd.DataFrame( + [ + _raw_sipp_row( + ssuid=1, + month=12, + age=31, + monthly_income=2_500.0, + tips=(10.0,), + allocation_flags=(1,), + ), + _raw_sipp_row( + ssuid=2, + month=12, + age=42, + monthly_income=3_500.0, + tips=(20.0,), + allocation_flags=(9,), + ), + ] + ) + path = tmp_path / "pu2023_slim.csv" + raw.to_csv(path, index=False) + + donor = load_sipp_2023_tip_donor(path) + + assert sorted(donor["age"].tolist()) == [31.0, 42.0] + assert sorted(donor["tip_income"].tolist()) == [120.0, 240.0] + + +def test_treasury_tipped_occupation_mapping() -> None: + derived = derive_treasury_tipped_occupation_code( + np.array([4040, 4110, 4230, 2770, -1, 9999, np.nan]) + ) + np.testing.assert_array_equal(derived, [101, 102, 304, 208, 0, 0, 0]) + + +# --------------------------------------------------------------------------- # +# QRF imputation +# --------------------------------------------------------------------------- # +def test_impute_is_deterministic_for_a_seed() -> None: + person = _person_rows(90) + donor = _donor_table() + first = impute_us_sipp_tips(person, donor, seed=7, n_estimators=8) + second = impute_us_sipp_tips(person, donor, seed=7, n_estimators=8) + pd.testing.assert_frame_equal(first, second) + + +def test_impute_caps_large_donor_with_retired_named_seed(monkeypatch) -> None: + captured: dict[str, np.ndarray] = {} + + class FakeFittedQRF: + def predict(self, features: pd.DataFrame) -> pd.DataFrame: + return pd.DataFrame( + {"tip_income": np.zeros(len(features), dtype=np.float64)}, + index=features.index, + ) + + class FakeQRF: + def __init__(self, **kwargs) -> None: + pass + + def fit(self, frame: pd.DataFrame, **kwargs) -> FakeFittedQRF: + captured["row_ids"] = frame["employment_income"].to_numpy() + return FakeFittedQRF() + + monkeypatch.setattr("populace.fit.QRF", FakeQRF) + donor = _donor_table(10_050) + donor["employment_income"] = np.arange(len(donor), dtype=np.float64) + + impute_us_sipp_tips(_person_rows(20), donor, seed=7, n_estimators=3) + + expected_positions = np.sort( + np.random.default_rng(5_559_651_045_748_063_828).choice( + len(donor), size=10_000, replace=False + ) + ) + np.testing.assert_array_equal(captured["row_ids"], expected_positions) + + +def test_impute_carries_code_without_erasing_tips_outside_list(monkeypatch) -> None: + class FakeFittedQRF: + def predict(self, features: pd.DataFrame) -> pd.DataFrame: + return pd.DataFrame( + {"tip_income": np.arange(len(features), dtype=np.float64) + 100.0}, + index=features.index, + ) + + class FakeQRF: + def __init__(self, **kwargs) -> None: + pass + + def fit(self, *args, **kwargs) -> FakeFittedQRF: + return FakeFittedQRF() + + monkeypatch.setattr("populace.fit.QRF", FakeQRF) + person = _person_rows(90) + result = impute_us_sipp_tips( + person, + _donor_table(), + seed=42, + n_estimators=8, + ) + assert tuple(result.columns) == SIPP_TIP_OUTPUT_COLUMNS + expected_codes = derive_treasury_tipped_occupation_code(person["PEIOOCC"]) + np.testing.assert_array_equal( + result["treasury_tipped_occupation_code"], expected_codes + ) + tips = result["tip_income"].to_numpy() + assert (tips >= 0).all() + assert np.all(tips[expected_codes == 0] > 0.0) + assert tips[expected_codes > 0].sum() > 0.0 + + +def test_impute_missing_donor_column_raises() -> None: + with pytest.raises(ValueError, match="donor table missing column"): + impute_us_sipp_tips( + _person_rows(20), + _donor_table().drop(columns=["tip_income"]), + seed=0, + n_estimators=5, + ) + + +# --------------------------------------------------------------------------- # +# Frame integration and healing +# --------------------------------------------------------------------------- # +def test_with_inputs_heals_default_surface_then_is_idempotent() -> None: + person = _person_rows(90) + person["tip_income"] = 0.0 + person["treasury_tipped_occupation_code"] = 0 + frame = _us_frame(person) + donor = _donor_table() + + healed = with_us_sipp_tip_inputs( + frame, + seed=42, + time_period=TIME_PERIOD, + sipp_donor=donor, + ) + healed_person = healed.table("person") + assert healed_person["tip_income"].sum() > 0.0 + assert (healed_person["treasury_tipped_occupation_code"] > 0).any() + + repeated = with_us_sipp_tip_inputs( + healed, + seed=99, + time_period=TIME_PERIOD, + sipp_donor=donor, + ) + for column in SIPP_TIP_OUTPUT_COLUMNS: + np.testing.assert_array_equal( + healed_person[column].to_numpy(), + repeated.table("person")[column].to_numpy(), + ) + + +# --------------------------------------------------------------------------- # +# Gate and summary +# --------------------------------------------------------------------------- # +def test_signal_gate_passes_plausible_surface() -> None: + gate = us_sipp_tips_signal_gate(_frame_with_tip_surface()) + assert gate.passed, gate.failures + + +def test_signal_gate_fails_when_columns_missing() -> None: + gate = us_sipp_tips_signal_gate(_us_frame(_person_rows(20))) + assert not gate.passed + assert any("missing" in failure for failure in gate.failures) + + +def test_signal_gate_fails_on_constant_default_surface() -> None: + frame = _frame_with_tip_surface(tipped_count=0, positive_tip_count=0) + gate = us_sipp_tips_signal_gate(frame) + assert not gate.passed + assert any("constant" in failure for failure in gate.failures) + + +def test_signal_gate_fails_implausible_nonzero_shares() -> None: + frame = _frame_with_tip_surface( + n=100, + tipped_count=50, + positive_tip_count=40, + ) + gate = us_sipp_tips_signal_gate(frame) + assert not gate.passed + assert any("tip-income share" in failure for failure in gate.failures) + assert any("tipped-occupation share" in failure for failure in gate.failures) + + +def test_summary_reports_weighted_shares_bands_and_unique_counts() -> None: + summary = us_sipp_tips_summary(_frame_with_tip_surface()) + assert summary["tip_income_nonzero_share"] == pytest.approx(0.008) + assert summary["tipped_occupation_share"] == pytest.approx(0.07) + assert summary["tip_income_nonzero_share_band"] == [0.001, 0.03] + assert summary["tipped_occupation_share_band"] == [0.02, 0.15] + assert summary["unique_counts"]["tip_income"] > 1 + assert summary["unique_counts"]["treasury_tipped_occupation_code"] > 1 + + +# --------------------------------------------------------------------------- # +# SHA-verified provisioning cache +# --------------------------------------------------------------------------- # +def test_fetch_reuses_cache_only_when_sha_matches(tmp_path, monkeypatch) -> None: + payload = b"synthetic pinned SIPP donor" + cached = tmp_path / "pu2023_slim.csv" + cached.write_bytes(payload) + digest = hashlib.sha256(payload).hexdigest() + + def unexpected_network_call(*args, **kwargs): + raise AssertionError("matching cache entry must not hit the network") + + monkeypatch.setattr("urllib.request.urlopen", unexpected_network_call) + result = fetch_sipp_2023_tip_donor( + cache_dir=tmp_path, + expected_sha256=digest, + ) + assert result == cached + assert result.read_bytes() == payload diff --git a/packages/populace-build/tests/test_us_sipp_vehicles.py b/packages/populace-build/tests/test_us_sipp_vehicles.py new file mode 100644 index 00000000..5b338abc --- /dev/null +++ b/packages/populace-build/tests/test_us_sipp_vehicles.py @@ -0,0 +1,649 @@ +"""Tests for the SHA-pinned SIPP household-vehicle source stage.""" + +from __future__ import annotations + +import hashlib +import io +import urllib.request +from pathlib import Path + +import numpy as np +import pandas as pd +import pytest +from sklearn.ensemble import RandomForestClassifier + +from populace.build.us_runtime.sipp_vehicles import ( + _OWNED_OBSERVED_COLUMN, + _VALUE_OBSERVED_COLUMN, + ARCHIVED_SIPP_VEHICLE_IMPUTE_URL, + ARCHIVED_SIPP_VEHICLE_RECEIVER_URL, + ARCHIVED_SIPP_VEHICLE_SOURCE_URL, + ARCHIVED_SIPP_VEHICLE_TRANSFORM_URL, + SIPP_2023_VEHICLE_DONOR_REVISION, + SIPP_2023_VEHICLE_DONOR_SHA256, + SIPP_2023_VEHICLE_DONOR_SIZE_BYTES, + SIPP_2023_VEHICLE_DONOR_URL, + SIPP_VEHICLE_MODEL_PREDICTORS, + SIPP_VEHICLE_SOURCE_COLUMNS, + US_SIPP_VEHICLE_NONCONSTANT_HOUSEHOLD_COLUMNS, + US_SIPP_VEHICLE_OUTPUT_COLUMNS, + _append_owned_dummies, + _predictor_encoding, + _recipient_household_predictor_table, + fetch_sipp_2023_vehicle_donor, + impute_us_sipp_vehicles, + load_sipp_2023_vehicle_donor, + us_sipp_vehicles_signal_gate, + us_sipp_vehicles_stage_spec, + us_sipp_vehicles_summary, + with_us_sipp_vehicle_inputs, +) +from populace.frame import US_SCHEMA, Frame, WeightKind, Weights + +TIME_PERIOD = 2026 + + +def _source_row( + household_id: int, + person_number: int, + *, + month: int = 12, + weight: float = 100.0, + age: float = 40.0, + sex: int = 1, + marital_status: int = 1, + total_income: float = 1_000.0, + bank_income: float = 10.0, + stock_income: float = 5.0, + bond_income: float = 2.0, + rental_income: float = 3.0, + vehicles_owned: float = 2.0, + vehicle_value: float = 30_000.0, + home_value: float = 100_000.0, + owned_status: int = 1, + household_value_status: int = 1, + vehicle_1_status: int = 1, + vehicle_2_status: int = 1, + vehicle_3_status: int = 1, +) -> dict[str, float | int]: + return { + "SSUID": household_id, + "PNUM": person_number, + "MONTHCODE": month, + "WPFINWGT": weight, + "TAGE": age, + "ESEX": sex, + "EMS": marital_status, + "TPTOTINC": total_income, + "TINC_BANK": bank_income, + "TINC_STMF": stock_income, + "TINC_BOND": bond_income, + "TINC_RENT": rental_income, + "TVEH_NUM": vehicles_owned, + "THVAL_VEH": vehicle_value, + "THVAL_HOME": home_value, + "AVEH_NUM": owned_status, + "AHVAL_VEH": household_value_status, + "AVEH1VAL": vehicle_1_status, + "AVEH2VAL": vehicle_2_status, + "AVEH3VAL": vehicle_3_status, + } + + +def _write_sipp_source(tmp_path: Path) -> Path: + rows = [ + # A November record with extreme values must not enter the household. + _source_row( + 10, + 1, + month=11, + total_income=999_999.0, + vehicles_owned=5, + vehicle_value=185_150.0, + ), + _source_row(10, 1, household_value_status=5), + _source_row( + 10, + 2, + age=10, + sex=2, + marital_status=2, + total_income=100.0, + bank_income=1.0, + stock_income=0.0, + bond_income=0.0, + rental_income=0.0, + household_value_status=5, + ), + _source_row( + 20, + 1, + age=30, + sex=2, + marital_status=2, + vehicles_owned=0, + vehicle_value=0, + home_value=0, + owned_status=0, + household_value_status=0, + vehicle_1_status=0, + vehicle_2_status=0, + vehicle_3_status=0, + ), + # Count is allocated, but value is observed. + _source_row( + 30, + 1, + age=70, + vehicles_owned=3, + vehicle_value=45_000, + owned_status=2, + ), + # Count is observed, but a component value is allocated with status 5. + _source_row( + 40, + 1, + vehicles_owned=1, + vehicle_value=9_000, + vehicle_1_status=5, + ), + # Invalid survey weight: removed even though both targets are observed. + _source_row(50, 1, weight=0.0), + ] + source = pd.DataFrame(rows, columns=SIPP_VEHICLE_SOURCE_COLUMNS) + path = tmp_path / "pu2023.csv" + source.to_csv(path, sep="|", index=False) + return path + + +def _ready_donor(n: int = 320) -> pd.DataFrame: + rng = np.random.default_rng(71) + income = rng.gamma(2.0, 22_000.0, n) + homeowner = rng.integers(0, 2, n).astype(np.float64) + count = np.clip((income // 28_000).astype(int) + homeowner.astype(int), 0, 5) + # Retain a sizeable structural-zero class for both output gates. + count[income < 18_000] = 0 + value = np.where( + count > 0, + count * 7_500.0 + rng.gamma(1.5, 2_000.0, n), + 0.0, + ) + donor = pd.DataFrame( + { + "household_employment_income": income, + "household_interest_income": rng.gamma(1.0, 500.0, n), + "household_dividend_income": rng.gamma(1.0, 400.0, n), + "household_rental_income": np.where( + rng.random(n) < 0.12, rng.gamma(1.0, 3_000.0, n), 0.0 + ), + "reference_age": rng.integers(19, 90, n).astype(np.float64), + "reference_is_female": rng.integers(0, 2, n).astype(np.float64), + "reference_is_married": rng.integers(0, 2, n).astype(np.float64), + "count_under_18": rng.integers(0, 5, n).astype(np.float64), + "household_size": rng.integers(1, 7, n).astype(np.float64), + "is_homeowner": homeowner, + "household_vehicles_owned": count.astype(np.float64), + "household_vehicles_value": value, + "household_weight": rng.uniform(500.0, 2_500.0, n), + _OWNED_OBSERVED_COLUMN: True, + _VALUE_OBSERVED_COLUMN: True, + } + ) + return donor + + +def _recipient_frame(n_households: int = 90) -> Frame: + rng = np.random.default_rng(72) + people: list[dict[str, object]] = [] + person_id = 1 + for household_id in range(1, n_households + 1): + homeowner = household_id % 3 != 0 + head_income = float((household_id % 9) * 9_000 + 3_000) + # Put line 2 first to prove that A_LINENO, not input order, selects head. + people.append( + { + "person_id": person_id, + "person_household_id": household_id, + "person_spm_unit_id": household_id + 2_000, + "A_LINENO": 2, + "age": float(rng.integers(2, 17)), + "is_female": bool(rng.integers(0, 2)), + "A_MARITL": 7, + "employment_income_before_lsr": 0.0, + "taxable_interest_income": 0.0, + "tax_exempt_interest_income": 0.0, + "qualified_dividend_income": 0.0, + "non_qualified_dividend_income": 0.0, + "rental_income": 0.0, + "SPM_TENMORTSTATUS": 1 if homeowner else 3, + } + ) + person_id += 1 + people.append( + { + "person_id": person_id, + "person_household_id": household_id, + "person_spm_unit_id": household_id + 2_000, + "A_LINENO": 1, + "age": float(25 + household_id % 50), + "is_female": bool(household_id % 2), + "A_MARITL": 1 if household_id % 2 else 5, + "employment_income_before_lsr": head_income, + "taxable_interest_income": float(household_id % 5 * 80), + "tax_exempt_interest_income": float(household_id % 5 * 20), + "qualified_dividend_income": float(household_id % 4 * 50), + "non_qualified_dividend_income": float(household_id % 4 * 25), + "rental_income": float(household_id % 7 == 0) * 2_000.0, + "SPM_TENMORTSTATUS": 1 if homeowner else 3, + } + ) + person_id += 1 + person = pd.DataFrame(people) + household_ids = np.arange(1, n_households + 1, dtype=np.int64) + person["person_tax_unit_id"] = person["person_household_id"] + 1_000 + person["person_family_id"] = person["person_household_id"] + 3_000 + # Duplicate marital-unit ids exercise the archived pairing fallback. + person["person_marital_unit_id"] = person["person_household_id"] + 4_000 + household = pd.DataFrame( + { + "household_id": household_ids, + "net_worth": np.linspace(50_000.0, 150_000.0, n_households), + } + ) + tables = { + "person": person, + "household": household, + "tax_unit": pd.DataFrame( + {"tax_unit_id": np.arange(1, n_households + 1) + 1_000} + ), + "spm_unit": pd.DataFrame( + {"spm_unit_id": np.arange(1, n_households + 1) + 2_000} + ), + "family": pd.DataFrame( + {"family_id": np.arange(1, n_households + 1) + 3_000} + ), + "marital_unit": pd.DataFrame( + {"marital_unit_id": np.arange(1, n_households + 1) + 4_000} + ), + } + return Frame( + tables, + US_SCHEMA, + { + "household": Weights( + np.full(n_households, 1_000_000.0), WeightKind.DESIGN + ) + }, + ) + + +def _replace_household(frame: Frame, **columns: np.ndarray) -> Frame: + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + for column, values in columns.items(): + tables["household"][column] = values + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + ) + + +def test_immutable_artifact_and_archive_coordinates_are_exact() -> None: + assert SIPP_2023_VEHICLE_DONOR_REVISION == ( + "21280dca5995e978d706740a8a4b9b7860cfd7b6" + ) + assert SIPP_2023_VEHICLE_DONOR_SHA256 == ( + "5c30439e365fc26483318ef61d1d8f4bb2f0e9d6bb47c22c06756a7698733ee2" + ) + assert SIPP_2023_VEHICLE_DONOR_SIZE_BYTES == 3_726_010_471 + assert SIPP_2023_VEHICLE_DONOR_REVISION in SIPP_2023_VEHICLE_DONOR_URL + assert SIPP_2023_VEHICLE_DONOR_URL.endswith("/pu2023.csv") + assert "/resolve/" in SIPP_2023_VEHICLE_DONOR_URL + for url in ( + ARCHIVED_SIPP_VEHICLE_SOURCE_URL, + ARCHIVED_SIPP_VEHICLE_TRANSFORM_URL, + ARCHIVED_SIPP_VEHICLE_IMPUTE_URL, + ARCHIVED_SIPP_VEHICLE_RECEIVER_URL, + ): + assert "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe" in url + retired_repository = "policyengine-" + "us-data" + assert f"github.com/PolicyEngine/{retired_repository}/blob/" in url + + +def test_exact_source_columns_predictors_outputs_and_stage_spec() -> None: + assert SIPP_VEHICLE_SOURCE_COLUMNS == ( + "SSUID", + "PNUM", + "MONTHCODE", + "WPFINWGT", + "TAGE", + "ESEX", + "EMS", + "TPTOTINC", + "TINC_BANK", + "TINC_STMF", + "TINC_BOND", + "TINC_RENT", + "TVEH_NUM", + "THVAL_VEH", + "THVAL_HOME", + "AVEH_NUM", + "AHVAL_VEH", + "AVEH1VAL", + "AVEH2VAL", + "AVEH3VAL", + ) + assert SIPP_VEHICLE_MODEL_PREDICTORS == ( + "household_employment_income", + "household_interest_income", + "household_dividend_income", + "household_rental_income", + "reference_age", + "reference_is_female", + "reference_is_married", + "count_under_18", + "household_size", + "is_homeowner", + ) + assert US_SIPP_VEHICLE_OUTPUT_COLUMNS == ( + "household_vehicles_owned", + "household_vehicles_value", + ) + assert US_SIPP_VEHICLE_NONCONSTANT_HOUSEHOLD_COLUMNS == ( + US_SIPP_VEHICLE_OUTPUT_COLUMNS + ) + spec = us_sipp_vehicles_stage_spec() + assert spec.stage == "vehicle_assets" + assert set(US_SIPP_VEHICLE_OUTPUT_COLUMNS) <= set(spec.outputs) + + +def test_loader_applies_exact_december_household_transform_and_masks(tmp_path) -> None: + donor = load_sipp_2023_vehicle_donor( + _write_sipp_source(tmp_path), expected_size_bytes=None, chunksize=2 + ).set_index("household_id") + + assert donor.index.tolist() == [10, 20, 30, 40] + household_10 = donor.loc[10] + assert household_10["household_weight"] == pytest.approx(100.0) + assert household_10["household_employment_income"] == pytest.approx(13_200.0) + assert household_10["household_interest_income"] == pytest.approx(156.0) + assert household_10["household_dividend_income"] == pytest.approx(60.0) + assert household_10["household_rental_income"] == pytest.approx(36.0) + assert household_10["reference_age"] == pytest.approx(40.0) + assert household_10["reference_is_female"] == pytest.approx(0.0) + assert household_10["reference_is_married"] == pytest.approx(1.0) + assert household_10["count_under_18"] == pytest.approx(1.0) + assert household_10["household_size"] == pytest.approx(2.0) + assert household_10["is_homeowner"] == pytest.approx(1.0) + assert household_10["household_vehicles_owned"] == pytest.approx(2.0) + assert household_10["household_vehicles_value"] == pytest.approx(30_000.0) + # AHVAL_VEH=5 does not invalidate value; archived masking uses AVEH1-3. + assert bool(household_10[_OWNED_OBSERVED_COLUMN]) + assert bool(household_10[_VALUE_OBSERVED_COLUMN]) + + assert not bool(donor.loc[30, _OWNED_OBSERVED_COLUMN]) + assert bool(donor.loc[30, _VALUE_OBSERVED_COLUMN]) + assert bool(donor.loc[40, _OWNED_OBSERVED_COLUMN]) + assert not bool(donor.loc[40, _VALUE_OBSERVED_COLUMN]) + # The November extreme must not leak into either target or predictor. + assert donor.loc[10, "household_vehicles_owned"] != 5 + assert donor.loc[10, "household_employment_income"] < 100_000 + + +def test_loader_rejects_missing_columns_and_bad_hash(tmp_path) -> None: + path = _write_sipp_source(tmp_path) + with pytest.raises(ValueError, match="sha-256 verification"): + load_sipp_2023_vehicle_donor( + path, + expected_sha256="0" * 64, + expected_size_bytes=None, + ) + + missing_path = tmp_path / "missing.csv" + pd.DataFrame({"SSUID": [1], "MONTHCODE": [12]}).to_csv( + missing_path, sep="|", index=False + ) + with pytest.raises(ValueError, match="missing column"): + load_sipp_2023_vehicle_donor(missing_path, expected_size_bytes=None) + + +def test_cached_full_donor_matches_pinned_household_support() -> None: + snapshot = ( + Path.home() + / ".cache/huggingface/hub" + / ("models--policyengine--policyengine-" + "us-data") + / "snapshots" + / SIPP_2023_VEHICLE_DONOR_REVISION + / "pu2023.csv" + ) + if not snapshot.is_file(): + pytest.skip("the 3.73 GB pinned SIPP donor is not mounted") + + donor = load_sipp_2023_vehicle_donor( + snapshot, + expected_sha256=SIPP_2023_VEHICLE_DONOR_SHA256, + expected_size_bytes=SIPP_2023_VEHICLE_DONOR_SIZE_BYTES, + ) + owned_observed = donor[_OWNED_OBSERVED_COLUMN].astype(bool) + value_observed = donor[_VALUE_OBSERVED_COLUMN].astype(bool) + assert len(donor) == 16_841 + assert int(owned_observed.sum()) == 16_840 + assert int(value_observed.sum()) == 8_829 + assert int((owned_observed & value_observed).sum()) == 8_828 + assert set(donor.loc[owned_observed, "household_vehicles_owned"].unique()) == { + 0.0, + 1.0, + 2.0, + 3.0, + 4.0, + 5.0, + } + assert int( + ( + (donor["household_vehicles_value"] > 0) + & (donor["household_vehicles_owned"] == 0) + ).sum() + ) == 87 + + +class _ChunkedResponse(io.BytesIO): + def __init__(self, payload: bytes) -> None: + super().__init__(payload) + self.read_sizes: list[int] = [] + + def read(self, size: int = -1) -> bytes: + self.read_sizes.append(size) + return super().read(size) + + def __enter__(self): + return self + + def __exit__(self, exc_type, exc_value, traceback) -> None: + self.close() + + +def test_fetch_streams_verifies_atomically_and_reuses_cache(tmp_path, monkeypatch) -> None: + payload = b"small synthetic pinned donor payload" + digest = hashlib.sha256(payload).hexdigest() + response = _ChunkedResponse(payload) + monkeypatch.setattr(urllib.request, "urlopen", lambda *_args, **_kwargs: response) + + path = fetch_sipp_2023_vehicle_donor( + tmp_path, + expected_sha256=digest, + expected_size_bytes=len(payload), + chunk_size=5, + ) + assert path.read_bytes() == payload + assert len(response.read_sizes) > 2 + assert all(size == 5 for size in response.read_sizes) + assert not (tmp_path / "pu2023.csv.part").exists() + + def unexpected_download(*_args, **_kwargs): + raise AssertionError("valid cached donor should not be downloaded again") + + monkeypatch.setattr(urllib.request, "urlopen", unexpected_download) + assert ( + fetch_sipp_2023_vehicle_donor( + tmp_path, + expected_sha256=digest, + expected_size_bytes=len(payload), + chunk_size=5, + ) + == path + ) + + +def test_fetch_failure_removes_partial_without_replacing_existing(tmp_path, monkeypatch) -> None: + target = tmp_path / "pu2023.csv" + target.write_bytes(b"existing invalid cache") + response = _ChunkedResponse(b"new but wrong payload") + monkeypatch.setattr(urllib.request, "urlopen", lambda *_args, **_kwargs: response) + + with pytest.raises(ValueError, match="sha-256 verification"): + fetch_sipp_2023_vehicle_donor( + tmp_path, + expected_sha256="f" * 64, + expected_size_bytes=len(b"new but wrong payload"), + chunk_size=4, + ) + assert target.read_bytes() == b"existing invalid cache" + assert not (tmp_path / "pu2023.csv.part").exists() + + +def test_receiver_uses_line_one_head_household_sums_and_raw_tenure() -> None: + frame = _recipient_frame(6) + receiver = _recipient_household_predictor_table(frame) + assert receiver.index.tolist() == [1, 2, 3, 4, 5, 6] + # Head is line 1 even though line 2 appears first in the person table. + assert receiver.loc[1, "reference_age"] == pytest.approx(26.0) + assert receiver.loc[1, "reference_is_female"] == pytest.approx(1.0) + assert receiver.loc[1, "reference_is_married"] == pytest.approx(1.0) + assert receiver.loc[1, "household_employment_income"] == pytest.approx(12_000.0) + assert receiver.loc[1, "count_under_18"] == pytest.approx(1.0) + assert receiver.loc[1, "household_size"] == pytest.approx(2.0) + assert receiver.loc[1, "is_homeowner"] == pytest.approx(1.0) + assert receiver.loc[3, "is_homeowner"] == pytest.approx(0.0) + + +def test_count_is_weighted_deterministic_classifier_and_value_chains_on_dummies( + monkeypatch, +) -> None: + frame = _recipient_frame(45) + donor = _ready_donor() + calls: list[dict[str, object]] = [] + real_fit = RandomForestClassifier.fit + + def recording_fit(self, x, y, sample_weight=None): + calls.append( + { + "model": self, + "columns": tuple(x.columns), + "sample_weight": np.asarray(sample_weight), + "classes": tuple(sorted(pd.unique(y))), + } + ) + return real_fit(self, x, y, sample_weight=sample_weight) + + monkeypatch.setattr(RandomForestClassifier, "fit", recording_fit) + first = impute_us_sipp_vehicles(frame, donor, seed=42, n_estimators=12) + second = impute_us_sipp_vehicles(frame, donor, seed=42, n_estimators=12) + + assert calls + assert all(isinstance(call["model"], RandomForestClassifier) for call in calls) + assert all(call["model"].random_state == 42 for call in calls) + assert all(call["model"].n_estimators == 12 for call in calls) + assert all(np.all(call["sample_weight"] > 0) for call in calls) + assert calls[0]["classes"] == tuple(sorted(donor.household_vehicles_owned.unique())) + pd.testing.assert_frame_equal(first, second) + assert np.issubdtype(first["household_vehicles_owned"].dtype, np.integer) + assert np.issubdtype(first["household_vehicles_value"].dtype, np.floating) + assert (first["household_vehicles_owned"] >= 0).all() + assert (first["household_vehicles_value"] >= 0).all() + + encoding = _predictor_encoding(donor) + encoded = encoding.transform(donor) + _, owned_dummy_columns = _append_owned_dummies( + encoded, + donor["household_vehicles_owned"], + levels=tuple(sorted(donor["household_vehicles_owned"].unique())), + ) + assert owned_dummy_columns + assert all(name.startswith("household_vehicles_owned__") for name in owned_dummy_columns) + + +def test_frame_application_is_household_grain_idempotent_and_preserves_net_worth() -> None: + frame = _recipient_frame() + original_net_worth = frame.table("household")["net_worth"].copy() + restored = with_us_sipp_vehicle_inputs( + frame, + seed=42, + time_period=TIME_PERIOD, + sipp_donor=_ready_donor(), + n_estimators=16, + ) + household = restored.table("household") + assert set(US_SIPP_VEHICLE_OUTPUT_COLUMNS) <= set(household.columns) + pd.testing.assert_series_equal(household["net_worth"], original_net_worth) + assert len(household["household_vehicles_owned"]) == len( + frame.table("household") + ) + assert restored.weights_for("household") == frame.weights_for("household") + + passed_through = with_us_sipp_vehicle_inputs( + restored, + seed=999, + time_period=TIME_PERIOD, + sipp_donor=_ready_donor(40), + n_estimators=4, + ) + assert passed_through is restored + + +def test_summary_and_gate_require_broad_real_signal() -> None: + frame = with_us_sipp_vehicle_inputs( + _recipient_frame(), + seed=42, + time_period=TIME_PERIOD, + sipp_donor=_ready_donor(), + n_estimators=16, + ) + summary = us_sipp_vehicles_summary(frame) + gate = us_sipp_vehicles_signal_gate(frame) + assert gate.passed, gate.failures + assert summary["owned_noninteger_count"] == 0 + assert summary["weighted_totals"]["household_vehicles_owned"] > 0 + assert summary["weighted_totals"]["household_vehicles_value"] > 0 + assert summary["positive_value_with_positive_owned_weighted_share"] > 0 + + +def test_gate_rejects_missing_constant_negative_and_fractional_surfaces() -> None: + base = _recipient_frame(10) + missing = us_sipp_vehicles_signal_gate(base) + assert not missing.passed + assert "missing" in missing.failures[0] + + constant = _replace_household( + base, + household_vehicles_owned=np.zeros(10), + household_vehicles_value=np.zeros(10), + ) + constant_gate = us_sipp_vehicles_signal_gate(constant) + assert not constant_gate.passed + assert any("constant" in failure for failure in constant_gate.failures) + + bad = _replace_household( + base, + household_vehicles_owned=np.array( + [0.0, 1.5, 1.0, 2.0, 1.0, 2.0, 0.0, 1.0, 2.0, 1.0] + ), + household_vehicles_value=np.array( + [0.0, 5_000.0, -1.0, 8_000.0, 4_000.0, 7_000.0, 0.0, 3_000.0, 9_000.0, 2_000.0] + ), + ) + bad_gate = us_sipp_vehicles_signal_gate(bad) + assert not bad_gate.passed + assert any("non-integer" in failure for failure in bad_gate.failures) + assert any("negative" in failure for failure in bad_gate.failures) diff --git a/packages/populace-build/tests/test_us_source_runtime.py b/packages/populace-build/tests/test_us_source_runtime.py index 255a634b..32eb108b 100644 --- a/packages/populace-build/tests/test_us_source_runtime.py +++ b/packages/populace-build/tests/test_us_source_runtime.py @@ -15,14 +15,27 @@ from populace.build.us_runtime.puf_aggregate_records import ( AGGREGATE_RECIDS, SYNTHETIC_RECID_START, + derive_puf_policyengine_variables, disaggregate_puf_aggregate_records, ) from populace.build.us_runtime.source_runtime import ( aggregate_us_person_to_tax_unit_from_manifest, calibrate_us_binary_assignment_from_manifest, calibrate_us_binary_assignment_joint_targets_from_manifest, + derive_us_child_support_from_manifest, + derive_us_disability_benefits_from_manifest, + derive_us_energy_subsidy_from_manifest, + derive_us_other_health_insurance_from_manifest, + derive_us_prior_year_income_from_manifest, derive_us_puf_policyengine_variables_from_manifest, + derive_us_weeks_unemployed_from_manifest, disaggregate_us_puf_aggregate_records_from_manifest, + impute_us_child_support_to_puf_support_from_manifest, + impute_us_disability_benefits_to_puf_support_from_manifest, + impute_us_energy_subsidy_to_puf_support_from_manifest, + impute_us_other_health_insurance_to_puf_support_from_manifest, + impute_us_prior_year_income_to_puf_support_from_manifest, + impute_us_weeks_unemployed_to_puf_support_from_manifest, us_source_operation_handlers, ) @@ -59,6 +72,7 @@ def _make_runtime_mini_puf() -> pd.DataFrame: "E26270": abs_agi * rng.uniform(0.01, 0.15) * sign, "E00900": abs_agi * 0.03 * sign, "E02100": abs_agi * 0.01 * sign, + "E27200": abs_agi * 0.005 * sign, "E00400": abs_agi * 0.01, "E00600": abs_agi * 0.05, } @@ -91,11 +105,98 @@ def _make_runtime_mini_puf() -> pd.DataFrame: "E26270": abs_agi * 0.10 * sign, "E00900": abs_agi * 0.03 * sign, "E02100": abs_agi * 0.01 * sign, + "E27200": abs_agi * 0.005 * sign, "E00400": abs_agi * 0.01, "E00600": abs_agi * 0.05, } ) - return pd.DataFrame(rows) + result = pd.DataFrame(rows) + result["E03230"] = np.where(result["RECID"] % 5 == 0, 1_500.0, 0.0) + result["E87530"] = np.where(result["RECID"] % 7 == 0, 4_000.0, 0.0) + result["E00800"] = np.where(result["RECID"] % 13 == 0, 3_000.0, 0.0) + result["E03500"] = np.where(result["RECID"] % 17 == 0, 2_000.0, 0.0) + result["E20500"] = np.where(result["RECID"] % 11 == 0, 8_000.0, 0.0) + result["E03240"] = np.where(result["RECID"] % 9 == 0, 6_000.0, 0.0) + result["E03220"] = np.where(result["RECID"] % 7 == 0, 300.0, 0.0) + result["E20400"] = np.where(result["RECID"] % 3 == 0, 2_500.0, 0.0) + result["E58990"] = np.where(result["RECID"] % 19 == 0, 5_000.0, 0.0) + result["E00700"] = np.where(result["RECID"] % 5 == 0, 1_200.0, 0.0) + result["E24518"] = np.where(result["RECID"] % 23 == 0, 4_000.0, 0.0) + result["E24515"] = np.where(result["RECID"] % 29 == 0, 6_000.0, 0.0) + return result + + +def test_us_child_support_handlers_are_in_shared_registry() -> None: + handlers = us_source_operation_handlers() + + assert ( + handlers["derive_child_support_inputs"] is derive_us_child_support_from_manifest + ) + assert ( + handlers["impute_child_support_to_puf_support"] + is impute_us_child_support_to_puf_support_from_manifest + ) + + +def test_us_disability_benefits_handlers_are_in_shared_registry() -> None: + handlers = us_source_operation_handlers() + + assert ( + handlers["derive_disability_benefits"] + is derive_us_disability_benefits_from_manifest + ) + assert ( + handlers["impute_disability_benefits_to_puf_support"] + is impute_us_disability_benefits_to_puf_support_from_manifest + ) + + +def test_us_energy_subsidy_handlers_are_in_shared_registry() -> None: + handlers = us_source_operation_handlers() + + assert handlers["derive_energy_subsidy"] is derive_us_energy_subsidy_from_manifest + assert ( + handlers["impute_energy_subsidy_to_puf_support"] + is impute_us_energy_subsidy_to_puf_support_from_manifest + ) + + +def test_us_other_health_insurance_handlers_are_in_shared_registry() -> None: + handlers = us_source_operation_handlers() + + assert ( + handlers["derive_other_health_insurance_premiums"] + is derive_us_other_health_insurance_from_manifest + ) + assert ( + handlers["impute_other_health_insurance_premiums_to_puf_support"] + is impute_us_other_health_insurance_to_puf_support_from_manifest + ) + + +def test_us_prior_year_income_handlers_are_in_shared_registry() -> None: + handlers = us_source_operation_handlers() + + assert ( + handlers["derive_prior_year_income"] + is derive_us_prior_year_income_from_manifest + ) + assert ( + handlers["impute_prior_year_income_to_puf_support"] + is impute_us_prior_year_income_to_puf_support_from_manifest + ) + + +def test_us_weeks_unemployed_handlers_are_in_shared_registry() -> None: + handlers = us_source_operation_handlers() + + assert ( + handlers["derive_weeks_unemployed"] is derive_us_weeks_unemployed_from_manifest + ) + assert ( + handlers["impute_weeks_unemployed_to_puf_support"] + is impute_us_weeks_unemployed_to_puf_support_from_manifest + ) def _make_aca_people() -> pd.DataFrame: @@ -180,7 +281,26 @@ def test_us_puf_manifest_prefix_runs_aggregate_disaggregation() -> None: stop_after="disaggregate_aggregate_records", ) - expected = disaggregate_puf_aggregate_records(mini_puf, seed=42) + expected = disaggregate_puf_aggregate_records( + derive_puf_policyengine_variables( + mini_puf, + qualified_tuition_primary_source="E03230", + qualified_tuition_optional_source="E87530", + alimony_income_source="E00800", + alimony_expense_source="E03500", + casualty_loss_source="E20500", + domestic_production_ald_source="E03240", + educator_expense_source="E03220", + unreimbursed_business_employee_expenses_source="E20400", + farm_operations_income_source="E02100", + farm_rent_income_source="E27200", + investment_income_elected_form_4952_source="E58990", + salt_refund_income_source="E00700", + collectibles_capital_gain_source="E24518", + unrecaptured_section_1250_gain_source="E24515", + ), + seed=42, + ) pd.testing.assert_frame_equal(result, expected) assert not result["RECID"].isin(AGGREGATE_RECIDS).any() assert (result["RECID"] >= SYNTHETIC_RECID_START).any() @@ -193,6 +313,32 @@ def test_us_puf_manifest_prefix_runs_aggregate_disaggregation() -> None: result["non_qualified_dividend_income"], result["E00600"] - result["E00650"], ) + assert np.allclose( + result["qualified_tuition_expenses"], + np.maximum(result["E03230"], result["E87530"]), + ) + assert (result["qualified_tuition_expenses"] >= 0.0).all() + assert np.allclose(result["alimony_income"], result["E00800"]) + assert np.allclose(result["alimony_expense"], result["E03500"]) + assert np.allclose(result["casualty_loss"], result["E20500"]) + assert (result["casualty_loss"] >= 0.0).all() + assert np.allclose(result["domestic_production_ald"], result["E03240"]) + assert (result["domestic_production_ald"] >= 0.0).all() + assert np.allclose(result["educator_expense"], result["E03220"]) + assert (result["educator_expense"] >= 0.0).all() + assert np.allclose( + result["unreimbursed_business_employee_expenses"], result["E20400"] + ) + assert np.allclose(result["farm_operations_income"], result["E02100"]) + assert np.allclose(result["farm_rent_income"], result["E27200"]) + assert np.allclose(result["investment_income_elected_form_4952"], result["E58990"]) + assert (result["investment_income_elected_form_4952"] >= 0.0).all() + assert np.allclose(result["salt_refund_income"], result["E00700"]) + assert (result["salt_refund_income"] >= 0.0).all() + assert np.allclose( + result["long_term_capital_gains_on_collectibles"], result["E24518"] + ) + assert np.allclose(result["unrecaptured_section_1250_gain"], result["E24515"]) def test_us_puf_manifest_prefix_uses_build_seed() -> None: diff --git a/packages/populace-build/tests/test_us_ssi_disability_criteria.py b/packages/populace-build/tests/test_us_ssi_disability_criteria.py new file mode 100644 index 00000000..cbe4206c --- /dev/null +++ b/packages/populace-build/tests/test_us_ssi_disability_criteria.py @@ -0,0 +1,672 @@ +"""Source-faithful SIPP SSI disability-criteria stage tests.""" + +from __future__ import annotations + +import importlib.util +from importlib.metadata import version +from pathlib import Path +from typing import Any + +import numpy as np +import pandas as pd +import pytest + +import populace.build.us_runtime.ssi_disability_criteria as module +from populace.build.us_runtime.puf_support import clone_us_frame_for_puf_support +from populace.build.us_runtime.ssi_disability_criteria import ( + SIPP_2023_SSI_DISABILITY_DONOR_REVISION, + SIPP_2023_SSI_DISABILITY_DONOR_SHA256, + SIPP_2023_SSI_DISABILITY_DONOR_SIZE_BYTES, + SIPP_2023_SSI_DISABILITY_DONOR_URL, + SIPP_SSI_DISABILITY_DIFFICULTY_PREDICTORS, + SIPP_SSI_DISABILITY_FIT_PARAMETERS, + SIPP_SSI_DISABILITY_MODEL_PREDICTORS, + SIPP_SSI_DISABILITY_READ_PARAMETERS, + SIPP_SSI_DISABILITY_SOURCE_COLUMNS, + SSI_DISABILITY_ARCHIVED_CPS_URL, + SSI_DISABILITY_ARCHIVED_EXTENDED_CPS_URL, + SSI_DISABILITY_ARCHIVED_SIPP_URL, + SSI_DISABILITY_ARCHIVED_SOURCE_IMPUTE_URL, + US_SSI_DISABILITY_CRITERIA_NONCONSTANT_PERSON_COLUMNS, + US_SSI_DISABILITY_CRITERIA_OUTPUT_COLUMNS, + US_SSI_DISABILITY_CRITERIA_STAGE_NAME, + impute_us_ssi_disability_criteria, + load_sipp_2023_ssi_disability_donor, + us_ssi_disability_criteria_signal_gate, + us_ssi_disability_criteria_stage_spec, + us_ssi_disability_criteria_summary, + with_us_ssi_disability_criteria, +) +from populace.frame import US_SCHEMA, Frame, WeightKind, Weights + +_OUTPUT = US_SSI_DISABILITY_CRITERIA_OUTPUT_COLUMNS[0] +_policyengine_us_installed = importlib.util.find_spec("policyengine_us") is not None +requires_us = pytest.mark.skipif( + not _policyengine_us_installed, + reason="requires the policyengine-us [us] extra", +) + + +def _source_row( + ssuid: str, + pnum: int, + *, + month: int = 12, + age: float = 40.0, + received_ssi: float = 2.0, + reason: float = np.nan, + receipt_allocation: float = 1.0, + reason_allocation: float = 1.0, + weight: float = 100.0, + assets: float = 0.0, + monthly_earnings: float = 0.0, + difficulty_seeing: bool = False, +) -> dict[str, object]: + row: dict[str, object] = { + column: 0.0 for column in SIPP_SSI_DISABILITY_SOURCE_COLUMNS + } + row.update( + { + "SSUID": ssuid, + "PNUM": pnum, + "MONTHCODE": month, + "SPANEL": 2023, + "SWAVE": 1, + "WPFINWGT": weight, + "TAGE": age, + "ESEX": 2, + "EMS": 2, + "TVAL_BANK": assets, + "TVAL_STMF": 0.0, + "TVAL_BOND": 0.0, + "TINC_BANK": 0.0, + "TINC_STMF": 0.0, + "TINC_BOND": 0.0, + "TINC_RENT": 0.0, + "TPTOTINC": monthly_earnings, + "TJB1_MSUM": monthly_earnings, + "TSSSAMT": 0.0, + "RSSI_YRYN": received_ssi, + "ESSI_BRSN": reason, + "ASSI_YRYN": receipt_allocation, + "ASSI_BRSN": reason_allocation, + "ESSRSN2YN": 2.0, + "EDISANY": 2.0, + "ESELFCARE": 2.0, + "EHEARING": 2.0, + "ESEEING": 1.0 if difficulty_seeing else 2.0, + "EERRANDS": 2.0, + "EAMBULAT": 2.0, + "ECOGNIT": 2.0, + } + ) + return row + + +def _write_source(tmp_path: Path, rows: list[dict[str, object]]) -> Path: + path = tmp_path / "pu2023.csv" + pd.DataFrame(rows, columns=SIPP_SSI_DISABILITY_SOURCE_COLUMNS).to_csv( + path, + sep="|", + index=False, + ) + return path + + +def _donor(n: int = 120) -> pd.DataFrame: + rng = np.random.default_rng(747) + donor = pd.DataFrame(index=np.arange(n)) + donor["age"] = np.arange(n, dtype=np.float64) + donor["is_female"] = rng.integers(0, 2, n) + donor["is_married"] = rng.integers(0, 2, n) + donor["employment_income"] = rng.gamma(2.0, 5_000.0, n) + donor["interest_income"] = rng.gamma(1.0, 100.0, n) + donor["dividend_income"] = rng.gamma(1.0, 100.0, n) + donor["rental_income"] = rng.normal(0.0, 500.0, n) + donor["bank_account_assets"] = rng.gamma(1.0, 1_000.0, n) + donor["stock_assets"] = rng.gamma(1.0, 1_000.0, n) + donor["bond_assets"] = rng.gamma(1.0, 200.0, n) + donor["count_under_18"] = rng.integers(0, 5, n) + for predictor in SIPP_SSI_DISABILITY_DIFFICULTY_PREDICTORS: + donor[predictor] = rng.integers(0, 2, n) + donor["social_security_disability"] = rng.integers(0, 2, n) * 6_000.0 + donor["has_disability_income"] = rng.integers(0, 2, n) + donor[_OUTPUT] = np.arange(n) % 5 == 0 + donor["household_weight"] = np.arange(1, n + 1, dtype=np.float64) + return donor.loc[ + :, + [*SIPP_SSI_DISABILITY_MODEL_PREDICTORS, _OUTPUT, "household_weight"], + ] + + +def _frame(n: int = 20) -> Frame: + ids = np.arange(1, n + 1, dtype=np.int64) + person = pd.DataFrame( + { + "person_id": ids, + "person_household_id": ids, + "person_tax_unit_id": ids + 100, + "person_spm_unit_id": ids + 200, + "person_family_id": ids + 300, + "person_marital_unit_id": ids + 400, + "age": np.full(n, 40.0), + "is_female": ids % 2 == 0, + "A_MARITL": np.full(n, 5), + "employment_income_before_lsr": np.zeros(n), + # Deliberately omit the aggregate leaves: the receiver must use + # the complete measured PolicyEngine component pairs. + "taxable_interest_income": np.arange(n, dtype=np.float64), + "tax_exempt_interest_income": np.full(n, 2.0), + "qualified_dividend_income": np.arange(n, dtype=np.float64) * 3.0, + "non_qualified_dividend_income": np.full(n, 4.0), + "rental_income": np.zeros(n), + "bank_account_assets": np.where(np.isin(ids, [2, 3]), 100.0, 0.0), + "stock_assets": np.zeros(n), + "bond_assets": np.zeros(n), + "PEDISDRS": np.where(ids == 2, 1, 2), + "PEDISEAR": np.full(n, 2), + "PEDISEYE": np.full(n, 2), + "PEDISOUT": np.full(n, 2), + "PEDISPHY": np.full(n, 2), + "PEDISREM": np.full(n, 2), + # The archived signal coercion is strictly >0, not generic truthiness. + "social_security_disability": np.where(ids == 3, -10.0, 0.0), + "disability_benefits": np.zeros(n), + "SSI_VAL": np.where(ids == 1, 1_200.0, 0.0), + } + ) + tables = { + "person": person, + "household": pd.DataFrame({"household_id": ids}), + "tax_unit": pd.DataFrame({"tax_unit_id": ids + 100}), + "spm_unit": pd.DataFrame({"spm_unit_id": ids + 200}), + "family": pd.DataFrame({"family_id": ids + 300}), + "marital_unit": pd.DataFrame({"marital_unit_id": ids + 400}), + } + return Frame( + tables, + US_SCHEMA, + {"household": Weights(np.ones(n), WeightKind.DESIGN)}, + ) + + +def _replace_person(frame: Frame, **columns: Any) -> Frame: + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + for column, values in columns.items(): + tables["person"][column] = values + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + ) + + +class _FakeQRF: + instances: list[_FakeQRF] = [] + predict_receivers: list[pd.DataFrame] = [] + predict_start_offsets: list[int] = [] + + def __init__(self, *, n_estimators: int, seed: int) -> None: + self.n_estimators = n_estimators + self.seed = seed + self.training: pd.DataFrame | None = None + self.weights: object = None + self.receiver: pd.DataFrame | None = None + self.draw_offset = 0 + self.__class__.instances.append(self) + + def fit( + self, + training: pd.DataFrame, + *, + predictors: list[str], + targets: list[str], + weights: object, + ) -> _FakeQRF: + assert predictors == list(SIPP_SSI_DISABILITY_MODEL_PREDICTORS) + assert targets == [_OUTPUT] + self.training = training.copy() + self.weights = weights + return self + + def predict(self, receiver: pd.DataFrame) -> pd.DataFrame: + self.receiver = receiver.copy() + self.__class__.predict_receivers.append(receiver.copy()) + self.__class__.predict_start_offsets.append(self.draw_offset) + self.draw_offset += len(receiver) + return pd.DataFrame( + {_OUTPUT: receiver["bank_account_assets"].to_numpy() > 0.0}, + index=receiver.index, + ) + + +@pytest.fixture(autouse=True) +def _clear_fake_instances() -> None: + _FakeQRF.instances.clear() + _FakeQRF.predict_receivers.clear() + _FakeQRF.predict_start_offsets.clear() + + +def test_archived_coordinates_exact_predictors_and_pinned_artifact() -> None: + assert US_SSI_DISABILITY_CRITERIA_STAGE_NAME == "ssi_disability_criteria" + assert US_SSI_DISABILITY_CRITERIA_OUTPUT_COLUMNS == ( + "meets_ssi_disability_criteria", + ) + assert US_SSI_DISABILITY_CRITERIA_NONCONSTANT_PERSON_COLUMNS == ( + "meets_ssi_disability_criteria", + ) + assert SIPP_2023_SSI_DISABILITY_DONOR_REVISION == ( + "21280dca5995e978d706740a8a4b9b7860cfd7b6" + ) + assert SIPP_2023_SSI_DISABILITY_DONOR_SHA256 == ( + "5c30439e365fc26483318ef61d1d8f4bb2f0e9d6bb47c22c06756a7698733ee2" + ) + assert SIPP_2023_SSI_DISABILITY_DONOR_SIZE_BYTES == 3_726_010_471 + assert SIPP_2023_SSI_DISABILITY_DONOR_REVISION in ( + SIPP_2023_SSI_DISABILITY_DONOR_URL + ) + assert SSI_DISABILITY_ARCHIVED_SIPP_URL.endswith("datasets/sipp/sipp.py#L63-L105") + assert SSI_DISABILITY_ARCHIVED_CPS_URL.endswith("datasets/cps/cps.py#L2853-L2886") + assert SSI_DISABILITY_ARCHIVED_SOURCE_IMPUTE_URL.endswith( + "calibration/source_impute.py#L869-L990" + ) + assert SSI_DISABILITY_ARCHIVED_EXTENDED_CPS_URL.endswith( + "datasets/cps/extended_cps.py#L392-L424" + ) + assert len(SIPP_SSI_DISABILITY_MODEL_PREDICTORS) == 19 + assert SIPP_SSI_DISABILITY_MODEL_PREDICTORS == ( + "age", + "is_female", + "is_married", + "employment_income", + "interest_income", + "dividend_income", + "rental_income", + "bank_account_assets", + "stock_assets", + "bond_assets", + "count_under_18", + *SIPP_SSI_DISABILITY_DIFFICULTY_PREDICTORS, + "social_security_disability", + "has_disability_income", + ) + + +def test_stage_manifest_pins_exact_runtime_contract() -> None: + spec = us_ssi_disability_criteria_stage_spec() + + assert spec.grain == "person" + assert spec.outputs == US_SSI_DISABILITY_CRITERIA_OUTPUT_COLUMNS + assert [operation.kind for operation in spec.operations] == [ + "read_table", + "fit_weighted_qrf", + ] + assert dict(spec.operations[0].parameters) == SIPP_SSI_DISABILITY_READ_PARAMETERS + assert dict(spec.operations[1].parameters) == SIPP_SSI_DISABILITY_FIT_PARAMETERS + assert SIPP_SSI_DISABILITY_FIT_PARAMETERS["training_sample_seed"] == ( + 8_386_123_572_872_638_692 + ) + assert SIPP_SSI_DISABILITY_FIT_PARAMETERS["model_seed"] == 42 + assert SIPP_SSI_DISABILITY_FIT_PARAMETERS["seed_from_build_config"] is False + + +def test_loader_applies_observation_allocation_and_financial_screens( + tmp_path: Path, + monkeypatch: pytest.MonkeyPatch, +) -> None: + monkeypatch.setattr( + module, + "_ssi_policy_screen_values", + lambda _year: { + "individual_resource_limit": 2_000.0, + "couple_resource_limit": 3_000.0, + "individual_fbr": 943.0, + "couple_fbr": 1_415.0, + "general_exclusion": 20.0, + "earned_exclusion": 65.0, + "earned_share_excluded": 0.5, + "non_blind_sga": 1_550.0, + }, + ) + rows = [ + _source_row("month11", 1, month=11), + # A positive observed label survives otherwise failing finances. + _source_row( + "positive", + 1, + received_ssi=1, + reason=1, + assets=100_000, + monthly_earnings=10_000, + ), + # Nonrecipients do not need an observed reason. + _source_row("negative", 1, reason=np.nan), + _source_row("high_assets", 1, assets=10_000), + _source_row( + "allocated_reason", + 1, + received_ssi=1, + reason=1, + reason_allocation=2, + ), + _source_row("allocated_receipt", 1, receipt_allocation=2), + _source_row("aged_negative", 1, age=70), + # $1,600 monthly passes countable-income but fails nonblind SGA. + _source_row("sga", 1, monthly_earnings=1_600), + _source_row( + "blind_sga", + 1, + monthly_earnings=1_600, + difficulty_seeing=True, + ), + # A reported aged reason is a clean false label for an under-65 row. + _source_row("aged_reason", 1, received_ssi=1, reason=2), + ] + donor = load_sipp_2023_ssi_disability_donor( + _write_source(tmp_path, rows), + expected_size_bytes=None, + chunksize=3, + ) + + assert len(donor) == 4 + assert donor[_OUTPUT].tolist() == [True, False, False, False] + assert donor["difficulty_seeing"].tolist() == [False, False, True, False] + audit = donor.attrs["source_audit"] + assert audit["december_rows"] == 9 + assert audit["training_rows"] == 4 + assert audit["positive_rows"] == 1 + assert audit["negative_rows"] == 3 + assert audit["pinned_transform"] is False + + +def test_loader_rejects_missing_allocation_flags(tmp_path: Path) -> None: + path = _write_source(tmp_path, [_source_row("one", 1)]) + source = pd.read_csv(path, sep="|").drop(columns=["ASSI_BRSN"]) + source.to_csv(path, sep="|", index=False) + + with pytest.raises(ValueError, match="ASSI_BRSN"): + load_sipp_2023_ssi_disability_donor( + path, + expected_size_bytes=None, + ) + + +def test_pinned_full_file_audit_contract_is_exact() -> None: + assert module._PINNED_DECEMBER_ROWS == 39_513 + assert module._PINNED_TRAINING_ROWS == 9_346 + assert module._PINNED_POSITIVE_ROWS == 577 + assert module._PINNED_NEGATIVE_ROWS == 8_769 + assert module._PINNED_WEIGHT_SUM == pytest.approx(88_690_359.47893329) + assert module._PINNED_POSITIVE_WEIGHT_SUM == pytest.approx(4_937_167.914119501) + assert module._PINNED_WEIGHTED_TRUE_SHARE == pytest.approx(0.05566746987074994) + assert module._PINNED_RESAMPLE_UNIQUE_SOURCE_ROWS == 5_314 + assert module._PINNED_RESAMPLE_POSITIVE_ROWS == 524 + assert module._PINNED_RESAMPLE_TRUE_SHARE == pytest.approx(0.05606676653113631) + + +def test_imputer_uses_exact_weighted_replacement_draw_and_fixed_model_seed( + monkeypatch: pytest.MonkeyPatch, +) -> None: + donor = _donor(120) + monkeypatch.setattr(module, "QRF", _FakeQRF) + + first = impute_us_ssi_disability_criteria(_frame(), donor, seed=1) + first_fit = _FakeQRF.instances[-1] + second = impute_us_ssi_disability_criteria(_frame(), donor, seed=999) + second_fit = _FakeQRF.instances[-1] + + probability = donor["household_weight"].to_numpy(dtype=np.float64).copy() + probability /= probability.sum() + expected_positions = np.random.default_rng(8_386_123_572_872_638_692).choice( + 120, size=120, replace=True, p=probability + ) + expected_ages = donor.iloc[expected_positions]["age"].to_numpy() + assert first_fit.n_estimators == 100 + assert first_fit.seed == 42 + assert first_fit.weights == "none" + np.testing.assert_array_equal(first_fit.training["age"], expected_ages) + np.testing.assert_array_equal(second_fit.training["age"], expected_ages) + np.testing.assert_array_equal(first, second) + + +def test_receiver_uses_complete_income_components_and_archived_signal_screen( + monkeypatch: pytest.MonkeyPatch, +) -> None: + monkeypatch.setattr(module, "QRF", _FakeQRF) + result = impute_us_ssi_disability_criteria(_frame(), _donor(), seed=7) + receiver = _FakeQRF.instances[-1].receiver + assert receiver is not None + + np.testing.assert_array_equal( + receiver["interest_income"], + np.arange(20, dtype=np.float64) + 2.0, + ) + np.testing.assert_array_equal( + receiver["dividend_income"], + np.arange(20, dtype=np.float64) * 3.0 + 4.0, + ) + # Person 1 is preserved by direct reported SSI despite no model/signal; + # person 2 has both the positive model draw and a difficulty. Person 3 has + # a positive model draw but negative SSDI, which is not a disability signal. + assert np.flatnonzero(result.to_numpy()).tolist() == [0, 1] + + +def test_asec_reporter_anchor_is_not_copied_to_puf_and_rows_predict_separately( + monkeypatch: pytest.MonkeyPatch, +) -> None: + monkeypatch.setattr(module, "QRF", _FakeQRF) + expanded = clone_us_frame_for_puf_support(_frame()) + person = expanded.table("person") + puf = person["person_support_channel"].astype(str).eq("puf_tax_detail") + source_three = person["person_source_id"].eq(3) + # Give only the PUF row of source person 3 a positive model draw and signal. + person.loc[puf & source_three, "bank_account_assets"] = 100.0 + person.loc[puf & source_three, "PEDISDRS"] = 1 + + result = impute_us_ssi_disability_criteria(expanded, _donor(), seed=7) + assert len(_FakeQRF.predict_receivers) == 2 + assert _FakeQRF.predict_start_offsets == [0, 0] + assert [len(receiver) for receiver in _FakeQRF.predict_receivers] == [20, 20] + assert [ + expanded.table("person") + .loc[receiver.index, "person_support_channel"] + .unique() + .tolist() + for receiver in _FakeQRF.predict_receivers + ] == [["asec"], ["puf_tax_detail"]] + rows = pd.DataFrame( + { + "source": person["person_source_id"].to_numpy(), + "channel": person["person_support_channel"].astype(str).to_numpy(), + "value": result.to_numpy(), + } + ) + reporter = rows[rows["source"] == 1].set_index("channel")["value"] + source_three_values = rows[rows["source"] == 3].set_index("channel")["value"] + + assert bool(reporter["asec"]) + assert not bool(reporter["puf_tax_detail"]) + assert not bool(source_three_values["asec"]) + assert bool(source_three_values["puf_tax_detail"]) + + +def test_support_validation_allows_puf_only_source_ids( + monkeypatch: pytest.MonkeyPatch, +) -> None: + monkeypatch.setattr(module, "QRF", _FakeQRF) + expanded = clone_us_frame_for_puf_support(_frame()) + person = expanded.table("person") + puf = person["person_support_channel"].astype(str).eq("puf_tax_detail") + person.loc[puf, "person_source_id"] += 10_000 + + result = impute_us_ssi_disability_criteria(expanded, _donor(), seed=0) + + assert len(result) == len(person) + + +def test_support_validation_rejects_unknown_or_missing_channel( + monkeypatch: pytest.MonkeyPatch, +) -> None: + monkeypatch.setattr(module, "QRF", _FakeQRF) + expanded = clone_us_frame_for_puf_support(_frame()) + expanded.table("person").loc[0, "person_support_channel"] = "mystery" + + with pytest.raises(ValueError, match="unsupported support channel"): + impute_us_ssi_disability_criteria(expanded, _donor(), seed=0) + + +def test_wrapper_heals_stale_output_and_is_idempotent( + monkeypatch: pytest.MonkeyPatch, +) -> None: + monkeypatch.setattr(module, "QRF", _FakeQRF) + monkeypatch.setattr(module, "us_ssi_disability_criteria_stage_spec", lambda: None) + stale = _replace_person(_frame(), **{_OUTPUT: np.ones(20, dtype=bool)}) + + healed = with_us_ssi_disability_criteria( + stale, + seed=123, + time_period=2024, + sipp_donor=_donor(), + ) + twice = with_us_ssi_disability_criteria( + healed, + seed=999, + time_period=2024, + sipp_donor=_donor(), + ) + + assert healed.table("person")[_OUTPUT].sum() == 2 + assert twice is healed + + +def test_signal_gate_requires_each_channel_but_allows_clone_divergence( + monkeypatch: pytest.MonkeyPatch, +) -> None: + monkeypatch.setattr(module, "QRF", _FakeQRF) + expanded = clone_us_frame_for_puf_support(_frame()) + person = expanded.table("person") + puf = person["person_support_channel"].astype(str).eq("puf_tax_detail") + source_three = person["person_source_id"].eq(3) + person.loc[puf & source_three, "bank_account_assets"] = 100.0 + person.loc[puf & source_three, "PEDISDRS"] = 1 + values = impute_us_ssi_disability_criteria(expanded, _donor(), seed=0) + valid = _replace_person(expanded, **{_OUTPUT: values.to_numpy()}) + + summary = us_ssi_disability_criteria_summary(valid) + gate = us_ssi_disability_criteria_signal_gate(valid) + assert gate.passed, gate.failures + assert summary["clone_divergence_source_people"] == 2 + assert summary["channels"]["asec"]["unique_count"] == 2 + assert summary["channels"]["puf_tax_detail"]["unique_count"] == 2 + + dead_values = values.copy() + dead_values.loc[puf.to_numpy()] = False + dead = _replace_person(valid, **{_OUTPUT: dead_values.to_numpy()}) + dead_gate = us_ssi_disability_criteria_signal_gate(dead) + assert not dead_gate.passed + assert any("puf_tax_detail" in failure for failure in dead_gate.failures) + + implausible_values = values.copy() + implausible_values.loc[puf.to_numpy()] = True + first_puf = int(np.flatnonzero(puf.to_numpy())[0]) + implausible_values.iloc[first_puf] = False + implausible = _replace_person( + valid, + **{_OUTPUT: implausible_values.to_numpy()}, + ) + implausible_gate = us_ssi_disability_criteria_signal_gate(implausible) + assert not implausible_gate.passed + assert any( + "puf_tax_detail" in failure and "plausibility band" in failure + for failure in implausible_gate.failures + ) + + +def test_gate_requires_complete_support_provenance( + monkeypatch: pytest.MonkeyPatch, +) -> None: + monkeypatch.setattr(module, "QRF", _FakeQRF) + expanded = clone_us_frame_for_puf_support(_frame()) + values = impute_us_ssi_disability_criteria(expanded, _donor(), seed=0) + tables = {entity: expanded.table(entity).copy() for entity in expanded.entities} + tables["person"][_OUTPUT] = values.to_numpy() + tables["person"] = tables["person"].drop(columns=["person_source_id"]) + without_source_id = Frame( + tables, + expanded.schema, + {entity: expanded.weights_for(entity) for entity in expanded.weighted_entities}, + expanded.strata, + mass_log=expanded.mass_log, + ) + + gate = us_ssi_disability_criteria_signal_gate(without_source_id) + assert not gate.passed + assert gate.details["support_provenance_missing"] is True + assert any("provenance" in failure for failure in gate.failures) + + tables = {entity: expanded.table(entity).copy() for entity in expanded.entities} + tables["person"][_OUTPUT] = values.to_numpy() + tables["person"] = tables["person"].drop(columns=["person_support_channel"]) + without_channel = Frame( + tables, + expanded.schema, + {entity: expanded.weights_for(entity) for entity in expanded.weighted_entities}, + expanded.strata, + mass_log=expanded.mass_log, + ) + channel_gate = us_ssi_disability_criteria_signal_gate(without_channel) + assert not channel_gate.passed + assert channel_gate.details["support_provenance_missing"] is True + assert any( + "support channel 'asec' is missing" in failure + for failure in channel_gate.failures + ) + assert any( + "support channel 'puf_tax_detail' is missing" in failure + for failure in channel_gate.failures + ) + + +@requires_us +def test_policyengine_us_1_764_6_ssi_is_positive_then_zero_when_neutralized() -> None: + from policyengine_us import CountryTaxBenefitSystem, Simulation + + assert version("policyengine-us") == "1.764.6" + variable = CountryTaxBenefitSystem().variables[_OUTPUT] + assert variable.is_input_variable() + assert variable.entity.key == "person" + assert variable.value_type is bool + assert variable.default_value is False + + def situation(criterion: bool) -> dict[str, object]: + return { + "people": { + "adult": { + "age": {"2024": 40}, + _OUTPUT: {"2024": criterion}, + } + }, + "tax_units": { + "unit": { + "members": ["adult"], + "filing_status": {"2024": "SINGLE"}, + } + }, + "families": {"family": {"members": ["adult"]}}, + "spm_units": {"spm": {"members": ["adult"]}}, + "households": { + "household": { + "members": ["adult"], + "state_code": {"2024": "CA"}, + } + }, + "marital_units": {"marital": {"members": ["adult"]}}, + } + + active = Simulation(situation=situation(True)) + neutralized = Simulation(situation=situation(False)) + + assert active.calculate("ssi", "2024-01")[0] == pytest.approx(943.0) + assert neutralized.calculate("ssi", "2024-01")[0] == 0.0 diff --git a/packages/populace-build/tests/test_us_ssi_take_up.py b/packages/populace-build/tests/test_us_ssi_take_up.py new file mode 100644 index 00000000..fcd22a0a --- /dev/null +++ b/packages/populace-build/tests/test_us_ssi_take_up.py @@ -0,0 +1,468 @@ +"""Reporter-anchored, SSA count-calibrated SSI take-up tests.""" + +from __future__ import annotations + +import copy +import importlib.util +import json +from importlib.metadata import version +from pathlib import Path + +import numpy as np +import pandas as pd +import pytest + +from populace.build.us_runtime.ssi_take_up import ( + SSI_TAKE_UP_ARCHIVED_DERIVATION_URL, + SSI_TAKE_UP_ARCHIVED_RANDOMNESS_URL, + SSI_TAKE_UP_ARCHIVED_TARGETS_URL, + SSI_TAKE_UP_SSA_SOURCE_URL, + US_SSI_TAKE_UP_AGE_TARGETS, + US_SSI_TAKE_UP_ANCHOR, + US_SSI_TAKE_UP_OUTPUT_COLUMNS, + US_SSI_TAKE_UP_STAGE_NAME, + US_SSI_TAKE_UP_TARGET_TABLE_NAME, + us_ssi_take_up_diagnostics, + us_ssi_take_up_gate, + us_ssi_take_up_reporter_source_ids, + us_ssi_take_up_stage_spec, + with_us_ssi_take_up, + write_us_ssi_take_up_diagnostics, +) +from populace.frame import US_SCHEMA, Frame, WeightKind, Weights + +_OUTPUT = US_SSI_TAKE_UP_OUTPUT_COLUMNS[0] +_TARGETS = {target.key: 50.0 for target in US_SSI_TAKE_UP_AGE_TARGETS} +_AGES = {"under_18": 12.0, "18_64": 40.0, "65_plus": 72.0} +_policyengine_us_installed = importlib.util.find_spec("policyengine_us") is not None +requires_us = pytest.mark.skipif( + not _policyengine_us_installed, + reason="requires the policyengine-us [us] extra", +) + + +def _frame(*, stale_output: bool = False) -> tuple[Frame, np.ndarray]: + """Build three age bands with ASEC/PUF clones and PUF-only support.""" + + rows: list[dict[str, object]] = [] + potential: list[float] = [] + person_id = 0 + for band, age in _AGES.items(): + for source_number in range(9): + if source_number <= 5 or source_number == 8: + channels = ("asec", "puf_tax_detail") + elif source_number == 6: + channels = ("asec",) + else: + channels = ("puf_tax_detail",) + candidate = source_number <= 5 or source_number == 7 + reporter = source_number in {0, 6} + for channel in channels: + rows.append( + { + "person_id": person_id, + "person_household_id": person_id, + "person_tax_unit_id": person_id, + "person_spm_unit_id": person_id, + "person_family_id": person_id, + "person_marital_unit_id": person_id, + "age": age, + # PUF copies can carry SSI_VAL, but only direct ASEC + # rows are independent reporter anchors. + US_SSI_TAKE_UP_ANCHOR: 1_200.0 if reporter else 0.0, + "person_source_id": f"{band}:{source_number}", + "person_support_channel": channel, + _OUTPUT: bool((person_id % 2) == 0) if stale_output else False, + } + ) + potential.append(100.0 if candidate else 0.0) + person_id += 1 + + person = pd.DataFrame(rows) + ids = person["person_id"].to_numpy() + tables = { + "person": person, + "household": pd.DataFrame({"household_id": ids}), + "tax_unit": pd.DataFrame({"tax_unit_id": ids}), + "spm_unit": pd.DataFrame({"spm_unit_id": ids}), + "family": pd.DataFrame({"family_id": ids}), + "marital_unit": pd.DataFrame({"marital_unit_id": ids}), + } + frame = Frame( + tables, + US_SCHEMA, + { + "household": Weights( + values=np.full(len(person), 10.0), + kind=WeightKind.DESIGN, + ) + }, + ) + return frame, np.asarray(potential, dtype=np.float64) + + +def _replace_person(frame: Frame, person: pd.DataFrame) -> Frame: + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + tables["person"] = person + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + ) + + +def _assigned( + *, + seed: int = 17, + targets: dict[str, float] | None = None, + stale_output: bool = False, +) -> tuple[Frame, Frame, np.ndarray, dict[str, object]]: + frame, potential = _frame(stale_output=stale_output) + result, diagnostics = with_us_ssi_take_up( + frame, + uncapped_ssi=potential, + seed=seed, + targets=targets or _TARGETS, + ) + return frame, result, potential, diagnostics + + +def test_stage_contract_pins_archived_method_and_official_age_targets() -> None: + spec = us_ssi_take_up_stage_spec() + assert spec.stage == US_SSI_TAKE_UP_STAGE_NAME + assert spec.source == SSI_TAKE_UP_SSA_SOURCE_URL + assert spec.outputs == (_OUTPUT,) + assert US_SSI_TAKE_UP_TARGET_TABLE_NAME in spec.operations[2].parameters["targets"] + assert "42ed5d45" in SSI_TAKE_UP_ARCHIVED_DERIVATION_URL + assert "cps.py#L650-L657" in SSI_TAKE_UP_ARCHIVED_DERIVATION_URL + assert "takeup.py#L10-L35" in SSI_TAKE_UP_ARCHIVED_RANDOMNESS_URL + assert "ssi_targets.py#L41-L74" in SSI_TAKE_UP_ARCHIVED_TARGETS_URL + assert [target.person_count for target in US_SSI_TAKE_UP_AGE_TARGETS] == [ + 1_001_922.0, + 3_905_779.0, + 2_382_142.0, + ] + + +def test_assignment_preserves_asec_reporters_and_fans_source_decisions() -> None: + _, result, _, diagnostics = _assigned() + person = result.table("person") + flag = person[_OUTPUT].to_numpy(dtype=bool) + direct_reporter = ( + person["person_support_channel"].eq("asec") + & person[US_SSI_TAKE_UP_ANCHOR].gt(0) + ).to_numpy() + assert flag[direct_reporter].all() + assert diagnostics["reporter_anchor_lost_count"] == 0 + assert diagnostics["source_identity_mismatch_count"] == 0 + assert person.groupby("person_source_id")[_OUTPUT].nunique().max() == 1 + + +def test_puf_only_ssi_value_is_not_promoted_to_reporter_anchor() -> None: + frame, potential = _frame() + baseline, baseline_diagnostics = with_us_ssi_take_up( + frame, + uncapped_ssi=potential, + seed=17, + targets={key: 20.0 for key in _TARGETS}, + ) + person = frame.table("person").copy() + puf_only = person["person_source_id"].str.endswith(":7") + person.loc[puf_only, US_SSI_TAKE_UP_ANCHOR] = 9_999.0 + result, diagnostics = with_us_ssi_take_up( + _replace_person(frame, person), + uncapped_ssi=potential, + seed=17, + targets={key: 20.0 for key in _TARGETS}, + ) + np.testing.assert_array_equal( + result.table("person")[_OUTPUT], baseline.table("person")[_OUTPUT] + ) + assert diagnostics["age_bands"] == baseline_diagnostics["age_bands"] + + +def test_reporter_lineage_survives_when_l0_keeps_only_the_puf_clone() -> None: + full, potential = _frame() + reporter_source_ids = us_ssi_take_up_reporter_source_ids(full) + person = full.table("person") + dropped = person["person_source_id"].eq("under_18:0") & person[ + "person_support_channel" + ].eq("asec") + sparse = full.select(~dropped.to_numpy()) + sparse_potential = potential[~dropped.to_numpy()] + result, diagnostics = with_us_ssi_take_up( + sparse, + uncapped_ssi=sparse_potential, + seed=17, + targets=_TARGETS, + reporter_source_ids=reporter_source_ids, + ) + surviving_clone = result.table("person")["person_source_id"].eq("under_18:0") + assert result.table("person").loc[surviving_clone, _OUTPUT].all() + child = diagnostics["age_bands"][0] + assert child["reporter_source_identity_count"] == 2 + + +def test_non_candidate_reporter_remains_anchored_but_not_in_recipient_count() -> None: + _, result, _, diagnostics = _assigned() + person = result.table("person") + noncandidate_reporter = person["person_source_id"].str.endswith(":6") + assert person.loc[noncandidate_reporter, _OUTPUT].all() + for band in diagnostics["age_bands"]: + assert band["reporter_source_identity_count"] == 2 + assert band["reporter_candidate_floor"] == 20.0 + + +def test_reachable_targets_hit_within_one_source_identity_weight() -> None: + _, _, _, diagnostics = _assigned() + for band in diagnostics["age_bands"]: + assert not band["saturated"] + assert ( + abs(band["selected_recipient_weight"] - band["target"]) + <= band["max_source_candidate_weight"] + ) + assert us_ssi_take_up_gate(diagnostics, targets=_TARGETS).passed + + +def test_unreachable_target_saturates_only_the_counted_candidate_domain() -> None: + targets = {"under_18": 1_000.0, "18_64": 50.0, "65_plus": 50.0} + _, result, potential, diagnostics = _assigned(targets=targets) + flag = result.table("person")[_OUTPUT].to_numpy(dtype=bool) + child_rows = result.table("person")["age"].lt(18).to_numpy() + assert flag[child_rows & (potential > 0)].all() + child, adult, aged = diagnostics["age_bands"] + assert child["saturated"] + assert child["selected_recipient_weight"] == child["candidate_capacity"] + assert child["target_shortfall"] > 0 + assert not adult["saturated"] + assert not aged["saturated"] + assert us_ssi_take_up_gate(diagnostics, targets=targets).passed + + +def test_reporter_floor_above_target_never_drops_an_anchor() -> None: + targets = {key: 5.0 for key in _TARGETS} + _, result, _, diagnostics = _assigned(targets=targets) + person = result.table("person") + direct_reporter = person["person_support_channel"].eq("asec") & person[ + US_SSI_TAKE_UP_ANCHOR + ].gt(0) + assert person.loc[direct_reporter, _OUTPUT].all() + for band in diagnostics["age_bands"]: + assert band["reachable_goal"] == band["reporter_candidate_floor"] + + +def test_assignment_is_deterministic_source_keyed_and_seed_sensitive() -> None: + _, first, _, first_diagnostics = _assigned(seed=17) + _, repeat, _, repeat_diagnostics = _assigned(seed=17) + _, alternative, _, _ = _assigned(seed=18) + np.testing.assert_array_equal( + first.table("person")[_OUTPUT], repeat.table("person")[_OUTPUT] + ) + assert first_diagnostics == repeat_diagnostics + assert not np.array_equal( + first.table("person")[_OUTPUT], alternative.table("person")[_OUTPUT] + ) + + +def test_stale_output_is_healed_and_exact_rerun_returns_same_frame() -> None: + _, result, potential, diagnostics = _assigned(stale_output=True) + again, again_diagnostics = with_us_ssi_take_up( + result, + uncapped_ssi=potential, + seed=17, + targets=_TARGETS, + ) + assert again is result + assert again_diagnostics == diagnostics + assert us_ssi_take_up_gate(again_diagnostics, targets=_TARGETS).passed + + +def test_matching_numeric_output_is_rewritten_to_canonical_boolean() -> None: + _, result, potential, _ = _assigned() + person = result.table("person").copy() + person[_OUTPUT] = person[_OUTPUT].astype(np.float64) + numeric = _replace_person(result, person) + healed, diagnostics = with_us_ssi_take_up( + numeric, + uncapped_ssi=potential, + seed=17, + targets=_TARGETS, + ) + assert healed is not numeric + assert pd.api.types.is_bool_dtype(healed.table("person")[_OUTPUT]) + assert us_ssi_take_up_gate(diagnostics, targets=_TARGETS).passed + + +@pytest.mark.parametrize( + ("mutation", "message"), + [ + ("missing_column", "missing person source"), + ("unknown_channel", "exact ASEC/PUF"), + ("missing_channel", "complete support provenance"), + ("cross_age_band", "cross SSA age bands"), + ("blank_source", "nonblank"), + ("duplicate_channel", "at most one row"), + ], +) +def test_source_provenance_failures_are_rejected(mutation: str, message: str) -> None: + frame, potential = _frame() + person = frame.table("person").copy() + if mutation == "missing_column": + person = person.drop(columns=["person_source_id"]) + elif mutation == "unknown_channel": + person.loc[person.index[0], "person_support_channel"] = "synthetic" + elif mutation == "missing_channel": + person.loc[person.index[0], "person_support_channel"] = None + elif mutation == "cross_age_band": + adult = person["person_source_id"].eq("18_64:7") + person.loc[adult, "person_source_id"] = "under_18:6" + elif mutation == "blank_source": + person.loc[person.index[0], "person_source_id"] = " " + else: + person.loc[person.index[1], "person_support_channel"] = "asec" + with pytest.raises(ValueError, match=message): + with_us_ssi_take_up( + _replace_person(frame, person), + uncapped_ssi=potential, + seed=17, + targets=_TARGETS, + ) + + +@pytest.mark.parametrize( + ("field", "value", "failure_fragment"), + [ + ("schema_version", 999, "schema version"), + ("target_source", "https://example.org", "target source"), + ("candidate_definition", "is_ssi_eligible", "candidate definition"), + ("unique_count", 1, "constant"), + ("reporter_anchor_lost_count", 1, "reporter anchors"), + ("source_identity_mismatch_count", 1, "source identity"), + ], +) +def test_gate_rejects_tampered_top_level_diagnostics( + field: str, value: object, failure_fragment: str +) -> None: + _, _, _, diagnostics = _assigned() + tampered = copy.deepcopy(diagnostics) + tampered[field] = value + gate = us_ssi_take_up_gate(tampered, targets=_TARGETS) + assert not gate.passed + assert any(failure_fragment in failure for failure in gate.failures) + + +def test_gate_rejects_hidden_saturation_and_large_reachable_miss() -> None: + _, _, _, diagnostics = _assigned() + hidden = copy.deepcopy(diagnostics) + hidden["age_bands"][0]["saturated"] = True + hidden_gate = us_ssi_take_up_gate(hidden, targets=_TARGETS) + assert not hidden_gate.passed + assert any("saturation status" in failure for failure in hidden_gate.failures) + + missed = copy.deepcopy(diagnostics) + missed["age_bands"][0]["selected_recipient_weight"] = 0.0 + missed_gate = us_ssi_take_up_gate(missed, targets=_TARGETS) + assert not missed_gate.passed + assert any("misses reachable goal" in failure for failure in missed_gate.failures) + + +def test_gate_rejects_duplicate_age_band_diagnostics() -> None: + _, _, _, diagnostics = _assigned() + duplicated = copy.deepcopy(diagnostics) + duplicated["age_bands"].append(copy.deepcopy(duplicated["age_bands"][0])) + gate = us_ssi_take_up_gate(duplicated, targets=_TARGETS) + assert not gate.passed + assert any("exactly one diagnostic row" in failure for failure in gate.failures) + + +def test_existing_assignment_diagnostics_do_not_reassign_flags() -> None: + _, result, potential, _ = _assigned() + original = result.table("person")[_OUTPUT].to_numpy(dtype=bool).copy() + candidate_selected = original & (potential > 0) + drifted_weights = np.where(candidate_selected, 100.0, 1.0) + reweighted = Frame( + {entity: result.table(entity).copy() for entity in result.entities}, + result.schema, + { + "household": Weights( + drifted_weights, + WeightKind.CALIBRATED, + ) + }, + result.strata, + mass_log=result.mass_log, + ) + diagnostics = us_ssi_take_up_diagnostics( + reweighted, + uncapped_ssi=potential, + seed=17, + targets=_TARGETS, + ) + np.testing.assert_array_equal(reweighted.table("person")[_OUTPUT], original) + assert not us_ssi_take_up_gate(diagnostics, targets=_TARGETS).passed + + recomputed, _ = with_us_ssi_take_up( + reweighted, + uncapped_ssi=potential, + seed=17, + targets=_TARGETS, + ) + assert not np.array_equal(recomputed.table("person")[_OUTPUT], original) + + +def test_writer_emits_strict_json_and_refuses_nan(tmp_path: Path) -> None: + _, _, _, diagnostics = _assigned() + path = write_us_ssi_take_up_diagnostics(diagnostics, tmp_path / "ssi.json") + assert json.loads(path.read_text()) == diagnostics + invalid = copy.deepcopy(diagnostics) + invalid["weighted_flag_true_share"] = np.nan + with pytest.raises(ValueError, match="Out of range float"): + write_us_ssi_take_up_diagnostics(invalid, tmp_path / "invalid.json") + + +@requires_us +def test_policyengine_us_1_764_6_take_up_flag_controls_positive_ssi() -> None: + from policyengine_us import CountryTaxBenefitSystem, Simulation + + assert version("policyengine-us") == "1.764.6" + variable = CountryTaxBenefitSystem().variables[_OUTPUT] + assert variable.is_input_variable() + assert variable.entity.key == "person" + assert variable.value_type is bool + assert variable.default_value is True + + def situation(takes_up: bool) -> dict[str, object]: + return { + "people": { + "adult": { + "age": {"2024": 40}, + "meets_ssi_disability_criteria": {"2024": True}, + _OUTPUT: {"2024": takes_up}, + } + }, + "tax_units": { + "unit": { + "members": ["adult"], + "filing_status": {"2024": "SINGLE"}, + } + }, + "families": {"family": {"members": ["adult"]}}, + "spm_units": {"spm": {"members": ["adult"]}}, + "households": { + "household": { + "members": ["adult"], + "state_code": {"2024": "CA"}, + } + }, + "marital_units": {"marital": {"members": ["adult"]}}, + } + + active = Simulation(situation=situation(True)) + neutralized = Simulation(situation=situation(False)) + period = "2024-12" + assert active.calculate("uncapped_ssi", period)[0] == pytest.approx(943.0) + assert neutralized.calculate("uncapped_ssi", period)[0] == pytest.approx(943.0) + assert active.calculate("ssi", period)[0] == pytest.approx(943.0) + assert neutralized.calculate("ssi", period)[0] == 0.0 diff --git a/packages/populace-build/tests/test_us_survivor_benefits_exclusion.py b/packages/populace-build/tests/test_us_survivor_benefits_exclusion.py new file mode 100644 index 00000000..ea18a9ed --- /dev/null +++ b/packages/populace-build/tests/test_us_survivor_benefits_exclusion.py @@ -0,0 +1,177 @@ +"""Evidence contract for the irreducible survivor-benefits exclusion.""" + +from __future__ import annotations + +import importlib.util +import json +from hashlib import sha256 +from importlib.metadata import version +from importlib.resources import files +from pathlib import Path + +import numpy as np +import pytest + +from populace.build.us_runtime.release_input_coverage import ( + load_release_input_coverage_manifest, +) + +ROOT = Path(__file__).resolve().parents[3] +policyengine_us_installed = importlib.util.find_spec("policyengine_us") is not None +requires_us = pytest.mark.skipif( + not policyengine_us_installed, + reason="requires the policyengine-us [us] extra (build environment)", +) + + +def _entry() -> dict[str, object]: + payload = json.loads( + files("populace.build.us").joinpath("ecps_parity_known_gaps.json").read_text() + ) + return payload["known_gaps"]["survivor_benefits"] + + +def _sha256(path: Path) -> str: + digest = sha256() + with path.open("rb") as stream: + for chunk in iter(lambda: stream.read(1024 * 1024), b""): + digest.update(chunk) + return digest.hexdigest() + + +def test_exclusion_pins_exact_archived_person_source_and_qrf_dependency() -> None: + entry = _entry() + evidence = entry["evidence"] + + assert entry["reason"].startswith("SOURCE UNAVAILABILITY WITH EVIDENCE:") + assert evidence["classification"] == "source_unavailability" + assert evidence["retired_derivation"] == { + "repository_owner": "PolicyEngine", + "repository_name_parts": ["policyengine-", "us-data"], + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": [ + "policyengine_", + "us_data", + "datasets", + "cps", + "cps.py", + ], + "lines": "1495", + } + assert evidence["required_person_source_declaration"]["lines"] == ("39-58,306-359") + assert evidence["puf_clone_qrf"]["lines"] == "140-194,639-745" + assert evidence["puf_clone_qrf"]["target_line"] == 167 + assert evidence["no_independent_puf_source"] == { + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": [ + "policyengine_", + "us_data", + "calibration", + "puf_impute.py", + ], + "tax_detail_target_lines": "90-149,158-198", + "survivor_benefits_occurrences": 0, + } + assert evidence["required_person_columns"] == ["SRVS_VAL"] + assert evidence["aggregate_only_columns"] == { + "family": ["FSURVAL", "FINC_SUR"], + "household": ["HSURVAL", "HSUR_YN"], + } + assert "synthesize" in evidence["semantic_rejection"] + + +def test_exclusion_pins_all_sha_locked_hermetic_inputs() -> None: + evidence = _entry()["evidence"] + evidence_hashes = { + item["filename"]: item["sha256"] for item in evidence["hermetic_inputs"] + } + build_summary = json.loads( + (ROOT / "experiments/build_j_recert/base_j.summary.json").read_text() + ) + recorded_hashes = { + Path(item["path"]).name: item["sha256"] + for item in build_summary["base_source"]["sources"] + } + + assert evidence_hashes == recorded_hashes + assert set(evidence_hashes) == { + "census_cps_2022.h5", + "census_cps_2023.h5", + "census_cps_2024.h5", + } + for item in evidence["hermetic_inputs"]: + assert item["missing_person_columns"] == ["SRVS_VAL"] + assert item["present_family_columns"] == ["FSURVAL", "FINC_SUR"] + assert item["present_household_columns"] == ["HSURVAL", "HSUR_YN"] + assert item["positive_multi_person_families"] > 0 + + assert ( + "HSURVAL equals the household sum of FSURVAL" in evidence["aggregate_identity"] + ) + + build_script = (ROOT / "experiments/build_j_recert/buildj_base.sh").read_text() + for year in (2022, 2023, 2024): + assert f'--asec-h5 {year}="$USD/census_cps_{year}.h5"' in build_script + assert "buildj_base.sh lines 65-69" in evidence["hermetic_build_contract"] + assert "base_j.summary.json lines 55-75" in evidence["hermetic_build_contract"] + + +@requires_us +def test_sha_locked_artifact_schemas_match_the_recorded_absence() -> None: + from policyengine_us.data import USSingleYearDataset + + evidence = _entry()["evidence"] + build_summary = json.loads( + (ROOT / "experiments/build_j_recert/base_j.summary.json").read_text() + ) + paths = { + Path(item["path"]).name: Path(item["path"]) + for item in build_summary["base_source"]["sources"] + } + if not all(path.is_file() for path in paths.values()): + pytest.skip("SHA-locked ASEC artifacts are not mounted in this environment") + + for item in evidence["hermetic_inputs"]: + path = paths[item["filename"]] + assert _sha256(path) == item["sha256"] + dataset = USSingleYearDataset(file_path=str(path)) + assert set(item["missing_person_columns"]).isdisjoint(dataset.person.columns) + assert set(item["present_family_columns"]) <= set(dataset.family.columns) + assert set(item["present_household_columns"]) <= set(dataset.household.columns) + positive_multi_person_families = int( + ((dataset.family["FSURVAL"] > 0) & (dataset.family["FPERSONS"] > 1)).sum() + ) + assert positive_multi_person_families == item["positive_multi_person_families"] + family_amount_by_household = dataset.family.groupby("FH_SEQ", sort=False)[ + "FSURVAL" + ].sum() + household_amount = dataset.household.set_index("H_SEQ")["HSURVAL"] + np.testing.assert_array_equal( + family_amount_by_household.reindex( + household_amount.index, + fill_value=0, + ).to_numpy(), + household_amount.to_numpy(), + ) + + +@requires_us +def test_policyengine_1_764_6_requires_a_person_year_input() -> None: + from policyengine_us import CountryTaxBenefitSystem + + assert version("policyengine-us") == "1.764.6" + variable = CountryTaxBenefitSystem().variables["survivor_benefits"] + assert variable.is_input_variable() + assert variable.entity.key == "person" + assert str(variable.definition_period).lower() == "year" + assert variable.documentation == ( + "Survivor benefits other than Social Security survivor benefits." + ) + + +def test_generated_release_manifest_preserves_the_evidenced_exclusion() -> None: + reason = load_release_input_coverage_manifest().reviewed_exclusions[ + "survivor_benefits" + ] + assert reason == _entry()["reason"] + assert reason.startswith("SOURCE UNAVAILABILITY WITH EVIDENCE:") diff --git a/packages/populace-build/tests/test_us_take_up_contract.py b/packages/populace-build/tests/test_us_take_up_contract.py index 89261fbf..ce90fdbc 100644 --- a/packages/populace-build/tests/test_us_take_up_contract.py +++ b/packages/populace-build/tests/test_us_take_up_contract.py @@ -18,6 +18,7 @@ TAKE_UP_CONTRACT_ENGINE_FACT_KEYS, assert_take_up_contract_current, assert_take_up_treatments_consistent, + count_calibrated_take_up_programs, load_take_up_contract, seeded_take_up_programs, ) @@ -50,6 +51,76 @@ def test_every_program_has_a_valid_treatment(self) -> None: "near_universal", } + def test_ssi_uses_reporter_anchored_ssa_age_count_calibration(self) -> None: + program = load_take_up_contract().program_map()["takes_up_ssi_if_eligible"] + calibration = program.raw["calibration"] + + assert program in count_calibrated_take_up_programs() + assert calibration == { + "anchor": "SSI_VAL", + "targets": ["ssa_ssi_federal_payment_recipients_by_age"], + "target_table": "ssa_ssi_federal_payment_recipients_by_age", + "target_source": ( + "https://www.ssa.gov/policy/docs/statcomps/" + "ssi_monthly/2024-12/table01.html" + ), + "target_period": "2024-12", + "target_measure": "Total with—Federal payment", + "target_values": { + "under_18": 1_001_922, + "18_64": 3_905_779, + "65_plus": 2_382_142, + }, + "aggregate_target": 7_289_843, + "age_bands": { + "under_18": "age < 18", + "18_64": "18 <= age < 65", + "65_plus": "age >= 65", + }, + "semantics": ( + "SSA SSI Monthly Statistics December 2024 Table 1 recipients " + "in the Total with—Federal payment row, calibrated within " + "uncapped_ssi > 0 by source-person identity; unreachable age " + "bands saturate without assigning outside modeled eligibility" + ), + } + assert program.raw["scope_owner"] == ( + "ssi_take_up source stage (eCPS exported-input coverage)" + ) + + def test_head_start_is_owned_by_measured_sipp_stage(self) -> None: + program = load_take_up_contract().program_map()[ + "takes_up_head_start_if_eligible" + ] + + assert program.populace_treatment == "out_of_scope" + assert program.raw["scope_owner"] == ( + "sipp_head_start source stage (measured SIPP enrollment response)" + ) + assert program.rate == {"status": "not_used_measured_source"} + notes = program.raw["notes"] + assert "EEDHEADST" in notes + assert "direct December age-3--5" in notes + assert "QRF" in notes + assert "retired NIEER scalar" in notes + + def test_early_head_start_records_irreducible_source_unavailability(self) -> None: + program = load_take_up_contract().program_map()[ + "takes_up_early_head_start_if_eligible" + ] + + assert program.populace_treatment == "rate_unsourced" + assert program.rate["status"] == "source_unavailable" + followup = program.raw["followup"].lower() + assert "locked individual-level source" in followup + assert "infants" in followup + assert "toddlers" in followup + assert "pregnant" in followup + notes = program.raw["notes"].lower() + assert "ecps_parity_known_gaps.json" in notes + assert "aggregate" in notes + assert "synthesize" in notes + class TestEngineAssertion: def test_checked_in_table_matches_installed_engine(self) -> None: diff --git a/packages/populace-build/tests/test_us_voluntary_filing.py b/packages/populace-build/tests/test_us_voluntary_filing.py new file mode 100644 index 00000000..1311bb72 --- /dev/null +++ b/packages/populace-build/tests/test_us_voluntary_filing.py @@ -0,0 +1,748 @@ +"""Measured SIPP voluntary-filing source-stage tests.""" + +from __future__ import annotations + +import hashlib +import importlib.util +import io +import urllib.request +from importlib.metadata import version +from pathlib import Path + +import numpy as np +import pandas as pd +import pytest + +import populace.build.us_runtime.voluntary_filing as module +from populace.build.us_runtime.puf_support import clone_us_frame_for_puf_support +from populace.build.us_runtime.release_input_coverage import ( + us_release_reform_coverage_probes, +) +from populace.build.us_runtime.voluntary_filing import ( + SIPP_2023_VOLUNTARY_FILING_DONOR_REVISION, + SIPP_2023_VOLUNTARY_FILING_DONOR_SHA256, + SIPP_2023_VOLUNTARY_FILING_DONOR_SIZE_BYTES, + SIPP_2023_VOLUNTARY_FILING_DONOR_URL, + SIPP_VOLUNTARY_FILING_MODEL_PREDICTORS, + SIPP_VOLUNTARY_FILING_SOURCE_COLUMNS, + US_VOLUNTARY_FILING_NONCONSTANT_TAX_UNIT_COLUMNS, + US_VOLUNTARY_FILING_OUTPUT_COLUMNS, + US_VOLUNTARY_FILING_STAGE_NAME, + VOLUNTARY_FILING_ARCHIVED_DERIVATION_URL, + VOLUNTARY_FILING_ARCHIVED_PARAMETERS_URL, + VOLUNTARY_FILING_SIPP_DICTIONARY_URL, + fetch_sipp_2023_voluntary_filing_donor, + impute_us_voluntary_filing, + load_sipp_2023_voluntary_filing_donor, + us_voluntary_filing_signal_gate, + us_voluntary_filing_stage_spec, + us_voluntary_filing_summary, + with_us_voluntary_filing_input, +) +from populace.frame import US_SCHEMA, Frame, WeightKind, Weights + +_OUTPUT = US_VOLUNTARY_FILING_OUTPUT_COLUMNS[0] +_policyengine_us_installed = importlib.util.find_spec("policyengine_us") is not None +requires_us = pytest.mark.skipif( + not _policyengine_us_installed, + reason="requires the policyengine-us [us] extra", +) + + +def _source_row( + ssuid: int, + pnum: int, + *, + month: int = 12, + weight: float = 10.0, + age: float = 40.0, + sex: int = 1, + spouse: float = np.nan, + filing: float = 1.0, + filing_status: int = 1, + will_file: float = np.nan, + will_file_status: int = 0, + dependent: float = np.nan, + monthly_wages: float = 1_000.0, +) -> dict[str, object]: + row: dict[str, object] = { + "SSUID": ssuid, + "PNUM": pnum, + "MONTHCODE": month, + "WPFINWGT": weight, + "TAGE": age, + "ESEX": sex, + "EPNSPOUSE": spouse, + "AFILING": filing_status, + "EFILING": filing, + "AWILLFILE": will_file_status, + "EWILLFILE": will_file, + "EDEPCLM": dependent, + } + for job in range(1, 8): + row[f"TJB{job}_MSUM"] = monthly_wages if job == 1 else 0.0 + return row + + +def _write_source(tmp_path: Path, rows: list[dict[str, object]]) -> Path: + path = tmp_path / "pu2023.csv" + pd.DataFrame(rows, columns=SIPP_VOLUNTARY_FILING_SOURCE_COLUMNS).to_csv( + path, sep="|", index=False + ) + return path + + +def _synthetic_source_rows() -> list[dict[str, object]]: + return [ + # A non-December duplicate must not enter either target or predictors. + _source_row(1, 101, month=11, monthly_wages=999_999.0), + # Reciprocal spouses report the same filing target and collapse once. + _source_row(1, 101, spouse=102, age=40, monthly_wages=1_000.0), + _source_row( + 1, + 102, + spouse=101, + age=38, + sex=2, + monthly_wages=500.0, + ), + # Household context is counted before the filing-response filter. + _source_row( + 1, + 103, + age=10, + filing=np.nan, + filing_status=0, + monthly_wages=0.0, + ), + # A directly reported no/not-planning response supplies the false class. + _source_row( + 2, + 101, + age=70, + sex=2, + filing=2, + will_file=2, + will_file_status=1, + monthly_wages=200.0, + ), + # Claimed dependents are not standalone receiver tax units. + _source_row(3, 101, age=21, dependent=1, monthly_wages=300.0), + # Imputed filing and will-file answers are not measured targets. + _source_row(4, 101, filing_status=2), + _source_row( + 5, + 101, + filing=2, + will_file=1, + will_file_status=2, + ), + # Invalid survey weights are excluded after source-unit construction. + _source_row(6, 101, weight=0.0), + ] + + +def _donor(n: int = 80) -> pd.DataFrame: + rng = np.random.default_rng(918) + income = rng.gamma(2.0, 18_000.0, n) + target = (np.arange(n) % 4) != 0 + return pd.DataFrame( + { + "source_tax_unit_key": [f"unit:{i}" for i in range(n)], + "employment_income": income, + "reference_age": rng.integers(18, 90, n).astype(np.float64), + "reference_is_female": rng.integers(0, 2, n).astype(np.float64), + "reference_is_married": rng.integers(0, 2, n).astype(np.float64), + "count_under_18": rng.integers(0, 5, n).astype(np.float64), + _OUTPUT: target, + "tax_unit_weight": rng.uniform(100.0, 2_000.0, n), + } + ) + + +def _frame(n_households: int = 12) -> Frame: + people: list[dict[str, object]] = [] + person_id = 1 + for household_id in range(1, n_households + 1): + tax_unit_id = household_id + 100 + people.append( + { + "person_id": person_id, + "person_household_id": household_id, + "person_tax_unit_id": tax_unit_id, + "person_spm_unit_id": household_id + 200, + "person_family_id": household_id + 300, + "person_marital_unit_id": household_id + 400, + "age": float(25 + household_id % 50), + "is_female": bool(household_id % 2), + "tax_unit_role_input": "HEAD", + "employment_income_before_lsr": float(household_id * 3_000), + "A_LINENO": 1, + } + ) + person_id += 1 + if household_id % 3 == 0: + people.append( + { + "person_id": person_id, + "person_household_id": household_id, + "person_tax_unit_id": tax_unit_id, + "person_spm_unit_id": household_id + 200, + "person_family_id": household_id + 300, + "person_marital_unit_id": household_id + 400, + "age": float(22 + household_id % 40), + "is_female": not bool(household_id % 2), + "tax_unit_role_input": "SPOUSE", + "employment_income_before_lsr": 2_000.0, + "A_LINENO": 2, + } + ) + person_id += 1 + if household_id % 2 == 0: + people.append( + { + "person_id": person_id, + "person_household_id": household_id, + "person_tax_unit_id": tax_unit_id, + "person_spm_unit_id": household_id + 200, + "person_family_id": household_id + 300, + "person_marital_unit_id": household_id + 400, + "age": 8.0, + "is_female": False, + "tax_unit_role_input": "DEPENDENT", + "employment_income_before_lsr": 0.0, + "A_LINENO": 3, + } + ) + person_id += 1 + person = pd.DataFrame(people) + household_ids = np.arange(1, n_households + 1, dtype=np.int64) + tables = { + "person": person, + "household": pd.DataFrame({"household_id": household_ids}), + "tax_unit": pd.DataFrame( + { + "tax_unit_id": household_ids + 100, + "filing_status_input": np.where( + household_ids % 3 == 0, "JOINT", "SINGLE" + ), + } + ), + "spm_unit": pd.DataFrame({"spm_unit_id": household_ids + 200}), + "family": pd.DataFrame({"family_id": household_ids + 300}), + "marital_unit": pd.DataFrame({"marital_unit_id": household_ids + 400}), + } + return Frame( + tables, + US_SCHEMA, + {"household": Weights(np.linspace(1.0, 2.0, n_households), WeightKind.DESIGN)}, + ) + + +def _replace_tax_unit(frame: Frame, **columns: np.ndarray) -> Frame: + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + for column, values in columns.items(): + tables["tax_unit"][column] = values + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + ) + + +def test_archived_and_pinned_source_coordinates_are_exact() -> None: + assert US_VOLUNTARY_FILING_STAGE_NAME == "voluntary_filing_input" + assert SIPP_2023_VOLUNTARY_FILING_DONOR_REVISION == ( + "21280dca5995e978d706740a8a4b9b7860cfd7b6" + ) + assert SIPP_2023_VOLUNTARY_FILING_DONOR_SHA256 == ( + "5c30439e365fc26483318ef61d1d8f4bb2f0e9d6bb47c22c06756a7698733ee2" + ) + assert SIPP_2023_VOLUNTARY_FILING_DONOR_SIZE_BYTES == 3_726_010_471 + assert SIPP_2023_VOLUNTARY_FILING_DONOR_REVISION in ( + SIPP_2023_VOLUNTARY_FILING_DONOR_URL + ) + assert VOLUNTARY_FILING_ARCHIVED_DERIVATION_URL.endswith( + "datasets/cps/cps.py#L726-L747" + ) + assert VOLUNTARY_FILING_ARCHIVED_PARAMETERS_URL.endswith( + "parameters/take_up/voluntary_filing.yaml#L1-L43" + ) + assert VOLUNTARY_FILING_SIPP_DICTIONARY_URL.endswith( + "2023_SIPP_Data_Dictionary.pdf" + ) + + +def test_exact_source_columns_predictors_outputs_and_manifest_stage() -> None: + assert SIPP_VOLUNTARY_FILING_SOURCE_COLUMNS == ( + "SSUID", + "PNUM", + "MONTHCODE", + "WPFINWGT", + "TAGE", + "ESEX", + "EPNSPOUSE", + "AFILING", + "EFILING", + "AWILLFILE", + "EWILLFILE", + "EDEPCLM", + "TJB1_MSUM", + "TJB2_MSUM", + "TJB3_MSUM", + "TJB4_MSUM", + "TJB5_MSUM", + "TJB6_MSUM", + "TJB7_MSUM", + ) + assert SIPP_VOLUNTARY_FILING_MODEL_PREDICTORS == ( + "employment_income", + "reference_age", + "reference_is_female", + "reference_is_married", + "count_under_18", + ) + assert US_VOLUNTARY_FILING_OUTPUT_COLUMNS == (_OUTPUT,) + assert US_VOLUNTARY_FILING_NONCONSTANT_TAX_UNIT_COLUMNS == (_OUTPUT,) + spec = us_voluntary_filing_stage_spec() + assert spec.stage == "voluntary_filing_input" + assert spec.grain == "tax_unit" + assert spec.outputs == (_OUTPUT,) + assert [operation.kind for operation in spec.operations] == [ + "read_table", + "fit_weighted_qrf", + ] + + +def test_loader_uses_reported_answers_drops_dependents_and_pairs_spouses( + tmp_path: Path, +) -> None: + path = _write_source(tmp_path, _synthetic_source_rows()) + + donor = load_sipp_2023_voluntary_filing_donor( + path, expected_size_bytes=None + ).set_index("source_tax_unit_key") + + assert len(donor) == 2 + married_key = next(key for key in donor.index if str(key).startswith("1:")) + singleton_key = next(key for key in donor.index if str(key).startswith("2:")) + assert bool(donor.loc[married_key, _OUTPUT]) + assert not bool(donor.loc[singleton_key, _OUTPUT]) + assert donor.loc[married_key, "employment_income"] == pytest.approx(18_000.0) + assert donor.loc[married_key, "reference_age"] == pytest.approx(40.0) + assert donor.loc[married_key, "reference_is_female"] == pytest.approx(0.0) + assert donor.loc[married_key, "reference_is_married"] == pytest.approx(1.0) + assert donor.loc[married_key, "count_under_18"] == pytest.approx(1.0) + assert donor.loc[married_key, "tax_unit_weight"] == pytest.approx(10.0) + # Total income is deliberately absent: only TJB*_MSUM feeds wages. + assert "TPTOTINC" not in SIPP_VOLUNTARY_FILING_SOURCE_COLUMNS + # The November extreme and the dependent/imputed/zero-weight rows vanished. + assert donor["employment_income"].max() < 100_000.0 + + +def test_loader_rejects_reciprocal_spouse_target_disagreement(tmp_path: Path) -> None: + rows = [ + _source_row(1, 101, spouse=102, filing=1), + _source_row( + 1, + 102, + spouse=101, + filing=2, + will_file=2, + will_file_status=1, + ), + ] + path = _write_source(tmp_path, rows) + with pytest.raises(ValueError, match="spouses disagree"): + load_sipp_2023_voluntary_filing_donor(path, expected_size_bytes=None) + + +def test_loader_reference_is_minimum_pnum_before_response_filter( + tmp_path: Path, +) -> None: + rows = [ + _source_row( + 1, + 101, + spouse=102, + age=63, + sex=2, + weight=31, + filing=np.nan, + filing_status=0, + monthly_wages=100, + ), + _source_row( + 1, + 102, + spouse=101, + age=41, + sex=1, + weight=97, + filing=1, + monthly_wages=200, + ), + _source_row( + 2, + 101, + filing=2, + will_file=2, + will_file_status=1, + ), + ] + donor = load_sipp_2023_voluntary_filing_donor( + _write_source(tmp_path, rows), expected_size_bytes=None + ).set_index("source_tax_unit_key") + married = donor.loc[next(key for key in donor.index if str(key).startswith("1:"))] + + assert married["reference_age"] == pytest.approx(63) + assert married["reference_is_female"] == pytest.approx(1) + assert married["tax_unit_weight"] == pytest.approx(31) + assert married["employment_income"] == pytest.approx((100 + 200) * 12) + + +def test_loader_rejects_missing_columns_bad_hash_and_constant_target( + tmp_path: Path, +) -> None: + path = _write_source(tmp_path, _synthetic_source_rows()) + with pytest.raises(ValueError, match="sha-256 verification"): + load_sipp_2023_voluntary_filing_donor( + path, + expected_sha256="0" * 64, + expected_size_bytes=None, + ) + + missing = tmp_path / "missing.csv" + pd.DataFrame({"SSUID": [1], "MONTHCODE": [12]}).to_csv( + missing, sep="|", index=False + ) + with pytest.raises(ValueError, match="missing column"): + load_sipp_2023_voluntary_filing_donor(missing, expected_size_bytes=None) + + constant = tmp_path / "constant.csv" + _write_source( + tmp_path, + [_source_row(10, 101), _source_row(20, 101)], + ).replace(constant) + with pytest.raises(ValueError, match="target is constant"): + load_sipp_2023_voluntary_filing_donor(constant, expected_size_bytes=None) + + +def test_cached_full_donor_matches_locked_response_and_weight_facts() -> None: + snapshot = ( + Path.home() + / ".cache" + / "huggingface" + / "hub" + / ("models--policyengine--policyengine-" + "us-data") + / "snapshots" + / SIPP_2023_VOLUNTARY_FILING_DONOR_REVISION + / "pu2023.csv" + ) + if not snapshot.is_file(): + pytest.skip("the 3.73 GB pinned SIPP donor is not mounted") + + donor = load_sipp_2023_voluntary_filing_donor( + snapshot, + expected_sha256=SIPP_2023_VOLUNTARY_FILING_DONOR_SHA256, + expected_size_bytes=SIPP_2023_VOLUNTARY_FILING_DONOR_SIZE_BYTES, + ) + weights = donor["tax_unit_weight"].to_numpy(dtype=np.float64) + target = donor[_OUTPUT].to_numpy(dtype=bool) + audit = donor.attrs["source_audit"] + assert audit["december_rows"] == 39_513 + assert audit["observed_response_rows"] == 30_510 + assert audit["observed_response_true_rows"] == 24_473 + assert audit["claimed_dependent_observed_rows"] == 526 + assert audit["claimed_dependent_observed_true_rows"] == 526 + assert audit["spouse_target_disagreement_units"] == 0 + assert audit["canonical_preweight_units"] == 22_313 + assert audit["canonical_preweight_true_units"] == 16_820 + assert audit["positive_finite_weight_units"] == 22_296 + assert audit["positive_finite_weight_true_units"] == 16_817 + assert len(donor) == 22_296 + assert int(target.sum()) == 16_817 + assert float(weights.sum()) == pytest.approx(178_696_583.4878655, abs=1e-4) + assert float(weights[target].sum() / weights.sum()) == pytest.approx( + 0.760308456312741, abs=1e-12 + ) + assert audit["positive_finite_weight_sum"] == pytest.approx( + 178_696_583.4878655, abs=1e-4 + ) + assert audit["weighted_true_share"] == pytest.approx(0.760308456312741, abs=1e-12) + + +class _ChunkedResponse(io.BytesIO): + def __init__(self, payload: bytes) -> None: + super().__init__(payload) + self.read_sizes: list[int] = [] + + def read(self, size: int = -1) -> bytes: + self.read_sizes.append(size) + return super().read(size) + + def __enter__(self): + return self + + def __exit__(self, exc_type, exc_value, traceback) -> None: + self.close() + + +def test_fetch_streams_verifies_atomically_and_reuses_cache( + tmp_path: Path, monkeypatch: pytest.MonkeyPatch +) -> None: + payload = b"small synthetic pinned filing donor" + digest = hashlib.sha256(payload).hexdigest() + response = _ChunkedResponse(payload) + monkeypatch.setattr(urllib.request, "urlopen", lambda *_args: response) + + path = fetch_sipp_2023_voluntary_filing_donor( + tmp_path, + expected_sha256=digest, + expected_size_bytes=len(payload), + chunk_size=4, + ) + assert path.read_bytes() == payload + assert len(response.read_sizes) > 2 + assert all(size == 4 for size in response.read_sizes) + assert not (tmp_path / "pu2023.csv.part").exists() + + monkeypatch.setattr( + urllib.request, + "urlopen", + lambda *_args: (_ for _ in ()).throw( + AssertionError("valid cache must be reused") + ), + ) + assert ( + fetch_sipp_2023_voluntary_filing_donor( + tmp_path, + expected_sha256=digest, + expected_size_bytes=len(payload), + chunk_size=4, + ) + == path + ) + + +def test_receiver_uses_unit_wages_head_spouse_and_full_household_children() -> None: + frame = _frame(6) + receiver = module._recipient_tax_unit_predictor_table(frame) + + # Household/tax unit 6 has a head, spouse, and child. + row = receiver.loc[106] + assert row["employment_income"] == pytest.approx(20_000.0) + assert row["reference_age"] == pytest.approx(31.0) + assert row["reference_is_female"] == pytest.approx(0.0) + assert row["reference_is_married"] == pytest.approx(1.0) + assert row["count_under_18"] == pytest.approx(1.0) + # Household 5 has neither spouse nor child. + assert receiver.loc[105, "reference_is_married"] == pytest.approx(0.0) + assert receiver.loc[105, "count_under_18"] == pytest.approx(0.0) + + +def test_qrf_predicts_once_per_source_unit_and_fans_out_identical_clones( + monkeypatch: pytest.MonkeyPatch, +) -> None: + expanded = clone_us_frame_for_puf_support(_frame(10)) + calls: dict[str, object] = {} + + class FakeFitted: + def predict(self, receiver: pd.DataFrame) -> pd.DataFrame: + calls["receiver"] = receiver.copy() + return pd.DataFrame( + {_OUTPUT: (np.arange(len(receiver)) % 4) != 0}, + index=receiver.index, + ) + + class FakeQRF: + def __init__(self, **kwargs: object) -> None: + calls["init"] = kwargs + + def fit( + self, + training: pd.DataFrame, + *, + predictors: list[str], + targets: list[str], + weights: np.ndarray, + ) -> FakeFitted: + calls["training"] = training.copy() + calls["predictors"] = predictors + calls["targets"] = targets + calls["weights"] = weights.copy() + return FakeFitted() + + monkeypatch.setattr(module, "QRF", FakeQRF) + predicted = impute_us_voluntary_filing(expanded, _donor(), seed=17, n_estimators=9) + + assert calls["init"] == {"n_estimators": 9, "seed": 17} + assert calls["predictors"] == list(SIPP_VOLUNTARY_FILING_MODEL_PREDICTORS) + assert calls["targets"] == [_OUTPUT] + assert len(calls["receiver"]) == 10 + assert len(predicted) == 20 + tax_unit = expanded.table("tax_unit") + by_source = pd.DataFrame( + { + "source": tax_unit["tax_unit_source_id"], + "predicted": predicted.to_numpy(), + } + ).groupby("source")["predicted"] + assert (by_source.nunique() == 1).all() + + +def test_real_qrf_recomputation_is_deterministic() -> None: + frame = _frame(14) + donor = _donor(120) + first = impute_us_voluntary_filing(frame, donor, seed=31, n_estimators=8) + second = impute_us_voluntary_filing( + frame, + donor.sample(frac=1.0, random_state=99), + seed=31, + n_estimators=8, + ) + pd.testing.assert_series_equal(first, second) + + +def test_wrapper_recomputes_stale_signal_and_is_idempotent( + monkeypatch: pytest.MonkeyPatch, +) -> None: + frame = _replace_tax_unit( + _frame(10), + **{_OUTPUT: np.asarray([True, False] * 5)}, + ) + expected = pd.Series( + np.asarray([False, True, True, True, True] * 2), + index=frame.table("tax_unit").index, + name=_OUTPUT, + ) + monkeypatch.setattr(module, "us_voluntary_filing_stage_spec", lambda: object()) + monkeypatch.setattr( + module, + "impute_us_voluntary_filing", + lambda *_args, **_kwargs: expected, + ) + + restored = with_us_voluntary_filing_input( + frame, + seed=1, + time_period=2024, + sipp_donor=_donor(), + ) + assert np.array_equal(restored.table("tax_unit")[_OUTPUT], expected) + assert restored is not frame + repeated = with_us_voluntary_filing_input( + restored, + seed=1, + time_period=2024, + sipp_donor=_donor(), + ) + assert repeated is restored + + +def test_signal_gate_requires_boolean_plausible_and_clone_consistent() -> None: + expanded = clone_us_frame_for_puf_support(_frame(10)) + values = np.tile( + np.asarray([False, True, True, True, True, False, True, True, False, True]), + 2, + ) + valid = _replace_tax_unit(expanded, **{_OUTPUT: values}) + summary = us_voluntary_filing_summary(valid) + assert summary["clone_source_units"] == 10 + assert summary["clone_mismatch_source_units"] == 0 + gate = us_voluntary_filing_signal_gate(valid) + assert gate.passed, gate.failures + + mismatched_values = values.copy() + mismatched_values[10] = ~mismatched_values[0] + mismatch = _replace_tax_unit(expanded, **{_OUTPUT: mismatched_values}) + mismatch_gate = us_voluntary_filing_signal_gate(mismatch) + assert not mismatch_gate.passed + assert any("clones disagree" in failure for failure in mismatch_gate.failures) + + constant = _replace_tax_unit(_frame(10), **{_OUTPUT: np.ones(10, dtype=bool)}) + constant_gate = us_voluntary_filing_signal_gate(constant) + assert not constant_gate.passed + assert any("constant" in failure for failure in constant_gate.failures) + + +@requires_us +def test_policyengine_1_764_6_contract_and_aca_ptc_neutralization() -> None: + from policyengine_us import CountryTaxBenefitSystem, Simulation + + from populace.build.us_runtime.reform_coverage_smoke import _build_reform + + assert version("policyengine-us") == "1.764.6" + variable = CountryTaxBenefitSystem().variables[_OUTPUT] + assert variable.is_input_variable() + assert variable.entity.key == "tax_unit" + assert variable.value_type is bool + assert variable.default_value is False + + situation = { + "people": { + "adult": { + "age": {2024: 45}, + # Isolate the filing gate from unrelated ACA eligibility facts. + "is_aca_ptc_eligible": {2024: True}, + } + }, + "tax_units": { + "unit": { + "members": ["adult"], + "filing_status": {2024: "SINGLE"}, + _OUTPUT: {2024: True}, + "tax_unit_is_required_to_file": {2024: False}, + "eligible_for_refundable_credits": {2024: False}, + "would_file_if_eligible_for_refundable_credit": {2024: False}, + "slcsp": {"2024-01": 11_600}, + "aca_magi": {2024: 0}, + "aca_required_contribution_percentage": {2024: 0}, + } + }, + "families": {"family": {"members": ["adult"]}}, + "spm_units": {"spm": {"members": ["adult"]}}, + "households": { + "household": { + "members": ["adult"], + "state_code_str": {2024: "CA"}, + } + }, + "marital_units": {"marital": {"members": ["adult"]}}, + } + probe = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "voluntary_filing_aca_ptc_neutralization" + ) + assert probe.neutralized_variable == _OUTPUT + assert probe.binding_inputs == (_OUTPUT,) + assert probe.budget_measure == "aca_ptc" + assert probe.min_abs_effect == 100_000_000.0 + baseline = Simulation(situation=situation) + neutralized = Simulation(situation=situation, reform=_build_reform(probe)) + assert baseline.calculate("tax_unit_is_filer", 2024)[0] + assert not neutralized.calculate("tax_unit_is_filer", 2024)[0] + assert baseline.calculate("aca_ptc", 2024)[0] == pytest.approx(11_600.0) + assert neutralized.calculate("aca_ptc", 2024)[0] == 0.0 + + +@pytest.mark.parametrize( + ("column", "value", "message"), + [ + ("employment_income", np.nan, "finite and nonnegative"), + ("tax_unit_weight", 0.0, "finite and positive"), + (_OUTPUT, 2, "target must be boolean"), + ], +) +def test_imputer_fails_closed_on_invalid_donor( + column: str, value: float, message: str +) -> None: + donor = _donor() + if column == _OUTPUT: + donor[column] = donor[column].astype(np.int8) + donor.loc[0, column] = value + with pytest.raises(ValueError, match=message): + impute_us_voluntary_filing(_frame(), donor, seed=0, n_estimators=4) diff --git a/packages/populace-build/tests/test_us_weeks_unemployed.py b/packages/populace-build/tests/test_us_weeks_unemployed.py new file mode 100644 index 00000000..a5bff070 --- /dev/null +++ b/packages/populace-build/tests/test_us_weeks_unemployed.py @@ -0,0 +1,770 @@ +"""Exact ASEC repair and PUF-half treatment for weeks unemployed.""" + +from __future__ import annotations + +import hashlib +import io +import os +import zipfile +from pathlib import Path + +import numpy as np +import pandas as pd +import pytest + +import populace.build.us_runtime.weeks_unemployed as module +from populace.build.source_manifest import SourceOperationSpec, SourceStageSpec +from populace.build.source_runtime import ( + SourceRuntimeConfig, + SourceRuntimeContext, + SourceRuntimeError, +) +from populace.build.us_runtime.weeks_unemployed import ( + ASEC_2023_WEEKS_UNEMPLOYED_MEMBER, + ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_CRC32, + ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_SHA256, + ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_SIZE_BYTES, + ASEC_2023_WEEKS_UNEMPLOYED_POSITIVE_ROWS, + ASEC_2023_WEEKS_UNEMPLOYED_RAW_ROWS, + ASEC_2023_WEEKS_UNEMPLOYED_SOURCE_COLUMNS, + ASEC_2023_WEEKS_UNEMPLOYED_UNIQUE_KEYS, + ASEC_2023_WEEKS_UNEMPLOYED_WEIGHTED_SOURCE_SHARE, + ASEC_2023_WEEKS_UNEMPLOYED_WEIGHTED_WEEKS, + ASEC_2023_WEEKS_UNEMPLOYED_ZIP_SHA256, + ASEC_2023_WEEKS_UNEMPLOYED_ZIP_SIZE_BYTES, + WEEKS_UNEMPLOYED_DERIVE_PARAMETERS, + WEEKS_UNEMPLOYED_PUF_IMPUTATION_PARAMETERS, + WEEKS_UNEMPLOYED_READ_PARAMETERS, + derive_us_weeks_unemployed_from_manifest, + fetch_asec_2023_weeks_unemployed_source, + fill_asec_2022_weeks_unemployed_source, + impute_us_weeks_unemployed_to_puf_support_from_manifest, + load_asec_2023_weeks_unemployed_source, + us_weeks_unemployed_signal_gate, + us_weeks_unemployed_stage_spec, + with_us_weeks_unemployed, +) +from populace.frame import US_SCHEMA, Frame, WeightKind, Weights + +_OUTPUT = "weeks_unemployed" +_PREFIX = "weeks_unemployed_predictor_" +_REQUIRED_PREDICTORS = ( + "age", + "is_male", + "tax_unit_is_joint", + "is_tax_unit_head", + "is_tax_unit_spouse", + "is_tax_unit_dependent", +) + + +def _person_csv() -> bytes: + source = pd.DataFrame( + { + "PH_SEQ": [101, 102, 103], + "P_SEQ": [1, 1, 1], + "A_LINENO": [1, 1, 1], + "PERIDNUM": [f"{value:022d}" for value in (1, 2, 3)], + "LKWEEKS": [0, 4, -1], + "A_FNLWGT": [100, 200, 300], + } + ) + return source.to_csv(index=False).encode() + + +def _zip_bytes(member: bytes | None = None) -> bytes: + payload = io.BytesIO() + with zipfile.ZipFile(payload, "w", compression=zipfile.ZIP_DEFLATED) as archive: + archive.writestr(ASEC_2023_WEEKS_UNEMPLOYED_MEMBER, member or _person_csv()) + return payload.getvalue() + + +def _pins(payload: bytes, member: bytes) -> dict[str, object]: + with zipfile.ZipFile(io.BytesIO(payload)) as archive: + info = archive.getinfo(ASEC_2023_WEEKS_UNEMPLOYED_MEMBER) + return { + "expected_zip_size_bytes": len(payload), + "expected_zip_sha256": hashlib.sha256(payload).hexdigest(), + "expected_member_size_bytes": len(member), + "expected_member_crc32": f"{info.CRC:08x}", + "expected_member_sha256": hashlib.sha256(member).hexdigest(), + } + + +def _mini_source(tmp_path: Path) -> tuple[Path, dict[str, object]]: + member = _person_csv() + payload = _zip_bytes(member) + path = tmp_path / "asecpub23csv.zip" + path.write_bytes(payload) + return path, _pins(payload, member) + + +def _stage_spec() -> SourceStageSpec: + return SourceStageSpec( + stage="weeks_unemployed_input", + survey="Census CPS ASEC", + source="official fixture", + grain="person", + artifacts=(), + operations=( + SourceOperationSpec("read_table", WEEKS_UNEMPLOYED_READ_PARAMETERS), + SourceOperationSpec( + "derive_weeks_unemployed", WEEKS_UNEMPLOYED_DERIVE_PARAMETERS + ), + SourceOperationSpec( + "impute_weeks_unemployed_to_puf_support", + WEEKS_UNEMPLOYED_PUF_IMPUTATION_PARAMETERS, + ), + ), + outputs=(_OUTPUT,), + nonnegative_outputs=(_OUTPUT,), + ) + + +def _operation(kind: str) -> SourceOperationSpec: + spec = _stage_spec() + return next(operation for operation in spec.operations if operation.kind == kind) + + +def _frame(*, channels: bool = True) -> Frame: + person_count = 4 if channels else 2 + person = pd.DataFrame( + { + "person_id": np.arange(1, person_count + 1, dtype=np.int64), + "person_household_id": np.arange(11, 11 + person_count, dtype=np.int64), + "person_tax_unit_id": np.arange(21, 21 + person_count, dtype=np.int64), + "person_spm_unit_id": np.arange(31, 31 + person_count, dtype=np.int64), + "person_family_id": np.arange(41, 41 + person_count, dtype=np.int64), + "person_marital_unit_id": np.arange(51, 51 + person_count, dtype=np.int64), + "source_year": np.resize([2022, 2023], person_count), + "source_household_id": np.arange(101, 101 + person_count, dtype=np.int64), + "P_SEQ": np.ones(person_count, dtype=np.int64), + "A_LINENO": np.ones(person_count, dtype=np.int64), + "PERIDNUM": [f"{value:022d}" for value in range(1, person_count + 1)], + "LKWEEKS": np.resize([2, 4], person_count), + "age": np.arange(30, 30 + person_count, dtype=np.int64), + "is_female": np.resize([False, True], person_count), + "tax_unit_role_input": np.resize(["HEAD", "SPOUSE"], person_count), + "unemployment_compensation": np.resize([100.0, 0.0], person_count), + } + ) + if channels: + person["person_support_channel"] = [ + "asec", + "asec", + "puf_tax_detail", + "puf_tax_detail", + ] + entity_ids = { + "household": person["person_household_id"].to_numpy(), + "tax_unit": person["person_tax_unit_id"].to_numpy(), + "spm_unit": person["person_spm_unit_id"].to_numpy(), + "family": person["person_family_id"].to_numpy(), + "marital_unit": person["person_marital_unit_id"].to_numpy(), + } + tables = { + entity: pd.DataFrame({f"{entity}_id": ids}) + for entity, ids in entity_ids.items() + } + tables["person"] = person + tables["tax_unit"]["filing_status_input"] = np.resize( + ["SINGLE", "JOINT"], person_count + ) + return Frame( + tables, + US_SCHEMA, + { + "household": Weights( + np.ones(person_count, dtype=np.float64), WeightKind.DESIGN + ) + }, + ) + + +def _gate_frame() -> Frame: + asec_rows = 1_000 + puf_rows = 2_000 + person_count = asec_rows + puf_rows + asec_weeks = np.zeros(asec_rows, dtype=np.float64) + asec_weeks[:30] = 17.0 + puf_weeks = np.zeros(puf_rows, dtype=np.float64) + puf_weeks[:12] = 16.0 + unemployment_compensation = np.zeros(person_count, dtype=np.float64) + unemployment_compensation[asec_rows : asec_rows + 12] = 100.0 + person = pd.DataFrame( + { + "person_id": np.arange(1, person_count + 1, dtype=np.int64), + "person_household_id": np.arange( + 10_001, 10_001 + person_count, dtype=np.int64 + ), + "person_tax_unit_id": np.arange( + 20_001, 20_001 + person_count, dtype=np.int64 + ), + "person_spm_unit_id": np.arange( + 30_001, 30_001 + person_count, dtype=np.int64 + ), + "person_family_id": np.arange( + 40_001, 40_001 + person_count, dtype=np.int64 + ), + "person_marital_unit_id": np.arange( + 50_001, 50_001 + person_count, dtype=np.int64 + ), + "person_support_channel": ["asec"] * asec_rows + + ["puf_tax_detail"] * puf_rows, + "LKWEEKS": np.concatenate( + [asec_weeks, np.zeros(puf_rows, dtype=np.float64)] + ), + _OUTPUT: np.concatenate([asec_weeks, puf_weeks]), + "unemployment_compensation": unemployment_compensation, + } + ) + entity_links = { + "household": "person_household_id", + "tax_unit": "person_tax_unit_id", + "spm_unit": "person_spm_unit_id", + "family": "person_family_id", + "marital_unit": "person_marital_unit_id", + } + tables = { + entity: pd.DataFrame( + {f"{entity}_id": person[column].to_numpy(dtype=np.int64)} + ) + for entity, column in entity_links.items() + } + tables["person"] = person + return Frame( + tables, + US_SCHEMA, + { + "household": Weights( + np.ones(person_count, dtype=np.float64), WeightKind.DESIGN + ) + }, + ) + + +def test_public_stage_contract_is_exactly_manifest_pinned() -> None: + spec = us_weeks_unemployed_stage_spec() + + assert spec.stage == "weeks_unemployed_input" + assert [operation.kind for operation in spec.operations] == [ + "read_table", + "derive_weeks_unemployed", + "impute_weeks_unemployed_to_puf_support", + ] + assert dict(spec.operations[0].parameters) == WEEKS_UNEMPLOYED_READ_PARAMETERS + assert dict(spec.operations[1].parameters) == WEEKS_UNEMPLOYED_DERIVE_PARAMETERS + assert ( + dict(spec.operations[2].parameters) + == WEEKS_UNEMPLOYED_PUF_IMPUTATION_PARAMETERS + ) + artifact = next( + item + for item in spec.artifacts + if item.get("member") == ASEC_2023_WEEKS_UNEMPLOYED_MEMBER + ) + assert artifact["size_bytes"] == ASEC_2023_WEEKS_UNEMPLOYED_ZIP_SIZE_BYTES + assert artifact["sha256"] == ASEC_2023_WEEKS_UNEMPLOYED_ZIP_SHA256 + assert artifact["member_size_bytes"] == ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_SIZE_BYTES + assert artifact["member_crc32"] == ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_CRC32 + assert artifact["member_sha256"] == ASEC_2023_WEEKS_UNEMPLOYED_MEMBER_SHA256 + + +def test_loader_verifies_zip_member_identity_and_attaches_audit(tmp_path: Path) -> None: + path, pins = _mini_source(tmp_path) + + source = load_asec_2023_weeks_unemployed_source( + path, + **pins, + expected_rows=3, + expected_unique_keys=3, + ) + + assert tuple(source.columns) == ASEC_2023_WEEKS_UNEMPLOYED_SOURCE_COLUMNS + assert source["PERIDNUM"].str.len().eq(22).all() + assert source["LKWEEKS"].tolist() == [0, 4, -1] + assert source.attrs["source_audit"] == { + "raw_rows": 3, + "unique_keys": 3, + "positive_rows": 1, + "niu_rows": 1, + "minimum": -1, + "maximum": 4, + "weighted_source_share": pytest.approx(1 / 3), + "weighted_weeks": pytest.approx(8.0), + "pinned_transform": 0, + } + + +def test_fetch_verifies_both_parent_and_member_and_reuses_cache( + monkeypatch: pytest.MonkeyPatch, + tmp_path: Path, +) -> None: + member = _person_csv() + payload = _zip_bytes(member) + pins = _pins(payload, member) + calls = 0 + + def _urlopen(_request: object, *, timeout: int) -> io.BytesIO: + nonlocal calls + assert timeout == 180 + calls += 1 + return io.BytesIO(payload) + + monkeypatch.setattr(module.urllib.request, "urlopen", _urlopen) + result = fetch_asec_2023_weeks_unemployed_source(tmp_path, **pins, chunk_size=17) + assert result.read_bytes() == member + assert calls == 1 + + result_again = fetch_asec_2023_weeks_unemployed_source( + tmp_path, **pins, chunk_size=17 + ) + assert result_again == result + assert calls == 1 + + +@pytest.mark.parametrize("tamper", ["parent", "member"]) +def test_loader_rejects_tampered_source( + tmp_path: Path, + tamper: str, +) -> None: + path, pins = _mini_source(tmp_path) + if tamper == "parent": + path.write_bytes(path.read_bytes() + b"tamper") + else: + pins["expected_member_sha256"] = "0" * 64 + + with pytest.raises(ValueError, match="SHA-256|byte length"): + load_asec_2023_weeks_unemployed_source( + path, + **pins, + expected_rows=3, + expected_unique_keys=3, + ) + + +def test_loader_rejects_duplicate_or_non_fixed_width_person_keys( + tmp_path: Path, +) -> None: + raw = pd.read_csv(io.BytesIO(_person_csv()), dtype={"PERIDNUM": "string"}) + raw.loc[1, "PERIDNUM"] = raw.loc[0, "PERIDNUM"] + member = raw.to_csv(index=False).encode() + payload = _zip_bytes(member) + path = tmp_path / "duplicate.zip" + path.write_bytes(payload) + + with pytest.raises(ValueError, match="unique|duplicate"): + load_asec_2023_weeks_unemployed_source( + path, + **_pins(payload, member), + expected_rows=3, + expected_unique_keys=3, + ) + + raw.loc[1, "PERIDNUM"] = "123" + member = raw.to_csv(index=False).encode() + payload = _zip_bytes(member) + path.write_bytes(payload) + with pytest.raises(ValueError, match="22-digit"): + load_asec_2023_weeks_unemployed_source( + path, + **_pins(payload, member), + expected_rows=3, + expected_unique_keys=3, + ) + + +def test_fill_repairs_only_2022_and_accepts_cloned_duplicate_peridnum() -> None: + source = pd.DataFrame( + { + "PH_SEQ": [101, 102], + "P_SEQ": [1, 1], + "A_LINENO": [1, 1], + "PERIDNUM": [f"{value:022d}" for value in (1, 2)], + "LKWEEKS": [7, 9], + } + ) + person = pd.DataFrame( + { + "source_year": [2022, 2023, 2022], + "source_household_id": [101, 999, 101], + "P_SEQ": [1, 1, 1], + "A_LINENO": [1, 1, 1], + "PERIDNUM": [f"{value:022d}" for value in (1, 3, 1)], + "LKWEEKS": [np.nan, 11.0, np.nan], + "person_support_channel": ["asec", "asec", "puf_tax_detail"], + } + ) + + result = fill_asec_2022_weeks_unemployed_source(person, source) + + assert result["LKWEEKS"].tolist() == [7.0, 11.0, 7.0] + assert person["LKWEEKS"].isna().sum() == 2 + + +def test_fill_prefers_raw_source_household_id_over_transformed_ph_seq() -> None: + source = pd.DataFrame( + { + "PH_SEQ": [32], + "P_SEQ": [1], + "A_LINENO": [1], + "PERIDNUM": ["0000000000000000000001"], + "LKWEEKS": [7], + } + ) + person = pd.DataFrame( + { + "source_year": [2022], + # Pooled PH_SEQ is globally transformed; source_household_id is raw. + "PH_SEQ": [112_033], + "source_household_id": [32], + "P_SEQ": [1], + "A_LINENO": [1], + "PERIDNUM": ["0000000000000000000001"], + "LKWEEKS": [np.nan], + } + ) + + result = fill_asec_2022_weeks_unemployed_source(person, source) + + assert result["LKWEEKS"].tolist() == [7.0] + + +@pytest.mark.parametrize( + ("mutation", "message"), + [ + ({"PERIDNUM": "0000000000000000000099"}, "does not cover"), + ({"source_household_id": 999}, "identity mismatch"), + ({"LKWEEKS": 8.0}, "disagrees"), + ], +) +def test_fill_fails_closed_on_key_identity_or_existing_value_mismatch( + mutation: dict[str, object], + message: str, +) -> None: + source = pd.DataFrame( + { + "PH_SEQ": [101], + "P_SEQ": [1], + "A_LINENO": [1], + "PERIDNUM": ["0000000000000000000001"], + "LKWEEKS": [7], + } + ) + row: dict[str, object] = { + "source_year": 2022, + "source_household_id": 101, + "P_SEQ": 1, + "A_LINENO": 1, + "PERIDNUM": "0000000000000000000001", + "LKWEEKS": np.nan, + } + row.update(mutation) + + with pytest.raises(ValueError, match=message): + fill_asec_2022_weeks_unemployed_source(pd.DataFrame([row]), source) + + +def test_direct_carry_maps_only_minus_one_and_rejects_invalid_values() -> None: + operation = _operation("derive_weeks_unemployed") + source = pd.DataFrame({"LKWEEKS": [-1, 0, 1, 51, 52]}) + + result = derive_us_weeks_unemployed_from_manifest(source, operation, None) + + assert result[_OUTPUT].tolist() == [0, 0, 1, 51, 52] + for bad in (np.nan, np.inf, -2, 53, 1.5): + with pytest.raises(SourceRuntimeError, match="integer -1 or 0--52"): + derive_us_weeks_unemployed_from_manifest( + pd.DataFrame({"LKWEEKS": [bad]}), operation, None + ) + + +class _CapturingQRF: + instances: list[_CapturingQRF] = [] + prediction = pd.DataFrame({_OUTPUT: [1.6, 80.4]}) + + def __init__(self, *, n_estimators: int, seed: int) -> None: + self.n_estimators = n_estimators + self.seed = seed + self.training: pd.DataFrame | None = None + self.predictors: list[str] = [] + self.weights: np.ndarray | None = None + _CapturingQRF.instances.append(self) + + def fit( + self, + training: pd.DataFrame, + predictors: list[str], + targets: list[str], + *, + weights: np.ndarray, + ) -> _CapturingQRF: + assert targets == [_OUTPUT] + self.training = training.copy() + self.predictors = predictors + self.weights = weights.copy() + return self + + def predict(self, test: pd.DataFrame) -> pd.DataFrame: + assert list(test.columns) == self.predictors + return self.prediction.copy() + + +def _imputation_table(*, include_uc: bool = True) -> pd.DataFrame: + frame = pd.DataFrame( + { + "person_support_channel": [ + "asec", + "asec", + "puf_tax_detail", + "puf_tax_detail", + ], + "person_weight": [2.0, 3.0, 4.0, 5.0], + _OUTPUT: [0.0, 6.0, 0.0, 6.0], + } + ) + for index, predictor in enumerate(_REQUIRED_PREDICTORS): + frame[_PREFIX + predictor] = np.arange(4) + index + if include_uc: + frame[_PREFIX + "unemployment_compensation"] = [0.0, 100.0, 50.0, 0.0] + return frame + + +def test_puf_qrf_uses_seed_weights_round_clip_and_uc_zero_rule( + monkeypatch: pytest.MonkeyPatch, +) -> None: + _CapturingQRF.instances.clear() + _CapturingQRF.prediction = pd.DataFrame({_OUTPUT: [1.6, 80.4]}) + monkeypatch.setattr(module, "QRF", _CapturingQRF) + context = SourceRuntimeContext( + SourceRuntimeConfig(seed=918, target_year=2024), tables={} + ) + + result = impute_us_weeks_unemployed_to_puf_support_from_manifest( + _imputation_table(), + _operation("impute_weeks_unemployed_to_puf_support"), + context, + ) + + assert result[_OUTPUT].tolist() == [0.0, 6.0, 2.0, 0.0] + fitted = _CapturingQRF.instances[-1] + assert (fitted.n_estimators, fitted.seed) == (100, 918) + assert fitted.predictors == [*_REQUIRED_PREDICTORS, "unemployment_compensation"] + assert fitted.weights.tolist() == [2.0, 3.0] + + +def test_puf_qrf_omits_uc_predictor_and_rule_when_source_is_unavailable( + monkeypatch: pytest.MonkeyPatch, +) -> None: + _CapturingQRF.instances.clear() + _CapturingQRF.prediction = pd.DataFrame({_OUTPUT: [-2.4, 99.0]}) + monkeypatch.setattr(module, "QRF", _CapturingQRF) + + result = impute_us_weeks_unemployed_to_puf_support_from_manifest( + _imputation_table(include_uc=False), + _operation("impute_weeks_unemployed_to_puf_support"), + SourceRuntimeContext(SourceRuntimeConfig(seed=1), tables={}), + ) + + assert result[_OUTPUT].tolist() == [0.0, 6.0, 0.0, 52.0] + assert _CapturingQRF.instances[-1].predictors == list(_REQUIRED_PREDICTORS) + + +@pytest.mark.parametrize( + ("prediction", "message"), + [ + (pd.DataFrame({_OUTPUT: [np.nan, 2.0]}), "nonfinite"), + (pd.DataFrame({"wrong": [1.0, 2.0]}), "missing"), + (pd.DataFrame({_OUTPUT: [1.0]}), "expected 2"), + ], +) +def test_puf_qrf_rejects_adversarial_predictions( + monkeypatch: pytest.MonkeyPatch, + prediction: pd.DataFrame, + message: str, +) -> None: + _CapturingQRF.prediction = prediction + monkeypatch.setattr(module, "QRF", _CapturingQRF) + + with pytest.raises(SourceRuntimeError, match=message): + impute_us_weeks_unemployed_to_puf_support_from_manifest( + _imputation_table(), + _operation("impute_weeks_unemployed_to_puf_support"), + SourceRuntimeContext(SourceRuntimeConfig(seed=1), tables={}), + ) + + +def test_puf_qrf_caps_training_at_5000_with_build_seed( + monkeypatch: pytest.MonkeyPatch, +) -> None: + row_count = 5_003 + frame = pd.DataFrame( + { + "person_support_channel": ["asec"] * row_count + ["puf_tax_detail"] * 2, + "person_weight": np.arange(1, row_count + 3, dtype=np.float64), + _OUTPUT: np.resize([0.0, 4.0], row_count + 2), + } + ) + for index, predictor in enumerate(_REQUIRED_PREDICTORS): + frame[_PREFIX + predictor] = np.arange(row_count + 2) + index + _CapturingQRF.instances.clear() + _CapturingQRF.prediction = pd.DataFrame({_OUTPUT: [1.0, 2.0]}) + monkeypatch.setattr(module, "QRF", _CapturingQRF) + + impute_us_weeks_unemployed_to_puf_support_from_manifest( + frame, + _operation("impute_weeks_unemployed_to_puf_support"), + SourceRuntimeContext(SourceRuntimeConfig(seed=42), tables={}), + ) + + fitted = _CapturingQRF.instances[-1] + assert fitted.training is not None + assert len(fitted.training) == 5_000 + assert fitted.weights is not None + assert len(fitted.weights) == 5_000 + + +def test_wrapper_runs_before_and_after_support_cloning( + monkeypatch: pytest.MonkeyPatch, +) -> None: + monkeypatch.setattr(module, "us_weeks_unemployed_stage_spec", _stage_spec) + direct = with_us_weeks_unemployed(_frame(channels=False), seed=7, time_period=2024) + assert direct.table("person")[_OUTPUT].tolist() == [2, 4] + + _CapturingQRF.prediction = pd.DataFrame({_OUTPUT: [3.2, 9.8]}) + monkeypatch.setattr(module, "QRF", _CapturingQRF) + supported = with_us_weeks_unemployed(_frame(), seed=7, time_period=2024) + assert supported.table("person")[_OUTPUT].tolist() == [2.0, 4.0, 3.0, 0.0] + + +def test_signal_gate_requires_exact_asec_and_nondefault_integer_both_channels() -> ( + None +): + frame = _gate_frame() + assert us_weeks_unemployed_signal_gate(frame).passed + + person = frame.table("person").copy() + person.loc[0, _OUTPUT] = 3.0 + bad = module._replace_person_table(frame, person) + gate = us_weeks_unemployed_signal_gate(bad) + assert not gate.passed + assert any("reconciliation" in failure for failure in gate.failures) + + person = frame.table("person").copy() + person.loc[1_000, _OUTPUT] = 1.5 + bad = module._replace_person_table(frame, person) + gate = us_weeks_unemployed_signal_gate(bad) + assert not gate.passed + assert any("noninteger" in failure for failure in gate.failures) + + person = frame.table("person").copy() + person.loc[person["person_support_channel"].eq("puf_tax_detail"), _OUTPUT] = 0.0 + bad = module._replace_person_table(frame, person) + gate = us_weeks_unemployed_signal_gate(bad) + assert not gate.passed + assert any("puf_tax_detail" in failure for failure in gate.failures) + + +def test_signal_gate_rejects_collapsed_puf_share_and_weighted_weeks() -> None: + frame = _gate_frame() + person = frame.table("person").copy() + puf = person["person_support_channel"].eq("puf_tax_detail") + person.loc[puf, _OUTPUT] = 0.0 + first_puf = person.index[puf][0] + person.loc[first_puf, _OUTPUT] = 16.0 + collapsed_share = module._replace_person_table(frame, person) + + gate = us_weeks_unemployed_signal_gate(collapsed_share) + + assert not gate.passed + assert any("positive share" in failure for failure in gate.failures) + + person.loc[puf, _OUTPUT] = 0.0 + first_four = person.index[puf][:4] + person.loc[first_four, _OUTPUT] = 1.0 + person.loc[first_four, "unemployment_compensation"] = 100.0 + collapsed_weeks = module._replace_person_table(frame, person) + + gate = us_weeks_unemployed_signal_gate(collapsed_weeks) + + assert not gate.passed + assert not any("positive share" in failure for failure in gate.failures) + assert any("weighted mean weeks" in failure for failure in gate.failures) + + +def test_live_engine_graph_has_only_structurally_blocked_pa_uc_consumer() -> None: + policyengine_us = pytest.importorskip("policyengine_us") + system = policyengine_us.CountryTaxBenefitSystem() + + consumers = { + name + for name, variable in system.variables.items() + for formula in variable.formulas.values() + if _OUTPUT in formula.__code__.co_consts + } + assert consumers == {"pa_uc"} + weeks = system.variables[_OUTPUT] + assert weeks.default_value == 0 + assert not weeks.formulas + + # Pennsylvania UC cannot support an honest monetary neutralization probe + # until these independent direct inputs are also restored. All four are + # formula-less default-zero leaves in the pinned PolicyEngine-US graph. + for name in ( + "pa_uc_base_year_wages", + "pa_uc_highest_quarter_wages", + "pa_uc_credit_weeks", + "pa_uc_gross_weekly_earnings", + ): + variable = system.variables[name] + assert variable.default_value == 0 + assert not variable.formulas + + +def test_pinned_weighted_weeks_tolerates_only_binary_float_residue() -> None: + audit = { + "raw_rows": ASEC_2023_WEEKS_UNEMPLOYED_RAW_ROWS, + "unique_keys": ASEC_2023_WEEKS_UNEMPLOYED_UNIQUE_KEYS, + "positive_rows": ASEC_2023_WEEKS_UNEMPLOYED_POSITIVE_ROWS, + "weighted_source_share": ASEC_2023_WEEKS_UNEMPLOYED_WEIGHTED_SOURCE_SHARE, + "weighted_weeks": ASEC_2023_WEEKS_UNEMPLOYED_WEIGHTED_WEEKS + 2e-8, + } + module._assert_pinned_source_audit(audit) + audit["weighted_weeks"] = ASEC_2023_WEEKS_UNEMPLOYED_WEIGHTED_WEEKS + 2e-6 + with pytest.raises(ValueError, match="weighted_weeks"): + module._assert_pinned_source_audit(audit) + + +def test_optional_full_pinned_source_audit() -> None: + candidates = [ + Path(value).expanduser() + for value in ( + os.environ.get("POPULACE_ASEC_2023_WEEKS_SOURCE", ""), + str( + Path.home() + / ".cache" + / "populace" + / "cps" + / "asec_2023" + / "asecpub23csv.zip" + ), + "/private/tmp/asecpub23csv.zip", + ) + if value + ] + path = next((candidate for candidate in candidates if candidate.is_file()), None) + if path is None: + pytest.skip("pinned ASEC 2023 archive is not mounted") + + source = load_asec_2023_weeks_unemployed_source(path) + audit = source.attrs["source_audit"] + assert audit["pinned_transform"] + assert audit["raw_rows"] == ASEC_2023_WEEKS_UNEMPLOYED_RAW_ROWS + assert audit["unique_keys"] == ASEC_2023_WEEKS_UNEMPLOYED_UNIQUE_KEYS + assert audit["positive_rows"] == ASEC_2023_WEEKS_UNEMPLOYED_POSITIVE_ROWS + assert audit["weighted_source_share"] == pytest.approx( + ASEC_2023_WEEKS_UNEMPLOYED_WEIGHTED_SOURCE_SHARE, abs=1e-12 + ) + assert audit["weighted_weeks"] == pytest.approx( + ASEC_2023_WEEKS_UNEMPLOYED_WEIGHTED_WEEKS, abs=1e-6 + ) diff --git a/packages/populace-build/tests/test_us_wic_claim.py b/packages/populace-build/tests/test_us_wic_claim.py new file mode 100644 index 00000000..5b780f7a --- /dev/null +++ b/packages/populace-build/tests/test_us_wic_claim.py @@ -0,0 +1,585 @@ +"""Source-backed WIC claim-input restoration tests.""" + +from __future__ import annotations + +import importlib.util +from copy import deepcopy +from importlib.metadata import version +from importlib.resources import files + +import numpy as np +import pandas as pd +import pytest + +from populace.build.source_manifest import SourceOperationSpec +from populace.build.source_runtime import ( + SourceRuntimeConfig, + SourceRuntimeContext, + SourceRuntimeError, +) +from populace.build.us_runtime import ( + PUF_TAX_DETAIL_SUPPORT_CHANNEL, + US_DONORS, + US_PREGNANCY_STAGE_NAME, + US_PUF_SUPPORT_STAGE_NAME, + US_STAGE_NAMES, + US_WIC_CLAIM_NONCONSTANT_PERSON_COLUMNS, + US_WIC_CLAIM_OUTPUT_COLUMNS, + US_WIC_CLAIM_REQUIRED_SOURCE_COLUMNS, + US_WIC_CLAIM_STAGE_NAME, + WIC_CLAIM_ARCHIVED_DERIVATION_URL, + WIC_CLAIM_ARCHIVED_PARAMETERS_URL, + WIC_CLAIM_ARCHIVED_RANDOMNESS_URL, + WIC_CLAIM_FNS_SOURCE_URL, + clone_us_frame_for_puf_support, + derive_us_wic_claim_from_manifest, + load_release_input_coverage_manifest, + us_release_reform_coverage_probes, + us_wic_claim_signal_gate, + us_wic_claim_stage_spec, + us_wic_claim_summary, + with_us_wic_claim_input, +) +from populace.build.us_runtime.l0_refit_export import ( + US_RELEASE_REQUIRED_PERSON_SOURCE_COLUMNS, +) +from populace.build.us_runtime.release_input_coverage import ( + RESTORED_REFERENCE_ECPS_REQUIRED_INPUTS, +) +from populace.build.us_runtime.source_runtime import us_source_operation_handlers +from populace.frame import US_SCHEMA, EntitySchema, Frame, WeightKind, Weights + +policyengine_us_installed = importlib.util.find_spec("policyengine_us") is not None +requires_us = pytest.mark.skipif( + not policyengine_us_installed, + reason="requires the policyengine-us [us] extra (build environment)", +) + +_OUTPUT = "would_claim_wic" +_RATES = { + "pregnant": 0.456, + "postpartum": 0.689, + "breastfeeding": 0.663, + "infant": 0.784, + "child": 0.460, + "none": 0.0, +} + + +def _frame(rows: list[dict[str, object]]) -> Frame: + records: list[dict[str, object]] = [] + for index, overrides in enumerate(rows, start=1): + record: dict[str, object] = { + "person_id": index, + "person_household_id": index, + "person_tax_unit_id": index + 10_000, + "person_spm_unit_id": index + 20_000, + "person_family_id": index + 30_000, + "person_marital_unit_id": index + 40_000, + "age": 30.0, + "is_female": False, + "is_pregnant": False, + "own_children_in_household": 0, + "source_year": 2024, + "source_household_id": index, + "source_person_id": index, + } + record.update(overrides) + records.append(record) + person = pd.DataFrame.from_records(records) + tables = { + "person": person, + "household": pd.DataFrame( + {"household_id": np.sort(person["person_household_id"].unique())} + ), + "tax_unit": pd.DataFrame( + {"tax_unit_id": np.sort(person["person_tax_unit_id"].unique())} + ), + "spm_unit": pd.DataFrame( + {"spm_unit_id": np.sort(person["person_spm_unit_id"].unique())} + ), + "family": pd.DataFrame( + {"family_id": np.sort(person["person_family_id"].unique())} + ), + "marital_unit": pd.DataFrame( + {"marital_unit_id": np.sort(person["person_marital_unit_id"].unique())} + ), + } + return Frame( + tables, + US_SCHEMA, + { + "household": Weights( + np.ones(len(tables["household"]), dtype=np.float64), + WeightKind.DESIGN, + ) + }, + ) + + +def _replace_person(frame: Frame, person: pd.DataFrame) -> Frame: + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + tables["person"] = person + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + ) + + +def _context(seed: int = 0) -> SourceRuntimeContext: + return SourceRuntimeContext( + config=SourceRuntimeConfig(seed=seed, target_year=2024), + tables={}, + ) + + +def _operation() -> SourceOperationSpec: + return next( + operation + for operation in us_wic_claim_stage_spec().operations + if operation.kind == "derive_wic_claim" + ) + + +def _derive(person: pd.DataFrame, *, seed: int = 0) -> pd.DataFrame: + return derive_us_wic_claim_from_manifest( + person, + _operation(), + _context(seed), + ) + + +def _plausible_rows() -> list[dict[str, object]]: + rows: list[dict[str, object]] = [] + next_id = 1 + + for _ in range(120): + rows.append( + { + "person_id": next_id, + "is_female": True, + "is_pregnant": True, + "age": 27, + } + ) + next_id += 1 + + for family_number in range(120): + family_id = 100_000 + family_number + rows.append( + { + "person_id": next_id, + "person_family_id": family_id, + "is_female": True, + "own_children_in_household": 1, + "age": 27, + } + ) + next_id += 1 + rows.append( + { + "person_id": next_id, + "person_family_id": family_id, + "age": 0, + } + ) + next_id += 1 + + for _ in range(240): + rows.append({"person_id": next_id, "age": 3}) + next_id += 1 + + for _ in range(7_500): + rows.append({"person_id": next_id, "age": 35}) + next_id += 1 + return rows + + +class TestManifestAndProvenance: + def test_stage_locks_the_exact_official_category_rate_contract(self) -> None: + spec = us_wic_claim_stage_spec() + + assert spec.stage == US_WIC_CLAIM_STAGE_NAME + assert tuple(spec.outputs) == US_WIC_CLAIM_OUTPUT_COLUMNS == (_OUTPUT,) + assert US_WIC_CLAIM_NONCONSTANT_PERSON_COLUMNS == (_OUTPUT,) + assert US_WIC_CLAIM_REQUIRED_SOURCE_COLUMNS == ( + "age", + "is_female", + "is_pregnant", + "own_children_in_household", + "person_family_id", + ) + assert [operation.kind for operation in spec.operations] == [ + "read_table", + "derive_wic_claim", + ] + assert _operation().parameters == { + "seed_from_build_config": True, + "category_rates": { + "source": WIC_CLAIM_FNS_SOURCE_URL, + "vintage": "CY2022", + "values": _RATES, + }, + } + assert "after pregnancy" in spec.notes + assert "no breastfeeding assessment" in spec.notes + + def test_source_urls_lock_archived_files_lines_and_fns_pdf(self) -> None: + commit = "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe" + assert commit in WIC_CLAIM_ARCHIVED_DERIVATION_URL + assert WIC_CLAIM_ARCHIVED_DERIVATION_URL.endswith( + "/datasets/cps/cps.py#L684-L691" + ) + assert WIC_CLAIM_ARCHIVED_PARAMETERS_URL.endswith( + "/parameters/take_up/wic_takeup.yaml#L1-L33" + ) + assert WIC_CLAIM_ARCHIVED_RANDOMNESS_URL.endswith("/utils/randomness.py#L5-L28") + assert WIC_CLAIM_FNS_SOURCE_URL == ( + "https://fns-prod.azureedge.us/sites/default/files/resource-files/" + "wic-eer-2022-summary.pdf" + ) + + def test_handler_and_plan_run_wic_after_pregnancy_before_support(self) -> None: + assert ( + us_source_operation_handlers()["derive_wic_claim"] + is derive_us_wic_claim_from_manifest + ) + assert US_WIC_CLAIM_STAGE_NAME in US_DONORS + assert ( + US_STAGE_NAMES.index(US_PREGNANCY_STAGE_NAME) + < US_STAGE_NAMES.index(US_WIC_CLAIM_STAGE_NAME) + < US_STAGE_NAMES.index(US_PUF_SUPPORT_STAGE_NAME) + ) + + +class TestDerivation: + def test_categories_follow_pe_order_with_collapsed_postpartum(self) -> None: + rows = [ + { + "person_family_id": 1, + "age": 28, + "is_female": True, + "is_pregnant": True, + "own_children_in_household": 1, + }, + {"person_family_id": 1, "age": 0}, + { + "person_family_id": 2, + "age": 30, + "is_female": True, + "own_children_in_household": 1, + }, + {"person_family_id": 2, "age": 0}, + { + "person_family_id": 3, + "age": 30, + "is_female": False, + "own_children_in_household": 1, + }, + {"person_family_id": 3, "age": 0}, + {"age": 4}, + {"age": 5}, + ] + derived = with_us_wic_claim_input(_frame(rows), seed=11, time_period=2024) + summary = us_wic_claim_summary(derived) + + assert summary["category_assignment_order"] == [ + "pregnant", + "postpartum", + "infant", + "child", + "none", + ] + assert summary["category_counts"] == { + "pregnant": 1, + "postpartum": 1, + "infant": 3, + "child": 1, + "none": 2, + } + assert summary["breastfeeding_source_available"] is False + assert summary["breastfeeding_rate_validated_but_unassigned"] == 0.663 + + def test_draws_are_reproducible_and_keyed_by_source_identity(self) -> None: + rows = [ + { + "person_id": 1, + "age": 3, + "source_year": 2024, + "source_household_id": 50, + "source_person_id": 7, + }, + { + "person_id": 2, + "age": 3, + "source_year": 2024, + "source_household_id": 50, + "source_person_id": 7, + }, + ] + person = _frame(rows).table("person") + first = _derive(person, seed=9) + second = _derive(person, seed=9) + + assert first[_OUTPUT].tolist() == second[_OUTPUT].tolist() + assert first[_OUTPUT].iloc[0] == first[_OUTPUT].iloc[1] + + many = _frame([{"age": 3} for _ in range(500)]).table("person") + assert not np.array_equal( + _derive(many, seed=1)[_OUTPUT].to_numpy(), + _derive(many, seed=2)[_OUTPUT].to_numpy(), + ) + + @pytest.mark.parametrize("column", US_WIC_CLAIM_REQUIRED_SOURCE_COLUMNS) + def test_missing_source_columns_fail_closed(self, column: str) -> None: + person = _frame([{}]).table("person").drop(columns=[column]) + with pytest.raises(SourceRuntimeError, match=column): + _derive(person) + + @pytest.mark.parametrize( + ("updates", "message"), + [ + ({"age": np.nan}, "age"), + ({"age": -1}, "age"), + ({"is_female": 2}, "is_female"), + ({"is_pregnant": None}, "is_pregnant"), + ({"own_children_in_household": 1.5}, "own_children"), + ({"is_female": False, "is_pregnant": True}, "nonfemale"), + ], + ) + def test_invalid_category_sources_fail_closed( + self, + updates: dict[str, object], + message: str, + ) -> None: + with pytest.raises(SourceRuntimeError, match=message): + _derive(_frame([updates]).table("person")) + + def test_missing_family_membership_fails_closed_at_raw_stage_boundary(self) -> None: + person = _frame([{}]).table("person").copy() + person.loc[0, "person_family_id"] = np.nan + with pytest.raises(SourceRuntimeError, match="person_family_id"): + _derive(person) + + def test_partial_stable_identity_fails_instead_of_changing_clone_key(self) -> None: + person = _frame([{"age": 3}]).table("person").drop(columns=["source_person_id"]) + with pytest.raises(SourceRuntimeError, match="partial"): + _derive(person) + + def test_wrong_operation_missing_frame_and_parameter_drift_are_rejected( + self, + ) -> None: + person = _frame([{}]).table("person") + with pytest.raises(SourceRuntimeError, match="unexpected operation"): + derive_us_wic_claim_from_manifest( + person, + SourceOperationSpec(kind="wrong", parameters={}), + _context(), + ) + with pytest.raises(SourceRuntimeError, match="person table"): + derive_us_wic_claim_from_manifest(None, _operation(), _context()) + + mutations = [] + for mutate in ("source", "vintage", "child", "extra"): + parameters = deepcopy(dict(_operation().parameters)) + if mutate == "source": + parameters["category_rates"]["source"] = "https://example.com" + elif mutate == "vintage": + parameters["category_rates"]["vintage"] = "CY2021" + elif mutate == "child": + parameters["category_rates"]["values"]["child"] = 0.99 + else: + parameters["category_rates"]["values"]["extra"] = 0.1 + mutations.append(parameters) + for parameters in mutations: + with pytest.raises(SourceRuntimeError): + derive_us_wic_claim_from_manifest( + person, + SourceOperationSpec( + kind="derive_wic_claim", + parameters=parameters, + ), + _context(), + ) + + +class TestFrameAndGate: + def test_wrapper_recomputes_stale_surface_and_is_idempotent_by_equality( + self, + ) -> None: + frame = _frame(_plausible_rows()) + derived = with_us_wic_claim_input(frame, seed=0, time_period=2024) + assert us_wic_claim_signal_gate(derived).passed + assert with_us_wic_claim_input(derived, seed=0, time_period=2024) is derived + + stale_person = derived.table("person").copy() + stale_person[_OUTPUT] = ~stale_person[_OUTPUT] + stale = _replace_person(derived, stale_person) + healed = with_us_wic_claim_input(stale, seed=0, time_period=2024) + np.testing.assert_array_equal( + healed.table("person")[_OUTPUT], + derived.table("person")[_OUTPUT], + ) + + def test_constant_default_true_is_recomputed(self) -> None: + frame = _frame([{**row, _OUTPUT: True} for row in _plausible_rows()]) + derived = with_us_wic_claim_input(frame, seed=0, time_period=2024) + assert derived.table("person")[_OUTPUT].nunique() == 2 + assert us_wic_claim_signal_gate(derived).passed + + def test_support_clones_keep_identical_claims_and_channel_signal(self) -> None: + derived = with_us_wic_claim_input( + _frame(_plausible_rows()), seed=0, time_period=2024 + ) + cloned = clone_us_frame_for_puf_support(derived) + refreshed = with_us_wic_claim_input(cloned, seed=0, time_period=2024) + summary = us_wic_claim_summary(refreshed) + + assert refreshed is cloned + assert summary["clone_group_count"] == len(derived.table("person")) + assert summary["clone_claim_mismatch_count"] == 0 + assert summary["clone_category_mismatch_count"] == 0 + assert set(summary["channel_weighted_claim_shares"]) == { + "asec", + PUF_TAX_DETAIL_SUPPORT_CHANNEL, + } + assert us_wic_claim_signal_gate(refreshed).passed + + def test_gate_allows_unpaired_selected_channels_but_rejects_clone_mismatch( + self, + ) -> None: + derived = with_us_wic_claim_input( + _frame(_plausible_rows()), seed=0, time_period=2024 + ) + selected_person = derived.table("person").copy() + selected_person["person_support_channel"] = np.where( + np.arange(len(selected_person)) % 2, + "asec", + PUF_TAX_DETAIL_SUPPORT_CHANNEL, + ) + selected = _replace_person(derived, selected_person) + assert us_wic_claim_signal_gate(selected).passed + + cloned = clone_us_frame_for_puf_support(derived) + broken_person = cloned.table("person").copy() + duplicate = broken_person["source_person_id"].duplicated(keep="first") + row = int(np.flatnonzero(duplicate.to_numpy())[0]) + broken_person.loc[row, _OUTPUT] = not bool(broken_person.loc[row, _OUTPUT]) + broken = _replace_person(cloned, broken_person) + gate = us_wic_claim_signal_gate(broken) + assert not gate.passed + assert any("clone claim mismatch" in failure for failure in gate.failures) + + def test_gate_rejects_missing_constant_and_none_category_claim(self) -> None: + assert not us_wic_claim_signal_gate(_frame([{}])).passed + + constant = _frame([{_OUTPUT: True} for _ in range(10)]) + assert not us_wic_claim_signal_gate(constant).passed + + derived = with_us_wic_claim_input( + _frame(_plausible_rows()), seed=0, time_period=2024 + ) + person = derived.table("person").copy() + none_row = person.index[person["age"] == 35][0] + person.loc[none_row, _OUTPUT] = True + gate = us_wic_claim_signal_gate(_replace_person(derived, person)) + assert not gate.passed + assert any("none weighted claim share" in failure for failure in gate.failures) + + invalid = derived.table("person").copy() + invalid[_OUTPUT] = invalid[_OUTPUT].astype(object) + invalid.loc[invalid.index[0], _OUTPUT] = "not-a-boolean" + gate = us_wic_claim_signal_gate(_replace_person(derived, invalid)) + assert not gate.passed + assert any("nonnumeric" in failure for failure in gate.failures) + + def test_non_us_schema_is_rejected(self) -> None: + frame = Frame( + { + "person": pd.DataFrame({"person_id": [1], "person_household_id": [1]}), + "household": pd.DataFrame({"household_id": [1]}), + }, + EntitySchema(group_entities=("household",)), + {"household": Weights(np.ones(1), WeightKind.DESIGN)}, + ) + with pytest.raises(ValueError, match="US WIC"): + with_us_wic_claim_input(frame, seed=0, time_period=2024) + + +@requires_us +def test_policyengine_contract_and_live_wic_neutralization() -> None: + import inspect + + from policyengine_us import CountryTaxBenefitSystem, Simulation + + from populace.build.us_runtime.reform_coverage_smoke import _build_reform + + assert version("policyengine-us") == "1.764.6" + system = CountryTaxBenefitSystem() + variable = system.variables[_OUTPUT] + assert variable.is_input_variable() + assert variable.entity.key == "person" + assert variable.value_type is bool + assert bool(variable.default_value) is True + assert str(variable.definition_period).lower() == "month" + mother_formula = inspect.getsource(system.variables["is_mother"].formula) + assert 'person("is_parent", period)' in mother_formula + assert "female & has_children" in mother_formula + + probe = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "wic_claim_neutralization" + ) + situation = { + "people": { + "mother": { + "age": {"2024": 27}, + "is_pregnant": {"2024": True}, + "is_wic_at_nutritional_risk": {"2024": True}, + _OUTPUT: {"2024": True}, + } + }, + "tax_units": {"tax_unit": {"members": ["mother"]}}, + "families": {"family": {"members": ["mother"]}}, + "spm_units": {"spm_unit": {"members": ["mother"]}}, + "households": { + "household": { + "members": ["mother"], + "state_code": {"2024": "CA"}, + } + }, + "marital_units": {"marital_unit": {"members": ["mother"]}}, + } + baseline = Simulation(situation=situation) + neutralized = Simulation(situation=situation, reform=_build_reform(probe)) + assert baseline.calculate("wic", 2024)[0] > 0 + assert neutralized.calculate("wic", 2024)[0] == 0 + + +def test_release_promotion_probe_and_retired_gap_removal() -> None: + manifest = load_release_input_coverage_manifest() + assert _OUTPUT in RESTORED_REFERENCE_ECPS_REQUIRED_INPUTS + assert _OUTPUT in manifest.required_columns + assert _OUTPUT not in manifest.reviewed_exclusions + assert _OUTPUT in US_RELEASE_REQUIRED_PERSON_SOURCE_COLUMNS + + probe = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "wic_claim_neutralization" + ) + assert probe.neutralized_variable == _OUTPUT + assert probe.binding_inputs == (_OUTPUT,) + assert probe.budget_measure == "wic" + assert probe.effect_direction == "baseline_minus_reform" + assert probe.expected_sign == "positive" + assert probe.min_abs_effect == 25_000_000.0 + + known_gaps = __import__("json").loads( + files("populace.build.us").joinpath("ecps_parity_known_gaps.json").read_text() + )["known_gaps"] + assert _OUTPUT not in known_gaps diff --git a/packages/populace-build/tests/test_us_wic_nutritional_risk_exclusion.py b/packages/populace-build/tests/test_us_wic_nutritional_risk_exclusion.py new file mode 100644 index 00000000..50f416e3 --- /dev/null +++ b/packages/populace-build/tests/test_us_wic_nutritional_risk_exclusion.py @@ -0,0 +1,232 @@ +"""Evidence contract for the irreducible WIC nutritional-risk exclusion.""" + +from __future__ import annotations + +import importlib.util +import inspect +import json +from hashlib import sha256 +from importlib.metadata import version +from importlib.resources import files +from pathlib import Path + +import pytest + +from populace.build.us_runtime.asec_pool import load_asec_h5_tables +from populace.build.us_runtime.release_input_coverage import ( + load_release_input_coverage_manifest, +) + +ROOT = Path(__file__).resolve().parents[3] +policyengine_us_installed = importlib.util.find_spec("policyengine_us") is not None +requires_us = pytest.mark.skipif( + not policyengine_us_installed, + reason="requires the policyengine-us [us] extra (build environment)", +) + + +def _entry() -> dict[str, object]: + payload = json.loads( + files("populace.build.us").joinpath("ecps_parity_known_gaps.json").read_text() + ) + return payload["known_gaps"]["is_wic_at_nutritional_risk"] + + +def _sha256(path: Path) -> str: + digest = sha256() + with path.open("rb") as stream: + for chunk in iter(lambda: stream.read(1024 * 1024), b""): + digest.update(chunk) + return digest.hexdigest() + + +def _wic_like_columns(columns) -> list[str]: + return sorted( + column + for column in columns + if any(token in str(column).upper() for token in ("WIC", "NUTRI", "RISK")) + ) + + +def test_exclusion_pins_exact_archived_stochastic_derivation() -> None: + entry = _entry() + evidence = entry["evidence"] + + assert entry["reason"].startswith("SOURCE UNAVAILABILITY WITH EVIDENCE:") + assert evidence["classification"] == "source_unavailability" + assert evidence["retired_derivation"] == { + "repository_owner": "PolicyEngine", + "repository_name_parts": ["policyengine-", "us-data"], + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": [ + "policyengine_", + "us_data", + "datasets", + "cps", + "cps.py", + ], + "lines": "684-701", + "method": "receives_wic OR category-specific seeded Bernoulli draw", + } + assert evidence["retired_risk_rates"]["lines"] == "1-13" + assert evidence["retired_risk_rates"]["values"] == { + "PREGNANT": 0.913, + "POSTPARTUM": 0.933, + "BREASTFEEDING": 0.889, + "INFANT": 0.95, + "CHILD": 0.752, + "NONE": 0, + } + assert evidence["category_rate_source"]["lines"] == "38-50" + assert evidence["retired_receipt_mapping"] == { + "commit": "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe", + "path_parts": [ + "policyengine_", + "us_data", + "datasets", + "cps", + "cps.py", + ], + "lines": "1553-1559", + "mapping": "receives_wic = person.WICYN == 1", + } + assert evidence["retired_source_declaration"]["lines"] == "293-365" + assert evidence["retired_source_declaration"]["available_wic_fields"] == [ + "SPM_WICVAL", + "WICYN", + ] + + +def test_exclusion_pins_official_assessment_semantics_and_no_substitute() -> None: + evidence = _entry()["evidence"] + + census = evidence["official_census_dictionary"] + assert census["field"] == "WICYN" + assert census["definition"] == "Who received WIC?" + assert census["universe"] == "Adult female" + assert "on behalf of a child" in census["questionnaire_context"] + assert census["codes"] == { + "0": "Not in universe", + "1": "Received WIC", + "2": "Did not receive WIC", + } + assert census["url"].endswith("/cpsmar24.pdf") + + certification = evidence["official_certification_source"] + assert certification["url"].startswith("https://www.fns.usda.gov/") + assert "competent professional authority" in certification["finding"] + assert "state administrative records" in certification["finding"] + assert "inference" not in certification["finding"] + assert "Therefore" in certification["record_linkage_inference"] + assert ( + "neither an assessment field nor a link key" + in certification["record_linkage_inference"] + ) + + current_method = evidence["current_fns_estimation_method"] + assert current_method["lines"] == "PDF page 84 (printed report page 70)" + assert current_method["adopted_when"] == "while producing the CY2020 estimates" + assert current_method["applied_period"] == ( + "revised CY2016-CY2021 estimates in the cited report" + ) + assert current_method["risk_adjustment"] == 1.0 + assert current_method["scope"] == "all participant categories" + + substitutes = evidence["semantic_non_substitutes"] + assert "does not identify the assessed person" in substitutes["WICYN"] + assert "nonreceipt also does not prove" in substitutes["WICYN"] + assert "neither an assessed person nor a negative" in substitutes["SPM_WICVAL"] + assert "synthesize" in substitutes["rejection"] + + +def test_exclusion_pins_all_sha_locked_hermetic_inputs() -> None: + evidence = _entry()["evidence"] + evidence_hashes = { + item["filename"]: item["sha256"] for item in evidence["hermetic_inputs"] + } + build_summary = json.loads( + (ROOT / "experiments/build_j_recert/base_j.summary.json").read_text() + ) + recorded_hashes = { + Path(item["path"]).name: item["sha256"] + for item in build_summary["base_source"]["sources"] + } + + assert evidence_hashes == recorded_hashes + assert set(evidence_hashes) == { + "census_cps_2022.h5", + "census_cps_2023.h5", + "census_cps_2024.h5", + } + for item in evidence["hermetic_inputs"]: + assert item["person_wic_columns"] == ["SPM_WICVAL", "WICYN"] + assert item["household_wic_columns"] == ["HRNUMWIC", "HRWICYN"] + assert sum(item["wicyn_counts"].values()) > 100_000 + assert item["wicyn_counts"]["1"] > 1_000 + + build_script = (ROOT / "experiments/build_j_recert/buildj_base.sh").read_text() + for year in (2022, 2023, 2024): + assert f'--asec-h5 {year}="$USD/census_cps_{year}.h5"' in build_script + assert "buildj_base.sh lines 65-69" in evidence["hermetic_build_contract"] + assert "base_j.summary.json lines 55-75" in evidence["hermetic_build_contract"] + asec_pool = ( + ROOT / "packages/populace-build/src/populace/build/us_runtime/asec_pool.py" + ).read_text() + assert 'pd.HDFStore(path, mode="r")' in asec_pool + + +def test_mounted_artifacts_contain_receipt_but_no_risk_assessment() -> None: + evidence = _entry()["evidence"] + build_summary = json.loads( + (ROOT / "experiments/build_j_recert/base_j.summary.json").read_text() + ) + paths = { + Path(item["path"]).name: Path(item["path"]) + for item in build_summary["base_source"]["sources"] + } + if not all(path.is_file() for path in paths.values()): + pytest.skip("SHA-locked ASEC artifacts are not mounted in this environment") + + for item in evidence["hermetic_inputs"]: + path = paths[item["filename"]] + assert _sha256(path) == item["sha256"] + tables = load_asec_h5_tables(path) + assert _wic_like_columns(tables["person"].columns) == sorted( + item["person_wic_columns"] + ) + assert _wic_like_columns(tables["household"].columns) == sorted( + item["household_wic_columns"] + ) + observed_counts = { + str(int(code)): int(count) + for code, count in tables["person"]["WICYN"].value_counts().items() + } + assert observed_counts == item["wicyn_counts"] + + +@requires_us +def test_policyengine_1_764_6_defaults_risk_true_and_uses_it_for_eligibility() -> None: + from policyengine_us import CountryTaxBenefitSystem + + assert version("policyengine-us") == "1.764.6" + system = CountryTaxBenefitSystem() + variable = system.variables["is_wic_at_nutritional_risk"] + assert variable.is_input_variable() + assert variable.entity.key == "person" + assert str(variable.definition_period).lower() == "month" + assert variable.value_type is bool + assert variable.default_value is True + + eligibility_formula = system.variables["is_wic_eligible"].get_formula("2024-01") + assert eligibility_formula is not None + assert 'person("is_wic_at_nutritional_risk", period)' in inspect.getsource( + eligibility_formula + ) + + +def test_generated_release_manifest_preserves_the_evidenced_exclusion() -> None: + reason = load_release_input_coverage_manifest().reviewed_exclusions[ + "is_wic_at_nutritional_risk" + ] + assert reason == _entry()["reason"] + assert reason.startswith("SOURCE UNAVAILABILITY WITH EVIDENCE:") diff --git a/packages/populace-build/tests/test_us_workers_compensation.py b/packages/populace-build/tests/test_us_workers_compensation.py new file mode 100644 index 00000000..78bf0b2b --- /dev/null +++ b/packages/populace-build/tests/test_us_workers_compensation.py @@ -0,0 +1,652 @@ +"""ASEC workers' compensation restoration and PUF-half QRF treatment.""" + +from __future__ import annotations + +import importlib.util +import json +from hashlib import sha256 +from importlib.metadata import version +from pathlib import Path + +import numpy as np +import pandas as pd +import pytest + +import populace.build.us_runtime.workers_compensation as module +from populace.build.source_runtime import ( + SourceRuntimeConfig, + SourceRuntimeContext, + SourceRuntimeError, +) +from populace.build.us_runtime.l0_refit_export import ( + US_RELEASE_REQUIRED_PERSON_SOURCE_COLUMNS, +) +from populace.build.us_runtime.puf_support import clone_us_frame_for_puf_support +from populace.build.us_runtime.release_input_coverage import ( + RESTORED_REFERENCE_ECPS_REQUIRED_INPUTS, + load_release_input_coverage_manifest, + us_release_reform_coverage_probes, +) +from populace.build.us_runtime.source_runtime import us_source_operation_handlers +from populace.build.us_runtime.workers_compensation import ( + US_WORKERS_COMPENSATION_OUTPUT_COLUMNS, + US_WORKERS_COMPENSATION_REQUIRED_SOURCE_COLUMNS, + WORKERS_COMPENSATION_ARCHIVED_DERIVATION_URL, + WORKERS_COMPENSATION_ARCHIVED_PUF_IMPUTATION_URL, + WORKERS_COMPENSATION_ARCHIVED_PUF_OUTPUTS_URL, + WORKERS_COMPENSATION_ARCHIVED_SOURCE_COLUMNS_URL, + derive_us_workers_compensation_from_manifest, + impute_us_workers_compensation_to_puf_support_from_manifest, + us_workers_compensation_signal_gate, + us_workers_compensation_stage_spec, + with_us_workers_compensation, +) +from populace.frame import US_SCHEMA, Frame, WeightKind, Weights + +policyengine_us_installed = importlib.util.find_spec("policyengine_us") is not None +requires_us = pytest.mark.skipif( + not policyengine_us_installed, + reason="requires the policyengine-us [us] extra (build environment)", +) + +_OUTPUT = US_WORKERS_COMPENSATION_OUTPUT_COLUMNS[0] +_PREDICTORS = ( + "age", + "is_male", + "has_esi", + "tax_unit_is_joint", + "tax_unit_count_dependents", + "employment_income", + "self_employment_income", + "social_security", +) +ROOT = Path(__file__).resolve().parents[3] + + +def _sha256(path: Path) -> str: + digest = sha256() + with path.open("rb") as stream: + for chunk in iter(lambda: stream.read(1024 * 1024), b""): + digest.update(chunk) + return digest.hexdigest() + + +def _person_source() -> pd.DataFrame: + count = 100 + workers_compensation = np.zeros(count) + workers_compensation[0] = 3_600.0 + return pd.DataFrame( + { + "person_id": np.arange(1, count + 1, dtype="int64"), + "person_household_id": np.arange(1, count + 1, dtype="int64") * 10, + "person_tax_unit_id": np.arange(1, count + 1, dtype="int64") * 100, + "person_spm_unit_id": np.arange(1, count + 1, dtype="int64") * 1_000, + "person_family_id": np.arange(1, count + 1, dtype="int64") * 10_000, + "person_marital_unit_id": ( + np.arange(1, count + 1, dtype="int64") * 100_000 + ), + "WC_VAL": workers_compensation, + "WSAL_VAL": np.linspace(0.0, 99_000.0, count), + "SEMP_VAL": np.zeros(count), + "employment_income_before_lsr": np.linspace(0.0, 99_000.0, count), + "self_employment_income_before_lsr": np.zeros(count), + "age": np.arange(20, 20 + count), + "is_female": np.tile([False, True], count // 2), + "has_esi": np.tile([True, False], count // 2), + "tax_unit_role_input": ["HEAD"] * count, + "social_security_retirement": np.zeros(count), + "social_security_disability": np.zeros(count), + "social_security_dependents": np.zeros(count), + "social_security_survivors": np.zeros(count), + } + ) + + +def _frame() -> Frame: + person = _person_source() + count = len(person) + ids = { + "household": person["person_household_id"].to_numpy(), + "tax_unit": person["person_tax_unit_id"].to_numpy(), + "spm_unit": person["person_spm_unit_id"].to_numpy(), + "family": person["person_family_id"].to_numpy(), + "marital_unit": person["person_marital_unit_id"].to_numpy(), + } + tables = { + entity: pd.DataFrame({f"{entity}_id": values}) for entity, values in ids.items() + } + tables["person"] = person + tables["tax_unit"]["filing_status_input"] = ["SINGLE"] * count + return Frame( + tables, + US_SCHEMA, + { + "household": Weights( + np.ones(count, dtype=np.float64), + WeightKind.DESIGN, + ) + }, + ) + + +def _derive(person: pd.DataFrame) -> pd.DataFrame: + operation = next( + operation + for operation in us_workers_compensation_stage_spec().operations + if operation.kind == "derive_workers_compensation" + ) + return derive_us_workers_compensation_from_manifest(person, operation, None) + + +def test_archived_sources_are_sha_and_line_pinned_and_sources_available() -> None: + commit = "42ed5d45c56df80d754fbe24cce21cfeb8d05cbe" + urls = ( + WORKERS_COMPENSATION_ARCHIVED_DERIVATION_URL, + WORKERS_COMPENSATION_ARCHIVED_SOURCE_COLUMNS_URL, + WORKERS_COMPENSATION_ARCHIVED_PUF_OUTPUTS_URL, + WORKERS_COMPENSATION_ARCHIVED_PUF_IMPUTATION_URL, + ) + assert all(commit in url for url in urls) + assert WORKERS_COMPENSATION_ARCHIVED_DERIVATION_URL.endswith( + "datasets/cps/cps.py#L1559-L1571" + ) + assert WORKERS_COMPENSATION_ARCHIVED_SOURCE_COLUMNS_URL.endswith( + "datasets/cps/census_cps.py#L306-L381" + ) + assert WORKERS_COMPENSATION_ARCHIVED_PUF_OUTPUTS_URL.endswith( + "datasets/cps/extended_cps.py#L135-L194" + ) + assert WORKERS_COMPENSATION_ARCHIVED_PUF_IMPUTATION_URL.endswith( + "datasets/cps/extended_cps.py#L639-L745" + ) + assert US_WORKERS_COMPENSATION_REQUIRED_SOURCE_COLUMNS == ("WC_VAL",) + + +@requires_us +def test_all_sha_locked_asec_sources_have_exact_wc_val_signal() -> None: + from policyengine_us.data import USSingleYearDataset + + expected = { + "census_cps_2022.h5": { + "sha256": "7ccca976284bb47815d84460cc4f75a0a65d26d7754ab0a0f417de351b3d474e", + "positive": 328, + "raw_sum": 3_821_784.0, + "maximum": 99_999.0, + "weighted_share": 0.0021741514, + "weighted_total": 7_785_631_296.0, + }, + "census_cps_2023.h5": { + "sha256": "cb57817327799f42b741caed5f9be94d04021c2e6809c1ad7bd0686da5428d88", + "positive": 379, + "raw_sum": 4_710_504.0, + "maximum": 99_999.0, + "weighted_share": 0.0026142319, + "weighted_total": 10_850_596_150.0, + }, + "census_cps_2024.h5": { + "sha256": "ec36604cb735a660b51b0b2f90be27d803b5878f3464fb30d0eacead59c1260d", + "positive": 391, + "raw_sum": 4_160_758.0, + "maximum": 72_000.0, + "weighted_share": 0.0028173010, + "weighted_total": 9_146_153_896.0, + }, + } + summary = json.loads( + (ROOT / "experiments/build_j_recert/base_j.summary.json").read_text() + ) + paths = { + Path(item["path"]).name: Path(item["path"]) + for item in summary["base_source"]["sources"] + } + if not all(path.is_file() for path in paths.values()): + pytest.skip("SHA-locked ASEC artifacts are not mounted") + + assert set(paths) == set(expected) + for filename, facts in expected.items(): + path = paths[filename] + assert _sha256(path) == facts["sha256"] + person = USSingleYearDataset(file_path=str(path)).person + values = pd.to_numeric(person["WC_VAL"], errors="coerce").to_numpy( + dtype=np.float64 + ) + weights = ( + pd.to_numeric(person["A_FNLWGT"], errors="coerce").to_numpy( + dtype=np.float64 + ) + / 100.0 + ) + positive = values > 0.0 + assert np.isfinite(values).all() + assert (values >= 0.0).all() + assert np.array_equal(values, np.floor(values)) + assert int(np.count_nonzero(positive)) == facts["positive"] + assert float(values.sum()) == facts["raw_sum"] + assert float(values.max()) == facts["maximum"] + assert float(weights[positive].sum() / weights.sum()) == pytest.approx( + facts["weighted_share"], abs=1e-9 + ) + assert float((values * weights).sum()) == pytest.approx( + facts["weighted_total"], rel=1e-8 + ) + + +def test_stage_manifest_pins_direct_wc_val_formula_and_one_output_qrf() -> None: + spec = us_workers_compensation_stage_spec() + + assert spec.stage == "workers_compensation_input" + assert spec.survey == "Census CPS ASEC" + assert spec.grain == "person" + assert tuple(spec.outputs) == (_OUTPUT,) + assert tuple(spec.nonnegative_outputs) == (_OUTPUT,) + assert [operation.kind for operation in spec.operations] == [ + "read_table", + "derive_workers_compensation", + "impute_workers_compensation_to_puf_support", + ] + assert spec.operations[0].parameters == { + "table": "person", + "weight": "person_weight", + } + assert spec.operations[1].parameters == { + "source": "WC_VAL", + "output": _OUTPUT, + } + assert spec.operations[2].parameters == { + "predictors": list(_PREDICTORS), + "max_train_samples": 5_000, + "n_estimators": 100, + "seed_from_build_config": True, + "weight": "person_weight", + } + + handlers = us_source_operation_handlers() + assert ( + handlers["derive_workers_compensation"] + is derive_us_workers_compensation_from_manifest + ) + assert ( + handlers["impute_workers_compensation_to_puf_support"] + is impute_us_workers_compensation_to_puf_support_from_manifest + ) + + +def test_direct_formula_carries_wc_val_and_preserves_topcodes() -> None: + source = pd.DataFrame({"WC_VAL": [0.0, 99_999.0, 500.0, 800.0, 300.0]}) + original = source.copy(deep=True) + + result = _derive(source) + + assert result[_OUTPUT].tolist() == [0.0, 99_999.0, 500.0, 800.0, 300.0] + pd.testing.assert_frame_equal(source, original) + + +def test_direct_formula_does_not_substitute_disability_code_one_slots() -> None: + source = pd.DataFrame( + { + "WC_VAL": [1_000.0, 0.0, 2_000.0], + "DIS_VAL1": [1_000.0, 5_000.0, 0.0], + "DIS_SC1": [1, 1, 0], + "DIS_VAL2": [0.0, 4_000.0, 8_000.0], + "DIS_SC2": [0, 1, 1], + } + ) + + result = _derive(source) + + assert result[_OUTPUT].tolist() == [1_000.0, 0.0, 2_000.0] + + +@pytest.mark.parametrize("missing", US_WORKERS_COMPENSATION_REQUIRED_SOURCE_COLUMNS) +def test_direct_formula_fails_closed_when_source_is_missing(missing: str) -> None: + with pytest.raises(SourceRuntimeError, match=missing): + _derive(_person_source().drop(columns=[missing])) + + +@pytest.mark.parametrize( + ("column", "bad_value", "message"), + [ + ("WC_VAL", np.nan, "nonfinite"), + ("WC_VAL", np.inf, "nonfinite"), + ("WC_VAL", -1.0, "negative"), + ], +) +def test_direct_formula_rejects_invalid_sources( + column: str, + bad_value: float, + message: str, +) -> None: + person = _person_source() + person.loc[0, column] = bad_value + + with pytest.raises(SourceRuntimeError, match=message): + _derive(person) + + +def test_with_inputs_materializes_exact_asec_values_without_mutation() -> None: + frame = _frame() + original = frame.table("person").copy(deep=True) + + result = with_us_workers_compensation(frame, seed=0, time_period=2024) + + pd.testing.assert_frame_equal(frame.table("person"), original) + expected = original["WC_VAL"] + np.testing.assert_allclose(result.table("person")[_OUTPUT], expected) + gate = us_workers_compensation_signal_gate(result) + assert gate.passed, gate.failures + + +def test_release_requires_opt_in_to_preserve_valid_existing_surface() -> None: + materialized = with_us_workers_compensation(_frame(), seed=0, time_period=2024) + tables = { + entity: materialized.table(entity).copy() for entity in materialized.entities + } + tables["person"] = tables["person"].drop( + columns=list(US_WORKERS_COMPENSATION_REQUIRED_SOURCE_COLUMNS) + ) + release = Frame( + tables, + materialized.schema, + { + entity: materialized.weights_for(entity) + for entity in materialized.weighted_entities + }, + materialized.strata, + mass_log=materialized.mass_log, + ) + + with pytest.raises(ValueError, match="cannot heal.*without measured"): + with_us_workers_compensation(release, seed=0, time_period=2024) + + assert ( + with_us_workers_compensation( + release, + seed=0, + time_period=2024, + allow_existing_without_source=True, + ) + is release + ) + + +def test_puf_half_uses_one_output_qrf_and_preserves_asec( + monkeypatch: pytest.MonkeyPatch, +) -> None: + direct = with_us_workers_compensation(_frame(), seed=0, time_period=2024) + expanded = clone_us_frame_for_puf_support(direct) + original = expanded.table("person").copy(deep=True) + calls: dict[str, object] = {} + + class FakeFitted: + def predict(self, test: pd.DataFrame) -> pd.DataFrame: + calls["test"] = test.copy() + predicted = np.zeros(len(test), dtype=np.float64) + predicted[0] = 7_200.0 + return pd.DataFrame({_OUTPUT: predicted}, index=test.index) + + class FakeQRF: + def __init__(self, **kwargs: object) -> None: + calls["init"] = kwargs + + def fit( + self, + training: pd.DataFrame, + predictors: list[str], + targets: list[str], + *, + weights: np.ndarray, + ) -> FakeFitted: + calls["fit_count"] = int(calls.get("fit_count", 0)) + 1 + calls["training"] = training.copy() + calls["predictors"] = predictors + calls["targets"] = targets + calls["weights"] = weights.copy() + return FakeFitted() + + monkeypatch.setattr(module, "QRF", FakeQRF) + + result = with_us_workers_compensation(expanded, seed=7, time_period=2024) + + assert calls["fit_count"] == 1 + assert calls["init"] == {"n_estimators": 100, "seed": 7} + assert calls["predictors"] == list(_PREDICTORS) + assert calls["targets"] == [_OUTPUT] + training = calls["training"] + assert isinstance(training, pd.DataFrame) + assert list(training.columns) == [*_PREDICTORS, _OUTPUT] + asec_mask = original["person_support_channel"] == "asec" + expected_weights = expanded.resolve_weights("person").values[asec_mask] + np.testing.assert_allclose(calls["weights"], expected_weights) + + person = result.table("person") + asec = person[person["person_support_channel"] == "asec"] + puf = person[person["person_support_channel"] == "puf_tax_detail"] + assert asec[_OUTPUT].tolist() == [3_600.0, *([0.0] * 99)] + assert puf[_OUTPUT].tolist() == [7_200.0, *([0.0] * 99)] + pd.testing.assert_frame_equal(expanded.table("person"), original) + gate = us_workers_compensation_signal_gate(result) + assert gate.passed, gate.failures + + +def test_puf_qrf_caps_training_at_5000_and_keeps_weights_aligned( + monkeypatch: pytest.MonkeyPatch, +) -> None: + operation = us_workers_compensation_stage_spec().operations[2] + asec_rows = 5_010 + puf_rows = 2 + rows = asec_rows + puf_rows + frame = pd.DataFrame( + { + "person_support_channel": ["asec"] * asec_rows + + ["puf_tax_detail"] * puf_rows, + "person_weight": np.arange(1.0, rows + 1.0), + _OUTPUT: np.tile([0.0, 500.0], (rows + 1) // 2)[:rows], + **{ + f"workers_compensation_predictor_{predictor}": np.arange( + rows, dtype=np.float64 + ) + for predictor in _PREDICTORS + }, + } + ) + calls: dict[str, object] = {} + + class FakeFitted: + def predict(self, test: pd.DataFrame) -> pd.DataFrame: + calls["test_rows"] = len(test) + return pd.DataFrame({_OUTPUT: np.zeros(len(test))}, index=test.index) + + class FakeQRF: + def __init__(self, **kwargs: object) -> None: + calls["init"] = kwargs + + def fit( + self, + training: pd.DataFrame, + predictors: list[str], + targets: list[str], + *, + weights: np.ndarray, + ) -> FakeFitted: + calls["training_rows"] = len(training) + calls["training_index"] = training.index.to_numpy() + calls["weights"] = weights.copy() + return FakeFitted() + + monkeypatch.setattr(module, "QRF", FakeQRF) + context = SourceRuntimeContext( + config=SourceRuntimeConfig(seed=19, target_year=2024), + tables={}, + ) + + impute_us_workers_compensation_to_puf_support_from_manifest( + frame, + operation, + context, + ) + + assert calls["init"] == {"n_estimators": 100, "seed": 19} + assert calls["training_rows"] == 5_000 + assert calls["test_rows"] == 2 + np.testing.assert_allclose( + calls["weights"], + frame.loc[calls["training_index"], "person_weight"].to_numpy(), + ) + + +def test_signal_gate_rejects_missing_default_and_invalid_surfaces() -> None: + valid = with_us_workers_compensation(_frame(), seed=0, time_period=2024) + assert us_workers_compensation_signal_gate(valid).passed + + candidates: list[Frame] = [] + for replacement in ( + None, + np.zeros(100), + np.asarray([-1.0, *([0.0] * 99)]), + np.asarray([np.nan, *([0.0] * 99)]), + ): + candidate = with_us_workers_compensation(_frame(), seed=0, time_period=2024) + if replacement is None: + candidate.table("person").drop(columns=[_OUTPUT], inplace=True) + else: + candidate.table("person")[_OUTPUT] = replacement + candidates.append(candidate) + + assert all( + not us_workers_compensation_signal_gate(candidate).passed + for candidate in candidates + ) + + +@pytest.mark.parametrize("dead_channel", ["asec", "puf_tax_detail"]) +def test_signal_gate_rejects_either_dead_support_channel(dead_channel: str) -> None: + direct = with_us_workers_compensation(_frame(), seed=0, time_period=2024) + expanded = clone_us_frame_for_puf_support(direct) + assert us_workers_compensation_signal_gate(expanded).passed + + channel = expanded.table("person")["person_support_channel"] + expanded.table("person").loc[channel == dead_channel, _OUTPUT] = 0.0 + gate = us_workers_compensation_signal_gate(expanded) + + assert not gate.passed + assert any(dead_channel in failure for failure in gate.failures) + + +@requires_us +def test_policyengine_us_1_764_6_contract_and_positive_annual_behavior() -> None: + from policyengine_us import CountryTaxBenefitSystem, Simulation + + assert version("policyengine-us") == "1.764.6" + variable = CountryTaxBenefitSystem().variables[_OUTPUT] + assert variable.is_input_variable() + assert variable.entity.key == "person" + assert str(variable.definition_period).lower() == "year" + assert variable.default_value == 0 + + situation = { + "people": { + "adult": { + "age": {"2024": 40}, + _OUTPUT: {"2024": 6_000.0}, + } + }, + "tax_units": { + "tax_unit": { + "members": ["adult"], + "filing_status": {"2024": "SINGLE"}, + } + }, + "spm_units": {"spm_unit": {"members": ["adult"]}}, + "households": { + "household": { + "members": ["adult"], + "state_code": {"2024": "CA"}, + } + }, + } + simulation = Simulation(situation=situation) + + assert simulation.calculate(_OUTPUT, 2024)[0] == pytest.approx(6_000.0) + assert simulation.calculate(_OUTPUT, "2024-01")[0] == pytest.approx(500.0) + assert simulation.calculate("snap_unearned_income", "2024-01")[0] == pytest.approx( + 500.0 + ) + + +@requires_us +def test_shipped_snap_exclusion_probe_binds_with_positive_sign() -> None: + from policyengine_core.reforms import Reform + from policyengine_us import CountryTaxBenefitSystem, Simulation + + probe = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "workers_compensation_snap_exclusion" + ) + reform = Reform.from_dict(dict(probe.parameter_changes), country_id="us") + situation = { + "people": { + "adult": { + "age": {"2024": 40}, + "employment_income": {"2024": 12_000.0}, + _OUTPUT: {"2024": 6_000.0}, + } + }, + "tax_units": { + "tax_unit": { + "members": ["adult"], + "filing_status": {"2024": "SINGLE"}, + } + }, + "spm_units": {"spm_unit": {"members": ["adult"]}}, + "households": { + "household": { + "members": ["adult"], + "state_code": {"2024": "CA"}, + } + }, + } + baseline = Simulation(situation=situation) + reformed = Simulation( + tax_benefit_system=CountryTaxBenefitSystem(reform=(reform,)), + situation=situation, + ) + + effect = reformed.calculate("snap", 2024)[0] - baseline.calculate("snap", 2024)[0] + assert effect > 1_000.0 + + +def test_release_promotion_plan_builders_and_probe_are_wired() -> None: + from populace.build.us_runtime import US_DONORS, US_STAGE_NAMES + + manifest = load_release_input_coverage_manifest() + assert _OUTPUT in RESTORED_REFERENCE_ECPS_REQUIRED_INPUTS + assert _OUTPUT in manifest.required_columns + assert _OUTPUT not in manifest.reviewed_exclusions + assert _OUTPUT in US_RELEASE_REQUIRED_PERSON_SOURCE_COLUMNS + assert module.US_WORKERS_COMPENSATION_STAGE_NAME in US_DONORS + assert module.US_WORKERS_COMPENSATION_STAGE_NAME in US_STAGE_NAMES + + probe = next( + probe + for probe in us_release_reform_coverage_probes() + if probe.id == "workers_compensation_snap_exclusion" + ) + assert probe.binding_inputs == (_OUTPUT,) + assert probe.budget_measure == "snap" + assert probe.effect_direction == "reform_minus_baseline" + assert probe.expected_sign == "positive" + assert probe.min_abs_effect == 10_000_000.0 + sources = probe.parameter_changes["gov.usda.snap.income.sources.unearned"][ + "2024-01-01.2024-12-31" + ] + assert _OUTPUT not in sources + assert "disability_benefits" in sources + + support_builder = (ROOT / "tools/build_us_puf_support_base.py").read_text() + fiscal_builder = (ROOT / "tools/build_us_fiscal_refresh_release.py").read_text() + cache_driver = (ROOT / "experiments/build_j_recert/buildj_base.sh").read_text() + assert support_builder.count("with_us_workers_compensation(") == 2 + assert "us_workers_compensation_signal_gate(" in support_builder + assert "us_workers_compensation_signal_gate(" in fiscal_builder + assert f'"{_OUTPUT}"' in cache_driver diff --git a/packages/populace-build/tests/test_validation_input_coverage.py b/packages/populace-build/tests/test_validation_input_coverage.py index a60b793a..f551e8dd 100644 --- a/packages/populace-build/tests/test_validation_input_coverage.py +++ b/packages/populace-build/tests/test_validation_input_coverage.py @@ -15,6 +15,7 @@ import pytest from populace.build.us_runtime import ( + US_QBI_OUTPUT_COLUMNS, US_VALIDATION_PROVISION_INPUT_LEAVES, ValidationInputLeaf, assert_validation_leaf_registry_current, @@ -39,23 +40,68 @@ def test_reads_shipped_manifest_outputs(self) -> None: # A representative PUF income leaf declared by the tax-detail stage. assert "employment_income_before_lsr" in outputs assert "student_loan_interest" in outputs - # The two invisible-gap leaves are NOT declared by any stage — the - # whole reason the gate needs their reviewed exclusions. - assert "qualified_tuition_expenses" not in outputs - assert "qualified_passenger_vehicle_loan_interest" not in outputs + assert "fsla_overtime_premium" in outputs + # Tuition and qualifying auto-loan interest are declared source-stage + # outputs and may no longer hide behind exclusions. + assert "qualified_tuition_expenses" in outputs + assert "qualified_passenger_vehicle_loan_interest" in outputs + assert "household_vehicles_owned" in outputs + assert "household_vehicles_value" in outputs + assert "traditional_401k_contributions_desired" in outputs + assert "roth_401k_contributions_desired" in outputs + assert "traditional_ira_contributions_desired" in outputs + assert "roth_ira_contributions_desired" in outputs + assert "self_employed_pension_contributions_desired" in outputs + assert "takes_up_ssi_if_eligible" in outputs + assert "takes_up_head_start_if_eligible" in outputs + assert "weeks_unemployed" in outputs + assert "casualty_loss" in outputs + assert "domestic_production_ald" in outputs + assert "unreimbursed_business_employee_expenses" in outputs + assert "spm_unit_pre_subsidy_childcare_expenses" in outputs + assert "spm_unit_energy_subsidy" in outputs + assert "child_support_received" in outputs + assert "child_support_expense" in outputs + assert "disability_benefits" in outputs + assert "educator_expense" in outputs + assert "other_health_insurance_premiums" in outputs + assert "farm_operations_income" in outputs + assert "farm_rent_income" in outputs + assert set(US_QBI_OUTPUT_COLUMNS) <= outputs class TestUsValidationInputCoverageGate: - def test_shipped_config_passes_with_known_gaps_allowlisted(self) -> None: - # The shipped state: both critical leaves are missing but each carries a - # reviewed exclusion with its tracking issue, so the gate passes. + def test_shipped_config_passes_without_reviewed_exclusions(self) -> None: result = us_validation_input_coverage_gate() assert result.passed, result.failures - assert set(result.details["reviewed_exclusions"]) == { - "qualified_tuition_expenses", - "qualified_passenger_vehicle_loan_interest", - } + assert result.details["reviewed_exclusions"] == {} assert result.details["missing"] == [] + requirements = us_validation_input_leaf_requirements() + assert requirements["tip_income"] == ["obbba_no_tax_on_tips"] + assert requirements["treasury_tipped_occupation_code"] == [ + "obbba_no_tax_on_tips" + ] + assert requirements["fsla_overtime_premium"] == ["obbba_no_tax_on_overtime"] + assert requirements["qualified_passenger_vehicle_loan_interest"] == [ + "obbba_auto_loan_interest" + ] + for leaf in ( + "traditional_401k_contributions_desired", + "roth_401k_contributions_desired", + "traditional_ira_contributions_desired", + "roth_ira_contributions_desired", + "self_employed_pension_contributions_desired", + ): + assert requirements[leaf] == ["soi_savers_credit"] + assert requirements["casualty_loss"] == ["obbba_casualty_loss_limit"] + assert requirements["unreimbursed_business_employee_expenses"] == [ + "obbba_misc_itemized_deductions" + ] + assert requirements["spm_unit_pre_subsidy_childcare_expenses"] == [ + "obbba_cdcc", + "soi_cdcc", + "te_cdcc", + ] def test_planted_missing_leaf_fails_loudly(self) -> None: # Plant a NEW validation row whose provision keys on an un-imputed, @@ -96,49 +142,70 @@ def test_shipped_gate_fails_when_registry_gains_an_unimputed_leaf( # No reason: this leaf is expected to be present, not a tracked gap. ), ) - monkeypatch.setattr( - module, "US_VALIDATION_PROVISION_INPUT_LEAVES", planted - ) + monkeypatch.setattr(module, "US_VALIDATION_PROVISION_INPUT_LEAVES", planted) result = us_validation_input_coverage_gate() assert not result.passed assert result.details["missing"] == ["some_new_unimputed_input"] - assert any( - "obbba_new_untested_provision" in line for line in result.failures + assert any("obbba_new_untested_provision" in line for line in result.failures) + + def test_removing_tuition_from_outputs_makes_the_row_fail(self) -> None: + requirements = us_validation_input_leaf_requirements() + + from populace.build.gates import source_stage_input_coverage_gate + + result = source_stage_input_coverage_gate( + requirements, + declared_outputs=us_source_stage_outputs() - {"qualified_tuition_expenses"}, + reviewed_exclusions={}, + name="us_validation_input_coverage", ) + assert not result.passed + assert set(result.details["missing"]) == {"qualified_tuition_expenses"} + + def test_removing_casualty_loss_makes_the_obbba_row_fail(self) -> None: + requirements = us_validation_input_leaf_requirements() - def test_leaf_becoming_a_declared_output_flags_stale_exclusion(self) -> None: - # If a known-gap leaf is later produced by a stage, its reviewed - # exclusion is stale and the gate flags it so the register cannot rot. - result = us_validation_input_coverage_gate( - source_stage_outputs=[ - *us_source_stage_outputs(), - "qualified_tuition_expenses", - ], + from populace.build.gates import source_stage_input_coverage_gate + + result = source_stage_input_coverage_gate( + requirements, + declared_outputs=us_source_stage_outputs() - {"casualty_loss"}, + reviewed_exclusions={}, + name="us_validation_input_coverage", ) assert not result.passed - assert any("Stale reviewed exclusions" in line for line in result.failures) - assert "qualified_tuition_expenses" in result.details["stale_exclusions"] + assert set(result.details["missing"]) == {"casualty_loss"} - def test_removing_the_gap_reason_makes_the_row_fail(self) -> None: - # A registry entry with no reason (not a tracked gap) that is still - # unproduced must fail: this is the guard against re-shipping the #252 - # silent zero once the leaf is expected to be present. + def test_removing_misc_expense_makes_the_obbba_row_fail(self) -> None: requirements = us_validation_input_leaf_requirements() + leaf = "unreimbursed_business_employee_expenses" from populace.build.gates import source_stage_input_coverage_gate result = source_stage_input_coverage_gate( requirements, - declared_outputs=us_source_stage_outputs(), - reviewed_exclusions={}, # no leaf allowlisted + declared_outputs=us_source_stage_outputs() - {leaf}, + reviewed_exclusions={}, + name="us_validation_input_coverage", + ) + assert not result.passed + assert set(result.details["missing"]) == {leaf} + + def test_removing_childcare_makes_cdcc_rows_fail(self) -> None: + requirements = us_validation_input_leaf_requirements() + leaf = "spm_unit_pre_subsidy_childcare_expenses" + + from populace.build.gates import source_stage_input_coverage_gate + + result = source_stage_input_coverage_gate( + requirements, + declared_outputs=us_source_stage_outputs() - {leaf}, + reviewed_exclusions={}, name="us_validation_input_coverage", ) assert not result.passed - assert set(result.details["missing"]) == { - "qualified_tuition_expenses", - "qualified_passenger_vehicle_loan_interest", - } + assert set(result.details["missing"]) == {leaf} class TestValidationInputLeafRegistry: diff --git a/packages/populace-data/tests/test_release.py b/packages/populace-data/tests/test_release.py index f138b85a..38ad4d32 100644 --- a/packages/populace-data/tests/test_release.py +++ b/packages/populace-data/tests/test_release.py @@ -778,6 +778,39 @@ def test_publish_uploads_manifest_release_diagnostics_from_release_dir( ) +def test_publish_uploads_ssi_take_up_diagnostics_without_extra_files( + hub: FakeHub, release_dir: Path, artifact_root: Path +) -> None: + diagnostics_path = release_dir / "us_ssi_take_up.json" + diagnostics_path.write_text('{"schema_version": 1}') + manifest_path = release_dir / "release_manifest.json" + manifest = json.loads(manifest_path.read_text()) + manifest["artifacts"]["us_ssi_take_up"] = { + "kind": "diagnostics", + "path": diagnostics_path.name, + "repo_id": "policyengine/populace-us", + "revision": RELEASE_ID, + "sha256": _sha256(diagnostics_path), + } + manifest_path.write_text(json.dumps(manifest)) + + publish_release( + release_dir, + "policyengine/populace-us", + api=hub, + artifact_root=artifact_root, + updated_at="2026-06-11T13:53:15+00:00", + ) + + uploaded_paths = [path for path, _ in hub.uploads] + release_path = f"releases/{RELEASE_ID}/{diagnostics_path.name}" + assert diagnostics_path.name not in uploaded_paths + assert release_path in uploaded_paths + assert uploaded_paths.index(release_path) < uploaded_paths.index( + LATEST_POINTER_PATH + ) + + def test_publish_requires_artifact_root_for_root_artifacts( hub: FakeHub, release_dir: Path ) -> None: diff --git a/tools/build_us_fiscal_refresh_release.py b/tools/build_us_fiscal_refresh_release.py index adcc0394..09cdd981 100644 --- a/tools/build_us_fiscal_refresh_release.py +++ b/tools/build_us_fiscal_refresh_release.py @@ -22,8 +22,9 @@ import sys import time import tomllib -from collections.abc import Iterable, Mapping, Sequence +from collections.abc import Collection, Iterable, Mapping, Sequence from contextlib import contextmanager +from dataclasses import replace from datetime import UTC, datetime from pathlib import Path from typing import Any @@ -43,14 +44,26 @@ from populace.build.source_runtime import SourceRuntimeConfig, run_source_stage from populace.build.staging import StagingTelemetry from populace.build.us_runtime import ( + ASEC_2023_WEEKS_UNEMPLOYED_SOURCE_SHA256, CONGRESSIONAL_DISTRICT_VINTAGE_CROSSWALK_SHA256_ATTR, CONGRESSIONAL_DISTRICT_VINTAGE_TARGET_ATTR, CURRENT_CONGRESSIONAL_DISTRICT_VINTAGE, + ORG_2024_DONOR_CONTENT_SHA256, + SIPP_2023_HEAD_START_DONOR_SHA256, + SIPP_2023_HEAD_START_DONOR_SIZE_BYTES, + SIPP_2023_SSI_DISABILITY_DONOR_SHA256, + SIPP_2023_SSI_DISABILITY_DONOR_SIZE_BYTES, + SIPP_2023_TIP_DONOR_SHA256, + SIPP_2023_VEHICLE_DONOR_SHA256, + SIPP_2023_VEHICLE_DONOR_SIZE_BYTES, + SIPP_2023_VOLUNTARY_FILING_DONOR_SHA256, + SIPP_2023_VOLUNTARY_FILING_DONOR_SIZE_BYTES, SOI_VARIABLE_MAP, US_FISCAL_TARGET_COVERAGE_REQUIREMENTS, US_FISCAL_TARGET_SUPPORT_EXCLUSIONS, US_JCT_TAX_EXPENDITURE_REFORMS, US_MEDICAID_ENROLLMENT_TARGET_TABLE, + US_MEDICAID_TAKE_UP_VARIABLE, US_SOURCE_MANIFEST, apply_us_medicaid_enrollment_substitutions, assert_release_input_coverage_manifest_current, @@ -58,40 +71,110 @@ assert_take_up_treatments_consistent, assert_validation_leaf_registry_current, compile_us_fiscal_target_registry, + fetch_asec_2023_weeks_unemployed_source, + fetch_org_2024_donor, + fetch_scf_2022_full_extract, fetch_scf_2022_summary_extract, + fetch_sipp_2023_tip_donor, + fetch_sipp_2023_vehicle_donor, hard_target_package_aliases, + load_asec_2023_weeks_unemployed_source, load_congressional_district_vintage_crosswalk, + load_org_2024_donor, + load_scf_2022_auto_loan_donor, load_scf_2022_financial_asset_donor, + load_sipp_2023_head_start_donor, + load_sipp_2023_ssi_disability_donor, + load_sipp_2023_tip_donor, + load_sipp_2023_vehicle_donor, + load_sipp_2023_voluntary_filing_donor, + us_alimony_signal_gate, + us_capital_gain_details_signal_gate, + us_casualty_loss_signal_gate, + us_child_support_signal_gate, + us_childcare_signal_gate, + us_disability_benefits_signal_gate, + us_domestic_production_ald_signal_gate, + us_education_inputs_signal_gate, + us_educator_expense_signal_gate, us_eligibility_inputs_signal_gate, + us_energy_subsidy_signal_gate, + us_farm_business_income_signal_gate, + us_form_4952_election_signal_gate, us_hours_worked_signal_gate, + us_housing_inputs_signal_gate, us_immigration_composition_gate, + us_medicaid_source_person_table, + us_medicaid_take_up_diagnostics, us_medicaid_take_up_gate, + us_medicare_take_up_signal_gate, + us_misc_itemized_signal_gate, + us_org_wages_signal_gate, + us_other_health_insurance_signal_gate, us_pregnancy_signal_gate, + us_prior_year_income_signal_gate, + us_qbi_inputs_signal_gate, us_reform_coverage_smoke_gate, us_register_consistency_gate, + us_relationship_inputs_signal_gate, us_release_input_coverage_gate, + us_retirement_contributions_signal_gate, + us_retirement_distributions_signal_gate, + us_salt_refund_income_signal_gate, + us_scf_auto_loans_signal_gate, us_scf_wealth_signal_gate, + us_sipp_head_start_signal_gate, + us_sipp_tips_signal_gate, + us_sipp_vehicles_signal_gate, us_snap_discretionary_exemption_signal_gate, us_snap_state_take_up_gate, us_snap_take_up_signal_gate, us_source_coverage_diagnostics, us_source_operation_handlers, + us_ssi_disability_criteria_signal_gate, + us_ssi_take_up_diagnostics, + us_ssi_take_up_gate, + us_ssi_take_up_reporter_source_ids, us_take_up_participation_diagnostics, us_take_up_signal_gate, us_validation_input_coverage_gate, + us_voluntary_filing_signal_gate, + us_weeks_unemployed_signal_gate, + us_wic_claim_signal_gate, + us_workers_compensation_signal_gate, + with_us_childcare_inputs, + with_us_education_inputs, with_us_eligibility_inputs, + with_us_energy_subsidy_input, with_us_hours_worked_inputs, with_us_immigration_inputs, with_us_medicaid_take_up, + with_us_medicare_take_up_input, + with_us_org_wages_inputs, + with_us_other_health_insurance_inputs, with_us_pregnancy_inputs, + with_us_qbi_input_reconciliation, + with_us_relationship_inputs, + with_us_retirement_contribution_inputs, + with_us_retirement_distribution_inputs, + with_us_scf_auto_loan_inputs, with_us_scf_wealth_inputs, + with_us_sipp_head_start_input, + with_us_sipp_tip_inputs, + with_us_sipp_vehicle_inputs, with_us_snap_discretionary_exemption_inputs, with_us_snap_state_take_up, with_us_snap_take_up_inputs, + with_us_ssi_disability_criteria, + with_us_ssi_take_up, with_us_take_up_inputs, + with_us_voluntary_filing_input, + with_us_weeks_unemployed, + with_us_wic_claim_input, write_us_medicaid_take_up_diagnostics, write_us_snap_state_take_up_diagnostics, write_us_source_coverage_diagnostics, + write_us_ssi_take_up_diagnostics, write_us_take_up_participation_diagnostics, ) from populace.build.us_runtime.demographics import ( @@ -161,6 +244,7 @@ # while a change to any of these keys still invalidates the entry (no stale reuse). REFORM_VECTOR_CACHE_CONTEXT_KEYS: tuple[str, ...] = ( "base_dataset_sha256", + "weeks_unemployed_source_sha256", "policyengine_us_version", "target_period", "congressional_district_vintage_crosswalk_sha256", @@ -171,10 +255,23 @@ # medicaid_enrolled target columns differ from version-1 checkpoints; the # checkpoint identity hashes the on-disk base dataset, not the staged frame, # and would otherwise silently reuse pre-stage frames. -TARGET_FRAME_CHECKPOINT_MATERIALIZER_VERSION = 2 +# 3: the post-base SIPP SSI-disability stage restores +# meets_ssi_disability_criteria after SCF assets, changing SSI eligibility and +# target vectors without changing that on-disk base hash. +# 4: reporter-anchored SSI take-up now replaces the engine-default universal +# flag after the disability stage, changing SSI and its target vectors while +# the on-disk base hash remains unchanged. +# 5: the measured-SIPP Head Start stage now replaces the engine-default +# universal take-up flag before target materialization and must be present on +# every restored checkpoint even though the on-disk base hash is unchanged. +# 6: the official-ASEC sidecar restores 2022 LKWEEKS before target +# materialization. The source is external to the on-disk base hash, so old +# checkpoints must not survive the new measured input. +TARGET_FRAME_CHECKPOINT_MATERIALIZER_VERSION = 6 DEFAULT_MAXIMUM_MICROSIM_BATCH_SIZE = 5_000 DEFAULT_L0_REFIT_LAMBDA_SHARE = 0.8 DEFAULT_US_FISCAL_CALIBRATION_EPOCHS = 1_500 +SSI_TAKE_UP_RECONCILIATION_MAX_PASSES = 3 def _collect_batch_garbage() -> None: @@ -428,12 +525,6 @@ def _automatic_gc_suspended(): # degenerate column, or one of these becoming non-degenerate, fails the # default-valued-columns gate so this list cannot rot. US_DEGENERATE_INPUT_REVIEWED_EXCLUSIONS = { - "takes_up_ssi_if_eligible": ( - "SSI take-up imputation backlog; constant True forces 100% take-up." - ), - "takes_up_medicare_if_eligible": ( - "Medicare take-up imputation backlog (PolicyEngine/populace#98)." - ), "takes_up_dc_ptc": ("DC PTC take-up imputation backlog; constant True."), "second_home_mortgage_balance": ( "Second-home mortgage decomposition not imputed; constant at the" @@ -447,53 +538,23 @@ def _automatic_gc_suspended(): "Second-home mortgage decomposition not imputed; constant at the" " engine default (PolicyEngine/populace#38)." ), - "takes_up_head_start_if_eligible": ( - "Head Start take-up imputation backlog; constant True." - ), "takes_up_early_head_start_if_eligible": ( - "Early Head Start take-up imputation backlog; constant True." + "Early Head Start person enrollment is absent from every locked source; " + "see the archived-derivation and source-domain evidence in the parity " + "gap register (PolicyEngine/populace#312)." ), # ssn_card_type and immigration_status_str are intentionally NOT excluded: # PR #266 imputes them from CPS ASEC citizenship, so a base where they are # still constant at CITIZEN skipped that stage and should fail this gate. - "spm_unit_tenure_type": ( - "SPM tenure input not yet carried through (PolicyEngine/populace#32); " - "constant RENTER misstates SNAP shelter deductions and SPM housing." - ), "is_wic_at_nutritional_risk": ( - "WIC inputs not yet carried through (PolicyEngine/populace#32)." - ), - "would_claim_wic": ( - "WIC take-up inputs not yet carried through (PolicyEngine/populace#32)." + "Person-level nutritional-risk assessments are absent from all locked " + "sources; see the archived-derivation evidence in the parity gap " + "register (PolicyEngine/populace#312)." ), "s_corp_income": ( "Combined partnership/S-corp income is carried in partnership_income " "in pre-PUF-support bases; the S-corp leaf is constant zero there." ), - "estate_income_would_be_qualified": ( - "QBI qualification flags default True pending formula-constrained " - "leaf imputation (PolicyEngine/populace#186)." - ), - "farm_operations_income_would_be_qualified": ( - "QBI qualification flags default True pending formula-constrained " - "leaf imputation (PolicyEngine/populace#186)." - ), - "farm_rent_income_would_be_qualified": ( - "QBI qualification flags default True pending formula-constrained " - "leaf imputation (PolicyEngine/populace#186)." - ), - "partnership_s_corp_income_would_be_qualified": ( - "QBI qualification flags default True pending formula-constrained " - "leaf imputation (PolicyEngine/populace#186)." - ), - "rental_income_would_be_qualified": ( - "QBI qualification flags default True pending formula-constrained " - "leaf imputation (PolicyEngine/populace#186)." - ), - "self_employment_income_would_be_qualified": ( - "QBI qualification flags default True pending formula-constrained " - "leaf imputation (PolicyEngine/populace#186)." - ), } #: Person inputs SNAP work-requirement rules read that have NO CPS ASEC @@ -869,16 +930,63 @@ def _parse_args() -> argparse.Namespace: ), ) parser.add_argument("--seed", type=int, default=0) + parser.add_argument( + "--asec-2023-weeks-unemployed-source", + type=Path, + help=( + "Optional local path to the SHA-pinned official 2023 ASEC CSV ZIP " + "used to restore income-year-2022 LKWEEKS. When omitted the " + "official Census archive is fetched and verified." + ), + ) parser.add_argument( "--scf-summary-extract", dest="scf_summary_extract", default=None, help=( "Path to the Federal Reserve SCF 2022 public summary extract " - "(rscfp2022.dta) that feeds the SSI countable-resource asset " - "imputation (scf_wealth stage, populace#356/#368). When omitted the " - "fixed-vintage extract is fetched and cached from the Federal " - "Reserve." + "(rscfp2022.dta) that feeds the signed household net-worth and SSI " + "countable-resource asset imputations (scf_wealth stage, " + "populace#49/#356/#368). When omitted the fixed-vintage extract is " + "fetched and cached from the Federal Reserve." + ), + ) + parser.add_argument( + "--scf-full-extract", + type=Path, + help=( + "Optional path to the Federal Reserve SCF 2022 full public Stata " + "extract (p22i6.dta) used for household auto-loan balance and " + "interest. When omitted, scf2022s.zip is fetched and cached." + ), + ) + parser.add_argument( + "--sipp-tip-donor", + type=Path, + help=( + "Optional local path to the sha-pinned SIPP 2023 slim CSV that " + "feeds tip_income and Treasury tipped-occupation coverage. When " + "omitted the immutable donor revision is fetched and verified." + ), + ) + parser.add_argument( + "--sipp-vehicle-donor", + type=Path, + help=( + "Optional local path to the sha-pinned full SIPP 2023 public-use " + "file that feeds SSI disability criteria, household vehicle " + "count/value, and measured voluntary tax filing. When omitted the " + "immutable donor revision is fetched and verified." + ), + ) + parser.add_argument( + "--org-wages-donor", + type=Path, + help=( + "Optional local path to the canonical transformed 2024 CPS ORG " + "donor cache. When omitted, the twelve official 2024 CPS " + "basic-month ORG files are fetched, transformed, and verified " + "against the pinned canonical donor-content SHA-256." ), ) parser.add_argument( @@ -1300,6 +1408,7 @@ def _target_frame_checkpoint_identity( seed: int, target_period: int, target_registry_version: str, + weeks_unemployed_source_sha256: str, congressional_district_vintage_crosswalk_sha256: object, ) -> dict[str, object]: return { @@ -1308,6 +1417,7 @@ def _target_frame_checkpoint_identity( "kind": "us_fiscal_refresh_target_frame", "country": "us", "base_dataset_sha256": str(base_dataset_sha256), + "weeks_unemployed_source_sha256": str(weeks_unemployed_source_sha256), "policyengine_us_version": str(policyengine_us_version), "seed": int(seed), "target_period": int(target_period), @@ -2380,6 +2490,87 @@ def _medicaid_person_eligibility( return eligibility +def _ssi_person_uncapped_amount( + frame: Frame, + *, + simulation=None, + maximum_microsim_batch_size: int | None = DEFAULT_MAXIMUM_MICROSIM_BATCH_SIZE, +) -> np.ndarray: + """December person-level potential federal SSI, batched like Medicaid. + + SSA's age-band recipient counts are December 2024 point-in-time stocks. + ``uncapped_ssi > 0`` is the PolicyEngine-US 1.764.6 current-benefit + candidate mask and does not depend on the take-up input being assigned. + """ + + period = f"{PERIOD}-12" + + def calculate(active_simulation) -> np.ndarray: + values = np.asarray( + active_simulation.calculate( + "uncapped_ssi", + period=period, + map_to="person", + ), + dtype=np.float64, + ) + if not np.isfinite(values).all(): + raise RuntimeError( + "SSI take-up materialization produced nonfinite uncapped_ssi values." + ) + return values + + if simulation is not None: + return calculate(simulation) + + from policyengine_us import Microsimulation + + person_ids = frame.table("person")["person_id"].to_numpy() + uncapped = np.zeros(len(person_ids), dtype=np.float64) + person_positions = pd.Series( + np.arange(len(person_ids), dtype=np.int64), index=person_ids + ) + n_households = frame.n("household") + batches = tuple( + _household_position_batches(n_households, maximum_microsim_batch_size) + ) + if len(batches) > 1: + print( + "Materializing December SSI candidate amounts in " + f"{len(batches)} batches of up to " + f"{maximum_microsim_batch_size:,} households.", + flush=True, + ) + for household_positions in batches: + with _automatic_gc_suspended(): + full_batch = len(household_positions) == n_households + batch_frame = ( + frame + if full_batch + else _select_households_by_position(frame, household_positions) + ) + batch_simulation = Microsimulation( + dataset=_dataset_from_frame( + batch_frame, + assert_no_formula_owned_columns=False, + ) + ) + batch_uncapped = calculate(batch_simulation) + positions = person_positions.reindex( + batch_frame.table("person")["person_id"].to_numpy() + ).to_numpy() + if np.isnan(positions).any(): + raise RuntimeError( + "SSI candidate batch produced person_id values not present " + "in the full person table." + ) + uncapped[positions.astype(np.int64)] = batch_uncapped + batch_simulation._invalidate_all_caches() + del batch_frame, batch_simulation + _collect_batch_garbage() + return uncapped + + def _with_medicaid_take_up_outputs( frame: Frame, target_specs: tuple, @@ -2419,6 +2610,48 @@ def _with_medicaid_take_up_outputs( ) +def _medicaid_diagnostics_for_existing_output( + frame: Frame, + target_specs: tuple, + *, + seed: int, + substitutions: Sequence[dict[str, object]] = (), + maximum_microsim_batch_size: int | None = DEFAULT_MAXIMUM_MICROSIM_BATCH_SIZE, +) -> dict[str, object]: + """Diagnose persisted Medicaid flags on actual release weights.""" + + target_table = _medicaid_source_target_table(target_specs) + if target_table.empty: + raise RuntimeError( + "Final Medicaid take-up diagnostics require CMS state targets." + ) + eligibility = _medicaid_person_eligibility( + frame, + maximum_microsim_batch_size=maximum_microsim_batch_size, + ) + assigned = us_medicaid_source_person_table( + frame, + is_medicaid_eligible=eligibility, + state_fips=_person_state_fips(frame), + seed=seed, + ) + person = frame.table("person") + if US_MEDICAID_TAKE_UP_VARIABLE not in person: + raise RuntimeError( + f"Final release is missing person.{US_MEDICAID_TAKE_UP_VARIABLE}." + ) + takes_up = person[US_MEDICAID_TAKE_UP_VARIABLE] + if not pd.api.types.is_bool_dtype(takes_up.dtype) or takes_up.isna().any(): + raise RuntimeError("Final Medicaid take-up output must be complete boolean.") + assigned[US_MEDICAID_TAKE_UP_VARIABLE] = takes_up.to_numpy(dtype=bool) + return us_medicaid_take_up_diagnostics( + assigned, + target_table, + substitutions=substitutions, + weights_basis="final_release_weights", + ) + + def _snap_state_target_table(target_specs: tuple) -> pd.DataFrame: """FNS state household caseload counts as the take-up calibration table. @@ -4352,6 +4585,321 @@ def _fiscal_target_loss_weights(registry: TargetRegistry) -> np.ndarray: return weights / weights.mean() +class _SSITakeUpReconciliationResult: + """Fixed assignments and calibration state after SSI dependency replay.""" + + __slots__ = ( + "export_frame", + "calibration_result", + "registry", + "compilation", + "ssi_diagnostics", + "medicaid_diagnostics", + "health_input_gate", + "other_health_insurance_gate", + "passes", + ) + + def __init__( + self, + *, + export_frame: Frame, + calibration_result: Any, + registry: TargetRegistry, + compilation: Mapping[str, object], + ssi_diagnostics: Mapping[str, object], + medicaid_diagnostics: Mapping[str, object], + health_input_gate: GateResult, + other_health_insurance_gate: GateResult, + passes: int, + ) -> None: + self.export_frame = export_frame + self.calibration_result = calibration_result + self.registry = registry + self.compilation = compilation + self.ssi_diagnostics = ssi_diagnostics + self.medicaid_diagnostics = medicaid_diagnostics + self.health_input_gate = health_input_gate + self.other_health_insurance_gate = other_health_insurance_gate + self.passes = passes + + +def _frame_with_reconciliation_weight_basis( + frame: Frame, + household_weights: np.ndarray, +) -> Frame: + """Copy input tables onto the fixed pre-refit household weight basis.""" + + values = np.asarray(household_weights, dtype=np.float64) + if values.shape != (frame.n("household"),): + raise ValueError( + "SSI take-up reconciliation weight basis must align to households: " + f"{values.shape} != {(frame.n('household'),)}." + ) + if not np.isfinite(values).all() or (values <= 0).any(): + raise ValueError( + "SSI take-up reconciliation requires finite strictly positive " + "prior weights." + ) + weights = {entity: frame.weights_for(entity) for entity in frame.weighted_entities} + weights["household"] = Weights(values, WeightKind.CALIBRATED) + return Frame( + {entity: frame.table(entity).copy() for entity in frame.entities}, + frame.schema, + weights, + frame.strata, + mass_log=frame.mass_log, + ) + + +def _assert_reconciliation_support_unchanged( + reference: Frame, + candidate: Frame, +) -> None: + """Fail if a dependency replay changes selected entity IDs or order.""" + + if reference.schema != candidate.schema or reference.entities != candidate.entities: + raise RuntimeError("SSI take-up reconciliation changed the support schema.") + for entity in reference.entities: + id_column = ( + reference.schema.person_id_column + if entity == reference.schema.person_entity + else reference.schema.id_column(entity) + ) + expected = reference.table(entity)[id_column].to_numpy() + actual = candidate.table(entity)[id_column].to_numpy() + if not np.array_equal(expected, actual): + raise RuntimeError( + "SSI take-up reconciliation changed selected support IDs or " + f"order for {entity!r}." + ) + + +def _reconcile_ssi_take_up_and_refit( + base_frame: Frame, + initial_result, + target_specs: tuple, + *, + dense_default_dataset: bool, + seed: int, + epochs: int, + learning_rate: float, + max_weight_ratio: float | None, + l2_lambda: float, + target_loss_cap: float, + reporter_source_ids: Collection[str] | None = None, + medicaid_enrollment_substitutions: Sequence[Mapping[str, object]] = (), + maximum_microsim_batch_size: int | None = DEFAULT_MAXIMUM_MICROSIM_BATCH_SIZE, + gate_congressional_district_targets: bool = True, + progress_callback=None, + max_passes: int = SSI_TAKE_UP_RECONCILIATION_MAX_PASSES, +) -> _SSITakeUpReconciliationResult: + """Reconcile SSI on final weights before replaying dependent target inputs. + + The initial dense/L0 fit supplies an actual release-weight surface. Each + bounded pass then fixes SSI on those weights, replays ACA, Medicaid, and + other-health inputs that can depend on SSI, rematerializes every fiscal + target from those exact fixed inputs, and performs an ordinary refit on the + already-selected support. The returned-weight check diagnoses the persisted + SSI flags without rewriting them; a rewrite after optimization would make + both SSI and Medicaid target vectors stale. + """ + + if max_passes <= 0: + raise ValueError("SSI take-up reconciliation requires at least one pass.") + reporter_source_ids = ( + us_ssi_take_up_reporter_source_ids(base_frame) + if reporter_source_ids is None + else frozenset(str(value) for value in reporter_source_ids) + ) + if dense_default_dataset: + current_support = _with_calibrated_weights( + base_frame, + np.asarray(initial_result.weights, dtype=np.float64), + ) + else: + current_support = _with_l0_refit_weights(base_frame, initial_result) + prior_weights = np.asarray(initial_result.initial_weights, dtype=np.float64) + if prior_weights.shape != (current_support.n("household"),): + raise RuntimeError( + "SSI take-up reconciliation prior weights do not align to the " + "selected release support." + ) + selected_support = current_support + + last_failures: tuple[str, ...] = () + for pass_number in range(1, max_passes + 1): + uncapped_ssi = _ssi_person_uncapped_amount( + current_support, + maximum_microsim_batch_size=maximum_microsim_batch_size, + ) + assigned_support, stage_diagnostics = with_us_ssi_take_up( + current_support, + uncapped_ssi=uncapped_ssi, + seed=seed, + reporter_source_ids=reporter_source_ids, + ) + stage_gate = us_ssi_take_up_gate(stage_diagnostics) + if not stage_gate.passed: + raise RuntimeError( + "SSI take-up reconciliation assignment failed: " + + "; ".join(stage_gate.failures) + ) + + # SSI recipient status can alter Marketplace eligibility, Medicaid + # eligibility/take-up, and the ASEC private-premium residual. Replay + # that full dependency tail before any target vector is materialized. + assigned_support = _with_aca_marketplace_source_outputs( + assigned_support, + target_specs, + seed=seed, + maximum_microsim_batch_size=maximum_microsim_batch_size, + ) + health_gate = _health_input_signal_gate(assigned_support) + if not health_gate.passed: + raise RuntimeError( + "SSI take-up reconciliation ACA input gate failed: " + + "; ".join(health_gate.failures) + ) + assigned_support, medicaid_diagnostics = _with_medicaid_take_up_outputs( + assigned_support, + target_specs, + seed=seed, + substitutions=medicaid_enrollment_substitutions, + maximum_microsim_batch_size=maximum_microsim_batch_size, + ) + medicaid_gate = us_medicaid_take_up_gate(dict(medicaid_diagnostics)) + if not medicaid_gate.passed: + raise RuntimeError( + "SSI take-up reconciliation Medicaid gate failed: " + + "; ".join(medicaid_gate.failures) + ) + assigned_support = with_us_other_health_insurance_inputs( + assigned_support, + seed=seed, + time_period=PERIOD, + maximum_microsim_batch_size=maximum_microsim_batch_size, + ) + other_health_gate = us_other_health_insurance_signal_gate(assigned_support) + if not other_health_gate.passed: + raise RuntimeError( + "SSI take-up reconciliation other-health gate failed: " + + "; ".join(other_health_gate.failures) + ) + _assert_reconciliation_support_unchanged(selected_support, assigned_support) + + calibration_input = _frame_with_reconciliation_weight_basis( + assigned_support, + prior_weights, + ) + target_frame, registry, compilation = _materialize_target_frame( + calibration_input, + target_specs, + maximum_microsim_batch_size=maximum_microsim_batch_size, + gate_congressional_district_targets=gate_congressional_district_targets, + ) + current_weights = np.asarray( + assigned_support.weights_for("household").values, + dtype=np.float64, + ) + reconciled_result = calibrate( + target_frame, + registry.to_target_set(), + epochs=epochs, + learning_rate=learning_rate, + max_weight_ratio=max_weight_ratio, + seed=seed, + mass="conserve", + l2_lambda=l2_lambda, + target_loss_weights=_fiscal_target_loss_weights(registry), + target_loss_cap=target_loss_cap, + warm_start_weights=current_weights, + progress_callback=progress_callback, + ) + export_frame = _with_calibrated_weights( + calibration_input, + np.asarray(reconciled_result.weights, dtype=np.float64), + ) + _assert_reconciliation_support_unchanged(selected_support, export_frame) + final_uncapped_ssi = _ssi_person_uncapped_amount( + export_frame, + maximum_microsim_batch_size=maximum_microsim_batch_size, + ) + final_diagnostics = us_ssi_take_up_diagnostics( + export_frame, + uncapped_ssi=final_uncapped_ssi, + seed=seed, + reporter_source_ids=reporter_source_ids, + ) + final_gate = us_ssi_take_up_gate(final_diagnostics) + final_health_gate = _health_input_signal_gate(export_frame) + final_medicaid_diagnostics = _medicaid_diagnostics_for_existing_output( + export_frame, + target_specs, + seed=seed, + substitutions=medicaid_enrollment_substitutions, + maximum_microsim_batch_size=maximum_microsim_batch_size, + ) + final_medicaid_gate = us_medicaid_take_up_gate(final_medicaid_diagnostics) + final_other_health_gate = us_other_health_insurance_signal_gate(export_frame) + if ( + final_gate.passed + and final_health_gate.passed + and final_medicaid_gate.passed + and final_other_health_gate.passed + ): + reconciliation_compilation = { + **dict(compilation), + "target_frame_checkpoint": { + "enabled": False, + "status": "recomputed_after_ssi_take_up_reconciliation", + }, + "ssi_take_up_reconciliation": { + "passes": pass_number, + "max_passes": max_passes, + "reporter_source_identity_count": len(reporter_source_ids), + "dependency_replay": [ + "aca_marketplace", + "medicaid_take_up", + "other_health_insurance", + "fiscal_target_materialization", + "ordinary_refit", + ], + }, + } + calibration_result = ( + reconciled_result + if dense_default_dataset + else replace(initial_result, refit=reconciled_result) + ) + return _SSITakeUpReconciliationResult( + export_frame=export_frame, + calibration_result=calibration_result, + registry=registry, + compilation=reconciliation_compilation, + ssi_diagnostics=final_diagnostics, + medicaid_diagnostics=final_medicaid_diagnostics, + health_input_gate=final_health_gate, + other_health_insurance_gate=final_other_health_gate, + passes=pass_number, + ) + last_failures = tuple( + [f"SSI: {failure}" for failure in final_gate.failures] + + [f"ACA: {failure}" for failure in final_health_gate.failures] + + [f"Medicaid: {failure}" for failure in final_medicaid_gate.failures] + + [ + f"Other health: {failure}" + for failure in final_other_health_gate.failures + ] + ) + current_support = export_frame + + raise RuntimeError( + "SSI take-up reconciliation did not remain count-faithful on returned " + f"weights after {max_passes} pass(es): " + "; ".join(last_failures) + ) + + def _fiscal_target_concept_budget_weights(registry: TargetRegistry) -> np.ndarray: weights = _fiscal_target_value_basis_weights(registry) group_indices: dict[tuple[object, ...], list[int]] = {} @@ -5549,6 +6097,12 @@ def _build_manifests( kind="diagnostics", revision=release_id, ), + "us_ssi_take_up": _artifact_entry( + "us_ssi_take_up.json", + _sha256(release_dir / "us_ssi_take_up.json"), + kind="diagnostics", + revision=release_id, + ), **( { "reform_validation": _artifact_entry( @@ -5879,6 +6433,52 @@ def main() -> None: if telemetry is not None: telemetry.stage("load_base_frame", message="Loading base population H5.") base_frame = _load_frame(base_h5) + weeks_unemployed_source_path = ( + args.asec_2023_weeks_unemployed_source + if args.asec_2023_weeks_unemployed_source is not None + else fetch_asec_2023_weeks_unemployed_source() + ) + weeks_unemployed_source = load_asec_2023_weeks_unemployed_source( + weeks_unemployed_source_path + ) + base_frame = with_us_weeks_unemployed( + base_frame, + seed=args.seed, + time_period=PERIOD, + asec_2023_source=weeks_unemployed_source, + ) + weeks_unemployed_gate = us_weeks_unemployed_signal_gate(base_frame) + if not weeks_unemployed_gate.passed: + if telemetry is not None: + telemetry.stage( + "weeks_unemployed_input_gate", + status="failed", + message="Weeks-unemployed input signal gate failed.", + failures=list(weeks_unemployed_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + "Weeks-unemployed input signal failed: " + failure + for failure in weeks_unemployed_gate.failures + ) + ) + if telemetry is not None: + telemetry.stage( + "weeks_unemployed_input", + message=( + "Restored measured ASEC LKWEEKS before frozen-support " + "selection and target materialization." + ), + source_path=str(Path(weeks_unemployed_source_path).resolve()), + source_sha256=ASEC_2023_WEEKS_UNEMPLOYED_SOURCE_SHA256, + source_rows=len(weeks_unemployed_source), + ) + # Capture direct ASEC reporter lineage on the FULL clone-aware support. + # Frozen-support recovery may retain only a PUF clone; deriving anchors + # after that prune would erase the underlying measured ASEC reporter. + ssi_reporter_source_ids = us_ssi_take_up_reporter_source_ids(base_frame) # Frozen-support recovery (populace#328): if a selection source is supplied, # reduce the base pool to exactly that support by stable source identity, @@ -5913,6 +6513,28 @@ def main() -> None: n_base_candidates=selection_report.n_base_candidates, n_unmapped=selection_report.n_unmapped, ) + post_selection_weeks_unemployed_gate = us_weeks_unemployed_signal_gate( + base_frame + ) + if not post_selection_weeks_unemployed_gate.passed: + if telemetry is not None: + telemetry.stage( + "post_selection_weeks_unemployed_input_gate", + status="failed", + message=( + "Frozen-support selection collapsed weeks-unemployed " + "input signal." + ), + failures=list(post_selection_weeks_unemployed_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + "Post-selection weeks-unemployed input signal failed: " + failure + for failure in post_selection_weeks_unemployed_gate.failures + ) + ) base_frame, base_population_repair = _with_base_population_mass_repair(base_frame) base_frame, social_security_component_repair = ( @@ -5953,15 +6575,310 @@ def main() -> None: for failure in base_population_gate.failures ) ) - if telemetry is not None: - telemetry.stage( - "immigration_inputs", - message="Deriving SSN card type and immigration status inputs.", - ) - base_frame = with_us_immigration_inputs( - base_frame, - seed=args.seed, - time_period=PERIOD, + base_frame = with_us_qbi_input_reconciliation(base_frame) + qbi_inputs_gate = us_qbi_inputs_signal_gate(base_frame) + if not qbi_inputs_gate.passed: + if telemetry is not None: + telemetry.stage( + "qbi_input_gate", + status="failed", + message="QBI-input signal gate failed.", + failures=list(qbi_inputs_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + "QBI-input signal failed: " + failure + for failure in qbi_inputs_gate.failures + ) + ) + farm_business_income_gate = us_farm_business_income_signal_gate(base_frame) + if not farm_business_income_gate.passed: + if telemetry is not None: + telemetry.stage( + "farm_business_income_gate", + status="failed", + message="Farm-business-income signal gate failed.", + failures=list(farm_business_income_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + "Farm-business-income signal failed: " + failure + for failure in farm_business_income_gate.failures + ) + ) + domestic_production_ald_gate = us_domestic_production_ald_signal_gate(base_frame) + if not domestic_production_ald_gate.passed: + if telemetry is not None: + telemetry.stage( + "domestic_production_ald_gate", + status="failed", + message="Domestic-production-ALD signal gate failed.", + failures=list(domestic_production_ald_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + "Domestic-production-ALD signal failed: " + failure + for failure in domestic_production_ald_gate.failures + ) + ) + child_support_gate = us_child_support_signal_gate(base_frame) + if not child_support_gate.passed: + if telemetry is not None: + telemetry.stage( + "child_support_input_gate", + status="failed", + message="Child-support input signal gate failed.", + failures=list(child_support_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + "Child-support input signal failed: " + failure + for failure in child_support_gate.failures + ) + ) + disability_benefits_gate = us_disability_benefits_signal_gate(base_frame) + if not disability_benefits_gate.passed: + if telemetry is not None: + telemetry.stage( + "disability_benefits_input_gate", + status="failed", + message="Disability-benefits input signal gate failed.", + failures=list(disability_benefits_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + "Disability-benefits input signal failed: " + failure + for failure in disability_benefits_gate.failures + ) + ) + workers_compensation_gate = us_workers_compensation_signal_gate(base_frame) + if not workers_compensation_gate.passed: + if telemetry is not None: + telemetry.stage( + "workers_compensation_input_gate", + status="failed", + message="Workers-compensation input signal gate failed.", + failures=list(workers_compensation_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + "Workers-compensation input signal failed: " + failure + for failure in workers_compensation_gate.failures + ) + ) + educator_expense_gate = us_educator_expense_signal_gate(base_frame) + if not educator_expense_gate.passed: + if telemetry is not None: + telemetry.stage( + "educator_expense_input_gate", + status="failed", + message="Educator-expense input signal gate failed.", + failures=list(educator_expense_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + "Educator-expense input signal failed: " + failure + for failure in educator_expense_gate.failures + ) + ) + form_4952_election_gate = us_form_4952_election_signal_gate(base_frame) + if not form_4952_election_gate.passed: + if telemetry is not None: + telemetry.stage( + "form_4952_election_input_gate", + status="failed", + message="Form 4952 election input signal gate failed.", + failures=list(form_4952_election_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + "Form 4952 election signal failed: " + failure + for failure in form_4952_election_gate.failures + ) + ) + salt_refund_income_gate = us_salt_refund_income_signal_gate(base_frame) + if not salt_refund_income_gate.passed: + if telemetry is not None: + telemetry.stage( + "salt_refund_income_input_gate", + status="failed", + message="SALT-refund-income signal gate failed.", + failures=list(salt_refund_income_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + "SALT-refund-income signal failed: " + failure + for failure in salt_refund_income_gate.failures + ) + ) + capital_gain_details_gate = us_capital_gain_details_signal_gate(base_frame) + if not capital_gain_details_gate.passed: + if telemetry is not None: + telemetry.stage( + "capital_gain_details_input_gate", + status="failed", + message="Capital-gain details input signal gate failed.", + failures=list(capital_gain_details_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + "Capital-gain details signal failed: " + failure + for failure in capital_gain_details_gate.failures + ) + ) + base_frame = with_us_childcare_inputs( + base_frame, + seed=args.seed, + time_period=PERIOD, + allow_existing_without_source=True, + ) + childcare_gate = us_childcare_signal_gate(base_frame) + if not childcare_gate.passed: + if telemetry is not None: + telemetry.stage( + "childcare_input_gate", + status="failed", + message="Childcare-input signal gate failed.", + failures=list(childcare_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + "Childcare-input signal failed: " + failure + for failure in childcare_gate.failures + ) + ) + base_frame = with_us_energy_subsidy_input( + base_frame, + seed=args.seed, + time_period=PERIOD, + allow_existing_without_source=True, + ) + energy_subsidy_gate = us_energy_subsidy_signal_gate(base_frame) + if not energy_subsidy_gate.passed: + if telemetry is not None: + telemetry.stage( + "energy_subsidy_gate", + status="failed", + message="Energy-subsidy signal gate failed.", + failures=list(energy_subsidy_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + "Energy-subsidy signal failed: " + failure + for failure in energy_subsidy_gate.failures + ) + ) + alimony_gate = us_alimony_signal_gate(base_frame) + if not alimony_gate.passed: + if telemetry is not None: + telemetry.stage( + "alimony_input_gate", + status="failed", + message="Alimony-input signal gate failed.", + failures=list(alimony_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + "Alimony-input signal failed: " + failure + for failure in alimony_gate.failures + ) + ) + casualty_loss_gate = us_casualty_loss_signal_gate(base_frame) + if not casualty_loss_gate.passed: + if telemetry is not None: + telemetry.stage( + "casualty_loss_input_gate", + status="failed", + message="Casualty-loss signal gate failed.", + failures=list(casualty_loss_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + "Casualty-loss signal failed: " + failure + for failure in casualty_loss_gate.failures + ) + ) + misc_itemized_gate = us_misc_itemized_signal_gate(base_frame) + if not misc_itemized_gate.passed: + if telemetry is not None: + telemetry.stage( + "misc_itemized_input_gate", + status="failed", + message="Miscellaneous-itemized signal gate failed.", + failures=list(misc_itemized_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + "Miscellaneous-itemized signal failed: " + failure + for failure in misc_itemized_gate.failures + ) + ) + if telemetry is not None: + telemetry.stage( + "retirement_contribution_inputs", + message=("Verifying ASEC-sourced desired retirement-contribution inputs."), + ) + base_frame = with_us_retirement_contribution_inputs( + base_frame, + seed=args.seed, + time_period=PERIOD, + ) + retirement_contributions_gate = us_retirement_contributions_signal_gate(base_frame) + if not retirement_contributions_gate.passed: + if telemetry is not None: + telemetry.stage( + "retirement_contribution_inputs_gate", + status="failed", + message="Retirement-contribution signal gate failed.", + failures=list(retirement_contributions_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + "Retirement-contribution signal failed: " + failure + for failure in retirement_contributions_gate.failures + ) + ) + if telemetry is not None: + telemetry.stage( + "immigration_inputs", + message="Deriving SSN card type and immigration status inputs.", + ) + base_frame = with_us_immigration_inputs( + base_frame, + seed=args.seed, + time_period=PERIOD, ) immigration_gate = us_immigration_composition_gate(base_frame) if not immigration_gate.passed: @@ -6060,6 +6977,135 @@ def main() -> None: for failure in snap_take_up_gate.failures ) ) + if telemetry is not None: + telemetry.stage( + "relationship_inputs", + message=("Deriving household-head and marital-status inputs from ASEC."), + ) + base_frame = with_us_relationship_inputs( + base_frame, + seed=args.seed, + time_period=PERIOD, + ) + relationship_inputs_gate = us_relationship_inputs_signal_gate(base_frame) + if not relationship_inputs_gate.passed: + if telemetry is not None: + telemetry.stage( + "relationship_inputs_gate", + status="failed", + message="Relationship-input signal gate failed.", + failures=list(relationship_inputs_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + f"Relationship-input signal failed: {failure}" + for failure in relationship_inputs_gate.failures + ) + ) + if telemetry is not None: + telemetry.stage( + "medicare_take_up_input", + message="Deriving measured Medicare enrollment from ASEC MCARE.", + ) + base_frame = with_us_medicare_take_up_input( + base_frame, + seed=args.seed, + time_period=PERIOD, + ) + medicare_take_up_gate = us_medicare_take_up_signal_gate(base_frame) + if not medicare_take_up_gate.passed: + if telemetry is not None: + telemetry.stage( + "medicare_take_up_input_gate", + status="failed", + message="Medicare take-up input signal gate failed.", + failures=list(medicare_take_up_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + f"Medicare take-up input signal failed: {failure}" + for failure in medicare_take_up_gate.failures + ) + ) + prior_year_income_gate = us_prior_year_income_signal_gate(base_frame) + if not prior_year_income_gate.passed: + if telemetry is not None: + telemetry.stage( + "prior_year_income_gate", + status="failed", + message="Prior-year-income signal gate failed.", + failures=list(prior_year_income_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + f"Prior-year-income signal failed: {failure}" + for failure in prior_year_income_gate.failures + ) + ) + if telemetry is not None: + telemetry.stage( + "prior_year_income_gate", + message="Verified restored adjacent-ASEC prior-year income inputs.", + details=dict(prior_year_income_gate.details), + ) + housing_inputs_gate = us_housing_inputs_signal_gate(base_frame) + if not housing_inputs_gate.passed: + if telemetry is not None: + telemetry.stage( + "housing_inputs_gate", + status="failed", + message="Housing/tenure input signal gate failed.", + failures=list(housing_inputs_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + f"Housing/tenure input signal failed: {failure}" + for failure in housing_inputs_gate.failures + ) + ) + if telemetry is not None: + telemetry.stage( + "housing_inputs_gate", + message="Verified restored CPS/ACS housing and tenure inputs.", + details=dict(housing_inputs_gate.details), + ) + if telemetry is not None: + telemetry.stage( + "retirement_distribution_inputs", + message=( + "Carrying measured ASEC retirement distributions by account type." + ), + ) + base_frame = with_us_retirement_distribution_inputs( + base_frame, + seed=args.seed, + time_period=PERIOD, + ) + retirement_distributions_gate = us_retirement_distributions_signal_gate(base_frame) + if not retirement_distributions_gate.passed: + if telemetry is not None: + telemetry.stage( + "retirement_distribution_inputs_gate", + status="failed", + message="Retirement-distribution signal gate failed.", + failures=list(retirement_distributions_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + "Retirement-distribution signal failed: " + failure + for failure in retirement_distributions_gate.failures + ) + ) if telemetry is not None: telemetry.stage( "eligibility_inputs", @@ -6087,6 +7133,36 @@ def main() -> None: for failure in eligibility_inputs_gate.failures ) ) + if telemetry is not None: + telemetry.stage( + "education_inputs", + message=( + "Carrying ASEC educational assistance and deriving AOTC " + "factual inputs from PUF qualified tuition." + ), + ) + base_frame = with_us_education_inputs( + base_frame, + seed=args.seed, + time_period=PERIOD, + ) + education_inputs_gate = us_education_inputs_signal_gate(base_frame) + if not education_inputs_gate.passed: + if telemetry is not None: + telemetry.stage( + "education_inputs_gate", + status="failed", + message="Education-input signal gate failed.", + failures=list(education_inputs_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + f"Education-input signal failed: {failure}" + for failure in education_inputs_gate.failures + ) + ) if telemetry is not None: telemetry.stage( "pregnancy_inputs", @@ -6114,6 +7190,33 @@ def main() -> None: for failure in pregnancy_gate.failures ) ) + if telemetry is not None: + telemetry.stage( + "wic_claim_input", + message="Assigning WIC claims from USDA FNS category coverage rates.", + ) + base_frame = with_us_wic_claim_input( + base_frame, + seed=args.seed, + time_period=PERIOD, + ) + wic_claim_gate = us_wic_claim_signal_gate(base_frame) + if not wic_claim_gate.passed: + if telemetry is not None: + telemetry.stage( + "wic_claim_input_gate", + status="failed", + message="WIC-claim input signal gate failed.", + failures=list(wic_claim_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + f"WIC-claim input signal failed: {failure}" + for failure in wic_claim_gate.failures + ) + ) if telemetry is not None: telemetry.stage( "snap_discretionary_exemption_inputs", @@ -6147,15 +7250,16 @@ def main() -> None: telemetry.stage( "scf_wealth_inputs", message=( - "Imputing SSI countable-resource assets (bank/stock/bond) from " - "the Federal Reserve SCF 2022 summary extract." + "Imputing signed household net worth and SSI countable-resource " + "assets (bank/stock/bond) from the Federal Reserve SCF 2022 " + "summary extract." ), ) - # populace#356/#368 Deliverable 2: restore the three SSI countable-resource - # asset inputs (bank_account_assets / stock_assets / bond_assets). Without - # them ssi_countable_resources is 0 for every record and every SSI - # resource-limit reform silently scores $0 — the failure the #368 column and - # reform-coverage gates are RED on. A CLI-supplied extract path is used when + # populace#49/#356/#368: restore signed household net_worth plus the three + # SSI countable-resource asset inputs (bank_account_assets / stock_assets / + # bond_assets) from their direct SCF summary-extract targets. Without the + # latter, ssi_countable_resources is 0 for every record and SSI resource- + # limit reforms silently score $0. A CLI-supplied extract path is used when # given; otherwise the fixed-vintage public extract is fetched and cached. scf_summary_extract_path = ( Path(args.scf_summary_extract) @@ -6186,6 +7290,319 @@ def main() -> None: for failure in scf_wealth_gate.failures ) ) + # The SSI criterion, Head Start, vehicle, and filing families share one + # immutable full SIPP artifact. Resolve it once here because the criterion's + # receiver also needs the SCF asset leaves materialized immediately above. + sipp_vehicle_donor_path = ( + Path(args.sipp_vehicle_donor) + if args.sipp_vehicle_donor is not None + else fetch_sipp_2023_vehicle_donor() + ) + if telemetry is not None: + telemetry.stage( + "ssi_disability_criteria", + message=( + "Imputing SSI-specific disability criteria from the " + "sha-pinned full SIPP 2023 donor after SCF assets." + ), + ) + ssi_disability_donor = load_sipp_2023_ssi_disability_donor( + sipp_vehicle_donor_path, + expected_sha256=SIPP_2023_SSI_DISABILITY_DONOR_SHA256, + expected_size_bytes=SIPP_2023_SSI_DISABILITY_DONOR_SIZE_BYTES, + time_period=PERIOD, + ) + base_frame = with_us_ssi_disability_criteria( + base_frame, + # The retired weighted bootstrap and MicroImpute forest are fixed. + seed=42, + time_period=PERIOD, + sipp_donor=ssi_disability_donor, + ) + ssi_disability_gate = us_ssi_disability_criteria_signal_gate(base_frame) + if not ssi_disability_gate.passed: + if telemetry is not None: + telemetry.stage( + "ssi_disability_criteria_gate", + status="failed", + message="SSI disability-criteria signal gate failed.", + failures=list(ssi_disability_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + f"SSI disability-criteria signal failed: {failure}" + for failure in ssi_disability_gate.failures + ) + ) + if telemetry is not None: + telemetry.stage( + "sipp_head_start", + message=( + "Imputing age-3--5 Head Start take-up from direct and strict " + "structural December responses in the sha-pinned full SIPP " + "2023 donor." + ), + ) + head_start_donor = load_sipp_2023_head_start_donor( + sipp_vehicle_donor_path, + expected_sha256=SIPP_2023_HEAD_START_DONOR_SHA256, + expected_size_bytes=SIPP_2023_HEAD_START_DONOR_SIZE_BYTES, + ) + base_frame = with_us_sipp_head_start_input( + base_frame, + seed=args.seed, + time_period=PERIOD, + sipp_donor=head_start_donor, + ) + head_start_gate = us_sipp_head_start_signal_gate(base_frame) + if not head_start_gate.passed: + if telemetry is not None: + telemetry.stage( + "sipp_head_start_gate", + status="failed", + message="SIPP Head Start signal gate failed.", + failures=list(head_start_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + f"SIPP Head Start signal failed: {failure}" + for failure in head_start_gate.failures + ) + ) + if telemetry is not None: + telemetry.stage( + "ssi_take_up", + message=( + "Assigning SSI take-up from ASEC reporter anchors and SSA " + "December 2024 federal-payment recipient counts by age." + ), + ) + ssi_uncapped_amount = _ssi_person_uncapped_amount( + base_frame, + maximum_microsim_batch_size=args.maximum_microsim_batch_size, + ) + base_frame, ssi_take_up_stage_diagnostics = with_us_ssi_take_up( + base_frame, + uncapped_ssi=ssi_uncapped_amount, + seed=args.seed, + reporter_source_ids=ssi_reporter_source_ids, + ) + ssi_take_up_gate = us_ssi_take_up_gate(ssi_take_up_stage_diagnostics) + if not ssi_take_up_gate.passed: + if telemetry is not None: + telemetry.stage( + "ssi_take_up_gate", + status="failed", + message="SSI take-up count-calibration gate failed.", + failures=list(ssi_take_up_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + f"SSI take-up failed: {failure}" + for failure in ssi_take_up_gate.failures + ) + ) + if telemetry is not None: + telemetry.stage( + "scf_auto_loan_inputs", + message=( + "Imputing household auto-loan balance and interest from the " + "full Federal Reserve SCF 2022 extract, then deriving the " + "OBBBA qualifying-interest proxy." + ), + ) + scf_full_extract_path = ( + Path(args.scf_full_extract) + if args.scf_full_extract is not None + else fetch_scf_2022_full_extract() + ) + scf_auto_loan_donor = load_scf_2022_auto_loan_donor( + scf_summary_extract_path, + scf_full_extract_path, + ) + base_frame = with_us_scf_auto_loan_inputs( + base_frame, + seed=args.seed, + time_period=PERIOD, + scf_auto_loan_donor=scf_auto_loan_donor, + ) + scf_auto_loan_gate = us_scf_auto_loans_signal_gate(base_frame) + if not scf_auto_loan_gate.passed: + if telemetry is not None: + telemetry.stage( + "scf_auto_loan_gate", + status="failed", + message="SCF auto-loan signal gate failed.", + failures=list(scf_auto_loan_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + f"SCF auto-loan signal failed: {failure}" + for failure in scf_auto_loan_gate.failures + ) + ) + if telemetry is not None: + telemetry.stage( + "sipp_vehicle_inputs", + message=( + "Imputing household vehicle count and value from the " + "sha-pinned full SIPP 2023 donor." + ), + ) + sipp_vehicle_donor = load_sipp_2023_vehicle_donor( + sipp_vehicle_donor_path, + expected_sha256=SIPP_2023_VEHICLE_DONOR_SHA256, + expected_size_bytes=SIPP_2023_VEHICLE_DONOR_SIZE_BYTES, + ) + base_frame = with_us_sipp_vehicle_inputs( + base_frame, + # The retired MicroImpute model pins both forests to seed 42. + seed=42, + time_period=PERIOD, + sipp_donor=sipp_vehicle_donor, + ) + sipp_vehicles_gate = us_sipp_vehicles_signal_gate(base_frame) + if not sipp_vehicles_gate.passed: + if telemetry is not None: + telemetry.stage( + "sipp_vehicles_gate", + status="failed", + message="SIPP-vehicle signal gate failed.", + failures=list(sipp_vehicles_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + f"SIPP-vehicle signal failed: {failure}" + for failure in sipp_vehicles_gate.failures + ) + ) + if telemetry is not None: + telemetry.stage( + "voluntary_filing_input", + message=( + "Imputing measured voluntary tax-filing responses from the " + "same sha-pinned full SIPP 2023 donor." + ), + ) + voluntary_filing_donor = load_sipp_2023_voluntary_filing_donor( + sipp_vehicle_donor_path, + expected_sha256=SIPP_2023_VOLUNTARY_FILING_DONOR_SHA256, + expected_size_bytes=SIPP_2023_VOLUNTARY_FILING_DONOR_SIZE_BYTES, + ) + base_frame = with_us_voluntary_filing_input( + base_frame, + seed=args.seed, + time_period=PERIOD, + sipp_donor=voluntary_filing_donor, + ) + voluntary_filing_gate = us_voluntary_filing_signal_gate(base_frame) + if not voluntary_filing_gate.passed: + if telemetry is not None: + telemetry.stage( + "voluntary_filing_gate", + status="failed", + message="Voluntary-filing signal gate failed.", + failures=list(voluntary_filing_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + f"Voluntary-filing signal failed: {failure}" + for failure in voluntary_filing_gate.failures + ) + ) + if telemetry is not None: + telemetry.stage( + "sipp_tip_inputs", + message=( + "Imputing tip income from the sha-pinned SIPP donor and " + "carrying Treasury tipped-occupation codes from ASEC." + ), + ) + sipp_tip_donor_path = ( + Path(args.sipp_tip_donor) + if args.sipp_tip_donor is not None + else fetch_sipp_2023_tip_donor() + ) + sipp_tip_donor = load_sipp_2023_tip_donor( + sipp_tip_donor_path, + expected_sha256=SIPP_2023_TIP_DONOR_SHA256, + ) + base_frame = with_us_sipp_tip_inputs( + base_frame, + seed=args.seed, + time_period=PERIOD, + sipp_donor=sipp_tip_donor, + ) + sipp_tips_gate = us_sipp_tips_signal_gate(base_frame) + if not sipp_tips_gate.passed: + if telemetry is not None: + telemetry.stage( + "sipp_tips_gate", + status="failed", + message="SIPP-tip signal gate failed.", + failures=list(sipp_tips_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + f"SIPP-tip signal failed: {failure}" + for failure in sipp_tips_gate.failures + ) + ) + if telemetry is not None: + telemetry.stage( + "org_wages_inputs", + message=( + "Imputing CPS ORG hourly-pay inputs, carrying ASEC occupation " + "groups, assigning BLS union coverage, and deriving the FLSA " + "overtime premium." + ), + ) + org_wages_donor_path = ( + Path(args.org_wages_donor) + if args.org_wages_donor is not None + else fetch_org_2024_donor() + ) + org_wages_donor = load_org_2024_donor( + org_wages_donor_path, + expected_content_sha256=ORG_2024_DONOR_CONTENT_SHA256, + ) + base_frame = with_us_org_wages_inputs( + base_frame, + seed=args.seed, + time_period=PERIOD, + org_donor=org_wages_donor, + ) + org_wages_gate = us_org_wages_signal_gate(base_frame) + if not org_wages_gate.passed: + if telemetry is not None: + telemetry.stage( + "org_wages_gate", + status="failed", + message="CPS ORG/FLSA signal gate failed.", + failures=list(org_wages_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + f"CPS ORG/FLSA signal failed: {failure}" + for failure in org_wages_gate.failures + ) + ) if telemetry is not None: telemetry.stage( "source_inputs", @@ -6274,6 +7691,37 @@ def main() -> None: for failure in snap_state_take_up_gate.failures ) ) + if telemetry is not None: + telemetry.stage( + "other_health_insurance_inputs", + message=( + "Deriving the ASEC non-Medicare premium residual after ACA, " + "CHIP, and Medicaid premiums, then imputing the PUF support." + ), + ) + base_frame = with_us_other_health_insurance_inputs( + base_frame, + seed=args.seed, + time_period=PERIOD, + maximum_microsim_batch_size=args.maximum_microsim_batch_size, + ) + other_health_insurance_gate = us_other_health_insurance_signal_gate(base_frame) + if not other_health_insurance_gate.passed: + if telemetry is not None: + telemetry.stage( + "other_health_insurance_input_gate", + status="failed", + message="Other-health-insurance premium signal gate failed.", + failures=list(other_health_insurance_gate.failures), + force_upload=True, + ) + raise RuntimeError( + "Release gates failed: " + + "; ".join( + "Other-health-insurance premium signal failed: " + failure + for failure in other_health_insurance_gate.failures + ) + ) if telemetry is not None and args.input_mass_reference_h5 is not None: telemetry.stage( "input_mass_reference_gate", @@ -6356,6 +7804,7 @@ def main() -> None: seed=args.seed, target_period=PERIOD, target_registry_version=active_target_registry.version, + weeks_unemployed_source_sha256=(ASEC_2023_WEEKS_UNEMPLOYED_SOURCE_SHA256), congressional_district_vintage_crosswalk_sha256=( congressional_district_vintage_crosswalk_metadata or {} ).get("sha256"), @@ -6369,6 +7818,9 @@ def main() -> None: target_materialization_cache_dir=target_materialization_cache_dir, target_materialization_cache_context={ "base_dataset_sha256": base_dataset_sha256, + "weeks_unemployed_source_sha256": ( + ASEC_2023_WEEKS_UNEMPLOYED_SOURCE_SHA256 + ), "build_commit": full_commit, "policyengine_us_version": policyengine_us_version, "seed": args.seed, @@ -6508,6 +7960,64 @@ def main() -> None: "refit_initial_loss": float(result.initial_loss), "refit_final_loss": float(result.final_loss), } + if telemetry is not None: + telemetry.stage( + "ssi_take_up_reconciliation", + message=( + "Reconciling SSI on release weights, replaying dependent health " + "inputs, and rematerializing fiscal targets before final refit." + ), + max_passes=SSI_TAKE_UP_RECONCILIATION_MAX_PASSES, + ) + pre_reconciliation_final_loss = float(result.final_loss) + reconciliation_l2_lambda = float( + args.l2_lambda + if args.dense_default_dataset or args.refit_l2_lambda is None + else args.refit_l2_lambda + ) + reconciliation = _reconcile_ssi_take_up_and_refit( + base_frame, + result, + target_specs, + dense_default_dataset=bool(args.dense_default_dataset), + seed=args.seed, + epochs=args.epochs, + learning_rate=args.learning_rate, + max_weight_ratio=args.max_weight_ratio, + l2_lambda=reconciliation_l2_lambda, + target_loss_cap=US_FISCAL_TARGET_LOSS_CAP, + reporter_source_ids=ssi_reporter_source_ids, + medicaid_enrollment_substitutions=medicaid_enrollment_substitutions, + maximum_microsim_batch_size=args.maximum_microsim_batch_size, + gate_congressional_district_targets=args.gate_congressional_district_targets, + progress_callback=( + telemetry.calibration_progress if telemetry is not None else None + ), + ) + export_frame = reconciliation.export_frame + result = reconciliation.calibration_result + registry = reconciliation.registry + compilation = dict(reconciliation.compilation) + ssi_take_up_diagnostics = dict(reconciliation.ssi_diagnostics) + medicaid_take_up_diagnostics = dict(reconciliation.medicaid_diagnostics) + health_input_gate = reconciliation.health_input_gate + other_health_insurance_gate = reconciliation.other_health_insurance_gate + if congressional_district_vintage_crosswalk_metadata is not None: + compilation = { + **compilation, + "congressional_district_vintage_crosswalk": ( + congressional_district_vintage_crosswalk_metadata + ), + } + default_dataset = { + **default_dataset, + "pre_ssi_reconciliation_final_loss": pre_reconciliation_final_loss, + "ssi_take_up_reconciliation_passes": reconciliation.passes, + "final_loss": float(result.final_loss), + } + if default_dataset["sparse"]: + default_dataset["refit_initial_loss"] = float(result.initial_loss) + default_dataset["refit_final_loss"] = float(result.final_loss) timing["calibration_seconds"] = time.perf_counter() - calibration_started timing["elapsed_through_calibration_seconds"] = time.perf_counter() - build_started if telemetry is not None: @@ -6613,11 +8123,6 @@ def main() -> None: if telemetry is not None: telemetry.stage("export_dataset", message="Writing PolicyEngine-US H5.") - export_frame = ( - _with_calibrated_weights(base_frame, result.weights) - if args.dense_default_dataset - else _with_l0_refit_weights(base_frame, result) - ) release_engine = PolicyEngineUSEngine() # populace#368: full eCPS input-column coverage as a HARD release gate. # Every input column the reference eCPS exports must be persisted by the @@ -6923,6 +8428,10 @@ def main() -> None: medicaid_take_up_diagnostics, release_dir / "us_medicaid_take_up.json", ) + write_us_ssi_take_up_diagnostics( + ssi_take_up_diagnostics, + release_dir / "us_ssi_take_up.json", + ) write_us_snap_state_take_up_diagnostics( snap_state_take_up_diagnostics, release_dir / "us_snap_state_take_up.json", @@ -6952,6 +8461,10 @@ def main() -> None: "us_medicaid_take_up", release_dir / "us_medicaid_take_up.json", ) + telemetry.attach_artifact( + "us_ssi_take_up", + release_dir / "us_ssi_take_up.json", + ) telemetry.attach_artifact( "us_snap_state_take_up", release_dir / "us_snap_state_take_up.json", diff --git a/tools/build_us_puf_support_base.py b/tools/build_us_puf_support_base.py index cd69fafd..cc9e50af 100644 --- a/tools/build_us_puf_support_base.py +++ b/tools/build_us_puf_support_base.py @@ -21,6 +21,7 @@ from populace.build.ledger_artifact import load_ledger_consumer_artifact from populace.build.source_manifest import SupportSpineSpec, load_support_spine_manifest from populace.build.us_runtime import ( + ASEC_2023_WEEKS_UNEMPLOYED_SOURCE_SHA256, BASE_ASEC_SUPPORT_CHANNEL, CONGRESSIONAL_DISTRICT_VINTAGE_CROSSWALK_SHA256_ATTR, CONGRESSIONAL_DISTRICT_VINTAGE_TARGET_ATTR, @@ -37,18 +38,66 @@ congressional_district_assignment_summary, congressional_district_distribution_from_ledger_facts, derive_us_cps_carried_inputs, + fetch_asec_2023_weeks_unemployed_source, + impute_us_housing_assistance_to_puf_support, impute_us_puf_tax_detail_support, + load_acs_2022_rent_donor, + load_asec_2023_weeks_unemployed_source, load_congressional_district_vintage_crosswalk, load_us_block_ladder, puf_tax_unit_donor_from_arrays, support_channel_column, translate_congressional_district_facts_to_current_vintage, + us_alimony_signal_gate, + us_capital_gain_details_signal_gate, + us_casualty_loss_signal_gate, + us_child_support_signal_gate, + us_childcare_signal_gate, + us_disability_benefits_signal_gate, + us_domestic_production_ald_signal_gate, + us_education_inputs_signal_gate, + us_educator_expense_signal_gate, + us_eligibility_inputs_signal_gate, + us_energy_subsidy_signal_gate, + us_farm_business_income_signal_gate, + us_form_4952_election_signal_gate, us_geography_ladder_assignment_summary, us_geography_ladder_gate, + us_housing_inputs_signal_gate, us_immigration_composition_summary, + us_medicare_take_up_signal_gate, + us_misc_itemized_signal_gate, + us_pregnancy_signal_gate, + us_prior_year_income_signal_gate, + us_prior_year_income_source_reconciliation_gate, + us_qbi_inputs_signal_gate, + us_relationship_inputs_signal_gate, + us_retirement_contributions_signal_gate, + us_retirement_distributions_signal_gate, + us_salt_refund_income_signal_gate, + us_weeks_unemployed_signal_gate, + us_wic_claim_signal_gate, + us_workers_compensation_signal_gate, with_household_congressional_districts, with_household_us_geography_ladder, + with_us_child_support_inputs, + with_us_childcare_inputs, + with_us_disability_benefits, + with_us_education_inputs, + with_us_eligibility_inputs, + with_us_energy_subsidy_input, + with_us_housing_inputs, with_us_immigration_inputs, + with_us_medicare_take_up_input, + with_us_pregnancy_inputs, + with_us_prior_year_income_inputs, + with_us_qbi_input_reconciliation, + with_us_relationship_inputs, + with_us_retirement_contribution_inputs, + with_us_retirement_distribution_inputs, + with_us_weeks_unemployed, + with_us_wic_claim_input, + with_us_workers_compensation, ) from populace.build.us_runtime.puf_support import PUF_TAX_DETAIL_DEFAULT_PREDICTORS from populace.frame import Frame, WeightKind, Weights @@ -85,6 +134,24 @@ def _parse_args(argv: list[str] | None = None) -> argparse.Namespace: ), ) parser.add_argument("--puf-h5", required=True, type=Path) + parser.add_argument( + "--asec-2023-weeks-unemployed-source", + type=Path, + help=( + "Optional local path to the SHA-pinned official 2023 ASEC CSV ZIP " + "used to restore income-year-2022 LKWEEKS. When omitted the " + "official Census archive is fetched and verified." + ), + ) + parser.add_argument( + "--acs-h5", + type=Path, + help=( + "SHA-pinned processed ACS 2022 ARRAYS artifact used by the " + "housing/rent stage. Required unless --base-h5 already carries " + "a green housing-input surface." + ), + ) parser.add_argument("--out", required=True, type=Path) parser.add_argument("--seed", default=0, type=int) parser.add_argument("--n-estimators", default=32, type=int) @@ -191,7 +258,138 @@ def main() -> None: summary_path = out_dir / _summary_filename(args.target_year) raw_base, base_source = _load_base_frame_from_args(args) + weeks_unemployed_source_path = ( + args.asec_2023_weeks_unemployed_source + if args.asec_2023_weeks_unemployed_source is not None + else fetch_asec_2023_weeks_unemployed_source() + ) + weeks_unemployed_source = load_asec_2023_weeks_unemployed_source( + weeks_unemployed_source_path + ) base = derive_us_cps_carried_inputs(raw_base) + base = with_us_prior_year_income_inputs( + base, + seed=args.seed, + time_period=args.target_year, + ) + base = with_us_relationship_inputs( + base, + seed=args.seed, + time_period=args.target_year, + ) + relationship_inputs_gate = us_relationship_inputs_signal_gate(base) + if not relationship_inputs_gate.passed: + raise SystemExit( + "Relationship-input signal gate failed:\n " + + "\n ".join(relationship_inputs_gate.failures) + ) + base = with_us_medicare_take_up_input( + base, + seed=args.seed, + time_period=args.target_year, + ) + medicare_take_up_gate = us_medicare_take_up_signal_gate(base) + if not medicare_take_up_gate.passed: + raise SystemExit( + "Medicare take-up input signal gate failed before support cloning:\n " + + "\n ".join(medicare_take_up_gate.failures) + ) + housing_inputs_gate = us_housing_inputs_signal_gate(base) + acs_rent_donor: pd.DataFrame | None = None + if not housing_inputs_gate.passed: + if args.acs_h5 is None: + raise SystemExit( + "Housing-input signal gate is not already green and --acs-h5 " + "was not provided; exact pre_subsidy_rent restoration requires " + "the pinned ACS 2022 donor." + ) + acs_rent_donor = load_acs_2022_rent_donor(args.acs_h5) + base = with_us_housing_inputs( + base, + seed=args.seed, + time_period=args.target_year, + acs_rent_donor=acs_rent_donor, + ) + housing_inputs_gate = us_housing_inputs_signal_gate(base) + if not housing_inputs_gate.passed: + raise SystemExit( + "Housing-input signal gate failed before support cloning:\n " + + "\n ".join(housing_inputs_gate.failures) + ) + base = with_us_eligibility_inputs( + base, + seed=args.seed, + time_period=args.target_year, + ) + eligibility_inputs_gate = us_eligibility_inputs_signal_gate(base) + if not eligibility_inputs_gate.passed: + raise SystemExit( + "Eligibility-input signal gate failed before support cloning:\n " + + "\n ".join(eligibility_inputs_gate.failures) + ) + base = with_us_pregnancy_inputs( + base, + seed=args.seed, + time_period=args.target_year, + ) + pregnancy_gate = us_pregnancy_signal_gate(base) + if not pregnancy_gate.passed: + raise SystemExit( + "Pregnancy signal gate failed before support cloning:\n " + + "\n ".join(pregnancy_gate.failures) + ) + base = with_us_wic_claim_input( + base, + seed=args.seed, + time_period=args.target_year, + ) + wic_claim_gate = us_wic_claim_signal_gate(base) + if not wic_claim_gate.passed: + raise SystemExit( + "WIC-claim signal gate failed before support cloning:\n " + + "\n ".join(wic_claim_gate.failures) + ) + base = with_us_child_support_inputs( + base, + seed=args.seed, + time_period=args.target_year, + ) + base = with_us_disability_benefits( + base, + seed=args.seed, + time_period=args.target_year, + ) + base = with_us_workers_compensation( + base, + seed=args.seed, + time_period=args.target_year, + ) + base = with_us_weeks_unemployed( + base, + seed=args.seed, + time_period=args.target_year, + asec_2023_source=weeks_unemployed_source, + ) + base = with_us_childcare_inputs( + base, + seed=args.seed, + time_period=args.target_year, + ) + base = with_us_energy_subsidy_input( + base, + seed=args.seed, + time_period=args.target_year, + ) + base = with_us_retirement_contribution_inputs( + base, + seed=args.seed, + time_period=args.target_year, + ) + base = with_us_retirement_distribution_inputs( + base, + seed=args.seed, + time_period=args.target_year, + ) base = with_us_immigration_inputs( base, seed=args.seed, @@ -206,6 +404,211 @@ def main() -> None: seed=args.seed, n_estimators=args.n_estimators, ) + imputed = with_us_qbi_input_reconciliation(imputed) + imputed = with_us_wic_claim_input( + imputed, + seed=args.seed, + time_period=args.target_year, + ) + wic_claim_gate = us_wic_claim_signal_gate(imputed) + if not wic_claim_gate.passed: + raise SystemExit( + "WIC-claim signal gate failed after support cloning:\n " + + "\n ".join(wic_claim_gate.failures) + ) + medicare_take_up_gate = us_medicare_take_up_signal_gate(imputed) + if not medicare_take_up_gate.passed: + raise SystemExit( + "Medicare take-up input signal gate failed after support cloning:\n " + + "\n ".join(medicare_take_up_gate.failures) + ) + imputed = impute_us_housing_assistance_to_puf_support( + imputed, + seed=args.seed, + ) + imputed = with_us_prior_year_income_inputs( + imputed, + seed=args.seed, + time_period=args.target_year, + ) + prior_year_income_gate = us_prior_year_income_signal_gate(imputed) + if not prior_year_income_gate.passed: + raise SystemExit( + "Prior-year-income signal gate failed:\n " + + "\n ".join(prior_year_income_gate.failures) + ) + prior_year_income_reconciliation_gate = ( + us_prior_year_income_source_reconciliation_gate(imputed) + ) + if not prior_year_income_reconciliation_gate.passed: + raise SystemExit( + "Prior-year-income source reconciliation failed:\n " + + "\n ".join(prior_year_income_reconciliation_gate.failures) + ) + housing_inputs_gate = us_housing_inputs_signal_gate(imputed) + if not housing_inputs_gate.passed: + raise SystemExit( + "Housing-input signal gate failed after PUF-support imputation:\n " + + "\n ".join(housing_inputs_gate.failures) + ) + qbi_inputs_gate = us_qbi_inputs_signal_gate(imputed) + if not qbi_inputs_gate.passed: + raise SystemExit( + "QBI-input signal gate failed:\n " + "\n ".join(qbi_inputs_gate.failures) + ) + farm_business_income_gate = us_farm_business_income_signal_gate(imputed) + if not farm_business_income_gate.passed: + raise SystemExit( + "Farm-business-income signal gate failed:\n " + + "\n ".join(farm_business_income_gate.failures) + ) + domestic_production_ald_gate = us_domestic_production_ald_signal_gate(imputed) + if not domestic_production_ald_gate.passed: + raise SystemExit( + "Domestic-production-ALD signal gate failed:\n " + + "\n ".join(domestic_production_ald_gate.failures) + ) + imputed = with_us_child_support_inputs( + imputed, + seed=args.seed, + time_period=args.target_year, + ) + child_support_gate = us_child_support_signal_gate(imputed) + if not child_support_gate.passed: + raise SystemExit( + "Child-support signal gate failed:\n " + + "\n ".join(child_support_gate.failures) + ) + imputed = with_us_disability_benefits( + imputed, + seed=args.seed, + time_period=args.target_year, + ) + disability_benefits_gate = us_disability_benefits_signal_gate(imputed) + if not disability_benefits_gate.passed: + raise SystemExit( + "Disability-benefits signal gate failed:\n " + + "\n ".join(disability_benefits_gate.failures) + ) + imputed = with_us_workers_compensation( + imputed, + seed=args.seed, + time_period=args.target_year, + ) + workers_compensation_gate = us_workers_compensation_signal_gate(imputed) + if not workers_compensation_gate.passed: + raise SystemExit( + "Workers-compensation signal gate failed:\n " + + "\n ".join(workers_compensation_gate.failures) + ) + imputed = with_us_weeks_unemployed( + imputed, + seed=args.seed, + time_period=args.target_year, + asec_2023_source=weeks_unemployed_source, + ) + weeks_unemployed_gate = us_weeks_unemployed_signal_gate(imputed) + if not weeks_unemployed_gate.passed: + raise SystemExit( + "Weeks-unemployed signal gate failed:\n " + + "\n ".join(weeks_unemployed_gate.failures) + ) + educator_expense_gate = us_educator_expense_signal_gate(imputed) + if not educator_expense_gate.passed: + raise SystemExit( + "Educator-expense signal gate failed:\n " + + "\n ".join(educator_expense_gate.failures) + ) + form_4952_election_gate = us_form_4952_election_signal_gate(imputed) + if not form_4952_election_gate.passed: + raise SystemExit( + "Form 4952 election signal gate failed:\n " + + "\n ".join(form_4952_election_gate.failures) + ) + salt_refund_income_gate = us_salt_refund_income_signal_gate(imputed) + if not salt_refund_income_gate.passed: + raise SystemExit( + "SALT-refund-income signal gate failed:\n " + + "\n ".join(salt_refund_income_gate.failures) + ) + capital_gain_details_gate = us_capital_gain_details_signal_gate(imputed) + if not capital_gain_details_gate.passed: + raise SystemExit( + "Capital-gain details signal gate failed:\n " + + "\n ".join(capital_gain_details_gate.failures) + ) + imputed = with_us_childcare_inputs( + imputed, + seed=args.seed, + time_period=args.target_year, + ) + childcare_gate = us_childcare_signal_gate(imputed) + if not childcare_gate.passed: + raise SystemExit( + "Childcare-input signal gate failed:\n " + + "\n ".join(childcare_gate.failures) + ) + imputed = with_us_energy_subsidy_input( + imputed, + seed=args.seed, + time_period=args.target_year, + ) + energy_subsidy_gate = us_energy_subsidy_signal_gate(imputed) + if not energy_subsidy_gate.passed: + raise SystemExit( + "Energy-subsidy signal gate failed:\n " + + "\n ".join(energy_subsidy_gate.failures) + ) + alimony_gate = us_alimony_signal_gate(imputed) + if not alimony_gate.passed: + raise SystemExit( + "Alimony-input signal gate failed:\n " + "\n ".join(alimony_gate.failures) + ) + casualty_loss_gate = us_casualty_loss_signal_gate(imputed) + if not casualty_loss_gate.passed: + raise SystemExit( + "Casualty-loss signal gate failed:\n " + + "\n ".join(casualty_loss_gate.failures) + ) + misc_itemized_gate = us_misc_itemized_signal_gate(imputed) + if not misc_itemized_gate.passed: + raise SystemExit( + "Miscellaneous-itemized signal gate failed:\n " + + "\n ".join(misc_itemized_gate.failures) + ) + imputed = with_us_retirement_contribution_inputs( + imputed, + seed=args.seed, + time_period=args.target_year, + ) + retirement_contributions_gate = us_retirement_contributions_signal_gate(imputed) + if not retirement_contributions_gate.passed: + raise SystemExit( + "Retirement-contribution signal gate failed:\n " + + "\n ".join(retirement_contributions_gate.failures) + ) + imputed = with_us_retirement_distribution_inputs( + imputed, + seed=args.seed, + time_period=args.target_year, + ) + retirement_distributions_gate = us_retirement_distributions_signal_gate(imputed) + if not retirement_distributions_gate.passed: + raise SystemExit( + "Retirement-distribution signal gate failed:\n " + + "\n ".join(retirement_distributions_gate.failures) + ) + imputed = with_us_education_inputs( + imputed, + seed=args.seed, + time_period=args.target_year, + ) + education_inputs_gate = us_education_inputs_signal_gate(imputed) + if not education_inputs_gate.passed: + raise SystemExit( + "Education-input signal gate failed:\n " + + "\n ".join(education_inputs_gate.failures) + ) congressional_district_assignment = {"applied": False} if args.assign_congressional_districts: ledger_facts = load_ledger_consumer_artifact(args.ledger_facts).facts @@ -313,6 +716,18 @@ def main() -> None: "base_sha256": _sha256(args.base_h5) if args.base_h5 is not None else None, "puf_h5": str(args.puf_h5.resolve()), "puf_sha256": _sha256(args.puf_h5), + "acs_h5": str(args.acs_h5.resolve()) if args.acs_h5 is not None else None, + "acs_sha256": _sha256(args.acs_h5) if args.acs_h5 is not None else None, + "acs_rent_donor_rows": ( + int(len(acs_rent_donor)) if acs_rent_donor is not None else None + ), + "weeks_unemployed_source": { + "path": str(Path(weeks_unemployed_source_path).resolve()), + "sha256": _sha256(weeks_unemployed_source_path), + "upstream_archive_sha256": (ASEC_2023_WEEKS_UNEMPLOYED_SOURCE_SHA256), + "rows": int(len(weeks_unemployed_source)), + "audit": dict(weeks_unemployed_source.attrs.get("source_audit", {})), + }, "output_h5": str(output_h5), "output_sha256": _sha256(output_h5), "seed": args.seed, @@ -327,6 +742,141 @@ def main() -> None: "puf_donor_rows": int(len(donor)), "puf_donor_columns": sorted(donor.columns.tolist()), "weights_audit": weights_audit, + "qbi_inputs_signal": { + "passed": qbi_inputs_gate.passed, + "failures": list(qbi_inputs_gate.failures), + "details": dict(qbi_inputs_gate.details), + }, + "farm_business_income_signal": { + "passed": farm_business_income_gate.passed, + "failures": list(farm_business_income_gate.failures), + "details": dict(farm_business_income_gate.details), + }, + "domestic_production_ald_signal": { + "passed": domestic_production_ald_gate.passed, + "failures": list(domestic_production_ald_gate.failures), + "details": dict(domestic_production_ald_gate.details), + }, + "child_support_signal": { + "passed": child_support_gate.passed, + "failures": list(child_support_gate.failures), + "details": dict(child_support_gate.details), + }, + "disability_benefits_signal": { + "passed": disability_benefits_gate.passed, + "failures": list(disability_benefits_gate.failures), + "details": dict(disability_benefits_gate.details), + }, + "workers_compensation_signal": { + "passed": workers_compensation_gate.passed, + "failures": list(workers_compensation_gate.failures), + "details": dict(workers_compensation_gate.details), + }, + "weeks_unemployed_signal": { + "passed": weeks_unemployed_gate.passed, + "failures": list(weeks_unemployed_gate.failures), + "details": dict(weeks_unemployed_gate.details), + }, + "eligibility_inputs_signal": { + "passed": eligibility_inputs_gate.passed, + "failures": list(eligibility_inputs_gate.failures), + "details": dict(eligibility_inputs_gate.details), + }, + "pregnancy_signal": { + "passed": pregnancy_gate.passed, + "failures": list(pregnancy_gate.failures), + "details": dict(pregnancy_gate.details), + }, + "wic_claim_signal": { + "passed": wic_claim_gate.passed, + "failures": list(wic_claim_gate.failures), + "details": dict(wic_claim_gate.details), + }, + "educator_expense_signal": { + "passed": educator_expense_gate.passed, + "failures": list(educator_expense_gate.failures), + "details": dict(educator_expense_gate.details), + }, + "form_4952_election_signal": { + "passed": form_4952_election_gate.passed, + "failures": list(form_4952_election_gate.failures), + "details": dict(form_4952_election_gate.details), + }, + "salt_refund_income_signal": { + "passed": salt_refund_income_gate.passed, + "failures": list(salt_refund_income_gate.failures), + "details": dict(salt_refund_income_gate.details), + }, + "capital_gain_details_signal": { + "passed": capital_gain_details_gate.passed, + "failures": list(capital_gain_details_gate.failures), + "details": dict(capital_gain_details_gate.details), + }, + "childcare_inputs_signal": { + "passed": childcare_gate.passed, + "failures": list(childcare_gate.failures), + "details": dict(childcare_gate.details), + }, + "energy_subsidy_signal": { + "passed": energy_subsidy_gate.passed, + "failures": list(energy_subsidy_gate.failures), + "details": dict(energy_subsidy_gate.details), + }, + "alimony_inputs_signal": { + "passed": alimony_gate.passed, + "failures": list(alimony_gate.failures), + "details": dict(alimony_gate.details), + }, + "casualty_loss_signal": { + "passed": casualty_loss_gate.passed, + "failures": list(casualty_loss_gate.failures), + "details": dict(casualty_loss_gate.details), + }, + "misc_itemized_signal": { + "passed": misc_itemized_gate.passed, + "failures": list(misc_itemized_gate.failures), + "details": dict(misc_itemized_gate.details), + }, + "education_inputs_signal": { + "passed": education_inputs_gate.passed, + "failures": list(education_inputs_gate.failures), + "details": dict(education_inputs_gate.details), + }, + "retirement_contributions_signal": { + "passed": retirement_contributions_gate.passed, + "failures": list(retirement_contributions_gate.failures), + "details": dict(retirement_contributions_gate.details), + }, + "retirement_distributions_signal": { + "passed": retirement_distributions_gate.passed, + "failures": list(retirement_distributions_gate.failures), + "details": dict(retirement_distributions_gate.details), + }, + "relationship_inputs_signal": { + "passed": relationship_inputs_gate.passed, + "failures": list(relationship_inputs_gate.failures), + "details": dict(relationship_inputs_gate.details), + }, + "medicare_take_up_input_signal": { + "passed": medicare_take_up_gate.passed, + "failures": list(medicare_take_up_gate.failures), + "details": dict(medicare_take_up_gate.details), + }, + "housing_inputs_signal": { + "passed": housing_inputs_gate.passed, + "failures": list(housing_inputs_gate.failures), + "details": dict(housing_inputs_gate.details), + }, + "prior_year_income_signal": { + "passed": prior_year_income_gate.passed, + "failures": list(prior_year_income_gate.failures), + "details": dict(prior_year_income_gate.details), + }, + "prior_year_income_source_reconciliation": { + "passed": prior_year_income_reconciliation_gate.passed, + "failures": list(prior_year_income_reconciliation_gate.failures), + "details": dict(prior_year_income_reconciliation_gate.details), + }, "congressional_district_assignment": congressional_district_assignment, "geography_ladder_assignment": geography_ladder_assignment, "channel_output_totals": _channel_output_totals(imputed), @@ -395,9 +945,7 @@ def impute_and_audit_us_puf_support( ) report = weights_audit_gate(fit_records) if not report.passed: - raise SystemExit( - "Weights audit failed:\n " + "\n ".join(report.failures) - ) + raise SystemExit("Weights audit failed:\n " + "\n ".join(report.failures)) weights_audit = { "passed": report.passed, "failures": list(report.failures), diff --git a/tools/build_us_release_input_coverage_manifest.py b/tools/build_us_release_input_coverage_manifest.py index 83bb1dfa..d28e617d 100644 --- a/tools/build_us_release_input_coverage_manifest.py +++ b/tools/build_us_release_input_coverage_manifest.py @@ -7,7 +7,8 @@ Derivation (fully from checked-in, sha-pinned facts — no transient artifact): - Required surface = every input-variable column the pinned reference eCPS - populates, i.e. the ``nonzero_shares`` keys of ``ecps_parity_reference.json`` + populates, i.e. the ``nonzero_shares`` keys of ``ecps_parity_reference.json``, + plus explicit later inputs needed by shipped reform probes (computed once from the sha-verified ``enhanced_cps_2024.h5``; an input the incumbent exports but leaves all-zero is not a coverage requirement, the same rule the parity gate uses). @@ -59,7 +60,167 @@ "bond_assets", ) -#: The single pinned reform-coverage probe: raising the SSI resource limit from +# The pinned reference H5 predates the retired pipeline's FLSA-premium export, +# OBBBA's distinct qualifying passenger-vehicle interest leaf, its five final +# desired retirement-contribution inputs, and its final SIPP-imputed SSI +# disability criterion. These later inputs are hard requirements because the +# shipped validation provisions must bind. +POST_REFERENCE_ECPS_REQUIRED_INPUTS = ( + "fsla_overtime_premium", + "qualified_passenger_vehicle_loan_interest", + "traditional_401k_contributions_desired", + "roth_401k_contributions_desired", + "traditional_ira_contributions_desired", + "roth_ira_contributions_desired", + "self_employed_pension_contributions_desired", + "meets_ssi_disability_criteria", +) + +QBI_INPUTS = ( + "estate_income_would_be_qualified", + "farm_operations_income_would_be_qualified", + "farm_rent_income_would_be_qualified", + "partnership_s_corp_income_would_be_qualified", + "rental_income_would_be_qualified", + "self_employment_income_would_be_qualified", + "sstb_self_employment_income_would_be_qualified", + "business_is_sstb", + "qualified_bdc_income", + "qualified_reit_and_ptp_income", + "sstb_self_employment_income_before_lsr", + "sstb_unadjusted_basis_qualified_property", + "sstb_w2_wages_from_qualified_business", + "unadjusted_basis_qualified_property", + "w2_wages_from_qualified_business", +) + +CHILD_SUPPORT_INPUTS = ( + "child_support_received", + "child_support_expense", +) + +DISABILITY_BENEFITS_INPUTS = ("disability_benefits",) + +WORKERS_COMPENSATION_INPUTS = ("workers_compensation",) + +WEEKS_UNEMPLOYED_INPUTS = ("weeks_unemployed",) + +WIC_CLAIM_INPUTS = ("would_claim_wic",) + +EDUCATOR_EXPENSE_INPUTS = ("educator_expense",) + +OTHER_HEALTH_INSURANCE_INPUTS = ("other_health_insurance_premiums",) + +PRIOR_YEAR_INCOME_INPUTS = ( + "self_employment_income_last_year", + "previous_year_income_available", +) + +FARM_BUSINESS_INCOME_INPUTS = ( + "farm_operations_income", + "farm_rent_income", +) + +SIPP_VEHICLE_INPUTS = ( + "household_vehicles_owned", + "household_vehicles_value", +) + +VOLUNTARY_FILING_INPUTS = ("would_file_taxes_voluntarily",) + +SSI_TAKE_UP_INPUTS = ("takes_up_ssi_if_eligible",) + +HEAD_START_INPUTS = ("takes_up_head_start_if_eligible",) + +SCF_NET_WORTH_INPUTS = ("net_worth",) + +FORM_4952_INPUTS = ("investment_income_elected_form_4952",) + +CAPITAL_GAIN_DETAIL_INPUTS = ( + "long_term_capital_gains_on_collectibles", + "unrecaptured_section_1250_gain", +) + +SALT_REFUND_INPUTS = ("salt_refund_income",) + +ENERGY_SUBSIDY_INPUTS = ("spm_unit_energy_subsidy",) + +RELATIONSHIP_INPUTS = ( + "is_household_head", + "is_separated", + "is_surviving_spouse", +) + +HOUSING_INPUTS = ( + "pre_subsidy_rent", + "receives_housing_assistance", + "takes_up_housing_assistance_if_eligible", + "spm_unit_tenure_type", + "tenure_type", +) + +RETIREMENT_DISTRIBUTION_INPUTS = ( + "taxable_401k_distributions", + "taxable_403b_distributions", + "tax_exempt_ira_distributions", + "taxable_ira_distributions", + "keogh_distributions", + "taxable_sep_distributions", +) + +# Reference-populated inputs whose primary-source restoration has shipped. +# They remain hard requirements even if a stale parity-gap entry is +# accidentally reintroduced later. +RESTORED_REFERENCE_ECPS_REQUIRED_INPUTS = ( + "alimony_expense", + "alimony_income", + "casualty_loss", + *CHILD_SUPPORT_INPUTS, + *DISABILITY_BENEFITS_INPUTS, + *WORKERS_COMPENSATION_INPUTS, + *WEEKS_UNEMPLOYED_INPUTS, + *WIC_CLAIM_INPUTS, + *EDUCATOR_EXPENSE_INPUTS, + *OTHER_HEALTH_INSURANCE_INPUTS, + *FARM_BUSINESS_INCOME_INPUTS, + *SCF_NET_WORTH_INPUTS, + *SIPP_VEHICLE_INPUTS, + *VOLUNTARY_FILING_INPUTS, + *HEAD_START_INPUTS, + *SSI_TAKE_UP_INPUTS, + *FORM_4952_INPUTS, + *CAPITAL_GAIN_DETAIL_INPUTS, + *SALT_REFUND_INPUTS, + *ENERGY_SUBSIDY_INPUTS, + *RELATIONSHIP_INPUTS, + *HOUSING_INPUTS, + *PRIOR_YEAR_INCOME_INPUTS, + *RETIREMENT_DISTRIBUTION_INPUTS, + "domestic_production_ald", + "household_weight", + "spm_unit_pre_subsidy_childcare_expenses", + "unreimbursed_business_employee_expenses", + *QBI_INPUTS, +) + +RETIREMENT_CONTRIBUTION_INPUTS = ( + "traditional_401k_contributions_desired", + "roth_401k_contributions_desired", + "traditional_ira_contributions_desired", + "roth_ira_contributions_desired", + "self_employed_pension_contributions_desired", +) + +AOTC_EDUCATION_INPUTS = ( + "qualified_tuition_expenses", + "is_pursuing_credential_for_american_opportunity_credit", + "attends_eligible_educational_institution_for_american_opportunity_credit", + "is_enrolled_at_least_half_time_for_american_opportunity_credit", + "has_american_opportunity_credit_1098_t_or_exception", + "has_american_opportunity_credit_institution_ein", +) + +#: Pinned reform-coverage probes. Raising the SSI resource limit from #: the 2024 statutory $2,000 individual / $3,000 couple to $10,000 / $20,000 is #: a pure relaxation that binds only through ``ssi_countable_resources``. Nonzero #: iff the asset inputs are restored. Dense-native reference magnitudes: +$1.6B @@ -67,6 +228,119 @@ #: a floor far below the plausible effect but far above simulation noise, so a #: structural $0 fails while a real (even conservative) score passes. REFORM_COVERAGE_PROBES = [ + { + "id": "prior_year_self_employment_neutralization", + "name": "Prior-year self-employment support neutralization", + "parameter_changes": {}, + "neutralized_variable": "self_employment_income_last_year", + "budget_measure": "tax_unit_earned_income_last_year", + "period": 2024, + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "binding_inputs": ["self_employment_income_last_year"], + "min_abs_effect": 1_000_000_000.0, + "reason": ( + "PolicyEngine-US adds self_employment_income_last_year to " + "earned_income_last_year and then aggregates it over tax-unit " + "nondependents. The Wyden-Smith ACTC lookback is the downstream " + "policy consumer, but its parameter also reads formula-owned " + "prior-year wages, so neutralizing this one leaf is the unique " + "coverage probe. The SHA-locked strict equal-share 2022-2024 ASEC " + "pool carries $357.24 billion of weighted net source amount; " + "without the adjacent-year carry the neutralization is a " + "structural zero. The distinct " + "previous_year_income_available flag has no formula consumer in " + "PolicyEngine-US 1.764.6 and remains protected by the hard " + "non-default column gate." + ), + "issue": "PolicyEngine/populace#38", + }, + { + "id": "keogh_distribution_neutralization", + "name": "Keogh distribution neutralization", + "parameter_changes": {}, + "neutralized_variable": "keogh_distributions", + "budget_measure": "income_tax", + "period": 2024, + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "binding_inputs": ["keogh_distributions"], + "min_abs_effect": 1_000_000.0, + "reason": ( + "PolicyEngine-US includes keogh_distributions directly in taxable " + "retirement distributions and federal gross income. Neutralizing " + "only the measured code-5 ASEC leaf lowers federal income tax. On " + "the 865,046-person staged Build-J support, the locked ASEC source " + "carries $148.97 million of weighted Keogh distributions and the " + "baseline-minus-neutralized income-tax effect is +$24.47 million. " + "Without the restored DST_SC*/DST_VAL* mapping the effect is a " + "structural zero." + ), + "issue": "PolicyEngine/populace#38", + }, + { + "id": "household_head_childcare_cap_neutralization", + "name": "Household-head childcare earned-cap neutralization", + "parameter_changes": {}, + "neutralized_variable": "is_household_head", + "budget_measure": "spm_unit_capped_work_childcare_expenses", + "period": 2024, + "effect_direction": "baseline_minus_reform", + "expected_sign": "negative", + "binding_inputs": ["is_household_head"], + "min_abs_effect": 1_000_000.0, + "reason": ( + "PolicyEngine-US uses is_household_head to identify whose earnings " + "cap SPM work-childcare expenses when an SPM unit contains multiple " + "tax units. Neutralizing only the measured head flag falls back to " + "tax-unit roles and changes the cap. On the 865,046-person staged " + "Build-J artifact, baseline-minus-neutralized capped expenses are " + "-$265.50 million. Without the restored P_SEQ input the " + "neutralization is a structural zero." + ), + "issue": "PolicyEngine/populace#38", + }, + { + "id": "spm_unit_energy_subsidy_neutralization", + "name": "SPM energy-subsidy neutralization", + "parameter_changes": {}, + "neutralized_variable": "spm_unit_energy_subsidy", + "budget_measure": "spm_unit_benefits", + "period": 2024, + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "binding_inputs": ["spm_unit_energy_subsidy"], + "min_abs_effect": 100_000_000.0, + "reason": ( + "PolicyEngine-US adds this measured LIHEAP resource dollar-for-dollar " + "to spm_unit_benefits. Neutralizing the leaf must therefore lower " + "benefits by its weighted source mass; without the restored " + "SPM_ENGVAL carry, the effect is a structural zero. No OBBBA " + "provision consumes this SPM resource, so the direct neutralization " + "is the uniquely isolating policy-engine probe." + ), + "issue": "PolicyEngine/populace#32", + }, + { + "id": "medicare_take_up_neutralization", + "name": "Measured Medicare enrollment neutralization", + "parameter_changes": {}, + "neutralized_variable": "takes_up_medicare_if_eligible", + "budget_measure": "medicare_cost", + "period": 2024, + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "binding_inputs": ["takes_up_medicare_if_eligible"], + "min_abs_effect": 1_000_000_000.0, + "reason": ( + "PolicyEngine-US computes medicare_enrolled from the measured " + "takes_up_medicare_if_eligible leaf and modeled eligibility, then " + "gates Medicare costs on enrollment. Neutralizing only the " + "restored MCARE == 1 leaf must reduce aggregate Medicare cost; " + "without the measured carry the probe is a structural zero." + ), + "issue": "PolicyEngine/populace#312", + }, { "id": "ssi_asset_limit_10k_20k", "name": "SSI asset limits raised to $10k individual / $20k couple", @@ -79,7 +353,9 @@ }, }, "budget_measure": "ssi", + "period": 2024, "effect_direction": "reform_minus_baseline", + "expected_sign": "positive", "binding_inputs": list(SSI_COUNTABLE_RESOURCE_ASSETS), "min_abs_effect": 1_000_000_000.0, "reason": ( @@ -93,7 +369,839 @@ "limit (PolicyEngine/populace#356)." ), "issue": "PolicyEngine/populace#356", - } + }, + { + "id": "ssi_disability_criteria_neutralization", + "name": "SSI disability-criteria neutralization", + "parameter_changes": {}, + "neutralized_variable": "meets_ssi_disability_criteria", + "budget_measure": "ssi", + "period": 2024, + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "binding_inputs": ["meets_ssi_disability_criteria"], + "min_abs_effect": 100_000_000.0, + "reason": ( + "PolicyEngine-US requires the person-level disability criterion " + "for non-aged SSI eligibility. Neutralizing only the restored " + "SIPP-imputed criterion must therefore remove SSI from otherwise " + "eligible disabled or blind people, so baseline-minus-neutralized " + "SSI is positive. If the post-reference exported input is absent " + "or degenerate, this isolated eligibility channel scores exactly " + "$0. A deterministic 6,000-source-household Build-J smoke with " + "the pinned SIPP and SCF donors scored +$586.393 million " + "baseline-minus-neutralized SSI; the $100 million floor retains " + "ample sampling margin while rejecting a materially weakened " + "criterion channel." + ), + "issue": "PolicyEngine/populace#312", + }, + { + "id": "ssi_take_up_neutralization", + "name": "SSI take-up neutralization", + "parameter_changes": {}, + "neutralized_variable": "takes_up_ssi_if_eligible", + "budget_measure": "ssi", + "period": 2024, + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "binding_inputs": ["takes_up_ssi_if_eligible"], + "min_abs_effect": 10_000_000_000.0, + "reason": ( + "PolicyEngine-US gates SSI benefits on the restored person-level " + "take-up leaf after eligibility. Neutralizing only that leaf must " + "therefore remove SSI from source-reported and SSA-count-calibrated " + "recipients, so baseline-minus-neutralized SSI is positive. If the " + "restored exported input is absent, all false, or not persisted, " + "this isolated channel scores exactly $0. A production-ingredient " + "sparse smoke (staged artifact sha256 c5939dad81153da51b2cc57081" + "ddb3e729700366144868df742b3ad86eafcd7c; restored artifact sha256 " + "9269360d3409fdc15c90c43dda394ada6c91eff5bb64c12ccd9def7d670dd077) " + "measured +$57,114,569,526.38 of baseline-minus-neutralized 2024 " + "SSI. The $10 billion floor retains over 5.7x observed margin while " + "rejecting a materially degenerate persisted flag." + ), + "issue": "PolicyEngine/populace#312", + }, + { + "id": "head_start_take_up_neutralization", + "name": "Measured Head Start take-up neutralization", + "parameter_changes": {}, + "neutralized_variable": "takes_up_head_start_if_eligible", + "budget_measure": "head_start", + "period": 2024, + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "binding_inputs": list(HEAD_START_INPUTS), + "min_abs_effect": 100_000_000.0, + "reason": ( + "PolicyEngine-US gates Head Start benefits on the person-level " + "take-up leaf after modeled age, income, and categorical " + "eligibility. Neutralizing only the restored SIPP response model " + "must therefore remove Head Start from measured-proxy recipients, " + "so baseline-minus-neutralized Head Start is positive. If the " + "restored export is absent, all false, or not persisted, this " + "isolated channel scores exactly $0. A production-ingredient sparse " + "smoke (staged artifact sha256 " + "67ad74b9ad9222ed342a0279dfc8175e872966fa59f86aeecb7fad52021ba500) " + "measured +$4,575,181,976.69 of baseline-minus-neutralized 2024 " + "Head Start. The $100 million floor retains over 45x observed " + "margin while remaining far above numerical noise." + ), + "issue": "PolicyEngine/populace#312", + }, + { + "id": "aotc_abolition", + "name": "American Opportunity Tax Credit abolition", + "parameter_changes": { + "gov.irs.credits.education.american_opportunity_credit.abolition": { + "2024-01-01.2100-12-31": True + } + }, + "budget_measure": "american_opportunity_credit", + "period": 2024, + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "binding_inputs": list(AOTC_EDUCATION_INPUTS), + "min_abs_effect": 100_000_000.0, + "reason": ( + "Abolishing the American Opportunity Tax Credit sets the credit " + "to zero, so baseline-minus-reform AOTC must be positive. With " + "qualified tuition or any of the five affirmative AOTC factual " + "inputs absent or degenerate, the baseline credit is a structural " + "zero and the abolition scores exactly $0." + ), + "issue": "PolicyEngine/populace#253", + }, + { + "id": "savers_credit_abolition", + "name": "Retirement Saver's Credit abolition", + "parameter_changes": { + "gov.irs.credits.retirement_saving.contributions_cap": { + "2024-01-01.2100-12-31": 0 + } + }, + "budget_measure": "savers_credit", + "period": 2024, + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "binding_inputs": list(RETIREMENT_CONTRIBUTION_INPUTS), + "min_abs_effect": 100_000_000.0, + "reason": ( + "Setting the Saver's Credit contribution cap to zero abolishes " + "the credit, so baseline-minus-reform Saver's Credit must be " + "positive. PolicyEngine-US builds qualified contributions from " + "the realized forms of all five desired retirement-contribution " + "inputs. If the desired-input family is absent or degenerate, " + "the baseline credit is a structural zero and abolition scores $0." + ), + "issue": "PolicyEngine/populace#278", + }, + { + "id": "qbi_reit_ptp_rate_abolition", + "name": "Section 199A qualified REIT/PTP component abolition", + "parameter_changes": { + "gov.irs.deductions.qbi.max.reit_ptp_rate": {"2024-01-01.2100-12-31": 0} + }, + "budget_measure": "qualified_business_income_deduction", + "period": 2024, + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "binding_inputs": ["qualified_reit_and_ptp_income"], + "min_abs_effect": 1_000_000.0, + "reason": ( + "Setting only the qualified REIT/PTP component rate to zero " + "removes that component from the Section 199A deduction, so " + "baseline-minus-reform QBID must be positive. Without populated " + "qualified_reit_and_ptp_income the change is a structural zero." + ), + "issue": "PolicyEngine/populace#298", + }, + { + "id": "qbi_wage_property_guardrails_zeroed", + "name": "Section 199A W-2 wage and UBIA guardrails zeroed", + "parameter_changes": { + "gov.irs.deductions.qbi.max.w2_wages.rate": {"2024-01-01.2100-12-31": 0}, + "gov.irs.deductions.qbi.max.w2_wages.alt_rate": { + "2024-01-01.2100-12-31": 0 + }, + "gov.irs.deductions.qbi.max.business_property.rate": { + "2024-01-01.2100-12-31": 0 + }, + }, + "budget_measure": "qualified_business_income_deduction", + "period": 2024, + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "binding_inputs": [ + "w2_wages_from_qualified_business", + "unadjusted_basis_qualified_property", + ], + "min_abs_effect": 1_000_000.0, + "reason": ( + "Zeroing all W-2 wage and UBIA cap rates tightens the Section " + "199A deduction for high-income qualified businesses, so " + "baseline-minus-reform QBID must be positive. If the total W-2 " + "and UBIA inputs are absent, both baseline and reform guardrails " + "are zero and the change is a structural zero. The archived " + "all-or-nothing SSTB routing leaves its SSTB-allocable copies to " + "the hard signal gate rather than overclaiming reform coverage." + ), + "issue": "PolicyEngine/populace#298", + }, + { + "id": "qbi_farm_operations_income_exclusion", + "name": "Exclude farm-operations income from Section 199A QBI", + "parameter_changes": { + "gov.irs.deductions.qbi.income_definition": { + "2026-01-01.2026-12-31": [ + "self_employment_income", + "partnership_s_corp_income", + "farm_rent_income", + "rental_income", + "estate_income", + ] + } + }, + "budget_measure": "qualified_business_income_deduction", + "period": 2026, + "effect_direction": "baseline_minus_reform", + "expected_sign": "negative", + "binding_inputs": ["farm_operations_income"], + "min_abs_effect": 1_000_000.0, + "reason": ( + "Removing only farm_operations_income from the 2026 Section 199A " + "income definition isolates the restored signed Schedule F leaf. " + "The real staged candidate is loss-heavy, so excluding it raises " + "QBID and baseline-minus-reform is negative (-$4.16M). Without " + "farm_operations_income the reform is a structural zero." + ), + "issue": "PolicyEngine/populace#298", + }, + { + "id": "qbi_farm_rent_income_exclusion", + "name": "Exclude farm-rent income from Section 199A QBI", + "parameter_changes": { + "gov.irs.deductions.qbi.income_definition": { + "2026-01-01.2026-12-31": [ + "self_employment_income", + "partnership_s_corp_income", + "farm_operations_income", + "rental_income", + "estate_income", + ] + } + }, + "budget_measure": "qualified_business_income_deduction", + "period": 2026, + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "binding_inputs": ["farm_rent_income"], + "min_abs_effect": 1_000_000.0, + "reason": ( + "Removing only farm_rent_income from the 2026 Section 199A income " + "definition isolates the restored signed E27200 leaf. The real " + "staged candidate produces +$9.14M baseline-minus-reform QBID. " + "Without farm_rent_income the reform is a structural zero." + ), + "issue": "PolicyEngine/populace#298", + }, + { + "id": "domestic_production_ald_reactivation", + "name": "Former Section 199 domestic-production deduction reactivation", + "parameter_changes": { + "gov.irs.ald.deductions": { + "2024-01-01.2024-12-31": [ + "loss_ald", + "self_employment_tax_ald", + "student_loan_interest_ald", + "early_withdrawal_penalty", + "alimony_expense_ald", + "educator_expense", + "health_savings_account_ald", + "self_employed_health_insurance_ald", + "self_employed_pension_contribution_ald", + "traditional_ira_contributions", + "qualified_adoption_assistance_expense", + "us_bonds_for_higher_ed", + "specified_possession_income", + "puerto_rico_income", + "domestic_production_ald", + ] + } + }, + "budget_measure": "income_tax", + "period": 2024, + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "binding_inputs": ["domestic_production_ald"], + "min_abs_effect": 1_000_000.0, + "reason": ( + "PolicyEngine-US 1.764.6 excludes the former Section 199 deduction " + "from current-law above-the-line deductions. This probe preserves " + "the exact 2024 list and adds only domestic_production_ald, so " + "baseline-minus-reform income tax must be positive. Without the " + "restored E03240 input, reactivation is a structural zero." + ), + "issue": "PolicyEngine/populace#298", + }, + { + "id": "form_4952_election_neutralization", + "name": "Form 4952 elected investment income neutralization", + "parameter_changes": {}, + "neutralized_variable": "investment_income_elected_form_4952", + "budget_measure": "income_tax", + "period": 2024, + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "binding_inputs": ["investment_income_elected_form_4952"], + "min_abs_effect": 1_000_000.0, + "reason": ( + "PolicyEngine-US subtracts the tax-unit sum of " + "investment_income_elected_form_4952 from net capital gain. " + "Neutralizing only that leaf increases preferential net capital " + "gain and lowers income tax, so baseline-minus-reform income tax " + "must be positive. Without the restored E58990 input the " + "neutralization is a structural zero." + ), + "issue": "PolicyEngine/populace#274", + }, + { + "id": "salt_refund_income_neutralization", + "name": "State and local tax refund income neutralization", + "parameter_changes": {}, + "neutralized_variable": "salt_refund_income", + "budget_measure": "state_income_tax", + "period": 2024, + "effect_direction": "baseline_minus_reform", + "expected_sign": "negative", + "binding_inputs": ["salt_refund_income"], + "min_abs_effect": 1_000_000.0, + "reason": ( + "PolicyEngine-US 1.764.6 includes salt_refund_income in the " + "South Carolina, Idaho, and West Virginia subtraction lists. " + "Neutralizing only that leaf removes the state subtraction and " + "raises state income tax, so baseline-minus-reform state income " + "tax must be negative. Without the restored E00700 input the " + "neutralization is a structural zero." + ), + "issue": "PolicyEngine/populace#38", + }, + { + "id": "collectibles_gain_neutralization", + "name": "Long-term collectibles gain neutralization", + "parameter_changes": {}, + "neutralized_variable": "long_term_capital_gains_on_collectibles", + "budget_measure": "income_tax", + "period": 2024, + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "binding_inputs": ["long_term_capital_gains_on_collectibles"], + "min_abs_effect": 1_000_000.0, + "reason": ( + "PolicyEngine-US includes collectibles in capital_gains_28_percent_" + "rate_gain. Neutralizing only the E24518 memo leaf reclassifies " + "those gains from the special 28-percent bucket to the ordinary " + "preferential capital-gain schedule, so baseline-minus-reform " + "income tax must be positive. Without the restored leaf the " + "neutralization is a structural zero." + ), + "issue": "PolicyEngine/populace#274", + }, + { + "id": "unrecaptured_section_1250_gain_neutralization", + "name": "Unrecaptured section 1250 gain neutralization", + "parameter_changes": {}, + "neutralized_variable": "unrecaptured_section_1250_gain", + "budget_measure": "income_tax", + "period": 2024, + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "binding_inputs": ["unrecaptured_section_1250_gain"], + "min_abs_effect": 1_000_000.0, + "reason": ( + "PolicyEngine-US taxes the E24515 memo leaf at the special " + "unrecaptured-section-1250 rate. Neutralizing only that leaf " + "reclassifies the same net gain onto the ordinary preferential " + "capital-gain schedule, so baseline-minus-reform income tax must " + "be positive. Without the restored leaf the neutralization is a " + "structural zero." + ), + "issue": "PolicyEngine/populace#274", + }, + { + "id": "child_support_received_snap_exclusion", + "name": "Exclude child-support receipts from SNAP unearned income", + "parameter_changes": { + "gov.usda.snap.income.sources.unearned": { + "2024-01-01.2024-12-31": [ + "ssi", + "tanf", + "general_assistance", + "pension_income", + "veterans_benefits", + "unemployment_compensation", + "disability_benefits", + "workers_compensation", + "social_security", + "retirement_distributions", + "rental_income", + "alimony_income", + "financial_assistance", + "survivor_benefits", + "dividend_income", + "interest_income", + "miscellaneous_income", + ] + } + }, + "budget_measure": "snap", + "period": 2024, + "effect_direction": "reform_minus_baseline", + "expected_sign": "positive", + "binding_inputs": ["child_support_received"], + "min_abs_effect": 1_000_000.0, + "reason": ( + "Removing only child_support_received from SNAP unearned-income " + "sources lowers countable income and must increase SNAP for some " + "recipients. Without the measured/QRF child-support receipt leaf, " + "the source-list reform is a structural zero." + ), + "issue": "PolicyEngine/populace#32", + }, + { + "id": "child_support_expense_snap_deduction_abolition", + "name": "Abolish the SNAP child-support expense deduction", + "parameter_changes": { + "gov.usda.snap.income.deductions.allowed": { + "2024-01-01.2024-12-31": [ + "snap_standard_deduction", + "snap_earned_income_deduction", + "snap_dependent_care_deduction", + "snap_excess_medical_expense_deduction", + "snap_excess_shelter_expense_deduction", + ] + } + }, + "budget_measure": "snap", + "period": 2024, + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "binding_inputs": ["child_support_expense"], + "min_abs_effect": 1_000_000.0, + "reason": ( + "Removing only snap_child_support_deduction raises countable net " + "income and must reduce SNAP in states that take the expense as a " + "net-income deduction. Without the measured/QRF positive expense " + "leaf, abolition is a structural zero." + ), + "issue": "PolicyEngine/populace#32", + }, + { + "id": "disability_benefits_snap_exclusion", + "name": "Exclude disability benefits from SNAP unearned income", + "parameter_changes": { + "gov.usda.snap.income.sources.unearned": { + "2024-01-01.2024-12-31": [ + "ssi", + "tanf", + "general_assistance", + "pension_income", + "veterans_benefits", + "unemployment_compensation", + "workers_compensation", + "social_security", + "retirement_distributions", + "rental_income", + "child_support_received", + "alimony_income", + "financial_assistance", + "survivor_benefits", + "dividend_income", + "interest_income", + "miscellaneous_income", + ] + } + }, + "budget_measure": "snap", + "period": 2024, + "effect_direction": "reform_minus_baseline", + "expected_sign": "positive", + "binding_inputs": ["disability_benefits"], + "min_abs_effect": 1_000_000.0, + "reason": ( + "Removing only disability_benefits from SNAP unearned-income " + "sources lowers countable income and must increase SNAP for some " + "recipients. Without the measured/QRF non-workers-compensation " + "benefit leaf, the source-list reform is a structural zero." + ), + "issue": "PolicyEngine/populace#38", + }, + { + "id": "workers_compensation_snap_exclusion", + "name": "Exclude workers' compensation from SNAP unearned income", + "parameter_changes": { + "gov.usda.snap.income.sources.unearned": { + "2024-01-01.2024-12-31": [ + "ssi", + "tanf", + "general_assistance", + "pension_income", + "veterans_benefits", + "unemployment_compensation", + "disability_benefits", + "social_security", + "retirement_distributions", + "rental_income", + "child_support_received", + "alimony_income", + "financial_assistance", + "survivor_benefits", + "dividend_income", + "interest_income", + "miscellaneous_income", + ] + } + }, + "budget_measure": "snap", + "period": 2024, + "effect_direction": "reform_minus_baseline", + "expected_sign": "positive", + "binding_inputs": ["workers_compensation"], + "min_abs_effect": 10_000_000.0, + "reason": ( + "Removing only workers_compensation from SNAP unearned-income " + "sources lowers countable income and must increase SNAP for some " + "recipients. A production-ingredient 30,000-household smoke scored " + "+$28.26M reform-minus-baseline; without the measured WC_VAL carry " + "and PUF-half QRF, the source-list reform is a structural zero." + ), + "issue": "PolicyEngine/populace#32", + }, + { + "id": "wic_claim_neutralization", + "name": "WIC claim neutralization", + "parameter_changes": {}, + "neutralized_variable": "would_claim_wic", + "budget_measure": "wic", + "period": 2024, + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "binding_inputs": ["would_claim_wic"], + "min_abs_effect": 25_000_000.0, + "reason": ( + "PolicyEngine-US multiplies each eligible person's monthly WIC " + "food package by would_claim_wic. A 6,000-household production-" + "ingredient smoke with the FNS category-rate stage scored " + "+$57.19M baseline-minus-neutralized; without the restored claim " + "surface the probe is a structural zero." + ), + "issue": "PolicyEngine/populace#312", + }, + { + "id": "educator_expense_ald_abolition", + "name": "Abolish the educator-expense above-the-line deduction", + "parameter_changes": { + "gov.irs.ald.deductions": { + "2024-01-01.2024-12-31": [ + "loss_ald", + "self_employment_tax_ald", + "student_loan_interest_ald", + "early_withdrawal_penalty", + "alimony_expense_ald", + "health_savings_account_ald", + "self_employed_health_insurance_ald", + "self_employed_pension_contribution_ald", + "traditional_ira_contributions", + "qualified_adoption_assistance_expense", + "us_bonds_for_higher_ed", + "specified_possession_income", + "puerto_rico_income", + ] + } + }, + "budget_measure": "income_tax", + "period": 2024, + "effect_direction": "reform_minus_baseline", + "expected_sign": "positive", + "binding_inputs": ["educator_expense"], + "min_abs_effect": 1_000_000.0, + "reason": ( + "Removing only educator_expense from the above-the-line deduction " + "source list raises taxable income and must increase income tax for " + "some filers. Without the restored PUF E03220 leaf, abolition is a " + "structural zero." + ), + "issue": "PolicyEngine/populace#32", + }, + { + "id": "alimony_expense_ald_abolition", + "name": "Alimony expense above-the-line deduction abolition", + "parameter_changes": { + "gov.irs.ald.alimony_expense.divorce_year_threshold[0].amount": { + "2024-01-01.2100-12-31": False + } + }, + "budget_measure": "income_tax", + "period": 2024, + "effect_direction": "baseline_minus_reform", + "expected_sign": "negative", + "binding_inputs": ["alimony_expense"], + "min_abs_effect": 1_000_000.0, + "reason": ( + "The retired export has no nondefault divorce_year input, so " + "PolicyEngine-US applies its default year 0 through the first " + "eligibility bracket. Setting that bracket's amount to false " + "abolishes the alimony-expense above-the-line deduction on the " + "release, so baseline-minus-reform income tax must be negative. " + "With alimony_expense absent or degenerate, the abolition scores " + "exactly $0." + ), + "issue": "PolicyEngine/populace#38", + }, + { + "id": "obbba_casualty_loss_limit", + "name": "OBBBA casualty-loss deduction reactivation", + "parameter_changes": { + "gov.irs.deductions.itemized.casualty.active": { + "2026-01-01.2026-12-31": True + } + }, + "budget_measure": "income_tax", + "period": 2026, + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "binding_inputs": ["casualty_loss"], + "min_abs_effect": 1_000_000.0, + "reason": ( + "Reactivating the casualty-loss deduction lowers income tax only " + "for tax units with casualty_loss above the statutory AGI floor, " + "so baseline-minus-reform income tax must be positive. With the " + "casualty-loss input absent or degenerate, the reactivation scores " + "exactly $0." + ), + "issue": "PolicyEngine/populace#32", + }, + { + "id": "obbba_misc_itemized_deductions", + "name": "OBBBA miscellaneous-itemized deduction reactivation", + "parameter_changes": { + "gov.irs.deductions.itemized.misc.applies": {"2026-01-01.2026-12-31": True} + }, + "budget_measure": "income_tax", + "period": 2026, + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "binding_inputs": ["unreimbursed_business_employee_expenses"], + "min_abs_effect": 100_000_000.0, + "reason": ( + "Reactivating the miscellaneous itemized deduction lowers income " + "tax only for tax units with qualifying expenses above its AGI " + "floor, so baseline-minus-reform income tax must be positive. " + "Without unreimbursed_business_employee_expenses, the retired " + "pipeline's only populated miscellaneous-expense input, the " + "reactivation is a structural zero on the export." + ), + "issue": "PolicyEngine/populace#32", + }, + { + "id": "obbba_cdcc", + "name": "OBBBA Child and Dependent Care Credit reversion", + "parameter_changes": { + "gov.irs.credits.cdcc.phase_out.max": {"2026-01-01.2026-12-31": 0.35}, + "gov.irs.credits.cdcc.phase_out.min": {"2026-01-01.2026-12-31": 0.2}, + "gov.irs.credits.cdcc.phase_out.amended_structure.applies": { + "2026-01-01.2026-12-31": False + }, + }, + "budget_measure": "income_tax", + "period": 2026, + "effect_direction": "baseline_minus_reform", + "expected_sign": "negative", + "binding_inputs": ["spm_unit_pre_subsidy_childcare_expenses"], + "min_abs_effect": 1_000_000.0, + "reason": ( + "Reverting the OBBBA CDCC enhancement raises income tax, so " + "baseline-minus-reform income tax must be negative. With measured " + "pre-subsidy childcare expenses absent or degenerate, no filer has " + "qualifying care expenses and the reversion scores exactly $0." + ), + "issue": "PolicyEngine/populace#278", + }, + { + "id": "obbba_no_tax_on_tips", + "name": "OBBBA no-tax-on-tips deduction", + "parameter_changes": { + "gov.irs.deductions.tip_income.cap": {"2026-01-01.2026-12-31": 0} + }, + "budget_measure": "income_tax", + "period": 2026, + "effect_direction": "baseline_minus_reform", + "expected_sign": "negative", + "binding_inputs": [ + "tip_income", + "treasury_tipped_occupation_code", + ], + "min_abs_effect": 100_000_000.0, + "reason": ( + "Setting the OBBBA tip-deduction cap to zero removes the deduction, " + "so baseline-minus-reform income tax must be negative in 2026. " + "With tip_income or the Treasury tipped-occupation code absent, " + "qualified tip income is zero and the repeal scores exactly $0." + ), + "issue": "PolicyEngine/populace#38", + }, + { + "id": "obbba_no_tax_on_overtime", + "name": "OBBBA no-tax-on-overtime deduction", + "parameter_changes": { + "gov.irs.deductions.overtime_income.cap.JOINT": { + "2026-01-01.2026-12-31": 0 + }, + "gov.irs.deductions.overtime_income.cap.SINGLE": { + "2026-01-01.2026-12-31": 0 + }, + "gov.irs.deductions.overtime_income.cap.HEAD_OF_HOUSEHOLD": { + "2026-01-01.2026-12-31": 0 + }, + "gov.irs.deductions.overtime_income.cap.SURVIVING_SPOUSE": { + "2026-01-01.2026-12-31": 0 + }, + "gov.irs.deductions.overtime_income.cap.SEPARATE": { + "2026-01-01.2026-12-31": 0 + }, + }, + "budget_measure": "income_tax", + "period": 2026, + "effect_direction": "baseline_minus_reform", + "expected_sign": "negative", + "binding_inputs": ["fsla_overtime_premium"], + "min_abs_effect": 100_000_000.0, + "reason": ( + "Setting every OBBBA overtime-deduction cap to zero removes the " + "deduction, so reform income tax rises and baseline-minus-reform " + "must be negative in 2026. With fsla_overtime_premium absent or " + "degenerate, qualified overtime is zero and the repeal scores $0." + ), + "issue": "PolicyEngine/populace#242", + }, + { + "id": "obbba_auto_loan_interest", + "name": "OBBBA no-tax-on-auto-loan-interest deduction", + "parameter_changes": { + "gov.irs.deductions.auto_loan_interest.cap": {"2026-01-01.2026-12-31": 0} + }, + "budget_measure": "income_tax", + "period": 2026, + "effect_direction": "baseline_minus_reform", + "expected_sign": "negative", + "binding_inputs": ["qualified_passenger_vehicle_loan_interest"], + "min_abs_effect": 100_000_000.0, + "reason": ( + "Setting the OBBBA auto-loan-interest deduction cap to zero " + "removes the deduction, so reform income tax rises and " + "baseline-minus-reform must be negative in 2026. With qualified " + "passenger-vehicle loan interest absent or degenerate, the repeal " + "scores exactly $0." + ), + "issue": "PolicyEngine/populace#252", + }, + { + "id": "tx_snap_additional_vehicle_exemption_abolition", + "name": "Texas SNAP additional-vehicle exemption abolition", + "parameter_changes": { + "gov.hhs.tanf.non_cash.tx_additional_vehicle_exemption": { + "2026-01-01.2100-12-31": 0 + } + }, + "budget_measure": "snap", + "period": 2026, + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "binding_inputs": list(SIPP_VEHICLE_INPUTS), + "min_abs_effect": 1_000_000.0, + "reason": ( + "Setting the Texas TANF non-cash additional-vehicle exemption " + "to zero tightens the asset test used by Texas SNAP categorical " + "eligibility, so baseline-minus-reform SNAP must be positive. " + "PolicyEngine-US computes the exemption from both household " + "vehicle count and value; if either restored SIPP vehicle input " + "is absent or degenerate, this vehicle-specific reform loses its " + "intended binding channel. A persisted 30,000-household Populace " + "smoke scored +$3.58 million in 2026 SNAP." + ), + "issue": "PolicyEngine/populace#49", + }, + { + "id": "voluntary_filing_aca_ptc_neutralization", + "name": "Voluntary tax filing ACA PTC neutralization", + "parameter_changes": {}, + "neutralized_variable": "would_file_taxes_voluntarily", + "budget_measure": "aca_ptc", + "period": 2024, + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "binding_inputs": list(VOLUNTARY_FILING_INPUTS), + "min_abs_effect": 100_000_000.0, + "reason": ( + "PolicyEngine-US includes would_file_taxes_voluntarily in " + "tax_unit_is_filer alongside required and credit filers. " + "Neutralizing only the restored SIPP filing-response leaf therefore " + "removes ACA premium tax credits from otherwise eligible voluntary " + "filers, so baseline-minus-neutralized aca_ptc must be positive. " + "With the filing leaf absent or degenerate, this isolated response " + "channel scores exactly $0." + ), + "issue": "PolicyEngine/populace#312", + }, + { + "id": "pre_subsidy_rent_neutralization", + "name": "Pre-subsidy rent neutralization", + "parameter_changes": {}, + "neutralized_variable": "pre_subsidy_rent", + "budget_measure": "snap", + "period": 2024, + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "binding_inputs": ["pre_subsidy_rent"], + "min_abs_effect": 1_000_000.0, + "reason": ( + "PolicyEngine-US includes pre_subsidy_rent in the SNAP shelter " + "deduction. Neutralizing only the restored ACS rent leaf therefore " + "reduces SNAP; on the pinned retired small eCPS artifact under " + "PolicyEngine-US 1.764.6, baseline-minus-neutralized SNAP is " + "+$11.731 billion. Without nondefault ACS rent, that effect is a " + "structural zero. The " + "other source-mapped housing leaves are enforced by their exact " + "ASEC mappings and signal gate; household tenure_type has no " + "standalone PolicyEngine-US 1.764.6 formula consumer." + ), + "issue": "PolicyEngine/populace#32", + }, + { + "id": "housing_assistance_take_up_neutralization", + "name": "Measured housing-assistance take-up neutralization", + "parameter_changes": {}, + "neutralized_variable": "takes_up_housing_assistance_if_eligible", + "budget_measure": "housing_assistance", + "period": 2024, + "effect_direction": "baseline_minus_reform", + "expected_sign": "positive", + "binding_inputs": ["takes_up_housing_assistance_if_eligible"], + "min_abs_effect": 100_000_000.0, + "reason": ( + "PolicyEngine-US multiplies HUD HAP by the restored SPM-unit " + "take-up leaf after eligibility. Populace keeps that leaf exactly " + "equal to source-backed housing-assistance receipt, so neutralizing " + "it must remove the assistance paid to measured/imputed recipients. " + "A 6,000-household production-ingredient smoke scored $202.795 " + "million baseline-minus-neutralized; the $100 million floor is " + "below that observed subset effect but far above numerical noise. " + "A default-only or absent carry makes the source reconciliation " + "or this uniquely isolating probe fail." + ), + "issue": "PolicyEngine/populace#312", + }, ] @@ -107,7 +1215,7 @@ def build_manifest() -> dict: populated_layers = { name for name, share in parity["nonzero_shares"].items() if float(share) > 0.0 - } + } | set(POST_REFERENCE_ECPS_REQUIRED_INPUTS) ssi_assets = set(SSI_COUNTABLE_RESOURCE_ASSETS) missing_assets = sorted(ssi_assets - populated_layers) @@ -116,6 +1224,19 @@ def build_manifest() -> dict: "SSI countable-resource asset inputs are not in the reference eCPS " f"populated surface, cannot pin them as required: {missing_assets}." ) + restored_inputs = set(RESTORED_REFERENCE_ECPS_REQUIRED_INPUTS) + missing_restored = sorted(restored_inputs - populated_layers) + if missing_restored: + raise ValueError( + "Restored reference inputs are absent from the reference populated " + f"surface: {missing_restored}." + ) + stale_restored_gaps = sorted(restored_inputs & set(known_gaps)) + if stale_restored_gaps: + raise ValueError( + "Restored reference inputs cannot remain in the parity-gap register: " + f"{stale_restored_gaps}." + ) columns: dict[str, dict] = {} for name in sorted(populated_layers): @@ -182,12 +1303,20 @@ def build_manifest() -> dict: ), "reference": reference, "derivation": ( - "Required surface = ecps_parity_reference.json populated layers " - "(input columns the pinned, sha-verified reference eCPS populates). " + "Required surface = input columns in the pinned, sha-verified " + "ecps_parity_reference.json populated layers, plus the documented " + "post-reference fsla_overtime_premium, " + "qualified_passenger_vehicle_loan_interest, five desired " + "retirement-contribution inputs, and " + "meets_ssi_disability_criteria required by shipped validation " + "probes. " "status='reviewed_exclusion' for ecps_parity_known_gaps.json entries " - "(reason+issue from that register); EXCEPT the SSI countable-resource " - "asset inputs (bank_account_assets, stock_assets, bond_assets), which " - "are status='required' with NO exclusion per PolicyEngine/populace#368 " + "(reason+issue from that register); EXCEPT every primary-source " + "restoration pinned by RESTORED_REFERENCE_ECPS_REQUIRED_INPUTS " + "(including the Section 199A QBI family), and the SSI countable-" + "resource asset inputs (bank_account_assets, " + "stock_assets, bond_assets), which are status='required' with NO " + "exclusion per PolicyEngine/populace#368 " "so the gate fails on today's artifacts and asset restoration " "(Deliverable 2) turns it green. All other populated layers are " "'required'. Regenerate with "