Skip to content

Resolve 2035+ TOB calibration instability before broad static rerun #115

Description

@MaxGhenis

Summary

Do not launch the broad static/conventional dashboard rerun until this is resolved.

We fixed several code-level causes of the 2034→2035 discontinuity, but the latest one-year sentinel shows a remaining methodology problem: when the 2035 H5 is recalibrated exactly under the corrected Trustees tax-threshold assumption, the current-law baseline matches Trustees, but the option impact jumps because the calibration weights become much more concentrated.

This issue lays out the current state so someone who was not on the debugging thread can pick it up.

User-visible symptom

The dashboard showed odd discontinuities around 2035, especially for the $700 credit / option11 case and for measures shown as percent of taxable payroll.

Old dashboard static option11:

year baseline TOB total option11 TOB impact
2034 $243.308B $15.300B
2035 $257.528B $20.694B

The baseline grows by about 5.8%, but the option impact grows by about 35%, which is suspicious.

Code issues already found/fixed locally

1. Tax assumption was applied before it should be

modal_batch/compute.py used the baseline tax-assumption reform unconditionally. That let the 2035-start Trustees tax-threshold assumption affect pre-2035 runs.

Local fix:

  • add year-gated tax-assumption loading
  • do not require tax-assumption-tagged H5 metadata before 2035
  • do not use stale baseline artifacts when the tax assumption is inactive

Relevant file: modal_batch/compute.py

2. The tax-threshold reform anchored 2035 to stale 2026 values

The long-run tax-threshold reform read raw parameter values before normal PolicyEngine uprating had materialized future CPI-indexed values. Result: when the 2035 reform was active, 2035 thresholds were wage-indexed from stale 2026 levels rather than from the 2034 current-law level.

Example found during debugging:

  • old broken 2035 bracket2 single: $52,325
  • fixed 2035 bracket2 single: $61,800

Local fix:

  • preserve default 2027-2034 threshold path
  • anchor 2035 wage indexing to the 2034 default value

Relevant file: policyengine_us_data/datasets/cps/long_term/tax_assumptions.py

3. Static runs accidentally activated labor-supply response machinery

Applying the baseline tax-assumption reform makes PolicyEngine treat the sim as a reform scenario. In policyengine-us-6830-port, labor_supply_behavioral_response checked p.elasticities.income == 0, but p.elasticities.income is a parameter node, not a scalar. That made static baseline computations enter the LSR branch even when all elasticities were zero.

Symptom: even a 50-household 2035 sample hung at income_tax start until interrupted.

Local fix:

  • check scalar elasticity leaves instead of the parameter node
  • keep conventional LSR active when income base or substitution decile elasticities are nonzero

Relevant file: policyengine_us/variables/gov/simulation/labor_supply_response/labor_supply_behavioral_response.py

Sentinel runs after those fixes

A. Corrected 2034 full sentinel, old H5

File: results/pre2035_tax_assumption_fix_static_sentinel_static/year_2034.csv

metric value
baseline TOB total $243.439B
option11 TOB impact $15.928B

B. Fixed tax assumption, old 2035 H5

File: results/tax_assumption_anchor_fix_static_sentinel_2035.csv

metric value
baseline TOB total $254.730B
Trustees 2035 target $257.528B
baseline miss -$2.798B (-1.087%)
option11 TOB impact $15.894B

This is smooth with 2034, but the baseline misses Trustees because the existing 2035 H5 weights were calibrated under the old/broken tax-threshold assumption.

C. Rebuilt exact 2035 H5 under corrected tax assumption

H5 built locally at:

/Users/maxghenis/.codex-worktrees/us-data-calibration-contract/tmp/trustees_core_threshold_2035_fixed/2035.h5

Metadata says the current-law baseline hits Trustees almost exactly:

target achieved error
OASDI TOB -$0.0008M
HI TOB +$0.0006M

Modal sentinel file:

results/tax_assumption_anchor_fix_static_sentinel_2035_recalibrated.csv

metric value
baseline TOB total $257.528B
option11 TOB impact $21.969B

This restores the baseline target but brings back the discontinuity.

Weight diagnostics

The exact recalibration appears to create a much more concentrated 2035 microdata support.

dataset positive HHs effective N top-10 weight share top-100 weight share
2034 old H5 6,682 811 5.44% 25.79%
2035 old H5 6,682 802 5.39% 25.98%
2035 rebuilt exact H5 6,264 352 12.58% 36.82%

That weight concentration is the leading explanation for why exact recalibration makes option11 jump even though the tax-threshold code is now fixed.

Current interpretation

There are two separable problems:

  1. Code bugs: tax-assumption timing, 2035 threshold anchoring, accidental static→LSR execution. These are locally fixed and tested.
  2. Calibration/methodology issue: exact 2035 TOB hard targeting under corrected thresholds may over-concentrate weights and amplify reform impacts. This must be resolved before a broad rerun.

Required next decision

Pick and validate a stable 2035+ calibration approach before rerunning the dashboard. Candidate approaches:

  1. regularized / approximate TOB calibration with explicit weight stability thresholds;
  2. preserve the old stable weights under the corrected tax assumption and align Trustees baseline post hoc;
  3. support augmentation for high-TOB households, then exact calibration;
  4. another documented method that preserves both baseline alignment and impact stability.

Acceptance criteria

Before rerunning all policy-year cells:

  • 2035 current-law OASDI and HI TOB match Trustees, or any residual baseline adjustment is explicit and documented.
  • 2034→2035 option11 impact is not driven by a sudden weight concentration artifact.
  • Weight diagnostics for 2035+ stay within agreed stability bounds, not just the loose current pass/fail thresholds.
  • Sentinel table includes at least 2034, 2035, and 2036 for option11 and full repeal, shown in dollars and percent of taxable payroll.
  • Modal runner uses explicit controls to avoid stale baseline artifacts during validation.
  • Only after those checks pass should we run the broad static panel and then apply the behavioral multipliers.

Verification already run

CRFB repo:

python -m pytest tests/test_modal_compute_tax_assumption.py tests/test_modal_compute_helpers.py -q
python -m ruff check modal_batch/compute.py tests/test_modal_compute_tax_assumption.py

Result: 10 passed, 1 skipped; ruff passed.

Additional smoke:

  • local 50-household fixed-tax-assumption income_tax calculation now completes after the LSR guard fix;
  • full Modal option11 2035 sentinel completes on both old and rebuilt 2035 H5s.

Related issue

This supersedes/refines the diagnosis in #79. The original symptom is still relevant, but the latest evidence shows the remaining blocker is calibration stability under the corrected tax assumption, not just release-artifact baseline stitching.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions