Skip to main content
PolicyBench

Notes

Dated records of board changes and findings. Each note names the data release its numbers come from; a test in the repository checks every number against the frozen snapshot of that release.

Numbers from release dashboard-data-20260905c (board snapshot 2026-09-05)

GPT-6 Astra debuts second: the two rules it invented

GPT-6 Astra reached the API on 2026-09-04 and joined the board on 2026-09-05 at 88.0% of answers within $1, #2 of 39, behind GPT-5.6 Sol (89.2%) and ahead of Claude Fable 5.1 (86.9%, #3). Every row on this board is scored on 1,973 of its 1,984 requested outputs: 11 outputs whose reference depends on an engine input the household facts never list are excluded for every model.

Row by row, the two OpenAI models agree far more than they differ. Both are right on 1,816 scored outputs and both wrong on 95. The gap is 33 outputs Astra misses and Sol gets right, against 29 the other way.

Astra's solo misses cluster in two invented rules. In 13 rows it treats a listed employer-sponsored insurance premium as a pre-tax salary reduction and subtracts it from wages before computing payroll tax, federal income tax or state income tax; the household facts state gross wages, and the reference keeps them. In 10 rows it marks a household member under 65 who is flagged disabled as Medicare-eligible; the reference grants Medicare before 65 only with 24 months of Social Security disability receipt recorded, and none of these households carry that input. The remaining 10 solo misses are one-off rule errors with no shared mechanism. One further Medicare row of the same kind, on scenario_074, sits on an excluded output and is outside every count here.

The judge's row notes name an employer-sponsored insurance premium on 26 of Astra's annotated rows and on 4 of Sol's. Sol netted the same premium once, on scenario_045, which cost it 2 of its 29 solo misses; the rest of Sol's solo misses share no mechanism.

The row list, with both predictions, the reference and the judge's note for every disagreement, is committed beside this note. Click any scenario in the explorer to see the household facts and each model's answer.

Data

Numbers from release dashboard-data-20260901c (board snapshot 2026-09-01)

Six SNAP households the top three models deny

20 of the 100 households qualify for SNAP in the PolicyEngine reference. The top three models on the board, GPT-5.6 Sol, Claude Fable 5.1, and Kimi K3, each predict $0 for the same 6: scenario_027, scenario_030, scenario_045, scenario_073, scenario_108, scenario_112. The reference for each is $2881 for the year, the minimum allotment (8% of the one-person maximum, $23.84 a month, for households of one or two).

Five of the six qualify through broad-based categorical eligibility: Connecticut, Michigan, Texas, and Wisconsin confer SNAP eligibility on households receiving a TANF-funded non-cash benefit, with gross-income limits of 200% of poverty (CT, MI, WI) and 165% (TX) and the net-income and asset tests waived. The sixth, scenario_112 in Texas, passes the federal tests; the models counted farm-rent income the rules exclude.

10 of the 20 eligible households qualify only through categorical eligibility: 6 because income exceeds the federal limits (one of those also exceeds the asset limit), 4 because savings alone exceed the federal asset limit. On the 4 asset cases the models compute benefits near the reference: Sol within 1% on three and 10% on one, Fable 5.1 within 1% on all four.

Across the 100 households, Sol's SNAP explanations mention categorical eligibility in 1 and assets in 36; Fable 5.1's mention broad-based categorical eligibility in 28; Kimi K3's mention categorical eligibility in 17.

Pathways were recomputed with policyengine-us 1.723.0 from the frozen scenarios and agree with the frozen reference SNAP values for all 100 households; the reference itself was generated with policyengine-us 1.755.4.

1 Unrounded frozen reference: $287.68.

Data

Numbers from release dashboard-data-20260901c (board snapshot 2026-09-01)

Claude Fable 5.1 added

Claude Fable 5.1 was released on 2026-09-01 and joined the board the same day. Board row: 86.3% of answers within $1, #2 of 33. 1,984 of 1,984 answers parsed.

The model's API rejects forced tool calls, which is how the board collects most models' answers, so its row answers as a JSON object and reasons at the provider default. The serving-configuration table in the paper records the transport for every row.

Sensitivity: with the answer tool declared under tool_choice auto, Fable 5.1 scores 87.5% (would rank #2). Under the same request Claude Fable 5 scores 86.9%. Fable 5's board row (79.9%) ran chunked under a forced tool call, which switches off its thinking, so the 6.4-point gap between the two board rows is mostly request shape.

The next board version moves every model to tool_choice auto so each provider's default reasoning engages under the recorded request shape (policybench#139).

Data