Claude Sonnet 5.5 debuts fifth as the references move to policyengine-us 2.15.17
Claude Sonnet 5.5, Grok 4.7 and DeepSeek V4.1 Flash joined the board on 2026-09-29, bringing it to 45 models. Claude Sonnet 5.5 scores 92.1% of answers within $1, weighted by household impact, #5 of 45. GPT-6 Luna, #4, also rounds to 92.1% and sits about level with it, less than 0.02 points ahead. Grok 4.7 scores 88.3% (#11) and DeepSeek V4.1 Flash 87.1% (#15). GPT-6 Sol still leads at 95.0%, ahead of Claude Opus 5.5 (93.7%) and GPT-5.6 Sol (93.6%).
Claude Sonnet 5.5 costs $0.0343 a household, against $0.0019 for GPT-6 Luna. DeepSeek V4.1 Flash costs $0.0242 at DeepSeek's peak list price, and Grok 4.7 costs $0.2597.
Claude Sonnet 5.5's API rejects forced tool calls, as Claude Opus 5.5's and Claude Fable 5.1's do, so its row answers as a JSON object and reasons with adaptive thinking, the provider default. Grok 4.7 answers through the forced tool call with a 1,800-second request timeout, because its whole-household onboarding probe ran past a 600-second timeout. DeepSeek serves V4.1 Flash under the alias deepseek-flash, and all 1,984 of the row's answers report that alias and one system fingerprint; none reports a version. PolicyBench labels the row from DeepSeek's September 10 release note, which put V4.1 Flash on that alias.
The release also moves PolicyBench's scored references from policyengine-us 1.755.4 to 2.15.17, the newest release when PolicyBench began sweeping the references on 2026-09-29 (uploaded at 00:23 UTC). The newer version includes the eight upstream fixes that PolicyBench applied as sandbox fixes for the September 22 references, and it encodes law the older version lacked. PolicyBench still builds each scored reference from the stated facts and law published before it froze the references on 2026-07-03, so it ported its nine publication conventions to the new version. With those conventions and an adapter that keeps Maryland county tax out of state income tax, policyengine-us 2.17.0, the newest release when PolicyBench checked PyPI on 2026-09-29 at 14:58 UTC, gives the same value as 2.15.17 for all 1,984 outputs.
The move changes four scored references. New Jersey's child tax credit schedule for 2026 to 2028, which the state approved on June 30, raises one household's state refundable credits from $5,342.40 to $5,842.40. Arizona raised the income limit for its broad-based categorical eligibility from 185% to 200% of the poverty guideline effective March, and the new limit gives an Arizona household $240 of SNAP where the reference was $0. policyengine-us now counts child support received as school-meal income, which ends a Pennsylvania household's eligibility for reduced-price meals. A fix to the rounding in New York's Empire State child credit phase-out (policyengine-us #9425) raises a New York household's state refundable credits from $650.50 to $667.
PolicyBench stops scoring three federal income tax outputs. policyengine-us 2.15.17 counts the whole of a listed state and local tax refund as income. Federal law counts the refund only to the extent the refunded tax lowered the household's federal tax in the year the household paid it, and the prompts do not say whether it did. PolicyBench also stops scoring one California household head's Medicaid eligibility, which its re-review of the household's excluded SNAP output flagged. On disability, the prompt says only that the head is disabled. The head's income, 141% of the poverty guideline, is above the 138% limit for the adult expansion group, so only a disability pathway leads to Medi-Cal, California's Medicaid program. Medi-Cal's Working Disabled Program requires SSI's definition of disability, which the prompt does not state, and policyengine-us 2.15.17 tests the general disability flag instead. The same unstated fact already keeps the household's SNAP out of scoring. PolicyBench now scores every model on 1,928 of its 1,984 requested outputs and excludes 56. Another two references move by less than $1, and PolicyBench re-reviewed the 19 excluded outputs whose values moved; all stay excluded.
Together, the new references and the Medicaid exclusion raise the exact rate of every one of the 42 earlier models, by 0.14 to 0.83 points. None of the 42 matched any of the three federal outputs now excluded, and all 42 had matched the Arizona household's old $0 SNAP reference, which none matches now. On the Medicaid output, 20 of the 42 had matched the reference and 22 had not. Among the 42, five models move up in the order: GLM-5.3-Flash (preview) moves above Inkling, Claude Opus 5 above Claude Fable 5, DeepSeek V4 Pro 0813 above DeepSeek V4 Flash 0731 and Gemini 3.7 Flash, Grok Build 0.1 above Claude Sonnet 4.6 and Gemini 3.5 Flash, and DeepSeek V4 Pro above Gemini 3.1 Flash Lite Preview. Every other pair keeps its order.
The Arizona household's income keeps it from qualifying for SNAP under the program's ordinary tests, so it qualifies only through broad-based categorical eligibility, and all 45 models answer $0 for it. The September 29 update to PolicyBench's September 23 note on such households adds it as a fifth household held back by income. The update also gives the three new models' answers for the four other households held back by income, in Connecticut, Texas, Michigan and Wisconsin, and for the four held back by savings.
Data
- Dashboard data release
- Reference upgrade record
- Reference sidecar with every change
- Exclusion record
- Paper serving-configuration table
- DeepSeek V4.1 Flash release note (September 10)
- Claude Sonnet 5.5 model page
- Grok 4.7 model page
- DeepSeek V4.1 Flash model page
- Note on the SNAP households that qualify through BBCE, with its September 29 update