Five SNAP households almost every model gets wrong
On the September 22 board, 13 scored SNAP outputs have a positive reference. For 5 of them (scenario_027, scenario_030, scenario_045, scenario_073, scenario_108), the reference is $288 for 2026, twelve months of the $24 minimum benefit. Each of these households has one or two people, lives in Connecticut, Michigan, Texas or Wisconsin, and qualifies for SNAP only through broad-based categorical eligibility (BBCE). For each, 30% of net income exceeds the maximum allotment, so the regular formula gives nothing and the minimum applies.
No model on the 42-model board gets all five within $1. Of the 210 answers requested, 3 came back with no value and no explanation, and 5 are within $1: 1 for scenario_027, 2 for scenario_073 and 2 for scenario_108; none are for scenario_030 or scenario_045. 20 models answer $0 for all five. GPT-6 Astra gets the most, 3 of the five, and answers $0 for the other two. The most common wrong amount other than $0 is $276, twelve months of $23, given 8 times; 7 of those explanations call it the minimum benefit. 26 of the 207 explanations, from 16 models, mention categorical eligibility, and 14 of those still answer $0.
The top three models on the board get at most one each. GPT-6 Sol answers $0 for four households and $1,208.40 for scenario_030. Claude Opus 5.5 answers $288 for scenario_108, $1,208 for scenario_030 and $0 for the other three. GPT-5.6 Sol answers $0 for all five. On scenario_030 both GPT-6 Sol and Claude Opus 5.5 compute the benefit from wages alone; the engine also counts the household's $12,000 of financial assistance as income, and with it 30% of net income exceeds the $298 maximum. Across all 100 households, the SNAP explanations of GPT-6 Sol mention categorical eligibility in 1, Claude Opus 5.5's in 18 and GPT-5.6 Sol's in 1.
In policyengine-us 1.755.4, the engine version behind the references, a household is categorically eligible for SNAP if every member receives SSI, if it receives TANF cash assistance, or if it is eligible for a TANF-funded non-cash benefit. On the last route the engine checks eligibility for the benefit, not receipt of it. In these four states that eligibility turns on gross income: at most 200% of the poverty guideline in Connecticut, Michigan and Wisconsin and 165% in Texas, with no net income test and no asset test apart from a $5,000 limit in Texas. All five households pass, none receives TANF cash assistance, and each fails at least one ordinary SNAP test. All five have gross income above 130% of the poverty guideline, the ordinary SNAP gross limit, but only 2 fail the gross income test: the other 3 have a member the engine treats as elderly or disabled, which exempts them from it. Of the five, 4 fail the net income test and 1 the asset test.
The prompt does not mention categorical eligibility or a non-cash benefit; TANF appears only as a requested output, the household's annual TANF benefit. It tells models to treat "any unlisted numeric input as 0 and any other unlisted household fact, boolean, or status input as false," to "assume tax filing and program take-up when required" and not to "infer unlisted income, expenses, assets, benefit receipt, rent, or health coverage." The engine's rule is consistent with the take-up instruction if the non-cash benefit counts as a program the household takes up. A model that reads the first or last instruction as ruling out a benefit the prompt does not list would apply only the ordinary tests, and each of the five fails at least one. Of the 172 answers of $0, 1 mentions TANF: Claude Sonnet 4.6, on scenario_030, writes that Texas's categorical eligibility covers households receiving TANF or SSI and that this household receives neither.
The September 3 note, on the September 1 board of 33 models, counted six households at $287.68. That figure was the frozen engine's unrounded minimum: $23.84 a month through September and a projected $24.37 from October. The September 22 references hold October to December at the FY2026 schedule, the last USDA published before the reference freeze, and round the minimum to the nearest dollar, which policyengine-us has done since #9162, merged after the freeze. The sixth household, scenario_112 in Texas, passes the ordinary tests but is no longer scored: its reference assumed 40 hours of work a week, which the prompt does not list, and under the prompt's rule that unlisted numbers are 0, the time limit for able-bodied adults without dependents applies. That note also said these states confer eligibility on households receiving the non-cash benefit and waive the asset test; the engine tests eligibility for the benefit, and Texas keeps an asset limit.
Data
- scenario_027
- scenario_030
- scenario_045
- scenario_073
- scenario_108
- Every model's answer for the five households
- SNAP pathways under the September 22 references
- Dashboard data release
- Six SNAP households note (September 3)
- Reference audit note
- Exclusion record
- Prompt preface
- policyengine-us #9162 (SNAP minimum rounding)