GPT-6.1 Sol
United States benchmark
Exact match by programall 18 output groups, best first
Bar length is the unweighted exact-match rate on that program's outputs; the right column is the program's share of the headline weight, which follows household dollars, so the bars do not average to the headline.
- Free school meals eligibility100.0%0.7%
- Local income tax100.0%0.3%
- Person-level CHIP eligibility100.0%0.2%
- Person-level Early Head Start eligibility100.0%3.1%
- Person-level Head Start eligibility100.0%1.2%
- Person-level WIC eligibility100.0%0.3%
- Reduced-price school meals eligibility100.0%0.1%
- Self-employment tax100.0%2.1%
- SSI100.0%2.0%
- TANF99.0%0.3%
- Federal refundable credits99.0%3.7%
- Person-level Medicaid eligibility96.5%30%
- Person-level Medicare eligibility91.9%11%
- SNAP91.4%4.1%
- State refundable credits89.8%0.6%
- Payroll tax86.9%16%
- Federal tax before refundable credits84.1%19%
- State tax before refundable credits74.4%6.0%
Table view (exact, within 1%, outputs), hardest first
| Program | Exact | Within 1% | n |
|---|---|---|---|
| State tax before refundable credits | 74.4% | 76.7% | 86 |
| Federal tax before refundable credits | 84.1% | 85.4% | 82 |
| Payroll tax | 86.9% | 86.9% | 99 |
| State refundable credits | 89.8% | 89.8% | 98 |
| SNAP | 91.4% | 92.5% | 93 |
| Person-level Medicare eligibility | 91.9% | 91.9% | 172 |
| Person-level Medicaid eligibility | 96.5% | 96.5% | 173 |
| Federal refundable credits | 99.0% | 99.0% | 98 |
| TANF | 99.0% | 99.0% | 100 |
| Free school meals eligibility | 100.0% | 100.0% | 100 |
| Local income tax | 100.0% | 100.0% | 100 |
| Person-level CHIP eligibility | 100.0% | 100.0% | 177 |
| Person-level Early Head Start eligibility | 100.0% | 100.0% | 38 |
| Person-level Head Start eligibility | 100.0% | 100.0% | 38 |
| Person-level WIC eligibility | 100.0% | 100.0% | 177 |
| Reduced-price school meals eligibility | 100.0% | 100.0% | 100 |
| Self-employment tax | 100.0% | 100.0% | 100 |
| SSI | 100.0% | 100.0% | 97 |
Hardest casesworst misses on positive references
- SNAP100% offReference $240 · predicted $0Inspect household #013
- Federal tax before refundable credits100% offReference $643 · predicted $0Inspect household #023
- State tax before refundable credits100% offReference $13 · predicted $0Inspect household #023
- SNAP100% offReference $288 · predicted $0Inspect household #027
- SNAP100% offReference $288 · predicted $0Inspect household #030
- Federal tax before refundable credits100% offReference $2,362 · predicted $0Inspect household #042
Scores are from the frozen manuscript snapshot under the AI-alone condition: structured responses over the same household facts (whole-scenario or in output subsets, per the serving-configuration table), no tools, graded against PolicyEngine reference outputs. See the leaderboard and paper for methodology. The frozen US annotations cover 7,860 rows whose legacy threshold score is below 1. That universe contains 7,856 of 7,856 exact-match misses and 4 exact hits. Another 2,107 rows with bounded score below 100 were not selected and have no audit annotation.