PolicyBench
Claude Fable 5.1
Anthropic · AI alone, no tools
Current board, snapshot 2026-09-01
United States rank
#2 of 33
US exact
86.3%
United States
United States benchmark
Exact match
86.3%
Bounded score
95.3%
Parse rate
100.0%
Eligibility flags
958/984
Score by programhardest first
| Program | Exact | Within 1% | n |
|---|---|---|---|
| State tax before refundable credits | 63.0% | 70.0% | 100 |
| Federal tax before refundable credits | 69.0% | 79.0% | 100 |
| SNAP | 79.0% | 84.0% | 100 |
| Payroll tax | 88.0% | 90.0% | 100 |
| State refundable credits | 89.0% | 90.0% | 100 |
| Person-level Medicare eligibility | 93.8% | 93.8% | 177 |
| SSI | 96.0% | 96.0% | 100 |
| Person-level Medicaid eligibility | 96.6% | 96.6% | 177 |
| Federal refundable credits | 97.0% | 97.0% | 100 |
| Person-level CHIP eligibility | 97.2% | 97.2% | 177 |
| Free school meals eligibility | 98.0% | 98.0% | 100 |
| Reduced-price school meals eligibility | 98.0% | 98.0% | 100 |
| TANF | 99.0% | 99.0% | 100 |
| Local income tax | 100.0% | 100.0% | 100 |
| Person-level Early Head Start eligibility | 100.0% | 100.0% | 38 |
| Person-level Head Start eligibility | 100.0% | 100.0% | 38 |
| Person-level WIC eligibility | 100.0% | 100.0% | 177 |
| Self-employment tax | 100.0% | 100.0% | 100 |
Hardest casesworst misses on positive references
- SNAP675% offReference $461 · predicted $3,576Inspect household #023
- State refundable credits311% offReference $19 · predicted $78Inspect household #043
- Federal tax before refundable credits282% offReference $1,346 · predicted $5,145Inspect household #039
- State tax before refundable credits186% offReference $514 · predicted $1,473Inspect household #039
- Federal tax before refundable credits123% offReference $5,200 · predicted $11,600Inspect household #123
- Federal tax before refundable credits100% offReference $643 · predicted $0Inspect household #023
Scores are from the frozen manuscript snapshot under the AI-alone condition: structured responses over the same household facts (whole-scenario or in output subsets, per the serving-configuration table), no tools, graded against PolicyEngine reference outputs. See the leaderboard and paper for methodology. The frozen US annotations cover 7,840 rows whose legacy threshold score is below 1. That universe contains 7,838 of 7,838 exact-match misses and 2 exact hits. Another 1,324 rows with bounded score below 100 were not selected and have no audit annotation.