Skip to main content
PolicyBench

Claude Fable 5.1

Anthropic · AI alone, no tools
Current board, snapshot 2026-09-01
United States rank
#2 of 33
US exact
86.3%
United States

United States benchmark

Exact match
86.3%
Bounded score
95.3%
Parse rate
100.0%
Eligibility flags
958/984

Score by programhardest first

ProgramExactWithin 1%n
State tax before refundable credits63.0%70.0%100
Federal tax before refundable credits69.0%79.0%100
SNAP79.0%84.0%100
Payroll tax88.0%90.0%100
State refundable credits89.0%90.0%100
Person-level Medicare eligibility93.8%93.8%177
SSI96.0%96.0%100
Person-level Medicaid eligibility96.6%96.6%177
Federal refundable credits97.0%97.0%100
Person-level CHIP eligibility97.2%97.2%177
Free school meals eligibility98.0%98.0%100
Reduced-price school meals eligibility98.0%98.0%100
TANF99.0%99.0%100
Local income tax100.0%100.0%100
Person-level Early Head Start eligibility100.0%100.0%38
Person-level Head Start eligibility100.0%100.0%38
Person-level WIC eligibility100.0%100.0%177
Self-employment tax100.0%100.0%100

Hardest casesworst misses on positive references

Scores are from the frozen manuscript snapshot under the AI-alone condition: structured responses over the same household facts (whole-scenario or in output subsets, per the serving-configuration table), no tools, graded against PolicyEngine reference outputs. See the leaderboard and paper for methodology. The frozen US annotations cover 7,840 rows whose legacy threshold score is below 1. That universe contains 7,838 of 7,838 exact-match misses and 2 exact hits. Another 1,324 rows with bounded score below 100 were not selected and have no audit annotation.