Skip to main content
PolicyBench

Claude Opus 5

Anthropic · AI alone, no tools
United States rank
#8 of 27
US exact
79.8%
United States

United States benchmark

Exact match
79.8%
Bounded score
90.9%
Parse rate
100.0%
Eligibility flags
952/984

Score by programhardest first

ProgramExactWithin 1%n
Federal tax before refundable credits50.0%50.0%100
State tax before refundable credits59.0%61.0%100
SNAP73.0%73.0%100
State refundable credits79.0%80.0%100
Federal refundable credits83.0%84.0%100
Payroll tax85.0%90.0%100
Person-level Medicaid eligibility93.2%93.2%177
Person-level Medicare eligibility93.8%93.8%177
SSI96.0%96.0%100
Free school meals eligibility98.0%98.0%100
Reduced-price school meals eligibility98.0%98.0%100
Self-employment tax98.0%98.0%100
Person-level CHIP eligibility98.3%98.3%177
Person-level WIC eligibility98.9%98.9%177
TANF99.0%99.0%100
Local income tax100.0%100.0%100
Person-level Early Head Start eligibility100.0%100.0%38
Person-level Head Start eligibility100.0%100.0%38

Hardest casesworst misses on positive references

Scores are from the frozen manuscript snapshot under the AI-alone condition: one structured response per household, no tools, graded against PolicyEngine reference outputs. See the leaderboard and paper for methodology.