Skip to main content
PolicyBench

Claude Opus 5.5

Anthropic · AI alone, no tools
Current board, snapshot 2026-09-22
United States rank
#2 of 42
US exact
92.9%
United States

United States benchmark

Exact match
92.9%
Bounded score
97.4%
Parse rate
100.0%
Eligibility flags
960/976

Exact match by programall 18 output groups, best first

Bar length is the unweighted exact-match rate on that program's outputs; the right column is the program's share of the headline weight, which follows household dollars, so the bars do not average to the headline.

  1. Federal refundable credits100.0%3.7%
  2. Local income tax100.0%0.3%
  3. Person-level Early Head Start eligibility100.0%3.1%
  4. Person-level Head Start eligibility100.0%1.2%
  5. Person-level Medicare eligibility100.0%11%
  6. Person-level WIC eligibility100.0%0.3%
  7. Self-employment tax100.0%2.1%
  8. SSI100.0%2.0%
  9. TANF99.0%0.3%
  10. Free school meals eligibility98.0%0.7%
  11. Reduced-price school meals eligibility98.0%0.1%
  12. Person-level Medicaid eligibility97.1%30%
  13. Person-level CHIP eligibility96.0%0.2%
  14. Payroll tax92.9%16%
  15. State refundable credits90.8%0.6%
  16. SNAP89.2%4.1%
  17. State tax before refundable credits80.2%6.0%
  18. Federal tax before refundable credits80.0%19%
Table view (exact, within 1%, outputs), hardest first
ProgramExactWithin 1%n
Federal tax before refundable credits80.0%84.7%85
State tax before refundable credits80.2%83.7%86
SNAP89.2%91.4%93
State refundable credits90.8%90.8%98
Payroll tax92.9%94.9%99
Person-level CHIP eligibility96.0%96.0%177
Person-level Medicaid eligibility97.1%97.1%174
Free school meals eligibility98.0%98.0%100
Reduced-price school meals eligibility98.0%98.0%100
TANF99.0%100.0%100
Federal refundable credits100.0%100.0%98
Local income tax100.0%100.0%100
Person-level Early Head Start eligibility100.0%100.0%38
Person-level Head Start eligibility100.0%100.0%38
Person-level Medicare eligibility100.0%100.0%172
Person-level WIC eligibility100.0%100.0%177
Self-employment tax100.0%100.0%100
SSI100.0%100.0%97

Hardest casesworst misses on positive references

Scores are from the frozen manuscript snapshot under the AI-alone condition: structured responses over the same household facts (whole-scenario or in output subsets, per the serving-configuration table), no tools, graded against PolicyEngine reference outputs. See the leaderboard and paper for methodology. The frozen US annotations cover 7,583 rows whose legacy threshold score is below 1. That universe contains 7,579 of 7,579 exact-match misses and 4 exact hits. Another 1,843 rows with bounded score below 100 were not selected and have no audit annotation.