Skip to main content
PolicyBench

GLM-5.3

Z.ai · AI alone, no tools
Current board, snapshot 2026-09-05
United States rank
#22 of 39
US exact
77.8%
United States

United States benchmark

Exact match
77.8%
Bounded score
88.5%
Parse rate
96.1%
Eligibility flags
913/979

Exact match by programall 18 output groups, best first

Bar length is the unweighted exact-match rate on that program's outputs; the right column is the program's share of the headline weight, which follows household dollars, so the bars do not average to the headline.

  1. Local income tax97.0%0.3%
  2. Self-employment tax97.0%2.1%
  3. Reduced-price school meals eligibility96.0%0.1%
  4. SSI95.9%2.0%
  5. Free school meals eligibility95.0%0.7%
  6. Person-level WIC eligibility94.9%0.3%
  7. Person-level Medicare eligibility94.8%11%
  8. TANF94.0%0.3%
  9. Person-level CHIP eligibility93.2%0.2%
  10. Person-level Early Head Start eligibility92.1%3.1%
  11. Person-level Head Start eligibility92.1%1.2%
  12. Person-level Medicaid eligibility88.1%30%
  13. Federal refundable credits81.0%3.7%
  14. Payroll tax81.0%16%
  15. SNAP77.3%4.1%
  16. State refundable credits76.0%0.6%
  17. State tax before refundable credits49.0%6.0%
  18. Federal tax before refundable credits48.0%19%
Table view (exact, within 1%, outputs), hardest first
ProgramExactWithin 1%n
Federal tax before refundable credits48.0%53.0%100
State tax before refundable credits49.0%54.0%100
State refundable credits76.0%76.0%100
SNAP77.3%78.4%97
Federal refundable credits81.0%82.0%100
Payroll tax81.0%84.0%100
Person-level Medicaid eligibility88.1%88.1%177
Person-level Early Head Start eligibility92.1%92.1%38
Person-level Head Start eligibility92.1%92.1%38
Person-level CHIP eligibility93.2%93.2%177
TANF94.0%94.0%100
Person-level Medicare eligibility94.8%94.8%172
Person-level WIC eligibility94.9%94.9%177
Free school meals eligibility95.0%95.0%100
SSI95.9%96.9%97
Reduced-price school meals eligibility96.0%96.0%100
Local income tax97.0%97.0%100
Self-employment tax97.0%97.0%100

Hardest casesworst misses on positive references

Scores are from the frozen manuscript snapshot under the AI-alone condition: structured responses over the same household facts (whole-scenario or in output subsets, per the serving-configuration table), no tools, graded against PolicyEngine reference outputs. See the leaderboard and paper for methodology. The frozen US annotations cover 8,783 rows whose legacy threshold score is below 1. That universe contains 8,780 of 8,780 exact-match misses and 3 exact hits. Another 1,605 rows with bounded score below 100 were not selected and have no audit annotation.