Skip to main content
PolicyBench

Qwen 3.8 Max

Alibaba · AI alone, no tools
United States rank
#25 of 29
US exact
71.5%
United States

United States benchmark

Exact match
71.5%
Bounded score
80.7%
Parse rate
100.0%
Eligibility flags
906/984

Score by programhardest first

ProgramExactWithin 1%n
Federal tax before refundable credits45.0%45.0%100
State tax before refundable credits54.0%56.0%100
Federal refundable credits66.0%66.0%100
Payroll tax71.0%74.0%100
SNAP77.0%77.0%100
State refundable credits79.0%79.0%100
Person-level Medicaid eligibility82.5%82.5%177
Person-level Head Start eligibility86.8%86.8%38
Person-level WIC eligibility89.3%89.3%177
Self-employment tax92.0%92.0%100
Person-level CHIP eligibility92.1%92.1%177
SSI96.0%96.0%100
Person-level Medicare eligibility96.6%96.6%177
Free school meals eligibility98.0%98.0%100
TANF98.0%98.0%100
Reduced-price school meals eligibility99.0%99.0%100
Local income tax100.0%100.0%100
Person-level Early Head Start eligibility100.0%100.0%38

Hardest casesworst misses on positive references

Scores are from the frozen manuscript snapshot under the AI-alone condition: one structured response per household, no tools, graded against PolicyEngine reference outputs. See the leaderboard and paper for methodology.