Skip to main content
PolicyBench

Kimi K3

Moonshot AI · AI alone, no tools
United States rank
#2 of 25
US exact
86.2%
United States

United States benchmark

Exact match
86.2%
Bounded score
93.4%
Parse rate
96.8%
Eligibility flags
928/984

Score by programhardest first

ProgramExactWithin 1%n
State tax before refundable credits61.0%64.0%100
Federal tax before refundable credits71.0%79.0%100
SNAP77.0%81.0%100
State refundable credits78.0%79.0%100
Federal refundable credits90.0%92.0%100
Person-level Medicare eligibility91.0%91.0%177
Payroll tax92.0%93.0%100
SSI93.0%93.0%100
Person-level CHIP eligibility93.8%93.8%177
Person-level Medicaid eligibility94.4%94.4%177
Person-level Early Head Start eligibility94.7%94.7%38
Person-level Head Start eligibility94.7%94.7%38
Reduced-price school meals eligibility95.0%95.0%100
Free school meals eligibility96.0%96.0%100
TANF96.0%96.0%100
Person-level WIC eligibility96.6%96.6%177
Local income tax97.0%97.0%100
Self-employment tax97.0%97.0%100

Hardest casesworst misses on positive references

Scores are from the frozen manuscript snapshot under the AI-alone condition: one structured response per household, no tools, graded against PolicyEngine reference outputs. See the leaderboard and paper for methodology.