Skip to main content
PolicyBench

MiniMax M3

MiniMax · AI alone, no tools
Current board, snapshot 2026-09-30
United States rank
#40 of 46
US exact
76.9%
United States

United States benchmark

Exact match
76.9%
Bounded score
82.9%
Parse rate
100.0%
Eligibility flags
933/975

Exact match by programall 18 output groups, best first

Bar length is the unweighted exact-match rate on that program's outputs; the right column is the program's share of the headline weight, which follows household dollars, so the bars do not average to the headline.

  1. Local income tax100.0%0.3%
  2. Reduced-price school meals eligibility100.0%0.1%
  3. Free school meals eligibility99.0%0.7%
  4. TANF99.0%0.3%
  5. Person-level Medicare eligibility98.3%11%
  6. SSI97.9%2.0%
  7. Person-level CHIP eligibility97.7%0.2%
  8. Person-level Early Head Start eligibility97.4%3.1%
  9. Person-level WIC eligibility97.2%0.3%
  10. Person-level Head Start eligibility94.7%1.2%
  11. Self-employment tax93.0%2.1%
  12. SNAP86.0%4.1%
  13. Person-level Medicaid eligibility85.0%30%
  14. State refundable credits81.6%0.6%
  15. Federal refundable credits80.6%3.7%
  16. Payroll tax66.3%16%
  17. State tax before refundable credits61.2%6.0%
  18. Federal tax before refundable credits54.4%19%
Table view (exact, within 1%, outputs), hardest first
ProgramExactWithin 1%n
Federal tax before refundable credits54.4%57.0%79
State tax before refundable credits61.2%62.4%85
Payroll tax66.3%72.6%95
Federal refundable credits80.6%80.6%98
State refundable credits81.6%81.6%98
Person-level Medicaid eligibility85.0%85.0%173
SNAP86.0%86.0%93
Self-employment tax93.0%95.0%100
Person-level Head Start eligibility94.7%94.7%38
Person-level WIC eligibility97.2%97.2%177
Person-level Early Head Start eligibility97.4%97.4%38
Person-level CHIP eligibility97.7%97.7%177
SSI97.9%97.9%97
Person-level Medicare eligibility98.3%98.3%172
Free school meals eligibility99.0%99.0%100
TANF99.0%99.0%100
Local income tax100.0%100.0%100
Reduced-price school meals eligibility100.0%100.0%100

Hardest casesworst misses on positive references

Scores are from the frozen manuscript snapshot under the AI-alone condition: structured responses over the same household facts (whole-scenario or in output subsets, per the serving-configuration table), no tools, graded against PolicyEngine reference outputs. See the leaderboard and paper for methodology. The frozen US annotations cover 7,527 rows whose legacy threshold score is below 1. That universe contains 7,523 of 7,523 exact-match misses and 4 exact hits. Another 2,072 rows with bounded score below 100 were not selected and have no audit annotation.