Skip to main content
PolicyBench

Gemini 3 Flash Preview

Google · AI alone, no tools
Current board, snapshot 2026-09-30
United States rank
#29 of 46
US exact
82.2%
United States

United States benchmark

Exact match
82.2%
Bounded score
92.8%
Parse rate
100.0%
Eligibility flags
964/975

Exact match by programall 18 output groups, best first

Bar length is the unweighted exact-match rate on that program's outputs; the right column is the program's share of the headline weight, which follows household dollars, so the bars do not average to the headline.

  1. Free school meals eligibility100.0%0.7%
  2. Person-level Early Head Start eligibility100.0%3.1%
  3. Person-level Head Start eligibility100.0%1.2%
  4. Person-level CHIP eligibility99.4%0.2%
  5. Person-level WIC eligibility99.4%0.3%
  6. Person-level Medicare eligibility99.4%11%
  7. Reduced-price school meals eligibility99.0%0.1%
  8. Self-employment tax99.0%2.1%
  9. TANF99.0%0.3%
  10. Local income tax98.0%0.3%
  11. SSI97.9%2.0%
  12. Person-level Medicaid eligibility96.0%30%
  13. Federal refundable credits85.7%3.7%
  14. State refundable credits82.7%0.6%
  15. SNAP81.7%4.1%
  16. Payroll tax68.4%16%
  17. State tax before refundable credits63.5%6.0%
  18. Federal tax before refundable credits55.7%19%
Table view (exact, within 1%, outputs), hardest first
ProgramExactWithin 1%n
Federal tax before refundable credits55.7%57.0%79
State tax before refundable credits63.5%64.7%85
Payroll tax68.4%69.5%95
SNAP81.7%81.7%93
State refundable credits82.7%82.7%98
Federal refundable credits85.7%86.7%98
Person-level Medicaid eligibility96.0%96.0%173
SSI97.9%97.9%97
Local income tax98.0%98.0%100
Reduced-price school meals eligibility99.0%99.0%100
Self-employment tax99.0%99.0%100
TANF99.0%99.0%100
Person-level Medicare eligibility99.4%99.4%172
Person-level CHIP eligibility99.4%99.4%177
Person-level WIC eligibility99.4%99.4%177
Free school meals eligibility100.0%100.0%100
Person-level Early Head Start eligibility100.0%100.0%38
Person-level Head Start eligibility100.0%100.0%38

Hardest casesworst misses on positive references

Scores are from the frozen manuscript snapshot under the AI-alone condition: structured responses over the same household facts (whole-scenario or in output subsets, per the serving-configuration table), no tools, graded against PolicyEngine reference outputs. See the leaderboard and paper for methodology. The frozen US annotations cover 7,527 rows whose legacy threshold score is below 1. That universe contains 7,523 of 7,523 exact-match misses and 4 exact hits. Another 2,072 rows with bounded score below 100 were not selected and have no audit annotation.