Skip to main content
PolicyBench

Inkling

Thinking Machines · AI alone, no tools
Current board, snapshot 2026-09-05
United States rank
#7 of 39
US exact
84.1%
United States

United States benchmark

Exact match
84.1%
Bounded score
95.2%
Parse rate
100.0%
Eligibility flags
970/979

Exact match by programall 18 output groups, best first

Bar length is the unweighted exact-match rate on that program's outputs; the right column is the program's share of the headline weight, which follows household dollars, so the bars do not average to the headline.

  1. Person-level CHIP eligibility100.0%0.2%
  2. Person-level Early Head Start eligibility100.0%3.1%
  3. Person-level Head Start eligibility100.0%1.2%
  4. Person-level Medicare eligibility100.0%11%
  5. Free school meals eligibility99.0%0.7%
  6. Local income tax99.0%0.3%
  7. Self-employment tax99.0%2.1%
  8. TANF99.0%0.3%
  9. SSI99.0%2.0%
  10. Person-level WIC eligibility98.9%0.3%
  11. Reduced-price school meals eligibility98.0%0.1%
  12. Person-level Medicaid eligibility97.7%30%
  13. Federal refundable credits87.0%3.7%
  14. SNAP81.4%4.1%
  15. Payroll tax81.0%16%
  16. State refundable credits80.0%0.6%
  17. Federal tax before refundable credits60.0%19%
  18. State tax before refundable credits53.0%6.0%
Table view (exact, within 1%, outputs), hardest first
ProgramExactWithin 1%n
State tax before refundable credits53.0%56.0%100
Federal tax before refundable credits60.0%68.0%100
State refundable credits80.0%80.0%100
Payroll tax81.0%86.0%100
SNAP81.4%86.6%97
Federal refundable credits87.0%92.0%100
Person-level Medicaid eligibility97.7%97.7%177
Reduced-price school meals eligibility98.0%98.0%100
Person-level WIC eligibility98.9%98.9%177
SSI99.0%99.0%97
Free school meals eligibility99.0%99.0%100
Local income tax99.0%99.0%100
Self-employment tax99.0%99.0%100
TANF99.0%99.0%100
Person-level CHIP eligibility100.0%100.0%177
Person-level Early Head Start eligibility100.0%100.0%38
Person-level Head Start eligibility100.0%100.0%38
Person-level Medicare eligibility100.0%100.0%172

Hardest casesworst misses on positive references

Scores are from the frozen manuscript snapshot under the AI-alone condition: structured responses over the same household facts (whole-scenario or in output subsets, per the serving-configuration table), no tools, graded against PolicyEngine reference outputs. See the leaderboard and paper for methodology. The frozen US annotations cover 8,783 rows whose legacy threshold score is below 1. That universe contains 8,780 of 8,780 exact-match misses and 3 exact hits. Another 1,605 rows with bounded score below 100 were not selected and have no audit annotation.