Skip to main content
Numbers from release dashboard-data-20261010 (board snapshot 2026-10-10)

Claude Haiku 5.5 joins the board as the references move to policyengine-us 2.38.6

Claude Haiku 5.5 joined the board on 2026-10-10, bringing it to 47 models. It gets 79.6% of answers within $1, weighted by household impact, #37 of 47 and 4.8 points above Claude Haiku 4.5 (74.8%, #43). GPT-6 Sol still leads at 95.6%, ahead of Claude Opus 5.5 (94.4%) and GPT-5.6 Sol (93.5%).

Claude Haiku 5.5 cost $0.0013 a household, the least of any row with a cost above zero; Claude Haiku 4.5 cost $0.0554. Unlike Claude Opus 5.5 and Claude Sonnet 5.5, its API accepts a forced tool call, so its row takes the board's forced answer tool. A probe with that request returned no thinking block, where a request without the forced tool did, so the row answers without extended thinking, as Claude Opus 5's does. Re-run with tool_choice: auto, it gets 89.2%, which would rank #10.

The release also moves PolicyBench's references from policyengine-us 2.15.17 to 2.38.6, a release that fixes the engine defects behind 18 outputs. The previous release excluded 14 of them, and they return to scoring at the new version's values. The other four are the October 6 defects below, which stay scored at the new version's values. Each of the 18 lands within $1 of the corrected value PolicyBench's audits computed for it. Between them they take the upstream fixes for the IRA deduction's compensation limit and a dependent's contributions (policyengine-us #10032, #9801; three outputs), the IRA deduction's active-participant phase-out and California's itemized deduction conformity (policyengine-us #10032, #10034), New Jersey's worker unemployment and workforce contributions (policyengine-us #10031), Arizona's standard deduction indexing (policyengine-us #9928), California's itemized deduction conformity (policyengine-us #10034), Ohio's medical deduction for health insurance premiums (policyengine-us #10020), estate income in gross income (policyengine-us #10027, #9633; two outputs), Colorado's 2026 sales tax refund (policyengine-us #9946), the IRA deduction's active-participant phase-out (policyengine-us #10032; four outputs), New York's 2026 child and dependent care credit (policyengine-us #9948) and the IRA deduction's active-participant phase-out and estate income in gross income (policyengine-us #10032, #10027; two outputs). PolicyBench still builds each reference from the stated facts and from law published before it froze the references on 2026-07-03, under the same publication conventions.

Two rulings of October 6 apply too. PolicyBench's reference adversary flags the scored outputs where many models, or several of the strongest, agree on an answer the reference does not give, and a judge that cannot yet see the engine's derivation works each one from primary law. Its findings led to rulings to exclude eight outputs. Four were engine defects, in Arizona, Ohio, Colorado and New York; the new version fixes them, so PolicyBench keeps scoring those four instead. The other four are the federal and state income tax of a Pennsylvania and a Missouri household whose dependent earns $45,000 of wages and must file their own return. The output definitions do not say whether the household's income tax includes that return, so PolicyBench stops scoring them. It also stops scoring two Louisiana households' state income tax, whose reference rests on a 2026 standard deduction computed from published price indexes; Louisiana published the amount itself on September 28, after the freeze.

The new version counts Indiana county income tax in local income tax, and the prompt names no county, so PolicyBench stops scoring the local income tax of two Indiana households. It also counts Idaho's $10 permanent building fund tax in state income tax, which moves one household's reference from $6,818.34 to $6,828.34. Every model is now scored on 1,926 of its 1,984 requested outputs; PolicyBench excludes the other 58.

Against the previous release, the exact rate of every one of the 46 earlier models falls, by 0.24 to 1.37 points. The cause is the 14 restored outputs that the previous release did not score, which few models get right: half of them are answered within $1 by at most one of the 47 models, and four by none. Had they stayed excluded, every earlier model's rate would have risen, by 0.39 to 0.71 points.

Data