Skip to main content
Numbers from release dashboard-data-20260905c (board snapshot 2026-09-05)

GPT-6 Astra debuts second: the two rules it invented

GPT-6 Astra reached the API on 2026-09-04 and joined the board on 2026-09-05 at 88.0% of answers within $1, #2 of 39, behind GPT-5.6 Sol (89.2%) and ahead of Claude Fable 5.1 (86.9%, #3). Every row on this board is scored on 1,973 of its 1,984 requested outputs: 11 outputs whose reference depends on an engine input the household facts never list are excluded for every model.

Row by row, the two OpenAI models agree far more than they differ. Both are right on 1,816 scored outputs and both wrong on 95. The gap is 33 outputs Astra misses and Sol gets right, against 29 the other way.

Astra's solo misses cluster in two invented rules. In 13 rows it treats a listed employer-sponsored insurance premium as a pre-tax salary reduction and subtracts it from wages before computing payroll tax, federal income tax or state income tax; the household facts state gross wages, and the reference keeps them. In 10 rows it marks a household member under 65 who is flagged disabled as Medicare-eligible; the reference grants Medicare before 65 only with 24 months of Social Security disability receipt recorded, and none of these households carry that input. The remaining 10 solo misses are one-off rule errors with no shared mechanism. One further Medicare row of the same kind, on scenario_074, sits on an excluded output and is outside every count here.

The judge's row notes name an employer-sponsored insurance premium on 26 of Astra's annotated rows and on 4 of Sol's. Sol netted the same premium once, on scenario_045, which cost it 2 of its 29 solo misses; the rest of Sol's solo misses share no mechanism.

The row list, with both predictions, the reference and the judge's note for every disagreement, is committed beside this note. Click any scenario in the explorer to see the household facts and each model's answer.

Data