Skip to main content
Numbers from release dashboard-data-20260922 (board snapshot 2026-09-22)

GPT-6 Sol debuts first, Claude Opus 5.5 second

GPT-6 Sol, GPT-6 Luna and Claude Opus 5.5 joined the board on 2026-09-22. GPT-6 Sol leads at 94.2% of answers within $1, weighted by household impact, #1 of 42 and 1.3 points above Claude Opus 5.5 (92.9%, #2). GPT-5.6 Sol, the previous leader, is #3 at 92.7%.

GPT-6 Luna scores 91.4% (#4) at $0.0019 a household, against $0.0271 for GPT-6 Sol and $0.0675 for Claude Opus 5.5.

Claude Opus 5.5's API rejects forced tool calls, as Claude Fable 5.1's does, so its row answers as a JSON object and reasons at the provider default. Claude Opus 5's board row runs the forced tool call, which switches that model's thinking off, and scores 83.4%; its tool_choice auto re-run, where it reasons, scores 89.3%. The two Opus rows differ in serving shape as well as in model, so the 9.6-point gap between them is not a measure of the model change alone.

Every row is scored on 1,932 of its 1,984 requested outputs. This release regenerated 26 references and excludes 52 outputs from scoring for every model; the reference audit note explains both.

The next board version moves every model to tool_choice auto so each provider's default reasoning engages under the recorded request shape (policybench#139).

Data