# Changelog

Every released version of the dataset, newest first. Versions are
`YYYY.MM.PATCH`: `YYYY.MM` is the snapshot period the figures describe, and
`PATCH` increments when something in that snapshot is corrected after release.

A citation should name the version, not just the URL, because the figures
behind a given URL change. Corrections that change a published number are
listed here **and** on [corrections.html](https://benchmarks.santismm.com/corrections.html)
with the reason and the date.

The version in force is published in three places that a guard keeps in
agreement: `version` in [data.json](https://benchmarks.santismm.com/data.json),
`version` in [CITATION.cff](https://github.com/santismm/benchmarks/blob/main/CITATION.cff),
and the newest entry here.

---

## 2026.08.0 — 2026-08-22

Snapshot moved from **2026-06** to **2026-08**. Same 69 rows: no benchmark was
added or removed (FrontierCode was re-pointed, see below). Figures were
re-verified against public boards and trackers via web research on 2026-08-22;
evidence links are in the pull request that shipped this version.

### Context change affecting many rows

Claude Fable 5, suspended 2026-06-12 under a US export-control directive, was
**reinstated 2026-07-01** (controls lifted 06-30). Every "SUSPENDED / not
currently runnable" annotation from 2026.06.0 has been rewritten. Mythos 5
remains restricted-access. New frontier releases since June: GPT-5.6
(Luna/Terra/Sol, GA 07-09), Claude Opus 5 (07-24), Kimi K3, Grok 4.5/4.6,
Qwen3.8, Meta Muse Spark 1.1, GLM-5.x, DeepSeek V4.

### SOTA changes (12 rows)

- **SWE-bench Verified** 95 → 96 (Claude Opus 5; boards range 96–97).
- **OSWorld (Verified)** 85.4 → 86.1 (Qwen3.8 Max; top cluster within ~1.1pt).
- **BrowseComp** 90.1 → 92.2 (GPT-5.6 Sol Ultra setup).
- **Terminal-Bench 2.1** 83.4 → 89.5 (GPT-5.6 Sol, AA board; public snapshots differ).
- **BFCL v4** 73 → 75.0 (Qwen3.7 Max — the June row's notes already said 75; the cell now matches). Band active → near.
- **FrontierMath** 52.4 (v1) → 87.7 (GPT-5.5 Pro on Epoch's corrected v2; Fable 5 87.0 statistically tied). **v1 and v2 figures are not comparable.** Band active → near.
- **FrontierCode**: Cognition deprecated the Diamond subset (2026-07); the row now tracks **FrontierCode 1.1 (Main)** at 53.5 (Fable 5). Renamed from "FrontierCode (Diamond)". Band open → active.
- **ARC-AGI-3** 12.58 → 30.2 (Claude Opus 5 — first frontier LLM in the lead; aggregator-tracked, verify against arcprize.org).
- **Vending-Bench 2** $10,937 → $11,181 (Claude Opus 5, Andon board, with documented misalignment findings).
- **WebArena** 71.6 → 74.3 (WebTactix / DeepSeek v3.2; now carried by a live tracker, resolving June's skepticism). Source repointed to the steel.dev board.
- **WebDev / Code Arena** 1567 → ~1582 Elo (Opus 4.8; Elo re-based after the WebDev/Code Arena merge).
- **ExploitBench** holder changed at the same 78.0: now Fable 5 on the public board (was Mythos-5-unblocked, Anthropic-reported).

### Deliberately NOT changed, with the reasoning in each row's notes

- **ARC-AGI-2** stays 52.9: the 92.5 (GPT-5.6 Sol) / 90.4 (Opus 5) figures circulating are high-compute self-reports, not ARC-Prize-verified under efficiency limits — same bar the June row applied.
- **PaperBench** stays 24.4 (verified o1-high): the 93.0 aggregator snapshot for Qwen3.8 Max has unclear protocol/verification.
- **HLE** stays 64.7 tool-assisted; AA's text-only board now reads Fable 5 55.5 — protocol difference, both recorded.
- **GDPval** stays 41 win/tie%: the cited board moved to a re-based Elo (Opus 5 1862 #1); % and Elo eras are kept separate.
- **METR Time Horizon** stays 14.5h measured: no measured Fable 5 / Opus 5 horizons exist yet; ~61h figures in circulation are extrapolations.
- **tau2-bench (76.5)** was **not re-verified** against its source this snapshot; aggregate snapshots showing GLM-5.2 at 99.1 use a different protocol. Flagged in the row.
- Rows with no newer published figure keep their June values (Aider Polyglot, GAIA, SWE-Lancer, Konwinski Prize, Spider 2.0, BIRD-SQL, ScreenSpot-Pro, MLE-bench, ScienceAgentBench, Blueprint-Bench 2, AutomationBench, LAB, HealthBench Professional, BioMysteryBench, and the long tail of academic rows).

### Tooling

- `scripts/build-derived.mjs` added: regenerates `data.csv`, the copy embedded
  in `index.html`, `catalog.md` and `agents.html` from `data.json`, so the five
  representations cannot drift. The CI guards still verify the output.



First versioned release. The dataset already existed; this is the point from
which changes to it are tracked, so the snapshot can be cited unambiguously.

- 69 benchmarks and meta-leaderboards, figures as of the **2026-06** snapshot.
- Published under CC-BY-4.0 with an explicit `cite_as` string in `data.json`.
- Methodology published at
  [methodology.html](https://benchmarks.santismm.com/methodology.html):
  what saturation and priority-to-beat mean, how they are computed, where the
  figures come from, and what the dataset does not cover.
- Corrections log opened at
  [corrections.html](https://benchmarks.santismm.com/corrections.html). Empty
  at release; entries are added rather than figures being changed quietly.
- Embeddable table published at
  [embed.html](https://benchmarks.santismm.com/embed.html).

No benchmark figure changed in this release.

### Known limitations at this version

Stated here because a report citing this dataset inherits them. The full list,
with reasoning, is in [methodology.html](https://benchmarks.santismm.com/methodology.html).

- Several 2026-06 SOTA figures are vendor-reported and were produced on
  harnesses that are not standardised across rows.
- Some figures come from models that are not generally available, including
  ones suspended after publication. Each affected row says so in its `notes`.
- `p1`–`p4`, and therefore `priority`, are editorial 1–5 judgements, not
  measurements.
- Human baselines come from separate studies with different protocols, so
  `gap` is indicative rather than a like-for-like comparison.
- No repeatability or variance data: each row carries a single reported
  figure, with no error bars and no re-run.
