cff-version: 1.2.0
message: "If you use this dataset, please cite it as below."
title: "AI Agent Benchmarks — Saturation Tracker"
abstract: >-
  A curated catalog of AI agent (Model + Harness/scaffold) evaluation
  benchmarks: what each measures, the current state of the art, the remaining
  headroom to the metric ceiling, the gap to the human baseline where one
  exists, and a saturation band. Covers 69 benchmarks and meta-leaderboards as
  of the August 2026 snapshot. Curated reference, not a leaderboard of record —
  read methodology.html for what is and is not measured before citing.
type: dataset
version: "2026.08.0"
date-released: "2026-08-22"
license: CC-BY-4.0
url: "https://benchmarks.santismm.com/"
repository-code: "https://github.com/santismm/benchmarks"
authors:
  - family-names: "Santa María Morales"
    given-names: "Santiago"
    website: "https://santismm.com"
keywords:
  - AI agents
  - benchmarks
  - evaluation
  - saturation
  - harness engineering
  - state of the art
