evaluation
261
- vs previous scan
- +7
- recent daily average
- 215.00
- vs its average
- +21%
- cumulative
- 5,343
- also updated (not counted above)
- 85
Which difficult benchmarks have been tested most
Each mark is one benchmark record. Height counts distinct models with numeric scores. Source model IDs preserve evaluated configurations; repeated score rows add no models. Source documents counts distinct cited reports or registry pages through the same evidence collection. Both counts appear on hover.
The side view projects reported scores and the selected counts onto the left wall, leaving out time. The gold staircase compares only benchmarks with verified score scales. Within the dated 2024+ cohort, a benchmark is Pareto when no other eligible benchmark has an equal or lower normalized score and an equal or greater selected count, with at least one strict advantage. Dates select the cohort; they do not determine dominance. Scores and counts are never multiplied. The score slice cannot promote a dominated benchmark onto the frontier.
Height uses log1p(count); ticks, tooltips and Pareto use raw counts. The axis covers the full scored 2024+ cohort and stays fixed while filtering. Only declared percentage metrics with a known selected count enter Pareto; lower-is-better percentages become 100 minus the original value. Hollow points flag an unverified score scale or an unknown count, which is never treated as zero. Overlapping caps are spaced apart; hover or focus traces their exact coordinates. Arrow keys move through benchmarks. A shared scale does not establish equivalent test protocols or prove a benchmark is solved.
The timeline starts on January 1, 2024, inclusive. Dates use the benchmark's release first, then its earliest numeric LLM score report. When a source dates score entries by model release, the earliest entry supplies a clearly labelled model-release proxy; this is not a verified score-publication date. Crawl timestamps and adoption-only mentions never supply dates. Known dates before 2024 are excluded. Records with no date remain individually visible in a labelled area.
No scored benchmarks match these filters.
Each source document counts once per benchmark record. Repeated scores or mentions in the same document add no citations. This table covers all sources.
Source documents from every registry. Expand a document to inspect the benchmarks it records and open the original evidence.
Big picture
See what shows up most often and where it came from. Open the connections view when you want to look at a specific item.
There are a lot of dots here. Pick one to see what it connects to, or open the matching results.
Scores over time
Every value that could be read verbatim from a cited document, placed at the date that document was published rather than at any evaluation date. The line directly links actual observations that set a new reported record among the values shown; it neither holds a score between reports nor extends past the final record. Test versions and run conditions can differ, so this is a reported-record path, not a like-for-like trend. When no newer number could be read, the gap is marked rather than drawn through. Whether a benchmark has saturated stays a reading you make, not a score this panel prints.
261
208
121
36
10
Category tags overlap. Each bar is an independent count, not a part of a stacked total.
Releases only excludes records re-announced as an update to something already surfaced.
Source mix counts ranked evidence after scoring. Fetch health counts raw records returned before scoring, so a source can be ok and still empty.
| Date | Coverage (UTC) | Evidence | Source mix | Categories | Events | Attention | Fetch health |
|---|
Dashboard unavailable
Try refreshing, or inspect the latest daily Issue while the dashboard rebuilds.
Open daily Issues ↗