Benchmark Radar™
RSS Contact Star

What changed in AI evaluation, and why it matters

  1. Daily AI benchmark brief: 2026-09-18

    The new SonoCorpus and SonoBase release pairs 456,963 ultrasound images and 1,626,085 expert segmentation masks with an interactive segmentation model…

  2. Daily AI benchmark brief: 2026-09-17

    The new lexEN benchmark replaces disputed word-sense labels with a conservative, human-adjudicated correction layer: 211 labels changed and 56 removed…

  3. Daily AI benchmark brief: 2026-09-16

    A new audit of 254 SWE-bench coding-agent submissions reports that the leaderboard cannot reliably order its leading entries: the top two resolve the same…

  4. Daily AI benchmark brief: 2026-09-15

    The new Sophea release evaluates a production Greek-English speech recognizer against nine simultaneous gates covering both languages, language…

  5. Daily AI benchmark brief: 2026-09-14

    The new KNOWS Benchmark evaluates web agents on 110 tasks that combine web research with producing documents, spreadsheets, and slide decks in Google…

  6. Daily AI benchmark brief: 2026-09-13

    The newly captured MusicAI background-music benchmark tests audio-capable large language models on 2,000 questions under clean audio, white noise, and 55…

  7. Daily AI benchmark brief: 2026-09-12

    The newly released Benchmark Radar combines a searchable catalog of AI evaluations with links to datasets and code, mentions in model cards and technical…

  8. Daily AI benchmark brief: 2026-09-11

    The new oncology visual question answering benchmark builds test items from private, single-institution radiology reports paired with three-dimensional…

  9. Daily AI benchmark brief: 2026-09-10

    The new “Double Measurement Confound” paper identifies a recurring agent-evaluation problem in this feed: a fixed software scaffold may make execution…

  10. Daily AI benchmark brief: 2026-09-09

    The new “Style Over Substance” study tests automated safety judges by preserving an answer byte-for-byte while adding tone-only wrappers, such as…

  11. Daily AI benchmark brief: 2026-09-08

    The new URL Intelligence Benchmark evaluates agents and web-analysis tools on operational hazards that ordinary answer-accuracy tests miss: redirects…

  12. Daily AI benchmark brief: 2026-09-07

    TruthInsightBench is a new benchmark for scientific-discovery agents that replaces reproduction-style tasks with 40 blind tasks from peer-reviewed…

  13. Daily AI benchmark brief: 2026-09-06

    The new Amharic Automatic Speech Recognition benchmark evaluates open models on 1,548 clips that its publisher says were unavailable for training. It…

  14. Daily AI benchmark brief: 2026-09-04

    A new preregistered audit, “Clean Engineering, Unstable Measurement,” tests whether a black-box large language model judge behaves like a stable…

  15. Daily AI benchmark brief: 2026-09-03

    Among today’s captured releases, EarlyEval introduces early outcome prediction: it estimates an agent’s final result from intermediate behavior and stops…

  16. Daily AI benchmark brief: 2026-09-02

    The newly released InSight benchmark tests agents that must actively interact with visualizations to verify 21,349 claims, rather than answer once from a…

  17. Daily AI benchmark brief: 2026-09-01

    EleutherAI released Language Model Evaluation Harness v0.4.13 with fixes that can change prior scores: test questions could leak into their own few-shot…

  18. Daily AI benchmark brief: 2026-08-31

    PCFBench reflects a recurring push in this captured feed to inspect an agent’s process rather than only its final answer. It separately tests…

  19. Daily AI benchmark brief: 2026-08-30

    The new INSIDER LLM Detection Benchmark evaluates models that may take harmful actions by comparing the model’s self-reported action log with an…

  20. Daily AI benchmark brief: 2026-08-29

    The new NBPO benchmark-generations dataset publishes every decoded model response used in three judge-based comparisons, allowing another evaluator to…

  21. Daily AI benchmark brief: 2026-08-28

    The new Same Model, Different Harness study holds the coding model and tasks fixed while changing how the agent harness manages conversation history and…

  22. Daily AI benchmark brief: 2026-08-27

    OpenCompass v0.5.4 is a substantive harness update, adding native VLMEvalKit-based multimodal evaluation, multi-round inference with the Multi-IF…

  23. Daily AI benchmark brief: 2026-08-26

    Across this captured feed, three new agent benchmarks make the evaluated unit an interactive model-plus-runtime system rather than a final answer…

  24. Daily AI benchmark brief: 2026-08-25

    New release SUSVIBES evaluates 12 coding-agent settings on 186 real-world feature requests for which human developers previously committed vulnerable…

  25. Daily AI benchmark brief: 2026-08-24

    No material GPT insight: No category moved far enough, persistently enough, or across enough independent sources to support a decision-useful finding in…

  26. Daily AI benchmark brief: 2026-08-23

    No material GPT insight: No category moved far enough, persistently enough, or across enough independent sources to support a material finding in today’s…

  27. Daily AI benchmark brief: 2026-08-22

    No material GPT insight: No material pattern was supported: the captured items did not show a sufficiently large, persistent, cross-source shift. Only 19…

  28. Daily AI benchmark brief: 2026-08-21

    No material GPT insight: No material pattern cleared the feed’s persistence and cross-source thresholds today. Only 65 of 198 corpus evidence records were…

  29. Daily AI benchmark brief: 2026-08-20

    No material GPT insight: No category changed far enough, persistently enough, and across enough sources to support a decision-useful pattern in this…