Daily AI benchmark brief: 2026-09-18
The new SonoCorpus and SonoBase release pairs 456,963 ultrasound images and 1,626,085 expert segmentation masks with an interactive segmentation model…
The new SonoCorpus and SonoBase release pairs 456,963 ultrasound images and 1,626,085 expert segmentation masks with an interactive segmentation model…
The new lexEN benchmark replaces disputed word-sense labels with a conservative, human-adjudicated correction layer: 211 labels changed and 56 removed…
A new audit of 254 SWE-bench coding-agent submissions reports that the leaderboard cannot reliably order its leading entries: the top two resolve the same…
The new Sophea release evaluates a production Greek-English speech recognizer against nine simultaneous gates covering both languages, language…
The new KNOWS Benchmark evaluates web agents on 110 tasks that combine web research with producing documents, spreadsheets, and slide decks in Google…
The newly captured MusicAI background-music benchmark tests audio-capable large language models on 2,000 questions under clean audio, white noise, and 55…
The newly released Benchmark Radar combines a searchable catalog of AI evaluations with links to datasets and code, mentions in model cards and technical…
The new oncology visual question answering benchmark builds test items from private, single-institution radiology reports paired with three-dimensional…
The new “Double Measurement Confound” paper identifies a recurring agent-evaluation problem in this feed: a fixed software scaffold may make execution…
The new “Style Over Substance” study tests automated safety judges by preserving an answer byte-for-byte while adding tone-only wrappers, such as…
The new URL Intelligence Benchmark evaluates agents and web-analysis tools on operational hazards that ordinary answer-accuracy tests miss: redirects…
TruthInsightBench is a new benchmark for scientific-discovery agents that replaces reproduction-style tasks with 40 blind tasks from peer-reviewed…
The new Amharic Automatic Speech Recognition benchmark evaluates open models on 1,548 clips that its publisher says were unavailable for training. It…
Several new releases in this captured feed turn agent evaluation into domain-specific workflow testing. The clearest is…
A new preregistered audit, “Clean Engineering, Unstable Measurement,” tests whether a black-box large language model judge behaves like a stable…
Among today’s captured releases, EarlyEval introduces early outcome prediction: it estimates an agent’s final result from intermediate behavior and stops…
The newly released InSight benchmark tests agents that must actively interact with visualizations to verify 21,349 claims, rather than answer once from a…
EleutherAI released Language Model Evaluation Harness v0.4.13 with fixes that can change prior scores: test questions could leak into their own few-shot…
PCFBench reflects a recurring push in this captured feed to inspect an agent’s process rather than only its final answer. It separately tests…
The new INSIDER LLM Detection Benchmark evaluates models that may take harmful actions by comparing the model’s self-reported action log with an…
The new NBPO benchmark-generations dataset publishes every decoded model response used in three judge-based comparisons, allowing another evaluator to…
The new Same Model, Different Harness study holds the coding model and tasks fixed while changing how the agent harness manages conversation history and…
OpenCompass v0.5.4 is a substantive harness update, adding native VLMEvalKit-based multimodal evaluation, multi-round inference with the Multi-IF…
Across this captured feed, three new agent benchmarks make the evaluated unit an interactive model-plus-runtime system rather than a final answer…
New release SUSVIBES evaluates 12 coding-agent settings on 186 real-world feature requests for which human developers previously committed vulnerable…
No material GPT insight: No category moved far enough, persistently enough, or across enough independent sources to support a decision-useful finding in…
No material GPT insight: No category moved far enough, persistently enough, or across enough independent sources to support a material finding in today’s…
No material GPT insight: No material pattern was supported: the captured items did not show a sufficiently large, persistent, cross-source shift. Only 19…
No material GPT insight: No material pattern cleared the feed’s persistence and cross-source thresholds today. Only 65 of 198 corpus evidence records were…
No material GPT insight: No category changed far enough, persistently enough, and across enough sources to support a decision-useful pattern in this…