Search AI, LLM, agent, multimodal, and AI-for-science benchmarks
Today's radar
Benchmarks with this name
Daily briefing
Questions for today
Leaderboard
Latest releases · 30 days
i
Ranking attention-ranking-v1: 55% GitHub stars, 30% Hugging Face paper upvotes, 15% Hugging Face dataset downloads, last 30 days, each normalized as log1p(value) / log1p(window maximum) and summed to a 0 to 100 score. A release is ranked only when enough of its weight comes from fresh, durable signals; the rest are listed with limited signals. Dataset downloads are a rolling 30-day figure, never a cumulative total. Stars come from the benchmark's own repository, never a parent framework. Window Sep 11, 2026 to Oct 11, 2026, UTC.
Recently released benchmarks, ranked by the attention they are receiving now: GitHub stars, Hugging Face paper upvotes and Hugging Face dataset downloads over the last 30 days.
01openclaw 2026.8.34released Oct 2, 202655low
This is a gateway-only extended-stable release, which is our current equivalent to LTS. This release is OpenClaw from the end of August 2026, plus critical security updates, reliability and performance fixes, and features like new model support. The latest version of OpenClaw at the time of this release is 2026.9.7 2026.8.34 Highlights Extended-stable correctness rollup: backport 113 audit-selected fix units across upgrades, Doctor, authentication, sessions, channels, plugins, sandboxing, filesystem safety, model runtimes, and release packaging. Complete rescan: re-evaluate the full 2026.8.33 discovery range, large mixed-purpose pull requests, and the 2026.7.35 lineage instead of advancing only from the previous backport cursor. Upgrade and recovery hardening: preserve credentials, agent state, plugin inventory, transcripts, schedules, service ownership, and runtime links through upgrades, restarts, and Doctor repair. Boundary fixes: tighten remote filesystem mutation, plugin and skill scanning, browser authentication, channel reply identity, private skill ingress, and scoped runtime ownership. Changes Extended-stable release preparation: align OpenClaw, publishable plugins, native version metadata, generated channel catalogs, and npm package inventories to 2026.8.34. Fixes Plugins: Recover orphan installs without weakening ownership. Doctor: Repair keyed multi-agent rosters without ownership (#134706) Prevent device pairing migration crash on contaminated records. Doctor: Preserve per-agent memory search during upgrades. Doctor: Detect shared auth migration by provenance (#134808) Cron: Refuse stale doctor store rewrites. Gateway: Refresh model catalog after auth changes (#134361) Openai: Repair stale doctor route pins.
- GitHub stars
- 391,667normalized 1.00source ↗weight 55%
- Hugging Face paper upvotes
- not observednormalized not scoredweight 30%
- Hugging Face dataset downloads, last 30 days
- not observednormalized not scoredweight 15%
02GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplayreleased Sep 23, 202654high
Modern video games provide a measurable testbed for AI models, combining abilities of visual understanding, instruction decomposition, goal planning, and precise action control over multiple temporal horizons. Existing datasets and benchmarks, however, either cover a narrow range of games, lack language instructions, or rely on high-variance online rollouts. To address these challenges, we introduce GameHorizon, a unified data and evaluation suite that measures gameplay capabilities at different horizons for diverse model families. GameHorizon Suite consists of three components. First, GameHorizon-Annotator is a scalable and automated annotation pipeline for multi-horizon instructions. Second, utilizing the pipeline, we construct GameHorizon-Data, the first large-scale AAA gameplay dataset with temporally aligned videos, player actions, and multi-horizon instructions. It comprises 5,000 hours of recordings from 21 games, collected by 100 human expert players. Third, we build GameHorizon-Bench with reproducible offline and stepwise online testing. The offline track enables reproducible evaluation using thousands of standardized questions organized into three primary tasks and a series of diagnostic variants, while the online track tests whether offline scores reflect actual gameplay capabilities and localizes failures to specific steps within long-horizon gameplay. Based on our GameHorizon Suite, we evaluate 47 models through more than one million model invocations, revealing a meaningful hierarchy of task difficulty and pronounced differences in model capabilities. Our work can provide a standardized yardstick for evaluating gameplay capabilities across horizons and model families. We will release our dataset, annotator, and benchmark to facilitate future research.
03OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialoguereleased Sep 29, 202651high
We define OmniVChat (Omni Video Chat) as the task of native audio-visual dialogue between a user and an omni model. In OmniVChat, omni models directly and simultaneously receive audio and video from a user and return text. The user's query is embedded in the audio and video, without a separate text question, external captioning, or speech recognition. Direct audio-visual input reduces external latency and computation while preserving perceptual cues. However, research on OmniVChat faces two constraints: data availability and evaluation. Recordings of people using their own devices are scarce. Furthermore, good replies often depend on multimodal context and can be phrased in many ways, making keyword matching unreliable for evaluation. Recent progress in agent systems and video generation makes generation for comprehension viable, which means using synthesized dialogues for training and evaluation. Therefore, we present OmniVChat-Studio, a multi-agent data engine for synthesizing single- and multi-turn audio-visual dialogues. We use synthesized dialogues to build OmniVChat-Bench, an evaluation benchmark that evaluates omni models' basic dialogue abilities across five ability categories. Replies are judged by a large language model based on explicit scoring criteria. We also present OmniVChat-RL, a reinforcement learning reward design that jointly targets reply correctness, efficiency, and style in OmniVChat. Training Qwen3-Omni-Instruct with OmniVChat-RL on synthesized dialogues improves its performance on both OmniVChat-Bench and the human-recorded OmniVChat-Bench-Human. These gains validate the reward design and show transfer to real-world dialogues in training and evaluation.
04Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidencereleased Sep 28, 202647high
Latent visual reasoning (LVR) enables multimodal large language models (MLLMs) to perform intermediate computation in continuous latent tokens rather than expressing every reasoning step in words. However, unlike textual CoT, latent reasoning is not directly observable, making it difficult to supervise what latent tokens learn. In this work, we first conduct a thorough analysis of latent-token behavior and identify a latent evidence-credit gap: latent tokens respond only weakly to image perturbations that alter the correct answer. We hypothesize that this issue stems from the lack of explicit supervision during GRPO training. These findings suggest that a final-answer reward provides too little guidance on what visual evidence to preserve or how credit should be assigned across latent tokens. To bridge this gap, we propose ReaLVR, which brings visual-evidence supervision to the model's own free-running latent trajectories. ReaLVR contrasts correct and model-generated wrong answers to determine where stronger supervision is needed, and relevant and mismatched visual evidence to specify what to preserve. Across three model families, ReaLVR consistently outperforms evaluated LVR baselines, achieving the highest five-task average of 63.7% on Qwen2.5-VL-7B. Crucially, we are the first to scale visual reasoning in latent space, showing that our framework continues to deliver robust improvements at frontier model scales up to 235B. Further analyses show more question-sensitive latent-token positions, stronger alignment with relevant visual regions, and greater fixed-context dependence on the most attended latent tokens.
05SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialoguereleased Sep 23, 202646high
Long-term conversational memory in multi-party settings requires more than retrieving relevant content from long-term conversations: it must distinguish who said what, whom each statement concerns, how individuals perceive one another, what information is shared by the group, and how states change over time. Recent studies on multi-party dialogue benchmarks show that existing general-purpose LLM memory systems tend to lose person and group relations or struggle to integrate clues distributed across members, groups, and time. Together, these issues reveal two core bottlenecks: message attribution and relational understanding in multi-party dialogue, and state reconstruction from interleaved histories. To address both, we propose $\textbf{SpeakerMem-R1}$: its dual-track memory stores speaker-labeled verbatim messages and derived states organized into person-level and group-level views, then combines evidence from both tracks by entity, event, and time at query time. To reduce attribution and update errors during structured memory construction while enabling local deployment, we train Writer-R1 with SpeakerLevenshtein and speaker-conditioned GRPO. On GroupMemBench, SocialMemBench, and EverMemBench, SpeakerMem-R1 achieves binary accuracies of 47.9%, 69.2%, and 61.9%, respectively. On the publicly reported EverMemBench leaderboard from EverMind-AI, we achieves 62.33%, the best reported result among the latest state-of-the-art frameworks. It also achieves 70.85% on all 1,986 LoCoMo questions, which we use as a two-person long-term conversation boundary test. In a controlled evaluation of 305 questions, RL raises the SFT Writer's mean accuracy from 57.38% to 68.20%. We report both binary accuracy and token-F1, and ablations show that the verbatim and structured tracks, as well as person-level and group-level views, are complementary under the standardized evaluation interface.
06From Traces to Agentic Worlds: Agentic Language World Models for Interactive Environment Simulationreleased Oct 5, 202643high
Realistic environment replicas are increasingly valuable for training and evaluating LLM agents, yet the original systems may be inaccessible or impractical to reproduce. We explore agentic language world modeling: rather than rebuilding an executable environment, a world model agent serves as the environment for a task agent and supports faithful and stateful simulation. We instantiate this paradigm with Trace2Env, a learning-free framework for settings where the original system is unavailable but historical interaction traces remain accessible. Trace2Env reconstructs these traces into a reusable environment worldbook containing environment schemas, grounded evidence, and induced behavioral knowledge. At runtime, the world model agent actively consults the worldbook together with persistent episodic state to infer each action's observation and lasting state effects. Across nine environments, Trace2Env improves both next-observation fidelity and long-horizon interaction consistency over conventional prompt-based LWMs. In multi-turn interaction, task agent actions generated against Trace2Env remain valid more often when replayed in the real environment, indicating that its simulated dynamics better preserve the consequences of earlier actions across successive turns. These results establish agentic language world modeling as an alternative direction for building realistic environment replicas without reconstructing the original executable system.
07UltraText Bench: A Comprehensive Bilingual Benchmark for Evaluating Visual Text Rendering in Image Generationreleased Oct 8, 202643high
Dense visual text requires image generators to reproduce long strings across multiple regions with correct placement and legibility. As short-string rendering improves, evaluation must test sustained performance across more demanding scenes. We introduce UltraText Bench, a bilingual benchmark for prompt-only generation of dense visual text. It contains 432 prompts spanning 24 real-world scene categories and three difficulty levels, split equally between English and Chinese. Each human-reviewed prompt supplies exact strings for four to twelve text regions, paired with structured references for their content, placement, and visual attributes. We use the Q-Judger vision-language model to assess each image against the complete reference, reporting text fidelity, text clarity, spatial quality, and scene quality. Across 24 model configurations, these dimensions reveal different strengths: Z-Image-Turbo gains 3.81 clarity points over Z-Image-Base while losing 14.76 fidelity points under the reported settings. Performance also varies with workload; Qwen-Image-2512's English composite falls from 86.50 at L1 to 42.86 at L3. Ten participants took part in human evaluation of the automatic scores. Repository: https://github.com/LINs-lab/UltraText_Bench.
08UniWAM: Unified World-Action Modelreleased Oct 1, 202643high
Vision-language-action models benefit from the understanding and reasoning capabilities of pretrained vision-language models, but action-only supervision provides limited grounding in world dynamics. Conversely, world-action models inherit spatiotemporal priors from video generation models, yet remain limited in semantic understanding and reasoning under distribution shifts. We introduce UniWAM, a unified architecture that integrates a physical reasoner, a world generator, and an action predictor to jointly learn semantic understanding of the physical world, visual generation, and action prediction. To ensure the quality of the training data, we developed a rigorous data cleaning and annotation pipeline for both human egocentric data and robot data. To adapt the vision-language component to embodied tasks while preserving its inherited language capabilities, we represent low-level actions in natural language and introduce a pre-training recipe that assigns complementary supervision from visual question answering (VQA) data, human egocentric data, and robot demonstrations to the appropriate model components. During post-training, future visual noise augmentation reduces reliance on precise future predictions, while history-conditioned flow matching uses encoded action history to initialize action generation. Together, these designs significantly reduce denoising steps while maintaining performance. UniWAM achieves state-of-the-art (SOTA) performance across multiple evaluations, including in-distribution performance, robustness, generalization, instruction following, and long-horizon task execution. Furthermore, we uncover a log-linear scaling law of unified human-robot co-training, demonstrating the effectiveness of large-scale pre-training on a mixture of human and robot data.
09Introducing Jev in DeepEvalreleased Sep 22, 202642low
New to deepeval ? Get started here. 🧠 Introducing Jev: System One for LLM Evals Every LLM-as-a-judge metric has had the same awkward secret: somewhere inside the thing measuring your LLM is another LLM generating text. Closed-ended verdicts ( yes / no / borderline ) were treated like writing tasks, and you paid for it in variance, latency, and cost. deepeval 4.2.2 introduces support for Jev , TypeSafe AI's System One model. Jev is not a language model. It does not generate text. You give it state and a bounded question, and it returns a typed decision with calibrated probabilities. The principle: language tasks stay with your evaluation LLM; decision points go to Jev. What changes when Jev is on Less flaky — verdicts become probabilities mapped through fixed thresholds. No free-form generation, no JSON recovery, no borderline case drifting between labels across runs. Cheaper — $0.042 per million input tokens, no output token charge. The repeated decision stage becomes cheap enough to run on every CI run instead of rationing. Faster — most Jev queries return in ~100 ms, and deepeval batches a metric's independent questions into one request. Where it plugs in | Metric | LLM still handles | Jev now handles | | --- | --- | --- | | FaithfulnessMetric (and other QAG metrics) | Extracting claims and truths | One Noul ( P(yes) ) per claim | | GEval | Generating evaluation steps | Noul in strict / rubric-free mode, Score when a rubric is provided | | DAGMetric | Task node text | Noul for binary nodes, Choice for non-binary nodes | | Classifiers | — | Choice over the closed label set | Score calculations are unchanged. Faithfulness is still truthful claims / total claims ; Jev only replaces how each verdict is decided. Try it (opt-in, experimental) Nothing changes for existing users. Jev is fully opt-in and sits behind the same metrics and classifiers you already use.
- GitHub stars
- 18,740normalized 0.76source ↗weight 55%
- Hugging Face paper upvotes
- not observednormalized not scoredweight 30%
- Hugging Face dataset downloads, last 30 days
- not observednormalized not scoredweight 15%
10Does Learning Protein Folding Generalize to Broader Reasoning?released Sep 30, 202641high
Large language models rely heavily on human text, which often conveys surface answers rather than the spatial and structural logic behind them. Protein folding is a natural testbed, because one solved structure yields thousands of exactly checkable spatial and topological statements. We ask: can learning to fold proteins teach general models reusable reasoning capabilities? To answer this, we build FoldingCorpus, a protein-derived question-answer dataset, and Fold2Reason, a recipe that post-trains on it through two complementary signals: discrete structural answers predicted via the model's native language head, and continuous 3D geometry decoded from the same shared representations. On FoldBench, Fold2Reason achieves structure prediction scores 2.7 to 3.5 times those of Qwen3.5-9B. Beyond protein structure prediction, it improves performance on all 10 benchmarks spanning spatial, graph, scientific, and general reasoning, raising macro-average accuracy from 45.09% to 48.33% (+3.23 pp), with positive gains on all 10 benchmarks, while matched controls built from random, synthetic, and shuffled structure yield substantially smaller or negative gains. Our work shows that non-linguistic, structure-dense scientific data can systematically improve broad reasoning in language models, making a solved scientific problem a practical source of post-training supervision.
375 of 1,030 releases in this window are ranked; the rest are listed with limited signals. Data through Oct 11, 2026 · average signal coverage 32%.
Benchmark Frontier
Which difficult benchmarks have been tested most
Each mark is one benchmark record. Height counts distinct models with numeric scores. Source model IDs preserve evaluated configurations; repeated score rows add no models. Source documents counts distinct cited reports or registry pages through the same evidence collection. Both counts appear on hover.
The side view projects reported scores and the selected counts onto the left wall, leaving out time. The gold staircase compares only benchmarks with verified score scales. Within the dated 2024+ cohort, a benchmark is Pareto when no other eligible benchmark has an equal or lower normalized score and an equal or greater selected count, with at least one strict advantage. Dates select the cohort; they do not determine dominance. Scores and counts are never multiplied. The score slice cannot promote a dominated benchmark onto the frontier.
Height uses log1p(count); ticks, tooltips and Pareto use raw counts. The axis covers the full scored 2024+ cohort and stays fixed while filtering. Only declared percentage metrics with a known selected count enter Pareto; lower-is-better percentages become 100 minus the original value. Hollow points flag an unverified score scale or an unknown count, which is never treated as zero. Overlapping caps are spaced apart; hover or focus traces their exact coordinates. Arrow keys move through benchmarks. A shared scale does not establish equivalent test protocols or prove a benchmark is solved.
The timeline starts on January 1, 2024, inclusive. Dates use the benchmark's release first, then its earliest numeric LLM score report. When a source dates score entries by model release, the earliest entry supplies a clearly labelled model-release proxy; this is not a verified score-publication date. Crawl timestamps and adoption-only mentions never supply dates. Known dates before 2024 are excluded. Records with no date remain individually visible in a labelled area.
- 01577 data points
- 02577 data points
- 03510 data points
- 04492 data points
- 05489 data points
No scored benchmarks match these filters.
Most documented benchmarks
Counts distinct source documents that record each benchmark. Model reports and registry pages use the same rule. Each document counts once per benchmark record. This measures documentation coverage, not benchmark quality.
- 01GPQA Diamond29 source documents
- 02Humanity's Last Exam22 source documents
- 03Terminal-Bench21 source documents
- 04SWE-bench Verified19 source documents
- 05AIME18 source documents
i
Counts distinct source documents that record each benchmark. Model reports and registry pages use the same rule. Each document counts once per benchmark record. This measures documentation coverage, not benchmark quality. Open the source-document list below to trace each count to its citations.
What the evidence shows Stated findings
Benchmarks by source documents
Each source document counts once per benchmark record. Repeated scores or mentions in the same document add no citations. This table covers all sources.
Audit the counts Source documents in the catalog
Source documents from every registry. Expand a document to inspect the benchmarks it records and open the original evidence.
Big picture
What we found
See what shows up most often and where it came from. Open the connections view when you want to look at a specific item.
Want more detail? See how everything connects
There are a lot of dots here. Pick one to see what it connects to, or open the matching results.
Saturation
Scores over time
Benchmark reported scores over time
How to read this chart
Every value that could be read verbatim from a cited document, placed at the date that document was published rather than at any evaluation date. The line directly links actual observations that set a new reported record among the values shown; it neither holds a score between reports nor extends past the final record. Test versions and run conditions can differ, so this is a reported-record path, not a like-for-like trend. When no newer number could be read, the gap is marked rather than drawn through. Whether a benchmark has saturated stays a reading you make, not a score this panel prints.
Trend
New by domain
Daily evidence and attention volume
Category tags overlap. Each bar is an independent count, not a part of a stacked total.
Releases only excludes records re-announced as an update to something already surfaced.
Dev checker
Source mix counts ranked evidence after scoring. Fetch health counts raw records returned before scoring, so a source can be ok and still empty.
| Date | Coverage (UTC) | Evidence | Source mix | Categories | Events | Attention | Fetch health |
|---|
Dashboard unavailable
The validated data file could not be loaded.
Try refreshing, or inspect the latest daily Issue while the dashboard rebuilds.
Open daily Issues ↗