Benchmark Radar™
RSS Star —

Leaderboard

Latest releases · 30 days

i

Ranking attention-ranking-v1: 55% GitHub stars, 30% Hugging Face paper upvotes, 15% Hugging Face dataset downloads, last 30 days, each normalized as log1p(value) / log1p(window maximum) and summed to a 0 to 100 score. A release is ranked only when enough of its weight comes from fresh, durable signals; the rest are listed with limited signals. Dataset downloads are a rolling 30-day figure, never a cumulative total. Stars come from the benchmark's own repository, never a parent framework. Window Sep 11, 2026 to Oct 11, 2026, UTC.

Recently released benchmarks, ranked by the attention they are receiving now: GitHub stars, Hugging Face paper upvotes and Hugging Face dataset downloads over the last 30 days.

  1. 01openclaw 2026.8.34released Oct 2, 202655low

    This is a gateway-only extended-stable release, which is our current equivalent to LTS. This release is OpenClaw from the end of August 2026, plus critical security updates, reliability and performance fixes, and features like new model support. The latest version of OpenClaw at the time of this release is 2026.9.7 2026.8.34 Highlights Extended-stable correctness rollup: backport 113 audit-selected fix units across upgrades, Doctor, authentication, sessions, channels, plugins, sandboxing, filesystem safety, model runtimes, and release packaging. Complete rescan: re-evaluate the full 2026.8.33 discovery range, large mixed-purpose pull requests, and the 2026.7.35 lineage instead of advancing only from the previous backport cursor. Upgrade and recovery hardening: preserve credentials, agent state, plugin inventory, transcripts, schedules, service ownership, and runtime links through upgrades, restarts, and Doctor repair. Boundary fixes: tighten remote filesystem mutation, plugin and skill scanning, browser authentication, channel reply identity, private skill ingress, and scoped runtime ownership. Changes Extended-stable release preparation: align OpenClaw, publishable plugins, native version metadata, generated channel catalogs, and npm package inventories to 2026.8.34. Fixes Plugins: Recover orphan installs without weakening ownership. Doctor: Repair keyed multi-agent rosters without ownership (#134706) Prevent device pairing migration crash on contaminated records. Doctor: Preserve per-agent memory search during upgrades. Doctor: Detect shared auth migration by provenance (#134808) Cron: Refuse stale doctor store rewrites. Gateway: Refresh model catalog after auth changes (#134361) Openai: Repair stale doctor route pins.

    GitHub stars
    391,667normalized 1.00source ↗weight 55%
    Hugging Face paper upvotes
    not observednormalized not scoredweight 30%
    Hugging Face dataset downloads, last 30 days
    not observednormalized not scoredweight 15%

    Coverage 0.55 · confidence low

  2. 02GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplayreleased Sep 23, 202654high

    Modern video games provide a measurable testbed for AI models, combining abilities of visual understanding, instruction decomposition, goal planning, and precise action control over multiple temporal horizons. Existing datasets and benchmarks, however, either cover a narrow range of games, lack language instructions, or rely on high-variance online rollouts. To address these challenges, we introduce GameHorizon, a unified data and evaluation suite that measures gameplay capabilities at different horizons for diverse model families. GameHorizon Suite consists of three components. First, GameHorizon-Annotator is a scalable and automated annotation pipeline for multi-horizon instructions. Second, utilizing the pipeline, we construct GameHorizon-Data, the first large-scale AAA gameplay dataset with temporally aligned videos, player actions, and multi-horizon instructions. It comprises 5,000 hours of recordings from 21 games, collected by 100 human expert players. Third, we build GameHorizon-Bench with reproducible offline and stepwise online testing. The offline track enables reproducible evaluation using thousands of standardized questions organized into three primary tasks and a series of diagnostic variants, while the online track tests whether offline scores reflect actual gameplay capabilities and localizes failures to specific steps within long-horizon gameplay. Based on our GameHorizon Suite, we evaluate 47 models through more than one million model invocations, revealing a meaningful hierarchy of task difficulty and pronounced differences in model capabilities. Our work can provide a standardized yardstick for evaluating gameplay capabilities across horizons and model families. We will release our dataset, annotator, and benchmark to facilitate future research.

    GitHub stars
    567normalized 0.49source ↗weight 55%
    Hugging Face paper upvotes
    132normalized 0.91source ↗weight 30%
    Hugging Face dataset downloads, last 30 days
    not observednormalized not scoredweight 15%

    Coverage 0.85 · confidence high

  3. 03OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialoguereleased Sep 29, 202651high

    We define OmniVChat (Omni Video Chat) as the task of native audio-visual dialogue between a user and an omni model. In OmniVChat, omni models directly and simultaneously receive audio and video from a user and return text. The user's query is embedded in the audio and video, without a separate text question, external captioning, or speech recognition. Direct audio-visual input reduces external latency and computation while preserving perceptual cues. However, research on OmniVChat faces two constraints: data availability and evaluation. Recordings of people using their own devices are scarce. Furthermore, good replies often depend on multimodal context and can be phrased in many ways, making keyword matching unreliable for evaluation. Recent progress in agent systems and video generation makes generation for comprehension viable, which means using synthesized dialogues for training and evaluation. Therefore, we present OmniVChat-Studio, a multi-agent data engine for synthesizing single- and multi-turn audio-visual dialogues. We use synthesized dialogues to build OmniVChat-Bench, an evaluation benchmark that evaluates omni models' basic dialogue abilities across five ability categories. Replies are judged by a large language model based on explicit scoring criteria. We also present OmniVChat-RL, a reinforcement learning reward design that jointly targets reply correctness, efficiency, and style in OmniVChat. Training Qwen3-Omni-Instruct with OmniVChat-RL on synthesized dialogues improves its performance on both OmniVChat-Bench and the human-recorded OmniVChat-Bench-Human. These gains validate the reward design and show transfer to real-world dialogues in training and evaluation.

    GitHub stars
    206normalized 0.41source ↗weight 55%
    Hugging Face paper upvotes
    153normalized 0.93source ↗weight 30%
    Hugging Face dataset downloads, last 30 days
    not observednormalized not scoredweight 15%

    Coverage 0.85 · confidence high

  4. 04Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidencereleased Sep 28, 202647high

    Latent visual reasoning (LVR) enables multimodal large language models (MLLMs) to perform intermediate computation in continuous latent tokens rather than expressing every reasoning step in words. However, unlike textual CoT, latent reasoning is not directly observable, making it difficult to supervise what latent tokens learn. In this work, we first conduct a thorough analysis of latent-token behavior and identify a latent evidence-credit gap: latent tokens respond only weakly to image perturbations that alter the correct answer. We hypothesize that this issue stems from the lack of explicit supervision during GRPO training. These findings suggest that a final-answer reward provides too little guidance on what visual evidence to preserve or how credit should be assigned across latent tokens. To bridge this gap, we propose ReaLVR, which brings visual-evidence supervision to the model's own free-running latent trajectories. ReaLVR contrasts correct and model-generated wrong answers to determine where stronger supervision is needed, and relevant and mismatched visual evidence to specify what to preserve. Across three model families, ReaLVR consistently outperforms evaluated LVR baselines, achieving the highest five-task average of 63.7% on Qwen2.5-VL-7B. Crucially, we are the first to scale visual reasoning in latent space, showing that our framework continues to deliver robust improvements at frontier model scales up to 235B. Further analyses show more question-sensitive latent-token positions, stronger alignment with relevant visual regions, and greater fixed-context dependence on the most attended latent tokens.

    GitHub stars
    47normalized 0.30source ↗weight 55%
    Hugging Face paper upvotes
    219normalized 1.00source ↗weight 30%
    Hugging Face dataset downloads, last 30 days
    not observednormalized not scoredweight 15%

    Coverage 0.85 · confidence high

  5. 05SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialoguereleased Sep 23, 202646high

    Long-term conversational memory in multi-party settings requires more than retrieving relevant content from long-term conversations: it must distinguish who said what, whom each statement concerns, how individuals perceive one another, what information is shared by the group, and how states change over time. Recent studies on multi-party dialogue benchmarks show that existing general-purpose LLM memory systems tend to lose person and group relations or struggle to integrate clues distributed across members, groups, and time. Together, these issues reveal two core bottlenecks: message attribution and relational understanding in multi-party dialogue, and state reconstruction from interleaved histories. To address both, we propose $\textbf{SpeakerMem-R1}$: its dual-track memory stores speaker-labeled verbatim messages and derived states organized into person-level and group-level views, then combines evidence from both tracks by entity, event, and time at query time. To reduce attribution and update errors during structured memory construction while enabling local deployment, we train Writer-R1 with SpeakerLevenshtein and speaker-conditioned GRPO. On GroupMemBench, SocialMemBench, and EverMemBench, SpeakerMem-R1 achieves binary accuracies of 47.9%, 69.2%, and 61.9%, respectively. On the publicly reported EverMemBench leaderboard from EverMind-AI, we achieves 62.33%, the best reported result among the latest state-of-the-art frameworks. It also achieves 70.85% on all 1,986 LoCoMo questions, which we use as a two-person long-term conversation boundary test. In a controlled evaluation of 305 questions, RL raises the SFT Writer's mean accuracy from 57.38% to 68.20%. We report both binary accuracy and token-F1, and ablations show that the verbatim and structured tracks, as well as person-level and group-level views, are complementary under the standardized evaluation interface.

    GitHub stars
    113normalized 0.37source ↗weight 55%
    Hugging Face paper upvotes
    102normalized 0.86source ↗weight 30%
    Hugging Face dataset downloads, last 30 days
    not observednormalized not scoredweight 15%

    Coverage 0.85 · confidence high

  6. 06From Traces to Agentic Worlds: Agentic Language World Models for Interactive Environment Simulationreleased Oct 5, 202643high

    Realistic environment replicas are increasingly valuable for training and evaluating LLM agents, yet the original systems may be inaccessible or impractical to reproduce. We explore agentic language world modeling: rather than rebuilding an executable environment, a world model agent serves as the environment for a task agent and supports faithful and stateful simulation. We instantiate this paradigm with Trace2Env, a learning-free framework for settings where the original system is unavailable but historical interaction traces remain accessible. Trace2Env reconstructs these traces into a reusable environment worldbook containing environment schemas, grounded evidence, and induced behavioral knowledge. At runtime, the world model agent actively consults the worldbook together with persistent episodic state to infer each action's observation and lasting state effects. Across nine environments, Trace2Env improves both next-observation fidelity and long-horizon interaction consistency over conventional prompt-based LWMs. In multi-turn interaction, task agent actions generated against Trace2Env remain valid more often when replayed in the real environment, indicating that its simulated dynamics better preserve the consequences of earlier actions across successive turns. These results establish agentic language world modeling as an alternative direction for building realistic environment replicas without reconstructing the original executable system.

    GitHub stars
    43normalized 0.29source ↗weight 55%
    Hugging Face paper upvotes
    129normalized 0.90source ↗weight 30%
    Hugging Face dataset downloads, last 30 days
    not observednormalized not scoredweight 15%

    Coverage 0.85 · confidence high

  7. 07UltraText Bench: A Comprehensive Bilingual Benchmark for Evaluating Visual Text Rendering in Image Generationreleased Oct 8, 202643high

    Dense visual text requires image generators to reproduce long strings across multiple regions with correct placement and legibility. As short-string rendering improves, evaluation must test sustained performance across more demanding scenes. We introduce UltraText Bench, a bilingual benchmark for prompt-only generation of dense visual text. It contains 432 prompts spanning 24 real-world scene categories and three difficulty levels, split equally between English and Chinese. Each human-reviewed prompt supplies exact strings for four to twelve text regions, paired with structured references for their content, placement, and visual attributes. We use the Q-Judger vision-language model to assess each image against the complete reference, reporting text fidelity, text clarity, spatial quality, and scene quality. Across 24 model configurations, these dimensions reveal different strengths: Z-Image-Turbo gains 3.81 clarity points over Z-Image-Base while losing 14.76 fidelity points under the reported settings. Performance also varies with workload; Qwen-Image-2512's English composite falls from 86.50 at L1 to 42.86 at L3. Ten participants took part in human evaluation of the automatic scores. Repository: https://github.com/LINs-lab/UltraText_Bench.

    GitHub stars
    34normalized 0.28source ↗weight 55%
    Hugging Face paper upvotes
    153normalized 0.93source ↗weight 30%
    Hugging Face dataset downloads, last 30 days
    not observednormalized not scoredweight 15%

    Coverage 0.85 · confidence high

  8. 08UniWAM: Unified World-Action Modelreleased Oct 1, 202643high

    Vision-language-action models benefit from the understanding and reasoning capabilities of pretrained vision-language models, but action-only supervision provides limited grounding in world dynamics. Conversely, world-action models inherit spatiotemporal priors from video generation models, yet remain limited in semantic understanding and reasoning under distribution shifts. We introduce UniWAM, a unified architecture that integrates a physical reasoner, a world generator, and an action predictor to jointly learn semantic understanding of the physical world, visual generation, and action prediction. To ensure the quality of the training data, we developed a rigorous data cleaning and annotation pipeline for both human egocentric data and robot data. To adapt the vision-language component to embodied tasks while preserving its inherited language capabilities, we represent low-level actions in natural language and introduce a pre-training recipe that assigns complementary supervision from visual question answering (VQA) data, human egocentric data, and robot demonstrations to the appropriate model components. During post-training, future visual noise augmentation reduces reliance on precise future predictions, while history-conditioned flow matching uses encoded action history to initialize action generation. Together, these designs significantly reduce denoising steps while maintaining performance. UniWAM achieves state-of-the-art (SOTA) performance across multiple evaluations, including in-distribution performance, robustness, generalization, instruction following, and long-horizon task execution. Furthermore, we uncover a log-linear scaling law of unified human-robot co-training, demonstrating the effectiveness of large-scale pre-training on a mixture of human and robot data.

    GitHub stars
    112normalized 0.37source ↗weight 55%
    Hugging Face paper upvotes
    55normalized 0.75source ↗weight 30%
    Hugging Face dataset downloads, last 30 days
    not observednormalized not scoredweight 15%

    Coverage 0.85 · confidence high

  9. 09Introducing Jev in DeepEvalreleased Sep 22, 202642low

    New to deepeval ? Get started here. 🧠 Introducing Jev: System One for LLM Evals Every LLM-as-a-judge metric has had the same awkward secret: somewhere inside the thing measuring your LLM is another LLM generating text. Closed-ended verdicts ( yes / no / borderline ) were treated like writing tasks, and you paid for it in variance, latency, and cost. deepeval 4.2.2 introduces support for Jev , TypeSafe AI's System One model. Jev is not a language model. It does not generate text. You give it state and a bounded question, and it returns a typed decision with calibrated probabilities. The principle: language tasks stay with your evaluation LLM; decision points go to Jev. What changes when Jev is on Less flaky — verdicts become probabilities mapped through fixed thresholds. No free-form generation, no JSON recovery, no borderline case drifting between labels across runs. Cheaper — $0.042 per million input tokens, no output token charge. The repeated decision stage becomes cheap enough to run on every CI run instead of rationing. Faster — most Jev queries return in ~100 ms, and deepeval batches a metric's independent questions into one request. Where it plugs in | Metric | LLM still handles | Jev now handles | | --- | --- | --- | | FaithfulnessMetric (and other QAG metrics) | Extracting claims and truths | One Noul ( P(yes) ) per claim | | GEval | Generating evaluation steps | Noul in strict / rubric-free mode, Score when a rubric is provided | | DAGMetric | Task node text | Noul for binary nodes, Choice for non-binary nodes | | Classifiers | — | Choice over the closed label set | Score calculations are unchanged. Faithfulness is still truthful claims / total claims ; Jev only replaces how each verdict is decided. Try it (opt-in, experimental) Nothing changes for existing users. Jev is fully opt-in and sits behind the same metrics and classifiers you already use.

    GitHub stars
    18,740normalized 0.76source ↗weight 55%
    Hugging Face paper upvotes
    not observednormalized not scoredweight 30%
    Hugging Face dataset downloads, last 30 days
    not observednormalized not scoredweight 15%

    Coverage 0.55 · confidence low

  10. 10Does Learning Protein Folding Generalize to Broader Reasoning?released Sep 30, 202641high

    Large language models rely heavily on human text, which often conveys surface answers rather than the spatial and structural logic behind them. Protein folding is a natural testbed, because one solved structure yields thousands of exactly checkable spatial and topological statements. We ask: can learning to fold proteins teach general models reusable reasoning capabilities? To answer this, we build FoldingCorpus, a protein-derived question-answer dataset, and Fold2Reason, a recipe that post-trains on it through two complementary signals: discrete structural answers predicted via the model's native language head, and continuous 3D geometry decoded from the same shared representations. On FoldBench, Fold2Reason achieves structure prediction scores 2.7 to 3.5 times those of Qwen3.5-9B. Beyond protein structure prediction, it improves performance on all 10 benchmarks spanning spatial, graph, scientific, and general reasoning, raising macro-average accuracy from 45.09% to 48.33% (+3.23 pp), with positive gains on all 10 benchmarks, while matched controls built from random, synthetic, and shuffled structure yield substantially smaller or negative gains. Our work shows that non-linguistic, structure-dense scientific data can systematically improve broad reasoning in language models, making a solved scientific problem a practical source of post-training supervision.

    GitHub stars
    65normalized 0.33source ↗weight 55%
    Hugging Face paper upvotes
    61normalized 0.77source ↗weight 30%
    Hugging Face dataset downloads, last 30 days
    not observednormalized not scoredweight 15%

    Coverage 0.85 · confidence high

375 of 1,030 releases in this window are ranked; the rest are listed with limited signals. Data through Oct 11, 2026 · average signal coverage 32%.