← All Radar guides

Radar

Benchmarks

Datasets and leaderboards shape what the industry optimizes for. This changelog tracks the ones that actually moved training, marketing, or model selection - from ImageNet to MMLU, Arena Elo, and SWE-bench.

10 entries

  1. benchmarklanguage

    Terminal-Bench 2.1

    Agentic coding and long-horizon CLI task suite, setting the bar for frontier reasoning models.

    Why it matteredIt measures how well models act as autonomous engineers rather than just code completers.

    ImpactGPT-5.6 Sol Ultra hit 91.9%, shifting the conversation from 'can it code' to 'can it operate'.

    Coding Agent Evals>90%
    • coding
    • agents
    Source →
  2. benchmarklanguage

    BrowseComp

    Long-horizon, high-difficulty information seeking and browser agent benchmark.

    Why it matteredCrucial for evaluating multi-step reasoning and tool use.

    ImpactKimi K3 scored 91.2, proving open-weights can match closed labs in agentic tooling.

    • agents
    • tool-use
    Source →
  3. benchmarklanguage

    Humanity's Last Exam (HLE)

    Extremely difficult test designed to be hard for AI, heavily testing test-time compute.

    Why it matteredPre-2026 models scored <10%; reasoning models (Deep Research, Sol) jumped the score by ~30% in one year.

    ImpactProved the scaling law for inference (test-time compute) over pretraining.

    HLE Score Progression~30% jump
    Pre-2025o1Sol
    • reasoning
    • hard
    Source →
  4. benchmarklanguage

    SWE-bench Verified

    A refined, human-validated subset of SWE-bench.

    Why it matteredReduced false-negatives in the original SWE-bench to accurately score top models.

    ImpactClaude 5 Fable set the current record at 95.0% in July 2026.

    Coding Agent Evals>90%
    • coding
    • agents
    Source →
  5. leaderboardlanguage

    MMLU-Pro / harder knowledge suites

    Harder, cleaner successors as classic MMLU loses discriminative power.

    Why it matteredSignal that the community is retiring saturated scoreboards.

    ImpactExpect more ‘Pro’ / contamination-aware suites in vendor cards.

    • mmlu
    • eval
    Source →
  6. benchmarklanguage

    Arena-Hard / hard-prompt suites

    Benchmarks built from harder Arena-style prompts to restore separation among top models.

    Why it matteredTop models clustered on easy chat evals.

    ImpactAnother sign leaderboards must keep getting harder to stay useful.

    • preference
    • hard
    Source →
  7. benchmarklanguage

    GPQA

    Graduate-level Google-proof Q&A designed to be hard even for experts.

    Why it matteredPushes past saturated undergrad-style suites.

    ImpactFavorite citation for ‘frontier reasoning’ claims.

    • reasoning
    • hard
    Source →
  8. benchmarklanguage

    SWE-bench

    Real GitHub issues → patch generation evaluated with tests.

    Why it matteredMoved coding eval from toy functions toward repository-level work.

    ImpactAgent and coding-model launches now often lead with SWE-bench variants.

    Coding Agent Evals>90%
    • coding
    • agents
    Source →
  9. benchmarklanguage

    MMLU

    Massive Multitask Language Understanding - 57-subject knowledge/reasoning suite.

    Why it matteredBecame the default slide for claiming ‘general’ LLM ability.

    ImpactNow saturated >92% by top models.

    • llm
    • knowledge
    Source →
  10. benchmarklanguage

    SQuAD

    Extractive QA benchmark that dominated late-2010s NLP progress charts.

    Why it matteredMade reading-comprehension a standard fine-tune target before LLMs ate QA.

    ImpactBERT-era models saturated it - a cautionary tale about benchmark expiry.