LLM LEAGUE · DAILY · AUDITABLE

The LLM research worth reading, ranked.

We read, compare, and score the latest arXiv LLM research every day, so you can focus on what matters.

36 screened5 full-text reads33 pairwise matchesShielded evaluator

LLM RESEARCH · LATEST COHORT

Today’s leaderboard

Ranked against today’s peers. Scores estimate research quality, not popularity.

AUTOMATION ACTIVE

After 00:00 UTC, the first arena visit reads the newest arXiv papers, scores both leagues, and freezes each daily cohort in the archive.

Showing the last bundled LLM cohort while the live arXiv feed reconnects. Evaluator firewall: 0 untrusted segments excluded.

LLM ELO · BENCHMARK V0.2

An Elo score you can inspect.

Every paper starts at 1500 and enters two confidence-weighted round-robin passes. Match outcomes are soft pairwise judgments from the weighted evidence rubric.

R′ = R + K × (S − E)K = 30 × lower confidence · no author, citation, or recency signal
24%

Evidence

Baselines, ablations, uncertainty and sample breadth

18%

Novelty

A new mechanism or finding, not new branding

17%

Reproducibility

Methods, artifacts, data and implementation detail

15%

Impact

Capability unlocked and plausible path to use

10%

Robustness

Stress tests, limitations, safety and failure analysis

8%

Efficiency

Quality gained per unit of compute, data and latency

8%

Clarity

Claim precision, structure and evidence traceability

EVALUATOR FIREWALL

The paper is evidence, never instruction.

01

Isolate

Paper HTML is handled as untrusted data. Scripts, metadata, code, prompts, quotations, and example blocks never enter scoring text.

02

Neutralize

Evaluator-directed language such as “ignore the rubric” or “grade this 100/100” is removed and logged before feature extraction.

03

Saturate

Each evidence marker has a hard occurrence cap. Repeating “ablation,” “proof,” or “threat model” cannot keyword-stuff a paper upward.

04

Constrain

The rating engine is deterministic and cannot execute instructions. Repeated manipulation also caps confidence and appears in the audit drawer.

Never used as quality signalsCitationsInstitutionAuthor identitySocial reachSelf-assigned scores

LLM RESEARCH HISTORY

Daily cohorts stay frozen.

These are saved runs, not sample rows. Each UTC day keeps its winning paper, Elo result, integrity exclusions, and full ranked cohort.

6 rankedZero-Mem: Zero-Token Memory Operations for LLM Agents1847