Evidence
Baselines, ablations, uncertainty and sample breadth
Reading LLM papers…
LLM LEAGUE · DAILY · AUDITABLE
We read, compare, and score the latest arXiv LLM research every day, so you can focus on what matters.
LLM RESEARCH · LATEST COHORT
Ranked against today’s peers. Scores estimate research quality, not popularity.
After 00:00 UTC, the first arena visit reads the newest arXiv papers, scores both leagues, and freezes each daily cohort in the archive.
Showing the last bundled LLM cohort while the live arXiv feed reconnects. Evaluator firewall: 0 untrusted segments excluded.
LLM ELO · BENCHMARK V0.2
Every paper starts at 1500 and enters two confidence-weighted round-robin passes. Match outcomes are soft pairwise judgments from the weighted evidence rubric.
Baselines, ablations, uncertainty and sample breadth
A new mechanism or finding, not new branding
Methods, artifacts, data and implementation detail
Capability unlocked and plausible path to use
Stress tests, limitations, safety and failure analysis
Quality gained per unit of compute, data and latency
Claim precision, structure and evidence traceability
EVALUATOR FIREWALL
Paper HTML is handled as untrusted data. Scripts, metadata, code, prompts, quotations, and example blocks never enter scoring text.
Evaluator-directed language such as “ignore the rubric” or “grade this 100/100” is removed and logged before feature extraction.
Each evidence marker has a hard occurrence cap. Repeating “ablation,” “proof,” or “threat model” cannot keyword-stuff a paper upward.
The rating engine is deterministic and cannot execute instructions. Repeated manipulation also caps confidence and appears in the audit drawer.
LLM RESEARCH HISTORY
These are saved runs, not sample rows. Each UTC day keeps its winning paper, Elo result, integrity exclusions, and full ranked cohort.