{"version":"1.0","title":"Research Paper Arena","description":"Daily evidence-ranked LLM, security, and cryptography research from arXiv.","generatedAt":"2026-10-01T22:19:56.346Z","filter":"llm","count":12,"filters":["llm","security"],"links":{"self":"https://papers.turen.io/json?league=llm","all":"https://papers.turen.io/json","llm":"https://papers.turen.io/json?league=llm","security":"https://papers.turen.io/json?league=security"},"cohorts":[{"league":"llm","source":"snapshot","generatedAt":"2026-10-01T22:19:56.346Z","screened":45,"integritySummary":{"flagged":5,"excluded":9},"refresh":{"mode":"automatic","cadence":"daily","boundary":"00:00 UTC"},"count":12,"latestArchiveDay":"2026-10-01"}],"papers":[{"id":"2609.40361","league":"llm","rank":1,"title":"Ranking-Aware Prompt Optimization for Multimodal Clinical Diagnosis","authors":"Tian Xia, Minghao Liu, Yiqing Liang, Laixi Shi, Jiayun Wang","published":"Sep 30, 2026","topics":["Evaluation","Memory"],"mode":"Empirical","elo":1548,"confidence":96,"abstract":"Multimodal large language models (MLLMs) are rapidly advancing clinical diagnosis, yet their adaptation pipelines remain anchored to accuracy-based objectives. Clinical data are heavily class-imbalanced: a constant-majority predictor can score above 90% accuracy while being clinically useless. We therefore evaluate and optimize for AUROC, a threshold-free score that ranks positives above negatives and is invariant to class balance. We focus on prompt optimization in MLLMs. Reflective methods such as GEPA use a binary scores matrix with one row per evaluation instance and one column per candidate prompt; cells record per-instance correctness, so the column average is accuracy and drives candidate selection. We introduce pair-level Pareto prompt evolution (Ranking-PE), which replaces each correctness row with a pairwise-ordering row over (positive, negative) instance pairs: the cell is 1 if the candidate scores the positive higher than the paired negative. The column average then equals empirical AUROC (by the Wilcoxon-Mann-Whitney identity). We apply this swap at all three layers the prompt evolution search reads from - the scores matrix that decides Pareto dominance, the per-example feedback to the reflection LM, and final candidate selection - at no extra model calls and with no surrogate loss. Across three diseases on MIMIC, accuracy-based prompt evolution can degrade ranking; Ranking-PE reverses this, beating the accuracy-based recipe by +5.8 AUROC pp on fine-tuned Qwen3-VL-8B and +16.2 pp on MedGemma-4B. Ablations examine each design component and show that a medical-grade visual backbone - via vision-encoder-tuned SFT or medical pretraining - is a prerequisite that prompt search cannot replace - our recipe extends reflective prompt evolution from text-only data to multimodal clinical decision-making.","verdict":"This empirical paper wins most clearly on evidence. Its relative weakness is efficiency.","coverage":"Full text","change":48,"matches":22,"scores":{"Evidence":95,"Novelty":73,"Reproducibility":87,"Impact":74,"Robustness":79,"Efficiency":68,"Clarity":86},"integrity":{"flagged":false,"excludedSegments":0,"reasons":[]}},{"id":"2609.40347","league":"llm","rank":2,"title":"Image Classifiers are Efficient Self-Supervised Video Representation Learners","authors":"Owais Iqbal, Sudipta Sarkar, Shyam Marjit, Omprakash Chakraborty, Anirban Chakraborty, Abir Das","published":"Sep 30, 2026","topics":["Memory","Efficiency"],"mode":"Empirical","elo":1548,"confidence":96,"abstract":"We introduce VideoMSN, a Masked Siamese Network framework for efficient self-supervised spatio-temporal representation learning in videos. Instead of relying on heavy 3D architectures or reconstruction-based autoencoders for learning with unlabeled data, we repurpose standard image Vision Transformers by representing videos as super images which are grids composed of frames sampled from videos. From each super image, we construct two views: one with spatial patch masking and the other with temporal frame masking, ensuring no information leakage across frames. A shared Vision Transformer (ViT) encoder aligns their embeddings using a masked Siamese loss, capturing both motion and appearance cues without reconstruction. Our decoder-free formulation leverages an image foundation model towards efficient video representation learning. Starting from pretrained DINO-v3 and DeiT-v3 image encoders, VideoMSN achieves state-of-the-art performance on Kinetics-400, UCF101, and HMDB51 while requiring up to $32\\times$ fewer and $160\\times$ fewer video pretraining epochs compared to prior video self-supervised learning methods. Our proposed approach also shows strong performance in low-shot classification, confirming the transferability of the learned representations in a label-scarce scenario. Project Page: https://cvir.github.io/projects/videomsn.","verdict":"This empirical paper wins most clearly on evidence. Its relative weakness is robustness.","coverage":"Full text","change":48,"matches":22,"scores":{"Evidence":93,"Novelty":75,"Reproducibility":83,"Impact":76,"Robustness":71,"Efficiency":80,"Clarity":90},"integrity":{"flagged":false,"excludedSegments":0,"reasons":[]}},{"id":"2609.40340","league":"llm","rank":3,"title":"EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery","authors":"Young-Jun Lee, Jinheon Baek, Soyeong Jeong, Minki Kang, Seungyeon Jwa, Jonghyun Choi, Seungho Han, Dongyeop Kang","published":"Sep 30, 2026","topics":["Evaluation","Memory"],"mode":"Empirical","elo":1531,"confidence":76,"abstract":"Evolutionary search with large language models (LLMs) can stall when progress requires external knowledge the model lacks. Supplying relevant documents helps, but simply adding web search tool can keep returning the same pages as solutions change. We introduce EvoDuet, a bi-level optimization method that co-evolves solutions and search queries with fixed model parameters. At each iteration, a retrieval gate lets the LLM assess its knowledge gap and choose to retrieve new documents, reuse stored ones, or proceed without them. An inner loop refines queries and ranks documents by the solution scores they are predicted to yield; an outer loop generates candidates in parallel from these documents and records the evaluated outcomes for later searches. Across 21 optimization tasks with one candidate per iteration, EvoDuet raises OpenEvolve's normalized discovery gain from 74.1% to 78.0% with GPT-5.6-Luna and from 61.3% to 82.3% with Gemini-3.8-Flash, whereas Qwen3.5-9B does not benefit. Our best runs surpass the previously reported best scores on eight tasks, including Swap Reduction on Q20 and Rosetta, and match them on three more. EvoDuet also improves with other scaffolds (e.g., Top-K, EvoX) on Sums/Diffs and Denoising, demonstrating its applicability across evolutionary search scaffolds.","verdict":"This empirical paper wins most clearly on evidence. Its relative weakness is impact. The evaluator firewall excluded 4 grader-directed segments before scoring.","coverage":"Full text","change":31,"matches":22,"scores":{"Evidence":97,"Novelty":73,"Reproducibility":83,"Impact":70,"Robustness":71,"Efficiency":78,"Clarity":86},"integrity":{"flagged":true,"excludedSegments":4,"reasons":["grader-targeting language","score manipulation"]}},{"id":"2609.40360","league":"llm","rank":4,"title":"Semifactual Credit-Augmented Policy Optimization","authors":"Junshu Pan, Zhizhang Fu, Shulin Huang, Yiran Ding, Zifan Cheng, Wenqi Shao, Qiaosheng Zhang, Yue Zhang","published":"Sep 30, 2026","topics":["Evaluation","Reasoning"],"mode":"Empirical","elo":1530,"confidence":90,"abstract":"Reinforcement learning with verifiable rewards (RLVR) has improved the reasoning capabilities of large language models (LLMs), yet their predictions remain sensitive to task-irrelevant prompt features. We investigate this sensitivity through semifactual prompt interventions that preserve the underlying problem and its answer. Our analysis reveals substantial variation in token-level sensitivity and shows that suppressing high-drift token candidates during decoding improves reasoning accuracy without updating model weights. These findings highlight a limitation of Group Relative Policy Optimization (GRPO), which assigns the same outcome-derived advantage to every response token and may reinforce potential spurious dependence alongside useful reasoning. Motivated by this observation, we introduce Semifactual Credit-Augmented Policy Optimization (SCAPO), a causally inspired variant of GRPO that incorporates semifactual stability into token-level credit assignment. SCAPO measures token probability drift for fixed responses under semifactual interventions and uses normalized stability scores to reduce advantages for relatively unstable tokens during early training, while granting no additional credit for stability alone. On Qwen3-4B-Base and Qwen3-1.7B-Base, SCAPO improves AIME 2024-2026 accuracy over GRPO by 5.63 and 4.17 percentage points, respectively. At both model scales, SCAPO achieves the best results on most evaluated mathematics benchmarks and all evaluated out-of-distribution benchmarks among the compared methods. These results suggest that semifactual stability provides an effective training signal for improving reasoning and generalization through finer-grained credit assignment in RLVR. The code is available at https://github.com/DtYXs/SCAPO.","verdict":"This empirical paper wins most clearly on evidence. Its relative weakness is impact. The evaluator firewall excluded 1 grader-directed segment before scoring.","coverage":"Full text","change":30,"matches":22,"scores":{"Evidence":97,"Novelty":71,"Reproducibility":85,"Impact":67,"Robustness":79,"Efficiency":70,"Clarity":88},"integrity":{"flagged":true,"excludedSegments":1,"reasons":["grader-targeting language"]}},{"id":"2609.40306","league":"llm","rank":5,"title":"DynaHarness: A Dynamic Physical Harness for Self-Evolving Robot Agents","authors":"Haoyuan Deng, Jiebin Liu, Tengxiao Zhang, Langning Yan, Hongye Cao, Ziwei Wang","published":"Sep 30, 2026","topics":["Agents","Reasoning"],"mode":"Empirical","elo":1521,"confidence":96,"abstract":"Pretrained robot policies provide useful action priors, but long-horizon manipulation still requires coordination between semantic reasoning and physical execution. Semantic reasoning operates at a coarser timescale than physical interaction, while episode-level failures provide limited guidance on which system component should be revised. We propose DynaHarness, a dynamic physical harness that couples semantic reasoning with physical governance through a shared execution contract and turns failure evidence into validated capability revisions. To be more specific, the slow brain proposes capabilities and symbolic arguments, while the fast brain grounds and monitors commands, refuses unresolved actions, substitutes capabilities, and requests replans when needed. The physical execution contract bounds each accepted command and records execution evidence across analytic skills, recovery skills, and the frozen VLA. Failure attribution localizes faults in these records and directs targeted revisions of reusable capabilities or execution mechanisms. Paired regression checks govern admission or rejection, closing the self-evolution loop. On LIBERO-Pro, DynaHarness achieves 75.2% on 800 newly sampled initial states, compared with 17.5% for the frozen policy. With the same capability library, full dynamic execution reaches 74.0% versus 63.9% under nominal one-step replanning. This demonstrates the value of DynaHarness as a dynamic physical harness that governs how existing capabilities are grounded, monitored, and coordinated during execution. Our project page is at https://denghaoyuan123.github.io/Dynaharness_page/.","verdict":"This empirical paper wins most clearly on evidence. Its relative weakness is novelty.","coverage":"Full text","change":21,"matches":22,"scores":{"Evidence":93,"Novelty":71,"Reproducibility":81,"Impact":74,"Robustness":79,"Efficiency":78,"Clarity":83},"integrity":{"flagged":false,"excludedSegments":0,"reasons":[]}},{"id":"2609.40316","league":"llm","rank":6,"title":"Scaling Laws for Looped Mixture of Experts","authors":"Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi","published":"Sep 30, 2026","topics":["Evaluation","Reasoning"],"mode":"Empirical","elo":1511,"confidence":96,"abstract":"Looped transformers and Mixture-of-Experts (MoE) offer complementary routes to efficient scaling: recurrence increases computational depth at fixed parameters, while MoE sparsity expands total capacity at fixed active compute. Yet existing scaling laws model recurrence or sparsity in isolation. In this work, we introduce Loop Scaling Laws, the first scaling law to jointly model recurrence and sparsity alongside model size and data. At its core is a bounded, sparsity-conditional recurrence mapping that characterizes the effective-parameter gain from looping and how sparsity raises this gain. The laws predict the held-out loss of looped models more accurately than prior alternatives, and recover the standard dense and MoE scaling laws as special cases. Beyond prediction, the fitted laws provide a principled foundation for designing looped MoE models under compute and memory constraints. Downstream evaluations further demonstrate the complementary benefits of the two axes: sparsity delivers ~3x active-parameter efficiency, recurrence yields ~2x total-parameter efficiency on reasoning, and joint scaling further advances the performance frontier. As a practical extension, we show these gains hold at trillion-token scale: at matched training compute, a looped MoE with law-derived recurrence matches a ~2x larger non-looped MoE on the reasoning benchmarks, while enabling test-time scaling through recurrence.","verdict":"This empirical paper wins most clearly on evidence. Its relative weakness is impact.","coverage":"Full text","change":11,"matches":22,"scores":{"Evidence":95,"Novelty":75,"Reproducibility":76,"Impact":72,"Robustness":73,"Efficiency":76,"Clarity":85},"integrity":{"flagged":false,"excludedSegments":0,"reasons":[]}},{"id":"2609.40335","league":"llm","rank":7,"title":"Is Weight Tying Still Beneficial for Decoder-Only LLMs in Private Settings Under DP-SGD?","authors":"Razan El Mais, Ali Chehab, Ibrahim Issa, Razane Tajeddine","published":"Sep 30, 2026","topics":["Memory","Efficiency"],"mode":"Empirical","elo":1495,"confidence":96,"abstract":"Differentially Private Stochastic Gradient Descent (DP-SGD) is a leading approach for privacy-preserving fine-tuning of large language models (LLMs). Many decoder-only LLMs employ weight tying between input and output embeddings, a design choice originally introduced for parameter efficiency and improved language modeling performance in the non-private setting. However, the impact of weight tying under differentially private training remains largely unexplored. In this work, we investigate the role of weight tying in the DP setting using GPT2 and DistilGPT2 as representative decoder-only architectures. Interestingly, we find that untied embeddings consistently outperform weight-tied models under DP-SGD, achieving gains of up to 4.74% points in accuracy on SST-2, QNLI, and QQP. Beyond improved utility, untying embeddings enables the use of memory-efficient ghost clipping for DP-SGD. By contrast, weight tying introduces shared-parameter interactions that complicate standard ghost norm computation and largely negate its computational advantages. As a result, untied models achieve over 60% lower memory usage while preserving the benefits of ghost clipping. Our results indicate that untied embeddings provide a more effective and scalable design for differentially private training of decoder-only LLMs and highlight the need to revisit standard LLM architectural choices in the privacy-preserving setting.","verdict":"This empirical paper wins most clearly on clarity. Its relative weakness is robustness.","coverage":"Full text","change":-5,"matches":22,"scores":{"Evidence":85,"Novelty":73,"Reproducibility":85,"Impact":72,"Robustness":71,"Efficiency":76,"Clarity":90},"integrity":{"flagged":false,"excludedSegments":0,"reasons":[]}},{"id":"2609.40330","league":"llm","rank":8,"title":"Turbo Harness: Instance-Adaptive Harness Optimization","authors":"Tunyu Zhang, Hao Wang, Kai Xu, Dimitris N. Metaxas","published":"Sep 30, 2026","topics":["Agents","Evaluation"],"mode":"Empirical","elo":1494,"confidence":90,"abstract":"Automating the search for effective harnesses is an important step toward enabling agents to recursively self-improve. Existing harness optimizations typically produce a single global harness that is applied uniformly across task instances. However, a harness that works well on average may not be optimal for every instance. We introduce Turbo Harness, a framework that can adapt a globally optimized harness to each instance by reusing information generated during the original optimization process. Specifically, Turbo Harness recycles artifacts produced during a completed global harness optimization run, and summarizes them into a structured playbook. We train a harness editor to leverage this prior optimization experience to generate instance-specific patches to the global harness. At inference time, the editor uses the instance and the playbook to construct a tailored harness in which the execution model operates. Through numerical experiments, we show that Turbo Harness consistently outperforms existing harness optimization baselines across seven benchmarks spanning interactive agent tasks, software engineering, and long-horizon terminal tasks.","verdict":"This empirical paper wins most clearly on evidence. Its relative weakness is robustness. The evaluator firewall excluded 1 grader-directed segment before scoring.","coverage":"Full text","change":-6,"matches":22,"scores":{"Evidence":93,"Novelty":75,"Reproducibility":74,"Impact":74,"Robustness":68,"Efficiency":74,"Clarity":86},"integrity":{"flagged":true,"excludedSegments":1,"reasons":["grader-targeting language"]}},{"id":"2609.40325","league":"llm","rank":9,"title":"WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents","authors":"Ziyan Jiang, Jingbo Yang, Jiabao Ji, Yujian Liu, Qiucheng Wu, Tommi Jaakkola, Yang Zhang, Shiyu Chang","published":"Sep 30, 2026","topics":["Agents","Evaluation"],"mode":"Empirical","elo":1479,"confidence":90,"abstract":"As interactive 3D worlds are increasingly used to study intelligent behavior, it becomes important to develop efficient pipelines for identifying anomalies in these simulated environments, such as floating objects, traversable walls, or objects inconsistent with the surrounding scene. Multimodal AI systems, including vision-language models (VLMs) and vision-language-action models (VLAs), have shown potential for automating this task. However, 3D world auditing is complex, requiring the close coupling of two distinct capabilities: action, to navigate the 3D world and search for anomalies systematically and efficiently; and visual reasoning, to understand the environment and identify anomalies from multimodal observations. It remains largely unexplored whether multimodal agents can effectively couple these two capabilities, using visual reasoning to identify potential anomalies while taking actions to validate them. In this paper, we introduce WorldAuditBench, a benchmark for 3D world auditing comprising 213 anomaly tasks across 13 environments built with Unreal Engine 5 and Three.js, spanning five anomaly families. We evaluate five frontier models under a fixed exploration budget using two auditing paradigms: VLA-based exploration followed by VLM-based anomaly identification, and an end-to-end VLM agent in which visual reasoning directly guides action selection. Across the evaluated models and two paradigms, success rates range from 6.6% to 42.3%, substantially below human performance (83.4%). Through the task of world auditing, WorldAuditBench provides a testbed for studying how multimodal agents couple action and visual reasoning in interactive 3D environments, while highlighting current limitations in their ability to gather and interpret evidence during exploration.","verdict":"This empirical paper wins most clearly on evidence. Its relative weakness is impact. The evaluator firewall excluded 1 grader-directed segment before scoring.","coverage":"Full text","change":-21,"matches":22,"scores":{"Evidence":93,"Novelty":75,"Reproducibility":81,"Impact":63,"Robustness":68,"Efficiency":70,"Clarity":86},"integrity":{"flagged":true,"excludedSegments":1,"reasons":["instruction override"]}},{"id":"2609.40322","league":"llm","rank":10,"title":"MatLoom: Layered Text-to-Material Generation in a Compact Program Space","authors":"Anson Y. Lam, Shuqing Li, Michael R. Lyu","published":"Sep 30, 2026","topics":["Evaluation","Memory"],"mode":"Empirical","elo":1479,"confidence":90,"abstract":"Material generation should produce not only an appearance, but also the rules that construct it. We introduce MatLoom, a compact, layer-oriented language for text-to-material generation with pretrained language models. Each program composes alpha-masked layers whose shared spatial expressions define coverage and physically based rendering (PBR) channels, making dependencies between patterns, color, and relief explicit. A standalone interpreter evaluates the program into material maps, while the source retains named fields and layer parameters for subsequent authoring. Without task-specific fine-tuning, our pipeline uses parser-guided repair and preview-based critique to revise material designs, then searches noise seeds while keeping each candidate's remaining source fixed. On a curated benchmark of 141 prompts evaluated with six backbones, our best-performing configuration achieves higher mean scores than three diffusion baselines on all four flat-layout prompt-alignment metrics. Its initial programs already exceed all three baselines on mean BLIPScore, before critique or seed search. Retained programs have a median length of 21 lines when pooled across backbones. In a blind four-way comparison involving 30 participants and 20 prompts, our renders receive 59.2% of choices, compared with 19.3% for the most-preferred baseline. Compact executable programs thus offer a way to generate prompt-aligned materials while retaining their construction as part of the asset.","verdict":"This empirical paper wins most clearly on evidence. Its relative weakness is impact. The evaluator firewall excluded 2 grader-directed segments before scoring.","coverage":"Full text","change":-21,"matches":22,"scores":{"Evidence":97,"Novelty":73,"Reproducibility":79,"Impact":65,"Robustness":68,"Efficiency":66,"Clarity":85},"integrity":{"flagged":true,"excludedSegments":2,"reasons":["grader-targeting language"]}},{"id":"2609.40359","league":"llm","rank":11,"title":"Removing Timing Shortcuts Improves Non-Invasive Brain-to-Text","authors":"Dulhan Jayalath, Oiwi Parker Jones","published":"Sep 30, 2026","topics":["Evaluation","Memory"],"mode":"Empirical","elo":1474,"confidence":93,"abstract":"We find that major reported improvements in decoding words from non-invasive brain recordings are largely reproducible without any brain data. In the influential work of d'Ascoli et al. (2025), time series of brain activity from subjects perceiving continuous speech are segmented into fixed-length windows starting at each word. A neural network then generates predictions for all of the words in a sentence together. Neighbouring windows partially overlap, implicitly revealing the interval between words. Since these intervals indicate the duration of the words spoken, and different words tend to have different durations - for example, \"the\" is much shorter than \"supercalifragilisticexpialidocious\" - the neural network can improve its predictions of words without relying on the underlying brain activity. Consistent with this, the method reaches 22.0% balanced accuracy on synthetic signals containing no brain information, compared with 22.3% on real brain recordings. To prevent the network from learning this shortcut, we make a single, simple change. Instead of jointly encoding all windows in a sentence, we process each independently. As a result, the neural network achieves better performance by learning underlying word-specific information from brain recordings. This makes two existing strategies become much more effective than before. Both aggregating predictions from distinct neural responses to the same word and using a pretrained LLM as a linguistic prior now substantially improve results. On our perceived speech benchmark, this simple recipe (SimpleB2T) achieves a word error rate of 36.6% with five observations per word, approaching past invasive speech decoding performance, albeit under different conditions. The results in this work expose an important shortcut in brain-to-text decoding and show that removing it leads to a simple and considerably more effective strategy.","verdict":"This empirical paper wins most clearly on evidence. Its relative weakness is robustness.","coverage":"Full text","change":-26,"matches":22,"scores":{"Evidence":93,"Novelty":66,"Reproducibility":85,"Impact":72,"Robustness":64,"Efficiency":70,"Clarity":83},"integrity":{"flagged":false,"excludedSegments":0,"reasons":[]}},{"id":"2609.40324","league":"llm","rank":12,"title":"Cogentic: Multi-Agent Orchestration for Automated Proof Discovery","authors":"Yang Cai, Vineet Gupta, Yanchen Jiang, Christopher Liaw, Aranyak Mehta, Grigoris Velegkas, Di Wang","published":"Sep 30, 2026","topics":["Agents"],"mode":"Empirical","elo":1392,"confidence":94,"abstract":"We present Cogentic, a multi-agent harness for automated proof discovery on open research problems. While frontier language models can generate strong mathematical ideas in a single shot, single-shot generation is often insufficient for open problems that require exploring multiple competing conjectures, overcoming subtle technical obstructions, and retaining intermediate progress over a long horizon. Cogentic addresses these challenges through an iterative prove--verify loop in which an orchestrator allocates a population of independent provers across distinct proof directions, subjects their output to adversarial verification by several specialized components, and promotes confirmed intermediate results into a persistent verified ledger that later rounds build on. The harness is designed to be able to solve research-level math and theoretical computer science problems. Using Gemini as the base model, Cogentic produced novel results on five open problems across online learning, auction theory, and mechanism design. Each result was independently verified by domain experts and is developed in full in companion papers. We list these results, and new ones as they are verified, at https://sites.google.com/view/cogentic .","verdict":"This empirical paper wins most clearly on clarity. Its relative weakness is impact.","coverage":"Full text","change":-108,"matches":22,"scores":{"Evidence":76,"Novelty":78,"Reproducibility":70,"Impact":65,"Robustness":73,"Efficiency":72,"Clarity":79},"integrity":{"flagged":false,"excludedSegments":0,"reasons":[]}}]}