📄 Notable* Recent AI/ML arXiv Papers

Last updated just now...

📄 Game Arena: Strategic LLM Evaluation in Competitive Environments
🗓️ Published: 9/25/2026
🔗 http://arxiv.org/abs/2609.31473v1
👥 Authors: Bovard Doerschuk-Tiberi, Yao Yan, Justin Chiu (possible past Carnegie Mellon University affiliation), Hann Wang, Timothy Chung, Martyna Plomecka, John Schultz, Jon Lipovetz, Clayton Drazner, Yuchen Zhuang, Jaimie Hwang, Nate Keating, Riley Jones, Andrew Lee, Oran Kelly, Ian Gemp, Michael Aaron, Laurel Prince, Kate Larson, Jeff Moser, Harrison Jobe, Chad Woodford, Siqi Liu (possible past University Of Oxford affiliation), Andrew Wang (possible past University Of California, Berkeley affiliation), Bo Chang, Christopher D'mello, Diane Chaleff, Addison Howard (possible past Google (United States) affiliation), Johnny Yip, Chuck Sugnet, Antonio Gulli, Meghan O'connell, Will Cukierski, Nenad Tomasev, Dima Yeroshenko, Kinjal Parekh, Roxanne Daniel, Marc Lanctot (possible past Google (United States) affiliation), Domino Weir, Elsa Dong, Daniel Hennes (possible past Deepmind (United Kingdom) affiliation), Melissa Nalubwama, Robert Fraser, Ryan Trostle, Jun Peng, Tom Mason, Lloyd Hightower, Chiamaka Chukwuka, Yuexiang Zhai, Phoebe Kirk, Yi Su, Yuting Han, Jie Ren (possible past Google (United States) affiliation), Chris Prichard, Sahand Sharifzadeh, Karim Hakimzadeh, Dj Sterling, Meg Risdal, Kate Olszewska, Ya Xu (possible past Stanford University affiliation), Orhan Firat, Minmin Chen (possible past Google (United States) affiliation)
Abstract

We introduce Kaggle Game Arena, an open and ever-expanding platform to evaluate large language models (LLMs) through competitive games. Different from static benchmarks, game arena enables models to play head-to-head matchups in structured environments where the gameplay strength naturally increases as models evolve, preventing performance saturation. This technical report details the infrastructure behind Game Arena and describes the three pilot game environments: Chess, Poker, and Werewolf. Th...

📄 ActKV: Efficient LLM Agents through Action-Guided KV Cache Management
🗓️ Published: 9/25/2026
🔗 http://arxiv.org/abs/2609.31395v1
👥 Authors: Zihan Wang (possible past Tsinghua University affiliation), Cheng Tang, Lei Gong, Chao Wang (possible past Google (United States) affiliation), Wenqi Lou, Teng Wang, Xuehai Zhou
Abstract

Agentic LLM inference accumulates long KV caches across iterative observation-reasoning-action loops, imposing substantial memory overhead and limiting serving throughput. Existing compression methods emphasize overall output quality, overlooking the asymmetric importance of actions in driving task progress. Our key idea is to establish a compression criterion that values KV entries by their contribution to action generation and prioritizes action quality. However, iterative execution, dynamic m...

📄 Which Influence Are We Estimating? The Role of Counterfactual Specifications in Data Attribution
🗓️ Published: 9/25/2026
🔗 http://arxiv.org/abs/2609.31214v1
👥 Authors: Zhe Li (possible past Google (United States) affiliation), Wei Zhao (possible past Tencent (China) affiliation), Peixin Zhang, Jun Sun
Abstract

Estimating the influence of training examples on model behavior is essential for data debugging, valuation, and attribution. Existing influence estimators often produce incompatible rankings, which are commonly ascribed to approximation error. We argue that a more fundamental source of disagreement is specification mismatch: influence depends on the behavior being attributed, the intervention applied to each training example, and the counterfactual training process that maps the intervention to ...

📄 JevAdvBench: A Benchmark and Black-Box Attacks for Reinforcement Learning for Calibrated Decisions Models
🗓️ Published: 9/25/2026
🔗 http://arxiv.org/abs/2609.31142v1
👥 Authors: Jianyi Hu, Hangtao Zhang, Yi Liu (possible past Google (United States) affiliation), Yeqi Zeng, Li Zeng, Xianlong Wang, Rui Wang (possible past Tencent (China) affiliation), Leo Yu Zhang
Abstract

Models trained with reinforcement learning for calibrated decisions (RLCD), such as Jev, answer a typed question about an input, the state, with a probability, a choice, or a score, and software acts on the answer without a person reading it. Their robustness has not been measured: adversarial benchmarks score what a model generates or executes, whereas a typed model generates nothing and returns a well-formed answer even when manipulated. Measurement is also hard, because identical requests can...

📄 G$^2$PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation
🗓️ Published: 9/25/2026
🔗 http://arxiv.org/abs/2609.31009v1
👥 Authors: Ruikang Liu, Haoli Bai, Yuxuan Sun, Qian Zhang (possible past University Of Washington affiliation), Wenzheng Cai, Yanqi Hao, Feiyu Wang, Weidong Zhong, Zhuang Wang, Tong Yang (possible past Peking University affiliation), Xiangsheng Zhou
Abstract

Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large language models (LLMs) without retraining. GPTQ-based methods have become the de facto standard, yet they suffer from two complementary limitations. Methods with local, layer-wise objectives lack global supervision; while methods with global objectives fix their Hessian estimates at the start and ignore first-order gradients, so their guidance grows stale as quantization proceeds....

📄 SciHorizon-eLab: An Agentic Protocol-to-Task Compiler for Scalable Benchmarking of Scientific Embodied Agents
🗓️ Published: 9/25/2026
🔗 http://arxiv.org/abs/2609.30971v1
👥 Authors: Maokai Qin, Chuan Qin (possible past Baidu (China) affiliation), Qi Zhang (possible past Tencent (China) affiliation), Dianyu Liu, Zirui Liu, Hongting Niu, Yuanchun Zhou, Hengshu Zhu (possible past Baidu (China) affiliation)
Abstract

Embodied agents offer a promising route to automating scientific experimentation, yet their progress is constrained by the lack of reliable and systematic evaluation environments. Existing simulation-based laboratory benchmarks rely heavily on manual task engineering, making it challenging to systematically compile diverse scientific protocols into executable and verifiable embodied tasks at scale. To address this challenge, we introduce SciHorizon-eLab, an agentic protocol-to-task compiler that...

📄 UltraG-Bench: A Multi-task Benchmark for assessing Large Vision-Language Models on Pixel-level Evidence Grounding in Ultrasound
🗓️ Published: 9/25/2026
🔗 http://arxiv.org/abs/2609.30928v1
👥 Authors: Quanhao Zhu, Bo Xu, Rui Lin, Chenyuan Wang, Yu Shao, Boling Zhu, Jiuyan Sun, Liang Zhao (possible past Baidu (China) affiliation), Hongfei Lin, Feng Xia (possible past Tencent (China) affiliation)
Abstract

Ultrasound is one of the most widely used medical imaging modalities, and recent large vision-language models(VLMs) have shown increasing capabilities in ultrasound image understanding. However, these models fail to provide pixel-level visual evidence aligned with their semantic predictions, and their fine-grained grounding capability in ultrasound remains largely unclear. We introduce UltraG-Bench, a large-scale multi-task benchmark for evaluating pixel-level evidence grounding in ultrasound. U...

📄 MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation
🗓️ Published: 9/25/2026
🔗 http://arxiv.org/abs/2609.30837v1
👥 Authors: Tianze Xu, Yanzhao Zheng, Zhentao Zhang, Yuanqiang Yu, Chao Ma (possible past Shanghai Jiao Tong University affiliation), Jihuai Zhu, Lelun Wu, Lyumanshan Ye, Pengfei Liu, Baohua Dong, Hangcheng Zhu, Ruohui Huang, Gang Yu (possible past Tencent (China) affiliation)
Abstract

Multi-teacher on-policy distillation (MOPD) integrates specialized capabilities into a single student, but existing practice typically hard-routes each prompt to a domain-matched teacher for the entire rollout. This dependence on prompt-level domain labels restricts using unlabeled training mixtures and leaves complementary signals from other teachers unused. We introduce MOPD-Router, a framework that routes supervision over the full teacher pool at each token, without domain labels or training ...

📄 Evaluation Is All You Need for Multi-Modal Autonomous Driving
🗓️ Published: 9/25/2026
🔗 http://arxiv.org/abs/2609.30818v1
👥 Authors: Zeyu He, Shiqi Liu, Ke Chen (possible past Tencent (China) affiliation), Yun Yan, Jinzi Wu, Dianqiao Lei, Sirui Wang, Shurui Peng, Tao Chen, Zhuo Huang, Yu Wu (possible past Baidu (China) affiliation), Yadong Shao, Zhichao Li (possible past Baidu (China) affiliation), Ke Sun, Yang Guan, Keqiang Li, Shengbo Eben Li (possible past Tsinghua University affiliation)
Abstract

Multi-modal planning is promising for autonomous driving by representing multiple plausible behaviors in ambiguous and long-tail scenarios. Existing methods mainly focus on improving trajectory multi-modality, enhancing trajectory representations, or reshaping the candidate distribution. Nevertheless, we identify a pronounced generation-evaluation asymmetry in multi-modal planning: despite strong oracle performance, existing planners often fail to reliably select the best available candidate, le...

📄 Analyzing and Mitigating Cost-Inefficient Behaviors in Coding Agents
🗓️ Published: 9/25/2026
🔗 http://arxiv.org/abs/2609.30725v1
👥 Authors: Yiran Hu, Nan Jiang (possible past Stanford University affiliation), Shanchao Liang, Anik Dey, Yi Wu (possible past University Of California, Berkeley affiliation), Lin Tan
Abstract

Although effective, coding agents often incur substantial monetary costs. Their recurring cost-inefficient behaviors remain underexplored. We conduct the first study of behavioral cost inefficiencies in coding agents, analyzing 1,200 trajectories from Claude Code and Mini-SWE-Agent across four configurations on SWE-bench Verified. We identify three cost-inefficient behaviors: subsumed retrieval, similar script generation, and test re-execution. We then evaluate three mitigation strategies: struc...

📄 T-RoPE: Time-Aware Rotary Position Embedding for Sequential Recommendation
🗓️ Published: 9/24/2026
🔗 http://arxiv.org/abs/2609.30576v1
👥 Authors: Yang Liu (possible past Tsinghua University affiliation), Noel Loo, Ali Khanafer, Shuying Sun (possible past Meta (United States) affiliation), Akshay Soni, Zhong Wu, Linjun Yang (possible past Microsoft (United States) affiliation)
Abstract

Large-scale recommenders increasingly adopt the sequential generative recipe behind large language models, bringing the Transformer into recommendation along with design choices made for text, including Rotary Position Embedding (RoPE). In language models, RoPE encodes token indices for relative position reasoning, but in recommendation, an interaction index records only event order, saying nothing about elapsed time, behavioral cycles across scales, or calendar phase. We revisit this choice and...

📄 A Benchmarking Framework for Context-aware XR Interfaces
🗓️ Published: 9/24/2026
🔗 http://arxiv.org/abs/2609.30466v1
👥 Authors: Hyunsung Cho, Sarah Yewon Yun, Nancy Ruonan Sun, Ben Lafreniere, Mark Parent, Kashyap Todi, Tanya R. Jonker (possible past Meta (United States) affiliation), Hrvoje Benko (possible past Meta (United States) affiliation), Sherry Tongshuang Wu, David Lindlbauer
Abstract

Everyday Extended Reality (XR) systems aim to provide context-aware access to the right functionalities at the right time and place, with minimal manual reconfiguration as users switch context. Yet these interfaces are hard to evaluate: current prototyping and user-study workflows offer no systematic, repeatable way to compare adaptation methods across users and scenarios. We present ContextXR, a novel benchmarking framework for context-aware XR interfaces. ContextXR represents an XR application...

📄 Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems
🗓️ Published: 9/24/2026
🔗 http://arxiv.org/abs/2609.30383v1
👥 Authors: Zihao Zhu, Siwei Lyu (possible past University Of Washington affiliation), Adel Bibi, Baoyuan Wu (possible past Tencent (China) affiliation)
Abstract

A skill is a modular package of natural-language instructions, executable scripts, and reference resources that an agent can load at runtime to extend its capabilities for a specific task. Skill-based agent systems therefore enable flexible reuse of third-party capabilities, but the openness of this skill ecosystem also opens up a new attack surface. Prior work has focused on vulnerabilities within individual skills, but little attention has been paid to risks that arise from interactions across...

📄 TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations
🗓️ Published: 9/24/2026
🔗 http://arxiv.org/abs/2609.30222v1
👥 Authors: Ayush Jain (possible past Google (United States) affiliation), Sreeharsha Paruchuri, Ishita Gupta, Fan Zhang, Tanner Schmidt (possible past University Of Washington affiliation), Jakob Engel (possible past Meta (United States) affiliation), Katerina Fragkiadaki (possible past University Of California, Berkeley affiliation), Adam W. Harley (possible past Carnegie Mellon University affiliation)
Abstract

Existing point tracking models face a fundamental tradeoff: they can either track a sparse set of query points over long horizons, or track all points across only short clips. We introduce TrackEverything, a 3D point tracker that breaks this trade-off by representing videos as persistent 3D scene tracks in world coordinates. Grounded in the insight that videos are 2D projections of an underlying 3D world, TrackEverything decouples model complexity from video duration, allowing it to scale with u...

📄 ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds
🗓️ Published: 9/24/2026
🔗 http://arxiv.org/abs/2609.30199v1
👥 Authors: Ming Zhang (possible past Peking University affiliation), Zhenghao Xiang, Peizhong Gao, Yujiong Shen, Yuhui Wang, Zhonghan Yue, Shihan Dou, Zhangyue Yin, Junjie Ye, Shichun Liu, Weihuang Zheng, Jiahao Chen, Jiayi Chen, Hongzhang Liu, Jiaqi Shao, Tao Gui, Qi Zhang (possible past Tencent (China) affiliation), Xuanjing Huang, Suncong Zheng, Maxm Pan
Abstract

Scientific discovery begins where known problems end. There, AI systems must engage in exploration: framing hypotheses, designing experiments, and iterating on the results. However, evaluating this ability is difficult: (1) how to verify whether a genuinely new hypothesis holds, and (2) how to determine whether a system has discovered it through exploration or merely recalled related knowledge from pre-training data. To this end, we introduce ExplorationBench, which turns the wicked problem of e...

📄 Probabilistic Robustness-driven Universal Adversarial Perturbations with Explainability against Deep Reinforcement Learning-based Intrusion Detection System
🗓️ Published: 9/24/2026
🔗 http://arxiv.org/abs/2609.30605v1
👥 Authors: Hongsen Zhang, Lu Zhang (possible past Tencent (China) affiliation), Mingjing Xu, Yi Zhang (possible past Google (United States) affiliation), Gregory Epiphaniou, Carsten Maple
Abstract

Deep reinforcement learning (DRL) enables adaptive intrusion detection in dynamic network environments but also exposes intrusion detection systems (IDS) to adversarial threats such as universal adversarial perturbations (UAPs), which apply a single input-agnostic perturbation to degrade detection performance across traffic. Probabilistic Robustness (PR), as a post-hoc evaluation metric, provides a principled, population-level measure of adversarial impact that conceptually aligns with the unive...

*Notable papers are those with at least two authors from a "big" AI/ML lab.