📄 Notable* Recent AI/ML arXiv Papers

Last updated just now...

📄 Which Influence Are We Estimating? The Role of Counterfactual Specifications in Data Attribution
🗓️ Published: 9/25/2026
🔗 http://arxiv.org/abs/2609.31214v1
👥 Authors: Zhe Li (possible past Google (United States) affiliation), Wei Zhao (possible past Tencent (China) affiliation), Peixin Zhang, Jun Sun
Abstract

Estimating the influence of training examples on model behavior is essential for data debugging, valuation, and attribution. Existing influence estimators often produce incompatible rankings, which are commonly ascribed to approximation error. We argue that a more fundamental source of disagreement is specification mismatch: influence depends on the behavior being attributed, the intervention applied to each training example, and the counterfactual training process that maps the intervention to ...

📄 JevAdvBench: A Benchmark and Black-Box Attacks for Reinforcement Learning for Calibrated Decisions Models
🗓️ Published: 9/25/2026
🔗 http://arxiv.org/abs/2609.31142v1
👥 Authors: Jianyi Hu, Hangtao Zhang, Yi Liu (possible past Google (United States) affiliation), Yeqi Zeng, Li Zeng, Xianlong Wang, Rui Wang (possible past Tencent (China) affiliation), Leo Yu Zhang
Abstract

Models trained with reinforcement learning for calibrated decisions (RLCD), such as Jev, answer a typed question about an input, the state, with a probability, a choice, or a score, and software acts on the answer without a person reading it. Their robustness has not been measured: adversarial benchmarks score what a model generates or executes, whereas a typed model generates nothing and returns a well-formed answer even when manipulated. Measurement is also hard, because identical requests can...

📄 G$^2$PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation
🗓️ Published: 9/25/2026
🔗 http://arxiv.org/abs/2609.31009v1
👥 Authors: Ruikang Liu, Haoli Bai, Yuxuan Sun, Qian Zhang (possible past University Of Washington affiliation), Wenzheng Cai, Yanqi Hao, Feiyu Wang, Weidong Zhong, Zhuang Wang, Tong Yang (possible past Peking University affiliation), Xiangsheng Zhou
Abstract

Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large language models (LLMs) without retraining. GPTQ-based methods have become the de facto standard, yet they suffer from two complementary limitations. Methods with local, layer-wise objectives lack global supervision; while methods with global objectives fix their Hessian estimates at the start and ignore first-order gradients, so their guidance grows stale as quantization proceeds....

📄 SciHorizon-eLab: An Agentic Protocol-to-Task Compiler for Scalable Benchmarking of Scientific Embodied Agents
🗓️ Published: 9/25/2026
🔗 http://arxiv.org/abs/2609.30971v1
👥 Authors: Maokai Qin, Chuan Qin (possible past Baidu (China) affiliation), Qi Zhang (possible past Tencent (China) affiliation), Dianyu Liu, Zirui Liu, Hongting Niu, Yuanchun Zhou, Hengshu Zhu (possible past Baidu (China) affiliation)
Abstract

Embodied agents offer a promising route to automating scientific experimentation, yet their progress is constrained by the lack of reliable and systematic evaluation environments. Existing simulation-based laboratory benchmarks rely heavily on manual task engineering, making it challenging to systematically compile diverse scientific protocols into executable and verifiable embodied tasks at scale. To address this challenge, we introduce SciHorizon-eLab, an agentic protocol-to-task compiler that...

📄 UltraG-Bench: A Multi-task Benchmark for assessing Large Vision-Language Models on Pixel-level Evidence Grounding in Ultrasound
🗓️ Published: 9/25/2026
🔗 http://arxiv.org/abs/2609.30928v1
👥 Authors: Quanhao Zhu, Bo Xu, Rui Lin, Chenyuan Wang, Yu Shao, Boling Zhu, Jiuyan Sun, Liang Zhao (possible past Baidu (China) affiliation), Hongfei Lin, Feng Xia (possible past Tencent (China) affiliation)
Abstract

Ultrasound is one of the most widely used medical imaging modalities, and recent large vision-language models(VLMs) have shown increasing capabilities in ultrasound image understanding. However, these models fail to provide pixel-level visual evidence aligned with their semantic predictions, and their fine-grained grounding capability in ultrasound remains largely unclear. We introduce UltraG-Bench, a large-scale multi-task benchmark for evaluating pixel-level evidence grounding in ultrasound. U...

📄 MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation
🗓️ Published: 9/25/2026
🔗 http://arxiv.org/abs/2609.30837v1
👥 Authors: Tianze Xu, Yanzhao Zheng, Zhentao Zhang, Yuanqiang Yu, Chao Ma (possible past Shanghai Jiao Tong University affiliation), Jihuai Zhu, Lelun Wu, Lyumanshan Ye, Pengfei Liu, Baohua Dong, Hangcheng Zhu, Ruohui Huang, Gang Yu (possible past Tencent (China) affiliation)
Abstract

Multi-teacher on-policy distillation (MOPD) integrates specialized capabilities into a single student, but existing practice typically hard-routes each prompt to a domain-matched teacher for the entire rollout. This dependence on prompt-level domain labels restricts using unlabeled training mixtures and leaves complementary signals from other teachers unused. We introduce MOPD-Router, a framework that routes supervision over the full teacher pool at each token, without domain labels or training ...

📄 Evaluation Is All You Need for Multi-Modal Autonomous Driving
🗓️ Published: 9/25/2026
🔗 http://arxiv.org/abs/2609.30818v1
👥 Authors: Zeyu He, Shiqi Liu, Ke Chen (possible past Tencent (China) affiliation), Yun Yan, Jinzi Wu, Dianqiao Lei, Sirui Wang, Shurui Peng, Tao Chen, Zhuo Huang, Yu Wu (possible past Baidu (China) affiliation), Yadong Shao, Zhichao Li (possible past Baidu (China) affiliation), Ke Sun, Yang Guan, Keqiang Li, Shengbo Eben Li (possible past Tsinghua University affiliation)
Abstract

Multi-modal planning is promising for autonomous driving by representing multiple plausible behaviors in ambiguous and long-tail scenarios. Existing methods mainly focus on improving trajectory multi-modality, enhancing trajectory representations, or reshaping the candidate distribution. Nevertheless, we identify a pronounced generation-evaluation asymmetry in multi-modal planning: despite strong oracle performance, existing planners often fail to reliably select the best available candidate, le...

📄 Analyzing and Mitigating Cost-Inefficient Behaviors in Coding Agents
🗓️ Published: 9/25/2026
🔗 http://arxiv.org/abs/2609.30725v1
👥 Authors: Yiran Hu, Nan Jiang (possible past Stanford University affiliation), Shanchao Liang, Anik Dey, Yi Wu (possible past University Of California, Berkeley affiliation), Lin Tan
Abstract

Although effective, coding agents often incur substantial monetary costs. Their recurring cost-inefficient behaviors remain underexplored. We conduct the first study of behavioral cost inefficiencies in coding agents, analyzing 1,200 trajectories from Claude Code and Mini-SWE-Agent across four configurations on SWE-bench Verified. We identify three cost-inefficient behaviors: subsumed retrieval, similar script generation, and test re-execution. We then evaluate three mitigation strategies: struc...

📄 T-RoPE: Time-Aware Rotary Position Embedding for Sequential Recommendation
🗓️ Published: 9/24/2026
🔗 http://arxiv.org/abs/2609.30576v1
👥 Authors: Yang Liu (possible past Tsinghua University affiliation), Noel Loo, Ali Khanafer, Shuying Sun (possible past Meta (United States) affiliation), Akshay Soni, Zhong Wu, Linjun Yang (possible past Microsoft (United States) affiliation)
Abstract

Large-scale recommenders increasingly adopt the sequential generative recipe behind large language models, bringing the Transformer into recommendation along with design choices made for text, including Rotary Position Embedding (RoPE). In language models, RoPE encodes token indices for relative position reasoning, but in recommendation, an interaction index records only event order, saying nothing about elapsed time, behavioral cycles across scales, or calendar phase. We revisit this choice and...

📄 A Benchmarking Framework for Context-aware XR Interfaces
🗓️ Published: 9/24/2026
🔗 http://arxiv.org/abs/2609.30466v1
👥 Authors: Hyunsung Cho, Sarah Yewon Yun, Nancy Ruonan Sun, Ben Lafreniere, Mark Parent, Kashyap Todi, Tanya R. Jonker (possible past Meta (United States) affiliation), Hrvoje Benko (possible past Meta (United States) affiliation), Sherry Tongshuang Wu, David Lindlbauer
Abstract

Everyday Extended Reality (XR) systems aim to provide context-aware access to the right functionalities at the right time and place, with minimal manual reconfiguration as users switch context. Yet these interfaces are hard to evaluate: current prototyping and user-study workflows offer no systematic, repeatable way to compare adaptation methods across users and scenarios. We present ContextXR, a novel benchmarking framework for context-aware XR interfaces. ContextXR represents an XR application...

📄 Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems
🗓️ Published: 9/24/2026
🔗 http://arxiv.org/abs/2609.30383v1
👥 Authors: Zihao Zhu, Siwei Lyu (possible past University Of Washington affiliation), Adel Bibi, Baoyuan Wu (possible past Tencent (China) affiliation)
Abstract

A skill is a modular package of natural-language instructions, executable scripts, and reference resources that an agent can load at runtime to extend its capabilities for a specific task. Skill-based agent systems therefore enable flexible reuse of third-party capabilities, but the openness of this skill ecosystem also opens up a new attack surface. Prior work has focused on vulnerabilities within individual skills, but little attention has been paid to risks that arise from interactions across...

📄 TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations
🗓️ Published: 9/24/2026
🔗 http://arxiv.org/abs/2609.30222v1
👥 Authors: Ayush Jain (possible past Google (United States) affiliation), Sreeharsha Paruchuri, Ishita Gupta, Fan Zhang, Tanner Schmidt (possible past University Of Washington affiliation), Jakob Engel (possible past Meta (United States) affiliation), Katerina Fragkiadaki (possible past University Of California, Berkeley affiliation), Adam W. Harley (possible past Carnegie Mellon University affiliation)
Abstract

Existing point tracking models face a fundamental tradeoff: they can either track a sparse set of query points over long horizons, or track all points across only short clips. We introduce TrackEverything, a 3D point tracker that breaks this trade-off by representing videos as persistent 3D scene tracks in world coordinates. Grounded in the insight that videos are 2D projections of an underlying 3D world, TrackEverything decouples model complexity from video duration, allowing it to scale with u...

📄 ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds
🗓️ Published: 9/24/2026
🔗 http://arxiv.org/abs/2609.30199v1
👥 Authors: Ming Zhang (possible past Peking University affiliation), Zhenghao Xiang, Peizhong Gao, Yujiong Shen, Yuhui Wang, Zhonghan Yue, Shihan Dou, Zhangyue Yin, Junjie Ye, Shichun Liu, Weihuang Zheng, Jiahao Chen, Jiayi Chen, Hongzhang Liu, Jiaqi Shao, Tao Gui, Qi Zhang (possible past Tencent (China) affiliation), Xuanjing Huang, Suncong Zheng, Maxm Pan
Abstract

Scientific discovery begins where known problems end. There, AI systems must engage in exploration: framing hypotheses, designing experiments, and iterating on the results. However, evaluating this ability is difficult: (1) how to verify whether a genuinely new hypothesis holds, and (2) how to determine whether a system has discovered it through exploration or merely recalled related knowledge from pre-training data. To this end, we introduce ExplorationBench, which turns the wicked problem of e...

📄 Accelerating Video Diffusion via Training-Free Trajectory Routing
🗓️ Published: 9/24/2026
🔗 http://arxiv.org/abs/2609.30096v1
👥 Authors: Mustafa Munir, Huy Vu, Shreyas Misra, Rohit Jena, Sajad Norouzi, Ali Taghibakhshi, Anis Ahmad, Anjul Patney (possible past Nvidia (United States) affiliation), Pavlo Molchanov (possible past Nvidia (United States) affiliation), Nima Tajbakhsh
Abstract

Video diffusion is computationally expensive, as it requires executing a large model across many denoising steps. Even with step-distillation, inference remains expensive because every distilled step still requires a costly model evaluation. We present TRACK: TRajectory-Aware Capacity routing via top-K selection, a heterogeneous denoising strategy that switches between compatible large and small models at selected steps, reducing the average cost per denoising evaluation. The switching steps are...

📄 SciWalker: Synthesizing Scientific Coding Problems with Operator Graphs and Execution Feedback
🗓️ Published: 9/24/2026
🔗 http://arxiv.org/abs/2609.30054v1
👥 Authors: Chenxi Li, Wenxuan Zeng, Yun Luo, Fangchen Yu, Peng Ye, Yu Cheng (possible past National University Of Singapore affiliation), Jun Zhang (possible past Tencent (China) affiliation)
Abstract

Improving the scientific coding capabilities of large language models (LLMs) requires high-quality training data. However, such data remain scarce because manually authoring realistic problems is costly and time-consuming, while systematically covering diverse scientific domains and algorithmic combinations remains challenging. To address this, we introduce SciWalker, a framework for synthesizing scientific coding problems through operator-chain sampling and execution feedback. The framework com...

📄 Advancing Model Research in AgentX: Long-Horizon Autonomy for Industrial Recommender Systems
🗓️ Published: 9/24/2026
🔗 http://arxiv.org/abs/2609.30001v1
👥 Authors: Shuang Yang, Zijie Zhuang, Changxin Lao, Pengbo Xu, Hanwen Xu, Yusheng Huang, Han Gao (possible past Tencent (China) affiliation), Guanchen Wang, Tianbao Ma, Linxun Chen, Peilin Song, Xuming Wang, Chen Li (possible past Tencent (China) affiliation), Fan Wu, Tao Wang (possible past Stanford University affiliation), Zibo Zhao, Xiangyu Wu, An Liu, Fei Pan, Peng Jiang, Chen Yang (possible past Tencent (China) affiliation), Zhaojie Liu, Wenwu Ou
Abstract

Sustaining industrial recommendation research requires using the results of one experiment to decide what to investigate next. We present AgentX-Model, the next generation of AgentX's model research framework, which connects proposal development and model experimentation within sandboxes defined by business inputs and prediction tasks. AgentX-Model adopts a dual-agent architecture comprising a Research Agent and a Model Agent. The Research Agent develops independently reviewed proposals from pap...

📄 From Interests to Semantic IDs: Retrieval-Grounded Credit Assignment for Generative Recommendation
🗓️ Published: 9/24/2026
🔗 http://arxiv.org/abs/2609.29983v1
👥 Authors: Mengdan Zhu, Yufan Zhao, Yao Zhao (possible past Microsoft (United States) affiliation), Sophie Di, Tao Di, Yulan Yan, Sridhar Iyer, Liang Zhao (possible past Baidu (China) affiliation)
Abstract

Semantic IDs (SIDs) encode each catalog item as a short token sequence, enabling generative recommenders to predict the next item autoregressively. Reasoning-enhanced variants, an increasingly common extension, first generate a textual trace and then decode a next-item SID by beam search. Such recommenders are commonly trained with group-relative policy optimization under an exact-match SID reward, which is sparse in large catalogs. Two failure modes follow. When all rollouts in a group miss the...

📄 Probabilistic Robustness-driven Universal Adversarial Perturbations with Explainability against Deep Reinforcement Learning-based Intrusion Detection System
🗓️ Published: 9/24/2026
🔗 http://arxiv.org/abs/2609.30605v1
👥 Authors: Hongsen Zhang, Lu Zhang (possible past Tencent (China) affiliation), Mingjing Xu, Yi Zhang (possible past Google (United States) affiliation), Gregory Epiphaniou, Carsten Maple
Abstract

Deep reinforcement learning (DRL) enables adaptive intrusion detection in dynamic network environments but also exposes intrusion detection systems (IDS) to adversarial threats such as universal adversarial perturbations (UAPs), which apply a single input-agnostic perturbation to degrade detection performance across traffic. Probabilistic Robustness (PR), as a post-hoc evaluation metric, provides a principled, population-level measure of adversarial impact that conceptually aligns with the unive...

*Notable papers are those with at least two authors from a "big" AI/ML lab.