📄 Notable* Recent AI/ML arXiv Papers

Last updated just now...

📄 TurboBias 2.0: Streaming Context-Biasing for Production-Efficient ASR Systems
🗓️ Published: 8/21/2026
🔗 http://arxiv.org/abs/2608.21343v1
👥 Authors: Vladimir Bataev, Lilit Grigoryan, Andrei Andrusenko, Nikolay Karpov, Vitaly Lavrukhin (possible past Nvidia (United States) affiliation), Boris Ginsburg (possible past Nvidia (United States) affiliation)
Abstract

Contextualization is essential for production automatic speech recognition (ASR) systems, where user-provided phrases must be recognized accurately under strict latency constraints. Although many context-biasing methods improve recognition accuracy, they often do not address the practical requirements of modern production ASR systems: streaming inference, efficient batched decoding, user-specific context lists, and low runtime overhead. We propose TurboBias 2.0, a production-oriented framework f...

📄 SRL-MPC: Shape-Aware Reinforcement Learned Model Predictive Control
🗓️ Published: 8/21/2026
🔗 http://arxiv.org/abs/2608.21175v1
👥 Authors: Ruihua Han, Rui Gao, Zhe Liu, Xinyi Wang (possible past Carnegie Mellon University affiliation), Chang Chen, Shuai Wang, Qi Hao, Jia Pan, Hengshuang Zhao (possible past University Of Oxford affiliation)
Abstract

Safe and efficient shape-aware navigation in heterogeneous crowds and robot fleets remains challenging. Traditional approaches often assume homogeneous robots, sparse workspaces, simplified geometry, offline computation, or handcrafted parameters to make the problem tractable, which limits their deployment in dense crowd scenarios. Toward this end, we propose Shape-Aware Reinforcement Learned Model Predictive Control (SRL-MPC), a method for safe, efficient, and adaptive navigation in crowds with...

📄 Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence
🗓️ Published: 8/21/2026
🔗 http://arxiv.org/abs/2608.21156v1
👥 Authors: Yuyuan Feng, Zhishang Xiang, Chaobin Yang, Qichao Ma, Zerui Chen, Yujing Zhang, Ke Huang, Chuanjie Wu, Zhaoxu Liu, Yili Wang, Xin He, Jiapu Wang, Zijin Hong, Hao Chen, Yuanchen Bei, Kun Wang, Shengyuan Chen, Ningyu Zhang (possible past Tencent (China) affiliation), Enyan Dai, Linhao Luo, Qingyi Pan, Qi Wang (possible past Tsinghua University affiliation), Wenqi Fan, Guangjing Wang, Na Zou, Yangqiu Song (possible past Tsinghua University affiliation), Xin Wang (possible past University Of Edinburgh affiliation), Zechao Li, Xia Hu, Qing Li, Xiao Huang, Zhihong Zhang, Jinsong Su, Qinggang Zhang, Yi Chang
Abstract

LLMs have evolved from language generators to autonomous agents capable of complex, long-horizon tasks. This evolution has produced paradigms including Prompt Engineering to elicit model capabilities, Context Engineering to manage information access, Harness Engineering to organize external tools and resources, and Loop Engineering to support continual reflection and self-improvement. Yet as tasks grow more complex, individual intelligence faces a fundamental limit: many tasks require heterogene...

📄 CIVA: Critic-Induced Value-Subspace Attacks on Visual World-Model Agents
🗓️ Published: 8/21/2026
🔗 http://arxiv.org/abs/2608.21114v1
👥 Authors: Jiancheng Wang, Mingli Zhu, Tong Zhang (possible past Tencent (China) affiliation), Jiaqi Ruan, Wei Wang (possible past University Of Oxford affiliation), Siyuan Liang, Dacheng Tao
Abstract

Visual world-model agents such as DreamerV3 act through a recurrent latent state rather than a single observation, which weakens frame-wise observation attacks and makes their perturbations vary sharply over time under a strict per-frame perturbation constraint. We study white-box, causal, online attacks on such agents and propose Critic-Induced Value-Subspace Attacks (\textbf{CIVA}). Our key observation is that, along a rollout, critic-guided perturbations concentrate in a low-dimensional subsp...

📄 Advantage-level Aggregation Reinforcement Learning for X-point Target Magnetic Configuration Control in an EXL-50U Experiment-Calibrated Simulation Environment
🗓️ Published: 8/21/2026
🔗 http://arxiv.org/abs/2608.20834v1
👥 Authors: Siqi Ding, Xuanhe Wang, Pei Guo, Guoyang Shi, Changquan Yu, Yiting Wang, Xianming Song, Xiang Gu, Zhengyuan Chen, Lei Xing (possible past Stanford University affiliation), Yapeng Zhang (possible past Baidu (China) affiliation), Jianguo Chen, Tianyuan Liu
Abstract

Managing divertor heat loads is a central challenge for compact, high-power tokamaks. To increase local flux expansion and decouple the dissipation volume from the core, EHL-2 adopts the X-point target (XPT) divertor. This requires the secondary X-point to remain on the divertor leg; displacement degrades the topology and exhaust geometry. Current experiments, including EXL-50U discharges, rely on precomputed feedforward waveforms with PID loops on global quantities. Lacking dedicated closed-loo...

📄 Scaling Muon for Diffusion Transformers
🗓️ Published: 8/21/2026
🔗 http://arxiv.org/abs/2608.20818v1
👥 Authors: Chenghao Li, Xiao Han (possible past Tencent (China) affiliation), Xinxin Huang, Wei Liu (possible past Tsinghua University affiliation), Boyang Li, Bing Xiao, Heran Zhang, Juanma Perez Rua, Ke Xu, Kangning Liu, Linjun Kuang, Na Li (possible past Tencent (China) affiliation), Tan Wang, Tian Xie, Wei Peng, Yang Pei, Yifan Xu (possible past Tencent (China) affiliation), Yuanhao Zhai, Yuwei Lin, Zhe Wang (possible past Deepmind (United Kingdom) affiliation), Zihao He, Daniel Li, Junbiao Tang, Ziyang Jiang, Dake Chen
Abstract

The matrix-aware optimizer Muon improves large model training by balancing updates across singular directions, yet its scaling behavior and end-to-end efficiency on large Diffusion Transformers (DiTs) remain unclear. We first establish Muon's scaling behavior on DiTs from 1.3B to 15B parameters, showing that its optimization and generative quality advantages over AdamW persist across model scales. However, at scale, the 5-step Newton--Schulz iteration (NS5) performed at every optimization step, ...

📄 Denoising the Future: Context-Aware Spectral Diffusion for Temporal Knowledge Graph Extrapolation
🗓️ Published: 8/21/2026
🔗 http://arxiv.org/abs/2608.20804v1
👥 Authors: Yanglei Gan, Peng He (possible past Tencent (China) affiliation), Run Lin, Peiyuan Jiang, Yifan Wang (possible past Stanford University affiliation), Qiao Liu
Abstract

Temporal Knowledge Graph (TKG) extrapolation seeks to infer future facts from time-varying relational histories. Recent diffusion-based approaches improve uncertainty modeling through generative denoising, but their aggregated conditioning on subject histories may insufficiently distinguish query-specific evidence from non-salient historical facts, thereby diluting target-discriminative signals. To bridge this gap, we propose FreqDiff, a Frequency-aware Diffusion framework for TKG extrapolation....

📄 Continuous-Time Quantum Walks based Graph Neural Network
🗓️ Published: 8/21/2026
🔗 http://arxiv.org/abs/2608.20738v1
👥 Authors: Yuliang Zhan, Zefeng Gao, Jian Li (possible past Tencent (China) affiliation), Yang Liu (possible past Tsinghua University affiliation), Hao Sun
Abstract

Graph Neural Networks (GNNs) are widely used on graph-structured data, but most suffer from two key weaknesses. First, message passing behaves as a low-pass filter under the homophily assumption, leading to poor performance on heterophilic graphs. Second, stacking layers drives node features toward constants, causing over-smoothing. Existing methods usually address these issues separately, while the few joint solutions rely largely on empirical heuristics, and many over-smoothing remedies sacrif...

📄 Volumetric Radiology AI in the Era of Multimodal Large Language Models
🗓️ Published: 8/20/2026
🔗 http://arxiv.org/abs/2608.20549v1
👥 Authors: Zanting Ye, Shengyuan Liu, Xin Liu, Chenhui Wang, Zhisong Wang, Jiashuai Liu, Zipei Wang, Cheng Wang, Wentao Pan (possible past Tsinghua University affiliation), Mengjie Fang, Di Dong, Mohammad Salmanpour, Arman Rahmim, Yu Gu, Yong Xia, Hongming Shan, Yixuan Yuan, Yefeng Zheng (possible past Tencent (China) affiliation), Lijun Lu
Abstract

Advances in multimodal large language models (MLLMs) are extending radiological artificial intelligence (AI) beyond task-specific image analysis toward multimodal understanding and reasoning. Volumetric radiology, however, presents a fundamental representational mismatch: clinical interpretation often requires full-volume spatial context and acquisition-dependent quantitative information, whereas current MLLMs are commonly conditioned on selected two-dimensional (2D) images, compressed visual re...

📄 Terminal Agents: A Survey of AI Agents in Command-Line Environments
🗓️ Published: 8/20/2026
🔗 http://arxiv.org/abs/2608.20485v1
👥 Authors: Yi Bin, Xiaoyang Yuan, Haoxi Zeng, Wencheng Ye, Wenqi Shao, Chen Qian (possible past Shanghai Jiao Tong University affiliation), Wei Ye (possible past Meta (United States) affiliation), Yujuan Ding, Zheng Wang, Pengpeng Zeng, Jingkuan Song, Heng Tao Shen
Abstract

Large language model agents increasingly act through terminals, yet existing surveys disperse terminal-mediated behavior across software engineering, tool use, and computer-use research. We regard terminal agents as systems whose dominant progress-bearing action--observation loop is mediated by terminal command execution, textual feedback, and stateful environment interaction. Using terminal-mediated execution as an organizing lens, this survey establishes workload-level boundaries and connects ...

📄 Pandora's AI Model Routing Box: Efficient Allocation with Costly Value Estimation
🗓️ Published: 8/20/2026
🔗 http://arxiv.org/abs/2608.20316v1
👥 Authors: Adam Fisch (possible past University Of Washington affiliation), Shubhendu Trivedi, Fantine Huot, William W. Cohen (possible past Google (United States) affiliation), Michael Kaisers, Mirella Lapata (possible past University Of Edinburgh affiliation), Kate Larson, Jacob Eisenstein (possible past Meta (United States) affiliation)
Abstract

Heterogeneous AI systems composed of multiple models, architectures, harnesses, or inference-time settings can improve quality and efficiency by routing queries to the specialist who can answer most effectively at the lowest cost. Routing requires estimating each specialist's expected return, but this value estimation has a cost. Cheap estimators (e.g., embedding-based predictors) are fast but noisy, while accurate estimators (e.g., fine-tuned models with access to retrieval results or partial r...

📄 Rule-Compliant Visual Spatial Planning for Multimodal Large Language Models
🗓️ Published: 8/20/2026
🔗 http://arxiv.org/abs/2608.20237v1
👥 Authors: Yu Chen (possible past Meta (United States) affiliation), Ting Lei, Yaoyi Li, Jia Cai, Zhecen Wu, Yang Liu (possible past Tsinghua University affiliation)
Abstract

Multimodal large language models (MLLMs) combine linguistic reasoning with visual perception, yet their ability to perform visual spatial planning under explicit or previously unseen rule constraints remains underexplored. This setting requires models to jointly understand spatial layouts, interpret natural-language rules, and plan valid actions accordingly. To address this gap, we introduce RuleMaze, a controllable benchmark in which MLLMs must navigate mazes while obeying natural-language rule...

📄 Towards Professional Tennis Styles for Humanoid Robots with Adaptive Motion Planning and Tracking
🗓️ Published: 8/20/2026
🔗 http://arxiv.org/abs/2608.20087v1
👥 Authors: Tao Huang, Ruofei Liu, Xuchen Tang, Xinyin Zhang, Junli Ren, Huayi Wang, Feiyu Jia, Yukai Qi, Kangning Yin, Weishuai Zeng, Lipeng Chen, Xi Li, Ting Wu, Kailin Li, Ruoli Dai, Jingbo Wang, Lei Han (possible past Tencent (China) affiliation), Jiangmiao Pang (possible past Shanghai Artificial Intelligence Laboratory affiliation)
Abstract

Humanoid robots have recently demonstrated promising capabilities in real-world ball sports. However, achieving professional motion styles while maintaining strong task performance remains challenging. In this work, we propose AdaPT, an Adaptive Motion Planning and Tracking framework that learns professional tennis serving and rally styles directly from broadcast videos. This hierarchical design is motivated by the key insight that the planner generates stylistic kinematic motions, while the tra...

📄 Asymmetric Capacity Allocation in Self-Refinement Pipelines
🗓️ Published: 8/21/2026
🔗 http://arxiv.org/abs/2608.21345v1
👥 Authors: Zhuoyi Yang, Ian G. Harris, Salar Hashemitaheri, Cassie Huang, Yuangang Li, Hyunwoo Oh, Paul Dourish (possible past Apple (United States) affiliation), Tony Givargis, Mohsen Imani, Li Zhang (possible past University Of Oxford affiliation)
Abstract

Self-refinement, typically structured as generation, critique, and revision, is a widely adopted paradigm for improving LLM generation and serves as a core mechanism in many LLM agents. While the three stages involve different cognitive demands, most existing approaches conveniently treat the model size as an implementation detail rather than a subject of study, which may lead to a waste of resources. Little work has systematically examined how model size affects each stage or whether effective ...

📄 COEC: Calibrated Orthogonal-Equivalence Compensation for Structured Pruning of Large Language Models
🗓️ Published: 8/21/2026
🔗 http://arxiv.org/abs/2608.21142v1
👥 Authors: Peiqi Yu, Nam Ling, Wei Wang (possible past University Of Oxford affiliation), Wei Jiang (possible past Apple (United States) affiliation)
Abstract

Structured pruning reduces the size and inference cost of large language models (LLMs) by removing weight columns, but the resulting output error can degrade accuracy. Existing training-free compensation methods use an additive bias or a single orthogonal rotation on the output side of the retained weight. These corrections leave its input singular frame unchanged and therefore limit how the retained weight can adapt after column removal. We propose COEC (Calibrated Orthogonal-Equivalence Compen...

📄 CubicSplat: Differentiable Vector Graphics via Error-Bounded Forward Relaxation
🗓️ Published: 8/21/2026
🔗 http://arxiv.org/abs/2608.20803v1
👥 Authors: Chenglong Liu, Xin Zhang (possible past Google (United States) affiliation), Yimeng Zhu, Liyang He, Yixiao Ma, Yu Su, Zhenya Huang, Qi Liu (possible past Tencent (China) affiliation)
Abstract

Vector graphics are prized for their resolution independence, compact storage, and direct editability, making differentiable optimization of their parametric primitives an attractive goal. Yet classical rasterization is discontinuous with respect to geometry, and existing remedies that smooth the forward pass demand increasingly elaborate heuristics as scene complexity grows. We trace this fragility to a gradient seesaw: design choices that improve forward geometric exactness can systematically ...

📄 Metag: A dataset to build agentic meta-reviewing capabilities
🗓️ Published: 8/20/2026
🔗 http://arxiv.org/abs/2608.20488v1
👥 Authors: Anirudh Sundar, Min Chen, Divya Tadimeti, Gemma Zhang, Alice Li, Nigel Boachie Kumankumah, Pavan Uttej Ravva, Sadid Hasan, Somya Chatterjee, Pruthvi Prakash Navada, Xiao Wang (possible past Google (United States) affiliation), Yue Kang, Sulaiman Vesal (possible past Stanford University affiliation), Larry Heck (possible past Google (United States) affiliation)
Abstract

AI tools increasingly support tasks across the scientific research cycle, from experiment design and manuscript preparation to peer review. At the same time, the continuing growth in conference submissions has increased the burden on meta-reviewers, who must synthesize reviewer feedback, author rebuttals, and manuscript revisions. To address this concern, this paper introduces Metag, a dataset to accelerate the development of meta-reviewing agents, specifically to identify changes made to scient...

📄 EnvHarness: Awakening Static Worlds for Agent Learning
🗓️ Published: 8/20/2026
🔗 http://arxiv.org/abs/2608.19880v1
👥 Authors: Chengsong Huang, Zifeng Wang, Rujun Han, Jun Yan, Yanfei Chen, Zoey Cuizhu, Ke Jiang, Peng Xia, Han Yu, Yufan Zhuang, Yifei Ming, Jiaqi Pan, Bhavana Dalvi Mishra (possible past Carnegie Mellon University affiliation), Jiaxin Huang, Burak Gokturk, Tomas Pfister (possible past University Of Oxford affiliation), Chen-Yu Lee (possible past Google (United States) affiliation)
Abstract

LLM agents learn by interacting with environments, yet these environments are hand-built and static: blind to an agent's weaknesses, and quickly left behind as it improves. While recent environment generation methods attempt to address this, they require domain-specific pipelines, rely on expensive or unreliable verifiers, and still produce static environments. To alleviate the engineering burden of rebuilding environments from scratch, we propose Environment Harness (EnvHarness), a programmable...

📄 Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning
🗓️ Published: 8/20/2026
🔗 http://arxiv.org/abs/2608.19669v1
👥 Authors: Haoqiang Kang, Yinpeng Chen, Luyang Liu (possible past Google (United States) affiliation), Jesper Sparre Andersen, Abhijit Ogale, Baochen Sun, Lichan Hong (possible past Google (United States) affiliation), Ed H. Chi
Abstract

Latent reasoning has advanced multimodal reasoning through a two-stage training paradigm: (1) a helper image is encoded into latent tokens to teach visual chain-of-thought during a supervised fine-tuning (SFT) stage, and (2) these latent tokens are further refined with reward feedback during a reinforcement learning (RL) stage. In this paper, we identify two key limitations of this framework, one in each stage. First, the SFT stage typically relies on an off-the-shelf vision encoder to encode th...

📄 Forking Fast: Efficiently Estimating Uncertainty Dynamics in Text Generation
🗓️ Published: 8/20/2026
🔗 http://arxiv.org/abs/2608.19611v1
👥 Authors: Eric Bigelow, Amir Zur, Satchel Grant, Tal Haklay, Can Rager, Owen Lewis, Thomas Mcgrath (possible past Google (United States) affiliation), Jack Merullo, Ekdeep Singh Lubana, Atticus Geiger (possible past Stanford University affiliation)
Abstract

LLM reasoning is stochastic, and so understanding a model requires grappling with the distribution of reasoning chains that it might produce for a given question, i.e., its uncertainty. Resampling-based analyses characterize this distribution, revealing which steps of a rollout determine how the model arrives at its answer. However, a major limitation of these approaches is that resampling text sequences at every token or sentence in a reasoning chain is very costly. Our work strives to make res...

*Notable papers are those with at least two authors from a "big" AI/ML lab.