πŸ“„ Notable* Recent AI/ML arXiv Papers

Last updated just now...

πŸ“„ Understanding Reasoning from Pretraining to Post-Training
πŸ—“οΈ Published: 7/17/2026
πŸ”— http://arxiv.org/abs/2607.16097v1
πŸ‘₯ Authors: Jingyan Shen, Ang Li (possible past Google (United States) affiliation), Salman Rahman, Yifan Sun (possible past Baidu (China) affiliation), Micah Goldblum, Matus Telgarsky, Pavel Izmailov
Abstract

Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, yet RL post-training is largely studied in isolation from the pretraining that precedes it. As a result, two basic questions remain open: (1) how do pretraining choices (model size, data) shape the returns to RL compute, and (2) what does RL actually do to the model? These questions are difficult to study in the standard LLM setting: pretraining corpora are vast and uncontrolled, ...

πŸ“„ JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models
πŸ—“οΈ Published: 7/17/2026
πŸ”— http://arxiv.org/abs/2607.16074v1
πŸ‘₯ Authors: Haoran Sun, Wentao Zhang (possible past Mila - Quebec Artificial Intelligence Institute affiliation), Junyang Hua, Hedan Yang, Yongjian Guo, Yifei Zhang, Xiaolong Xiang, Mingxi Luo, Jing Long, Chen Zhao (possible past Stanford University affiliation), Chen Zhou, Wanting Xu, Qiming Yang, Hui Zhang, Song Wang, Xiaodong Bai, Shuai Di, Xu Chu, Xiaotie Deng, Yicheng Gong, Junwu Xiong
Abstract

The post-training of Vision-Language-Action (VLA) models is essential due to the diversity of simulators, robot embodiments, and task objectives. Existing compute services, whether offered as direct accelerator rental or batch-workload submission, typically allocate an exclusive set of GPU and CPU resources to a single tenant. While this paradigm maximizes client flexibility, it burdens users with infrastructure adaptation, and the fixed card-hour accounting model renders short or bursty workloa...

πŸ“„ Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding
πŸ—“οΈ Published: 7/17/2026
πŸ”— http://arxiv.org/abs/2607.15778v1
πŸ‘₯ Authors: Wei Feng, Xin Wang (possible past University Of Edinburgh affiliation), Yu-Wei Zhan, Yuwei Zhou, Wenwu Zhu (possible past Tsinghua University affiliation)
Abstract

Video Large Language Models (Video LLMs) have made significant advancements in various video understanding tasks. However, long-video scenarios remain challenging due to the tension between limited visual token budgets and the need to capture multiple key events. Existing approaches typically process long videos in two stages, i.e., i) select keyframes and ii) perform detailed perception, which exhibit limitations: they lack a modular mechanism for adaptive capacity allocation and self-correctio...

πŸ“„ Scalable LLM Agent Tool Access in the Cloud
πŸ—“οΈ Published: 7/17/2026
πŸ”— http://arxiv.org/abs/2607.15593v1
πŸ‘₯ Authors: Mingxin Li, Enge Song, Yueshang Zuo, Xiaodong Liu, Rong Wen, Qiang Fu (possible past Tencent (China) affiliation), Gianni Antichi (possible past University Of Cambridge affiliation), Jian He, Jing Tie, Zhou Shao, Xiaobo Xue, Xiong Xiao, Luyao Zhong, Shaokai Zhang, Jiangu Zhao, Jianyuan Lu, Shize Zhang, Xiaoqing Sun, Changgang Zheng, Zihao Fan, Haonan Li, Tian Pan, Xiaomin Wu, Yang Song (possible past Stanford University affiliation), Xing Li, Biao Lyu, Meng Li (possible past Meta (United States) affiliation), Haipeng Dai, Guihai Chen (possible past Shanghai Jiao Tong University affiliation), Shunmin Zhu
Abstract

LLM agents increasingly rely on tool calling to act on external systems, and the Model Context Protocol (MCP) has quickly become its de facto interface. Operating MCP at cloud scale, however, becomes difficult. On the tool provider side, legacy services are not directly callable through MCP; the rapid protocol development also creates ongoing compatibility cost. On the agent side, the number of accessible tool is limited by the LLM context window and inference overhead; mounting a large tool set...

πŸ“„ MGDT: MLLM-Guided Diffusion Transformer with Relation-Adaptive Mixture-of-Experts for Multimodal Knowledge Graph Completion
πŸ—“οΈ Published: 7/17/2026
πŸ”— http://arxiv.org/abs/2607.15592v1
πŸ‘₯ Authors: Xu Hou, Meiyu Liang, Wei Huang (possible past Google (United States) affiliation), Yawen Li, Zhe Xue, Wu Liu, Guanhua Ye, Lei Shi (possible past Baidu (China) affiliation), Kangkang Lu
Abstract

Multimodal Knowledge Graph Completion (MKGC) requires inferring missing entities from structural, textual, and visual cues. Existing diffusion-based MKGC methods usually denoise directly on raw multimodal features. Such a design forces the denoiser to simultaneously perform relation-dependent cue selection, cross-modal semantic alignment, and structure-aware entity generation, which introduces noisy and semantically inconsistent conditions for diffusion and consequently leads to suboptimal compl...

πŸ“„ RoboTTT: Context Scaling for Robot Policies
πŸ—“οΈ Published: 7/16/2026
πŸ”— http://arxiv.org/abs/2607.15275v1
πŸ‘₯ Authors: Yunfan Jiang, Yevgen Chebotar (possible past Google (United States) affiliation), Ruijie Zheng, Fengyuan Hu, Yunhao Ge, Jimmy Wu, Tianyuan Dai, Scott Reed (possible past Google (United States) affiliation), Li Fei-Fei (possible past Stanford University affiliation), Yuke Zhu (possible past Stanford University affiliation), Linxi "jim" Fan
Abstract

Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude beyond state-of-the-art policies, without growing inference latency. At this context length, we unlock new robot capabilities: one-shot in-context imitation from human video demonstrations, on-the-fly policy improvement, robustness to perturb...

πŸ“„ SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper Revisions
πŸ—“οΈ Published: 7/16/2026
πŸ”— http://arxiv.org/abs/2607.15272v1
πŸ‘₯ Authors: Yasheng Sun, Zezi Zeng, Yifan Yang (possible past Tencent (China) affiliation), Chong Luo (possible past Google (United States) affiliation), Wenyi Wang, Ziwei Liu, JΓΌrgen Schmidhuber
Abstract

Editing the figures in a research paper is a routine and time-consuming part of everyday research practice: authors relabel components, rearrange panels, and restyle visuals as they revise their manuscripts. Automating this editing workflow under a natural-language instruction, however, is challenging, because a scientific figure is a dense infographic in which heterogeneous visual elements such as schematics, plots, photos, captions, and arrows are composed under a tight visual grammar to advan...

πŸ“„ Pretraining Data Can Be Poisoned through Computational Propaganda
πŸ—“οΈ Published: 7/16/2026
πŸ”— http://arxiv.org/abs/2607.15267v1
πŸ‘₯ Authors: Victoria Graf, Hannaneh Hajishirzi (possible past University Of Washington affiliation), Noah A. Smith (possible past University Of Washington affiliation), David Kohlbrenner, Kyle Lo
Abstract

Poisoning pretraining data can introduce harmful behaviors to LMs that are difficult to detect and mitigate. Prior work on poisoning pretraining data has largely exploited established data sources such as Wikipedia, which do not represent the large scale and heterogeneity typical of pretraining corpora, and has ignored the interaction between poisoned data and data curation pipelines. We demonstrate that poisoning attacks on pretraining data are feasible beyond this limited setting through an ex...

πŸ“„ Scaling Behavior Foundation Model for Humanoid Robots
πŸ—“οΈ Published: 7/16/2026
πŸ”— http://arxiv.org/abs/2607.15163v1
πŸ‘₯ Authors: Weishuai Zeng, Kangning Yin, Xiaojie Niu, Shunlin Lu, Weixiang Zhong, Jiahe Chen, Feiyu Jia, Xiao Chen, Zirui Wang, Furui Xu, Ming Zhou, Kailin Li, Weinan Zhang (possible past Shanghai Jiao Tong University affiliation), He Wang (possible past Stanford University affiliation), Li Yi (possible past Stanford University affiliation), Dahua Lin, Jiangmiao Pang (possible past Shanghai Artificial Intelligence Laboratory affiliation), Jingbo Wang
Abstract

Humanoid control requires natural whole-body coordination, precise real-time responses to control signals, and robust generalization across diverse environmental contexts, making it a cornerstone for generalist embodied agents. Behavior Foundation Models (BFMs) have recently emerged as a promising solution to address these challenges by leveraging large-scale behavioral data to achieve superior expressiveness, versatility and generalization. However, despite growing interest in scaling BFMs to f...

πŸ“„ InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring
πŸ—“οΈ Published: 7/16/2026
πŸ”— http://arxiv.org/abs/2607.14683v1
πŸ‘₯ Authors: Hao Yang (possible past Tencent (China) affiliation), Yanyan Zhao, Kewei Zhao, Hongbo Zhang, Tian Zheng, Yusheng Liu, Xing Fu, Bichen Wang, Yu Zhang (possible past Google (United States) affiliation), Hao He, Zhen Wu, Xuda Zhi, Yongbo Huang, Bing Qin
Abstract

Understanding driver emotion and state is critical for the next generation of intelligent in-cabin systems that ensure safety and enhance human-vehicle interaction. However, existing public datasets for in-cabin affective computing are largely limited to visual modalities and rarely include conversational information, making it difficult to capture the linguistic and interactive cues underlying driver emotion. To address these gaps, we introduce InCarEmo, a multimodal dataset for in-cabin emotio...

πŸ“„ Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization
πŸ—“οΈ Published: 7/16/2026
πŸ”— http://arxiv.org/abs/2607.14614v1
πŸ‘₯ Authors: Weiwen Xu, Jia Liu, Hou Pong Chan, Long Li, Deng Cai (possible past Shanghai Jiao Tong University affiliation), Min Chen, Hao Zhang (possible past Tencent (China) affiliation)
Abstract

Reinforcement learning with verifiable rewards (RLVR) commonly uses entropy for advantage shaping. However, entropy cannot distinguish useful uncertainty from detrimental confusion, limiting its effectiveness as a correctness signal. We propose Contrastive Policy Optimization (CPO), which uses token-level contrastive disagreement between reference-guided and vanilla generation distributions for correctness-aware advantage shaping. Both theoretical and empirical results show that this disagreemen...

πŸ“„ MathCoPilot: An Interactive System for Human-AI Symbiotic Paradigm of Mathematical Research
πŸ—“οΈ Published: 7/16/2026
πŸ”— http://arxiv.org/abs/2607.14582v1
πŸ‘₯ Authors: Junjie Zhang, Jiayu Liu, Wenbin Liu, Zhenya Huang, Doudou Wang, Yan Jiang, Leiye Xu, Tao Xiong, Wen Huang, Qi Liu (possible past Tencent (China) affiliation), Guoping Hu, Enhong Chen (possible past Baidu (China) affiliation), Mengping Zhang, Xiangdong Ye
Abstract

Existing LLM-based theorem provers have achieved impressive results on formal mathematics benchmarks, yet they remain confined to acting as autonomous agents that prove a stated proposition. In this paper, we propose MathCoPilot, a human-in-the-loop system that embodies a new human--AI symbiotic paradigm for mathematical research, in which the mathematician steers the high-level mathematical direction while AI agents carry out the detailed formalization and proof work under continuous human guid...

πŸ“„ Adaptive Multi-Step Lookahead Decoding for Diffusion Language Models
πŸ—“οΈ Published: 7/17/2026
πŸ”— http://arxiv.org/abs/2607.15655v1
πŸ‘₯ Authors: Yingqian Cui, Wei Deng (possible past Apple (United States) affiliation), Lantao Mei, Hang Li (possible past Huawei Technologies (China) affiliation), Charu C. Aggarwal, Hui Liu, Yue Xing
Abstract

Masked diffusion language models (DLMs) enable parallel text generation by iteratively refining masked tokens, offering a promising alternative to autoregressive decoding. Recent lookahead-based decoding methods improve the accuracy--efficiency trade-off by exploring future decoding states before committing token updates. However, existing approaches mainly rely on shallow one-step lookahead, which optimizes immediate information gain but can be suboptimal for longer-horizon decoding trajectorie...

πŸ“„ MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators
πŸ—“οΈ Published: 7/16/2026
πŸ”— http://arxiv.org/abs/2607.15273v1
πŸ‘₯ Authors: Yushi Huang, Xiangxin Zhou, Jun Zhang (possible past Tencent (China) affiliation), Liefeng Bo, Tianyu Pang (possible past Tsinghua University affiliation)
Abstract

MeanFlow generators achieve fast few-step sampling by predicting average velocities over time intervals, making them attractive for efficient generation. Reinforcement learning (RL) has become a powerful way to align diffusion and flow models with human preferences and task-specific objectives. In particular, DiffusionNFT offers an efficient forward-process RL framework that does not require reverse-process trajectories or likelihood estimation. However, applying such RL methods to MeanFlow rema...

πŸ“„ Online Neural Space Time Memory for Dynamic Novel View Synthesis
πŸ—“οΈ Published: 7/16/2026
πŸ”— http://arxiv.org/abs/2607.15271v1
πŸ‘₯ Authors: Baback Elmieh, Lynn Tsai, Zeman Li, Srinivas Kaza (possible past Google (United States) affiliation), Tiancheng Sun, Gabor Csapo, Ali Behrouz, Yuan Deng, Stephen Lombardi, Steven M. Seitz (possible past University Of Washington affiliation), Xuan Luo (possible past University Of Washington affiliation)
Abstract

Online novel view synthesis from multi-view streaming videos faces a fundamental trade-off: maintaining a persistent, long-horizon memory to reconstruct temporarily occluded regions while operating under strict real-time constraints. While Test-Time Training (TTT) offers a powerful memory mechanism, standard models mandate gradient-based memory updates at every frame to adapt to the changing motion in dynamic scenes. The computational cost of heavy memory updates precludes real-time application ...

πŸ“„ Causal Inference for Sequential Settings under Interference and Latent Confounding
πŸ—“οΈ Published: 7/16/2026
πŸ”— http://arxiv.org/abs/2607.14940v1
πŸ‘₯ Authors: Phevos Paschalidis, Constantinos Daskalakis (possible past University Of California, Berkeley affiliation), Devavrat Shah (possible past Massachusetts Institute Of Technology affiliation)
Abstract

We study causal inference under outcome interference for sequential, observational settings. Specifically, we consider settings where the binary outcomes over N units are Markovian across T time steps. At each time step, the outcomes of N units have dependencies captured through an Ising model; each outcome is also impacted through an external field capturing the effects of its treatment as well as latent confounders. Similar to panel data literature, these latent confounders are modeled to have...

πŸ“„ TIDE: Trustworthy and Interpretable Battery Degradation Estimation with Contextual Learning and Symbolic Distillation
πŸ—“οΈ Published: 7/16/2026
πŸ”— http://arxiv.org/abs/2607.14640v1
πŸ‘₯ Authors: Wen Yang Tan, Jiawei Li, Fang Liu (possible past Massachusetts Institute Of Technology affiliation), Wei Zhang (possible past Tsinghua University affiliation), Sumei Sun, Peng Cheng Wang, Elisa Y. M. Ang
Abstract

Battery health estimation is fundamental for battery management in battery-powered systems, where inaccurate health states may affect control, maintenance, and service life. It becomes even more critical in intelligent connected systems, where estimation errors can propagate across interconnected devices and downstream decisions. In this paper, we propose TIDE, a trustworthy and interpretable battery degradation estimator for reliable battery health estimation. TIDE jointly considers accuracy, t...

πŸ“„ xHC: Expanded Hyper-Connections
πŸ—“οΈ Published: 7/16/2026
πŸ”— http://arxiv.org/abs/2607.14530v1
πŸ‘₯ Authors: Xiangdong Zhang, Xiaohan Qin, Sunan Zou, Tuo Dai, Xiaoming Shi, Huaijin Wu, Yebin Yang, Zhuo Xia, Shaofeng Zhang, Lin Yao, Yuliang Liu, Yu Cheng (possible past National University Of Singapore affiliation), Junchi Yan (possible past Shanghai Jiao Tong University affiliation)
Abstract

Hyper-Connections (HC) expand the residual stream of Transformers into $N$ parallel streams, providing a form of memory scaling beyond model width and depth. Manifold-Constrained HC (mHC) stabilizes this formulation at scale. The large gains from $N{=}1$ to $N{=}4$ suggest residual-stream expansion as a promising scaling axis. However, existing HC-family methods typically stop at $N{=}4$. Our experiments reveal why: scaling mHC beyond this point yields diminishing performance gains and rapidly i...

πŸ“„ RetroAgent: Harnessing LLMs to Search Over Structured Memory for Agentic Retrosynthesis Planning
πŸ—“οΈ Published: 7/16/2026
πŸ”— http://arxiv.org/abs/2607.14512v1
πŸ‘₯ Authors: Yanqiao Zhu, Jingru Gan, Xiaoqi Sun, Fang Sun, Yidan Shi, Md Mofijul Islam, Chao Shang, Wenhao Gao (possible past Massachusetts Institute Of Technology affiliation), Connor W. Coley (possible past Massachusetts Institute Of Technology affiliation), Yizhou Sun, Wei Wang (possible past University Of Oxford affiliation)
Abstract

Multi-step retrosynthesis planning seeks to decompose a target molecule into commercially available building blocks through a sequence of feasible reactions. The vast combinatorial search space makes this task challenging even for expert chemists. Traditional methods combine tree search with offline-trained value networks that score candidates in isolation, without reasoning about complete multi-step routes. Recent work leverages Large Language Models (LLMs) for this task, but relies on simple i...

πŸ“„ Full-data accuracy with fewer labels for training and fine-tuning machine-learning force fields
πŸ—“οΈ Published: 7/16/2026
πŸ”— http://arxiv.org/abs/2607.14486v1
πŸ‘₯ Authors: Sheng Bi (possible past Google (United States) affiliation), Yi-Ze Wang, Jun Cheng (possible past Deepmind (United Kingdom) affiliation)
Abstract

Machine-learning force fields (MLFFs) are reliable only near their training distribution, making efficient construction of diverse training sets a major bottleneck for both train-from-scratch and foundation fine-tuning workflows. Active learning can reduce this cost, but standard model-committee uncertainty is impractical for foundation MLFFs because each committee member requires a separate fine-tuning run. We present an active-learning workflow based on last-layer-projection regression (LLPR),...

*Notable papers are those with at least two authors from a "big" AI/ML lab.