πŸ“„ Notable* Recent AI/ML arXiv Papers

Last updated just now...

πŸ“„ ExecCritic: Learn to Test, Test to Improve for Coding Agents
πŸ—“οΈ Published: 9/8/2026
πŸ”— http://arxiv.org/abs/2609.09133v1
πŸ‘₯ Authors: Leitian Tao, Baolin Peng, Haorui Wang, Hang Wang, Hao Cheng (possible past Tencent (China) affiliation), Wenlin Yao, Qianhui Wu, Tao Ge, Sharon Li, Jianfeng Gao (possible past Microsoft (United States) affiliation)
Abstract

Execution feedback can guide coding agents toward correct repository repairs, but only when the tests capture the behavior requested by the issue. Agent-generated tests can encode incomplete or incorrect behavioral targets; when the same trajectory writes both the patch and the test, their errors can agree and create false confidence. We introduce ExecCritic, combining a test--verify--revise scaffold with a role-specific reinforcement learning recipe for training agents within it. The scaffold s...

πŸ“„ PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving
πŸ—“οΈ Published: 9/8/2026
πŸ”— http://arxiv.org/abs/2609.08965v1
πŸ‘₯ Authors: Yuan Gao (possible past Tencent (China) affiliation), Sebastian MΓΌller, Mattia Piccinini, Marc Kaufeld, Yuchen Zhang (possible past University Of California, Berkeley affiliation), Finn Rasmus SchΓ€fer, Qunying Song, Johannes Betz
Abstract

Ensuring the safety of autonomous driving is a critical challenge. Scenario-based testing is a systematic process used to validate Autonomous Driving Systems (ADSs), but it remains a fragmented modular pipeline in which scenario generation, retrieval, modification, ADS execution, and results analysis are performed by separate tools with little interaction. Large Language Model (LLM) agents have shown promise across ADS sub-systems such as perception, planning, and control. However, no prior work...

πŸ“„ Kairos: A Dataset for Fine-Grained Video-Language Modeling over Space, Time, and Dynamics
πŸ—“οΈ Published: 9/8/2026
πŸ”— http://arxiv.org/abs/2609.08755v1
πŸ‘₯ Authors: Ruibo Ming, Lei Sun, Deheng Zhang, He Zhang, Jialu Li, Jian Wang (possible past Baidu (China) affiliation), Zhendong Li, Mengshun Hu, Danda Pani Paudel, Luc Van Gool (possible past Google (United States) affiliation), Jinjin Gu
Abstract

Many emerging video language modeling tasks require systems to move beyond clip-level abstraction and model visual content as it unfolds over extended time horizons. However, most existing video datasets rely on coarse or sparsely aligned supervision, which compresses temporal variation and limits the ability of models to learn reusable representations of continuous visual dynamics. We introduce Kairos, a video dataset for video-language modeling with time-resolved annotations. Kairos consists o...

πŸ“„ Same Values, Different Languages? From Multilingual Probing to Steering LLMs Toward Chinese Social Values
πŸ—“οΈ Published: 9/8/2026
πŸ”— http://arxiv.org/abs/2609.08515v1
πŸ‘₯ Authors: Yuemei Xu, Kexin Xu, Jian Zhou (possible past Tencent (China) affiliation), Haoyu Lu, Yequan Wang (possible past Tsinghua University affiliation), Aishan Liu
Abstract

As Large Language Models (LLMs) are increasingly integrated into human society, aligning them with pluralistic social values has become a critical priority. However, whether LLMs exhibit consistent value preferences across languages remains underexplored, particularly for culturally grounded values, which are more abstract and difficult to evaluate and align than safety-centric principles. We investigate this issue through Chinese Social Values (CSV), a value system rooted in Chinese culture and...

πŸ“„ Segment Any Motion with Radar: Robust Multimodal Moving-Object Segmentation and Tracking
πŸ—“οΈ Published: 9/8/2026
πŸ”— http://arxiv.org/abs/2609.08346v1
πŸ‘₯ Authors: Jue Wang (possible past Tencent (China) affiliation), Xuan Wang (possible past Tencent (China) affiliation), Hao Zhou, Ruixiang Zhou, Yixuan Zhou, Tianshuo Yuan, Jieming Ma, Jie Zhang, Fei Luo
Abstract

Moving-object perception must decide which image regions correspond to real motion and keep every instance identified over time. Methods that read motion from appearance, optical flow, or estimated trajectories lose that evidence under poor illumination, adverse weather, reflections, and occlusion. Radar is a natural remedy because it measures radial velocity directly instead of inferring it from photometric correspondence. However, existing benchmarks do not jointly provide radar measurements, ...

πŸ“„ MemForest: Efficient Agent Memory Management via EventTree Partitioning and Progressive Merging
πŸ—“οΈ Published: 9/8/2026
πŸ”— http://arxiv.org/abs/2609.08273v1
πŸ‘₯ Authors: Junxi Wang, Te Sun, Jiayi Zhu, Chen Zhang (possible past Peking University affiliation), Siyuan Li (possible past Tencent (China) affiliation), Xuyang Liu, Zichen Wen, Xiaobing Tu, Jinkui Ren, Xiantao Zhang, Ziqi Yuan, Linfeng Zhang
Abstract

Agent memory systems have demonstrated significant potential in long-term dialogue, personalized assistants, and video understanding. However, continuously accumulated memory introduces substantial storage and retrieval costs during inference. To address this issue, we propose \textbf{MemForest}, a general memory compression framework adaptable to various agent memory systems. Specifically, MemForest partitions historical memory into event-centric units by leveraging global semantic similarity a...

πŸ“„ Agentic ML Exploration (A-MLE) for Ads Ranking
πŸ—“οΈ Published: 9/8/2026
πŸ”— http://arxiv.org/abs/2609.08248v1
πŸ‘₯ Authors: Erwin Gao, Vinodh Kumar Sunkara, Jingyi Guan, Qinjin Jia, Hangjun Xu, Xiang Ji, Sherman Wong, Surya Teja Chavali, Pratik Vaishnavi, Aryan Pandhi, Xiaoyu Deng, Zhaodong Wang, Samarth Inani, Fan Yang (possible past Tencent (China) affiliation), Jakob Moberg, Zoe Zu, Nicolas Bievre, Sami Khenissi, Amit Jaspal, Ehsan Fakharizadi, Srinidhi Viswanathan, Dorothy Sun, Abishek Vanam, Sneha Iyer, Sheela Yadawad, Wenjie Chen, Gaby Nahum, Junhua Gu, Peter Chu, Yucheng Liu, Xin Zhao, Vitor Cid, Chaorong Chen, Vijay Pappu, Ashwin Kumar, Wenlin Chen (possible past Meta (United States) affiliation), Ben Schulte, Deepak Chandra (possible past Google (United States) affiliation), Ritwik Tewari
Abstract

Modern industrial ads ranking stacks are increasingly bottlenecked not by model capacity or training compute, but by the throughput of human ML iteration - the cycles of research, implementation, training, debugging, evaluation, and launch required to surface a single statistically significant improvement. A typical ranking stack contains numerous differentiated models with heterogeneous data, architectures, and infrastructure constraints, and each cycle takes days to weeks of senior engineer at...

πŸ“„ 3DWay: Generalizing Robot Manipulation via 3D Consistent Waypoints
πŸ—“οΈ Published: 9/8/2026
πŸ”— http://arxiv.org/abs/2609.08224v1
πŸ‘₯ Authors: Ziqin Huang, Yingyue Li, Chenyangguang Zhang, Ruida Zhang (possible past Tsinghua University affiliation), Yuxin Chen, Gu Wang (possible past Tsinghua University affiliation), Xingyu Liu (possible past Stanford University affiliation), Masayoshi Tomizuka (possible past University Of California, Berkeley affiliation), Xiangyang Ji (possible past Tsinghua University affiliation)
Abstract

Intermediate representations are key to bridging the modality gap between generalizable manipulation policies and large-scale pretrained vision-language models (VLMs). Among these, trajectory-based representations compactly represent motion-relevant cues, yet most existing approaches predict trajectories in 2D image space, resulting in intrinsic 3D ambiguity. Moreover, using 2D trajectories with depth still leaves the free-space waypoints ambiguous, limiting reliable 3D reasoning. To address thi...

πŸ“„ A*-Thought-V2: Efficient Latent Reasoning via Geometric Dynamics of LLM
πŸ—“οΈ Published: 9/7/2026
πŸ”— http://arxiv.org/abs/2609.07821v1
πŸ‘₯ Authors: Xiaoang Xu, Siyuan Liu, Shuo Wang (possible past Nvidia (United States) affiliation), Junlan Feng, Fanyu Meng, Zhu Zhang, Jixun Wang, Xiaorong Wang, Zihan Zhou, Xin Li (possible past Google (United States) affiliation), Chaojun Xiao, Yiming Zhang, Huijia Wu, Liuyu Xiang, Peipei Li, Zhaofeng He
Abstract

Chain-of-Thought (CoT) improves the reasoning ability of Large Language Models (LLMs) but incurs substantial computation and context costs. Existing methods either lose intermediate information through hard pruning or lack a principled criterion for continuous compression. We present A*-Thought-V2, a geometric dynamics of LLM guided framework that models CoT as a hidden-state trajectory and replaces hard deletion with an explicit-implicit interleaved latent architecture. After projecting questio...

πŸ“„ The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing
πŸ—“οΈ Published: 9/7/2026
πŸ”— http://arxiv.org/abs/2609.07713v1
πŸ‘₯ Authors: Chenguang Wang (possible past Amazon (United States) affiliation), Ming Li, Adebayo Braimah, Chenrui Fan, Tuo Wang, Weijie Guan, Ruiyi Zhang, Tianyi Zhou (possible past University Of Washington affiliation), Dawei Zhou
Abstract

Generative and agentic AI are reshaping both the production and evaluation of scientific research. These developments are often studied separately, as questions of how AI can produce research and how AI can review it. We argue that this separation misses an increasingly important feature of scholarly publishing: changes on one side alter the incentives, constraints, and behavior of the other. We synthesize 230 scholarly publications and institutional records using a taxonomy of six connected dyn...

πŸ“„ AgentIdeaBench: Benchmarking Scientific Ideation in the Agent Era
πŸ—“οΈ Published: 9/7/2026
πŸ”— http://arxiv.org/abs/2609.07611v1
πŸ‘₯ Authors: Yunxiang Mo, Tianshi Zheng, Yisen Gao, Rui Wang (possible past Tencent (China) affiliation), Newt Nguyen Kim Hue Nam, Kelvin Kiu Wai Tam, Jiaxin Bai, Yangqiu Song (possible past Tsinghua University affiliation), Ginny Wong, Simon See (possible past Nvidia (United States) affiliation)
Abstract

Scientific ideation is the capacity to formulate novel and testable hypotheses from scientific evidence, and autonomous AI scientists depend on it. Existing evaluations largely assess it by asking models to generate ideas from a static, curated set of reference papers. That passive setup departs from the retrieval-and-reasoning workflow of modern AI scientists, and it becomes less discriminative as models improve. We introduce AgentIdeaBench, a multidisciplinary benchmark that evaluates scientif...

πŸ“„ CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs
πŸ—“οΈ Published: 9/8/2026
πŸ”— http://arxiv.org/abs/2609.08345v1
πŸ‘₯ Authors: Nhat-Tan Bui, Varshini Elangovan, Arun Reddy Anugu, Sreyas Mohan, Wei Ye (possible past Meta (United States) affiliation), Dilin Wang (possible past Meta (United States) affiliation), Jq Huang, Rakesh Ranjan, Aviral Chharia, Fernando De La Torre (possible past Carnegie Mellon University affiliation)
Abstract

Representing a 3D scene as multi-view images allows 2D VLMs to reason in 3D by reusing priors from pre-training, sidestepping the scarcity of annotated 3D data. However, it produces thousands of redundant visual tokens whose cost grows with every view. Existing visual token pruners fall into two families, each limited in the 3D multi-view setting. Learned importance methods rank tokens by attention or encoder features; because redundancy here is fundamentally spatial, they keep near-duplicate to...

πŸ“„ Distillation as Probability Transport: Routed On-Policy Distillation
πŸ—“οΈ Published: 9/8/2026
πŸ”— http://arxiv.org/abs/2609.08337v1
πŸ‘₯ Authors: Tianle Xia, Lingxiang Hu, Yiding Sun, Linfang Shang, Ming Xu, Lan Xu, Ning Zheng, Wei Xu (possible past Tencent (China) affiliation), Jie Jiang (possible past Tencent (China) affiliation)
Abstract

On-policy distillation (OPD) transfers teacher knowledge on student-generated trajectories, but efficient sampled objectives reduce the teacher distribution to scalar credit on individual tokens. Such credit indicates whether a token should gain or lose probability, yet leaves the corresponding redistribution unspecified. We recast OPD as teacher-guided probability transport and propose RouteOPD (Routed On-Policy Distillation), which decomposes local teacher--student disagreement into student-ex...

πŸ“„ DRIFT: Removing Diffusion Watermarks by Deflecting the Generative Trajectory
πŸ—“οΈ Published: 9/8/2026
πŸ”— http://arxiv.org/abs/2609.08213v1
πŸ‘₯ Authors: Rui Bao, Zheng Gao, Xiaoyu Li (possible past Tencent (China) affiliation), Xiaoyan Feng, Yang Song (possible past Stanford University affiliation), Jiaojiao Jiang
Abstract

Diffusion watermarking embeds verifiable signals into the generative process and commonly verifies them by recovering trajectory-dependent evidence, making the marks robust to conventional pixel-space distortions. Existing removal attacks either regenerate along deterministic trajectories, which often preserve the watermark-bearing latent structure, or optimize every image separately. We identify the reliance on a recoverable generative trajectory as a common attack surface among the schemes we ...

πŸ“„ Learning Metamaterial Eigenmodes with Wavelet-Encoded Fourier Neural Operators
πŸ—“οΈ Published: 9/8/2026
πŸ”— http://arxiv.org/abs/2609.08102v1
πŸ‘₯ Authors: Han Zhang (possible past Tsinghua University affiliation), Alexander Ogren, Cynthia Rudin (possible past Massachusetts Institute Of Technology affiliation), Johann Guilleminot, L. Catherine Brinson
Abstract

Machine learning surrogates based on neural operators have shown broad applicability in solving forward PDE problems. However, eigenvalue problems, in which an eigenparameter and one of several valid eigenmodes must be simultaneously solved, remain difficult because standard operator learning formulations assume a unique input-output map. This work demonstrates that Fourier Neural Operators (FNOs), combined with wavelet-based encodings of PDE inputs, can learn and predict multiple eigenmodes of ...

*Notable papers are those with at least two authors from a "big" AI/ML lab.