πŸ“„ Notable* Recent AI/ML arXiv Papers

Last updated just now...

πŸ“„ GeoAAC: Geometry-Based Adaptive Action Chunking from Denoising Trajectories in VLA Policies
πŸ—“οΈ Published: 9/17/2026
πŸ”— http://arxiv.org/abs/2609.20776v1
πŸ‘₯ Authors: Xin Chen (possible past Tencent (China) affiliation), Sen Chen, Yujuan Ding, Jian Liu, Guoqing Wang, Wei Ye (possible past Meta (United States) affiliation), Heng Tao Shen, Yi Bin
Abstract

Action chunking is widely used for action generation and execution in Vision-Language-Action (VLA) policies, yet existing approaches commonly use a fixed action horizon. During a rollout, different task stages may require different levels of action continuity, control precision, and closed-loop feedback, making a fixed horizon unable to accommodate changing control requirements. We propose \textbf{GeoAAC}, a geometry-based adaptive action chunking method for flow-based VLA policies that adjusts ...

πŸ“„ HIL-UMI: Bringing Human-in-the-Loop Post-Training of Vision-Language-Action Models to Universal Manipulation Interface
πŸ—“οΈ Published: 9/17/2026
πŸ”— http://arxiv.org/abs/2609.20659v1
πŸ‘₯ Authors: Zimu Han, Yiming Zeng, Jiyao Zhang, Zihao Zhao (possible past Tsinghua University affiliation), Yuanfei Wang, Yixiang Jin, Shiqi Li, Shuangben Chen, Wei Huang (possible past Google (United States) affiliation), Ruodai Li, Hui Shen, Hao Dong
Abstract

Large-scale vision-language-action (VLA) models provide powerful priors for robot manipulation, yet adapting them to a specific deployment remains challenging. Supervised fine-tuning (SFT) on task-specific demonstrations provides a step toward deployment, but faces two persistent limitations: static data provide limited coverage of out-of-distribution states, and standard imitation objectives do not distinguish progressing behavior from less useful data. Interactive post-training can address the...

πŸ“„ SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness
πŸ—“οΈ Published: 9/17/2026
πŸ”— http://arxiv.org/abs/2609.20519v1
πŸ‘₯ Authors: Haozhe Liu, Tian Ye, Sensen Gao, Qihang Cao, Yitong Li, Mingchen Zhuge, Duomin Wang, Ruihua Zhang, Ping Luo (possible past Shanghai Artificial Intelligence Laboratory affiliation), Jiawang Bian, Lei Zhu, Ligeng Zhu, Enze Xie, Song Han (possible past Stanford University affiliation)
Abstract

As coding agents move from supervised code completion to unattended, around-the-clock exploration, their work expands from isolated predictions into long trajectories of reasoning, tool use, and feedback. Token efficiency therefore becomes important for scaling recursive self-improvement. We take an RSI-inspired approach at the harness layer, scaling auto-research loops across increasingly numerous and diverse environments for harness rollouts. At this scale, the process yields reusable improvem...

πŸ“„ greCAPTCHA: Assessing Understanding as Evidence of Research Authorship Under Generative AI
πŸ—“οΈ Published: 9/17/2026
πŸ”— http://arxiv.org/abs/2609.20481v1
πŸ‘₯ Authors: Justin Payan, BΓ‘lint GyevnΓ‘r, Atoosa Kasirzadeh (possible past University Of Toronto affiliation), Nihar B. Shah (possible past University Of California, Berkeley affiliation)
Abstract

Conferences, journals, funders, schools, and universities are struggling with a surge of potentially AI-generated submissions from ostensibly human authors, who may not have exercised sufficient human oversight for their manuscripts. In turn, institutions evaluating submissions can no longer reliably credit expertise based solely on authors' names on submitted work. To address this problem, we propose greCAPTCHA, a proctored assessment approach that measures authors' understanding of research ma...

πŸ“„ Local Sparsity Enables Unsupervised LLM Safety Detection
πŸ—“οΈ Published: 9/17/2026
πŸ”— http://arxiv.org/abs/2609.20129v1
πŸ‘₯ Authors: Xin Chen (possible past Tencent (China) affiliation), Gil Kur, Alexander Shevchenko, Andreas Krause (possible past Eth Zurich affiliation)
Abstract

Deployment-time safety methods for large language models (LLMs) are predominantly supervised and assume access to unsafe training data. Nevertheless, new attacks and harm categories regularly arise, not captured by models trained in such a supervised fashion. An alternative approach is to view this problem through the lens of anomaly detection, namely, to rely solely on modeling safe data and flagging out-of-distribution inputs. However, LLM activations lie in a high-dimensional space, raising c...

πŸ“„ Tailored to you: longitudinal effects of personalising language models
πŸ—“οΈ Published: 9/17/2026
πŸ”— http://arxiv.org/abs/2609.20077v1
πŸ‘₯ Authors: Canfer Akbulut, Justine Breuch, Arianna Manzini, Lujain Ibrahim, Matija Franklin, Roma Patel (possible past Google (United States) affiliation), Iason Gabriel (possible past Deepmind (United Kingdom) affiliation), Kristian Lum (possible past Google (United States) affiliation), Laura Weidinger (possible past Deepmind (United Kingdom) affiliation)
Abstract

Interest in developing personalised language models is rapidly growing. While personalisation is often viewed as a mechanism to better serve diverse user needs, the effects of sustained interactions with personalised models on people's perception of and behaviour toward AI remain poorly understood. Most critically, downstream consequences outside the immediate human--AI interaction loop, such as effects on users' self-perceptions and interpersonal relationships, remain largely unexamined. In thi...

πŸ“„ Past, Future, All at Once: Mitigating Stability-Plasticity Dilemma via Post-hoc JANUS Rectification
πŸ—“οΈ Published: 9/17/2026
πŸ”— http://arxiv.org/abs/2609.19985v1
πŸ‘₯ Authors: Zhilong Zheng, Letian Tao, Yang Guan, Yujie Yang, Wei Xiong, Kehua Sheng, Bo Zhang (possible past Tencent (China) affiliation), Jingliang Duan (possible past Tsinghua University affiliation), Keqiang Li, Shengbo Eben Li (possible past Tsinghua University affiliation)
Abstract

Fine-tuning foundation models on new tasks inevitably suffer from catastrophic forgetting. While existing works attempt to mitigate this on the basis of parameter-efficient fine-tuning methods, they adopted an overly restrictive Subspace Orthogonality condition. In this paper, we introduce a purely post-hoc and tuning-agnostic weight rectification framework that achieves Parameter Space Orthogonality, which is the necessary and sufficient condition for preserving historical performance to the fi...

πŸ“„ ClashBench: Conflicts Leading Agents to Seize and Harm
πŸ—“οΈ Published: 9/17/2026
πŸ”— http://arxiv.org/abs/2609.19892v1
πŸ‘₯ Authors: Yuejin Xie, Yu Li (possible past Tencent (China) affiliation), Dadi Guo, Qingyu Liu, Yuqian Fu, Yanwei Fu, Yujiu Yang (possible past Tsinghua University affiliation), Xia Hu, Dongrui Liu
Abstract

As agent systems become more widely used, multiple agent sessions increasingly run alongside pre-existing user tasks in the same environment, sharing resources with limited capacity or mutually exclusive states. This creates a safety risk: when granted sufficient privileges, an agent may resolve a resource conflict by terminating or otherwise disrupting an existing task rather than reporting it. In this work, we identify and formalize this failure mode, which we term destructive resource preempt...

πŸ“„ CoRELoop: Parameter-Efficient Controlled Recurrent Refinement for Audio Deepfake Detection
πŸ—“οΈ Published: 9/17/2026
πŸ”— http://arxiv.org/abs/2609.19818v1
πŸ‘₯ Authors: Kunyu Feng, Yuxiang Wang, Li Wang (possible past Tesla (United States) affiliation), Wan Lin, Zhizheng Wu (possible past University Of Edinburgh affiliation)
Abstract

Generalizing to unseen attacks remains challenging for audio deepfake detectors, and collecting training data covering all potential attacks is impractical. We explore recurrent refinement in an already-trained SSL-based detector without additional data or changes to its original parameters. However, directly recycling encoder outputs as inputs degrades detection in our diagnostic. We propose CoReLoop, which makes this reuse effective by adapting recurrent inputs to the frozen encoder, controlli...

πŸ“„ ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI
πŸ—“οΈ Published: 9/17/2026
πŸ”— http://arxiv.org/abs/2609.19644v1
πŸ‘₯ Authors: Jaehyun Nam, Jinsung Yoon (possible past Google (United States) affiliation), Yanzhou Pan, Yubo Wang, Rui Meng, Parthasarathy Ranganathan (possible past Google (United States) affiliation), Tomas Pfister (possible past University Of Oxford affiliation)
Abstract

Scientific discovery is defined by the ability to identify the boundaries of existing knowledge and venture into unexplored territory. The ultimate vision for AI in science is problem-driven autonomous research: given a fundamental challenge by a human expert, the AI independently navigates the scientific landscape, uncovers theoretical and empirical bottlenecks, and systematically expands the frontier of knowledge. In this paper, we introduce ScientistTwo, a fully autonomous multi-agent framewo...

πŸ“„ AURORA: A Natural Language-Driven Agentic Framework for Understanding, Reasoning, and Orchestrating Reliable Air-Ground Co-Simulation
πŸ—“οΈ Published: 9/17/2026
πŸ”— http://arxiv.org/abs/2609.19527v1
πŸ‘₯ Authors: Keshu Wu, Hao Zhang (possible past Tencent (China) affiliation), Rui Gan, Xiangbo Gao, Xiaopeng Li, Zhengzhong Tu (possible past Google (United States) affiliation), Yang Zhou
Abstract

Air-ground transportation research increasingly relies on co-simulation, yet constructing scenarios remains labor-intensive and difficult to validate. More importantly, a generated scenario may execute successfully while failing to realize the spatial, temporal, communication, or behavioral relationships requested by the user. This paper presents AURORA, a natural-language-driven agentic framework that treats air-ground scenario generation as a process of compilation with verification. Central t...

πŸ“„ OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher
πŸ—“οΈ Published: 9/17/2026
πŸ”— http://arxiv.org/abs/2609.20756v1
πŸ‘₯ Authors: Damiano Da Col, Maximilian Igl (possible past University Of Oxford affiliation), Peter Karkus, Kashyap Chitta, Boris Ivanovic, Marco Pavone (possible past Stanford University affiliation), Konrad Schindler, Christos Sakaridis
Abstract

As scaling pre-training data alone yields diminishing returns, post-training is becoming increasingly important across physical AI domains such as autonomous driving. End-to-end driving policies are pre-trained in open loop with behavior cloning on human demonstrations. However, compounding errors during closed-loop deployment can take the vehicle outside the training data distribution, increasing the risk of safety-critical incidents. Closed-loop post-training can mitigate this risk but require...

πŸ“„ Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation
πŸ—“οΈ Published: 9/17/2026
πŸ”— http://arxiv.org/abs/2609.20744v1
πŸ‘₯ Authors: Haocheng Xi, Yiming Xie, Hexu Zhao, Yiwen Zhang (possible past Shanghai Jiao Tong University affiliation), Michael Liu, Thomas Creavin, Kurt Keutzer (possible past University Of California, Berkeley affiliation), Xiuyu Li, Zhaoyang Lv, Chenfeng Xu (possible past University Of California, Berkeley affiliation), Haiwen Feng
Abstract

Video diffusion models repeatedly process long spatiotemporal token sequences during denoising, making attention a major computational bottleneck. Linear attention offers an appealing alternative and has been widely adopted in recent large language models, but directly applying it to video models often fails to preserve the fine-grained interactions required for high-quality generation. We present Video DeltaNet (VDN), which combines local Softmax attention with bidirectional linear memory for l...

πŸ“„ CellRFT: Reinforcement Fine-Tuning for Single-Cell Perturbation Modeling
πŸ—“οΈ Published: 9/17/2026
πŸ”— http://arxiv.org/abs/2609.19970v1
πŸ‘₯ Authors: Jie Yan, Li Liu (possible past National University Of Defense Technology affiliation), Hanze Guo, Jiaxin Hu, Houxin He, Xiaoning Qi, Haoran Wang, Cong Li (possible past Google (United States) affiliation), Zhong-Yuan Zhang, Yong Wang (possible past Baidu (China) affiliation)
Abstract

Predicting cellular responses to perturbations supports the study of gene function, disease mechanisms, and therapeutic strategies. Despite advances in single-cell perturbation modeling, existing models typically optimize surrogate losses that do not directly reflect the biological criteria used for evaluation, so better data fitting need not yield better biological predictions. To address this mismatch, we introduce \textbf{CellRFT}, a reinforcement fine-tuning framework that uses biological ev...

πŸ“„ EmbodiedMind: Adaptive Data Curation and Prefix-Tree Reinforcement Learning for Efficient Embodied Intelligence
πŸ—“οΈ Published: 9/17/2026
πŸ”— http://arxiv.org/abs/2609.19659v1
πŸ‘₯ Authors: Feifan Wang, Zongbing Zhang, Yu Zhang (possible past Google (United States) affiliation), Lingfeng Wang, Yurui Zhu, Jin Deng, Mingliang Zhang, Zhengguang Gao, Yongcheng Wang, Jin Xu (possible past Tencent (China) affiliation), Ri Yang
Abstract

Training embodied foundation models typically requires massive-scale datasets and extensive computational resources, yet often suffers from three critical limitations: (1) inefficient sample utilization due to low-informative samples; (2) imbalanced gradient contributions across heterogeneous tasks; and (3) severe credit assignment problem in long-horizon planning, where trajectory-level rewards indiscriminately penalize all tokens. To address these issues, we propose an efficient training parad...

πŸ“„ Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations
πŸ—“οΈ Published: 9/16/2026
πŸ”— http://arxiv.org/abs/2609.19101v1
πŸ‘₯ Authors: Leon Bergen, Usha Bhalla, Andrew Lee, Barak Widawsky, Linas Nasvytis, Connor Watts, Siddharth Boppana, Sidharth Baskaran, Dron Hazra, Michael Byun, Atticus Geiger (possible past Stanford University affiliation), Owen Lewis, Matthew Kowal, Vasudev Shyam, Thomas Fel, Thomas Mcgrath (possible past Google (United States) affiliation), Ekdeep Singh Lubana, Jack Merullo
Abstract

As models scale, reward hacking becomes more frequent, more sophisticated, and more consequential. Does it leave a telltale signature in model representations? This work analyzes how reward hacking is represented internally in frontier open source LLMs, and how those representations can be used to understand and discover the range of hacking behaviors a model displays. In particular, we find that simple difference of means vectors coherently represent reward hacking in Kimi K3, GLM 5.2, and Qwen...

πŸ“„ Physics-based prediction, uncertainty quantification and decision-making for IN718 crystallographic texture intensity across LPBF defocus regimes
πŸ—“οΈ Published: 9/16/2026
πŸ”— http://arxiv.org/abs/2609.18863v1
πŸ‘₯ Authors: Yisheng Lu, John Riris, Jie Song (possible past Eth Zurich affiliation), Yao Fu, Jie Chen (possible past Tencent (China) affiliation)
Abstract

Reliable prediction of crystallographic texture in laser powder bed fusion is critical for linking process conditions with anisotropic response and for qualification. However, black-box models may fail under shift and cannot distinguish weak data support from loss of physical validity. This study develops a two-stage physics-based model for <001> || BD (build direction) texture in Inconel 718. Stage 1 maps process variables to melting mode and melt pool geometry. Stage 2 predicts texture by comb...

πŸ“„ Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening
πŸ—“οΈ Published: 9/16/2026
πŸ”— http://arxiv.org/abs/2609.18708v1
πŸ‘₯ Authors: Yizhuo Li, Jianhao Yan, Yun Luo, Zhi Wang, Futing Wang, Rong-Xi Tan, Kanghui Tian, Ganqu Cui (possible past Tsinghua University affiliation), Ning Ding (possible past Tsinghua University affiliation), Peilin Zhao (possible past Tencent (China) affiliation), Yafu Li, Yu Cheng (possible past National University Of Singapore affiliation)
Abstract

In reinforcement learning for large language models, Proximal Policy Optimization (PPO) commonly uses a critic to estimate state values and reduce the variance of policy updates. However, we uncover a systematic failure mode in PPO critics, which we call Value Flattening: state values, estimated from multiple Monte Carlo continuations, change sharply across intermediate states while critic predictions remain comparatively flat. We further observe this phenomenon in a controlled FrozenLake enviro...

*Notable papers are those with at least two authors from a "big" AI/ML lab.