📄 Notable* Recent AI/ML arXiv Papers

Last updated just now...

📄 Local Sparsity Enables Unsupervised LLM Safety Detection
🗓️ Published: 9/17/2026
🔗 http://arxiv.org/abs/2609.20129v1
👥 Authors: Xin Chen (possible past Tencent (China) affiliation), Gil Kur, Alexander Shevchenko, Andreas Krause (possible past Eth Zurich affiliation)
Abstract

Deployment-time safety methods for large language models (LLMs) are predominantly supervised and assume access to unsafe training data. Nevertheless, new attacks and harm categories regularly arise, not captured by models trained in such a supervised fashion. An alternative approach is to view this problem through the lens of anomaly detection, namely, to rely solely on modeling safe data and flagging out-of-distribution inputs. However, LLM activations lie in a high-dimensional space, raising c...

📄 Tailored to you: longitudinal effects of personalising language models
🗓️ Published: 9/17/2026
🔗 http://arxiv.org/abs/2609.20077v1
👥 Authors: Canfer Akbulut, Justine Breuch, Arianna Manzini, Lujain Ibrahim, Matija Franklin, Roma Patel (possible past Google (United States) affiliation), Iason Gabriel (possible past Deepmind (United Kingdom) affiliation), Kristian Lum (possible past Google (United States) affiliation), Laura Weidinger (possible past Deepmind (United Kingdom) affiliation)
Abstract

Interest in developing personalised language models is rapidly growing. While personalisation is often viewed as a mechanism to better serve diverse user needs, the effects of sustained interactions with personalised models on people's perception of and behaviour toward AI remain poorly understood. Most critically, downstream consequences outside the immediate human--AI interaction loop, such as effects on users' self-perceptions and interpersonal relationships, remain largely unexamined. In thi...

📄 Past, Future, All at Once: Mitigating Stability-Plasticity Dilemma via Post-hoc JANUS Rectification
🗓️ Published: 9/17/2026
🔗 http://arxiv.org/abs/2609.19985v1
👥 Authors: Zhilong Zheng, Letian Tao, Yang Guan, Yujie Yang, Wei Xiong, Kehua Sheng, Bo Zhang (possible past Tencent (China) affiliation), Jingliang Duan (possible past Tsinghua University affiliation), Keqiang Li, Shengbo Eben Li (possible past Tsinghua University affiliation)
Abstract

Fine-tuning foundation models on new tasks inevitably suffer from catastrophic forgetting. While existing works attempt to mitigate this on the basis of parameter-efficient fine-tuning methods, they adopted an overly restrictive Subspace Orthogonality condition. In this paper, we introduce a purely post-hoc and tuning-agnostic weight rectification framework that achieves Parameter Space Orthogonality, which is the necessary and sufficient condition for preserving historical performance to the fi...

📄 ClashBench: Conflicts Leading Agents to Seize and Harm
🗓️ Published: 9/17/2026
🔗 http://arxiv.org/abs/2609.19892v1
👥 Authors: Yuejin Xie, Yu Li (possible past Tencent (China) affiliation), Dadi Guo, Qingyu Liu, Yuqian Fu, Yanwei Fu, Yujiu Yang (possible past Tsinghua University affiliation), Xia Hu, Dongrui Liu
Abstract

As agent systems become more widely used, multiple agent sessions increasingly run alongside pre-existing user tasks in the same environment, sharing resources with limited capacity or mutually exclusive states. This creates a safety risk: when granted sufficient privileges, an agent may resolve a resource conflict by terminating or otherwise disrupting an existing task rather than reporting it. In this work, we identify and formalize this failure mode, which we term destructive resource preempt...

📄 CoRELoop: Parameter-Efficient Controlled Recurrent Refinement for Audio Deepfake Detection
🗓️ Published: 9/17/2026
🔗 http://arxiv.org/abs/2609.19818v1
👥 Authors: Kunyu Feng, Yuxiang Wang, Li Wang (possible past Tesla (United States) affiliation), Wan Lin, Zhizheng Wu (possible past University Of Edinburgh affiliation)
Abstract

Generalizing to unseen attacks remains challenging for audio deepfake detectors, and collecting training data covering all potential attacks is impractical. We explore recurrent refinement in an already-trained SSL-based detector without additional data or changes to its original parameters. However, directly recycling encoder outputs as inputs degrades detection in our diagnostic. We propose CoReLoop, which makes this reuse effective by adapting recurrent inputs to the frozen encoder, controlli...

📄 ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI
🗓️ Published: 9/17/2026
🔗 http://arxiv.org/abs/2609.19644v1
👥 Authors: Jaehyun Nam, Jinsung Yoon (possible past Google (United States) affiliation), Yanzhou Pan, Yubo Wang, Rui Meng, Parthasarathy Ranganathan (possible past Google (United States) affiliation), Tomas Pfister (possible past University Of Oxford affiliation)
Abstract

Scientific discovery is defined by the ability to identify the boundaries of existing knowledge and venture into unexplored territory. The ultimate vision for AI in science is problem-driven autonomous research: given a fundamental challenge by a human expert, the AI independently navigates the scientific landscape, uncovers theoretical and empirical bottlenecks, and systematically expands the frontier of knowledge. In this paper, we introduce ScientistTwo, a fully autonomous multi-agent framewo...

📄 AURORA: A Natural Language-Driven Agentic Framework for Understanding, Reasoning, and Orchestrating Reliable Air-Ground Co-Simulation
🗓️ Published: 9/17/2026
🔗 http://arxiv.org/abs/2609.19527v1
👥 Authors: Keshu Wu, Hao Zhang (possible past Tencent (China) affiliation), Rui Gan, Xiangbo Gao, Xiaopeng Li, Zhengzhong Tu (possible past Google (United States) affiliation), Yang Zhou
Abstract

Air-ground transportation research increasingly relies on co-simulation, yet constructing scenarios remains labor-intensive and difficult to validate. More importantly, a generated scenario may execute successfully while failing to realize the spatial, temporal, communication, or behavioral relationships requested by the user. This paper presents AURORA, a natural-language-driven agentic framework that treats air-ground scenario generation as a process of compilation with verification. Central t...

📄 Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening
🗓️ Published: 9/16/2026
🔗 http://arxiv.org/abs/2609.18708v1
👥 Authors: Yizhuo Li, Jianhao Yan, Yun Luo, Zhi Wang, Futing Wang, Rong-Xi Tan, Kanghui Tian, Ganqu Cui (possible past Tsinghua University affiliation), Ning Ding (possible past Tsinghua University affiliation), Peilin Zhao (possible past Tencent (China) affiliation), Yafu Li, Yu Cheng (possible past National University Of Singapore affiliation)
Abstract

In reinforcement learning for large language models, Proximal Policy Optimization (PPO) commonly uses a critic to estimate state values and reduce the variance of policy updates. However, we uncover a systematic failure mode in PPO critics, which we call Value Flattening: state values, estimated from multiple Monte Carlo continuations, change sharply across intermediate states while critic predictions remain comparatively flat. We further observe this phenomenon in a controlled FrozenLake enviro...

📄 ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models
🗓️ Published: 9/16/2026
🔗 http://arxiv.org/abs/2609.18487v1
👥 Authors: Shijie Lian, Bin Yu, Zhaolong Shen, Xiaopeng Lin, Yichao Du, Zhirui Zhang (possible past Tencent (China) affiliation), Laurence T. Yang, Kai Chen (possible past Shanghai Jiao Tong University affiliation)
Abstract

Action tokenizers play a central role in autoregressive vision-language-action (VLA) models, determining both the targets for policy training and the executable commands recovered from predicted tokens. Their fidelity is commonly evaluated using pointwise reconstruction metrics such as mean squared error (MSE), yet small individual errors do not fully characterize how faithfully action adjustments across demonstrations are preserved. After compression, similar actions may still cluster around a ...

📄 Disentangling Long-Term Memory via Latent Neuro-Symbolic Reasoning
🗓️ Published: 9/16/2026
🔗 http://arxiv.org/abs/2609.18461v1
👥 Authors: Cai Ke, Xinghao Chen, Xiaoyu Shen, Keyu Chen, Siyu An, Junnan Dong, Ruifeng Xu, Ruizhi Qiao (possible past Tencent (China) affiliation), Xing Sun (possible past Tencent (China) affiliation)
Abstract

Personalized agents are required to reason over long-term history interactions to infer both explicit preferences and implicit behavioral evidence. While early flat retrieval methods score memory fragments independently and neglect the distributed information, current structured memory frameworks rely on query-agnostic static graphs that fail to capture the context-dependent relations. Crucially, raw textual memories are inherently entangled and noisy, making fine-grained personalization and cro...

📄 GeoAAC: Geometry-Based Adaptive Action Chunking from Denoising Trajectories in VLA Policies
🗓️ Published: 9/17/2026
🔗 http://arxiv.org/abs/2609.20776v1
👥 Authors: Xin Chen (possible past Tencent (China) affiliation), Sen Chen, Yujuan Ding, Jian Liu, Guoqing Wang, Wei Ye (possible past Meta (United States) affiliation), Heng Tao Shen, Yi Bin
Abstract

Action chunking is widely used for action generation and execution in Vision-Language-Action (VLA) policies, yet existing approaches commonly use a fixed action horizon. During a rollout, different task stages may require different levels of action continuity, control precision, and closed-loop feedback, making a fixed horizon unable to accommodate changing control requirements. We propose \textbf{GeoAAC}, a geometry-based adaptive action chunking method for flow-based VLA policies that adjusts ...

📄 OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher
🗓️ Published: 9/17/2026
🔗 http://arxiv.org/abs/2609.20756v1
👥 Authors: Damiano Da Col, Maximilian Igl (possible past University Of Oxford affiliation), Peter Karkus, Kashyap Chitta, Boris Ivanovic, Marco Pavone (possible past Stanford University affiliation), Konrad Schindler, Christos Sakaridis
Abstract

As scaling pre-training data alone yields diminishing returns, post-training is becoming increasingly important across physical AI domains such as autonomous driving. End-to-end driving policies are pre-trained in open loop with behavior cloning on human demonstrations. However, compounding errors during closed-loop deployment can take the vehicle outside the training data distribution, increasing the risk of safety-critical incidents. Closed-loop post-training can mitigate this risk but require...

📄 Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation
🗓️ Published: 9/17/2026
🔗 http://arxiv.org/abs/2609.20744v1
👥 Authors: Haocheng Xi, Yiming Xie, Hexu Zhao, Yiwen Zhang (possible past Shanghai Jiao Tong University affiliation), Michael Liu, Thomas Creavin, Kurt Keutzer (possible past University Of California, Berkeley affiliation), Xiuyu Li, Zhaoyang Lv, Chenfeng Xu (possible past University Of California, Berkeley affiliation), Haiwen Feng
Abstract

Video diffusion models repeatedly process long spatiotemporal token sequences during denoising, making attention a major computational bottleneck. Linear attention offers an appealing alternative and has been widely adopted in recent large language models, but directly applying it to video models often fails to preserve the fine-grained interactions required for high-quality generation. We present Video DeltaNet (VDN), which combines local Softmax attention with bidirectional linear memory for l...

📄 CellRFT: Reinforcement Fine-Tuning for Single-Cell Perturbation Modeling
🗓️ Published: 9/17/2026
🔗 http://arxiv.org/abs/2609.19970v1
👥 Authors: Jie Yan, Li Liu (possible past National University Of Defense Technology affiliation), Hanze Guo, Jiaxin Hu, Houxin He, Xiaoning Qi, Haoran Wang, Cong Li (possible past Google (United States) affiliation), Zhong-Yuan Zhang, Yong Wang (possible past Baidu (China) affiliation)
Abstract

Predicting cellular responses to perturbations supports the study of gene function, disease mechanisms, and therapeutic strategies. Despite advances in single-cell perturbation modeling, existing models typically optimize surrogate losses that do not directly reflect the biological criteria used for evaluation, so better data fitting need not yield better biological predictions. To address this mismatch, we introduce \textbf{CellRFT}, a reinforcement fine-tuning framework that uses biological ev...

📄 EmbodiedMind: Adaptive Data Curation and Prefix-Tree Reinforcement Learning for Efficient Embodied Intelligence
🗓️ Published: 9/17/2026
🔗 http://arxiv.org/abs/2609.19659v1
👥 Authors: Feifan Wang, Zongbing Zhang, Yu Zhang (possible past Google (United States) affiliation), Lingfeng Wang, Yurui Zhu, Jin Deng, Mingliang Zhang, Zhengguang Gao, Yongcheng Wang, Jin Xu (possible past Tencent (China) affiliation), Ri Yang
Abstract

Training embodied foundation models typically requires massive-scale datasets and extensive computational resources, yet often suffers from three critical limitations: (1) inefficient sample utilization due to low-informative samples; (2) imbalanced gradient contributions across heterogeneous tasks; and (3) severe credit assignment problem in long-horizon planning, where trajectory-level rewards indiscriminately penalize all tokens. To address these issues, we propose an efficient training parad...

📄 Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations
🗓️ Published: 9/16/2026
🔗 http://arxiv.org/abs/2609.19101v1
👥 Authors: Leon Bergen, Usha Bhalla, Andrew Lee, Barak Widawsky, Linas Nasvytis, Connor Watts, Siddharth Boppana, Sidharth Baskaran, Dron Hazra, Michael Byun, Atticus Geiger (possible past Stanford University affiliation), Owen Lewis, Matthew Kowal, Vasudev Shyam, Thomas Fel, Thomas Mcgrath (possible past Google (United States) affiliation), Ekdeep Singh Lubana, Jack Merullo
Abstract

As models scale, reward hacking becomes more frequent, more sophisticated, and more consequential. Does it leave a telltale signature in model representations? This work analyzes how reward hacking is represented internally in frontier open source LLMs, and how those representations can be used to understand and discover the range of hacking behaviors a model displays. In particular, we find that simple difference of means vectors coherently represent reward hacking in Kimi K3, GLM 5.2, and Qwen...

📄 Physics-based prediction, uncertainty quantification and decision-making for IN718 crystallographic texture intensity across LPBF defocus regimes
🗓️ Published: 9/16/2026
🔗 http://arxiv.org/abs/2609.18863v1
👥 Authors: Yisheng Lu, John Riris, Jie Song (possible past Eth Zurich affiliation), Yao Fu, Jie Chen (possible past Tencent (China) affiliation)
Abstract

Reliable prediction of crystallographic texture in laser powder bed fusion is critical for linking process conditions with anisotropic response and for qualification. However, black-box models may fail under shift and cannot distinguish weak data support from loss of physical validity. This study develops a two-stage physics-based model for <001> || BD (build direction) texture in Inconel 718. Stage 1 maps process variables to melting mode and melt pool geometry. Stage 2 predicts texture by comb...

*Notable papers are those with at least two authors from a "big" AI/ML lab.