πŸ“„ Notable* Recent AI/ML arXiv Papers

Last updated just now...

πŸ“„ CodeMidas: Scaling Agentic Coding RL Environments from Code Itself
πŸ—“οΈ Published: 9/18/2026
πŸ”— http://arxiv.org/abs/2609.22068v1
πŸ‘₯ Authors: Bowen Ye, Lei Li (possible past Carnegie Mellon University affiliation), Shicheng Li, Zihao Yue, Linghao Zhang, Hanglong Lv, Yuanxin Liu, Wenhan Ma, Hao Tian (possible past Baidu (China) affiliation), Rang Li, Jinhao Dong, Yikai Zhao, Xiangwei Deng, Hailin Zhang, Liang Zhao (possible past Baidu (China) affiliation), Qi Liu (possible past Tencent (China) affiliation), Lingpeng Kong (possible past Google (United States) affiliation), Tong Yang (possible past Peking University affiliation), Fuli Luo (possible past Peking University affiliation)
Abstract

Training capable coding agents via reinforcement learning (RL) requires diverse tasks with reliable verifiers. Open-source codebases offer a rich source of such tasks, while existing methods typically rely on development artifacts such as issues and commits, limiting the range of tasks that can be extracted. To better scale RL environments, we present CodeMidas, an agentic pipeline that turns implemented functionality in existing codebases into executable RL environments using source code as its...

πŸ“„ Value-Sensitive Delegation in Everyday AI Agent Use: Evidence from OpenClaw
πŸ—“οΈ Published: 9/18/2026
πŸ”— http://arxiv.org/abs/2609.22067v1
πŸ‘₯ Authors: Renkai Ma, Ruyuan Wan, Xuan Lu, Fan Yang (possible past Tencent (China) affiliation), Chen Chen (possible past Tencent (China) affiliation), Lingyao Li
Abstract

Users increasingly delegate work to autonomous AI agents, yet evaluations typically measure task completion rather than the values users prioritize. Using Value Sensitive Design, we analyzed, with LLM assistance, 73,093 first-person Reddit posts about using OpenClaw, each for its human value, agent aspect, value fulfillment, and user outcome. The 21 values form six value groups, including Autonomous, Dependable, and Affordable Operation, Bounded Reach, Reviewability, and Equitable Access. Relati...

πŸ“„ NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities
πŸ—“οΈ Published: 9/18/2026
πŸ”— http://arxiv.org/abs/2609.21967v1
πŸ‘₯ Authors: Jagadeesh Balam, Travis Bartley, Edresson Casanova, Sanjay Chauhan, Chen Chen (possible past Tencent (China) affiliation), Zhehuai Chen, Zijia Chen, Francesco Ciannella, Slyne Deng, Mikyas Desta, Harishchandra Dubey, Slim Essid, Nourchene Ferchichi, Boris Ginsburg (possible past Nvidia (United States) affiliation), Mariana Graterol Fuenmayor, Negar Habibi, Kevin Hu, Anand Joseph, Viraj Karandikar, Myungjong Kim, Viacheslav Klimkov, Seelan Lakshmi Narasimhan, Lily Lee, Jason Li (possible past Nvidia (United States) affiliation), Eileen Long, Ameya Mahabaleshwarkar, Aditya Malte, Adi Margolin, Sasha Meister, Valentin Mendelev, Oluwatobi Olabiyi, Ankita Pasad, Yifan Peng (possible past Stanford University affiliation), Elena Rastorgueva, Jayda Ritchie, Jason Roche, Nikhil Srihari, Yuanhang Su, Yoshi Suhara, Viet Anh Trinh, Jinhan Wang, Piotr Zelasko, Hui Wang, Puhui Meng, Chaosen Zhang, Yunsheng Liu, Shawn Wang, Wenjing Li, Zhonglei He
Abstract

We introduce NemotronLabs VoiceChat, an open full-duplex speech-to-speech model with native tool-calling capabilities. NemotronLabs VoiceChat combines a streaming speech encoder and decoder-only language model with parallel specialized output streams for agent text and structured function calls, an auxiliary RNN-T branch for incremental user transcription, and a streaming TTS decoder. This design enables the model to listen, transcribe, reason, invoke tools, and speak within a unified streaming ...

πŸ“„ CIBuzzBench: A Benchmark for Cross-Lingual Understanding of Chinese Internet Buzzwords
πŸ—“οΈ Published: 9/18/2026
πŸ”— http://arxiv.org/abs/2609.21722v1
πŸ‘₯ Authors: Yifan Wang (possible past Stanford University affiliation), Junyu Lu, Qifan Wang (possible past Google (United States) affiliation), Shun Zhang, Chaozhuo Li, Jiahao Liu, Zhijun Cao, Lingbin Bu, Fanliang Bu
Abstract

Chinese social media has generated a vast and continually evolving lexicon of internet buzzwords whose meanings are often non-literal and deeply rooted in local cultural and pragmatic contexts. Existing research has primarily focused on interpreting these buzzwords within Chinese, leaving largely unexplored whether LLMs can transfer such culturally grounded knowledge across languages and accurately convey the intended meanings in English. This cross-lingual capability is also critical for safety...

πŸ“„ One Prompt Does Not Fit All: Self-Meta-Evolve for Personalized Information Extraction
πŸ—“οΈ Published: 9/18/2026
πŸ”— http://arxiv.org/abs/2609.21626v1
πŸ‘₯ Authors: Hongliang Li, Lu Wang (possible past University Of Washington affiliation), Yong Xu (possible past Tencent (China) affiliation), Hanyang Chen, Zhitao Hou, Xiaoting Qin, Song Ge, Qingwei Lin, Dongmei Zhang
Abstract

Large language models (LLMs) are increasingly deployed for enterprise information extraction (IE), where the same document must be reorganized differently for each user. Existing prompt optimization methods, however, rely on a single prompt optimized against a global objective, which is misaligned with the inherent user heterogeneity of real workplaces. We formulate enterprise IE as per-user prompt adaptation under interaction feedback and propose Self-Meta-Evolve, a hierarchical framework that ...

πŸ“„ OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue
πŸ—“οΈ Published: 9/18/2026
πŸ”— http://arxiv.org/abs/2609.21465v1
πŸ‘₯ Authors: Haolin He, Yunfei Chu, Qi Chen (possible past Baidu (China) affiliation), Wen Huang, Yuan Feng, Muzhi Zhu, Zheqi Dai, Haoning Xu, Dongchao Yang, Chunyat Wu, Zining Liang, Zhengxi Liu, Xiquan Li, Xie Chen, Xize Cheng, Qize Yang, Jin Xu (possible past Tencent (China) affiliation), Qiuqiang Kong
Abstract

We define OmniVChat (Omni Video Chat) as the task of native audio-visual dialogue between a user and an omni model. In OmniVChat, omni models directly and simultaneously receive audio and video from a user and return text. The user's query is embedded in the audio and video, without a separate text question, external captioning, or speech recognition. Direct audio-visual input reduces external latency and computation while preserving perceptual cues. However, research on OmniVChat faces two cons...

πŸ“„ GVPO++: Group Variance Policy Optimization for LLM Post-Training and On-Policy Distillation
πŸ—“οΈ Published: 9/18/2026
πŸ”— http://arxiv.org/abs/2609.21432v1
πŸ‘₯ Authors: Kaichen Zhang, Yuzhong Hong, Junwei Bao, Hongfei Jiang, Yang Song (possible past Stanford University affiliation), Dingqian Hong, Hui Xiong (possible past Baidu (China) affiliation)
Abstract

Post-training plays a pivotal role in enhancing the reasoning capabilities and task-specific expertise of large language models (LLMs). Despite recent advances in post-training methods, such as Group Relative Policy Optimization (GRPO), their practical deployment remains impeded by training instability arising from the reliance on importance sampling. We introduce Group Variance Policy Optimization (GVPO), a novel post-training method that integrates the analytical solution of KL-constrained r...

πŸ“„ CogGym: Towards Large-Scale Comparative Evaluation of Human and Machine Cognition
πŸ—“οΈ Published: 9/18/2026
πŸ”— http://arxiv.org/abs/2609.21259v1
πŸ‘₯ Authors: Lance Ying, Jinzhou Wu, Yingshan Susan Wang, Shivam Aarya, Luca M. Schulze Buschoff, Harry Chen, Katherine M. Collins, Andrea De Varda, Shuhao Fu, Sean Dae Houlihan, Akshay K. Jagadish, Guangyuan Jiang, Samuel Kiegeland, Tetsu Kurumisawa, Rongzhi Liu, Ryan Liu, Ningshan Ma, Kathryn Mcgregor, Younes Strittmatter, Polina Tsvilodub, Jacob Hoover Vigly, Sarah Wu, Enjie Xu, Yiling Yun, Kelsey Allen, Tyler Brooke-Wilson, Brian Christian, Evelina Fedorenko, Michael C. Frank, Michael Franke, Tao Gao, Samuel J. Gershman, Robert D. Hawkins, Jennifer Hu, Julian Jara-Ettinger, Max Kleiman-Weiner, Sydney Levine, Tal Linzen (possible past Meta (United States) affiliation), Hongjing Lu, Timothy O'donnell, Desmond C. Ong, Steven T. Piantadosi, Rebecca Saxe (possible past Massachusetts Institute Of Technology affiliation), Eric Schulz, Tianmin Shu, Felix A. Sosa, Ilia Sucholutsky, Tan Zhi-Xuan, Tomer Ullman (possible past Massachusetts Institute Of Technology affiliation), Fei Xu, Ilker Yildirim, Jian-Qiao Zhu, Thomas L. Griffiths (possible past University Of California, Berkeley affiliation), Tobias Gerstenberg, Kevin Smith, Joshua B. Tenenbaum (possible past Massachusetts Institute Of Technology affiliation)
Abstract

Understanding and modeling human intelligence are parallel goals shared by artificial intelligence (AI) and cognitive science. As AI systems grow increasingly capable, in what ways do model responses resemble human responses, and where do they systematically diverge? The sheer breadth and diversity of the tasks humans can perform and think about pose a challenge for scalable and rigorous comparison between humans and models. We introduce CogGym, a scalable, unified framework grounded in cognitiv...

πŸ“„ PlantShade: Predicting Plant Shadows for Lighting-Aware Robotic Agricultural Operation
πŸ—“οΈ Published: 9/17/2026
πŸ”— http://arxiv.org/abs/2609.21059v1
πŸ‘₯ Authors: Longchao Da, Xiaoou Liu, Xingjian Li (possible past Baidu (China) affiliation), Lirong Xiang, Hua Wei (possible past Google (United States) affiliation)
Abstract

Plant growth and agricultural production form the foundation of a country's sustainable development and directly impact human livelihoods. Recent advances in frontier artificial intelligence have enabled scientific agriculture with strong potential to improve crop productivity. In this paper, we identify the importance and inherent complexity of plant shade simulation, as shading is a critical factor influencing plant growth. To advance this field and promote broader societal benefits, we focus ...

πŸ“„ CaLR: Causal Latent Revision for Robust Diffusion Reasoning
πŸ—“οΈ Published: 9/17/2026
πŸ”— http://arxiv.org/abs/2609.20981v1
πŸ‘₯ Authors: Wei Cai, Jian Zhao, Yuchen Yuan (possible past Baidu (China) affiliation), Xuelong Li (possible past Tencent (China) affiliation)
Abstract

Autoregressive (AR) models suffer from local greediness, while diffusion language models (DLMs) often lack the strict causal structure required for reasoning. To combine the advantages and overcome the drawbacks of the dual, we propose Causal Latent Revision (CaLR), a framework that reformulates reasoning as constrained latent optimization. By adopting a causal topology matrix (CTM) from an expert model and implicit differentiation, CaLR performs gradient-guided ``thought revision" to enforce lo...

πŸ“„ GeoAAC: Geometry-Based Adaptive Action Chunking from Denoising Trajectories in VLA Policies
πŸ—“οΈ Published: 9/17/2026
πŸ”— http://arxiv.org/abs/2609.20776v1
πŸ‘₯ Authors: Xin Chen (possible past Tencent (China) affiliation), Sen Chen, Yujuan Ding, Jian Liu, Guoqing Wang, Wei Ye (possible past Meta (United States) affiliation), Heng Tao Shen, Yi Bin
Abstract

Action chunking is widely used for action generation and execution in Vision-Language-Action (VLA) policies, yet existing approaches commonly use a fixed action horizon. During a rollout, different task stages may require different levels of action continuity, control precision, and closed-loop feedback, making a fixed horizon unable to accommodate changing control requirements. We propose \textbf{GeoAAC}, a geometry-based adaptive action chunking method for flow-based VLA policies that adjusts ...

πŸ“„ HIL-UMI: Bringing Human-in-the-Loop Post-Training of Vision-Language-Action Models to Universal Manipulation Interface
πŸ—“οΈ Published: 9/17/2026
πŸ”— http://arxiv.org/abs/2609.20659v1
πŸ‘₯ Authors: Zimu Han, Yiming Zeng, Jiyao Zhang, Zihao Zhao (possible past Tsinghua University affiliation), Yuanfei Wang, Yixiang Jin, Shiqi Li, Shuangben Chen, Wei Huang (possible past Google (United States) affiliation), Ruodai Li, Hui Shen, Hao Dong
Abstract

Large-scale vision-language-action (VLA) models provide powerful priors for robot manipulation, yet adapting them to a specific deployment remains challenging. Supervised fine-tuning (SFT) on task-specific demonstrations provides a step toward deployment, but faces two persistent limitations: static data provide limited coverage of out-of-distribution states, and standard imitation objectives do not distinguish progressing behavior from less useful data. Interactive post-training can address the...

πŸ“„ SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness
πŸ—“οΈ Published: 9/17/2026
πŸ”— http://arxiv.org/abs/2609.20519v1
πŸ‘₯ Authors: Haozhe Liu, Tian Ye, Sensen Gao, Qihang Cao, Yitong Li, Mingchen Zhuge, Duomin Wang, Ruihua Zhang, Ping Luo (possible past Shanghai Artificial Intelligence Laboratory affiliation), Jiawang Bian, Lei Zhu, Ligeng Zhu, Enze Xie, Song Han (possible past Stanford University affiliation)
Abstract

As coding agents move from supervised code completion to unattended, around-the-clock exploration, their work expands from isolated predictions into long trajectories of reasoning, tool use, and feedback. Token efficiency therefore becomes important for scaling recursive self-improvement. We take an RSI-inspired approach at the harness layer, scaling auto-research loops across increasingly numerous and diverse environments for harness rollouts. At this scale, the process yields reusable improvem...

πŸ“„ greCAPTCHA: Assessing Understanding as Evidence of Research Authorship Under Generative AI
πŸ—“οΈ Published: 9/17/2026
πŸ”— http://arxiv.org/abs/2609.20481v1
πŸ‘₯ Authors: Justin Payan, BΓ‘lint GyevnΓ‘r, Atoosa Kasirzadeh (possible past University Of Toronto affiliation), Nihar B. Shah (possible past University Of California, Berkeley affiliation)
Abstract

Conferences, journals, funders, schools, and universities are struggling with a surge of potentially AI-generated submissions from ostensibly human authors, who may not have exercised sufficient human oversight for their manuscripts. In turn, institutions evaluating submissions can no longer reliably credit expertise based solely on authors' names on submitted work. To address this problem, we propose greCAPTCHA, a proctored assessment approach that measures authors' understanding of research ma...

πŸ“„ Local Sparsity Enables Unsupervised LLM Safety Detection
πŸ—“οΈ Published: 9/17/2026
πŸ”— http://arxiv.org/abs/2609.20129v1
πŸ‘₯ Authors: Xin Chen (possible past Tencent (China) affiliation), Gil Kur, Alexander Shevchenko, Andreas Krause (possible past Eth Zurich affiliation)
Abstract

Deployment-time safety methods for large language models (LLMs) are predominantly supervised and assume access to unsafe training data. Nevertheless, new attacks and harm categories regularly arise, not captured by models trained in such a supervised fashion. An alternative approach is to view this problem through the lens of anomaly detection, namely, to rely solely on modeling safe data and flagging out-of-distribution inputs. However, LLM activations lie in a high-dimensional space, raising c...

πŸ“„ Tailored to you: longitudinal effects of personalising language models
πŸ—“οΈ Published: 9/17/2026
πŸ”— http://arxiv.org/abs/2609.20077v1
πŸ‘₯ Authors: Canfer Akbulut, Justine Breuch, Arianna Manzini, Lujain Ibrahim, Matija Franklin, Roma Patel (possible past Google (United States) affiliation), Iason Gabriel (possible past Deepmind (United Kingdom) affiliation), Kristian Lum (possible past Google (United States) affiliation), Laura Weidinger (possible past Deepmind (United Kingdom) affiliation)
Abstract

Interest in developing personalised language models is rapidly growing. While personalisation is often viewed as a mechanism to better serve diverse user needs, the effects of sustained interactions with personalised models on people's perception of and behaviour toward AI remain poorly understood. Most critically, downstream consequences outside the immediate human--AI interaction loop, such as effects on users' self-perceptions and interpersonal relationships, remain largely unexamined. In thi...

πŸ“„ Complex Problem Solving in Large Language Models: A Statistical Control Survey and Diagnostic Framework
πŸ—“οΈ Published: 9/17/2026
πŸ”— http://arxiv.org/abs/2609.20973v1
πŸ‘₯ Authors: Jiazhang Cai, Tao Wang (possible past Stanford University affiliation), Ruidong Zhang, Siyuan Li (possible past Tencent (China) affiliation), Terry Ma, Luyang Fang, Haoran Lu, Huimin Cheng, Yingchuan Zhang, Shushan Wu, Rui Xie, Lin Tang, Chao Huang (possible past Tencent (China) affiliation), Rongjie Liu, Ziyu Liu, Meizhi Yu, Yongkai Chen, Yifan Zhou, Zeliang Sun, Chang Liu, Zhen Xiang, Wei Xiao, Zixin Rao, Xinyi Liu, Yutong Hu, Mengrui Zhang, Jing Zhang (possible past University Of Washington affiliation), Weidi Luo, Jincheng Yu, Zhengliang Liu, Weihang You, Hanqi Jiang, Yi Pan, Junhao Chen, Xinliang Li, Tianming Liu, Wenxuan Zhong, Ping Ma
Abstract

Complex problem solving (CPS) with large language models (LLMs) is often framed as a matter of stronger reasoning or longer generation. Yet early-step error amplification, prompt brittleness, and failures to revise incorrect commitments are difficult to explain by missing knowledge or expressive capacity alone. This survey interprets CPS as a sequential estimation-and-decision problem over a latent solution state. A controller maintains a belief about an unobserved solution trajectory, updates i...

πŸ“„ OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher
πŸ—“οΈ Published: 9/17/2026
πŸ”— http://arxiv.org/abs/2609.20756v1
πŸ‘₯ Authors: Damiano Da Col, Maximilian Igl (possible past University Of Oxford affiliation), Peter Karkus, Kashyap Chitta, Boris Ivanovic, Marco Pavone (possible past Stanford University affiliation), Konrad Schindler, Christos Sakaridis
Abstract

As scaling pre-training data alone yields diminishing returns, post-training is becoming increasingly important across physical AI domains such as autonomous driving. End-to-end driving policies are pre-trained in open loop with behavior cloning on human demonstrations. However, compounding errors during closed-loop deployment can take the vehicle outside the training data distribution, increasing the risk of safety-critical incidents. Closed-loop post-training can mitigate this risk but require...

πŸ“„ Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation
πŸ—“οΈ Published: 9/17/2026
πŸ”— http://arxiv.org/abs/2609.20744v1
πŸ‘₯ Authors: Haocheng Xi, Yiming Xie, Hexu Zhao, Yiwen Zhang (possible past Shanghai Jiao Tong University affiliation), Michael Liu, Thomas Creavin, Kurt Keutzer (possible past University Of California, Berkeley affiliation), Xiuyu Li, Zhaoyang Lv, Chenfeng Xu (possible past University Of California, Berkeley affiliation), Haiwen Feng
Abstract

Video diffusion models repeatedly process long spatiotemporal token sequences during denoising, making attention a major computational bottleneck. Linear attention offers an appealing alternative and has been widely adopted in recent large language models, but directly applying it to video models often fails to preserve the fine-grained interactions required for high-quality generation. We present Video DeltaNet (VDN), which combines local Softmax attention with bidirectional linear memory for l...

πŸ“„ Past, Future, All at Once: Mitigating Stability-Plasticity Dilemma via Post-hoc JANUS Rectification
πŸ—“οΈ Published: 9/17/2026
πŸ”— http://arxiv.org/abs/2609.19985v1
πŸ‘₯ Authors: Zhilong Zheng, Letian Tao, Yang Guan, Yujie Yang, Wei Xiong, Kehua Sheng, Bo Zhang (possible past Tencent (China) affiliation), Jingliang Duan (possible past Tsinghua University affiliation), Keqiang Li, Shengbo Eben Li (possible past Tsinghua University affiliation)
Abstract

Fine-tuning foundation models on new tasks inevitably suffer from catastrophic forgetting. While existing works attempt to mitigate this on the basis of parameter-efficient fine-tuning methods, they adopted an overly restrictive Subspace Orthogonality condition. In this paper, we introduce a purely post-hoc and tuning-agnostic weight rectification framework that achieves Parameter Space Orthogonality, which is the necessary and sufficient condition for preserving historical performance to the fi...

*Notable papers are those with at least two authors from a "big" AI/ML lab.