📄 Notable* Recent AI/ML arXiv Papers

Last updated just now...

📄 Atria Dawn: The Dawn of Agentic Superintelligence
🗓️ Published: 9/14/2026
🔗 http://arxiv.org/abs/2609.15818v1
👥 Authors: Honglin Guo, Tao Gui, Yicheng Chen, Guanting Dong, Qiming Ge, Yuyang Hu, Zixian Huang, Jiajie Jin, Alexander Lam, Yining Li, Jiahang Lin, Yanjiang Liu, Xinyu Lu, Haijun Lv (possible past Baidu (China) affiliation), Junlin Shang, Qisheng Su, Guoqiang Wang, Rui Wang (possible past Tencent (China) affiliation), Zhecan Wang, Hao Xiang, Xinchen Xie, Shuhao Xing, Xiaoyu Xing, Wanghan Xu, Xinyu Yang, Yajie Yang, Chengfeng Zhao, Haoran Zhao, Ruojun Zhou, Yunhua Zhou, Yicheng Zou, Kun Cai, Qiye Cai, Xinmeng Che, Haodong Chen, Jiabei Chen, Jiahao Chen, Jiayi Chen, Yujia Chen, Lizhi Cui, Youheng Dai, Xin Deng, Yi Dong, Shihan Dou, Chenya Gu, Xu Guo, Ding Han, Feiyang Hao, Haotan He, Jie Hou, Binze Hu, Zijian Hu, Junhao Huang, Huicheng Jiang, Jiazhen Jiang, Shufan Jiang, Jiahao Kuang, Bowen Lai, Bo Li (possible past Tencent (China) affiliation), Jiaqiang Li, Peng Li (possible past Tsinghua University affiliation), Qilong Li, Zhuoqun Li, Jiaxiang Liu, Shuainan Liu, Tong Liu, Yi Liu (possible past Google (United States) affiliation), Zhonghang Lu, Jianwen Luo, Yanyi Luo, Huijie Lv, Ningsheng Ma, Zerun Ma, Houcheng Min, Chengjun Pan, Qiyuan Peng, Xiaoxuan Peng, Jianmin Qian, Jiantao Qiu, Wanying Ren, Huayu Sha, Jifei Shan, Zixin Shang, Bing Shao, Zhuohui Sheng, Jiayang Shi, Yang Shu, Aierpanjiang Simayi, Sirui Song, Yuxiao Song, Zhe Sun (possible past Tsinghua University affiliation), Zhichao Sun, Wenzhe Tan, Wenhui Tian, Zhongbo Tian, Hanchen Wang (possible past University Of Cambridge affiliation), Pengbo Wang, Rui Wang (possible past Tencent (China) affiliation), Yiding Wang, Yuhui Wang, Zhiheng Xi, Caijun Xu, Chao Xu, Yongfeng Xu, Xiaolei Yang, Zhixiong Yang, Qian Yao, Shihong Yi, Yuankai Ying, Jia Yu, Dingbo Yuan, Hao Yuan, Junjie Yuan, Bo Zhang (possible past Tencent (China) affiliation), Caixian Zhang, Qiuyinzhe Zhang, Jiyuan Zhao, Penghao Zhao, Ying Zhao (possible past Stanford University affiliation), Pujun Zheng, Xiaoxue Zhong, Xiaohao Zhou, Xinyu Zhou, Dongsheng Zhu, Guanru Zhu, Yulun Zhu, Yaojie Lu, Tao Ji, Hongyu Lin, Yutao Zhu, Pengfei Cao, Guoxiu He, Xianpei Han (possible past Tencent (China) affiliation), Ben He, Zhicheng Dou, Kang Liu, Qi Zhang (possible past Tencent (China) affiliation), Le Sun, Jun Zhao, Ji-Rong Wen, Xuanjing Huang, Yu-Gang Jiang, Bowen Zhou
Abstract

As AI agents become participants in the development of their successors, they reshape both the production of intelligence and the role of human researchers. We introduce Atria Dawn Preview, a foundation agentic language model designed for scientific research and engineering workflows, with the goal of expanding the frontier of agent productivity in the real world. This model is trained via a Verifiable Experience Pipeline that connects tool-mediated interactions to executable environments and ex...

📄 Navigating Sparse Evidence: Agentic Visual RAG via Explicit Context Selection and Consolidation
🗓️ Published: 9/14/2026
🔗 http://arxiv.org/abs/2609.15800v1
👥 Authors: Yucheng Shen, Lingyong Yan, Jiulong Wu, Shuaiqiang Wang (possible past Baidu (China) affiliation), Jianmin Wu, Dawei Yin (possible past Baidu (China) affiliation), Min Cao
Abstract

Visual Retrieval-Augmented Generation (VRAG) empowers models to navigate and answer queries about visually rich documents by retrieving relevant page images as visual evidence and reasoning over their content. However, effectively utilizing this visual evidence is usually impeded by two main challenges. First, answer-relevant evidence is sparse and may be concentrated in a small region of one page or dispersed across multiple pages. Second, existing agentic methods often generate answers based o...

📄 Bench2Dex: Benchmarking Visuo-Tactile Bimanual Dexterous Manipulation Across Dexterous Hands
🗓️ Published: 9/14/2026
🔗 http://arxiv.org/abs/2609.15726v1
👥 Authors: Zhenjie Yang, Yideng Zhang, Dongjie Zhang, Chenyu Jiang, Xianshuai Liu, Yufeng Li, Zuhao Ge, Xingyu Jiao, Zheng Zhang, Kaiyu He, He Wang (possible past Stanford University affiliation), Yuwen Zhong, Yi Deng, Muyun Jiang, Xianliang Huang, Haisheng Su, Donghang Zhang, Jian Zhang (possible past Tencent (China) affiliation), Xue Yang, Hongyang Li, Zuxuan Wu, Yu-Gang Jiang, Xiaosong Jia, Junchi Yan (possible past Shanghai Jiao Tong University affiliation)
Abstract

Tactile sensing provides contact information that can be difficult to infer from vision alone, but tactile hardware for dexterous hands has not converged to a common design. Dexterous hands differ in finger structure, contact surfaces, and sensor layouts, while simulated tactile signals still differ from measurements produced by physical sensors. These factors make it difficult to study visuo-tactile manipulation across diverse dexterous hands within a consistent experimental setting. We present...

📄 Circuit-MLLM: Topological Logic-Guided Latent-Space Visual Reasoning for Circuit Schematic Understanding
🗓️ Published: 9/14/2026
🔗 http://arxiv.org/abs/2609.15668v1
👥 Authors: Jinyuan Deng, Yuqi Jiang, Wenjing Huang, Xin Li (possible past Google (United States) affiliation), Qi Sun (possible past Google (United States) affiliation), Cheng Zhuo
Abstract

Through pre-training on extensive text and image datasets, current multi-modal large language models (MLLMs) achieve strong performance on general tasks. However, circuit schematics present a unique challenge for MLLMs due to their dense component layouts and distinct topological logic, demanding fine-grained structural parsing to extract the electrical semantics. To address this, we propose Circuit-MLLM, a multimodal reasoning framework that reformulates circuit topology analysis as a process o...

📄 A Unified Vision-Language Model for PSMA PET/CT Report Generation, Visual Question Answering, and Lesion Segmentation
🗓️ Published: 9/14/2026
🔗 http://arxiv.org/abs/2609.15603v1
👥 Authors: Yang Xing, Jiong Wu, Savas Ozdemir, Yang Zhou, Boxiao Yu, Ying Zhang (possible past Tencent (China) affiliation), Zheren Zhu, Chenyu You, Wei Shao (possible past Stanford University affiliation), Yang Lu (possible past Meta (United States) affiliation), Kang Wang, Tinsu Pan, Yang Yang (possible past Tencent (China) affiliation), Kuang Gong
Abstract

Accurate PSMA PET/CT interpretation is central to prostate cancer management, yet existing PET/CT AI models typically address isolated tasks. We propose a unified PSMA PET/CT vision-language model for report generation, visual question answering, and lesion segmentation. The framework adopts an LLaVA-style architecture, comprising a PET/CT vision encoder, an MLP-Mixer projection module, a LoRA-tuned large language model, and a 3D segmentation branch. Training followed a four-stage strategy: visi...

📄 Self-Evolving Memory for Generative Recommendation
🗓️ Published: 9/14/2026
🔗 http://arxiv.org/abs/2609.15598v1
👥 Authors: Xinyu Lin, Zhuosong Jiang, Zixiao Suo, Siqin Wang, Hanqing Zeng, Hanchao Yu, Yinglong Xia, Jiang Zhang (possible past Google (United States) affiliation), Aashu Singh, Fei Liu, Wenjie Wang, Fuli Feng (possible past National University Of Singapore affiliation), Yang Song (possible past Stanford University affiliation), Qifan Wang (possible past Google (United States) affiliation), Tat-Seng Chua
Abstract

Generative recommendation has emerged as a promising end-to-end paradigm for personalized recommendation. However, user preferences continuously evolve over time, making self-evolving an essential capability for generative recommender systems. Existing evolving strategies, such as continual retraining and distillation-based adaptation, directly update the shared model parameters using streaming interactions. Nevertheless, we find that directly applying such strategies to generative recommendatio...

📄 PIVOT: Physics-Grounded Verification for AI-Generated Audio-Video Detection
🗓️ Published: 9/14/2026
🔗 http://arxiv.org/abs/2609.15562v1
👥 Authors: Bo Zheng, Kangran Zhao, Xiaoyu Zhang, Weinan Guan, Zhiheng Li, Yize Chen, Haizhou Li, Qingshan Liu, Siwei Lyu (possible past University Of Washington affiliation), Baoyuan Wu (possible past Tencent (China) affiliation)
Abstract

As generative models continue to advance, AI-generated content (AIGC) is becoming increasingly realistic, weakening the artifact cues commonly exploited by existing detectors. Nevertheless, faithfully reproducing the physical behavior of real-world events remains challenging for current generators. We therefore explore detecting AIGC by assessing whether the depicted event satisfies measurable constraints derived from physical laws. We introduce PIVOT, a physics-grounded AIGC detector, instantia...

📄 Planning in the Backbone: DiffAdapterVLA for Native Continuous Trajectory Generation with Driving VLMs
🗓️ Published: 9/14/2026
🔗 http://arxiv.org/abs/2609.15322v1
👥 Authors: Changxin Lu, Xiaoliang Meng, Yu Wu (possible past Baidu (China) affiliation), Rui Huang (possible past Google (United States) affiliation), Honglin Li, Tao Chen, Kaixuan Zhou, Yadong Shao
Abstract

Pretrained driving vision-language models (VLMs) integrate visual, route, language, and driving context into rich driving priors, yet their representation objectives remain separated from continuous driving planning. Existing methods typically begin trajectory generation only after the VLM has formed a final condition, leaving depth-wise condition computation outside the stepwise formation of trajectory state. We introduce DiffAdapterVLA, which realizes Planning in the Backbone: it injects expli...

📄 PACE: Progressive Angular-to-Norm Contrastive Embedding
🗓️ Published: 9/14/2026
🔗 http://arxiv.org/abs/2609.15152v1
👥 Authors: Yanping Li, Wei Zhou, Yawen Liu, Yibo Wang, Ke Zhu, Guangda Huzhang, Qing-Guo Chen, Zhao Xu, Jun Zhang (possible past Tencent (China) affiliation), Wei Wei (possible past Google (United States) affiliation)
Abstract

Multimodal embedding models encode heterogeneous inputs into a shared embedding space, enabling efficient similarity computation across modalities and tasks. Most existing methods optimize cosine-based contrastive objectives, which promote stable training but restrict semantic compatibility to angular geometry, precluding embedding norms from serving as an additional semantic signal. However, directly optimizing the more expressive dot-product similarity, which leverages both angular and norm in...

📄 OpenAl4S: Code as Action, Science as Sessions
🗓️ Published: 9/14/2026
🔗 http://arxiv.org/abs/2609.15096v1
👥 Authors: Gongbo Zhang, Hao Li (possible past Tsinghua University affiliation), Yu Wang (possible past Tsinghua University affiliation), Mujie Lin, Liuzhenghao Lv, Yicheng Mao, Yimi Wang, Jun Zhu (possible past Tsinghua University affiliation), Minhan Tang, Zhengxiang Jiang, Yusong Wang, Jiayu Yao, Kunpeng Ning, Dawei Pang, Yonghong Tian (possible past Peking University affiliation), Openai4s Community, Yuyang Liu, Li Yuan (possible past National University Of Singapore affiliation)
Abstract

AI co-scientists could accelerate computational research, but over a long-running study the workflow also has to stay inspectable, resumable and reproducible, which requires persistent computational state and provenance. Here we present OpenAI4S, an open-source scientific research agent built around the principle of \emph{Code as Action, Science as Sessions}. OpenAI4S combines a persistent computing runtime with research-session management: orchestration is handled through structured tool calls,...

📄 Forty Shades of Blue: Quality-Diversity Alignment via Mode-Conditioned Reinforcement Learning
🗓️ Published: 9/14/2026
🔗 http://arxiv.org/abs/2609.14896v1
👥 Authors: Jiayi Yuan, Hangoo Kang, James Jihao Liu, Yejin Choi (possible past Allen Institute For Artificial Intelligence affiliation), Vikram Iyer, Liwei Jiang, Natasha Jaques (possible past University Of California, Berkeley affiliation)
Abstract

A notable byproduct of LLM alignment training is mode collapse: the progressive loss of output diversity that narrows a model's expressivity at inference time. This degradation is especially limiting for applications requiring open-ended exploration and pluralistic perspectives, such as scientific ideation and creative writing. We present MoDA (Mode-conditioned Diversity Alignment), an online post-training RL algorithm that jointly optimizes generation quality and diversity, inspired by the coor...

📄 LLMs as Oracles: Reliance on LLMs for Subjective Personal Questions
🗓️ Published: 9/13/2026
🔗 http://arxiv.org/abs/2609.14849v1
👥 Authors: Myra Cheng (possible past Deepmind (United Kingdom) affiliation), Lujain Ibrahim, Grace Liu, Michelle S. Lam (possible past Stanford University affiliation), Vishakh Padmakumar, Nick Madibekov, Diyi Yang (possible past Stanford University affiliation), Dan Jurafsky (possible past Stanford University affiliation)
Abstract

We characterize how people are turning to LLMs as oracles: all-knowing authorities on subjective personal questions. Motivated by risks to users' autonomy and well-being, we develop a typology and LLM-based methods to measure this form of AI reliance at scale and understand how people are offloading judgment and decision-making to AI. Applying our typology to public usage data (68K prompts from WildChat and ThoughtTrace), we find that LLM-as-oracle use has increased over time (2023-2026) and is ...

📄 MoveBench: A Benchmark for Global-Scale Wildlife Movement Forecasting
🗓️ Published: 9/14/2026
🔗 http://arxiv.org/abs/2609.15780v1
👥 Authors: Justin Kay, Shir Bar, Ellen O. Aikens, Martin Becker, Francesca Cagnacci, Juliet Cohen, Scott W. Forrest, Jessica Kendall-Bar, Madeleine Lucas, Macon Overcast, Meredith S. Palmer, Will Rogers, Nicholas J. Russo, Christian Rutz, Larissa T. Beumer, Michael Brown (possible past Google (United States) affiliation), Ying-Chi Chan, Sarah C. Davidson, Diego Ellis Soto, Anne G. Hertel, Roland Kays, Benjamin Koger, Guram Mikaberidze, Thomas Mueller, Ruth Oliver, Thorsten Papenbrock, Robert Patchett, Jared A. Stabach, Dane Taylor, Scott W. Yanco, Sara Beery (possible past Microsoft (United States) affiliation)
Abstract

Understanding and predicting wildlife movement is critical for ecology and conservation. While trajectory forecasting has advanced for human and vehicle movement, wildlife trajectories present distinct challenges: they are unconstrained in space, highly stochastic, and influenced by environmental conditions. We introduce MoveBench, the first large-scale benchmark for probabilistic wildlife movement forecasting, containing 2.6M GPS locations from 800+ individuals across 110 species in 127 countri...

📄 The Token Before the Value Is the Key: How Hybrid Architectures Organize Induction Circuits
🗓️ Published: 9/14/2026
🔗 http://arxiv.org/abs/2609.15545v1
👥 Authors: Ke Cheng, Xin Xu, Yixiao Chen, Lei Xin, Jianbo Zhao, Fanhu Zeng, Yue Liu, Jun Zhang (possible past Tencent (China) affiliation), Jie Jiang (possible past Tencent (China) affiliation)
Abstract

Hybrid language models can improve capability as well as efficiency, raising the question of how architectural complementarity becomes learned computation. We examine the established induction roles of Carrying predecessor information, Matching a source by content, and Copying its value. How are these position-sensitive and content-based computations allocated across heterogeneous layers? We introduce layer-type-agnostic paired probes that track Carrying and Matching through a common block-updat...

📄 Temporal Self-Distillation: Faster Inference in Discrete Diffusion Language Models
🗓️ Published: 9/14/2026
🔗 http://arxiv.org/abs/2609.15177v1
👥 Authors: Shijian Xu, Andrea Miele (possible past Nvidia (United States) affiliation), Metod Jazbec, Volker Roth, Eric Nalisnick (possible past Google (United States) affiliation), Ilija Bogunovic
Abstract

Diffusion language models (dLLMs) promise fast inference by generating multiple tokens in parallel, but suffer severe performance degradation when parallel decoding is pushed too aggressively. We introduce Temporal Self-Distillation (TSD), a simple on-policy method that trains dLLMs for fast inference by distilling predictions across time. Specifically, TSD distills the model's denoising distribution at earlier timesteps toward its distribution at the final timestep at which a token is committed...

📄 Know When to Stop, Where to Restart: Accelerating Multi-Turn Agentic On-Policy Distillation
🗓️ Published: 9/13/2026
🔗 http://arxiv.org/abs/2609.14636v1
👥 Authors: Zhiyu Gui, Kexin Huang (possible past Stanford University affiliation), Jia Guo, Junkang Wu, Zihao Wang, Zhiqiang Zhang, Jun Zhou, Jiancan Wu, Xiang Wang (possible past Tencent (China) affiliation)
Abstract

On-policy distillation (OPD) has become a standard approach for transferring capabilities from large teachers to compact students. Its cost, however, is dominated by autoregressive student rollouts and scales poorly in multi-turn agentic settings. Existing acceleration methods truncate or relocate the supervision signal according to fixed, offline budgets, despite substantial variation in teacher-signal reliability both within and across trajectories. Our empirical analysis on $τ^2$-bench reveal...

*Notable papers are those with at least two authors from a "big" AI/ML lab.