πŸ“„ Notable* Recent AI/ML arXiv Papers

Last updated just now...

πŸ“„ Agent-Editing World Model: Rethinking World Modeling for LLM Agents
πŸ—“οΈ Published: 9/23/2026
πŸ”— http://arxiv.org/abs/2609.28416v1
πŸ‘₯ Authors: Shuang Sun, Guoxin Chen, Fanzhe Meng, Jia Deng (possible past Google (United States) affiliation), Huatong Song, Jinhao Jiang, Wayne Xin Zhao (possible past Baidu (China) affiliation), Hongteng Xu, Ji-Rong Wen
Abstract

Recent advances in large language models (LLMs) have enabled agents to tackle long-horizon tasks across diverse environments. To further improve agent performance, existing language world models typically predict environment observations, yet reconstructing high-entropy, execution-dependent tool responses offers limited value when real feedback is available. Meanwhile, agents suffer from \emph{task-state contamination}, where unsupported assumptions and outdated plans persist in history and dist...

πŸ“„ InternW0: A Foundational Physical World Model for Efficient Real-World Interactions
πŸ—“οΈ Published: 9/23/2026
πŸ”— http://arxiv.org/abs/2609.27656v1
πŸ‘₯ Authors: Jisong Cai, Yao Mu, Ganlin Yang, Zhe Cao (possible past University Of California, Berkeley affiliation), Zhangzheng Tu, Xing Gao, Kailin Li, Xinyu Zhan, Lixin Yang, Yangkun Zhu, Haoxiang Ma, Ming Zhou, Qiaojun Yu, Yufei Xue, Liqun He, Yifei Yao, Yifan Zhu, Long Ling, Bingqi Jiang, Haoyu Guo, Xueyue Zhu, Bowen Zhou, Bin Zhao, Tianfan Xue (possible past Massachusetts Institute Of Technology affiliation), Chunhua Shen, Weinan Zhang (possible past Shanghai Jiao Tong University affiliation)
Abstract

Physical intelligence requires more than predicting how the world may evolve: predictions must remain actionable as the world continues to change. We introduce InternW0, the first instantiation of the InternW physical world model series from Shanghai AI Laboratory, built around omnimodal interfaces, asynchronous multi-frequency processing, and local physical modeling under partial observations and external influences. InternW0 jointly learns future visual dynamics and continuous robot control th...

πŸ“„ NV-Reason-CT: 3D Visual Language Model for CT Analysis
πŸ—“οΈ Published: 9/23/2026
πŸ”— http://arxiv.org/abs/2609.27511v1
πŸ‘₯ Authors: Andriy Myronenko (possible past Nvidia (United States) affiliation), Dong Yang (possible past Nvidia (United States) affiliation), Yucheng Tang, Baris Turkbey, Benjamin Simon, Stephanie Harmon, Rikhil Makwana, Mariam Aboian, Sena Azamat, Ibrahim Ethem Hamamci, Sezgin Er, Bjoern Menze, Marc Edgar, Yufan He (possible past Nvidia (United States) affiliation), Pengfei Guo, Daguang Xu (possible past Nvidia (United States) affiliation)
Abstract

We present NV-Reason-CT, a generative vision--language model for chest and abdominal CT combining native 3D visual encoding with radiologist-guided reasoning. The model couples a native 3D vision transformer with a language model, passing all visual tokens and their explicit 3D coordinates into language decoding without further spatial token merging. This retains volumetric spatial information within the vision encoder and through the language model's positional encoding during joint processing ...

πŸ“„ CART: Closed-Loop Adaptive Red Teaming for Large Language Models
πŸ—“οΈ Published: 9/23/2026
πŸ”— http://arxiv.org/abs/2609.27336v1
πŸ‘₯ Authors: Dongdong Zhang, Tengchao Lv, Yilin Jia, Yuzhong Zhao, Yupan Huang, Wenshan Wu, Xiangyang Zhou (possible past Baidu (China) affiliation), Shaohan Huang, Nan Yang, Li Dong, Lei Cui (possible past Tsinghua University affiliation), Furu Wei
Abstract

Automated red teaming often replays a fixed set of prompts, which measures known risks but cannot learn from failures found during testing. We present CART (Closed-Loop Adaptive Red Teaming), a framework that uses each result to guide what it tests next. CART begins with broad risk coverage, follows weaknesses that emerge, keeps new probes diverse, and records the evidence and source of every finding. It separates the Challenger that creates tests, the Target being tested, which may be a text-on...

πŸ“„ Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents
πŸ—“οΈ Published: 9/23/2026
πŸ”— http://arxiv.org/abs/2609.27334v1
πŸ‘₯ Authors: Yefan Zhou, Yang Li (possible past Google (United States) affiliation), Zeyu Leo Liu, Semih Yavuz (possible past Google (United States) affiliation), Shafiq Joty
Abstract

Agentic memory systems reuse past experience to improve future performance, yet most existing designs curate memory at write time: once a task is completed, its trajectory is distilled into a fixed artifact, such as a reflection, workflow, skill, or reasoning strategy, that is later retrieved by similarity. This forces the system to decide what is worth remembering before the future query is known, irreversibly discarding information and producing a query-independent summary that must serve many...

πŸ“„ Stable Geometry with Divergent Task Evidence for Efficient Long-Horizon Agent Compression
πŸ—“οΈ Published: 9/23/2026
πŸ”— http://arxiv.org/abs/2609.27332v1
πŸ‘₯ Authors: Mingxuan Wang (possible past Tencent (China) affiliation), Fei Luo, Bo Wang (possible past Tencent (China) affiliation), Guorun Yao, Yinglong Guo, Chao Ning, Hongyue Chen, Yanbiao Ma, Jungong Han
Abstract

Long horizon agents accumulate growing interaction histories that increase context and inference costs. We find that geometric redundancy alone is an insufficient criterion for safe compression. Although agent histories exhibit strong low dimensional structure, similar global geometry can preserve very different amounts of task evidence. At identical retained block counts, evidence aware selection raises next action Top 3 retention from 0.31 to 0.69, while centroid similarity remains 0.98. Contr...

πŸ“„ StateComp: Learning When to Compress History in Long Horizon Agents
πŸ—“οΈ Published: 9/23/2026
πŸ”— http://arxiv.org/abs/2609.27298v1
πŸ‘₯ Authors: Mingxuan Wang (possible past Tencent (China) affiliation), Hongyue Chen, Yinglong Guo, Fei Luo, Chao Ning, Bo Wang (possible past Tencent (China) affiliation), Guorun Yao, Yanbiao Ma, Jungong Han
Abstract

Long-horizon agents continuously accumulate interaction history during task execution, yet the importance of past interactions changes as the agent state evolves. Existing context management methods largely compress history based on fixed windows, periodic schedules, or current relevance, overlooking a more fundamental question: when has a past interaction become safe to replace? Premature compression may remove information still needed for future actions, while overly conservative retention lea...

πŸ“„ Large Knowledge Model: From Papers to a Scientific Reasoning Landscape
πŸ—“οΈ Published: 9/23/2026
πŸ”— http://arxiv.org/abs/2609.27297v1
πŸ‘₯ Authors: Yuan Huang, Sihan Hu, Hongyu Gu, Chao Ma (possible past Shanghai Jiao Tong University affiliation), Jiaxing Zhang, Zhiyong Zou, Caiyu Fan, Yan Xiao, Mingjun Xu, Chenyu Xie, Mingzhen Ju, Zhehao Ma, Qi Zhang (possible past Tencent (China) affiliation), Baozong Wang, Yu Li (possible past Tencent (China) affiliation), Zhiyuan Yao, Ruoxue Liao, Xinyu Li, Linfeng Zhang, Kun Chen, Weinan E
Abstract

Accumulated scientific knowledge advances inquiry when prior findings help researchers choose new questions, design investigations, and interpret results. Realizing this value at scale requires access to the reasoning that connects research problems, scientific procedures, conclusions, and evidence. We introduce the Large Knowledge Model (LKM), a scientific knowledge infrastructure that transforms the literature into a shared, computationally accessible reasoning resource. LKM represents papers ...

πŸ“„ KITE: KV-Invariant Transformer Expansion for Efficient Agentic LLM Scaling
πŸ—“οΈ Published: 9/23/2026
πŸ”— http://arxiv.org/abs/2609.27294v1
πŸ‘₯ Authors: Zhiheng Hu, Yixun Wei, Jian Zhou (possible past Tencent (China) affiliation), Yizhuang Zhou, Ji Li, Xing Chen, Yang Li (possible past Google (United States) affiliation), Bojun Wang, Yibo Zhu, Xiangyu Zhang, Daxin Jiang
Abstract

Scaling a language model is not only a question of final quality: the architectural choice determines how much computation is spent during training, prompt processing, and autoregressive decoding to achieve certain model quality. An ideal model architecture should lower all above computation costs to facilitate scaling to a larger model, while ensure the larger model indeed outperforms smaller baselines. We introduce KV-Invariant Transformer Expansion (KITE), a scaling paradigm that achieves thi...

πŸ“„ Ruby-ASR: Evidence-Preserving Supervision for Joint Orthographic and Lexical-Reading Recognition
πŸ—“οΈ Published: 9/23/2026
πŸ”— http://arxiv.org/abs/2609.27289v1
πŸ‘₯ Authors: Hao Shi, Yun Liu (possible past Google (United States) affiliation), Xuehao Yang, Jun Liu (possible past Tencent (China) affiliation), Chuanbo Hua, Xuanjun Chen, Lianbo Liu, Shiao Zhu, Zixiong Su
Abstract

Conventional Japanese automatic speech recognition (ASR) is supervised by an orthographic transcript, although the same written form can correspond to different lexical readings realized in speech. Such utterances receive an identical target, so their reading distinction is absent from the supervision interface and cannot be recovered reliably by post-hoc text-only grapheme-to-phoneme conversion. We present Ruby-ASR, which refines the conventional target into a span-bound orthographic--lexical-r...

πŸ“„ Memory Control Signals Emerge Before Action in Long Horizon Agents
πŸ—“οΈ Published: 9/23/2026
πŸ”— http://arxiv.org/abs/2609.27286v1
πŸ‘₯ Authors: Mingxuan Wang (possible past Tencent (China) affiliation), Guorun Yao, Fei Luo, Yinglong Guo, Chao Ning, Bo Wang (possible past Tencent (China) affiliation), Hongyue Chen, Yanbiao Ma, Jungong Han
Abstract

Long horizon language model agents continuously accumulate interaction history, increasing computational cost while making relevant information harder to preserve and reuse. Existing context management methods mainly focus on how to compress or retrieve history, but largely leave open whether the model itself already represents the need for these memory operations before they occur. We study the hidden state immediately before each agent action and find that compression and recall needs are alre...

πŸ“„ Hunyuan-A13B Technical Report
πŸ—“οΈ Published: 9/23/2026
πŸ”— http://arxiv.org/abs/2609.27284v1
πŸ‘₯ Authors: Tencent Hunyuan Team, Ao Liu (possible past Tencent (China) affiliation), Botong Zhou, Can Xu (possible past Google (United States) affiliation), Chayse Zhou, Chenchen Zhang, Chengcheng Xu, Chenhao Wang, Decheng Wu, Dengpeng Wu, Dian Jiao, Dong Du, Dong Wang (possible past Tsinghua University affiliation), Feng Zhang, Fengzong Lian, Guanghui Xu, Guanwei Zhang, Hai Wang, Haipeng Luo, Han Hu, Huilin Xu, Jiajia Wu, Jianchen Zhu, Jianfeng Yan (possible past Peking University affiliation), Jiaqi Zhu, Jihong Zhang, Jinbao Xue, Jun Xia, Junqiang Zheng, Kai Liu (possible past Baidu (China) affiliation), Kai Zhang, Kai Zheng, Kejiao Li, Keyao Wang, Lan Jiang, Lixin Liu, Lulu Wu, Mengyuan Huang, Peijie Yu, Peiqi Wang, Qian Wang, Qianbiao Xiang, Qibin Liu, Qingfeng Sun, Richard Guo, Ruobing Xie (possible past Tencent (China) affiliation), Saiyong Yang, Shaohua Chen, Shihui Hu, Shuai Li, Shuaipeng Li, Shuang Chen, Suncong Zheng, Tao Yang, Tian Zhang, Tinghao Yu, Weidong Han, Weijie Liu (possible past Tencent (China) affiliation), Weijin Zhou, Weikang Wang, Wesleye Chen, Xiao Feng, Xiaoqin Ren, Xingwu Sun (possible past Baidu (China) affiliation), Xiong Kuang, Xuemeng Huang, Xun Cao, Yanfeng Chen, Yang Du, Zhen Yang (possible past Tsinghua University affiliation), Yangyu Tao, Yaping Deng, Yi Shen (possible past Baidu (China) affiliation), Yigeng Hong, Yiqi Chen
Abstract

We present Hunyuan-A13B, an open-source large language model based on a Mixture-of-Experts architecture. It contains 80 billion total parameters but activates only 13 billion during inference, balancing model capability, computational efficiency, and deployment cost. The model is pretrained on a rigorously filtered 20T-token corpus with enhanced STEM data curation, improving factual reliability and reasoning ability. High-quality supervised fine-tuning and large-scale reinforcement learning furt...

πŸ“„ TimeEvo: Failure-Driven Self-Evolution of a Time Series Agent
πŸ—“οΈ Published: 9/23/2026
πŸ”— http://arxiv.org/abs/2609.27277v1
πŸ‘₯ Authors: Jie Yang (possible past Shanghai Jiao Tong University affiliation), Yan Zheng, Jiarui Sun (possible past Tencent (China) affiliation), Xiran Fan, Junpeng Wang, Liang Wang (possible past Tencent (China) affiliation), Zelin Xu, Qinghua Liu, Zhengyu Fang, Yiwei Cai, Philip S. Yu (possible past Tsinghua University affiliation)
Abstract

Time series agents answer analytical questions by calling external tools, and which tools they carry is decided by people before the agent runs. However, we identify two failures in this setup. Human-Agent Tool Misalignment: a library of 21 expert-curated tools helps on some tasks and hurts on others, dropping anomaly accuracy under every backbone we test. Silent Harm: one round of generic self-revision changes 147 answers and breaks 56 of them, while the final score moves by less than a point. ...

πŸ“„ DRSR: Learning Set-Level Deletion Risk for Efficient Long-Horizon Agents
πŸ—“οΈ Published: 9/23/2026
πŸ”— http://arxiv.org/abs/2609.27276v1
πŸ‘₯ Authors: Mingxuan Wang (possible past Tencent (China) affiliation), Bo Wang (possible past Tencent (China) affiliation), Fei Luo, Guorun Yao, Chao Ning, Yinglong Guo, Hongyue Chen, Yanbiao Ma, Jungong Han
Abstract

Long-horizon language-model agents accumulate reasoning traces, tool exchanges, and observations whose relevance changes with the current decision. Existing compression strategies often score historical units independently, but the safety of deleting several units is generally not determined by their singleton scores: redundant evidence, accumulated small effects, and the information that remains after deletion all matter. We introduce Direct Relational Set-Risk Pruning (DRSR), which formulates ...

πŸ“„ EMA: Elastic and Performance Transparent Memory Across GPUs
πŸ—“οΈ Published: 9/22/2026
πŸ”— http://arxiv.org/abs/2609.27040v1
πŸ‘₯ Authors: Yi Xu, Tian Xia (possible past Baidu (China) affiliation), Ion Stoica (possible past University Of California, Berkeley affiliation)
Abstract

Multi-GPU servers have become the standard building block of modern data centers, providing aggregated capacity through high-bandwidth interconnects. At the same time, workloads such as LLM inference exhibit highly dynamic memory demands, which can cause one GPU to exhaust its local memory while others remain underutilized. This mismatch motivates a model of elastic resource sharing across GPUs. We present EMA, a memory sharing system that allows GPUs within a server to borrow and reclaim memo...

πŸ“„ Train Where the Quantized Model Goes: On-Policy Distillation for Low-Bit Reasoning
πŸ—“οΈ Published: 9/22/2026
πŸ”— http://arxiv.org/abs/2609.26708v1
πŸ‘₯ Authors: Yuanteng Chen, Zhilei Liu, Peisong Wang, Yuantian Shao, Chuangyi Li, Weining Wang, Shuang Qiu, Gang Li (possible past Tsinghua University affiliation), Jing Liu (possible past Baidu (China) affiliation), Jian Cheng
Abstract

Quantization-aware distillation (QAD) restores much of the short-form question-answering performance lost to sub-3-bit quantization, yet leaves mathematical and code reasoning substantially impaired. Long generations often degenerate into repetitive loops, exhausting the decoding budget without completing a solution. We trace this gap to quantization-amplified exposure bias: QAD trains on fixed corpus prefixes, while quantization-induced deviations compound along the model's own autoregressive t...

πŸ“„ Safety Nudges: User-Facing Interventions for Real-Time AI Risk Awareness
πŸ—“οΈ Published: 9/22/2026
πŸ”— http://arxiv.org/abs/2609.26865v1
πŸ‘₯ Authors: Varshini Elangovan, James Wedgwood, Chhavi Yadav, William Agnew, Sauvik Das (possible past Carnegie Mellon University affiliation), Virginia Smith (possible past Carnegie Mellon University affiliation)
Abstract

Conversational AI systems can pose safety risks to their users such as hallucination, sycophancy, overconfidence, and anthropomorphism, but these risks are difficult for users to detect during everyday use. We introduce Safety Nudges, a browser-based tool that provides lightweight, in situ flags when concerning behavior is detected in chatbot conversations. We evaluated Safety Nudges in a two-week field study with 45 frequent chatbot users, collecting interaction logs, surveys, and feedback on i...

πŸ“„ QuantWM: Temporally Consistent 2-Bit KV Cache Quantization for World Models and Video Generation
πŸ—“οΈ Published: 9/22/2026
πŸ”— http://arxiv.org/abs/2609.26425v2
πŸ‘₯ Authors: Jiaqi Zhao, Xiaobin Hu (possible past Tencent (China) affiliation), Bo Yin, Junpeng Jiang, Miao Zhang (possible past Stanford University affiliation), Shuicheng Yan (possible past National University Of Singapore affiliation)
Abstract

KV cache memory has become a major deployment bottleneck for video generation and world models, which motivates low-bit quantization study for efficiency. Existing 2-bit KV cache quantization methods can achieve nearly lossless performance on video benchmarks such as VBench, however, we find that they still cause severe temporal flickering and visual degradation. Meanwhile, deeper investigates show that Key quantization produces smaller reconstruction errors than Value, but surprisingly leads to...

πŸ“„ Repairability of Inexact Solvers in Recursive State Estimation with Machine Learning
πŸ—“οΈ Published: 9/23/2026
πŸ”— http://arxiv.org/abs/2609.28425v1
πŸ‘₯ Authors: Yanjun Ji, Dennis Willsch, Orkun Şensebat, Priyanka Arkalgud Ganeshamurthy, Zhi Pei, M. Sahnawaz Alam, Ivelina Stoyanova, Frank K. Wilhelm, Bo Zhao (possible past National University Of Singapore affiliation), Chao Wang (possible past Google (United States) affiliation), Kristel Michielsen
Abstract

Recursive state estimation often executes approximate numerical solutions inside a feedback loop, where highly accurate local steps do not guarantee better overall results. For a fixed linear Kalman model, we characterize when a correction within a prescribed subspace and norm budget can meet a local admissibility tolerance, and how the defects actually executed affect the finite-horizon covariance response. Centering each defect on the exact gain for the implemented covariance separates current...

πŸ“„ Generalizable Robotic Insertion with World Models
πŸ—“οΈ Published: 9/23/2026
πŸ”— http://arxiv.org/abs/2609.28258v1
πŸ‘₯ Authors: Nicklas Hansen, Iretiayo Akinola, Yijie Guo, Jie Xu, Bingjie Tang, Hao Su, Xiaolong Wang (possible past Carnegie Mellon University affiliation), Abhishek Gupta (possible past University Of California, Berkeley affiliation), Dieter Fox (possible past University Of Washington affiliation), Yashraj Narang
Abstract

Robotic assembly in high-mixture settings requires adaptable systems that can handle diverse parts, yet current approaches typically rely on policies specialized to each insertion task. Although this can reach high success rates, it makes the process of deploying systems for new problems tedious and time consuming. We present a framework for generalizable insertion using world models that combine robot proprioceptive information with raw visual observations captured by a wrist-mounted camera. Ou...

πŸ“„ EBRL: Asynchronous Embodied RL by Multi-Grained Resource Management
πŸ—“οΈ Published: 9/23/2026
πŸ”— http://arxiv.org/abs/2609.27547v1
πŸ‘₯ Authors: Liang Mi, Weijun Wang (possible past Google (United States) affiliation), Bowen Gao, Tianze Yu, Zixu Hao, Han Xiao, Xin Ding, Mingzhe Huang, Xin He, Lu Shi, Hao Wu (possible past Tencent (China) affiliation), Haipeng Dai, Guihai Chen (possible past Shanghai Jiao Tong University affiliation), Yunxin Liu, Ting Cao
Abstract

Embodied reinforcement learning (RL) improves model capabilities with a pipeline of environment simulation, action generation, and model updates. These stages show heterogeneous CPU and GPU demands, making efficient resource utilization difficult. Recent systems overlap rollout (simulation and generation) with training for efficiency, but exclusive GPU allocation and synchronized barrier in rollout still leave substantial hardware resource waste. In this paper, we present EBRL, an asynchronous e...

πŸ“„ Prediction with Expert Advice: Anytime Regret with Many Experts Matches the Fixed-Time Constant
πŸ—“οΈ Published: 9/23/2026
πŸ”— http://arxiv.org/abs/2609.27206v1
πŸ‘₯ Authors: Yang Cai, Vineet Gupta (possible past Google (United States) affiliation), Yanchen Jiang, Christopher Liaw, Aranyak Mehta (possible past Google (United States) affiliation), Grigoris Velegkas, Di Wang
Abstract

Prediction with expert advice is a fundamental problem in online learning. When the time horizon $T$ is known in advance, the minimax cumulative regret over $n$ experts is asymptotically $\sqrt{\frac{T \ln n}{2}}$. This is achieved by the Multiplicative Weights Update algorithm with a learning rate tuned to $T$, and is known to be tight. If instead the regret bound is required to hold simultaneously at every time $t$, the best known guarantee has been $\sqrt{t \ln n}$---a factor of $\sqrt{2}$ wo...

πŸ“„ GTR: Gated Token Recurrence for Efficient Dense Prediction
πŸ—“οΈ Published: 9/22/2026
πŸ”— http://arxiv.org/abs/2609.26590v2
πŸ‘₯ Authors: Zhe Feng, Longfei Liu, Wei Liu (possible past Tsinghua University affiliation), Kai Chen (possible past Shanghai Jiao Tong University affiliation), Jiangang Kong, Wei Zhou, Yifeng Qian, Dexiong Chen, Xuanlong Yu, Xi Shen (possible past Tencent (China) affiliation)
Abstract

Self-attention-based vision backbones perform well on dense prediction, but the quadratic computational cost of global softmax attention limits their efficiency as image resolution increases. We introduce Gated Token Recurrence (GTR), a softmax-free recurrent vision backbone that combines gated linear attention, alternating spatial scan directions, and spatially enhanced SwiGLU blocks. GTR is distilled from a detection-specialized DINOv3 teacher using only final-layer patch-token alignment throu...

*Notable papers are those with at least two authors from a "big" AI/ML lab.