πŸ“„ Notable* Recent AI/ML arXiv Papers

Last updated just now...

πŸ“„ Base Models Can Reason By Taking a Cue From Training Data
πŸ—“οΈ Published: 10/5/2026
πŸ”— http://arxiv.org/abs/2610.06851v1
πŸ‘₯ Authors: Sophie L. Wang, Amil Dravid, Rulin Shao, Kevin Farhat, Sewon Min (possible past University Of Washington affiliation), Alexei A. Efros (possible past University Of California, Berkeley affiliation)
Abstract

In this paper, we study how training data creates associations between the tokens at the start of a base model's response and the reasoning behavior that follows. First, we demonstrate that fixing particular starting token cues makes a base model's performance competitive with that of its reinforcement learning (RL)-trained counterparts on math and coding. For instance, the cue ".\n\nOkay" raises Olmo-3-7B's MATH-500 pass@1 accuracy from 42% to 78%, while "Alright," raises Qwen3-14B's from 72% t...

πŸ“„ MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers
πŸ—“οΈ Published: 10/5/2026
πŸ”— http://arxiv.org/abs/2610.06801v1
πŸ‘₯ Authors: Jiarui Chen, Zeqiang Lai, Jiangshan Wang, Ziheng Ouyang, Ye Huang, Xiangyu Yue (possible past University Of California, Berkeley affiliation), Cewu Lu (possible past Shanghai Jiao Tong University affiliation), Chunchao Guo
Abstract

Sparse attention is a primary approach to reducing the latency of diffusion transformers in long-sequence generation tasks, such as video and high-resolution 3D asset generation. However, existing methods can degrade generation quality and fidelity at high sparsity levels. Through controlled oracle comparisons, we trace this degradation to three sources: constraints imposed by token grouping, inaccurate interaction selection, and the attention contributions lost when tokens are discarded. Guided...

πŸ“„ IdeaLens: Detecting AI Ideas in Long-form Writing
πŸ—“οΈ Published: 10/5/2026
πŸ”— http://arxiv.org/abs/2610.06778v1
πŸ‘₯ Authors: Rishanth Rajendhran, Minjoon Choi, Jenna Russell, Ramya Namuduri, Deniz BΓΆlΓΆni-Turgut, Marzena Karpinska, John Wieting (possible past Google (United States) affiliation), Mohit Iyyer (possible past Google (United States) affiliation)
Abstract

While modern AI detectors identify who wrote the words, emerging policies on AI use increasingly hinge on a different question: who came up with the ideas? We introduce IdeaLens, a detector that identifies whether a document's ideas came from a human or AI (idea provenance), regardless of who wrote its words. To focus IdeaLens on ideas rather than prose, we represent documents as outlines: lists of items that each pair a discourse role with a brief, paraphrased description of the content, minimi...

πŸ“„ BRANCH-MoE: Balance-Aware Tree Routing for Large Embedding Models
πŸ—“οΈ Published: 10/5/2026
πŸ”— http://arxiv.org/abs/2610.06725v1
πŸ‘₯ Authors: Gang Fu (possible past Google (United States) affiliation), Adel Javanmard, Mohammadhossein Bateni (possible past Google (United States) affiliation), Vahab Mirrokni (possible past Google (United States) affiliation)
Abstract

Mixture-of-experts (MoE) layers increase model capacity without a proportional increase in per-example computation. However, conventional flat routers can yield imbalanced expert utilization and treat experts as an unstructured collection, whose indices carry no topological meaning. We introduce {\bf BRANCH-MoE}, a routing architecture that places \(E\) experts at the leaves of a binary decision tree of depth \(\log_2 E\). At each internal node the branching probability is centered on the arriva...

πŸ“„ What Did the Agent Actually Do? Evidence-Grounded Oversight for Long-Horizon Agents
πŸ—“οΈ Published: 10/5/2026
πŸ”— http://arxiv.org/abs/2610.06406v1
πŸ‘₯ Authors: Zhongxiang Sun, Jiahao Yan, Hongkang Zhao, Haojie Ding, Boheng Zhang, Fan Yang (possible past Tencent (China) affiliation), Xiao Zhang (possible past Tsinghua University affiliation), Jun Xu (possible past Google (United States) affiliation)
Abstract

As agents take on long-horizon tasks, users shift from making individual decisions to overseeing autonomous execution. Yet the volume of agent activity and the fragmentation of supporting evidence make it difficult to determine which decisions warrant user verification. We study monitors that identify consequential decisions and locate evidence to help users assess their implications. We introduce AgentMonBench, a software-engineering benchmark comprising three subsets that cover two complementa...

πŸ“„ CVIF: A Criticality-Driven Visual Intervention Framework for Geometric Diagram Understanding in MLLMs
πŸ—“οΈ Published: 10/5/2026
πŸ”— http://arxiv.org/abs/2610.06399v1
πŸ‘₯ Authors: Jiahui Kang, Bifan Wei, Lingling Zhang (possible past Google (United States) affiliation), Tianwen Jiang, Qiuyong Xiao, Jihong Zhang, Jun Liu (possible past Tencent (China) affiliation)
Abstract

Despite significant progress in visual tasks by Multimodal Large Language Models (MLLMs), geometric diagram understanding remains challenging due to the presence of sparse visual cues and ambiguous symbol-primitive associations. MLLMs may therefore rely on textual priors, producing interpretations that conflict with visual evidence. We introduce the training-free Criticality-Driven Visual Intervention Framework (CVIF), an inference-time method that localizes critical layers and executes visual i...

πŸ“„ Capability-Driven Self-Evolution of Agent Memory
πŸ—“οΈ Published: 10/5/2026
πŸ”— http://arxiv.org/abs/2610.06361v1
πŸ‘₯ Authors: Yaoqi Chen, Yuru Feng, Qianxi Zhang, Baotong Lu, Jianan Lu, Zhirui Wang, Shusen Xu, Zewen Jin, Zengzhong Li, Cheng Li (possible past Google (United States) affiliation), Qi Chen (possible past Baidu (China) affiliation)
Abstract

Memory self-evolution uses task feedback to iteratively improve executable memory programs that store and retrieve information from past interactions. Existing approaches typically adopt holistic evolution, deriving revision directions from mixed feedback and judging progress by overall performance. This can obscure optimization directions and hide capability-specific gains offset by regressions elsewhere, leaving promising directions underexplored. We introduce capability-driven evolution, whic...

πŸ“„ Anlu: Enabling In-Context Time Series Anomaly Detection in Foundation Models via Counterfactual Supervision
πŸ—“οΈ Published: 10/5/2026
πŸ”— http://arxiv.org/abs/2610.06180v1
πŸ‘₯ Authors: Tian Lan, Yifei Gao, Yimeng Lu, Xuming An, Meng Wang (possible past Google (United States) affiliation), Yue Pan, Wenjun He, Chen Zhang (possible past Peking University affiliation)
Abstract

Whether a time-series pattern is anomalous often depends on the operating regime of the monitored process. A missing event can signal a fault in one regime and be routine in another, and the query alone may not reveal which regime applies. We study in-context learning (ICL) for time series anomaly detection (TSAD) through reference-conditioned detection, where a reference record provides evidence about expected behavior and model parameters remain fixed at inference. Supplying the reference is n...

πŸ“„ VERA: Scaling Verifiable Environments for Agentic co-Evolution
πŸ—“οΈ Published: 10/5/2026
πŸ”— http://arxiv.org/abs/2610.05923v1
πŸ‘₯ Authors: Junqi Liu, Yongyang Pan, Zhuosong Jiang, Dongbai Li, Bo Zhang (possible past Tencent (China) affiliation), Xitong Ling, Sheng Wang (possible past Tencent (China) affiliation), Hanrong Ye, Yufan He (possible past Nvidia (United States) affiliation), Can Zhao, Pengfei Guo, Dong Yang (possible past Nvidia (United States) affiliation), Andriy Myronenko (possible past Nvidia (United States) affiliation), Yuyin Zhou, Tianyu Liu, Daguang Xu (possible past Nvidia (United States) affiliation), Yucheng Tang
Abstract

Competent agents need precise and verifiable environments, such as sandboxes that are resumable at any stage and evolve from observable evidence. However, most long-horizon work exposes how rare these are: for example, an agent in medical research must ground a finding, classify it, and write a report over dozens of dependent steps, yet recent environments score only the outcome. To address the challenges in stable training, we present VERA, which builds such environments at scale and lets agent...

πŸ“„ Level-of-Token Diffusion
πŸ—“οΈ Published: 10/5/2026
πŸ”— http://arxiv.org/abs/2610.05816v1
πŸ‘₯ Authors: Kiyohiro Nakayama, Brian Chao, Jan Ackermann, Hansheng Chen, Federico Tombari (possible past Google (United States) affiliation), Leonidas Guibas (possible past Stanford University affiliation), Lior Yariv, Gordon Wetzstein (possible past Stanford University affiliation)
Abstract

Image and video diffusion models allocate equal computation to every region, even when the intended scene calls for varying levels of detail. The spatial distribution of detail can often be anticipated before generation, indicating where computation can be reduced. We introduce Level-of-Token (LoT) Diffusion, a framework that turns this knowledge into an explicit multiresolution token layout (Level-of-Token layout) for adaptive and efficient generation. Tokens represent rectangular patches of va...

πŸ“„ Beyond Semantic Similarity: Performance and Costs of Agentic Retrieval for Complex Tasks
πŸ—“οΈ Published: 10/5/2026
πŸ”— http://arxiv.org/abs/2610.05750v1
πŸ‘₯ Authors: Reza Esfandiarpoor, Radek Osmulski, Yauhen Babakhin, Gabriel De Souza P. Moreira, Oliver Holworthy, Jie He, Ronay Ak (possible past Nvidia (United States) affiliation), Jiarui Cai, Ryan Chesler, Bo Liu (possible past Meta (United States) affiliation), Even Oldridge
Abstract

Modern information systems, including many agentic workflows, use dense retrieval to explore large amounts of unstructured data. However, dense retrieval relies on surface-level semantic similarity, which is insufficient for increasingly complex search applications. Here, we investigate agentic retrieval that combines the reasoning capabilities of Large Language Models (LLMs) with the efficient corpus exploration of retrievers in a ReAct agentic loop to solve complex retrieval tasks. In our expe...

πŸ“„ CIPHER-MoE: Balancing Efficiency and Routing Fidelity in Trillion-Scale MoE Training
πŸ—“οΈ Published: 10/5/2026
πŸ”— http://arxiv.org/abs/2610.05744v1
πŸ‘₯ Authors: Jing Li (possible past Tencent (China) affiliation), Jian Meng, Yingmeng Gao, Suming Qiu, Linyuan Qiu, Dongfang Li, Baotian Hu, Binfan Zheng, Rongqian Zhao, Weijian Sun, Xin Chen (possible past Tencent (China) affiliation)
Abstract

Mixture-of-Experts (MoE) has been widely adopted in recent large language model (LLM) architectures. However, scaling up MoE in LLM training introduces system-level challenges on training, where non-uniform token routing can lead to highly imbalanced workloads across experts and devices, further destabilizing the training process. With trillion-scale LLMs, imbalanced expert workloads further amplify the resource cost of MoE training, resulting in degraded training efficiency and hardware utiliza...

πŸ“„ Towards Looped Models Done Right, Part II: Rethinking at Fixed Points
πŸ—“οΈ Published: 10/5/2026
πŸ”— http://arxiv.org/abs/2610.06833v1
πŸ‘₯ Authors: Benhao Huang, Chufan Shi, Junlin Chen, Shicheng Wen, Zhengzhong Liu (possible past Tencent (China) affiliation), Eric Xing, Xuezhe Ma (possible past Carnegie Mellon University affiliation)
Abstract

Every recurrence of a looped language model adds cost in training, decoding, prefill, and reinforcement learning (RL). The closer recurrent states get to fixed points, the less the path to them matters. This enables truncated backpropagation in training; terminal key-value (KV) sharing for decoding with almost no loss in accuracy; a distilled student that prefills up to 1.79x faster; and RL updates that compute gradients from saved rollout states, 2x faster than backpropagating through the repla...

πŸ“„ Finding Gaussian Structure in Bosonic States
πŸ—“οΈ Published: 10/5/2026
πŸ”— http://arxiv.org/abs/2610.06810v1
πŸ‘₯ Authors: Alvan Arulandu, Sitan Chen (possible past University Of California, Berkeley affiliation), Ziyun Chen, Jerry Li (possible past Microsoft (United States) affiliation), Eric Ma
Abstract

We study agnostic tomography of pure bosonic Gaussian states: given copies of an arbitrary $n$-mode bosonic state $ρ$, the goal is to output a pure Gaussian state whose infidelity with $ρ$ is at most $\mathrm{opt} + Ρ$, where $\mathrm{opt}$ is the minimum infidelity achievable by any pure Gaussian state. We give efficient protocols achieving this in both the high and low fidelity regimes. When $\mathrm{opt}$ is below some universal constant, our protocol has runtime and copy complexity which i...

πŸ“„ H-JEPA: End-to-End Learning of Hierarchical World Models for Visual Planning
πŸ—“οΈ Published: 10/5/2026
πŸ”— http://arxiv.org/abs/2610.06805v1
πŸ‘₯ Authors: Wancong Zhang, Basile Terver, Michael Rabbat (possible past Meta (United States) affiliation), Yann Lecun (possible past Meta (United States) affiliation), Randall Balestriero
Abstract

Long-horizon planning with latent world models requires reasoning across timescales and levels of abstraction. Existing task-agnostic JEPA world models predict and plan at a single timescale or with multiple horizons in one shared latent space. We introduce H-JEPA, an end-to-end recipe for training a hierarchy of action-conditioned JEPAs in which each level predicts farther ahead in its own learned latent space. Planning proceeds top-down: the top level optimizes progress toward the goal, and ea...

πŸ“„ Out-of-control Hamiltonian Learning
πŸ—“οΈ Published: 10/5/2026
πŸ”— http://arxiv.org/abs/2610.06709v1
πŸ‘₯ Authors: Weiyuan Gong, Muzhou Ma, Sitan Chen (possible past University Of California, Berkeley affiliation), Jordan Cotler, Hsin-Yuan Huang (possible past Google (United States) affiliation)
Abstract

Learning the Hamiltonian of a many-body system from its dynamics is a central task in quantum science, yet the algorithms with the strongest provable guarantees assume some level of quantum control--fast, arbitrary single-qubit gates interleaved with time evolution, and measurements in arbitrary bases--that is beyond the capabilities of near-term analog quantum simulators. Motivated by analog atom- and ion-based platforms, we study Hamiltonian learning under minimal access models. Uniform stat...

πŸ“„ EMG-FM-Bench: A Comprehensive Benchmark for Foundation Model Transfer and Adaptation on Electromyography
πŸ—“οΈ Published: 10/5/2026
πŸ”— http://arxiv.org/abs/2610.06450v1
πŸ‘₯ Authors: Tianhao Wu (possible past University Of Cambridge affiliation), Xu Wu, Amirmohammad Radmehr, Jiawei Yu, Yi Wu (possible past University Of California, Berkeley affiliation), Phuc Nguyen, Jian Liu
Abstract

Foundation models (FMs) are increasingly being developed for general time series and physiological signals, yet their transferability to downstream physiological tasks remains poorly understood. This question is particularly challenging for electromyography (EMG), where signal distributions vary substantially across users, sensing configurations, acquisition hardware, and downstream tasks. We introduce EMG-FM-Bench, a systematic benchmark for studying foundation-model transfer and adaptation on ...

πŸ“„ Stability-Shaped Deep Graph Learning
πŸ—“οΈ Published: 10/5/2026
πŸ”— http://arxiv.org/abs/2610.06344v1
πŸ‘₯ Authors: Junyou Zhu, Langzhou He, Fenying Cai, Christian Nauck, Ping Xiong, Chao Gao, Philip S. Yu (possible past Tsinghua University affiliation), Klaus-Robert MΓΌller, JΓΌrgen Kurths (possible past Google (United States) affiliation), Frank Hellmann
Abstract

In deep graph neural networks, increasing depth enlarges the receptive field but often leads to over-smoothing, where node representations tend to align. We develop a unified, mode-wise stability framework for deep GNN propagation that provides a principled characterization of over-smoothing. By interpreting layer depth as time and layer updates as graph-coupled dynamics, over-smoothing can be understood as an undesirable dynamical synchronization of features, for which the master stability curv...

πŸ“„ Encoded but Not in Control: Revealing the Grounding Gap in Vision-Language Robot Policies
πŸ—“οΈ Published: 10/5/2026
πŸ”— http://arxiv.org/abs/2610.06235v1
πŸ‘₯ Authors: Shaohan Jiang, Jiahang Cao, Qiduo He, Fengting Deng, Kun Wu, Jingkai Sun, Jiaxu Wang, Qiang Zhang (possible past Tsinghua University affiliation), Qihao Zheng, Chunfeng Song, Ping Luo (possible past Shanghai Artificial Intelligence Laboratory affiliation), Andrew F. Luo
Abstract

Instruction following is central to language-conditioned robot policies: language should determine what to do when the same scene permits multiple valid actions. Yet successful execution alone cannot establish whether a policy follows the instruction or infers the task from the scene. We study this ambiguity through scene-preserving instruction interventions, using valid target substitutions, arbitrary nouns, and unrelated sentences while holding the scene fixed. We evaluate vision-language-acti...

πŸ“„ ORCA: The Annealed Spectral Conditioning Optimizer for Faster, Better LLM Training
πŸ—“οΈ Published: 10/5/2026
πŸ”— http://arxiv.org/abs/2610.06116v1
πŸ‘₯ Authors: Yuanshi Liu, Boyuan Jiang (possible past Tencent (China) affiliation), Liang Hou, Xin Tao (possible past Tencent (China) affiliation), Pengfei Wan, Zhouchen Lin (possible past Peking University affiliation), Cong Fang
Abstract

Modern LLM optimizers such as Muon often produce weight matrices with higher effective rank than Adam, yet further spectral control has delivered only modest gains. We identify a tension behind this result: concentrated spectra can suppress gradient directions in coupled weight matrices and slow optimization, while constraints maintained throughout training can limit task-specific adaptation and raise the attainable loss floor. We introduce ORCA (Orthogonal Regularization, Cooled After), a minim...

πŸ“„ MEND: RL For Flow Models via Proximal Velocity Matching
πŸ—“οΈ Published: 10/5/2026
πŸ”— http://arxiv.org/abs/2610.05954v1
πŸ‘₯ Authors: Shreshth Saini, Neil Birkbeck (possible past Google (United States) affiliation), Yilin Wang (possible past Google (United States) affiliation), Balu Adsumilli (possible past Google (United States) affiliation), Alan C. Bovik
Abstract

Reward post-training of flow models either reweights the model's own samples under a KL penalty or a frozen reference, often for thousands of updates, or backpropagates the reward and moves every sample without checking that the move is worth its size. We introduce MEND, a reinforcement learning method built on proximal velocity matching. MEND caps rewards within each prompt group, so samples that already score well receive no move. Below the cap, it proposes moves along the reward gradient and ...

πŸ“„ Mulligan: Performance-Guided Data Collection for Efficient On-Robot Learning
πŸ—“οΈ Published: 10/5/2026
πŸ”— http://arxiv.org/abs/2610.05882v1
πŸ‘₯ Authors: Lars Ankile, Perry Dong, Rohan Bhowmik, Aneesh Muppidi, David D. Yuan, Shuran Song (possible past Google (United States) affiliation), Chelsea Finn (possible past University Of California, Berkeley affiliation)
Abstract

Learning from human demonstrations is a reliable way to teach robots new tasks, but the gains from each additional demonstration shrink as the policy improves. Continued improvement can instead come from supervised deployment, where an operator places the objects and intervenes when the policy fails. We ask how to maximize improvement from a fixed budget of supervised episodes on high-precision manipulation tasks with wide ranges of object placements. We observe that failures can concentrate in ...

*Notable papers are those with at least two authors from a "big" AI/ML lab.