📄 Notable* Recent AI/ML arXiv Papers

Last updated just now...

📄 Gacha Decoding: Eliciting Diverse Generations Through Instruction Following
🗓️ Published: 10/1/2026
🔗 http://arxiv.org/abs/2610.01382v1
👥 Authors: Scott Geng, Yufei Zhang, Joseph Lee, Jerry Li (possible past Microsoft (United States) affiliation), Marjan Ghazvininejad (possible past Meta (United States) affiliation), Pang Wei Koh
Abstract

We introduce Gacha Decoding, an inference-time method for eliciting diverse language model generations that scales with model capability. Across open-ended domains (in-the-wild chat, creative writing, planning for image generation, and protein design), Gacha Decoding significantly outperforms existing generation diversity approaches at equal quality (up to 2.4x Vendi over the next-best prior approach), reaching the same number of high-quality modes with over an order of magnitude fewer samples (...

📄 ProtoFlow: Prototype-Guided Flow Matching for Multivariate Time Series Forecasting
🗓️ Published: 10/1/2026
🔗 http://arxiv.org/abs/2610.01320v1
👥 Authors: Shibo Feng, Wanjin Feng, Yang Qiu, Deheng Ye (possible past Tencent (China) affiliation), Peilin Zhao (possible past Tencent (China) affiliation), Chunyan Miao
Abstract

Generative modeling has shown strong promise for multivariate time mseries (MTS) forecasting, especially scale to high-dimensional settings. Diffusion-based methods achieve competitive performance but typically require many sampling steps at inference. VAE-based non-iterative forecasting frameworks have therefore emerged as an efficient alternative. Within this line of work, vector quantization (VQ) enables controllable latent space modeling by mapping multivariate series into compact discrete r...

📄 Federated Agent Optimization
🗓️ Published: 10/1/2026
🔗 http://arxiv.org/abs/2610.01195v1
👥 Authors: Qiang Yang, Zhiqiang Kou, Xueyi Zhang, Dong-Dong Wu, Hanlin Gu, Jing Guo (possible past Tsinghua University affiliation), Yang Liu (possible past Tsinghua University affiliation), Di Jiang, Qian Xu (possible past Baidu (China) affiliation)
Abstract

Large language model (LLM) agents increasingly operate in private environments and accumulate valuable experience from task execution, tool use, feedback, and local knowledge. Yet such experience is distributed across organizations and cannot be directly shared because of privacy and proprietary constraints. Conventional federated learning is insufficient for this setting, as agent capabilities extend beyond model parameters to memory, tools, rewards, skills, and structured knowledge. In this pa...

📄 Can AI Scientists Coordinate at Runtime?
🗓️ Published: 10/1/2026
🔗 http://arxiv.org/abs/2610.00980v1
👥 Authors: Zijian Liu, Yangzhixin Luo, Junyu Lu, Yi Li (possible past University Of Washington affiliation), Yu Chen (possible past Meta (United States) affiliation), David Xu, William F. Shen, Xinchi Qiu, Xisen Wang
Abstract

Multi-agent AI scientists have shown improving performance across a diverse range of tasks. Yet a common approach is design-time agentic orchestration, which typically relies on fixed workflows. In contrast, human scientists coordinate and adjust their division of labor at runtime. We therefore ask: can AI scientists also coordinate at runtime? To this end, we introduce Runtime Agent Coordination (RAC), which selects agents from existing AI-scientist hosts during execution, assigns scoped work c...

📄 Generalist Representation, Specialist Detection: TS-Router for Time-Series Anomaly Detection
🗓️ Published: 10/1/2026
🔗 http://arxiv.org/abs/2610.00978v1
👥 Authors: Tian Lan, Yifei Gao, Yimeng Lu, Xuming An, Meng Wang (possible past Google (United States) affiliation), Yue Pan, Wenjun He, Chenghao Liu, Chen Zhang (possible past Peking University affiliation)
Abstract

Time-series anomaly detection (TSAD) is difficult to generalize across datasets because heterogeneous temporal dynamics imply different notions of normality and favor different detection criteria. While time-series foundation models provide transferable representations, coupling them with a fixed anomaly-scoring mechanism can overlook this variation. This motivates a different perspective on foundation-model-based TSAD: using foundation models to coordinate specialized anomaly criteria rather th...

📄 VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks
🗓️ Published: 10/1/2026
🔗 http://arxiv.org/abs/2610.00972v1
👥 Authors: Caiqi Zhang, Rujun Han, Zifeng Wang, Zoey Cuizhu, Nigel Collier, Tomas Pfister (possible past University Of Oxford affiliation), Chen-Yu Lee (possible past Google (United States) affiliation)
Abstract

As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to reference answers or grading rubrics at test time. Repeated sampling yields multiple rollouts that can contain complementary correct claims, but we need a reliable verification mechanism to determine which claims to trust. We first find that disagreement often exposes correct...

📄 Match the Distribution, Not the Compute: Post-Training Multi-Token Prediction Heads
🗓️ Published: 10/1/2026
🔗 http://arxiv.org/abs/2610.00888v1
👥 Authors: Prachi Badarayani, Aidan Jay, Chenghui Zhou, Dayquan Julienne, Yuan Gao (possible past Tencent (China) affiliation), Tianwei Chen, George Zerveas, Ishmam Zabir, Xiren Zhou, Chris Quirk (possible past Microsoft (United States) affiliation), Xia Song
Abstract

Multi-token prediction (MTP) improves the throughput of autoregressive generation by enabling the language model to draft multiple next tokens per forward pass, while a verification step over draft tokens ensures that token distribution of the backbone is preserved. Every open MTP-family release (MiMo-7B, DeepSeek-V3, Qwen3) trains its heads jointly with the backbone over the full pretraining run of tens of trillions of tokens, thus setting the drafter quality at pretraining time. We ask whether...

📄 Personalized Image Generation with Reasoning and Reflection
🗓️ Published: 9/30/2026
🔗 http://arxiv.org/abs/2610.00737v1
👥 Authors: Bo Ni, Ngoc N. Tran, Qinwen Ge, Franck Dernoncourt, Seunghyun Yoon, Samyadeep Basu, Sungchul Kim, Puneet Mathur, Nedim Lipka, Tong Yu (possible past Carnegie Mellon University affiliation), Yu Wang (possible past Tsinghua University affiliation), Ryan A. Rossi, Tyler Derr
Abstract

Personalized image generation has remained narrowly focused on conditional synthesis from curated visual exemplars, rather than capturing who a user is. In practice, however, a user's personal context is much richer, comprising reviews, posts, images, captions, and metadata accumulated over time. A truly personalized generator should leverage this history to produce images aligned with the user's lifestyle and aesthetic preferences. To this end, we introduce the first unified benchmark for perso...

📄 R-GroundBench: A Diagnostic Benchmark for R-Group Groundingin Markush Molecular Editing
🗓️ Published: 9/30/2026
🔗 http://arxiv.org/abs/2610.00700v1
👥 Authors: Xin Wang (possible past University Of Edinburgh affiliation), Zichuan Ying, Xinna Lin, Junqi Zhang, Hanyi Xiong, Tianyu Gao (possible past Tsinghua University affiliation), Hairong Zhang, Qixiang Hua, Botian Shi, Zhenhailong Wang, Kaicheng Yu
Abstract

Recent advances in AI for scientific discovery enable molecular understandingand design, yet reasoning over incomplete chemical representations remainsunclear.Markush structures, which encode molecular families through variable R-groupplaceholders (\textit{R\textsubscript{1}}, \textit{R\textsubscript{2}}, \textit{X}, etc.), are ubiquitous in pharmaceutical patents and requiregrounding across molecular, textual, and chemical information.However, existing molecule-language benchmarks focus on full...

📄 EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
🗓️ Published: 9/30/2026
🔗 http://arxiv.org/abs/2610.00492v1
👥 Authors: Jiayi Geng, Zhengxuan Wu (possible past Stanford University affiliation), Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev (possible past Carnegie Mellon University affiliation), Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig (possible past Carnegie Mellon University affiliation)
Abstract

When Isaac Newton discovered the law of gravitation, he did so through an iterative process of analyzing observed data such as planetary patterns, finding the underlying mechanisms by describing patterns in mathematical equations, and refining his theory against the Moon's orbit, revealing the startling insight that the same force governs both falling apples and orbiting planets. Would it be possible for AI agents to make similar discoveries? To measure this ability, we introduce EurekaBench, a ...

📄 Turbo Harness: Instance-Adaptive Harness Optimization
🗓️ Published: 9/30/2026
🔗 http://arxiv.org/abs/2609.40330v1
👥 Authors: Tunyu Zhang, Hao Wang (possible past Tsinghua University affiliation), Kai Xu (possible past National University Of Defense Technology affiliation), Dimitris N. Metaxas
Abstract

Automating the search for effective harnesses is an important step toward enabling agents to recursively self-improve. Existing harness optimizations typically produce a single global harness that is applied uniformly across task instances. However, a harness that works well on average may not be optimal for every instance. We introduce Turbo Harness, a framework that can adapt a globally optimized harness to each instance by reusing information generated during the original optimization process...

📄 WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents
🗓️ Published: 9/30/2026
🔗 http://arxiv.org/abs/2609.40325v1
👥 Authors: Ziyan Jiang, Jingbo Yang, Jiabao Ji, Yujian Liu, Qiucheng Wu, Tommi Jaakkola, Yang Zhang (possible past Tsinghua University affiliation), Shiyu Chang (possible past Tencent (China) affiliation)
Abstract

As interactive 3D worlds are increasingly used to study intelligent behavior, it becomes important to develop efficient pipelines for identifying anomalies in these simulated environments, such as floating objects, traversable walls, or objects inconsistent with the surrounding scene. Multimodal AI systems, including vision-language models (VLMs) and vision-language-action models (VLAs), have shown potential for automating this task. However, 3D world auditing is complex, requiring the close cou...

📄 Cogentic: Multi-Agent Orchestration for Automated Proof Discovery
🗓️ Published: 9/30/2026
🔗 http://arxiv.org/abs/2609.40324v1
👥 Authors: Yang Cai, Vineet Gupta (possible past Google (United States) affiliation), Yanchen Jiang, Christopher Liaw, Aranyak Mehta (possible past Google (United States) affiliation), Grigoris Velegkas, Di Wang
Abstract

We present Cogentic, a multi-agent harness for automated proof discovery on open research problems. While frontier language models can generate strong mathematical ideas in a single shot, single-shot generation is often insufficient for open problems that require exploring multiple competing conjectures, overcoming subtle technical obstructions, and retaining intermediate progress over a long horizon. Cogentic addresses these challenges through an iterative prove--verify loop in which an orchest...

📄 PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents
🗓️ Published: 9/30/2026
🔗 http://arxiv.org/abs/2609.40285v1
👥 Authors: Yinghui He, Yapei Chang, Khushi Bhardwaj, Daniele Molinari, Tugrul Konuk, Jan Kautz (possible past Nvidia (United States) affiliation), Ali Hatamizadeh (possible past Nvidia (United States) affiliation)
Abstract

On-policy distillation (OPD) is a promising approach for training language agents, providing dense teacher supervision on student-generated trajectories. However, in multi-turn interaction, an incorrect action changes the states the student encounters later, so errors compound across turns. In preliminary experiments across three Qwen3 models (8B to 235B), we find that more than half of the failed rollouts contain a pivotal mistake, an action that moves the agent farther from completing the task...

📄 cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents
🗓️ Published: 9/30/2026
🔗 http://arxiv.org/abs/2609.40284v1
👥 Authors: Pranjal Aggarwal, Lawrence Keunho Jang, Sean Welleck, Daniel Fried, Ruslan Salakhutdinov (possible past University Of Toronto affiliation), Jing Yu Koh (possible past Google (United States) affiliation)
Abstract

Computer use agents (CUAs), which use graphical user interfaces (GUIs) to complete tasks on a computer, have recently surpassed human performance on many standard benchmarks, including difficult long-horizon tasks. Their capabilities are undoubtedly impressive, however, a key barrier to the widespread adoption and deployment of CUAs remains their speed and cost. Progress towards faster yet capable CUAs requires reliable evaluation of their speed, but many CUA benchmarks currently face a reproduc...

📄 Belief-Aware Multi-Agent Path Finding under Map Uncertainty
🗓️ Published: 9/30/2026
🔗 http://arxiv.org/abs/2609.40269v1
👥 Authors: Viraj Parimi, Shao-Hung Chan, Han Zhang (possible past Tsinghua University affiliation), Jingkai Chen, Brian Williams (possible past Google (United States) affiliation)
Abstract

Multi-Agent Path Finding (MAPF) aims to find collision-free paths for multiple agents in a shared environment. Classical MAPF assumes that all static obstacles are known in advance, but real-world environments can change unexpectedly due to fallen objects, spills, or other local disturbances. When such changes are spatially correlated, an observation can inform traversability estimates beyond the observed location. Prior approaches address uncertainty in traversability through contingent plans o...

📄 Learning Skills from Historical Action Trajectories: Action Experience Dictionary for World Action Models
🗓️ Published: 9/30/2026
🔗 http://arxiv.org/abs/2609.40219v1
👥 Authors: Qi Lyu, Jiahua Dong, Hao Shen, Xudong Wang, Hongyuan Yu, Baichen Liu, Henghui Ding, Zhi Han, Nicu Sebe, Ivan Laptev, Fahad Shahbaz Khan (possible past Inception Institute Of Artificial Intelligence affiliation), Salman Khan (possible past Inception Institute Of Artificial Intelligence affiliation)
Abstract

World Action Models (WAMs) couple visual dynamics prediction with action generation, yet they do not explicitly support the reuse of action experience across manipulation tasks. Furthermore, existing WAMs struggle to capture underlying cross-task semantic relationships that could guide target action prediction, as redundant background elements interfere with the extraction of key visual information. To address these challenges, we develop a novel Action Experience Dictionary (AED) that encodes h...

📄 MemLife: Curating and Reasoning over Long-Term Egocentric Video Memories
🗓️ Published: 9/30/2026
🔗 http://arxiv.org/abs/2609.40195v1
👥 Authors: Guangzhi Xiong, Xinyuan Zhang, Xiao Yang (possible past Tencent (China) affiliation), Hyokun Yun, Kai Zhang, Shiun-Zu Kuo, Hyeonjeong Ha, Xilun Chen, Kai Sun, Lucas Liang, Guangqiang Dong, Ejaz Ahmed, Ahmed A Aly, Anuj Kumar (possible past Meta (United States) affiliation), Raffay Hamid, Aidong Zhang, Xin Luna Dong (possible past University Of Washington affiliation)
Abstract

Long-term egocentric video enables personalized AI assistants to reason about daily life. However, as video histories grow to hundreds of hours spanning months or years, reprocessing raw clips for every query becomes computationally prohibitive. Memory systems offer a scalable alternative by compacting videos into text representations, but often fail on practical benchmarks: either the memory does not preserve key evidence, or the retriever fails to locate relevant entries due to retrieval compe...

📄 Learning from Research: Toward Lifelong Agent Harness Evolution
🗓️ Published: 9/30/2026
🔗 http://arxiv.org/abs/2609.40169v1
👥 Authors: Jingbo Yang, Kwei-Herng Lai, Xiaowen Wang, Yaar Harari, Evgeniy Gabrilovich (possible past Google (United States) affiliation), Shiyu Chang (possible past Tencent (China) affiliation)
Abstract

Language agents are expected to solve increasingly complex tasks, creating a growing need for continual improvement. One promising approach is to evolve the agent harness, the software that governs tool use, memory management, and task execution, while keeping the underlying language model fixed. Recent methods automate this process by using a meta coding agent to modify the harness based on execution feedback. However, relying on that agent's existing knowledge and observed failures can restric...

📄 Game-Guided Skill Discovery through Self-Play for Playable Agent Control
🗓️ Published: 9/30/2026
🔗 http://arxiv.org/abs/2609.40137v1
👥 Authors: Seungeun Rho, Jeonghwan Kim, Xue Bin Peng (possible past University Of California, Berkeley affiliation), Sehoon Ha (possible past Google (United States) affiliation)
Abstract

We present Game-Guided Skill Discovery (GGSD), a framework that uses self-play in games to discover motor skills that are directly playable by humans. Playable skills provide a compact abstraction for controlling embodied agents through a small set of learned behaviors rather than low-level actions. To be effective, these skills should be semantically distinct, interpretable, and expressive; properties that existing unsupervised skill-discovery methods often fail to achieve simultaneously. GGSD ...

📄 Tactile Curiosity Drives Robot Interaction
🗓️ Published: 9/30/2026
🔗 http://arxiv.org/abs/2609.40134v1
👥 Authors: Klemens Iten, Alexander Proshkin, Bhavya Sukhija, Stelian Coros, Andreas Krause (possible past Eth Zurich affiliation), Pieter Abbeel (possible past University Of California, Berkeley affiliation), Carmelo Sferrazza
Abstract

Mastering robot manipulation skills via reinforcement learning (RL) remains largely sample-inefficient. The most common RL algorithms rely on random action sampling to discover new strategies, resulting in agents that allocate most of their training budget to motions in free space, away from the contacts from which manipulation skills emerge. Existing intrinsic motivation methods based on model disagreement or epistemic uncertainty improve on isotropic noise, but they can also reward uncertainty...

📄 HADRec: A Hierarchy-Aware Drug Recommendation Framework by Fusing Molecular Knowledge and Electronic Health Record
🗓️ Published: 10/1/2026
🔗 http://arxiv.org/abs/2610.00984v1
👥 Authors: Junke Wang, Hongshun Ling, Li Zhang (possible past University Of Oxford affiliation), Jinjing Wu, Tong Shao, Fang Wang (possible past Tencent (China) affiliation), Yuan Gao (possible past Tencent (China) affiliation)
Abstract

Accurate medication recommendation is central to clinical decision-making, directly determining therapeutic efficacy and patient safety. However, existing methods suffer from two key limitations: drugs are often abstracted as discrete tokens, ignoring their molecular structures and pharmacological mechanisms, and the commonly used "flat" recommendation paradigm fails to leverage the hierarchical logic of the internationally standardized Anatomical Therapeutic Chemical (ATC) classification system...

📄 Clock Diffusion: Efficient Semi-Autoregressive Continuous Diffusion Language Models
🗓️ Published: 10/1/2026
🔗 http://arxiv.org/abs/2610.00894v1
👥 Authors: Yair Schiff, Omer Belhasin, Roy Uziel, Matan Rusanovsky, Ran Zilberstein, Marianne Arriola, Gilad Turok, Guanghan Wang, Volodymyr Kuleshov (possible past Stanford University affiliation), Michael Elad (possible past Technion – Israel Institute Of Technology affiliation)
Abstract

Recent works on continuous diffusion for discrete data have demonstrated performance on par with comparable discrete diffusion models. However, these continuous counterparts lack key features that are essential to practical use as language models, namely variable-length generation and support for a key-value cache, and they still lag behind the frontier of autoregressive and discrete diffusion quality. In this work, we address these limitations. We do so by introducing a model parameterization t...

📄 Denoising Surface: Modeling and Predicting Inference Cost for Diffusion LLM Serving
🗓️ Published: 9/30/2026
🔗 http://arxiv.org/abs/2610.00499v1
👥 Authors: Haoyu Zheng, Fangcheng Fu (possible past Peking University affiliation), Binhang Yuan, Yongqiang Zhang, Liang Deng, Hao Wang (possible past Tsinghua University affiliation), Yuanyuan Zhu, Xiao Yan, Jiawei Jiang (possible past Tencent (China) affiliation)
Abstract

As diffusion large language models (dLLMs) become more capable, they are moving from research settings to real-world \textit{serving}, where request management (such as scheduling and resource allocation) relies on accurate estimation of per-request inference cost. However, common cost proxies fall short for dLLMs: output length ignores that one forward pass can unmask multiple tokens, and denoising-step count ignores the \textit{heterogeneous} per-step costs. We observe that the block-autoregre...

📄 Looped Diffusion Transformer
🗓️ Published: 9/30/2026
🔗 http://arxiv.org/abs/2609.40305v1
👥 Authors: Yong Xien Chng, Tianyi Chen, Wenwen Tong, Haiwen Diao, Zhongang Cai, Lei Yang (possible past Google (United States) affiliation), Ziwei Liu, Lewei Lu, Dahua Lin, Gao Huang (possible past Tsinghua University affiliation)
Abstract

Improving text-to-image models has traditionally relied on increasing model size or the number of denoising steps. In this work, we explore an alternative way to scale computation by repeatedly running shared Transformer blocks within each denoising step, effectively increasing computational depth while keeping the parameter count fixed. This looped computation enables iterative refinement of internal representations without explicit reasoning tokens. However, naive looping fails to consistently...

📄 How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
🗓️ Published: 9/30/2026
🔗 http://arxiv.org/abs/2609.40295v1
👥 Authors: Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting (possible past Google (United States) affiliation), Mohit Iyyer (possible past Google (United States) affiliation), Max Spero, Bradley Emi
Abstract

Web text makes up the majority of pretraining data and is increasingly AI-generated. After applying FineWeb quality filtering, we find that 27.5% of tokens from June 2026 web data are labeled as AI-generated by Pangram, rising to 31.1% by August. Unlike synthetic data or model-collapse setups, this *wild* AI text comes from many models, is written for human readers, and arrives unlabeled in pretraining corpora. How does AI text in the wild affect language model pretraining? To answer this questi...

📄 Every Batch Is Its Own Validation Set: Leave-One-Out Gradient Matching for Online Data Selection in LLM Fine-Tuning
🗓️ Published: 9/30/2026
🔗 http://arxiv.org/abs/2610.00436v1
👥 Authors: Hongyu Chen, Xinyi Luo, Ming Zhao (possible past Tencent (China) affiliation), Lin Tang, Zihan Xu, Jing Li (possible past Tencent (China) affiliation), Yuxuan Wang (possible past Google (United States) affiliation), Haoran Deng, Wei Zhang (possible past Tsinghua University affiliation)
Abstract

Online batch selection fine-tunes a language model on the most useful part of each candidate batch. Selectors that match the gradient of the candidate batch are attractive because they need no held-out data, yet they rarely beat training on the whole batch. We show why. In-sample gradient matching uses every example as part of its own target, so its objective credits each example with its own gradient noise. This is the covariance penalty that makes training error optimistic, now sitting on the ...

*Notable papers are those with at least two authors from a "big" AI/ML lab.