πŸ“„ Notable* Recent AI/ML arXiv Papers

Last updated just now...

πŸ“„ ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research
πŸ—“οΈ Published: 10/1/2026
πŸ”— http://arxiv.org/abs/2610.02202v1
πŸ‘₯ Authors: Sohyeon Kim, Yoonho Lee, Bo Liu (possible past Meta (United States) affiliation), Dayoon Ko, Rulin Shao, Seungone Kim, Graham Neubig (possible past Carnegie Mellon University affiliation), Pang Wei Koh, Aakanksha Chowdhery (possible past Stanford University affiliation), Akari Asai (possible past Tencent (China) affiliation), Omar Khattab, Yejin Choi (possible past Allen Institute For Artificial Intelligence affiliation), Gunhee Kim, Chelsea Finn (possible past University Of California, Berkeley affiliation)
Abstract

What makes great scientists great? Even as AI systems start to make progress on open problems, scientists remain far ahead of them at sensing which prior idea, buried in an ever-growing archive of research, a new problem needs. To study this skill, we draw on researchers who know firsthand which earlier work advanced their completed projects, with papers serving as pointers to the ideas within. Using our automated pipeline that makes author annotation scalable, we build ScholarCatalyst by having...

πŸ“„ DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation
πŸ—“οΈ Published: 10/1/2026
πŸ”— http://arxiv.org/abs/2610.02188v1
πŸ‘₯ Authors: Zhengming Yu, Junkun Yuan, Haotian Yang, Gordon Guocheng Qian, Yizhi Wang, Angtian Wang, Yiding Yang, Bo Liu (possible past Meta (United States) affiliation), Xin Li (possible past Google (United States) affiliation), Wenping Wang, Chongyang Ma
Abstract

Distribution Matching Distillation (DMD) trains a few-step student from the difference between separately estimated target and student scores, so it must keep an auxiliary diffusion model fitted to the student's evolving distribution at extra memory and computation cost. We introduce DMAD, Distribution Matching as Adversarial Distillation, which recasts distribution matching as classification and learns the required log-density ratios directly. Two discriminator heads on a shared backbone distin...

πŸ“„ Higher-Order Molecular Grammars for Generative and Foundation Models in Chemistry
πŸ—“οΈ Published: 10/1/2026
πŸ”— http://arxiv.org/abs/2610.02186v1
πŸ‘₯ Authors: Yiming Huang, Yujie Zeng, Vijay Prakash Dwivedi, Simone Foti, Jianmin Wang (possible past Tsinghua University affiliation), Jure Leskovec (possible past Stanford University affiliation), Tolga Birdal (possible past Stanford University affiliation)
Abstract

Molecular learning models are strongly shaped by their underlying representations. Yet standard sequential and graph formalisms struggle to explicitly encode higher-order topology, such as ring systems and recurring motifs. Existing higher-order representations can capture these structures directly, but they are often computationally demanding and difficult to decode into valid molecules. Here, we introduce Higher-order Grammar Representation (HGR), a principled, topology-aware framework that li...

πŸ“„ Finetuning with Sampling: SFT Learns Better Than You Think
πŸ—“οΈ Published: 10/1/2026
πŸ”— http://arxiv.org/abs/2610.02140v1
πŸ‘₯ Authors: Aayush Karan, Sitan Chen (possible past University Of California, Berkeley affiliation), Yilun Du (possible past Massachusetts Institute Of Technology affiliation)
Abstract

Introducing new capabilities to frontier models has long been the goal of posttraining, which predominantly employs supervised finetuning (SFT) and reinforcement learning (RL) to this end. Conventional wisdom dictates that RL enables strong generalization on new tasks without losing existing capabilities, while SFT is prone to weak generalization and catastrophic forgetting. At the same time, SFT can learn from off-policy expert data, whereas RL must rely on a model's ability to find successful ...

πŸ“„ Sharpening Tax in Post-Training
πŸ—“οΈ Published: 10/1/2026
πŸ”— http://arxiv.org/abs/2610.01509v1
πŸ‘₯ Authors: Changdae Oh, Qi Zeng, Qi Qi, Andrey Zhmoginov (possible past Google (United States) affiliation), Deren Lei, Yun He, Hoang Phan, Hangoo Kang, Azalia Mirhoseini (possible past Google (United States) affiliation), Sharon Li
Abstract

An emerging hypothesis about reinforcement learning (RL) post-training of large language models (LLMs) is that it merely sharpens existing behaviors of a base model, improving single-shot accuracy at the cost of solution coverage. Although this trade-off has been observed in math and coding tasks, it need not extend to agentic tasks, where multi-turn tool use and interaction may require capabilities newly acquired during post-training. Our surprising finding is that pre-trained LLMs, equipped wi...

πŸ“„ MWOP: Modality-aware Width-wise Operation Pruning for Efficient MLLMs
πŸ—“οΈ Published: 10/1/2026
πŸ”— http://arxiv.org/abs/2610.01434v1
πŸ‘₯ Authors: Xudong Wang, Hao Wu (possible past Tencent (China) affiliation), Haozhe Hu, Peiran Yin, Xinghao Chen, Yunpu Ma, Wei Zhang (possible past Tsinghua University affiliation), Xiaoyu Shen
Abstract

Multimodal large language models (MLLMs) incur substantial inference costs when processing long visual-textual sequences. While existing operation compression methods exploit modality-level redundancy, they largely treat computation within attention heads and shared feed-forward network (FFN) channels as unified units, leaving finer-grained redundancy underexplored. We find that redundancy varies both across modality-interaction paths within the same attention head and across visual and textual ...

πŸ“„ Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States
πŸ—“οΈ Published: 10/1/2026
πŸ”— http://arxiv.org/abs/2610.01415v1
πŸ‘₯ Authors: Yu Luo, Jiamin Jiang, Yimin Zuo, Xidao Wen, Rongchen Gao, Yongqian Sun (possible past Tsinghua University affiliation), Shenglin Zhang, Guiyang Liu, Cheng Zhang, Fang Situ, Qi Zhou, Dan Pei (possible past Tsinghua University affiliation)
Abstract

Large language model (LLM) agents can now undertake increasingly complex tasks, but the way they organize interaction history into memory does not ensure a coherent understanding of the current world. We introduce PoS, an inference-time framework that constructs and continually maintains explicit belief states as the agent's decision context. Each belief combines an estimate of the current world state with unresolved task requirements, making explicit what the agent still needs to learn and acco...

πŸ“„ Supervising Sound Localization by In-the-wild Egomotion
πŸ—“οΈ Published: 10/1/2026
πŸ”— http://arxiv.org/abs/2610.01388v1
πŸ‘₯ Authors: Anna Min, Ziyang Chen, Hang Zhao (possible past Nvidia (United States) affiliation), Andrew Owens (possible past Massachusetts Institute Of Technology affiliation)
Abstract

We present a method for learning binaural sound localization using egomotion as a supervisory signal. Over the course of a video, the cameras direction to a sound source will change as the camera moves. We train an audio model to predict sound directions that are consistent with visual estimates of camera motion, which we obtain using traditional methods from multi-view geometry. This provides a weak but plentiful form of supervision that we combine with traditional binaural cues. To evaluate th...

πŸ“„ Gacha Decoding: Eliciting Diverse Generations Through Instruction Following
πŸ—“οΈ Published: 10/1/2026
πŸ”— http://arxiv.org/abs/2610.01382v1
πŸ‘₯ Authors: Scott Geng, Yufei Zhang, Joseph Lee, Jerry Li (possible past Microsoft (United States) affiliation), Marjan Ghazvininejad (possible past Meta (United States) affiliation), Pang Wei Koh
Abstract

We introduce Gacha Decoding, an inference-time method for eliciting diverse language model generations that scales with model capability. Across open-ended domains (in-the-wild chat, creative writing, planning for image generation, and protein design), Gacha Decoding significantly outperforms existing generation diversity approaches at equal quality (up to 2.4x Vendi over the next-best prior approach), reaching the same number of high-quality modes with over an order of magnitude fewer samples (...

πŸ“„ ProtoFlow: Prototype-Guided Flow Matching for Multivariate Time Series Forecasting
πŸ—“οΈ Published: 10/1/2026
πŸ”— http://arxiv.org/abs/2610.01320v1
πŸ‘₯ Authors: Shibo Feng, Wanjin Feng, Yang Qiu, Deheng Ye (possible past Tencent (China) affiliation), Peilin Zhao (possible past Tencent (China) affiliation), Chunyan Miao
Abstract

Generative modeling has shown strong promise for multivariate time mseries (MTS) forecasting, especially scale to high-dimensional settings. Diffusion-based methods achieve competitive performance but typically require many sampling steps at inference. VAE-based non-iterative forecasting frameworks have therefore emerged as an efficient alternative. Within this line of work, vector quantization (VQ) enables controllable latent space modeling by mapping multivariate series into compact discrete r...

πŸ“„ Federated Agent Optimization
πŸ—“οΈ Published: 10/1/2026
πŸ”— http://arxiv.org/abs/2610.01195v1
πŸ‘₯ Authors: Qiang Yang, Zhiqiang Kou, Xueyi Zhang, Dong-Dong Wu, Hanlin Gu, Jing Guo (possible past Tsinghua University affiliation), Yang Liu (possible past Tsinghua University affiliation), Di Jiang, Qian Xu (possible past Baidu (China) affiliation)
Abstract

Large language model (LLM) agents increasingly operate in private environments and accumulate valuable experience from task execution, tool use, feedback, and local knowledge. Yet such experience is distributed across organizations and cannot be directly shared because of privacy and proprietary constraints. Conventional federated learning is insufficient for this setting, as agent capabilities extend beyond model parameters to memory, tools, rewards, skills, and structured knowledge. In this pa...

πŸ“„ The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models
πŸ—“οΈ Published: 10/1/2026
πŸ”— http://arxiv.org/abs/2610.02191v1
πŸ‘₯ Authors: Shuo Xing, Zilin Dai, Chengyuan Qian, Fangzhou Lin, Wenjing Chen, Ping He, Pan Lu (possible past Baidu (China) affiliation), Alvaro Velasquez, Mohit Bansal, Zhengzhong Tu (possible past Google (United States) affiliation)
Abstract

While Large Language Models (LLMs) have demonstrated striking capabilities on frontier mathematical problems, it remains unclear whether they possess the structural mathematical understanding underlying their solutions. In this paper, we take a first step toward systematically studying mathematical understanding in LLMs, from diagnosing its distinct capabilities to leveraging these findings to improve post-training. First, we introduce the notion of Mathematical Primitive to probe structural mat...

πŸ“„ Invent a Dataset: Measuring dataset generation abilities with zero seed
πŸ—“οΈ Published: 10/1/2026
πŸ”— http://arxiv.org/abs/2610.01674v1
πŸ‘₯ Authors: Shivalika Singh, Andrija Djurisic, Gbemileke Onilude, Sudip Roy (possible past Google (United States) affiliation), Sara Hooker (possible past Google (United States) affiliation)
Abstract

Building datasets remains one of the most manual and brittle parts of AI development. In this technical report, we focus on the most extreme but also most prevalent setting real world practitioners face: a zero data regime. Here, practitioners don't have any data for the capability they want to learn. We introduce Invent-A-Dataset which is a prompt based system to go from dataset description to realistic and large scale post-training datasets. We evaluate Invent-A-Dataset against five frontier m...

πŸ“„ Convergence Analysis of STORM Under Different Geometries
πŸ—“οΈ Published: 10/1/2026
πŸ”— http://arxiv.org/abs/2610.01599v1
πŸ‘₯ Authors: Wei Jiang (possible past Apple (United States) affiliation), Yibo Wang, Wenhao Yang, Rui Yan (possible past Peking University affiliation), Lijun Zhang, Zechao Li
Abstract

Stochastic recursive momentum (STORM) achieves fast convergence for nonconvex optimization via the variance reduction effect, but existing analyses rely on the strong average smoothness assumption. In this paper, we study the convergence of STORM for different objectives without average smoothness. We first revisit the results under average smoothness, obtaining the $O(T^{-1/3})$ bound for nonconvex objectives and the $O(Οƒ^2/(ΞΌT))$ bound for last-iterate output under the $ΞΌ$-Polyak--Łojasiewicz~...

πŸ“„ Optimal Momentum Methods for Stochastic Multilevel Compositional Optimization
πŸ—“οΈ Published: 10/1/2026
πŸ”— http://arxiv.org/abs/2610.01572v1
πŸ‘₯ Authors: Wei Jiang (possible past Apple (United States) affiliation), Rui Yan (possible past Peking University affiliation), Sifan Yang, Yuanyu Wan, Lijun Zhang, Zechao Li
Abstract

This paper investigates stochastic multi-level optimization where the objective is a nested composition of several smooth non-convex functions. We assume that only stochastic estimates of the gradient and function values for each level are accessible. Consequently, obtaining an accurate estimate of the overall gradient is challenging due to the nested structure. To address this, we employ a momentum-based estimator with mini-batches to track the function values of each level, which are subsequen...

*Notable papers are those with at least two authors from a "big" AI/ML lab.