📄 Notable* Recent AI/ML arXiv Papers

Last updated just now...

📄 VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning
🗓️ Published: 10/6/2026
🔗 http://arxiv.org/abs/2610.08761v1
👥 Authors: Zewei Zhou, Rachel Luo, Yulong Cao, Chaowei Xiao, Chensheng Peng, Boyi Li, Thomas Tian, Zheng Lian, Yan Wang (possible past Tencent (China) affiliation), Jiaqi Ma, Boris Ivanovic, Marco Pavone (possible past Stanford University affiliation), Wenhao Ding
Abstract

Self-improving policies continually expose new failure patterns, changing what their judges must be able to verify. However, current fixed judges constrain both optimization feedback and the discovery of useful training examples, limiting further self-improvement. This challenge is even more acute in embodied reasoning, where reliable evaluation must account for spatial grounding, causal reasoning, and safety-aware decision-making. We introduce VeriFine, an agent harness framework that scales ve...

📄 Agentic RCA for Internet-Scale Services Using Constrained Creativity
🗓️ Published: 10/6/2026
🔗 http://arxiv.org/abs/2610.08622v1
👥 Authors: Sayan Sinha, Vipul Harsh, B. Aditya Prakash (possible past Carnegie Mellon University affiliation), Vyas Sekar (possible past Carnegie Mellon University affiliation), Hui Zhang
Abstract

System administrators of Internet-scale services need to resolve failure incidents to maintain reliability of such services. Ideally, we want a troubleshooting system to be: (1) expressive to known and unknown incidents with high accuracy; (2) cost efficient at scale; (3) explainable to provide actionable insights operators can act on; and (4) entail low effort from the operators. Unfortunately, most existing systems, including emerging LLM-assisted agentic workflows and structured frameworks fo...

📄 VOMMI: Collecting and Leveraging Portable Demonstrations for Mobile Manipulation
🗓️ Published: 10/6/2026
🔗 http://arxiv.org/abs/2610.08220v1
👥 Authors: Yutian Zhang, Xingrui Xiong, Siyuan Ma, Yang Li (possible past Google (United States) affiliation), Jiawen Wen, Jiaqi Zhai, Liwen Yang, Ce Hao, Haozhen Chi, Yangkun Zhu, Yifan Zhu, Xiaowen Chu, Dong Wei (possible past Tencent (China) affiliation), Qiaojun Yu, Dibo Hou
Abstract

Portable mobile-manipulation demonstrations can help alleviate data scarcity for embodied intelligence, but obtaining reliable, low-cost, and robot-free motion supervision from RGB observations remains challenging. Existing approaches often rely on teleoperation or specialized devices equipped with additional sensing hardware, while directly using estimated visual odometry (VO) trajectories can introduce inconsistencies due to accumulated drift and imperfect motion supervision. We present the Vi...

📄 Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
🗓️ Published: 10/6/2026
🔗 http://arxiv.org/abs/2610.08077v1
👥 Authors: Haoxiang Zhang, Qinglin Chen, Hiroaki Hayashi, Zhuofeng Li, Siming Zhang, Jiaxin Zhang, Jixuan Chen, Fang Wu, Pan Lu (possible past Baidu (China) affiliation), Silvio Savarese (possible past Stanford University affiliation), Julian Mcauley, Chien-Sheng Wu (possible past Salesforce (United States) affiliation)
Abstract

Reinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction. For group-relative objectives, however, this signal vanishes when all rollouts receive the same reward, even though their trajectories may reveal useful information about what the task requires and how the agent fails. We ask a complementary question: can hindsight teach an agent what it could have anticipated before acting? We introduce p...

📄 VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs
🗓️ Published: 10/6/2026
🔗 http://arxiv.org/abs/2610.07987v1
👥 Authors: Yuan Feng, Qize Yang, Ruizhe Chen, Sibo Song, Haolin He, Muzhi Zhu, Zihan Liu, Yunfei Chu, Xize Cheng, Yuxuan Wang (possible past Google (United States) affiliation), Jin Xu (possible past Tencent (China) affiliation), Xike Xie
Abstract

Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed-size patch tokens. However, visual information is unevenly distributed: some regions require fine-grained detail, while others admit compact representations. Downsampling sacrifices this detail, while existing token pruning and adaptive approaches remain limited in content-adaptive granularity, task generalization, and integration with mode...

📄 SIGMA: Self-Improving Alignment Generalization from a Model Spec
🗓️ Published: 10/6/2026
🔗 http://arxiv.org/abs/2610.07935v1
👥 Authors: Jingyu Zhang, Shruti Palaskar (possible past Carnegie Mellon University affiliation), Daniel Khashabi, Benjamin Van Durme (possible past Google (United States) affiliation), Leon A. Gatys, Joseph Yitan Cheng
Abstract

LLM agents are increasingly capable of executing complex tasks and of recursively improving themselves on easy-to-verify objectives such as software engineering and mathematics. Since alignment is much harder to verify, this creates a growing risk of capabilities increasing without appropriate safety alignment, especially as capabilities expand to auto-research and cybersecurity. Existing approaches focus on capability self-improvement using verifiable feedback or on alignment training with supe...

📄 Visual Abstention in Unified Multimodal Models
🗓️ Published: 10/6/2026
🔗 http://arxiv.org/abs/2610.07887v1
👥 Authors: Chufan Shi, Cheng Yang (possible past Tsinghua University affiliation), Tiannuo Yang, Isadora White, Yiwei Chen, Taylor Berg-Kirkpatrick, Xuezhe Ma (possible past Carnegie Mellon University affiliation)
Abstract

Unified multimodal models (UMMs) integrate understanding and generation, yet their generative behavior is rarely governed by what they understand about the task. We formalize visual abstention: when a requested visual transformation is impossible under the task's rules, the model should recognize that no valid solution exists, state this, and decline to generate. We introduce Draw-or-Decline (DoD), a benchmark of 1,050 feasible-infeasible request pairs across 7 task categories that jointly measu...

📄 A self-learning scientific agent for X-ray diffraction
🗓️ Published: 10/6/2026
🔗 http://arxiv.org/abs/2610.07862v1
👥 Authors: Bin Cao (possible past Microsoft (United States) affiliation), Huichi Zhou, Runyu Yang, Jingsong Li, Shuchen Sun, Yan Song (possible past Tencent (China) affiliation), Hanyu Gao, Zhongwei Yu, Tong-Yi Zhang, Jun Wang (possible past Tencent (China) affiliation)
Abstract

A central challenge for scientific agents is to turn analytical experience into reusable expertise grounded in physical evidence. Here we introduce Gan Jiang, a self-learning agent for powder X-ray diffraction built on a diffraction-analysis ecosystem we developed: XMatcher, XQueryer, XDecomposer and WPEM. Together, these engines span phase identification, multiphase decomposition and physics-constrained whole-pattern modelling. Gan Jiang converts analytical experience into executable skills by ...

📄 The Geometry of Empowerment
🗓️ Published: 10/6/2026
🔗 http://arxiv.org/abs/2610.07796v1
👥 Authors: Catherine Ji, Vivek Myers, Sergey Levine (possible past University Of Washington affiliation), Benjamin Eysenbach (possible past Carnegie Mellon University affiliation)
Abstract

Empowerment captures the capacity for an agent to actively control its environment. While conceptually appealing as an information-theoretic quantity, the connection between empowerment and structurally central states that provide broad access to future outcomes has remained an open question. In this work, we link empowerment maximization and skill-learning methods to provide new geometries for interpreting and analyzing empowerment. Our analyses answer longstanding open questions on the connect...

📄 OTel: Open Telco AI Datasets, Benchmarks, and Models
🗓️ Published: 10/6/2026
🔗 http://arxiv.org/abs/2610.07766v1
👥 Authors: Farbod Tavakkoli, Gregory Diamos (possible past Baidu (China) affiliation), Kenneth Church (possible past Ibm (United States) affiliation), David Kanter, Mark Austin, Imtiaz Karim, Mirza Masfiqur Rahman, Merouane Abdelkader Debbah, Zeinab Nezami, Ali Maatouk, Leandros Tassiulas, Rex Ying (possible past Stanford University affiliation), Nick Sorros, Louis Powell, Nikolaos Vasiloglou, Ashish Vaswani (possible past Google (United States) affiliation), Somanshu Singla, Adarsh Chaluvaraju
Abstract

We present Open Telco (OTel), an open telecom AI resource that releases derived telecom datasets for retrieval, reranking, instruction tuning, and safety/abstention, together with 30 full-parameter post-trained baselines spanning 10 embedding models, 3 rerankers, and 17 language models. The community has already engaged substantially with the resource: as of May 3, 2026, the released models have been downloaded over 16 million times and the project has received 157+ pieces of media coverage worl...

📄 Disentangling Dual Image References in Frequency Aware Diffusion Models for Personalized Generation
🗓️ Published: 10/6/2026
🔗 http://arxiv.org/abs/2610.07684v1
👥 Authors: Haipeng Liu, Yang Wang (possible past Baidu (China) affiliation), Meng Wang (possible past Google (United States) affiliation)
Abstract

Personalized image generation aims to synthesize text-driven images conditioned on reference images, while mainly casting the generation as image customization for foreground and style transfer for background. Previous arts of diffusion models suffers from the text misalignment with background for image customization and foreground for style transfer during the denoising process. Such facts, as we observed, rooted from the entanglement among hybrid frequency bands during the denoising process. T...

📄 EigenDEXplore: Structured Exploration for Dexterous Manipulation with Human Priors
🗓️ Published: 10/6/2026
🔗 http://arxiv.org/abs/2610.07681v1
👥 Authors: Harsh Gupta, Tyler Ga Wei Lum, Changhao Wang, Chuer Pan, C. Karen Liu (possible past Stanford University affiliation), Jeannette Bohg (possible past Stanford University affiliation), Shuran Song (possible past Google (United States) affiliation)
Abstract

Dexterous manipulation poses a challenging high-dimensional optimization problem, as useful behaviors require coordinated motion across many hand joints. In reinforcement learning (RL) and sampling-based trajectory optimization, exploration commonly relies on independent robot joint perturbations, making coordinated behaviors difficult to discover. Prior work reduces this search space for grasp learning using low-dimensional spaces of coordinated joint motions learned from human hand data, but t...

📄 SMART: Zero-Shot Sim-to-Real Articulated Object Manipulation via Large-Scale Synthetic Pretraining
🗓️ Published: 10/6/2026
🔗 http://arxiv.org/abs/2610.07652v1
👥 Authors: Jicong Ao, Shuhan Jiang, Yuling Zhong, Yanwen Liu, Yuhan Gao, Jiangyuan Zhao, Yang Zhang (possible past Tsinghua University affiliation), Shiqiang Zhu, Chenjia Bai, Xuelong Li (possible past Tencent (China) affiliation)
Abstract

The ability to interact with articulated objects is essential for embodied intelligent systems, but collecting large-scale real-world demonstrations for these interactions remains challenging due to the precise contact and constraint-following motions involved. Although simulation provides a promising alternative, existing synthetic data efforts cover limited articulated-object categories, while general-purpose synthesis pipelines lack explicit designs for part-level semantics and articulation c...

📄 SkillPoison: Progressive Skill Poisoning via Successful Experiences
🗓️ Published: 10/6/2026
🔗 http://arxiv.org/abs/2610.07645v1
👥 Authors: Lizhi Zhang, Xin He, Dianxuan Fu, Yuyuan Feng, Jiatong Li, Qi Wang (possible past Tsinghua University affiliation), Xin Wang (possible past University Of Edinburgh affiliation), Qinggang Zhang
Abstract

Self-improving LLM agents increasingly distill successful experiences into persistent, reusable skills. Existing skill attack methods corrupt this learning pipeline by injecting malicious triggers, behaviors, or false facts into individual experiences or extracted skills. However, such attacks are easily detected, and the injected malicious behaviors often fail to accumulate as persistent skills. In this paper, we show that skill poisoning can arise even from verified successful experiences, wit...

📄 QF3: Fast Flow RL with Filtered Q-Gradients
🗓️ Published: 10/6/2026
🔗 http://arxiv.org/abs/2610.08789v1
👥 Authors: Chung Min Kim, Brent Yi, David Mcallister, Hongsuk Choi, Himanshu Gaurav Singh, Jinkun Cao (possible past Shanghai Jiao Tong University affiliation), Ken Goldberg (possible past University Of California, Berkeley affiliation), Pieter Abbeel (possible past University Of California, Berkeley affiliation), Carmelo Sferrazza, Angjoo Kanazawa (possible past University Of California, Berkeley affiliation)
Abstract

Flow policies have become a standard policy class for learning robot behaviors from demonstrations, but reinforcement learning is still critical for improving pre-trained flow policies or learning them from scratch through interaction. We introduce QF3 (Fast Flow RL with Filtered Q-Gradients), an online off-policy RL algorithm that trains a flow policy with flow matching plus the critic's action gradient, backpropagated through a one-step prediction of the flow's output. To keep updates where th...

📄 Optimal and Efficient Online Inverse Optimization
🗓️ Published: 10/6/2026
🔗 http://arxiv.org/abs/2610.08735v1
👥 Authors: Anupam Gupta (possible past Carnegie Mellon University affiliation), Guru Guruganesh (possible past Google (United States) affiliation), Honghao Lin, Vahab Mirrokni (possible past Google (United States) affiliation), Renato Paes Leme (possible past Google (United States) affiliation), David P. Woodruff
Abstract

In online inverse linear optimization, a learner recommends an action and then observes the choice of an expert who maximizes a fixed, unknown linear objective on $\mathbb{R}^{d}$; the goal is to learn to optimize this objective without observing it. Sakaue recently obtained the optimal regret $O(\sqrt d)$ with a randomized algorithm making $(dT)^{O(d)}$ linear optimizations per round, and asked whether it can be attained in polynomial time. We answer positively: our deterministic algorithm has ...

📄 VETTA: Coordinating Turn- and Token-Level Credit Assignment for Multi-Turn LLM Agents
🗓️ Published: 10/6/2026
🔗 http://arxiv.org/abs/2610.08402v1
👥 Authors: Jiaju Chen, Min Yang (possible past Baidu (China) affiliation), Jinghua Piao, Xiaochong Lan, Xu Xia, Xiangnan He (possible past National University Of Singapore affiliation), Yong Li (possible past Tsinghua University affiliation)
Abstract

Multi-turn LLM agents often receive sparse task feedback across several interactions, while generating each response token by token. This creates two related credit-assignment questions: which responses helped achieve the outcome, and which generation decisions mattered within each response? Existing methods typically focus on only one level: turn-level methods evaluate complete responses but do not distinguish the decisions within them; token-level methods can propagate feedback across turns bu...

📄 SIFT: Search Intent-to-Filter Transformer for Multi-Task Personalized Filter Ranking at Airbnb
🗓️ Published: 10/6/2026
🔗 http://arxiv.org/abs/2610.07810v1
👥 Authors: Shashank Dabriwal, Tanya Piplani, Hao Li (possible past Tsinghua University affiliation), Yiwei Wang (possible past Google (United States) affiliation), Ashish Jain, Kedar Bellare, Stephanie Moyerman
Abstract

Search filters help guests navigate vast catalogs in two-sided marketplaces like Airbnb, and recommending the right filters can meaningfully lift booking conversion. Many such production filter-ranking systems, however, represent the guest through hand-engineered, pre-aggregated features generated by ETL pipelines. This makes it expensive to maintain and difficult to extend for new filter types or contextual dimensions (trip length, group size). We present SIFT (Search Intent-to-Filter Transform...

📄 TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models
🗓️ Published: 10/6/2026
🔗 http://arxiv.org/abs/2610.07767v1
👥 Authors: Xin Wang (possible past University Of Edinburgh affiliation), Hao Yu, Zhengyang Zhuge, Bochao Mao, Zheng Li, Junda Feng, Yuyan Luo, Yi Zhang (possible past Google (United States) affiliation), Yizhong Cao, Mi Zhang, Dayiheng Liu, Jianwei Zhang
Abstract

Reinforcement learning (RL) for post-training large language models (LLMs) incurs substantial computation and memory overhead during rollout generation, which motivates low-precision rollout for efficient RL training. However, existing FP4 RL methods suffer from a key limitation: they primarily optimize quantization accuracy on the training and rollout paths independently rather than directly reducing the discrepancy between the two quantized execution paths. In this work, we propose TRACE (Trai...

*Notable papers are those with at least two authors from a "big" AI/ML lab.