📄 Notable* Recent AI/ML arXiv Papers

Last updated just now...

📄 Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams
🗓️ Published: 8/12/2026
🔗 http://arxiv.org/abs/2608.12262v1
👥 Authors: Weihao Bo, Shan Zhang, Yanpeng Sun, Jie Liu (possible past Tencent (China) affiliation), Yongke Yao, Jinhao Du, Wei He (possible past Baidu (China) affiliation), Kai Zou, Zechao Li, Jingdong Wang (possible past Baidu (China) affiliation)
Abstract

Multimodal Large Language Models (MLLMs) have been growing the capability for scientific writing and collaboration. For example, OpenAI Prism is a free workspace for scientific writing and collaboration. One important feature in Prism is turning scientific diagrams directly into LaTeX TikZ code. In this paper, we build a benchmark, Diagram-MMU, a multi-modal benchmark designed to assess MLLMs' ability for scientific diagram parsing and understanding. Diagram-MMU features 3.7k curated diagrams an...

📄 One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL
🗓️ Published: 8/12/2026
🔗 http://arxiv.org/abs/2608.12253v1
👥 Authors: Simon Yu, Nicholas Tomlin, Marwa Abdulhai, Ximing Lu, Derek Chong, Abe Hou, Dilara Soylu, Sergey Levine (possible past University Of Washington affiliation), Christopher D. Manning (possible past Stanford University affiliation), Weiyan Shi
Abstract

Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior. We show that this approach systematically fails to generalize, and trace the failure to simulator collapse: because the simulator LLM is mode-collapsed, an LLM policy trained against it overfits to narrow strategies that exploit the simulator's dominant mode, and such a policy transfers poorly to unseen simulators and real users. We formalize this collapse theo...

📄 Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
🗓️ Published: 8/12/2026
🔗 http://arxiv.org/abs/2608.12036v1
👥 Authors: Mengru Wang, Junfeng Fang, Shuofei Qiao, Zhenqian Xu, Haoming Xu, Haoxiong Wang, Shumin Deng (possible past Alibaba Group (China) affiliation), Linyi Yang, Zhixiang Cui, Xin Xu, Yunzhi Yao, Buqiang Xu, Fei Shen, Haozhe Luo, Yunxiang Wei, Ningyu Zhang (possible past Tencent (China) affiliation), Julian Mcauley, Tat Seng Chua, Huajun Chen (possible past Alibaba Group (China) affiliation)
Abstract

AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood. As AI development becomes faster and increasingly automated, mechanistic exploration remains largely manual, widening the gap between what models can do and our ability to understand and control them. To bridge this gap, we introduce Mechanist, an agentic system that uses AI as a scientific instrument for the autonomous discover...

📄 HCGRec: Hint-Conditioned Generative Recommendation with Semantic IDs
🗓️ Published: 8/12/2026
🔗 http://arxiv.org/abs/2608.11980v1
👥 Authors: Kangning Zhang, Haotian Fang, Xukun Luo, Hao Yin, Yang Gao (possible past Tencent (China) affiliation), Peng Yan, Weiwen Liu, Weinan Zhang (possible past Shanghai Jiao Tong University affiliation), Yong Yu (possible past Shanghai Jiao Tong University affiliation)
Abstract

Semantic-ID generative recommenders represent each item as a short sequence of discrete semantic tokens and predict the next item by autoregressively generating this token sequence. This paradigm enables a unified generation interface for item IDs, histories, and item text, but it also creates a structured optimization bottleneck during reward-based post-training: when an early semantic token enters the wrong branch of the item-token space, finite rollout groups rarely reach the ground-truth ite...

📄 Retry, Switch, or Abstain? Learning Strategy-Aware Tool-Use Policies via Controlled Error Injection
🗓️ Published: 8/12/2026
🔗 http://arxiv.org/abs/2608.11977v1
👥 Authors: Chaoran Chen, Vy Nguyen, Ziji Zhang, Abhinav Gullapalli, Ziyi Wang, Yuxuan Lu, Dakuo Wang (possible past Tencent (China) affiliation), Jing Huang (possible past Meta (United States) affiliation), Zhou Yu, Jin Lai
Abstract

Tool-using LLM agents are commonly trained and evaluated in environments where tool calls succeed reliably, yet deployed tools can fail transiently, persistently, or silently. Robust recovery therefore requires more than repeated retries: an agent may need to retry the same path, switch to an alternative, or recognize that no viable path remains. We present BENCH2ROBUST, a framework that converts failure-free tool-use benchmarks into controlled stochastic environments with scenario-controlled so...

📄 G0.5: One Autoregressive Stream for Robot Reasoning and Action
🗓️ Published: 8/12/2026
🔗 http://arxiv.org/abs/2608.11739v1
👥 Authors: Yicheng Liu, Zibin Dong, Baijun Ye, Tianyuan Yuan, Tao Jiang (possible past Alibaba Group (China) affiliation), Anqi Yang, Shicheng Cao, Haonan Liu, Yue Sun, Zihan Guo, Xiao Liu (possible past Baidu (China) affiliation), Dong Ke, Changxun Pan, Chenru Wu, Tailai Cheng, Xiaoshu Ren, Xinlei Zhang, Jianning Cui, Zijie Zhao, Haoyu Zhang, Kaiming Xu, Haodong Yang, Bowen Zhang, Jiahui Niu, Shaoting Zhu, Shiduo Zhang, Hang Zhao (possible past Nvidia (United States) affiliation)
Abstract

The prevailing recipe for Vision-Language-Action (VLA) models couples a pretrained VLM with a separately trained flow-matching action expert. This makes the VLM a context encoder rather than a decision-maker. We introduce G0.5, a pretrained autoregressive VLA in which a single transformer decoder emits reasoning and action tokens under a single objective. Three components make this tractable at foundation-model scale: a learnable cross-embodiment action tokenizer that maps heterogeneous robot ac...

📄 Fingerprinting Text-to-Image Diffusion Models via Collapsed Generation
🗓️ Published: 8/12/2026
🔗 http://arxiv.org/abs/2608.11732v1
👥 Authors: Yuanmin Huang, Chen Chen (possible past Tencent (China) affiliation), Geng Hong, Xiaoyu You, Hui Xue, Zhenxing Qian, Mi Zhang, Min Yang (possible past Baidu (China) affiliation)
Abstract

Proprietary text-to-image diffusion models are increasingly distributed as hosted services and downloadable checkpoints, making their intellectual property (IP) protection an increasingly critical concern when model leakage, copying, or unauthorized fine-tuning is disputed. In this work, we present a non-invasive model fingerprinting framework based on \emph{collapsed generation}, a phenomenon where certain input conditions produce highly consistent images across multiple stochastic seeds. We sh...

📄 XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication
🗓️ Published: 8/12/2026
🔗 http://arxiv.org/abs/2608.11676v1
👥 Authors: Wooseong Yang, Wei-Chieh Huang, Weizhi Zhang, Yu Wang (possible past Tsinghua University affiliation), Philip S. Yu (possible past Tsinghua University affiliation), Junhyun Lee
Abstract

Heterogeneous multi-agent LLM systems, where agents are powered by different model families, can outperform homogeneous configurations by reducing redundant reasoning patterns. Yet existing communication protocols either operate through text, discarding the sender's internal representations, or require architectural homogeneity for latent-level transfer. We identify the entity grounding problem in cross-architecture communication: cross-attention bridges that transfer continuous representations ...

📄 CoAdapt-GUI: Joint Workflow Context and Policy Adaptation for Unseen GUI Applications
🗓️ Published: 8/12/2026
🔗 http://arxiv.org/abs/2608.11588v1
👥 Authors: Linqiang Guo, Li Gu, Zihuan Jiang, Zhixiang Chi, Siobhan Reid, Ziqiang Wang, Yuanhao Yu, Wei Liu (possible past Tsinghua University affiliation), Yang Wang (possible past Baidu (China) affiliation), Tse-Hsun, Chen
Abstract

Mobile GUI agents remain brittle when deployed to applications absent from source training. We study novel-app generalization under a limited target interaction budget and without target demonstrations. We introduce CoAdapt-GUI, a test-time adaptation (TTA) framework that jointly adapts structured workflow context and policy from the agent's own target-app rollouts and rewards. The workflow context retains transferable procedures, failure modes, and verification rules while excluding app-bound s...

📄 EnterpriseRAG: Benchmarking LLM Instruction Adherence and Robustness under Non-Ideal Enterprise Retrieval
🗓️ Published: 8/12/2026
🔗 http://arxiv.org/abs/2608.11584v1
👥 Authors: Huiqi Miao, Xinbao Sun, Bo Wang (possible past Tencent (China) affiliation), Fanyu Meng, Lijun Mei, Na Wu, Di Jin (possible past Tencent (China) affiliation), Chao Deng, Junlan Feng
Abstract

Enterprise RAG deployments face a critical reliability gap: while LLMs satisfy 80% of individual constraints, only 26.8% of responses meet all requirements simultaneously, revealing a 57-point orchestration gap. Existing benchmarks assume clean retrieval with simple queries, failing to capture production conditions where noisy documents and multi-dimensional constraints coexist. We introduce EnterpriseRAG, a benchmark of 983 expert-validated samples across six domains that systematically simulat...

📄 Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence
🗓️ Published: 8/11/2026
🔗 http://arxiv.org/abs/2608.11341v1
👥 Authors: Brian Wang, Bin Feng, Xiaoman Pan, Chenyang An, Felix Liu, Tangqi Fang, Gongbo Sun, Lingfeng Shen, Ning Wang, Handuo Zhang, Feng Chen, Fuchao Yang, Xiang Wang (possible past Tencent (China) affiliation), Jiacheng Lin, Siting Li, Zixuan Liu, Chi Han, Zhenhailong Wang, Kunlun Zhu, Lawrence Zhao, Yueqi Guo, Kailong Wen, Feng Xing, Yiling Guo, Lidong Bing (possible past Tencent (China) affiliation), David Tan, Bo An, Heng Ji, Sheng Wang (possible past Tencent (China) affiliation)
Abstract

Apollo did not reach the Moon merely because its engineers could solve difficult equations. It succeeded by turning a distant ambition into a mission architecture of explicit objectives, simulation, verification, and repeated correction. AI now faces a similar transition: frontier models can solve difficult tasks once the problem, tools, and success criteria are specified, yet consequential real-world challenges rarely arrive in an executable or verifiable form. We introduce Apodex Discovery, ...

📄 ComBodied Agents: a New Paradigm of Human-Centric Agentic AI
🗓️ Published: 8/11/2026
🔗 http://arxiv.org/abs/2608.10915v2
👥 Authors: Qianggang Ding (possible past Tsinghua University affiliation), Xingyao Wang, Rui Feng, Zhibin Wang, Feixiang Yao, Kelong Mao, Hao Sun, Zhiyao Luo, Jiankai Tang, Lei Li (possible past Carnegie Mellon University affiliation), Jiadong Guo, Minheng Ni, Weicong Lin, Chenxi Yang, Hongxiang Gao, Zhenghua Chen, Yang Bai, Min Wu, Jun Cheng (possible past Deepmind (United Kingdom) affiliation), Huazhu Fu (possible past Inception Institute Of Artificial Intelligence affiliation), Dacheng Tao, Bang Liu
Abstract

After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication. Yet neither explains whether the person forgot, is confused, has side effects, or deliberately refused, nor what support is appropriate. This reveals a structural gap in Agentic AI: Digital Agents primarily transform software states, while Embodied Agents transform physical states; neither makes a person's evolving state and agency the primary object of modeli...

📄 LEMUR: Latent Entropy-aware Multimodal Unlearning via Visual-anchored Reasoning Redirection
🗓️ Published: 8/12/2026
🔗 http://arxiv.org/abs/2608.11691v1
👥 Authors: Xinhao Zhong, Yuxia Qiao, Junhao Li, Hao Fang (possible past University Of Washington affiliation), Yi Sun (possible past Google (United States) affiliation), Bin Chen
Abstract

Reinforcement-learning (RL) post-training equips multimodal large reasoning models (MLRMs) with exploratory chains of thought (CoT), substantially improving visual reasoning. However, we find that this capability introduces a distinct privacy vulnerability: even when a sensitive fact is successfully unlearned from the final answer, the model may still reproduce it in its reasoning trace. This leakage is substantially more pronounced in natively RL-trained MLRMs than in their non -reasoning base ...

📄 Adaptation of Generalist Robot Policies with Minimal Data
🗓️ Published: 8/11/2026
🔗 http://arxiv.org/abs/2608.11363v1
👥 Authors: Shreyas Kowshik, Sreyas Venkataraman, Leo Wang (possible past Tencent (China) affiliation), Niharika Pant, Max Simchowitz, Aviral Kumar (possible past University Of California, Berkeley affiliation)
Abstract

A central goal in robot learning is to move beyond task-specific human data collection toward robots that improve through autonomous interaction. Yet fully autonomous learning remains difficult with current policies: sparse rewards and weak zero-shot exploration make it unlikely that a robot will discover successful behavior from scratch. We study minimal-data adaptation, a regime in which a pre-trained robot policy must learn a new task from as little as one demonstration followed by autonomous...

📄 ReOrder-OPD:Reliability-Aware Prompt Ordering for On-Policy Distillation
🗓️ Published: 8/11/2026
🔗 http://arxiv.org/abs/2608.10905v1
👥 Authors: Ximo Zhu, Ruiqi Liu, Rong Wang, Ping Wu, Xiang Zheng, Wenzhuo Xu, Xubin Yao, Zhiyuan Yan, Bo Li (possible past Tencent (China) affiliation), Jun Gao (possible past Nvidia (United States) affiliation), Xiaolei Lv
Abstract

On-policy distillation (OPD) applies token-level teacher supervision to student-generated trajectories, but this supervision is not always reliable. Existing methods use local confidence or teacher-student agreement to weight, filter, or truncate the sampled trajectory. These signals do not directly determine whether the teacher can continue a student prefix to a correct answer, and trajectory-level interventions can conflate one rollout's unreliability with low expected training value of its pr...

*Notable papers are those with at least two authors from a "big" AI/ML lab.