πŸ“„ Notable* Recent AI/ML arXiv Papers

Last updated just now...

πŸ“„ ContextWeave: A Real-World Workflow Benchmark
πŸ—“οΈ Published: 8/5/2026
πŸ”— http://arxiv.org/abs/2608.04830v1
πŸ‘₯ Authors: Bo Wang (possible past Tencent (China) affiliation), Yuqian Yao, Enxi Wang, Luozhijie Jin, Yang Liu (possible past Tsinghua University affiliation), Yiran Suo, Yuxuan Cai, Enyu Zhou, Yufei Gao, Honglin Guo, Tianyu Huai, Li Ji, Zhikai Lei, Bufan Li, Lizhi Lin, Jinxiu Liu, Jie Yang (possible past Shanghai Jiao Tong University affiliation), Jiazheng Zhou, Maosen Zhou, Pengfang Qian, Shichun Liu, Guanshan Liu, Hao Zheng, Yunhao Yu, Hang Yan, Jihua Kang, Xinchi Chen, Xipeng Qiu
Abstract

Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce it to retrieval or question answering. We introduce ContextWeave, a longitudinal benchmark that evaluates whether recalled experience improves downstream agent performance in realistic office-work streams. ContextWeave reconstructs privacy-preserved, multi-month workflows of 14 participants into 1,005 executable tasks, including 568 core evaluation tasks, wit...

πŸ“„ NSF-HRPT: Neural Semantic Field meets Hierarchical Risk Perception Tree for Safety-Critical Scenario Assessment
πŸ—“οΈ Published: 8/5/2026
πŸ”— http://arxiv.org/abs/2608.04776v1
πŸ‘₯ Authors: Yu Zhao (possible past Tencent (China) affiliation), Jiangyu Pan, Tao Hu (possible past Baidu (China) affiliation), Ming Yin, Fan Yang (possible past Tencent (China) affiliation), Jiangfan Liu, Xiubo Liang
Abstract

The ability to accurately assess and anticipate risks in safety-critical scenarios is crucial for autonomous driving systems. While existing research has made progress in collision prediction, accurately quantifying risk levels from monocular vision inputs remains challenging due to the complex dynamics of multi-agent interactions and the inherent uncertainty in real-world environments. To address these challenges, we present NSF-HRPT, a novel framework that combines learning-based perception wi...

πŸ“„ InsightEmb: Learning Action-Intent Embeddings for Agentic Insight Retrieval
πŸ—“οΈ Published: 8/5/2026
πŸ”— http://arxiv.org/abs/2608.04761v1
πŸ‘₯ Authors: Tsz Ting Chung, Jiangnan Li, Jie Zhou (possible past Tsinghua University affiliation), Mo Yu (possible past Tencent (China) affiliation)
Abstract

Self-improving agents accumulate reusable insights from prior trajectories, making retrieval increasingly important for turning accumulated experience into actionable guidance. At each decision step, retrieving the right insight can help the agent progress toward its goal, a setting we refer to as agentic insight retrieval. However, existing retrieval methods primarily model semantic similarity, while overlooking whether a retrieved insight resolves the agent's current decision bottleneck. We pr...

πŸ“„ The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering
πŸ—“οΈ Published: 8/5/2026
πŸ”— http://arxiv.org/abs/2608.04589v1
πŸ‘₯ Authors: Yuqian Fu, Tianwen Qian, Yanjun Li, Yu Li (possible past Tencent (China) affiliation), Kunyu Peng, Xu Zheng, Yongqin Xian, Alessio Tonioni, Yanwei Fu, Xiaoling Wang, Danda Paudel, Federico Tombari (possible past Google (United States) affiliation), Luc Van Gool (possible past Google (United States) affiliation), Leyi Wu, Yifan Zhao, Jinjie Zhang, Yinchuan Li, Yingcong Chen, Zixu Li, Zhiwei Chen, Zhiheng Fu, Wenbo Wang, Yupeng Hu, Weili Guan, Liqiang Nie, Takuya Murakawa, Toru Tamaki, Yi Wen, Zhenglin Du, Zhengyang Li, Lingling Li, Licheng Jiao, Wenping Ma
Abstract

EgoCross is a cross-domain egocentric video question answering benchmark designed to evaluate whether multimodal large language models can generalize beyond common daily-life scenarios. The first EgoCross Challenge was hosted at the Third EgoVis Workshop at CVPR 2026 and evaluated models on first-person videos from four target domains: surgery, industrial assembly, extreme sports, and animal perspectives. Each test example consists of an egocentric video clip, a question, and four candidate answ...

πŸ“„ Breaking the Curse ofMultilinguality inMany-to-Many Speech-to-Text Translation via a Resource-AwareMixture of Speech Encoders
πŸ—“οΈ Published: 8/5/2026
πŸ”— http://arxiv.org/abs/2608.04586v1
πŸ‘₯ Authors: Yexing Du, Kaiyuan Liu, Youcheng Pan, Bo Yang (possible past Tencent (China) affiliation), Chengpeng Fu, Yu Wang (possible past Tsinghua University affiliation), Ming Liu
Abstract

Multimodal large language models (MLLMs) have achieved significant success in speech-to-text translation (S2TT). However, when processing multilingual speech inputs, a single speech encoder shared across all languages suffers from the curse of multilinguality: languages at different resource levels compete for limited representation capacity, leading to strong high-resource performance but substantial degradation on low-resource speech. To address this problem and improve multilingual consistenc...

πŸ“„ What Is a Skill Worth? Structure-Aware Shapley Valuation of Agent Skills
πŸ—“οΈ Published: 8/5/2026
πŸ”— http://arxiv.org/abs/2608.04562v1
πŸ‘₯ Authors: Tao Li (possible past Baidu (China) affiliation), Junfeng Liu, Qinghua Zhao, Yifan Li, Lei Wang (possible past Baidu (China) affiliation), Bo Shao, Xuejun Liu, Linjun Shou
Abstract

Agent skills are increasingly optimized by automated feedback loops, producing long structured artifacts whose internal value remains unclear. We study skill valuation: assigning credit to the internal units of a fixed skill, such as rules, examples, scripts, and heuristics, under a fixed agent and held-out task distribution. Skill valuation differs from data or prompt-span valuation because skill units are structured: they may depend on other units, belong to a document hierarchy, trigger agent...

πŸ“„ AudioScape-TTA: A Structured Soundscape Benchmark for Fine-Grained Text-to-Audio Evaluation
πŸ—“οΈ Published: 8/5/2026
πŸ”— http://arxiv.org/abs/2608.04479v1
πŸ‘₯ Authors: Jinting Wang, Yuguang Yang, Shengyu Li, Yan Rong, Shan Yang (possible past Google (United States) affiliation), Xiaoda Yang, Li Liu (possible past National University Of Defense Technology affiliation)
Abstract

Text-to-audio (TTA) generation has recently achieved remarkable progress in synthesizing realistic audio from natural language descriptions. However, determining whether generated audio faithfully satisfies complex textual instructions remains challenging. Existing benchmarks mainly rely on global similarity metrics, providing limited insight into fine-grained semantic failures. To address this limitation, we introduce \textbf{AudioScape-TTA}, a structured and complexity-aware benchmark for fine...

πŸ“„ EndoVLM: An Endoscopy Vision-Language Pre-training Model via Anatomy-Guided Sparsity and Progressive Alignment
πŸ—“οΈ Published: 8/5/2026
πŸ”— http://arxiv.org/abs/2608.04472v1
πŸ‘₯ Authors: Zhenyu Yi, Jianwei Xu, Yue Hu, Zhongwei Qiu, Sijing Li, Liang Huang (possible past Tsinghua University affiliation), Bin Lv, Ling Zhang (possible past Nvidia (United States) affiliation), Yingda Xia
Abstract

The development of foundation models (FMs) is crucial for advancing endoscopic image analysis. However, existing endoscopy FMs mainly rely on self-supervised learning from uni-modal images or videos, overlooking the rich semantic knowledge contained in clinical reports. Furthermore, effectively leveraging these records is hindered by a fundamental modality gap: structured anatomical descriptions are not naturally mapped to specific frames within the high-redundancy, uncurated visual streams. In ...

πŸ“„ Training-Free Hashing-Based Attention via Binary Principal Components
πŸ—“οΈ Published: 8/5/2026
πŸ”— http://arxiv.org/abs/2608.04405v1
πŸ‘₯ Authors: Daohai Yu, Zhanpeng Zeng, Keyu Chen, Wenhao Li, Zhifeng Shen, Luxi Lin, Ruizhi Qiao (possible past Tencent (China) affiliation), Xing Sun (possible past Tencent (China) affiliation), Rongrong Ji (possible past Tencent (China) affiliation)
Abstract

Long-context large language models (LLMs) are increasingly deployed in real-world applications, yet self-attention remains a major efficiency bottleneck -- especially during decoding -- due to the necessity of repeatedly processing ever-growing key-value (KV) caches. Existing sparse attention reduce computation by attending to fewer KV pairs, but often suffer from substantial accuracy degradation, require additional training, or rely on expensive hashing. In this work, we present BinaryPC, a tra...

πŸ“„ Combating Knowledge Corruption in Agent Systems: A Byzantine-Tolerant Secure Collaborative RAG Framework
πŸ—“οΈ Published: 8/5/2026
πŸ”— http://arxiv.org/abs/2608.04366v1
πŸ‘₯ Authors: Zhaoqi Wang, Daqing He, Zijian Zhang, Ye Liu, Jiamou Liu, Zhirui Zeng, Zhan Qin, Zhen Li (possible past Google (United States) affiliation), Xin Li (possible past Google (United States) affiliation), Hongwei Yao, Jincheng An, Yong Liu, Yi Li (possible past University Of Washington affiliation), Qi Sun (possible past Google (United States) affiliation), Xiulei Liu, Liehuang Zhu
Abstract

While retrieval-augmented generation systems partially address the hallucination issues in large language models, it also introduces new vulnerabilities to knowledge corruption attacks. Adversaries exploit these vulnerabilities by poisoning documents provided by RAG system to manipulate LLM outputs. To counter this threat, we propose SecureCollaRAG, a Byzantine-tolerant collaborative RAG framework leveraging Multi-source Knowledge Validation Mechanism. Our approach enables agent system to secure...

πŸ“„ COMPAS: Difficulty-Aware Joint Search for Optimizing Code Generation
πŸ—“οΈ Published: 8/5/2026
πŸ”— http://arxiv.org/abs/2608.04336v1
πŸ‘₯ Authors: Jingzhi Gong, Jie M. Zhang (possible past Peking University affiliation), Gunel Jahangirova, Dong Huang, Mohammad Reza Mousavi, Mark Harman (possible past Meta (United States) affiliation)
Abstract

Code generation systems make each LLM call with a model, a prompt, and decoding settings. However, existing optimization methods usually tune only part of these choices or use one fixed configuration for all tasks: global optimizers search one configuration for all tasks, routers choose only a model, and prompt optimizers keep the model and decoding settings fixed. This leaves their joint, group-specific interactions unclear. We therefore examine how these choices interact and observe that promp...

πŸ“„ MatrAIx: Simulating the World with 8.3 Billion Persona Agents
πŸ—“οΈ Published: 8/4/2026
πŸ”— http://arxiv.org/abs/2608.04205v1
πŸ‘₯ Authors: Xiaomin Li, Yuexing Hao, Jianheng Hou, Jintao Huang, Qianfeng Wen, Shirley Huang, Yifan Liu, Xiaoyi Liu, Yilan Fan, Yijun Wang, Koutian Wu, Ruoqi Gao, Muhammad Ahmed Mohsin, Jing Tang, Brihi Joshi, Heming Liu, Zheyuan Deng, Zonglin Di, Sankalp Jajee, Jiuyao Lu, Zhiwei Zhang, Saksham Kapoor, Ishan Gupta, Yunhan Zhao, Chanwoo Park, Yucheng Lu, Bing Hu (possible past Meta (United States) affiliation), Weihang Xiao, Aravind Mohan, Hanwen Xing, Runyu Zhang, Mihir Kulshreshtha, Yuanda Xu, Qianyu Zhu, Dianzhuo Wang, Yuxin Xiao, Bowen Jiang, Yongye Su, Wenhao Chai, Zuxin Liu, Lawrence Yunliang Chen, Xuandong Zhao, Ethan Ye, Shivam Patel, Jason Xie, Alex Martin Richmond, Weixiang Ding, Emre Okcular, Diya Mathew, Ziheng Wang, Rana M. Shahroz Khan, Zhejian Peng, Fang Wu, Fan Nie, Xinyang Han, Yubin Kim, Jiawei Zhang, Zhenting Qi, Huangyuan Su, Xu Pan, Abinitha Gourabathina, Hyewon Jeong, Hemanth Neelgund Ramesh, Kumail Alhamoud, Kimia Hamidieh, Zidi Xiong, Samuel Schmidgall, Pengrui Han, Yepeng Huang, Yongheng Wang, Bowen Yang, Alex Gu, Yuchu Wang, Akshay Paruchuri, Brenna Li, Hejie Cui, Jiayuan Ding, Chaosheng Dong, Jiahao Wang, Yixuan He, Chi Wang (possible past Microsoft (United States) affiliation), Pamela Bhattacharya, Tianyi Peng, Paul Pu Liang, Mitchell Gordon (possible past Stanford University affiliation), Yilun Du (possible past Massachusetts Institute Of Technology affiliation), Marinka Zitnik, James Zou, Prasanna Tambe, Philip Torr (possible past University Of Oxford affiliation), Emily Fox, Asu Ozdaglar, Dawn Song (possible past University Of California, Berkeley affiliation)
Abstract

Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. MatrAIx has three core components: First, Persona 8B contains 8.3 billion persona records represented by 1,290 categorical dimensions. Record...

πŸ“„ OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents
πŸ—“οΈ Published: 8/4/2026
πŸ”— http://arxiv.org/abs/2608.05013v1
πŸ‘₯ Authors: Jingsheng Zheng, Xinyuan Fang, Jintian Zhang, Zhengke Gui, Huajun Chen (possible past Alibaba Group (China) affiliation), Ningyu Zhang (possible past Tencent (China) affiliation)
Abstract

LLM agents are increasingly applied to open-ended everyday requests that span work, study, and life. These tasks are long-horizon, cross-environment, and multimodal, forcing the agent to preserve goals and constraints across many steps while navigating heterogeneous tools and attachments. While prior work has addressed individual failure modes such as goals drift, states loss, and context overflow, whether a single harness can manage them jointly and remain effective across backends has received...

πŸ“„ A game theory for foundation models shows new paths to rational cooperation through similarity inference
πŸ—“οΈ Published: 8/4/2026
πŸ”— http://arxiv.org/abs/2608.03958v1
πŸ‘₯ Authors: Alexander Meulemans, Maciej WoΕ‚czyk, Marissa A. Weis, Rajai Nasser, Roberta Rocca, Seijin Kobayashi (possible past Eth Zurich affiliation), Guillaume Lajoie, Angelika Steger, Blake Richards, Marcus Hutter, James Manyika, Rif A. Saurous (possible past Google (United States) affiliation), JoΓ£o Sacramento (possible past Eth Zurich affiliation), Blaise AgΓΌera Y Arcas (possible past Google (United States) affiliation)
Abstract

As autonomous agents powered by foundation models are increasingly integrated into social and economic systems, understanding the principles governing their collective behavior is essential for ensuring safety and cooperation. Classical game theory, the dominant framework for modeling rational interaction, is built upon the assumption of `decoupled agency,' where agents treat their own decision-making as independent of the environment and other actors. Modern AI agents, however, jointly predict ...

πŸ“„ When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding
πŸ—“οΈ Published: 8/4/2026
πŸ”— http://arxiv.org/abs/2608.03918v1
πŸ‘₯ Authors: Ke Li (possible past University Of California, Berkeley affiliation), Jiayu Chen, Maoliang Li, Zihao Zheng, Hailong Zou, Hengyi Zhang, Xuanzhe Liu, Xiang Chen (possible past Tencent (China) affiliation)
Abstract

Efficient long-video understanding requires vision--language models (VLMs) to reason over a small number of frames selected as sparse visual evidence. Existing relevance-based methods rely on static one-shot selection with fixed frame budgets and candidate pools, while agent-based schedulers achieve adaptivity through costly multi-round reasoning and interactive search. We propose EcoFrame, a training-free framework for low-overhead query-adaptive visual evidence scheduling. EcoFrame leverages t...

πŸ“„ Enhancing VLM Reward Models Through Structure-Aware Fine-Tuning
πŸ—“οΈ Published: 8/4/2026
πŸ”— http://arxiv.org/abs/2608.03875v1
πŸ‘₯ Authors: Pyrros Koussios, Chenhao Li, Xin Chen (possible past Tencent (China) affiliation), Andreas Krause (possible past Eth Zurich affiliation)
Abstract

Designing effective reward functions remains a major bottleneck in Reinforcement Learning (RL). Recent work uses large foundation Vision-Language Models (VLMs) as reward models, computing text-observation similarity to bypass manual reward engineering. Although promising, these rewards are often noisy and unreliable, limiting their direct utility during deployment. We present Structure-Aware Fine-Tuning (SAFT), a simple, self-supervised method that refines these imperfect reward signals online w...

πŸ“„ KnowHal: A Knowledge-Driven Benchmark for Comprehensive Multimodal Hallucination Evaluation
πŸ—“οΈ Published: 8/4/2026
πŸ”— http://arxiv.org/abs/2608.03782v1
πŸ‘₯ Authors: Ruihan Li, Jiyang Tan, Kailin Jiang, Huining Li, Hengyang Lu, Yu Huang (possible past Tencent (China) affiliation), Qian Li (possible past National University Of Defense Technology affiliation), Yuntao Du
Abstract

Hallucination remains a critical challenge for developing trustworthy Multimodal Large Language Models (MLLMs). While existing benchmarks mainly focus on entity, attribute, and relation hallucinations, knowledge-related failures are often investigated separately, lacking a unified evaluation framework across different hallucination dimensions. To overcome this, we propose \textbf{KnowHal}, a benchmark that explicitly incorporates knowledge hallucination into multimodal hallucination evaluation s...

πŸ“„ MDLMPE: Distribution Aware Positional Encoding for Masked Diffusion Language Models
πŸ—“οΈ Published: 8/4/2026
πŸ”— http://arxiv.org/abs/2608.03769v1
πŸ‘₯ Authors: Tong Ling, Hang Lei, Feng Xiao (possible past Google (United States) affiliation), Changhui Sun, Jiahang Xie, Hao Liu (possible past Tencent (China) affiliation), Lu Liu, Yanlong Du
Abstract

Masked diffusion language models (MDLMs) enable parallel generation and bidirectional context modeling, but their positional context differs fundamentally from that of autoregressive (AR) models. Whereas AR decoding exposes a contiguous prefix, MDLM denoising produces dynamic, non-contiguous configurations of revealed and masked tokens. Conventional positional encodings such as RoPE capture sequence order and pairwise displacement but remain insensitive to this evolving token-availability struct...

πŸ“„ Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
πŸ—“οΈ Published: 8/5/2026
πŸ”— http://arxiv.org/abs/2608.05000v1
πŸ‘₯ Authors: Junlin Han, Shengbang Tong, David Fan, Minghao Chen, Philip Torr (possible past University Of Oxford affiliation), Filippos Kokkinos, Mike Lewis (possible past Meta (United States) affiliation)
Abstract

Vision offers a critical axis for advancing foundation models, driving a shift towards natively unified multimodal pretraining. Despite this momentum, the design space and the fundamental mechanisms of how modalities interact during unified training remain underexplored. We provide empirical clarity through a systematic exploration of multimodal pretraining. Our controlled experiments on both synthetic and large-scale real-world datasets yield four key insights into the physics of multimodal pre...

πŸ“„ State2State: Environment-Derived Mid-Training for LLM Agents
πŸ—“οΈ Published: 8/5/2026
πŸ”— http://arxiv.org/abs/2608.04934v1
πŸ‘₯ Authors: Xuanyu Lei, Yiqi Zhu, Chenliang Li, Kaiming Liu, Peng Li (possible past Tsinghua University affiliation), Ming Yan, Jieping Ye, Ya-Qin Zhang, Yang Liu (possible past Tsinghua University affiliation)
Abstract

Training LLM agents commonly relies on supervised fine-tuning from expert trajectories or online reinforcement learning over human-specified tasks with handcrafted verifiers. Though effective, both remain bottlenecked by externally specified tasks and supervision signals, limiting the scalability and diversity of agent training. We study an environment learning paradigm in which agents acquire interaction and manipulation capabilities solely through environment interaction, without externally sp...

πŸ“„ Deltoris: Enabling Real-time VLA Inference in Embodied AI via Bit-level Sparsity and Speculative Inference
πŸ—“οΈ Published: 8/5/2026
πŸ”— http://arxiv.org/abs/2608.04428v1
πŸ‘₯ Authors: Zheng Liu, Zeyu Guo, Zihan Liu, Anbang Wu, Han Zhao, Fangxin Liu, Zhezhi He (possible past Shanghai Jiao Tong University affiliation), Yinhe Han, Jingwen Leng, Minyi Guo, Yiming Gan, Yu Feng (possible past University Of California, Berkeley affiliation)
Abstract

Vision-language-action (VLA) models have emerged as a key component in embodied AI. Among existing approaches, diffusion-based VLA models achieve superior motion quality and generalization. However, diffusion-based VLA models are compute-intensive and must run at high control frequency, e.g., 50-200 Hz. Thus, it imposes strict latency and energy constraints on edge devices. In this work, we present Deltoris, an algorithm-hardware co-design framework for efficient diffusion-based VLA inference....

*Notable papers are those with at least two authors from a "big" AI/ML lab.