📄 Notable* Recent AI/ML arXiv Papers

Last updated just now...

📄 AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
🗓️ Published: 7/30/2026
🔗 http://arxiv.org/abs/2607.28617v1
👥 Authors: Xiangning Lin, Shenzhe Zhu, Shu Yang, Zhenyu Zhang, Haoqian Zhang, Yipeng Zhao, Chengxuan Qian, Tianwei Wang, Ziheng Zhang, Zhenlong Yuan (possible past Tsinghua University affiliation), Dingcheng Wang, Juncheng Wu, Yuan Si, Jiaxin Liu, Baolong Bi, Robert Mahari, Tobin South, Dazza Greenwood, Zexue He, Rishi Bommasani, Sophia Kazinnik, Andreas Haupt, Samuele Marro, Erik Brynjolfsson, Alex Pentland (possible past Massachusetts Institute Of Technology affiliation), Jiaxin Pei
Abstract

System prompts are instructions configured by developers to govern the behaviors of foundation models in AI applications. They are used throughout commercial AI products, but are rarely disclosed to the public or regulators, creating a serious trust and accountability gap in the wide deployment of AI systems. In this paper, we introduce Artificial Intelligence System Prompt Assurance (AISPA), a user-centric framework for systematically auditing system prompts in AI systems. AISPA examines specif...

📄 OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
🗓️ Published: 7/30/2026
🔗 http://arxiv.org/abs/2607.28609v1
👥 Authors: Qiushi Sun, Kanzhi Cheng, Yian Wang, Bowen Yang, Hang Yan, Liheng Chen, Fangzhi Xu, Zichen Ding, Nuo Chen, Jialin Cao, Xingdong Gong, Zehao Li, Kaiming Jin, Xinfeng Yuan, Zhoumianze Liu, Jingyang Gong, Zhangyue Yin, Jiahui Gao, Zhiyong Wu (possible past Tsinghua University affiliation), Tianbao Xie, Jianbing Zhang, Ben Kao, Lingpeng Kong (possible past Google (United States) affiliation)
Abstract

Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the task instruction is central to CUA evaluation, data curation, and reinforcement learning. Neither human-written verifiers nor human annotators can provide such verification at scale, so the field increasingly turns to vision-language models (VLMs) as judges of CUA trajectories. But a fundamental question has long gone...

📄 A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports
🗓️ Published: 7/30/2026
🔗 http://arxiv.org/abs/2607.28466v1
👥 Authors: Jia Yu, Yan Zhu, Yili He, Zilong Wang, Xinyang Jiang (possible past Tencent (China) affiliation), Peiyao Fu, Ruijie Yang, Tianyi Chen, Siyuan Li (possible past Tencent (China) affiliation), Zhihua Wang, Fei Wu (possible past Google (United States) affiliation), Quanlin Li, Xian Yang, Pinghong Zhou, Shuo Wang (possible past Nvidia (United States) affiliation)
Abstract

Vision-language models remain underused in colonoscopy despite the rich expert descriptions recorded in routine reports. These reports document lesion appearance, size and location but summarise entire procedures rather than caption individual frames, leaving clinical findings only weakly linked to the corresponding images. Here we develop EndoCLIP, a colonoscopy vision-language foundation model trained on 125,756 lesion-level image-text pairs progressively recovered from 280,476 routine colonos...

📄 Teffic-Audio: Tell Fact from Fiction
🗓️ Published: 7/30/2026
🔗 http://arxiv.org/abs/2607.28351v1
👥 Authors: Wan Lin, Li Wang (possible past Tesla (United States) affiliation), Jindong Wang, Kunyu Feng, Zhizheng Wu (possible past University Of Edinburgh affiliation)
Abstract

Speech deepfake detection has expanded in scope with increasingly heterogeneous spoofing mechanisms, including speech synthesis, voice conversion, vocoder reconstruction, and neural-codec resynthesis. The resulting spoofing artifacts can be further shaped by variability in source speech, recording environments, and transmission channels. This variability makes robust generalization across heterogeneous conditions a central requirement for practical detection systems. This report presents Teffic-...

📄 Paying for Honesty Without Knowing the Truth: Reputation-Penalty Design for LLM Marketplace Agents
🗓️ Published: 7/30/2026
🔗 http://arxiv.org/abs/2607.28330v1
👥 Authors: Mingdai Yang, Shicheng Fan, Kejing Yu, Duohao Wang, Li Sun, Hao Peng (possible past Tsinghua University affiliation), Philip S. Yu (possible past Tsinghua University affiliation), Zhiwei Liu
Abstract

LLM agents increasingly act as autonomous merchants that write their own product listings, and under competitive pressure, they fabricate attributes to win sales. Even under instructions to be honest, they fabricate attributes in a majority of listings across models. A platform's obvious remedy---verifying each claim against the truth---is unavailable, because it observes only a noisy, biased complaint signal, never the ground truth. We design CARP, a reputation-penalty mechanism with a deadband...

📄 ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding
🗓️ Published: 7/30/2026
🔗 http://arxiv.org/abs/2607.28312v1
👥 Authors: Mingkang Dong, Muxin Pu, Jie Li, Bohan Guo, Songruo Chen, Bin Ren, Xu Zheng, Chen Zhao (possible past Stanford University affiliation), Tianwen Qian, Mohamed Elhoseiny (possible past Meta (United States) affiliation), Yuqian Fu
Abstract

Streaming video understanding requires models to continuously retain useful visual evidence before future questions are known. Existing approaches primarily manage the growing visual context according to token importance, temporal redundancy, or segment-level relevance, but rarely organize evidence around objects that persist and evolve over time. Thus, in this paper, we introduce ObjectStream, a training-free framework that treats latent objects as memory anchors for streaming video understandi...

📄 Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
🗓️ Published: 7/30/2026
🔗 http://arxiv.org/abs/2607.28227v1
👥 Authors: Hanzhang Zhou, Panrong Tong, Xu Zhang (possible past Tencent (China) affiliation), Quyu Kong, Chenglin Cai, Tianyu Xia, Gongjie Zhang, Jianan Zhang (possible past Google (United States) affiliation), Long Li, Long Chen (possible past Tencent (China) affiliation), Lei Wang (possible past Baidu (China) affiliation), Gaole Dai, Pengxiang Li, Liangyu Chen, Yue Wang, Steven Hoi
Abstract

GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them toward real-world use, we envision agents that operate reliably on real devices, execute workflows across platforms, combine GUI interaction with CLI execution, complete long-horizon tasks, proactively initiate useful services, and autonomously improve their capabilities with minimal human effort. Guided by this vision, we present Qwen-UI-Agent, a real-world centric foundation GUI ag...

📄 Where and When to Commit: Candidate-Aware Decoding for Diffusion Language Models
🗓️ Published: 7/30/2026
🔗 http://arxiv.org/abs/2607.28166v1
👥 Authors: Chia-Ming Lee, Ming-Ching Chang (possible past Nvidia (United States) affiliation), Xin Li (possible past Google (United States) affiliation), Yu-Lun Liu, Chih-Chung Hsu
Abstract

Diffusion language models (DLMs) expose a provisional prediction at every denoising step, creating an opportunity for generation-time early exit that stops decoding before the schedule is exhausted. Existing early-exit gates decide termination from fixed-region confidence statistics or schedule-dependent rules, evidence too coarse for a decision that freezes every remaining position at once, so they fire prematurely on long chain-of-thought outputs whose answers stabilize only near the end. Adap...

📄 Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale
🗓️ Published: 7/30/2026
🔗 http://arxiv.org/abs/2607.28074v1
👥 Authors: Yash Pandya, Sahil Gupta, Sarthak Harne, Archana Yadav, Kavyansh Chourasia, Hussein Mozannar, Vibhav Vineet (possible past University Of Oxford affiliation), Sara Abdali, Corby Rosset, Yash Lara, Ahmed Awadallah, Ece Kamar (possible past Microsoft (United States) affiliation), Akshay Nambi
Abstract

Computer-use agents learn from what their actions change, so training one needs applications it can act on, break and reset. The applications that matter most are login-gated and stateful, so synthetic environments stand in for them. Recent pipelines generate such environments in bulk, which moves the bottleneck from how many exist to what is inside each one. The returns, we find, come from three properties: how much behavioural depth an environment carries, whether it targets the interaction an...

📄 DataClawEval: A Benchmark for Data Engineering Agents in Real Industrial Harness
🗓️ Published: 7/30/2026
🔗 http://arxiv.org/abs/2607.28033v1
👥 Authors: Debin Meng, Jiaming Yang, Zefang Zong, Tengyue Xu, Haining Xie, Yang Li (possible past Google (United States) affiliation), Peng Chen (possible past Tencent (China) affiliation)
Abstract

Large language models (LLMs) and LLM-based agents are increasingly being deployed to automate complex workflows, promising to revolutionize data management and processing. However, existing benchmarks predominantly focus on simplified Text-to-SQL translation or data analysis, leaving the critical and complex domain of end-to-end data engineering largely unexplored. To bridge this gap, we introduce DataClawEval, the first comprehensive benchmark designed specifically to evaluate the end-to-end ta...

📄 Flux-OPD: On-Policy Distillation with Evolving Contexts
🗓️ Published: 7/30/2026
🔗 http://arxiv.org/abs/2607.28022v1
👥 Authors: Yuran Wang, Zekun Wang, Bohan Zeng, Ruixu Zhang, Wenxuan Liu, Liu Yang (possible past Google (United States) affiliation), Yifan Dai, Yang Shi, Bozhou Li, Chengzhuo Tong, Daili Hua, Yuanxing Zhang, Wentao Zhang (possible past Mila - Quebec Artificial Intelligence Institute affiliation)
Abstract

Large language model training in open-ended domains lacks verifiable rewards, making task preferences difficult to formalize as effective supervision. Contexts can convey such preferences, yet provide little additional supervision once distilled into the student, motivating contexts that evolve with student performance. However, directly using evolving contexts as in-training supervision results in an unstable distillation target and conflicting distributions, requiring mechanisms to stabilize t...

📄 From Scoring to Acting: Outcome-Verified Comparative Self-Distillation for LLM Agents
🗓️ Published: 7/30/2026
🔗 http://arxiv.org/abs/2607.27937v1
👥 Authors: Xu Xia, Jinghua Piao, Min Yang (possible past Baidu (China) affiliation), Xiaochong Lan, Jiaju Chen, Yong Li (possible past Tsinghua University affiliation)
Abstract

Recent work on LLM agents is shifting from external capability elicitation to capability internalization, enabling agents to retain useful skills without retrieval at inference time. On-policy self-distillation (OPSD) offers a promising direction, but many existing methods typically supervise students by scoring actions along student-generated trajectories. Such supervision has two limitations: teacher preferences are not validated by environment outcomes, and action-level scores underuse inform...

📄 FinanceHarness: Autonomous Financial Deep Research Framework
🗓️ Published: 7/30/2026
🔗 http://arxiv.org/abs/2607.27853v1
👥 Authors: Yijia Xiao, Rujun Han, Yanfei Chen, Zifeng Wang, Ke Jiang, Zhongying Cuizhu, Vishy Tirumalashetty, Wei Wang (possible past University Of Oxford affiliation), Burak Gokturk, Tomas Pfister (possible past University Of Oxford affiliation), Chen-Yu Lee (possible past Google (United States) affiliation)
Abstract

Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products. However, most deep research systems write general-purpose reports, which are inadequate for financial deep research. Financial research demands specialized knowledge to analyze historical patterns and forecast upcoming events. Automating financial deep research therefore requires both a layered harness to drive the research agent and a verifiable, point-in-time benchmark t...

📄 ROCS: Request-Oriented Compute Sharing for Efficient Large-Scale Recommendation
🗓️ Published: 7/30/2026
🔗 http://arxiv.org/abs/2607.27744v1
👥 Authors: Yuxin Chen, Liang Luo, Buyun Zhang, Jian Jiao, Boda Li, Haoyu Wang (possible past Tencent (China) affiliation), Tongyi Tang, Ao Cai, Zijian Shen, Zhengkai Zhang, Wenyi Xie, Ryan Dick, Han Liu (possible past Tsinghua University affiliation), Neng Shi, Bin Yu, Jianbo Xiao, Shuyao Bi, Hongtao Yu, Yuanwei Fang, Zhuoran Zhao, Sijia Chen, Yang Chen (possible past Tencent (China) affiliation), Shuqi Yang, Qianru Li, Zikun Liu, Wei Ling, Sihan Zeng, Longhao Jin, Jiaxin Lu, Yinbin Ma, Jiawei Li, Yichen Ruan, Yong Ler Lee, Birmingham Guan, Zijian Li, Jianbo Sun, Zhengyu Zhang, Zeliang Chen, Xiaohan Wei, Yuchen Hao, Gp Musumeci, Venkatesh Ranganathan, Yantao Yao, Chunqiang Tang, Wenlin Chen (possible past Meta (United States) affiliation), Santanu Kolay, Ellie Dingqiao Wen
Abstract

Modern recommendation models gain prediction quality by scaling feature-interaction and sequence modules, but production cost constraints cap how far systems can scale. In this work, we propose Request-Oriented Compute Sharing (ROCS), a modeling and inference paradigm that exploits a unique property of recommendation inference: each user request is evaluated against many candidates, while request-side features are shared across candidates. ROCS defers request-candidate interactions as late as ...

📄 JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles
🗓️ Published: 7/30/2026
🔗 http://arxiv.org/abs/2607.27670v1
👥 Authors: Shawn Li, Wei Yang (possible past Tencent (China) affiliation), Jike Zhong, Jiate Li, Jiawei Yang, You Qin, Ryan Rossi, Franck Dernoncourt, Roger Zimmermann, Yue Wang, Zhengzhong Tu (possible past Google (United States) affiliation), Vicente Ordonez, Mohit Bansal, Yue Zhao
Abstract

Jigsaw puzzle solving requires jointly reasoning about visual content and geometric constraints, yet existing benchmarks use rectangular cuts that create ambiguous ground truth in texture-repeated regions. We introduce \textit{\ours{}}, a benchmark with tab-and-blank interlocking pieces where geometric constraints provide strong local compatibility requirements that, combined with visual content, yield unambiguous ground truth. Across 95K instances at four grid densities (4$\times$4 to 16$\times...

📄 Change2Task: From Repository Changes to Executable Coding Agent Tasks and Environments
🗓️ Published: 7/30/2026
🔗 http://arxiv.org/abs/2607.28591v1
👥 Authors: Haomin Qi, Xingliang Wang, Xuanqi Gao, Baihui Sang, Xin Zhang (possible past Google (United States) affiliation), Minghua Ma, Pengfei Gao, Yu Kang, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang, Qi Zhang (possible past Tencent (China) affiliation)
Abstract

Scaling coding agents requires a continuing supply of executable data for training, benchmarking, and continuous evaluation. Each task must couple a realistic software state with a specification, development tools, and reliable verification. To expand this supply, we present Change2Task, a system grounded in repository history that converts merged pull requests into verified tasks on healthy modern revisions of the same repository. It aligns historical evidence with evolved code, reconstructs ta...

📄 Why Are GUI Agents Correct but Late? Decode on the Decision-Time Critical Path, Tested with Pre-Compiled Policy Trees
🗓️ Published: 7/30/2026
🔗 http://arxiv.org/abs/2607.28399v1
👥 Authors: Zihan Dong, Rui Qian (possible past Shanghai Jiao Tong University affiliation), Qishi Zhan, Dongshen Peng, Kaixin Li, Yu Li (possible past Tencent (China) affiliation)
Abstract

Computer-use agents often fail on transient GUI events because they produce the correct action only after the relevant window has already closed. We identify the main cause as expensive autoregressive decoding on the decision-time critical path. We propose Adaptive Anticipatory Policy Trees (AAPT), which eliminates this delay without modifying the underlying model. During idle screen periods, the same frozen multimodal model constructs a bounded conditional policy tree with observable guards, pr...

📄 What Makes Graph Unified? Principles and Generative Sliding-Window Transformer for Graph Foundation Models
🗓️ Published: 7/30/2026
🔗 http://arxiv.org/abs/2607.27966v1
👥 Authors: Dongxiao He, Siqi Liu (possible past University Of Oxford affiliation), Jitao Zhao, Yawen Li, Yi Wang, Di Jin (possible past Tencent (China) affiliation)
Abstract

Graph Foundation Models (GFMs) have recently emerged as a promising paradigm for general-purpose graph learning, aiming to learn reusable knowledge that generalizes across diverse graph domains and downstream tasks, reducing the need for specific model development. Achieving this goal requires reconciling the substantial heterogeneity in node features, graph structures, and semantic information across domains. Among them, heterogeneous node features constitute a fundamental input-level barrier, ...

📄 S-CEReBrO: Breaking the Memory Barrier in Continuous EEG Monitoring
🗓️ Published: 7/30/2026
🔗 http://arxiv.org/abs/2607.27913v1
👥 Authors: Glenn Anta Bucagu, Thorir Mar Ingolfsson, Yawei Li (possible past Google (United States) affiliation), Luca Benini (possible past Eth Zurich affiliation)
Abstract

Foundation models offer a promising paradigm for Electroencephalography (EEG) analysis, leveraging generalizable representations from vast unlabeled datasets. Yet, Transformer-based architectures face a critical bottleneck: global attention mechanisms couple the attention memory state to the signal duration, causing memory overflow during continuous monitoring. To address this, we introduce S-CEReBrO (Streaming CEReBrO), an evolution of the CEReBrO architecture designed for continuous monitoring...

📄 OneShot: Index-in-Ranking with Neural Scoring for Large-Scale Retrieval
🗓️ Published: 7/29/2026
🔗 http://arxiv.org/abs/2607.27475v1
👥 Authors: Ziwei Li, Shuyao Li, Xufeng Cai, Xue Zou, Yiming Ma, Huiting Lu, Wujie Yan, Zhichen Zhao, Yang Lu (possible past Meta (United States) affiliation), Zhe Wang (possible past Deepmind (United Kingdom) affiliation), Rui Luo, Zhengyu Su, Dan Zhang (possible past Google (United States) affiliation), Ji Liu (possible past Tencent (China) affiliation)
Abstract

In modern recommendation systems, retrieval serves as a primary stage responsible for filtering billions of candidate items down to thousands prior to refined ranking. To make this massive search effective and efficient, the system relies on ranking accuracy and indexing efficiency. However, these two objectives are traditionally misaligned: while the former optimizes for the alignment between ranking predictions and user behavior, the latter optimizes for a structural grouping of item represent...

📄 Scores Are Not Decisions: Cost-Aware Stopping for Tool Acquisition in LLM Agents
🗓️ Published: 7/29/2026
🔗 http://arxiv.org/abs/2607.27083v1
👥 Authors: Yicheng Feng, Yan Zhang, Yan Cheng (possible past Nvidia (United States) affiliation), Wei Qi (possible past Baidu (China) affiliation)
Abstract

As LLM agents increasingly depend on diverse external services such as search engines, databases, and connectors, agent harnesses face a fundamental tool-selection challenge: acquiring too few tools leaves the task under-informed, while too many adds cost, context load, and privacy exposure. Routers and retrievers can rank candidate tools by relevance, but a ranking alone does not determine how many are worth selecting. Existing approaches leave acquisition under heterogeneous costs unaddressed....

*Notable papers are those with at least two authors from a "big" AI/ML lab.