πŸ“„ Notable* Recent AI/ML arXiv Papers

Last updated just now...

πŸ“„ QR-Structured Thermal Triggers for Targeted Semantic Attacks on Infrared Vision-Language Models
πŸ—“οΈ Published: 7/31/2026
πŸ”— http://arxiv.org/abs/2607.29445v1
πŸ‘₯ Authors: Xiang Chen (possible past Tencent (China) affiliation), Yingying Zhao, Chao Li (possible past Baidu (China) affiliation), Jiaju Han, Ben Zhang, Ang Li (possible past Google (United States) affiliation), Jiahuan Long, Yiwei Wei, Jiujiang Guo, Chengyin Hu
Abstract

Infrared vision-language models (IR-VLMs) extend thermal perception to open-vocabulary classification, image captioning, and visual question answering. However, their robustness to structured thermal perturbations and the stability of cross-modal semantic alignment remain insufficiently studied. We propose QR-Structured Thermal Triggers (QR-STT), a stealthy, training-free, black-box framework for targeted semantic steering of IR-VLMs. QR-STT preserves the functional regions of a QR pattern while...

πŸ“„ Beyond Retrieval: Analytic Memory for Multimodal Agents
πŸ—“οΈ Published: 7/31/2026
πŸ”— http://arxiv.org/abs/2607.29440v1
πŸ‘₯ Authors: Zhoujin Tian, Yao Tian, Hao Zhang (possible past Tencent (China) affiliation), Cheng Chen (possible past Google (United States) affiliation), Yakun Li, Lei Zhang, Xiaofang Zhou
Abstract

Long-term multimodal memory must support not only retrieving relevant information but also computing over observations accumulated across interactions. Existing systems largely emphasize \emph{retrieval memory}, organizing interaction histories through summaries and indexes to return query-relevant information at multiple granularities, from high-level abstractions to underlying records. In this paper, we formulate \emph{analytic memory} as a complementary abstraction that organizes recurring mu...

πŸ“„ SeekBrain: An Autonomous Multi-Agent System for Accelerating Neuroscience Discovery
πŸ—“οΈ Published: 7/31/2026
πŸ”— http://arxiv.org/abs/2607.29347v1
πŸ‘₯ Authors: Jiamin Wu, Peishan Xiang, Jingyang Chen, Yuqing Zhu, Yuxi Li, Ling Luo, Qihao Zheng, Jialiang Zu, Yongchao Wu, Mindong Liu, Haitao Wu, Chaofan Hu, Yijie Sun, Yuqi Hang, Yu Zhu, Shuo Li, Yue Fan, Shiyang Feng, Wanghan Xu, Tianlei Zhang (possible past Baidu (China) affiliation), Jie Zhang, Wenlong Zhang, Bo Zhang (possible past Tencent (China) affiliation), Kai Wang, Lei Bai, Mianxin Liu, Wanli Ouyang, Jiulin Du, Chunfeng Song
Abstract

Modern neuroscience relies on integrating multi-scale, multimodal datasets to uncover the neural principles underlying intelligence. However, analytical challenges posed by highly heterogeneous data and fragmented workflows increasingly constrain discoveries. Here we introduce SeekBrain, an autonomous multi-agent framework designed to accelerate neuroscience discovery through domain-grounded hierarchical planning and cross-modal data analysis. SeekBrain dynamically constructs a repertoire of ana...

πŸ“„ CLIFT: Turning Gemini Robotics On-Device into Humanoid Specialists via Non-Invasive Closed-Loop Iterative Fine-Tuning
πŸ—“οΈ Published: 7/31/2026
πŸ”— http://arxiv.org/abs/2607.29172v1
πŸ‘₯ Authors: Yuxin Chen, Hari Srikanth, Nathan Jew, Menglin Wu, Pengcheng Wang, Junli Ren, Masayoshi Tomizuka (possible past University Of California, Berkeley affiliation), Peng Xu (possible past Google (United States) affiliation), Jinyu Xie, Thomas Tian
Abstract

While robot foundation models are growing increasingly capable, the strongest models are typically trained on proprietary data and remain closed-source, limiting downstream users' ability to adapt them to new tasks, embodiments, and deployment settings. Following the LLM community, an emerging access paradigm for closed-weight robot foundation models is the managed supervised fine-tuning (SFT) API, where users submit training data and receive a tuned policy without access to model weights, gradi...

πŸ“„ ActFovea: Runtime Safeguarding for VLA Policies via Spatiotemporal Visual-Action Consistency
πŸ—“οΈ Published: 7/31/2026
πŸ”— http://arxiv.org/abs/2607.29169v1
πŸ‘₯ Authors: Wenda Yu, Tianshi Wang, Fengling Li, Xin Li (possible past Google (United States) affiliation), Jingjing Li (possible past Google (United States) affiliation), Lei Zhu
Abstract

Vision-language-action (VLA) policies achieve strong performance in robotic manipulation but remain vulnerable to runtime disturbances that break the temporal alignment among visual observations, robot states, and executed actions. We introduce ActFovea, a plug-and-play safeguarding framework that detects and mitigates such failures without retraining or modifying the underlying VLA policy. ActFovea uses robot kinematics, proprioceptive states, and recent actions to construct action-conditioned ...

πŸ“„ Multi-Granularity Position Embedding of Graphs via Granular-Ball for Link Prediction
πŸ—“οΈ Published: 7/31/2026
πŸ”— http://arxiv.org/abs/2607.29115v1
πŸ‘₯ Authors: Sen Zhao, Cheng Liu, Shuyin Xia, Zhiyuan Liu (possible past Tsinghua University affiliation), Yi Liu (possible past Google (United States) affiliation), Yi Wang, Wei Wang (possible past University Of Oxford affiliation)
Abstract

Link prediction aims to identify potential or future connections within a given graph structure. Position information is essential for link prediction, as it distinguishes homogeneous nodes through their relative relationships, facilitating the accurate capture of structural patterns and implicit connections. Previous studies derive node positional information as distances to single-granularity landmarks, defined as the centers of homophilic regions, while neglecting the multi-granularity nature...

πŸ“„ Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions
πŸ—“οΈ Published: 7/30/2026
πŸ”— http://arxiv.org/abs/2607.28818v1
πŸ‘₯ Authors: Pranav Narayanan Venkit, Akshara Prabhakar, Yu Li (possible past Tencent (China) affiliation), Daniel Lee, Chien-Sheng Wu (possible past Salesforce (United States) affiliation)
Abstract

As AI companions increasingly mediate repeated social interaction, users may rely on a stable role and shared history, yet locally acceptable replies do not ensure that either persists. We study two observable long-horizon failures: 'persona collapse', the loss of a deployed role, boundaries, values, or style, and 'behavioral drift', the gradual or recurrent erosion of those properties. We introduce ANCHOR, a controlled synthetic audit that separately measures persona enactment and trajectory re...

πŸ“„ WaiT for the Signal: Simple Frequency-Aware Flow-Matching
πŸ—“οΈ Published: 7/30/2026
πŸ”— http://arxiv.org/abs/2607.28760v1
πŸ‘₯ Authors: Krunoslav Lehman Pavasovic, ThΓ©ophane Vallaeys, StΓ©phane Mallat (possible past Google (United States) affiliation), Giulio Biroli, Luke Zettlemoyer (possible past University Of Washington affiliation), Brian Karrer (possible past Meta (United States) affiliation), Jakob Verbeek
Abstract

As image generation models scale to ever higher resolutions, global coherence, local detail, and texture fidelity become critical axes for generation quality. However, standard flow matching treats all spatial frequencies uniformly, ignoring the natural frequency hierarchy where high-frequency bands become indistinguishable from pure noise far earlier than coarse structures. We introduce WaiT, a Wavelet-aware image Transformer that decomposes generation into coarse and fine bands via lossless wa...

πŸ“„ AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
πŸ—“οΈ Published: 7/30/2026
πŸ”— http://arxiv.org/abs/2607.28617v1
πŸ‘₯ Authors: Xiangning Lin, Shenzhe Zhu, Shu Yang, Zhenyu Zhang, Haoqian Zhang, Yipeng Zhao, Chengxuan Qian, Tianwei Wang, Ziheng Zhang, Zhenlong Yuan (possible past Tsinghua University affiliation), Dingcheng Wang, Juncheng Wu, Yuan Si, Jiaxin Liu, Baolong Bi, Robert Mahari, Tobin South, Dazza Greenwood, Zexue He, Rishi Bommasani, Sophia Kazinnik, Andreas Haupt, Samuele Marro, Erik Brynjolfsson, Alex Pentland (possible past Massachusetts Institute Of Technology affiliation), Jiaxin Pei
Abstract

System prompts are instructions configured by developers to govern the behaviors of foundation models in AI applications. They are used throughout commercial AI products, but are rarely disclosed to the public or regulators, creating a serious trust and accountability gap in the wide deployment of AI systems. In this paper, we introduce Artificial Intelligence System Prompt Assurance (AISPA), a user-centric framework for systematically auditing system prompts in AI systems. AISPA examines specif...

πŸ“„ OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
πŸ—“οΈ Published: 7/30/2026
πŸ”— http://arxiv.org/abs/2607.28609v1
πŸ‘₯ Authors: Qiushi Sun, Kanzhi Cheng, Yian Wang, Bowen Yang, Hang Yan, Liheng Chen, Fangzhi Xu, Zichen Ding, Nuo Chen, Jialin Cao, Xingdong Gong, Zehao Li, Kaiming Jin, Xinfeng Yuan, Zhoumianze Liu, Jingyang Gong, Zhangyue Yin, Jiahui Gao, Zhiyong Wu (possible past Tsinghua University affiliation), Tianbao Xie, Jianbing Zhang, Ben Kao, Lingpeng Kong (possible past Google (United States) affiliation)
Abstract

Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the task instruction is central to CUA evaluation, data curation, and reinforcement learning. Neither human-written verifiers nor human annotators can provide such verification at scale, so the field increasingly turns to vision-language models (VLMs) as judges of CUA trajectories. But a fundamental question has long gone...

πŸ“„ A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports
πŸ—“οΈ Published: 7/30/2026
πŸ”— http://arxiv.org/abs/2607.28466v1
πŸ‘₯ Authors: Jia Yu, Yan Zhu, Yili He, Zilong Wang, Xinyang Jiang (possible past Tencent (China) affiliation), Peiyao Fu, Ruijie Yang, Tianyi Chen, Siyuan Li (possible past Tencent (China) affiliation), Zhihua Wang, Fei Wu (possible past Google (United States) affiliation), Quanlin Li, Xian Yang, Pinghong Zhou, Shuo Wang (possible past Nvidia (United States) affiliation)
Abstract

Vision-language models remain underused in colonoscopy despite the rich expert descriptions recorded in routine reports. These reports document lesion appearance, size and location but summarise entire procedures rather than caption individual frames, leaving clinical findings only weakly linked to the corresponding images. Here we develop EndoCLIP, a colonoscopy vision-language foundation model trained on 125,756 lesion-level image-text pairs progressively recovered from 280,476 routine colonos...

πŸ“„ Teffic-Audio: Tell Fact from Fiction
πŸ—“οΈ Published: 7/30/2026
πŸ”— http://arxiv.org/abs/2607.28351v1
πŸ‘₯ Authors: Wan Lin, Li Wang (possible past Tesla (United States) affiliation), Jindong Wang, Kunyu Feng, Zhizheng Wu (possible past University Of Edinburgh affiliation)
Abstract

Speech deepfake detection has expanded in scope with increasingly heterogeneous spoofing mechanisms, including speech synthesis, voice conversion, vocoder reconstruction, and neural-codec resynthesis. The resulting spoofing artifacts can be further shaped by variability in source speech, recording environments, and transmission channels. This variability makes robust generalization across heterogeneous conditions a central requirement for practical detection systems. This report presents Teffic-...

πŸ“„ Paying for Honesty Without Knowing the Truth: Reputation-Penalty Design for LLM Marketplace Agents
πŸ—“οΈ Published: 7/30/2026
πŸ”— http://arxiv.org/abs/2607.28330v1
πŸ‘₯ Authors: Mingdai Yang, Shicheng Fan, Kejing Yu, Duohao Wang, Li Sun, Hao Peng (possible past Tsinghua University affiliation), Philip S. Yu (possible past Tsinghua University affiliation), Zhiwei Liu
Abstract

LLM agents increasingly act as autonomous merchants that write their own product listings, and under competitive pressure, they fabricate attributes to win sales. Even under instructions to be honest, they fabricate attributes in a majority of listings across models. A platform's obvious remedy---verifying each claim against the truth---is unavailable, because it observes only a noisy, biased complaint signal, never the ground truth. We design CARP, a reputation-penalty mechanism with a deadband...

πŸ“„ ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding
πŸ—“οΈ Published: 7/30/2026
πŸ”— http://arxiv.org/abs/2607.28312v1
πŸ‘₯ Authors: Mingkang Dong, Muxin Pu, Jie Li, Bohan Guo, Songruo Chen, Bin Ren, Xu Zheng, Chen Zhao (possible past Stanford University affiliation), Tianwen Qian, Mohamed Elhoseiny (possible past Meta (United States) affiliation), Yuqian Fu
Abstract

Streaming video understanding requires models to continuously retain useful visual evidence before future questions are known. Existing approaches primarily manage the growing visual context according to token importance, temporal redundancy, or segment-level relevance, but rarely organize evidence around objects that persist and evolve over time. Thus, in this paper, we introduce ObjectStream, a training-free framework that treats latent objects as memory anchors for streaming video understandi...

πŸ“„ Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
πŸ—“οΈ Published: 7/30/2026
πŸ”— http://arxiv.org/abs/2607.28227v1
πŸ‘₯ Authors: Hanzhang Zhou, Panrong Tong, Xu Zhang (possible past Tencent (China) affiliation), Quyu Kong, Chenglin Cai, Tianyu Xia, Gongjie Zhang, Jianan Zhang (possible past Google (United States) affiliation), Long Li, Long Chen (possible past Tencent (China) affiliation), Lei Wang (possible past Baidu (China) affiliation), Gaole Dai, Pengxiang Li, Liangyu Chen, Yue Wang, Steven Hoi
Abstract

GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them toward real-world use, we envision agents that operate reliably on real devices, execute workflows across platforms, combine GUI interaction with CLI execution, complete long-horizon tasks, proactively initiate useful services, and autonomously improve their capabilities with minimal human effort. Guided by this vision, we present Qwen-UI-Agent, a real-world centric foundation GUI ag...

πŸ“„ Where and When to Commit: Candidate-Aware Decoding for Diffusion Language Models
πŸ—“οΈ Published: 7/30/2026
πŸ”— http://arxiv.org/abs/2607.28166v1
πŸ‘₯ Authors: Chia-Ming Lee, Ming-Ching Chang (possible past Nvidia (United States) affiliation), Xin Li (possible past Google (United States) affiliation), Yu-Lun Liu, Chih-Chung Hsu
Abstract

Diffusion language models (DLMs) expose a provisional prediction at every denoising step, creating an opportunity for generation-time early exit that stops decoding before the schedule is exhausted. Existing early-exit gates decide termination from fixed-region confidence statistics or schedule-dependent rules, evidence too coarse for a decision that freezes every remaining position at once, so they fire prematurely on long chain-of-thought outputs whose answers stabilize only near the end. Adap...

πŸ“„ Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale
πŸ—“οΈ Published: 7/30/2026
πŸ”— http://arxiv.org/abs/2607.28074v1
πŸ‘₯ Authors: Yash Pandya, Sahil Gupta, Sarthak Harne, Archana Yadav, Kavyansh Chourasia, Hussein Mozannar, Vibhav Vineet (possible past University Of Oxford affiliation), Sara Abdali, Corby Rosset, Yash Lara, Ahmed Awadallah, Ece Kamar (possible past Microsoft (United States) affiliation), Akshay Nambi
Abstract

Computer-use agents learn from what their actions change, so training one needs applications it can act on, break and reset. The applications that matter most are login-gated and stateful, so synthetic environments stand in for them. Recent pipelines generate such environments in bulk, which moves the bottleneck from how many exist to what is inside each one. The returns, we find, come from three properties: how much behavioural depth an environment carries, whether it targets the interaction an...

πŸ“„ Change2Task: From Repository Changes to Executable Coding Agent Tasks and Environments
πŸ—“οΈ Published: 7/30/2026
πŸ”— http://arxiv.org/abs/2607.28591v1
πŸ‘₯ Authors: Haomin Qi, Xingliang Wang, Xuanqi Gao, Baihui Sang, Xin Zhang (possible past Google (United States) affiliation), Minghua Ma, Pengfei Gao, Yu Kang, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang, Qi Zhang (possible past Tencent (China) affiliation)
Abstract

Scaling coding agents requires a continuing supply of executable data for training, benchmarking, and continuous evaluation. Each task must couple a realistic software state with a specification, development tools, and reliable verification. To expand this supply, we present Change2Task, a system grounded in repository history that converts merged pull requests into verified tasks on healthy modern revisions of the same repository. It aligns historical evidence with evolved code, reconstructs ta...

πŸ“„ Why Are GUI Agents Correct but Late? Decode on the Decision-Time Critical Path, Tested with Pre-Compiled Policy Trees
πŸ—“οΈ Published: 7/30/2026
πŸ”— http://arxiv.org/abs/2607.28399v1
πŸ‘₯ Authors: Zihan Dong, Rui Qian (possible past Shanghai Jiao Tong University affiliation), Qishi Zhan, Dongshen Peng, Kaixin Li, Yu Li (possible past Tencent (China) affiliation)
Abstract

Computer-use agents often fail on transient GUI events because they produce the correct action only after the relevant window has already closed. We identify the main cause as expensive autoregressive decoding on the decision-time critical path. We propose Adaptive Anticipatory Policy Trees (AAPT), which eliminates this delay without modifying the underlying model. During idle screen periods, the same frozen multimodal model constructs a bounded conditional policy tree with observable guards, pr...

πŸ“„ Flux-OPD: On-Policy Distillation with Evolving Contexts
πŸ—“οΈ Published: 7/30/2026
πŸ”— http://arxiv.org/abs/2607.28022v1
πŸ‘₯ Authors: Yuran Wang, Zekun Wang, Bohan Zeng, Ruixu Zhang, Wenxuan Liu, Liu Yang (possible past Google (United States) affiliation), Yifan Dai, Yang Shi, Bozhou Li, Chengzhuo Tong, Daili Hua, Yuanxing Zhang, Wentao Zhang (possible past Mila - Quebec Artificial Intelligence Institute affiliation)
Abstract

Large language model training in open-ended domains lacks verifiable rewards, making task preferences difficult to formalize as effective supervision. Contexts can convey such preferences, yet provide little additional supervision once distilled into the student, motivating contexts that evolve with student performance. However, directly using evolving contexts as in-training supervision results in an unstable distillation target and conflicting distributions, requiring mechanisms to stabilize t...

πŸ“„ What Makes Graph Unified? Principles and Generative Sliding-Window Transformer for Graph Foundation Models
πŸ—“οΈ Published: 7/30/2026
πŸ”— http://arxiv.org/abs/2607.27966v1
πŸ‘₯ Authors: Dongxiao He, Siqi Liu (possible past University Of Oxford affiliation), Jitao Zhao, Yawen Li, Yi Wang, Di Jin (possible past Tencent (China) affiliation)
Abstract

Graph Foundation Models (GFMs) have recently emerged as a promising paradigm for general-purpose graph learning, aiming to learn reusable knowledge that generalizes across diverse graph domains and downstream tasks, reducing the need for specific model development. Achieving this goal requires reconciling the substantial heterogeneity in node features, graph structures, and semantic information across domains. Among them, heterogeneous node features constitute a fundamental input-level barrier, ...

*Notable papers are those with at least two authors from a "big" AI/ML lab.