πŸ“„ Notable* Recent AI/ML arXiv Papers

Last updated just now...

πŸ“„ Multimodal Model Diffing for Feature Discovery and Control
πŸ—“οΈ Published: 8/10/2026
πŸ”— http://arxiv.org/abs/2608.09928v1
πŸ‘₯ Authors: Hunar Batra, Lachin Naghashyar, Ashkan Khakzar, Philip Torr (possible past University Of Oxford affiliation), Christian Schroeder De Witt (possible past University Of Oxford affiliation), Constantin Venhoff, Ronald Clark
Abstract

Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to identify, audit, or control. While applicable to post-hoc inspection, hidden states that are decomposed into interpretable feature directions using sparse autoencoders (SAEs) neither readily isolate which features are changed by multimodal training, nor are they directly useful for targeted control. We introduce MMDiff, a multimodal model-diffing ...

πŸ“„ SHE: Trajectory-driven Safety Harness Evolution for LLM Agents
πŸ—“οΈ Published: 8/10/2026
πŸ”— http://arxiv.org/abs/2608.09885v1
πŸ‘₯ Authors: Wanying Qu, Qinghua Mao, Yu Li (possible past Tencent (China) affiliation), Jiyao Liu, Xin Zhang (possible past Google (United States) affiliation), Dadi Guo, Yanxu Zhu, Qingyu Liu, Leitao Yuan, Xi Lin, Shanfeng Zhu, Yanwei Fu, Jing Shao, Xia Hu, Dongrui Liu
Abstract

The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, permissions, and runtime control. Existing safety mechanisms often treat the harness as a fixed deployment artifact, limiting their ability to evolve with emerging risks. Moreover, coupled functions across harness components obscure safety responsibility attribution, making localized evolution difficult. We propose Safety Harness Evolution (SHE), a ...

πŸ“„ ArchAgent v2: A Case Study with the Data Prefetching Championship
πŸ—“οΈ Published: 8/10/2026
πŸ”— http://arxiv.org/abs/2608.09874v1
πŸ‘₯ Authors: Abraham Gonzalez, Raghav Gupta (possible past Google (United States) affiliation), Akanksha Jain, Hanna Alam, Alexander Novikov (possible past Google (United States) affiliation), Po-Sen Huang (possible past Google (United States) affiliation), Matej Balog (possible past Deepmind (United Kingdom) affiliation), Marvin Eisenberger, Sergey Shirobokov, NgΓ’n VΕ©, Hank Levy, Borivoje NikoliΔ‡ (possible past University Of California, Berkeley affiliation), Sagar Karandikar, Martin Dixon, Parthasarathy Ranganathan (possible past Google (United States) affiliation)
Abstract

Agentic artificial intelligence has shown great promise in automating algorithm design, but scaling similar techniques to computer microarchitecture discovery remains challenging due to vast search spaces, strict hardware budgets, and long simulation times. In this work, we present ArchAgent v2, a framework which scales automated microarchitecture search to multi-level data prefetching. While the original ArchAgent successfully discovered single-level cache replacement policies in competition se...

πŸ“„ Towards Expert-level Medical AI for Real-time Video Consultations
πŸ—“οΈ Published: 8/10/2026
πŸ”— http://arxiv.org/abs/2608.09861v1
πŸ‘₯ Authors: Mahvish Nagda, Jihyeon Lee, Matthew Thompson, Chunjong Park, Tim Strother, Valentin LiΓ©vin, Roma Ruparel, Akshay Goel, Teya Bergamaschi, Suhana Bedi, Meet Shah, Pavel Dubov, Liviu Panait, Toshiyuki Fukuzawa, Sam Schmidgall, Craig Schiff, Joseph Xu, Aliya Rysbek, Yana Lunts, Jan Freyberg (possible past Google (United States) affiliation), Rebecca Hemengway, Sunny Virmani (possible past Google (United States) affiliation), David Racz, Carey Radebaugh (possible past Google (United States) affiliation), JoΓ«lle Barral (possible past Google (United States) affiliation), Kavi Goel, Dale R. Webster (possible past Google (United States) affiliation), Katherine Chou (possible past Google (United States) affiliation), Avinatan Hassidim (possible past Google (United States) affiliation), Yossi Matias (possible past Google (United States) affiliation), James Manyika, Gregory Wayne (possible past Google (United States) affiliation), Tao Tu (possible past Google (United States) affiliation), Yun Liu (possible past Google (United States) affiliation), Ethan Goh, Christina Chen, Ryutaro Tanno (possible past Google (United States) affiliation), Po-Hsuan Cameron Chen (possible past Google (United States) affiliation), Mike Schaekermann (possible past Google (United States) affiliation), Anil Palepu (possible past Google (United States) affiliation)
Abstract

Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through non-verbal cues. While text-based AI has shown promise, it discards essential perceptual dimensions and limits patients who cannot articulate symptoms in writing. Early efforts to extend medical AI to audio-visual interaction have demonstrated feasibility but not reached clinician-level performance. Here, we provide the first demonstration of ex...

πŸ“„ AirFlow: Context Preserving and Multi-Rate State Modeling for Air Quality Forecasting
πŸ—“οΈ Published: 8/10/2026
πŸ”— http://arxiv.org/abs/2608.09775v1
πŸ‘₯ Authors: Fan Yang (possible past Tencent (China) affiliation), Nan Chen, Yijie Dong, Yuchen Zhang (possible past University Of California, Berkeley affiliation), Wei Zhang (possible past Tsinghua University affiliation)
Abstract

Accurate air quality forecasting is essential for public health and urban environmental management, but remains challenging because pollutant channels differ in periodicity and distribution drift, while their concentration trajectories contain both multi-scale dependencies and rapid changes. Recent methods have improved spatial dependency learning and meteorological covariate modeling. However, pollutant channels are still passed through the same normalization rule and temporal backbone, using a...

πŸ“„ Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models
πŸ—“οΈ Published: 8/10/2026
πŸ”— http://arxiv.org/abs/2608.09666v1
πŸ‘₯ Authors: Shulin Tian, Ziqi Huang, Fan Zhang, Hongyuan Zhu (possible past Tencent (China) affiliation), Yu Qiao (possible past Shanghai Artificial Intelligence Laboratory affiliation), Ziwei Liu
Abstract

Recent advances in visual generative models have enabled high-quality image and video generation, but evaluating these models often demands sampling hundreds or thousands of images or videos, which is computationally expensive. Existing evaluation methods also rely on rigid pipelines that overlook specific user needs and provide numerical results without clear explanations. Mimicking how humans quickly form impressions of a model's capabilities from only a few samples, we propose the Evaluation ...

πŸ“„ From Semantic Grounding to Decision Optimization: A Unified Framework for Long-Horizon UAV Vision-Language Navigation
πŸ—“οΈ Published: 8/10/2026
πŸ”— http://arxiv.org/abs/2608.09564v1
πŸ‘₯ Authors: Zeyuan Ma, Jiaxin Chen (possible past Inception Institute Of Artificial Intelligence affiliation), Di Huang (possible past Google (United States) affiliation)
Abstract

UAV vision-language navigation (UAV-VLN) focuses on enabling an aerial agent to follow natural-language instructions in open 3D environments from egocentric visual observations. Current approaches suffer from three coupled issues: weak grounding of instruction-relevant landmarks in visual observations, insufficient exploitation of long-horizon history, and unstable decisions under local traps or repeated exploration. To address these issues, we propose a unified semantic-to-decision framework. F...

πŸ“„ TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability
πŸ—“οΈ Published: 8/10/2026
πŸ”— http://arxiv.org/abs/2608.09538v1
πŸ‘₯ Authors: Vincent Cohen-Addad, Dimitris Paparas, Ernest Van Wijland, Max Springer, Julien Canitrot-Paradis, Honghao Lin, David Woodruff, Adarsh Kumarappan, Rajesh Jayaram, Rudrajit Das, Lalit Jain, Ola Svensson, Silvio Lattanzi (possible past Google (United States) affiliation), Mislav Balunovic, Theophane Weber, Vahab Mirrokni (possible past Google (United States) affiliation)
Abstract

We introduce TCS-Bench, a benchmark for evaluating Large Language Models (LLMs) on research-level Theoretical Computer Science (TCS) proof generation. TCS-Bench consists of theorem-proving tasks from papers published at top theoretical computer science venues (STOC, FOCS, and SODA). Each task provides the necessary context to derive a self-contained proof for a target result. We evaluate state-of-the-art models on this benchmark. We verify the correctness of generated proofs via a verification a...

πŸ“„ Carnot: Interpretable, Interactive, and Optimized Execution of Deep Research Queries
πŸ—“οΈ Published: 8/10/2026
πŸ”— http://arxiv.org/abs/2608.09532v1
πŸ‘₯ Authors: Matthew Russo, Yash Agarwal, Tianyu Li, Zhuohan Gu, Michael Cafarella (possible past University Of Washington affiliation), Omar Khattab, Tim Kraska (possible past Massachusetts Institute Of Technology affiliation), Samuel Madden (possible past Massachusetts Institute Of Technology affiliation)
Abstract

Enterprises increasingly seek to query data lakes using natural language via AI-driven tools like semantic operators or deep research agents. However, the latter operates as an opaque black box, hiding its intermediate reasoning and data retrieval steps, and failing to expose controls for managing API costs and execution latency. Meanwhile, the former can be prohibitively expensive for enterprise-scale data lakes. Consequently, analysts using these systems lack the agency to intercept hallucinat...

πŸ“„ VeinCast: Physics-Guided Dynamic Field Graphs with Graph-Conditioned Fusion for Global Medium-Range Weather Forecasting
πŸ—“οΈ Published: 8/10/2026
πŸ”— http://arxiv.org/abs/2608.09286v1
πŸ‘₯ Authors: Zhisheng Chen, Jinhan Li, Yuxuan Li, Yuan Gao (possible past Tencent (China) affiliation), Hao Wu (possible past Tencent (China) affiliation), Zheng Lu, Jinlong Du, Kun Wang, Bo An
Abstract

Global medium-range weather forecasting requires modeling structured yet state-dependent interactions among heterogeneous atmospheric fields. Existing data-driven models largely learn these interactions implicitly, whereas equation-level physical constraints may inherit approximation and model-form biases. We present VeinCast, a physics-guided dynamic field graph and graph-conditioned fusion framework that jointly forecasts 69 surface and upper-air fields. Within each local window, its Physics-G...

πŸ“„ Emotion2Skill: Model-Internal Emotion Signals for Adaptive Skill Selection and Evolution
πŸ—“οΈ Published: 8/10/2026
πŸ”— http://arxiv.org/abs/2608.09248v1
πŸ‘₯ Authors: Bohan Lin, Hejia Geng, Xinyi Xie, Heng Zhou, Qinghua Xing, Bo Liu (possible past Meta (United States) affiliation), Chen Zhang (possible past Peking University affiliation), Yudong Zhang
Abstract

Skill-based LLM agents select reusable procedures from an external library to solve complex tasks, yet their routing decisions rely entirely on text-level signals such as task descriptions, verbal reflections, and experience-derived rules, while the model's own internal representational state remains unobserved. Recent interpretability work has shown that LLMs maintain linear emotion representations that causally influence behavior; however, these representations have been exploited only for pos...

πŸ“„ RL-Native Distillation: Exploiting Scored Trajectories for Few-Step Image Generation
πŸ—“οΈ Published: 8/10/2026
πŸ”— http://arxiv.org/abs/2608.09226v1
πŸ‘₯ Authors: Yuhan Li, Fangao Zeng, Sicong Kang, Mengfei Xu, Hao Zhou, Wei Li (possible past Peking University affiliation), Pipei Huang, Bingbing Ni (possible past Shanghai Jiao Tong University affiliation)
Abstract

Efficient text-to-image generation requires both reinforcement-learning (RL)-based reward alignment and few-step distillation, yet these procedures are typically performed sequentially, increasing training cost and risking the loss of reward gains during compression. We instead take an RL-native perspective: diffusion RL already generates reward-scored finite-step trajectories, whose intermediate states provide a natural source of distillation supervision rather than a disposable byproduct of sa...

πŸ“„ SiriusDeliver: Automating Data Warehouse Delivery at Tencent
πŸ—“οΈ Published: 8/10/2026
πŸ”— http://arxiv.org/abs/2608.09185v1
πŸ‘₯ Authors: Haining Xie, Xiaokai Zhou, Jiaming Yang, Siqi Shen, Ziwei Wang, Yifeng Zheng, Tengyue Xu, Yipeng Shi, Zefang Zong, Yang Li (possible past Google (United States) affiliation), Peng Chen (possible past Tencent (China) affiliation), Jie Jiang (possible past Tencent (China) affiliation), Debiao He, Xiao Yan, Jiawei Jiang (possible past Tencent (China) affiliation)
Abstract

Enterprise data warehouses (DWs) support business-critical analytics, but warehouse task delivery remains a complicated production process involving context retrieval, workflow configuration, code generation, platform submission, and failure diagnosis. Although large language models (LLMs) and coding agents have improved software development, they are insufficient for production DW delivery, which requires dependency-aware orchestration, lifecycle-aware artifact control, and continuous adaptatio...

πŸ“„ From Relevance to Execution Utility: Reward-Aware Dynamic Execution Gating for Skill-Based LLM Agents
πŸ—“οΈ Published: 8/10/2026
πŸ”— http://arxiv.org/abs/2608.09168v1
πŸ‘₯ Authors: Liang He, Jingbo Wen, Hongyu Gu, Hao Li (possible past Tsinghua University affiliation), Haoyu Wang (possible past Tencent (China) affiliation), Yixiong Chen, Kangning Cui, Xilu Wang
Abstract

Agent skills are increasingly used to equip large language model (LLM) agents with reusable procedural knowledge. Although recent work has substantially improved skill retrieval due to the increasing skill libraries, retrieving a plausible skill bundle does not guarantee that executing it is worthwhile. Since every skill-conditioned rollout is computationally expensive, deciding whether a retrieved bundle should be executed has become an increasingly important challenge. To this end, we introduc...

πŸ“„ From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs
πŸ—“οΈ Published: 8/10/2026
πŸ”— http://arxiv.org/abs/2608.09158v1
πŸ‘₯ Authors: Yuanhe Zhang, Weiliu Wang, Jie Ren (possible past Google (United States) affiliation), Liang Lin, Zhenhong Zhou, Haoran Gao, Kun Wang, Chen Li (possible past Tencent (China) affiliation), Li Sun, Sen Su
Abstract

Large audio-language models (LALMs) have demonstrated strong capabilities in understanding diverse audio inputs. This diversity includes low-frequency signals that are inaudible to humans but can still enter the model and influence its generation. However, the practical impact of such low-frequency inputs on LALMs remains largely unexplored. In this paper, we propose Intermittent Low-Frequency Lockout (ILL), an inaudible red teaming method that evaluates this risk using a universal waveform temp...

πŸ“„ MARA: Flow-Matching-Guided Multi-Agent Resource Allocation for Computational Resource Efficient Learning
πŸ—“οΈ Published: 8/10/2026
πŸ”— http://arxiv.org/abs/2608.09130v1
πŸ‘₯ Authors: Hanye Zhao, Muning Wen, Yong Yu (possible past Shanghai Jiao Tong University affiliation), Weinan Zhang (possible past Shanghai Jiao Tong University affiliation)
Abstract

Allocating limited computation among concurrent learning tasks is difficult when each task must reach a target loss before a deadline but its required training effort is unknown. Existing approaches combine online loss prediction with adaptive resource allocation, yet commonly treat computation as continuously divisible throughput. We instead study a practical setting in which tasks arrive over time and computation is provided by discrete nodes. This setting introduces both uncertain demand and ...

πŸ“„ A Unified Issue Resolution Benchmark for Requirement Clarification, Planning, and Code Generation for Coding Agents
πŸ—“οΈ Published: 8/10/2026
πŸ”— http://arxiv.org/abs/2608.09072v1
πŸ‘₯ Authors: Xin Zhou (possible past Stanford University affiliation), Chun Yong Chong, Kisub Kim, Yun Peng, Rui Shu, Zihan Wu (possible past Tencent (China) affiliation), Xu Han (possible past Tsinghua University affiliation), Guowen Yuan, Zeyang Zhuang, Jounghoon Kim, Jeongjin Ju, Seongmin Ju, Taein Yoon, David Lo (possible past Stanford University affiliation)
Abstract

Large language model-powered coding agents are increasingly used to modify existing code repositories, for example, by adding features or fixing bugs. Yet existing repository-level benchmarks typically evaluate only whether the final patch passes tests. Satisfying a user request requires a long chain of interdependent reasoning and decisions: an agent must recover explicit and implicit requirements, formulate a repository-grounded implementation plan, and translate it into correct code. A pass/f...

πŸ“„ How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review
πŸ—“οΈ Published: 8/10/2026
πŸ”— http://arxiv.org/abs/2608.08975v1
πŸ‘₯ Authors: Ming Li, Chenguang Wang (possible past Amazon (United States) affiliation), Xirui Li, Xinyue Zeng, Dianqi Li, Peng Shi, Dawei Zhou, Tianyi Zhou (possible past University Of Washington affiliation)
Abstract

As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how these effects vary across evaluation conditions. We construct a controlled corpus of 4,200 full-paper manuscripts derived from 120 anonymized ICLR 2026 submissions. Two LLM rewriters transform six rhetorical dimensions in opposing directions, and five LLM reviewers eval...

πŸ“„ Idea Search: Guiding Tree Search with Ideas to Explore Diverse Scientific Methods
πŸ—“οΈ Published: 8/9/2026
πŸ”— http://arxiv.org/abs/2608.08958v1
πŸ‘₯ Authors: Xuefei Julie Wang, Hao Cui, Michael P. Brenner (possible past Google (United States) affiliation), Subhashini Venugopalan (possible past Google (United States) affiliation)
Abstract

Tree Search-based test-time scaling of LLMs is a powerful tool for automated scientific coding. However, pure Tree Search sometimes struggles with systematic exploration, becoming trapped in local optima, or unproductive loops, especially in the vast search space of scientific methods. To address this limitation, we propose Idea Search, a framework that systematically integrates a dynamic "Idea Bank" into Tree Search. Idea Search involves three steps: (1) decomposing existing methods into atomic...

πŸ“„ Full-bandwidth transformer
πŸ—“οΈ Published: 8/9/2026
πŸ”— http://arxiv.org/abs/2608.08888v1
πŸ‘₯ Authors: Xi Wang (possible past Tsinghua University affiliation), Ziyang Cai, Zheng Zhan, Harry Dong, Ying Fan, Gustavo De Rosa, Tim Pearce, John Langford (possible past Microsoft (United States) affiliation)
Abstract

Autoregressive transformers compute along two axes: horizontally across generated tokens, and vertically through model depth. Dense attention gives each token broad horizontal access to the past, but the vertical feedback channel between decoding steps remains narrow: only the sampled token returns to the bottom of the stack, while the top-layer hidden state is discarded. We introduce the \emph{full-bandwidth transformer}, which widens this channel with \emph{latent feedback}: at each decoding s...

πŸ“„ Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA
πŸ—“οΈ Published: 8/10/2026
πŸ”— http://arxiv.org/abs/2608.09819v1
πŸ‘₯ Authors: Mind Lab, :, Vin Bo, Asher Cai, Jingwei Cao, Song Cao, Vic Cao, Amelia Chen, Andrew Chen, Kaijie Chen, Cleon Cheng, Steven Chiang, Kaixuan Fan, Hera Feng, Huan Feng, Arthur Fu, Jun Gao (possible past Nvidia (United States) affiliation), Pyke Han, Nolan Ho, Ori Hong, Hailee Hou, Piers Hua, Charles Huang, Miles Jiang, Nora Jiang, Yuyi Jiang, Qiuyu Jin, Fancy Kong, Kuss Koo, Jaron Lee, Andrew Lei, Alexy Li, Dawn Li, Lucian Li, Ray Li, Ricardo Li, Smith Li, Theo Li, Allen Lin, Elliot Lin, Fan Lin, Chen Ling, Kairus Liu, Kieran Liu, Logan Liu, Neo Liu, Xiang Liu, Yuxin Lu, Maeve Luo, Pony Ma, Verity Niu, Cole Qiao, Guian Qiu, Vince Qu, Sentry, Niko Song, Vincent Wang, Bo Wu (possible past Tencent (China) affiliation), Rio Yang, Evelyn Ye, Fiona Ye, Ina Ye, Regis Ye, Josh Ying, Atlas Zeng, Danney Zeng, Salmon Zhan, Anya Zhang, Di Zhang, Mia Zhang, Sueky Zhang, Wei Zhao (possible past Tencent (China) affiliation), Ada Zhou, Adrian Zhou, Yuhua Zhou, Juno Zhu, Murphy Zhuang
Abstract

Macaron-V1 is an open agent-model family for experiential intelligence: learning from experience in real environments and continuing to learn after deployment. It is organized around two system goals. Adaptation is pursued through recursive improvement of versioned model-harness pairs, where experience from one configuration is evaluated under an external contract and used to construct its successor. Collaboration is pursued via the Mixture-of-LoRA (MoL) architecture that freezes a base model, c...

πŸ“„ Beyond Binary: Continuous State Optimization with Graph-Structured Objectives
πŸ—“οΈ Published: 8/10/2026
πŸ”— http://arxiv.org/abs/2608.09366v1
πŸ‘₯ Authors: Corinna Cortes (possible past Google (United States) affiliation), Yishay Mansour (possible past Google (United States) affiliation), Mehryar Mohri (possible past Google (United States) affiliation)
Abstract

Large-scale learning systems often face the challenge of balancing multiple, potentially competing objectives, such as fairness, accuracy, and latency. While recent work has formalized this as an optimization problem over binary states, many real-world control parameters, such as fairness thresholds, diversity mixing rates, or resource budgets, are continuous. In this work, we extend the framework to \emph{continuous state spaces}. We model the problem as minimizing a sum of linear o...

πŸ“„ Transfer Learning-Enabled Distortion Compensation for Amplitude-Phase-Time Block Modulation-Based Nonlinear Single-Carrier Wireless Communications
πŸ—“οΈ Published: 8/9/2026
πŸ”— http://arxiv.org/abs/2608.08554v1
πŸ‘₯ Authors: Guoxing Duan, Min Fan, Cheng Yi (possible past Google (United States) affiliation), Bensheng Yang, Wei Xu (possible past Tencent (China) affiliation), Haiming Wang, Xiaohu You
Abstract

Power amplifier (PA) nonlinearity and memory effects significantly limit the spectral compliance, reliability, and energy efficiency of communication systems. To address this, we propose a transfer-learning-enabled, fully digital transceiver-cooperative method for amplitude-phase-time block modulation (APTBM)-based nonlinear single-carrier transmission under adjacent channel leakage ratio (ACLR) constraints. At the transmitter, iterative clipping and filtering (ICAF) and static digital pre-disto...

*Notable papers are those with at least two authors from a "big" AI/ML lab.