πŸ“„ Notable* Recent AI/ML arXiv Papers

Last updated just now...

πŸ“„ ComBodied Agents: a New Paradigm of Human-Centric Agentic AI
πŸ—“οΈ Published: 8/11/2026
πŸ”— http://arxiv.org/abs/2608.10915v1
πŸ‘₯ Authors: Qianggang Ding (possible past Tsinghua University affiliation), Xingyao Wang, Rui Feng, Zhibin Wang, Feixiang Wang, Kelong Mao, Hao Sun, Zhiyao Luo, Jiankai Tang, Lei Li (possible past Carnegie Mellon University affiliation), Jiadong Guo, Minheng Ni, Weicong Lin, Chenxi Yang, Hongxiang Gao, Zhenghua Chen, Yang Bai, Min Wu, Jun Cheng (possible past Deepmind (United Kingdom) affiliation), Huazhu Fu (possible past Inception Institute Of Artificial Intelligence affiliation), Dacheng Tao, Bang Liu
Abstract

After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication. Yet neither explains whether the person forgot, is confused, has side effects, or deliberately refused, nor what support is appropriate. This reveals a structural gap in Agentic AI: Digital Agents primarily transform software states, while Embodied Agents transform physical states; neither makes a person's evolving state and agency the primary object of modeli...

πŸ“„ MIRA: Medical Image Reflection for Agentic Diagnosis
πŸ—“οΈ Published: 8/11/2026
πŸ”— http://arxiv.org/abs/2608.10827v1
πŸ‘₯ Authors: Shengzhi Wang, Jun Yang (possible past Tsinghua University affiliation), Kai Wu, Xiaozhong Ji, Yiwen Ye, Ziyang Chen, Mingliang Xiong, Wen Fang, Mingqing Liu, Mengyuan Xu, Miaoxuan Shan, Caiyan Liu, Bin He (possible past Baidu (China) affiliation), Qingwen Liu
Abstract

Medical visual agents can use tools to inspect images and retrieve external knowledge, but indiscriminate tool use may introduce noisy or misleading evidence. Reliable diagnosis therefore requires not only acquiring additional observations, but also verifying whether tool actions are necessary and whether the resulting evidence supports the current hypothesis. We introduce MIRA (Medical Image Reflection for Agentic Diagnosis), a medical visual diagnostic framework for autonomous evidence search ...

πŸ“„ Tree-of-Ideas: Automated Research Ideation via Cross-Trajectory Reasoning over Scholarly Evolution
πŸ—“οΈ Published: 8/11/2026
πŸ”— http://arxiv.org/abs/2608.10740v1
πŸ‘₯ Authors: Xun Li (possible past Meta (United States) affiliation), Yiying Yang, Pengtao Li, Xiao Yao, Suyu Liu, Xiaoyang Ye, Ziyu Lu, Yuan Yao (possible past Tsinghua University affiliation), Yangning Li, Yinghui Li, Wenhao Jiang (possible past Tencent (China) affiliation)
Abstract

Effective research ideation requires moving beyond a static understanding of prior work to trace how research problems and solutions evolve across the literature. Existing methods either treat papers as unstructured context or model scholarly evolution as isolated citation chains, overlooking interactions among research trajectories. We propose Tree-of-Ideas (ToI), a two-stage framework. EvoTrace reconstructs branching scholarly trajectories from citations, tracking evolving methods, resolved pr...

πŸ“„ SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information
πŸ—“οΈ Published: 8/11/2026
πŸ”— http://arxiv.org/abs/2608.10692v1
πŸ‘₯ Authors: Junjie Ye, Zhuohui Sheng, Shaofan Liu, Yulun Zhu, Wenjie Fu, Dingwei Zhu, Ming Zhang (possible past Peking University affiliation), Yujiong Shen, Weichao Wang, Xin Zhao, Shihan Dou, Tao Gui, Qi Zhang (possible past Tencent (China) affiliation), Xuanjing Huang, Pluto Zhou
Abstract

Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps) to complete user instructions. However, due to the lack of dedicated benchmarks, their capabilities remain poorly understood. To address this gap, we introduce SPIEval, a human-curated benchmark grounded in five cognitive capabilities (i.e., reasoning, disambiguation, integration, preference inference, and multi-intent...

πŸ“„ SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models
πŸ—“οΈ Published: 8/11/2026
πŸ”— http://arxiv.org/abs/2608.10538v1
πŸ‘₯ Authors: Chenhao Dang, Siyuan Xiong, Conghui He (possible past Tsinghua University affiliation), Weijia Li (possible past Tsinghua University affiliation)
Abstract

Agent skills represent a standardized format for packaging procedural knowledge and domain expertise, serving within agent harness systems as an essential mechanism to continually constrain a language model's behavior space for repeatable, high-quality task execution. However, because strong closed-source models entail high inference costs, current popular agent harnesses, such as Codex and OpenClaw, remain prohibitively expensive when deploying these skills to accomplish real-world tasks. The r...

πŸ“„ Critic-Free Pretraining for Efficient Online Reinforcement Learning Fine-Tuning
πŸ—“οΈ Published: 8/11/2026
πŸ”— http://arxiv.org/abs/2608.10473v1
πŸ‘₯ Authors: Daoyi Li, Yixian Zhang, Chao Yu, Wenbo Ding (possible past Tsinghua University affiliation), Yu Wang (possible past Tsinghua University affiliation)
Abstract

Offline-to-online (O2O) reinforcement learning aims to leverage policies pretrained on static datasets while improving them through online interaction. However, directly reusing an offline-trained critic can hinder online fine-tuning: as the policy and data distribution change rapidly, value estimates inherited from offline training may become misaligned with the online environment, leading to inaccurate policy improvement and inefficient exploration. To address this problem, we introduce \textb...

πŸ“„ From Reasoning Depth to Reasoning Breadth: Evaluating Multi-Point Associative Reasoning in Large Language Models
πŸ—“οΈ Published: 8/11/2026
πŸ”— http://arxiv.org/abs/2608.10444v1
πŸ‘₯ Authors: Si'an Xie, Jiaxun Liu, Biao Yang, Wei Yuan, Fan Yang (possible past Tencent (China) affiliation), Tingting Gao, Ming Wu (possible past Microsoft (United States) affiliation)
Abstract

Large language models (LLMs) have made substantial progress on reasoning tasks that require increasingly long and complex inferential chains. This progress primarily reflects reasoning depth. A complementary and comparatively unexamined capability is reasoning breadth: exploring multiple semantic directions in parallel and integrating the resulting clues into one coherent answer. We introduce MPAR-Bench, a bilingual English-Chinese benchmark that isolates reasoning breadth through multi-point as...

πŸ“„ Hidden in Plain Sight: Diffusion-Based Unrestricted Robotic Attacks on Vision-Language-Action Models
πŸ—“οΈ Published: 8/11/2026
πŸ”— http://arxiv.org/abs/2608.10393v1
πŸ‘₯ Authors: Jiahui Han, Yuhui Yao, Xin Wang (possible past University Of Edinburgh affiliation), Jiafei Cao, Mingxuan Zhang, Danfeng Shan, Huiqi Deng (possible past Shanghai Jiao Tong University affiliation), Guanchu Wang, Xia Hu
Abstract

Vision-Language-Action (VLA) models have shown strong capabilities in controlling robots across diverse manipulation tasks. However, their adversarial robustness remains largely underexplored, and exploiting this weakness can lead to physical-world harm. Existing attacks on VLA models often rely on pixel-space perturbations or white-box access, resulting in noticeable artifacts and limited deployability in real-world robotic systems. In this work, we propose DURA, a diffusion-based unrestricted ...

πŸ“„ Toward a Theory of Value in AI Alignment
πŸ—“οΈ Published: 8/10/2026
πŸ”— http://arxiv.org/abs/2608.10327v1
πŸ‘₯ Authors: Andrew Smart (possible past Google (United States) affiliation), Shazeda Ahmed, Jackie Kay (possible past Deepmind (United Kingdom) affiliation), Jimmy Tobin, Kris Shrishak, Abeba Birhane (possible past Deepmind (United Kingdom) affiliation)
Abstract

Can AI systems be aligned to human values? The popularization of large language models (LLMs) and multi-modal foundation models has seen a rise in harms spanning from toxic speech and hallucinations to AI agents executing unauthorized actions. Within the field of AI safety, these harmful instances are often framed as the alignment problem, or of models being misaligned with human values. Researchers have responded by pursuing applied and theoretical AI value alignment efforts, often without spec...

πŸ“„ FACT: Failure-Aware Causal Training for World-Action Models
πŸ—“οΈ Published: 8/10/2026
πŸ”— http://arxiv.org/abs/2608.10232v1
πŸ‘₯ Authors: Quanquan Peng, Yutong Liang, Rui Yan (possible past Peking University affiliation), Nicklas Hansen, Xiaolong Wang (possible past Carnegie Mellon University affiliation)
Abstract

Recent world-action models (WAMs) show that co-training policies with future prediction can provide physical priors for action generation. Building on the future-prediction ability of video models, many WAMs generate future videos and recover actions with inverse-dynamics models, or use these predicted videos as goal conditions for action generation. In both cases, the world model is trained mostly on successful demonstrations and has little reason to predict the consequences of bad actions. We ...

πŸ“„ Multimodal Model Diffing for Feature Discovery and Control
πŸ—“οΈ Published: 8/10/2026
πŸ”— http://arxiv.org/abs/2608.09928v1
πŸ‘₯ Authors: Hunar Batra, Lachin Naghashyar, Ashkan Khakzar, Philip Torr (possible past University Of Oxford affiliation), Christian Schroeder De Witt (possible past University Of Oxford affiliation), Constantin Venhoff, Ronald Clark
Abstract

Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to identify, audit, or control. While applicable to post-hoc inspection, hidden states that are decomposed into interpretable feature directions using sparse autoencoders (SAEs) neither readily isolate which features are changed by multimodal training, nor are they directly useful for targeted control. We introduce MMDiff, a multimodal model-diffing ...

πŸ“„ SHE: Trajectory-driven Safety Harness Evolution for LLM Agents
πŸ—“οΈ Published: 8/10/2026
πŸ”— http://arxiv.org/abs/2608.09885v1
πŸ‘₯ Authors: Wanying Qu, Qinghua Mao, Yu Li (possible past Tencent (China) affiliation), Jiyao Liu, Xin Zhang (possible past Google (United States) affiliation), Dadi Guo, Yanxu Zhu, Qingyu Liu, Leitao Yuan, Xi Lin, Shanfeng Zhu, Yanwei Fu, Jing Shao, Xia Hu, Dongrui Liu
Abstract

The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, permissions, and runtime control. Existing safety mechanisms often treat the harness as a fixed deployment artifact, limiting their ability to evolve with emerging risks. Moreover, coupled functions across harness components obscure safety responsibility attribution, making localized evolution difficult. We propose Safety Harness Evolution (SHE), a ...

πŸ“„ ArchAgent v2: A Case Study with the Data Prefetching Championship
πŸ—“οΈ Published: 8/10/2026
πŸ”— http://arxiv.org/abs/2608.09874v1
πŸ‘₯ Authors: Abraham Gonzalez, Raghav Gupta (possible past Google (United States) affiliation), Akanksha Jain, Hanna Alam, Alexander Novikov (possible past Google (United States) affiliation), Po-Sen Huang (possible past Google (United States) affiliation), Matej Balog (possible past Deepmind (United Kingdom) affiliation), Marvin Eisenberger, Sergey Shirobokov, NgΓ’n VΕ©, Hank Levy, Borivoje NikoliΔ‡ (possible past University Of California, Berkeley affiliation), Sagar Karandikar, Martin Dixon, Parthasarathy Ranganathan (possible past Google (United States) affiliation)
Abstract

Agentic artificial intelligence has shown great promise in automating algorithm design, but scaling similar techniques to computer microarchitecture discovery remains challenging due to vast search spaces, strict hardware budgets, and long simulation times. In this work, we present ArchAgent v2, a framework which scales automated microarchitecture search to multi-level data prefetching. While the original ArchAgent successfully discovered single-level cache replacement policies in competition se...

πŸ“„ Towards Expert-level Medical AI for Real-time Video Consultations
πŸ—“οΈ Published: 8/10/2026
πŸ”— http://arxiv.org/abs/2608.09861v1
πŸ‘₯ Authors: Mahvish Nagda, Jihyeon Lee, Matthew Thompson, Chunjong Park, Tim Strother, Valentin LiΓ©vin, Roma Ruparel, Akshay Goel, Teya Bergamaschi, Suhana Bedi, Meet Shah, Pavel Dubov, Liviu Panait, Toshiyuki Fukuzawa, Sam Schmidgall, Craig Schiff, Joseph Xu, Aliya Rysbek, Yana Lunts, Jan Freyberg (possible past Google (United States) affiliation), Rebecca Hemengway, Sunny Virmani (possible past Google (United States) affiliation), David Racz, Carey Radebaugh (possible past Google (United States) affiliation), JoΓ«lle Barral (possible past Google (United States) affiliation), Kavi Goel, Dale R. Webster (possible past Google (United States) affiliation), Katherine Chou (possible past Google (United States) affiliation), Avinatan Hassidim (possible past Google (United States) affiliation), Yossi Matias (possible past Google (United States) affiliation), James Manyika, Gregory Wayne (possible past Google (United States) affiliation), Tao Tu (possible past Google (United States) affiliation), Yun Liu (possible past Google (United States) affiliation), Ethan Goh, Christina Chen, Ryutaro Tanno (possible past Google (United States) affiliation), Po-Hsuan Cameron Chen (possible past Google (United States) affiliation), Mike Schaekermann (possible past Google (United States) affiliation), Anil Palepu (possible past Google (United States) affiliation)
Abstract

Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through non-verbal cues. While text-based AI has shown promise, it discards essential perceptual dimensions and limits patients who cannot articulate symptoms in writing. Early efforts to extend medical AI to audio-visual interaction have demonstrated feasibility but not reached clinician-level performance. Here, we provide the first demonstration of ex...

πŸ“„ ReOrder-OPD:Reliability-Aware Prompt Ordering for On-Policy Distillation
πŸ—“οΈ Published: 8/11/2026
πŸ”— http://arxiv.org/abs/2608.10905v1
πŸ‘₯ Authors: Ximo Zhu, Ruiqi Liu, Rong Wang, Ping Wu, Xiang Zheng, Wenzhuo Xu, Xubin Yao, Zhiyuan Yan, Bo Li (possible past Tencent (China) affiliation), Jun Gao (possible past Nvidia (United States) affiliation), Xiaolei Lv
Abstract

On-policy distillation (OPD) applies token-level teacher supervision to student-generated trajectories, but this supervision is not always reliable. Existing methods use local confidence or teacher-student agreement to weight, filter, or truncate the sampled trajectory. These signals do not directly determine whether the teacher can continue a student prefix to a correct answer, and trajectory-level interventions can conflate one rollout's unreliability with low expected training value of its pr...

πŸ“„ MARCO: Click-Intent Decomposition for Calibrated Ads Conversion Prediction
πŸ—“οΈ Published: 8/11/2026
πŸ”— http://arxiv.org/abs/2608.10562v1
πŸ‘₯ Authors: Shiwen Shen, Xiru Huang, Liang Luo, Jianbo Sun, He Lyu, Zihang Fu, Ivonne Xu, Zhizhuo Li, Zhengyu Zhang, Pei-Ju Sung, Yunmiao Wang, Zixuan Wang, Zhengli Zhao (possible past Google (United States) affiliation), Qiang Jin, Mike Jermann, Mingda Li, Yang Xiao, Bhavana Challa, Brooke Bian, Yang Li (possible past Google (United States) affiliation), Ashish Chamoli, Bibek Bhusal, Danning Di, Yuan Jin, Meet Raval, Zhiwen Chen, Boyao Sun, Shuguang Wang, Yunlong He, Yantao Yao, Sagar Chordia, Wenlin Chen (possible past Meta (United States) affiliation), Santanu Kolay, Qin Huang, Ellie Wen
Abstract

Not all clicks are equal. Industrial ads ranking decouples conversion probability into click-through rate (CTR) and post-click conversion rate (CVR), yet treats every click as the same event. In reality, users provide a free, self-generated signal of intent through their physical UI interactions. Different click types on the same ad exhibit a 4-fold difference in actual conversion rates. By conflating these signals, the standard CVR model under-predicts high-intent clicks and over-predicts low-i...

πŸ“„ TideRL: Boosting Agentic RL Goodput with Readiness-Aware Scheduling
πŸ—“οΈ Published: 8/11/2026
πŸ”— http://arxiv.org/abs/2608.10402v1
πŸ‘₯ Authors: Yanyu Ren, Xizheng Wang, Xiao Liu (possible past Baidu (China) affiliation), Bowen Lv, Hanchen Zhang, Shudan Zhang, Hanyu Lai, Shuai Wang, Li Chen, Dan Li, Jie Tang (possible past Tsinghua University affiliation)
Abstract

Reinforcement learning (RL) for large language models is moving toward multi-turn agentic workloads, where rollout tasks repeatedly pause for external environments, resume with growing contexts, and finish at highly variable times. In this setting, RL training goodput, measured by training throughput, matters more than raw GPU occupancy: GPU waiting and repeated prefill recomputation are pure overhead. We present TideRL, a readiness-aware elastic RL system with Continuous Task Batching, Resource...

πŸ“„ Generator-Guided Inverse Sampling for LΓ©vy-Driven Generative Models
πŸ—“οΈ Published: 8/11/2026
πŸ”— http://arxiv.org/abs/2608.10384v1
πŸ‘₯ Authors: Tianfu Qi, Jun Wang (possible past Tencent (China) affiliation), Jun Zhang (possible past Tencent (China) affiliation)
Abstract

This paper studies inverse sampling for LΓ©vy-driven generative models from the perspective of Markov generators. Unlike conventional diffusion models, LΓ©vy-driven dynamics involve infinite jump activities, which makes their reverse process nonlocal and difficult to characterize using score information alone. We address this challenge by analyzing the forward and reversed generators. It is derived that the reversed jump component generally becomes a state-dependent Markov jump process governed by...

πŸ“„ REATS: LLM Reasoning-based Ensemble Learning for Adaptive Time Series Forecasting
πŸ—“οΈ Published: 8/10/2026
πŸ”— http://arxiv.org/abs/2608.10149v1
πŸ‘₯ Authors: Xu Zhang (possible past Tencent (China) affiliation), Chang Xu, Hui Sun, Nan Ma, Zijian Zhang, Peng Wang (possible past Peking University affiliation), Wei Wang (possible past University Of Oxford affiliation), Li Zhao
Abstract

Due to the diversity of real-world time series, no single forecasting model consistently dominates across all samples. Ensemble learning addresses this by combining complementary model strengths, yet existing methods rely on fixed rules or black-box models based solely on numerical inputs, failing to leverage LLM reasoning for interpretable weighting decisions. We propose REATS, which leverages LLM reasoning capabilities as an intelligent ensemble router that jointly processes textual temporal p...

πŸ“„ Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA
πŸ—“οΈ Published: 8/10/2026
πŸ”— http://arxiv.org/abs/2608.09819v1
πŸ‘₯ Authors: Mind Lab, :, Vin Bo, Asher Cai, Jingwei Cao, Song Cao, Vic Cao, Amelia Chen, Andrew Chen, Kaijie Chen, Cleon Cheng, Steven Chiang, Kaixuan Fan, Hera Feng, Huan Feng, Arthur Fu, Jun Gao (possible past Nvidia (United States) affiliation), Pyke Han, Nolan Ho, Ori Hong, Hailee Hou, Piers Hua, Charles Huang, Miles Jiang, Nora Jiang, Yuyi Jiang, Qiuyu Jin, Fancy Kong, Kuss Koo, Jaron Lee, Andrew Lei, Alexy Li, Dawn Li, Lucian Li, Ray Li, Ricardo Li, Smith Li, Theo Li, Allen Lin, Elliot Lin, Fan Lin, Chen Ling, Kairus Liu, Kieran Liu, Logan Liu, Neo Liu, Xiang Liu, Yuxin Lu, Maeve Luo, Pony Ma, Verity Niu, Cole Qiao, Guian Qiu, Vince Qu, Sentry, Niko Song, Vincent Wang, Bo Wu (possible past Tencent (China) affiliation), Rio Yang, Evelyn Ye, Fiona Ye, Ina Ye, Regis Ye, Josh Ying, Atlas Zeng, Danney Zeng, Salmon Zhan, Anya Zhang, Di Zhang, Mia Zhang, Sueky Zhang, Wei Zhao (possible past Tencent (China) affiliation), Ada Zhou, Adrian Zhou, Yuhua Zhou, Juno Zhu, Murphy Zhuang
Abstract

Macaron-V1 is an open agent-model family for experiential intelligence: learning from experience in real environments and continuing to learn after deployment. It is organized around two system goals. Adaptation is pursued through recursive improvement of versioned model-harness pairs, where experience from one configuration is evaluated under an external contract and used to construct its successor. Collaboration is pursued via the Mixture-of-LoRA (MoL) architecture that freezes a base model, c...

πŸ“„ AirFlow: Context Preserving and Multi-Rate State Modeling for Air Quality Forecasting
πŸ—“οΈ Published: 8/10/2026
πŸ”— http://arxiv.org/abs/2608.09775v1
πŸ‘₯ Authors: Fan Yang (possible past Tencent (China) affiliation), Nan Chen, Yijie Dong, Yuchen Zhang (possible past University Of California, Berkeley affiliation), Wei Zhang (possible past Tsinghua University affiliation)
Abstract

Accurate air quality forecasting is essential for public health and urban environmental management, but remains challenging because pollutant channels differ in periodicity and distribution drift, while their concentration trajectories contain both multi-scale dependencies and rapid changes. Recent methods have improved spatial dependency learning and meteorological covariate modeling. However, pollutant channels are still passed through the same normalization rule and temporal backbone, using a...

*Notable papers are those with at least two authors from a "big" AI/ML lab.