πŸ“„ Notable* Recent AI/ML arXiv Papers

Last updated just now...

πŸ“„ WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution
πŸ—“οΈ Published: 8/27/2026
πŸ”— http://arxiv.org/abs/2608.27454v1
πŸ‘₯ Authors: Liyan Tang, Cyrus Rashtchian, Chun-Sung Ferng, Andrew Tomkins (possible past Google (United States) affiliation), Da-Cheng Juan (possible past Google (United States) affiliation), Tu Vu
Abstract

Agent skills package specialized knowledge and workflows into reusable resources that extend AI agent capabilities. Recent work automatically discovers such skills from agent experience, which enables agents to progressively adapt through interaction. However, the insights that guide skill development typically remain scattered across optimization histories, limiting their systematic reuse across iterations. We introduce WikiSkill, a framework that co-evolves agent skills with a persistent knowl...

πŸ“„ PAWBench: How Far Are We from Probabilistically Aligned World Modeling?
πŸ—“οΈ Published: 8/27/2026
πŸ”— http://arxiv.org/abs/2608.27345v1
πŸ‘₯ Authors: Yuandong Pu, Le Zhuo, Sayak Paul, Gabriel Jorge Menezes, Avram ĐorΔ‘eviΔ‡, Shiyang Li, Yifan Zhou, Bin Fu (possible past Tencent (China) affiliation), Wenlong Zhang, Junjun He, Yu Qiao (possible past Shanghai Artificial Intelligence Laboratory affiliation), Yihao Liu, Jingbo Xing, Xi Chen (possible past University Of California, Berkeley affiliation)
Abstract

Recent video generation models are increasingly framed as world models. Many physical processes can unfold in more than one valid way. Therefore, a world model should reproduce not only a plausible trajectory, but also the distribution of possible behaviors under the same initial observation and action. We call this distribution-level requirement probabilistic alignment. However, existing evaluations largely assess individual-video plausibility and do not test whether repeated generations recove...

πŸ“„ What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents
πŸ—“οΈ Published: 8/27/2026
πŸ”— http://arxiv.org/abs/2608.27260v1
πŸ‘₯ Authors: Xingshan Zeng, Zishan Xu, Boju Zhang, Yuzhou Wu, Lingzhi Wang, Jianghao Lin, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang, Weinan Zhang (possible past Shanghai Jiao Tong University affiliation), Yong Yu (possible past Shanghai Jiao Tong University affiliation), Qun Liu (possible past Huawei Technologies (China) affiliation), Weiwen Liu
Abstract

LLM agents increasingly rely on generated interaction data to learn how to interact with external environments. Agentic data generation must maintain consistency among environments, tasks, interactions, and success signals while producing experience that is useful rather than merely abundant. Existing work spans many agent domains, but domain-centered organization and heterogeneous evaluation often obscure common generation mechanisms and conflate candidate construction with verification and sel...

πŸ“„ When Tool Outputs Become Commands: Separating Action Induction from Runtime Authorization in Tool-Augmented LLM Agents
πŸ—“οΈ Published: 8/27/2026
πŸ”— http://arxiv.org/abs/2608.27146v1
πŸ‘₯ Authors: Xiaokun Guo, Zhen Xu (possible past Google (United States) affiliation), Dongdong Huo, Yanqiu Zhang, Wei Wang (possible past University Of Oxford affiliation), Qinfu Yang, Dongjin Yu, Yu Wang (possible past Tsinghua University affiliation)
Abstract

Tool-augmented LLM agents must rely on untrusted runtime Observations to complete open-ended tasks; however, when tool outputs no longer merely provide data but begin to specify concrete actions, they effectively become ``commands'' that can drive real-world side effects beyond user intent. We argue that this risk arises from conflating action induction with execution authorization. To address this distinction, we propose SARA, which treats action induction and execution authorization as distinc...

πŸ“„ From Atomic to Agentic: Towards Interpretable Evaluation of LLMs' Agentic Mathematical Capabilities
πŸ—“οΈ Published: 8/27/2026
πŸ”— http://arxiv.org/abs/2608.26950v1
πŸ‘₯ Authors: Jiayi Kuang, Yinghui Li, Yunze Song, Keyu Chen, Zhifeng Shen, Yangning Li, Yidong Wang, Di Yin, Ruizhi Qiao (possible past Tencent (China) affiliation), Xing Sun (possible past Tencent (China) affiliation), Kai Jin, Ying Shen, Liang Lin, Philip S. Yu (possible past Tsinghua University affiliation)
Abstract

Large Language Models (LLMs) are evolving from performing end-to-end mathematical reasoning to integrating agentic intelligence. However, most existing math benchmarks evaluate only final answers. This outcome-oriented evaluation provides limited diagnostic value for identifying process-level failures or rigorous logic, failing to guide the transformation of LLMs into robust agents. To bridge this gap, we present a process-level benchmark designed to evaluate the inherent agentic mathematical re...

πŸ“„ DEEPCHART: How Far are LLMs from Faithful Data-Science Chart Generation?
πŸ—“οΈ Published: 8/27/2026
πŸ”— http://arxiv.org/abs/2608.26757v1
πŸ‘₯ Authors: Jiahui Tang, Kuicai Dong, Dexun Li, Hongchao Gu, Haocheng Yu, Wei Han (possible past Google (United States) affiliation), Chen Zhang (possible past Peking University affiliation), Yong Liu, Hao Wang (possible past Tsinghua University affiliation), Enhong Chen (possible past Baidu (China) affiliation)
Abstract

Faithful chart generation in real-world data-science workflows requires grounding visualizations in scattered evidence, computing chart-ready quantities, and rendering them accurately. Modern LLMs can produce visually plausible, instruction-compliant charts, yet data-level hallucinations remain difficult to detect in long, noisy, and multimodal contexts. To measure this gap, we introduce DEEPCHART, an expert-annotated benchmark of 1,482 task-conditioned chart-generation instances drawn from real...

πŸ“„ Rethinking Message Passing as Retrieval for Text-Attributed Graph Learning
πŸ—“οΈ Published: 8/27/2026
πŸ”— http://arxiv.org/abs/2608.26732v1
πŸ‘₯ Authors: Jintang Li, Yuhong Chen, Ruofan Wu, Binli Luo, Jiayi Ji, Hui Li (possible past Baidu (China) affiliation), Rongrong Ji (possible past Tencent (China) affiliation)
Abstract

Graph neural networks (GNNs) are typically conceptualized as message-passing neural networks, yet it remains unclear why neighborhood aggregation reliably outperforms node-wise multilayer perceptrons (MLPs). Despite its empirical success, this paradigm can be computationally expensive and sensitive to imperfect graph structures. In this work, we present a retrieval-augmented view of GNNs: each layer makes predictions by applying an MLP to a node representation together with a permutation-invaria...

πŸ“„ LiveVVT: High-Fidelity Video Virtual Try-On in Real Time
πŸ—“οΈ Published: 8/27/2026
πŸ”— http://arxiv.org/abs/2608.26714v1
πŸ‘₯ Authors: Yushe Cao, Shikun Feng (possible past Baidu (China) affiliation), Ruxiang Duan, Liyong Wang, Dianxi Shi, Chun Yu, Junliang Xing (possible past Tsinghua University affiliation)
Abstract

Diffusion-based Video Virtual Try-On (VVT) achieves high visual fidelity through bidirectional spatio-temporal modeling, but complete-clip dependence incurs prohibitive latency and computational overhead in practical continuous deployment. Naively enforcing causality disrupts pretrained bidirectional priors and substantially degrades synthesis quality. We introduce LiveVVT, a rolling streaming diffusion framework that preserves bounded bidirectional modeling within causal recurrent generation. W...

πŸ“„ Accelerating Scientific Research with Gemini in the Real-World
πŸ—“οΈ Published: 8/27/2026
πŸ”— http://arxiv.org/abs/2608.26701v1
πŸ‘₯ Authors: Samuel Schmidgall, Xiaokai Zhu, Marian Shaw, Lin Yang, Valentin LiΓ©vin, Jingyun Yang, Yuchen Zhuang, Tim Strother, Alex Bijamov, Min Woo Sun, Anil Palepu (possible past Google (United States) affiliation), Justin Chen, David Steiner, Jacqueline Shreibati, Wei-Hung Weng, Yilin Zhao, Xingjian Hu, Nicholas Zahn, Sadhya Garg, Julia Kirby, Yuxiang Gan, Jiaoli Li, Divy Thakkar, Shekoofeh Azizi (possible past Google (United States) affiliation), David Racz, Juraj Gottweis (possible past Google (United States) affiliation), Vivek Natarajan (possible past Google (United States) affiliation), Chenglin Wu, Tal Danino, Keran Rong, Haozhe Wang, Benoit Schillings, Yong Cheng (possible past Tsinghua University affiliation), Quoc V. Le (possible past Stanford University affiliation), Tao Tu (possible past Google (United States) affiliation)
Abstract

We present an extension and comprehensive real-world validation of Co-Scientist, a Gemini-based multi-agent system designed to accelerate end-to-end scientific research across hypothesis generation, experimentation, and manuscript generation. Moving beyond in silico hypothesis generation, this specialized configuration transitions Co-Scientist into an execution-grounded research partner advancing closed-loop scientific workflows across materials science, biology, and computer science. In materia...

πŸ“„ DuMateBench: Evaluating Autonomous Agents in Complex Real-World Workflows
πŸ—“οΈ Published: 8/27/2026
πŸ”— http://arxiv.org/abs/2608.26546v1
πŸ‘₯ Authors: Zechun Niu, Yukun Zhao, Jiaxin Zhang, Xu Shen, Jinhua Si, Han Tian, Can Xu (possible past Google (United States) affiliation), Yunfan Song, Jiaxin Mao, Yansong Gao, Yuchen Li, Jianmin Wu, Lingyong Yan, Shuaiqiang Wang (possible past Baidu (China) affiliation), Dawei Yin (possible past Baidu (China) affiliation)
Abstract

Autonomous agents are increasingly adopted to complete complex, multi-tool workflows in real-world settings. However, existing benchmarks typically separate tasks by application or capability and evaluate agents in environments that are cleaner and more stable than those encountered in practice. We introduce DuMateBench, a real-session benchmark reconstructed from anonymized and privacy-screened user sessions collected from a large-scale production agent platform. Each task preserves the relevan...

πŸ“„ Fine-Tuning of Transformer models with Frames
πŸ—“οΈ Published: 8/26/2026
πŸ”— http://arxiv.org/abs/2608.26430v1
πŸ‘₯ Authors: Harshavardhan Adepu, Li Zhang (possible past University Of Oxford affiliation), Sanjiv Kumar (possible past Google (United States) affiliation), Vikas Singh
Abstract

Parameter-Efficient Fine-Tuning (PEFT) strategies such as Low-Rank Adaptation (LoRA) are effective solutions for fine-tuning large-scale pre-trained models; however, their memory requirements scale with the size of the model, $\mathcal{O}(dr)$, where $d$ is the model's hidden dimension and $r$ is the rank. Our proposal, FrameFT, models the parameter update $Ξ”W$ with a sparse coefficient matrix in a Fusion Frame basis. Fusion Frames can be generated algorithmically and shared across model layers,...

πŸ“„ VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning
πŸ—“οΈ Published: 8/26/2026
πŸ”— http://arxiv.org/abs/2608.26105v1
πŸ‘₯ Authors: Junxiang Xu, Ruisi Wang, Fanyi Pu, Maijunxian Wang, Ran Ji, Tongxi Zhou, Chenyang Gu, Jing Zuo, Hongcan Xiao, Yimeng Geng, Wanqi Yin, Wei Chen, Oscar Qian, Zhengan Yan, Ziqi Huang, Haiwen Diao, Liang Pan, Bo Li (possible past Tencent (China) affiliation), Xiangyu Fan, Dezhi Luo, Fengyuan Yu, Zehong Zhao, Qingying Gao, Tinghui Zhu, Yilan Zhang, Jingqi Tong, Pinyuan Feng, Zhengze Jiang, Letian Wang, Ziyu Guo, Renrui Zhang, Jieneng Chen, Sonia Joseph, Constantin Venhoff, Saman Motamed, Mengyue Yang, Chandra Sripada, Alan Yuille (possible past Google (United States) affiliation), Philip Torr (possible past University Of Oxford affiliation), Lvmin Zhang, Vikash Kumar (possible past University Of Washington affiliation), Daniel Khashabi, Nikolaus Kriegeskorte, RaphaΓ«l MilliΓ¨re, Vincent C. MΓΌller, Anyi Rao, Quan Wang (possible past Google (United States) affiliation), Ziwei Liu, Dahua Lin, Lei Yang (possible past Google (United States) affiliation), Hokin Deng, Zhongang Cai
Abstract

Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and videos) are not merely inputs to be understood or outputs to be rendered, but first-class substrates for problem solving beyond language. Yet progress remains bottlenecked by the lack of scalable training tasks, reliable feedback, and controlled comparisons across generative substrates. In this work, we introduce VBVR-Pro, a closed-loop testbed that makes native visual reasoning thr...

πŸ“„ Planetary Prediction Engine: Autonomous Geospatial Prediction via Intelligent Data Selection and Foundation Model Embeddings
πŸ—“οΈ Published: 8/26/2026
πŸ”— http://arxiv.org/abs/2608.26088v1
πŸ‘₯ Authors: Evelyn Ma, Rama Kumar Pasumarthi, Kishwar Shafin (possible past Google (United States) affiliation), Mandar Sharma, Mimi Sun (possible past Google (United States) affiliation), Hamed Sadeghi, Dav M. Ebengo, Mbulayi Onesime, Rouslan Solomakhin, John Wamburu, William Ogallo, Aisha Walcott-Bryant, Sanxing Chen, Arbaaz Muslim, Yael Mayer, Ronald Ho, Roy Lee, Ruth Alcantara, Abdoulaye Diack, Monica Bharel, Lambert Rosique, Jeremy Amez-Droz, Christopher Haire, James Manyika, Yossi Matias (possible past Google (United States) affiliation), Niv Efron, Gautam Prasad (possible past Google (United States) affiliation), Shravya Shetty (possible past Google (United States) affiliation)
Abstract

Addressing critical global challenges, from food security and disaster risk to disease outbreaks and socio-economic vulnerability, demands high-fidelity geospatial modeling. However, building predictive planetary models remains bottlenecked by a fragmented data ecosystem, requiring manual data retrieval, multimodal data curation and fusion along with iterative model selection. We present the Planetary Prediction Engine (PPE), an autonomous AI system that executes this end-to-end workflow directl...

πŸ“„ Prefix Sliding for efficient test-time scaling
πŸ—“οΈ Published: 8/26/2026
πŸ”— http://arxiv.org/abs/2608.26070v1
πŸ‘₯ Authors: Niklas Muennighoff, Zhengyang Wang, Zeyi Chen, Weijia Shi, Binyuan Hui, John Yang, Dapeng Jiang, Mika Senghaas, Fares Obeid, Johannes Hagemann, Sami Jaghouar, Ludwig Schmidt (possible past University Of Washington affiliation), Percy Liang (possible past Stanford University affiliation), Jason Wei (possible past Google (United States) affiliation), Andrew Y. Ng (possible past Stanford University affiliation), Luke Zettlemoyer (possible past University Of Washington affiliation), Yejin Choi (possible past Allen Institute For Artificial Intelligence affiliation), Mike Lewis (possible past Meta (United States) affiliation)
Abstract

Test-time scaling uses extra test-time compute to improve performance, such as letting language models reason longer when solving a problem. As models keep the entire reasoning trace in memory via full attention, hard tasks that need long thinking can be prohibitively expensive. However, we find most intermediate reasoning tokens lose importance as the model continues reasoning. This calls into question whether retaining them is worth the cost. Based on this insight, we propose Prefix Sliding, w...

πŸ“„ AsymSpec: Context-Asymmetric Speculative Decoding for Agentic LLMs
πŸ—“οΈ Published: 8/26/2026
πŸ”— http://arxiv.org/abs/2608.26004v1
πŸ‘₯ Authors: Sheng Liang, Yongyue Zhang, Nathanael Brian, Hang Lv, Hao Wang (possible past Tsinghua University affiliation), Chen Zhang (possible past Peking University affiliation), Yong Liu
Abstract

Agentic LLM pipelines face escalating inference costs as context accumulates across retrieval, tool use, and multi-turn interactions. To control latency, deployments routinely compress inputs, but this degrades task accuracy. Speculative decoding (SD) accelerates generation losslessly, yet it assumes the drafter and verifier share an identical context, preventing SD from resolving the accuracy-overhead trade-off. We propose AsymSpec, an asymmetric speculative decoding framework that breaks this ...

πŸ“„ CEDAR: Controlled and Event-Driven Demand Forecasting via Residual Decomposition
πŸ—“οΈ Published: 8/26/2026
πŸ”— http://arxiv.org/abs/2608.25871v1
πŸ‘₯ Authors: Junjie Meng, Ranxu Zhang, Zi-An Zhang, Shujun Liu, Xiaoning Qi, Xiaozhou Xu, Yanyong Zhang, Hui Xiong (possible past Baidu (China) affiliation), Chao Wang (possible past Google (United States) affiliation)
Abstract

Forecasting in large-scale e-commerce marketplaces is increasingly required to support planning: merchants need to evaluate sales outcomes under future action sequences such as budget schedules, rather than passively predicting what happens next. However, most existing time series forecasting (TSF) approaches remain inherently passive. Even when incorporating operational decisions as auxiliary covariates, they typically optimize for correlation-based extrapolation under historical policies. This...

πŸ“„ A Constitutive Markov Physics-Informed Neural Operator (MPNO) for Autoregressive Stability in Transient Dynamics
πŸ—“οΈ Published: 8/26/2026
πŸ”— http://arxiv.org/abs/2608.25744v1
πŸ‘₯ Authors: Wenpu Du, Peng Zhou (possible past Tencent (China) affiliation), Yunlong Xia, Sinuo Xin, Congcong Zhang, Boyang Zhang, Yi Zhang (possible past Google (United States) affiliation), Wenzheng Xu
Abstract

Neural operators applied to transient-dynamics PDEs with strong discontinuities exhibit autoregressive instability: in concrete-penetration stress-field prediction, the wavelet neural operator (WNO) diverges in autoregressive rollout, while MeshGraphNets collapse to zero predictions. WNO's instability stems from the lack of a structural constraint on the spectral radius of its propagation operator; the Fourier neural operator (FNO) is stable in these measurements but only emergently, not by cons...

πŸ“„ TRACE: Retrospective Streaming Generation of Physical Fields under Sparse Structured Sensing
πŸ—“οΈ Published: 8/26/2026
πŸ”— http://arxiv.org/abs/2608.26219v1
πŸ‘₯ Authors: Xinyu Zhang (possible past Baidu (China) affiliation), Lihao Chen, Panqi Chen, Lei Cheng, Ting Zhang (possible past Meta (United States) affiliation), Jianlong Li, Shikai Fang
Abstract

Reconstructing continuous physical fields from sparse measurements is central to scientific monitoring, inverse modeling, and digital-twin construction. Generative reconstruction has recently emerged as a promising paradigm for this task by learning data-driven physical priors that complete plausible full fields from limited observations. However, existing methods largely assume fixed, batch conditioning, whereas real sensing systems often produce structured streams: probes scan local regions, i...

*Notable papers are those with at least two authors from a "big" AI/ML lab.