πŸ“„ Notable* Recent AI/ML arXiv Papers

Last updated just now...

πŸ“„ VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning
πŸ—“οΈ Published: 8/26/2026
πŸ”— http://arxiv.org/abs/2608.26105v1
πŸ‘₯ Authors: Junxiang Xu, Ruisi Wang, Fanyi Pu, Maijunxian Wang, Ran Ji, Tongxi Zhou, Chenyang Gu, Jing Zuo, Hongcan Xiao, Yimeng Geng, Wanqi Yin, Wei Chen, Oscar Qian, Zhengan Yan, Ziqi Huang, Haiwen Diao, Liang Pan, Bo Li (possible past Tencent (China) affiliation), Xiangyu Fan, Dezhi Luo, Fengyuan Yu, Zehong Zhao, Qingying Gao, Tinghui Zhu, Yilan Zhang, Jingqi Tong, Pinyuan Feng, Zhengze Jiang, Letian Wang, Ziyu Guo, Renrui Zhang, Jieneng Chen, Sonia Joseph, Constantin Venhoff, Saman Motamed, Mengyue Yang, Chandra Sripada, Alan Yuille (possible past Google (United States) affiliation), Philip Torr (possible past University Of Oxford affiliation), Lvmin Zhang, Vikash Kumar (possible past University Of Washington affiliation), Daniel Khashabi, Nikolaus Kriegeskorte, RaphaΓ«l MilliΓ¨re, Vincent C. MΓΌller, Anyi Rao, Quan Wang (possible past Google (United States) affiliation), Ziwei Liu, Dahua Lin, Lei Yang (possible past Google (United States) affiliation), Hokin Deng, Zhongang Cai
Abstract

Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and videos) are not merely inputs to be understood or outputs to be rendered, but first-class substrates for problem solving beyond language. Yet progress remains bottlenecked by the lack of scalable training tasks, reliable feedback, and controlled comparisons across generative substrates. In this work, we introduce VBVR-Pro, a closed-loop testbed that makes native visual reasoning thr...

πŸ“„ Planetary Prediction Engine: Autonomous Geospatial Prediction via Intelligent Data Selection and Foundation Model Embeddings
πŸ—“οΈ Published: 8/26/2026
πŸ”— http://arxiv.org/abs/2608.26088v1
πŸ‘₯ Authors: Evelyn Ma, Rama Kumar Pasumarthi, Kishwar Shafin (possible past Google (United States) affiliation), Mandar Sharma, Mimi Sun (possible past Google (United States) affiliation), Hamed Sadeghi, Dav M. Ebengo, Mbulayi Onesime, Rouslan Solomakhin, John Wamburu, William Ogallo, Aisha Walcott-Bryant, Sanxing Chen, Arbaaz Muslim, Yael Mayer, Ronald Ho, Roy Lee, Ruth Alcantara, Abdoulaye Diack, Monica Bharel, Lambert Rosique, Jeremy Amez-Droz, Christopher Haire, James Manyika, Yossi Matias (possible past Google (United States) affiliation), Niv Efron, Gautam Prasad (possible past Google (United States) affiliation), Shravya Shetty (possible past Google (United States) affiliation)
Abstract

Addressing critical global challenges, from food security and disaster risk to disease outbreaks and socio-economic vulnerability, demands high-fidelity geospatial modeling. However, building predictive planetary models remains bottlenecked by a fragmented data ecosystem, requiring manual data retrieval, multimodal data curation and fusion along with iterative model selection. We present the Planetary Prediction Engine (PPE), an autonomous AI system that executes this end-to-end workflow directl...

πŸ“„ Prefix Sliding for efficient test-time scaling
πŸ—“οΈ Published: 8/26/2026
πŸ”— http://arxiv.org/abs/2608.26070v1
πŸ‘₯ Authors: Niklas Muennighoff, Zhengyang Wang, Zeyi Chen, Weijia Shi, Binyuan Hui, John Yang, Dapeng Jiang, Mika Senghaas, Fares Obeid, Johannes Hagemann, Sami Jaghouar, Ludwig Schmidt (possible past University Of Washington affiliation), Percy Liang (possible past Stanford University affiliation), Jason Wei (possible past Google (United States) affiliation), Andrew Y. Ng (possible past Stanford University affiliation), Luke Zettlemoyer (possible past University Of Washington affiliation), Yejin Choi (possible past Allen Institute For Artificial Intelligence affiliation), Mike Lewis (possible past Meta (United States) affiliation)
Abstract

Test-time scaling uses extra test-time compute to improve performance, such as letting language models reason longer when solving a problem. As models keep the entire reasoning trace in memory via full attention, hard tasks that need long thinking can be prohibitively expensive. However, we find most intermediate reasoning tokens lose importance as the model continues reasoning. This calls into question whether retaining them is worth the cost. Based on this insight, we propose Prefix Sliding, w...

πŸ“„ AsymSpec: Context-Asymmetric Speculative Decoding for Agentic LLMs
πŸ—“οΈ Published: 8/26/2026
πŸ”— http://arxiv.org/abs/2608.26004v1
πŸ‘₯ Authors: Sheng Liang, Yongyue Zhang, Nathanael Brian, Hang Lv, Hao Wang (possible past Tsinghua University affiliation), Chen Zhang (possible past Peking University affiliation), Yong Liu
Abstract

Agentic LLM pipelines face escalating inference costs as context accumulates across retrieval, tool use, and multi-turn interactions. To control latency, deployments routinely compress inputs, but this degrades task accuracy. Speculative decoding (SD) accelerates generation losslessly, yet it assumes the drafter and verifier share an identical context, preventing SD from resolving the accuracy-overhead trade-off. We propose AsymSpec, an asymmetric speculative decoding framework that breaks this ...

πŸ“„ Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models
πŸ—“οΈ Published: 8/26/2026
πŸ”— http://arxiv.org/abs/2608.25518v1
πŸ‘₯ Authors: Pengfei Zhou, Hexin Wang, Zhengfeiyang Zhang, Yixing Ma, Zhenglin Wan, Kaipeng Zhang (possible past Tencent (China) affiliation), Wangbo Zhao, Yang You (possible past University Of California, Berkeley affiliation)
Abstract

A common strategy for scaling world models is to train on more crawled video with more compute. We argue that this strategy is inefficient: scaling world models also requires a recursive data engine that offers grounded reward signals. The success of code agents illustrates why this matters. As code is executable, compilers and runtimes can provide high-quality rewards for Reinforcement Learning (RL) post-training of LLMs. By contrast, spatial generation still relies largely on fuzzy proxies suc...

πŸ“„ Where vs What: Decomposing Structural and Content Failures in LLM-Generated Structured Outputs
πŸ—“οΈ Published: 8/26/2026
πŸ”— http://arxiv.org/abs/2608.25358v1
πŸ‘₯ Authors: Yiwei Zhang (possible past Tsinghua University affiliation), Chengke Wu, Li Wang (possible past Tesla (United States) affiliation), Jianqiang Li
Abstract

Structured outputs such as JSON and tables are central to modern LLM-based systems, yet generation failures are evaluated monolithically, conflating two distinct error modes: placement errors (correct values at wrong positions) and value errors (wrong values at intended positions). We introduce Structure-Content Decomposition (SCD), a framework that independently measures structural fidelity and content accuracy. Applying SCD to nested JSON and table tasks across six models (7B to frontier), we ...

πŸ“„ Combining Self-Embedding Audio Watermarking with Ultra-Low-Bitrate Neural Codecs
πŸ—“οΈ Published: 8/26/2026
πŸ”— http://arxiv.org/abs/2608.25289v1
πŸ‘₯ Authors: Yigitcan Γ–zer, Xin Wang (possible past University Of Edinburgh affiliation), Zhe Zhang, Junichi Yamagishi (possible past University Of Edinburgh affiliation)
Abstract

Partial manipulation of speech recordings, where only localized segments of an utterance are altered, poses a significant challenge for content integrity verification, as reliable detection and localization of such edits becomes harder as the manipulated proportion decreases. Watermarking offers a proactive defense alternative by embedding auxiliary information prior to distribution; classical hash-based schemes achieve near-perfect detection and localization under ideal conditions, but the orig...

πŸ“„ A Training-Free Proactive Defense Against Partial Speech Manipulation via Self-Embedding Steganography
πŸ—“οΈ Published: 8/26/2026
πŸ”— http://arxiv.org/abs/2608.25285v1
πŸ‘₯ Authors: Yigitcan Γ–zer, Zhe Zhang, Wanying Ge, Xin Wang (possible past University Of Edinburgh affiliation), Junichi Yamagishi (possible past University Of Edinburgh affiliation)
Abstract

Partial deepfake speech, where only limited segments of an utterance are synthesized or manipulated, poses a significant challenge to existing deepfake detection systems. As the proportion of spoofed regions decreases, passive detectors become increasingly unreliable, and accurate detection and restoration remain challenging. In this paper, we revisit audio steganography from a new perspective and propose its use as a proactive defense against partially deepfaked audio. In particular, we conside...

πŸ“„ Tunable Tool-Call Rates in LLM Agents via Representation Steering
πŸ—“οΈ Published: 8/25/2026
πŸ”— http://arxiv.org/abs/2608.25198v1
πŸ‘₯ Authors: Yuqi Chen, Vincent Siu, Yang Liu (possible past Tsinghua University affiliation), Dawn Song (possible past University Of California, Berkeley affiliation), Chenguang Wang (possible past Amazon (United States) affiliation)
Abstract

Deciding whether to call a tool is a core competence of an LLM agent, and a costly one to get wrong: needless calls add latency, accrue cost, and may trigger irreversible side effects, while missing calls leave the model confidently wrong on questions it could only answer through tool-calls. Models manage this balance poorly, both over-using and under-using tools. Existing methods such as post-training and prompt engineering are expensive and difficult to modify at inference time. We show that w...

πŸ“„ FLINT: Efficiently Leveraging High Bandwidth Flash for Capacity-Scalable LLM Inference Acceleration
πŸ—“οΈ Published: 8/25/2026
πŸ”— http://arxiv.org/abs/2608.25062v1
πŸ‘₯ Authors: Geraldo F. Oliveira, Arash Tavakkol, Xiangyu Zhu, Ahmet Caner YΓΌzΓΌgΓΌler, Vamanan Arulchelvan, Lukas Cavigelli, Renzo Andri, Mohammad Sadrosadati, Jia Xinglei, Onur Mutlu (possible past Eth Zurich affiliation), Zhou Ke, Shai Bergman, Ji Zhang (possible past Nvidia (United States) affiliation)
Abstract

LLM inference is increasingly constrained by accelerator memory capacity rather than compute throughput. This constraint is especially acute in single-accelerator and small-node inference systems, where limited on-package memory capacity restricts the size of deployable models. HBF is an emerging 3D-stacked NAND flash technology that provides multi-terabyte near-accelerator capacity, making it a promising capacity tier for storing LLM weights. However, existing HBF-based proposals face three ado...

πŸ“„ Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses
πŸ—“οΈ Published: 8/25/2026
πŸ”— http://arxiv.org/abs/2608.24876v1
πŸ‘₯ Authors: Zhaochen Yu, Yingcheng Wu, Zhenfei Yin, Kaiyuan Chen, Zhe Zhao (possible past Tencent (China) affiliation), Mengdi Wang, Shuicheng Yan (possible past National University Of Singapore affiliation), Ling Yang
Abstract

Recursive self-improvement (RSI) remains hard in long-horizon tasks, where growing histories obscure the task state and misalign skill invocation. We introduce Recuris, a recursive Experiential-Working Memory architecture for long-horizon agent harnesses, in which Working Memory tracks task progress and guides skill selection from Experiential Memory, grounding skill use in current needs rather than the full history. This coupling also turns execution into structured evidence that localizes fail...

πŸ“„ StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing
πŸ—“οΈ Published: 8/25/2026
πŸ”— http://arxiv.org/abs/2608.24777v1
πŸ‘₯ Authors: Zhijie Zheng, Yu Li (possible past Tencent (China) affiliation), Chen Qian (possible past Shanghai Jiao Tong University affiliation), Yuqian Fu, Yanwei Fu, Lu Sheng, Jing Shao, Dongrui Liu
Abstract

LLM-based agents can interact with external environments through tool invocation, but this capability also introduces security risks such as file modification, information leakage, and unauthorized actions. Existing guardrails often evaluate completed trajectories, leaving pre-execution monitoring of step-level actions underexplored. We propose StepGuard, a step-level guard model that can audit completed agent trajectories and check tool actions before they are executed. To train StepGuard, we i...

πŸ“„ RACE: Scalable Statistical Estimation of Functional Consistency in LLM Neurons
πŸ—“οΈ Published: 8/25/2026
πŸ”— http://arxiv.org/abs/2608.24758v1
πŸ‘₯ Authors: Runyu Wang, Bo Liu (possible past Meta (United States) affiliation), Xiaxin Zhang, Yu Han, Jiawei Cao, Xiaoye Zhang, Zhe Zhang, Yifan Yang (possible past Tencent (China) affiliation), Peng Ping
Abstract

Discovering stable neuron behavior across entire domains remains a challenge in mechanistic interpretability. Existing methods often rely on instance-level point estimates or computationally expensive procedures, which either obscure population-level variability or limit scalable domain-wide analysis. We present RACE (Residual Alignment for Consistency Estimation), a forward-pass statistical framework that evaluates the domain-wide functional consistency of Transformer neurons. Perturbation expe...

πŸ“„ CEDAR: Controlled and Event-Driven Demand Forecasting via Residual Decomposition
πŸ—“οΈ Published: 8/26/2026
πŸ”— http://arxiv.org/abs/2608.25871v1
πŸ‘₯ Authors: Junjie Meng, Ranxu Zhang, Zi-An Zhang, Shujun Liu, Xiaoning Qi, Xiaozhou Xu, Yanyong Zhang, Hui Xiong (possible past Baidu (China) affiliation), Chao Wang (possible past Google (United States) affiliation)
Abstract

Forecasting in large-scale e-commerce marketplaces is increasingly required to support planning: merchants need to evaluate sales outcomes under future action sequences such as budget schedules, rather than passively predicting what happens next. However, most existing time series forecasting (TSF) approaches remain inherently passive. Even when incorporating operational decisions as auxiliary covariates, they typically optimize for correlation-based extrapolation under historical policies. This...

πŸ“„ A Constitutive Markov Physics-Informed Neural Operator (MPNO) for Autoregressive Stability in Transient Dynamics
πŸ—“οΈ Published: 8/26/2026
πŸ”— http://arxiv.org/abs/2608.25744v1
πŸ‘₯ Authors: Wenpu Du, Peng Zhou (possible past Tencent (China) affiliation), Yunlong Xia, Sinuo Xin, Congcong Zhang, Boyang Zhang, Yi Zhang (possible past Google (United States) affiliation), Wenzheng Xu
Abstract

Neural operators applied to transient-dynamics PDEs with strong discontinuities exhibit autoregressive instability: in concrete-penetration stress-field prediction, the wavelet neural operator (WNO) diverges in autoregressive rollout, while MeshGraphNets collapse to zero predictions. WNO's instability stems from the lack of a structural constraint on the spectral radius of its propagation operator; the Fourier neural operator (FNO) is stable in these measurements but only emergently, not by cons...

πŸ“„ JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution
πŸ—“οΈ Published: 8/26/2026
πŸ”— http://arxiv.org/abs/2608.25593v1
πŸ‘₯ Authors: Guibin Zhang, Leo Lu, Fangzhou Xie, Kang Zhu, Junhao Wang, Zhifei Xie, Zhaochen Yu, Zihang Liu, Zhongxiang Sun, Qiankun Li, Yue Liao, Heng Chang (possible past Tsinghua University affiliation), Xiaobin Hu (possible past Tencent (China) affiliation), Qibing Ren, Wangchunshu Zhou, Shuicheng Yan (possible past National University Of Singapore affiliation)
Abstract

Agent capability is not determined by the model alone. The agent harness, encompassing memory management, planning strategy, action protocol, and tool/skill orchestration, can dominate the contribution of the underlying foundation model. Yet harness design remains manual, task-specific, and fundamentally unscalable. We present JIT-Agent, a harness intelligence model trained to synthesize task-adaptive agent harnesses on the fly for arbitrary off-the-shelf agentic LLMs. We formalize the agent har...

πŸ“„ Scalable Self-Supervised Learning for Multiphase AC-OPF in Distribution Systems with Topology Reconfiguration
πŸ—“οΈ Published: 8/25/2026
πŸ”— http://arxiv.org/abs/2608.25095v1
πŸ‘₯ Authors: Hoang T. Nguyen, Shaohui Liu (possible past Tsinghua University affiliation), Reetam Sen Biswas, Varsha Pendyala, Nurali Virani, Deepjyoti Deka, Priya L. Donti (possible past Carnegie Mellon University affiliation)
Abstract

The proliferation of distributed energy resources (DERs) in distribution grids enables the active coordination of these assets to reduce costs and enable cleaner operations. Realizing this potential requires solving multiphase AC optimal power flow (AC-OPF) quickly across varying loads, DER availabilities, and topology reconfigurations, at much greater speed and scale than conventional nonlinear solvers. Learning-based surrogates can offer millisecond inference, yet existing methods target large...

πŸ“„ Maia 200: A Software Defined Dataflow System for Large-scale AI Acceleration
πŸ—“οΈ Published: 8/25/2026
πŸ”— http://arxiv.org/abs/2608.24664v1
πŸ‘₯ Authors: Sherry Xu, Marco Heddes, Jackson Peng, Tom Savell, Monica Tang, Prashant Ranjan, Jesse Benson, Ofer Dekel, Saurabh Dighe, Anupama Kurpad, Artour Levin, Matthew Mattina, George Petre, Cheng Tang, Yuan Yu (possible past Google (United States) affiliation), Li Zhang (possible past University Of Oxford affiliation), Torsten Hoefler (possible past Eth Zurich affiliation)
Abstract

We introduce Maia 200, an advanced AI accelerator delivering high performance-10 145 Tflop/s FP4 and 5072 Tflop/s FP8 within a 750W TDP and 7 TB/s HBM bandwidth. Maia exemplifies a new class of Software Defined Locally Accessed Dataflow Architectures (SDLA), which explicitly program dataflow engines to orchestrate highly specialized memories and data movement engines. This approach shifts the focus from today's thread-centric to data-movement-centric architecture, improving efficiency and scalab...

*Notable papers are those with at least two authors from a "big" AI/ML lab.