📄 Notable* Recent AI/ML arXiv Papers

Last updated just now...

📄 Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses
🗓️ Published: 8/25/2026
🔗 http://arxiv.org/abs/2608.24876v1
👥 Authors: Zhaochen Yu, Yingcheng Wu, Zhenfei Yin, Kaiyuan Chen, Zhe Zhao (possible past Tencent (China) affiliation), Mengdi Wang, Shuicheng Yan (possible past National University Of Singapore affiliation), Ling Yang
Abstract

Recursive self-improvement (RSI) remains hard in long-horizon tasks, where growing histories obscure the task state and misalign skill invocation. We introduce Recuris, a recursive Experiential-Working Memory architecture for long-horizon agent harnesses, in which Working Memory tracks task progress and guides skill selection from Experiential Memory, grounding skill use in current needs rather than the full history. This coupling also turns execution into structured evidence that localizes fail...

📄 StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing
🗓️ Published: 8/25/2026
🔗 http://arxiv.org/abs/2608.24777v1
👥 Authors: Zhijie Zheng, Yu Li (possible past Tencent (China) affiliation), Chen Qian (possible past Shanghai Jiao Tong University affiliation), Yuqian Fu, Yanwei Fu, Lu Sheng, Jing Shao, Dongrui Liu
Abstract

LLM-based agents can interact with external environments through tool invocation, but this capability also introduces security risks such as file modification, information leakage, and unauthorized actions. Existing guardrails often evaluate completed trajectories, leaving pre-execution monitoring of step-level actions underexplored. We propose StepGuard, a step-level guard model that can audit completed agent trajectories and check tool actions before they are executed. To train StepGuard, we i...

📄 RACE: Scalable Statistical Estimation of Functional Consistency in LLM Neurons
🗓️ Published: 8/25/2026
🔗 http://arxiv.org/abs/2608.24758v1
👥 Authors: Runyu Wang, Bo Liu (possible past Meta (United States) affiliation), Xiaxin Zhang, Yu Han, Jiawei Cao, Xiaoye Zhang, Zhe Zhang, Yifan Yang (possible past Tencent (China) affiliation), Peng Ping
Abstract

Discovering stable neuron behavior across entire domains remains a challenge in mechanistic interpretability. Existing methods often rely on instance-level point estimates or computationally expensive procedures, which either obscure population-level variability or limit scalable domain-wide analysis. We present RACE (Residual Alignment for Consistency Estimation), a forward-pass statistical framework that evaluates the domain-wide functional consistency of Transformer neurons. Perturbation expe...

📄 Maia 200: A Software Defined Dataflow System for Large-scale AI Acceleration
🗓️ Published: 8/25/2026
🔗 http://arxiv.org/abs/2608.24664v1
👥 Authors: Sherry Xu, Marco Heddes, Jackson Peng, Tom Savell, Monica Tang, Prashant Ranjan, Jesse Benson, Ofer Dekel, Saurabh Dighe, Anupama Kurpad, Artour Levin, Matthew Mattina, George Petre, Cheng Tang, Yuan Yu (possible past Google (United States) affiliation), Li Zhang (possible past University Of Oxford affiliation), Torsten Hoefler (possible past Eth Zurich affiliation)
Abstract

We introduce Maia 200, an advanced AI accelerator delivering high performance-10 145 Tflop/s FP4 and 5072 Tflop/s FP8 within a 750W TDP and 7 TB/s HBM bandwidth. Maia exemplifies a new class of Software Defined Locally Accessed Dataflow Architectures (SDLA), which explicitly program dataflow engines to orchestrate highly specialized memories and data movement engines. This approach shifts the focus from today's thread-centric to data-movement-centric architecture, improving efficiency and scalab...

📄 Parason: Revealing Subtask and Trial Parallelism in LLM Reasoning
🗓️ Published: 8/25/2026
🔗 http://arxiv.org/abs/2608.24658v1
👥 Authors: Zhengyang Zhang, Zijian Zhang, Jiaxuan Gao, Shusheng Xu, Yi Wu (possible past University Of California, Berkeley affiliation), Song Han (possible past Stanford University affiliation), Ligeng Zhu
Abstract

Scaling test-time reasoning has substantially improved the problem-solving ability of large language models (LLMs), but standard autoregressive decoding still executes long reasoning traces sequentially, creating severe latency for difficult tasks (up to days and weeks). Parallel reasoning offers a natural remedy. However, prior systems primarily focus on Subtask Parallelism, where the model learns to decompose a high-level task into smaller chunks that can be solved independently. This approach...

📄 Implicit Q-learning-bootstrapped ant colony optimization for maritime moving-target observation scheduling with agile satellites
🗓️ Published: 8/25/2026
🔗 http://arxiv.org/abs/2608.24471v1
👥 Authors: He Wang (possible past Stanford University affiliation), Junyu Wu, Yeye Liu, Yifan Zhou, Jie Zhang, Hui Li (possible past Baidu (China) affiliation), Yanjie Song, Liang Li
Abstract

Maritime moving-target observation scheduling with agile Earth observation satellites is a dynamic, sequence-dependent combinatorial optimization problem. Sea-surface targets move continuously, causing feasible observation windows to vary with target motion and satellite orbital geometry. The scheduler must jointly determine task selection, satellite assignment, observation-window selection, and observation ordering under time-window, attitude-maneuvering, onboard-resource, and cloud-affected av...

📄 Reinforcement Learning-Guided Evolutionary Policy Optimization for Preference-Adjustable Heterogeneous Agile Earth Observation Satellite Scheduling
🗓️ Published: 8/25/2026
🔗 http://arxiv.org/abs/2608.24470v1
👥 Authors: He Wang (possible past Stanford University affiliation), Junyu Wu, Hui Li (possible past Baidu (China) affiliation), Yanjie Song, Witold Pedrycz, Liang Li
Abstract

Heterogeneous agile Earth observation satellite (AEOS) scheduling requires task selection, satellite assignment, and observation sequencing under satellite-dependent visibility windows, attitude maneuvering requirements, energy consumption, and onboard storage constraints. Since satellites differ in orbital access, maneuvering capability, and payload resources, the same task may have different feasible windows, transition costs, and resource-consumption patterns on different platforms, which inc...

📄 Not All Tokens Are Equal: Region-Aware Consistency Repair of Backdoors in MLLMs
🗓️ Published: 8/25/2026
🔗 http://arxiv.org/abs/2608.24354v1
👥 Authors: Jiali Wei, Ming Fan, Mingkun Zhang, Haoyu Wang (possible past Tencent (China) affiliation), Jun Sun, Guoheng Sun, Xiaoning Ren, Haijun Wang, Ting Liu (possible past Google (United States) affiliation)
Abstract

MLLMs are increasingly deployed in user-facing applications, yet they inherit backdoor risks from the pipelines used to construct them: triggers may reside in images, texts, or both. Existing model-level backdoor removal methods, largely designed for conventional classifiers, show limited effectiveness on MLLMs, while MLLM-specific defenses mainly operate at inference time, filtering suspicious inputs without removing the backdoor embedded in the model. To address this gap and eliminate latent b...

📄 Contrastive Branch Policy Optimization
🗓️ Published: 8/25/2026
🔗 http://arxiv.org/abs/2608.24300v1
👥 Authors: Ying Wang (possible past Tsinghua University affiliation), Changlin Qiu, Bang Lin, Linbo Jin, Wen Jiang, Zhe Sun (possible past Tsinghua University affiliation), Jingli Yang
Abstract

Reinforcement learning with verifiable rewards (RLVR) enables language models to learn multi-turn interaction with external tools, yet its sparse outcome rewards provide no signal for identifying which intermediate decisions are responsible for success. Branch sampling induces local comparisons among alternative continuations, but existing methods tend to conflate two distinct problems: allocating a fixed rollout budget and translating branch outcomes into token-level credit. We introduce Contra...

📄 RePolicy: Reinforcement Learning for Safety-Policy Invocation in Agent Safeguards
🗓️ Published: 8/25/2026
🔗 http://arxiv.org/abs/2608.24275v1
👥 Authors: Houcheng Jiang, Boxuan Zhang, Qiyong Zhong, Junfeng Fang, Xiang Wang (possible past Tencent (China) affiliation), Xiangnan He (possible past National University Of Singapore affiliation)
Abstract

Safeguarding language model agents requires assessing complete execution trajectories under context-dependent safety policies. Existing policy-aware safeguards mainly rely on prompting or supervised fine-tuning, limiting their ability to adapt to unseen trajectories and changing policy contexts. We propose RePolicy, an agent safeguard that learns safety-policy invocation through reinforcement learning. Given an agent trajectory and a dynamic policy library, RePolicy invokes the applicable policy...

📄 Tlow: Flow-based Item Tokenizer for Recommendation
🗓️ Published: 8/25/2026
🔗 http://arxiv.org/abs/2608.24176v1
👥 Authors: Nian Li, Chonggang Song (possible past Tencent (China) affiliation), Jingtao Ding, Lingling Yi (possible past Tencent (China) affiliation), Yong Li (possible past Tsinghua University affiliation), Qingmin Liao
Abstract

Item tokenizer encodes semantic embeddings into token IDs to replace the randomly assigned item IDs used in traditional recommendation models, fundamentally addressing the problems of excessive parameters and cold starts. However, the most common tokenizer, RQ-VAE, suffers from low decoding efficiency due to the inherent dependencies among its codebooks. Meanwhile, efficient independent tokenizers such as optimized product quantization (OPQ) still struggle with dimensional correlations and distr...

📄 Task-Adaptive Rubrics for GUI Reward Modeling
🗓️ Published: 8/25/2026
🔗 http://arxiv.org/abs/2608.24174v1
👥 Authors: Tao Xiong, Xavier Hu, Wenkai Wang, Qinzhuo Wu, Changqiao Wu, Pengzhi Gao, Wei Liu (possible past Tsinghua University affiliation), Jian Luan, Shengyu Zhang (possible past Tencent (China) affiliation)
Abstract

Recent studies on GUI agents have increasingly focused on outcome reward modeling, which assigns outcome rewards by judging whether an executed trajectory satisfies the success criteria implied by the user instruction. Existing GUI reward verifiers, however, often under-specify how these criteria should be constructed for each task instance. Whether using generic rubric structures or implicit model reasoning, their judging criteria are not sufficiently task-adaptive: they can transfer checks acr...

📄 OmniJudge or OmniBias? Diagnosing Multimodal Judges through Balanced, Decoupled Lenses
🗓️ Published: 8/25/2026
🔗 http://arxiv.org/abs/2608.24160v1
👥 Authors: Guangzheng Hu, Ziyue Jiang, Weixu Qiao, Lixin Zhang, Jianye Kang, Yuru Wu, Rong Bao, Niantong Li, Wei Wang (possible past University Of Oxford affiliation), Ziyi Cheng, Xinfa Zhu, Hangrui Hu, Ting He, Bing Zhao, Lin Qu, Hu Wei, Jin Xu (possible past Tencent (China) affiliation)
Abstract

Multimodal understanding models that can jointly judge text-to-image (T2I), text-to-video (T2V) and text-to-speech (TTS) generation are increasingly used as "OmniJudges" for evaluation and automatic annotation. How reliably they understand what they score remains unclear, since existing benchmarks and training data tend to overemphasize positive examples and to conflate distinct failure modes, so a judge may score well without recognizing failures while its capability gaps stay hidden. Motivated...

📄 TransPhy: Visual In-Context Learning for Physically Grounded Image Editing
🗓️ Published: 8/25/2026
🔗 http://arxiv.org/abs/2608.24119v1
👥 Authors: Siyi Xie, Xuanke Shi, Jinsheng Quan, Haoran Tang (possible past University Of California, Berkeley affiliation), Zukai Chen, Lei Yang (possible past Google (United States) affiliation), Quan Wang (possible past Google (United States) affiliation)
Abstract

Visual demonstrations provide a natural interface for specifying image transformations that are difficult to describe exhaustively with text. However, existing visual in-context learning (VICL) methods primarily focus on appearance-level relation transfer and provide limited support for physically grounded transformations, whose outcomes depend on material properties, geometry, object interactions, and environmental conditions. Given a source--target exemplar pair and a query image, physically g...

📄 SAGE: From Direct Answering to Evidence-Grounded Inference for Chinese Ancient Document Understanding
🗓️ Published: 8/25/2026
🔗 http://arxiv.org/abs/2608.24011v1
👥 Authors: Yuchuan Wu, Xuan Luo (possible past University Of Washington affiliation), Yinglian Zhu, Meng Fang (possible past Tencent (China) affiliation), Xiangyang Xue, Bin Li
Abstract

Chinese ancient document understanding demands complex visual, linguistic, and historical reasoning. Current Large Vision-Language Models (LVLMs) typically rely on an opaque, single-pass generation paradigm, often producing overconfident and weakly grounded responses. To address this, we propose SAGE, an evidence-grounded multi-agent framework that reformulates Chinese ancient document understanding as evidence-grounded inference rather than direct answer generation. SAGE coordinates specialized...

📄 Incorporating Cognitive Load and Knowledge Transfer for Multi-Domain Knowledge Tracing
🗓️ Published: 8/25/2026
🔗 http://arxiv.org/abs/2608.24005v1
👥 Authors: Haotian Zhang (possible past Stanford University affiliation), Shucun Wang, Jinze Wu, Liang Ding, Shuochen Liu, Zhenya Huang, Jing Sha, Shijin Wang, Qi Liu (possible past Tencent (China) affiliation)
Abstract

Knowledge Tracing (KT) aims to assess students' dynamic knowledge states from their learning histories. While most existing KT methods focus on single-domain learning with notable success, real-world learning scenarios often involve multiple domains simultaneously, introducing two critical factors: 1) Cognitive load, arising from managing learning across domains in both temporal and knowledge dimensions. 2) Knowledge transfer, where knowledge states in one domain influence related states both wi...

📄 NeuralParker: A Reinforcement Learning Planner for Irregular Parking Environments
🗓️ Published: 8/25/2026
🔗 http://arxiv.org/abs/2608.24485v1
👥 Authors: Zihan Wang (possible past Tsinghua University affiliation), Bai Huang, Yang Guan, Xiao Li, Haoyu Xu, Naizheng Wang, Shengbo Eben Li (possible past Tsinghua University affiliation)
Abstract

Automated parking commonly assumes marked slots and short approach maneuvers. Delivery and service vehicles, however, may need to reach an operator-specified pose in an irregular bounded environment from a distant start. Existing learning-based parking planners often rely on local observations, which can restrict long-range route reasoning. To address this problem, we present NeuralParker, a reinforcement learning-based hybrid planner for arbitrary-pose parking. NeuralParker encodes full-environ...

📄 ConvergeFlow: Language Flow with Provable Convergence to Token Embeddings
🗓️ Published: 8/24/2026
🔗 http://arxiv.org/abs/2608.23551v1
👥 Authors: Na Li (possible past Tencent (China) affiliation), Yuchen Jiao, Changxiao Cai, Gen Li (possible past University Of Edinburgh affiliation)
Abstract

Recent advances in continuous diffusion and flow-based language models (LMs) have achieved performance competitive with discrete LMs. However, existing continuous frameworks still rely on decoders supervised with cross entropy (CE) because the flow trajectories are not guaranteed to terminate at valid token embeddings. Motivated by this limitation, we introduce \textbf{ConvergeFlow}, an embedding-space flow-based LM, which constrains the data predictor to the convex hull of token embeddings and ...

📄 Apodex 1.1: Scaling Agentic Intelligence for Complex Work
🗓️ Published: 8/24/2026
🔗 http://arxiv.org/abs/2608.23283v2
👥 Authors: B. An, B. Li, B. Wang, B. Zhang, B. L. Wang, C. Feng, C. Wei, C. Xue, C. Zhang, D. Ng, D. Ye, E. Min, F. Chen, F. Liu, F. Yang, F. Ye, G. Sun, H. Ji, H. Xu, H. Yang, H. Ye, H. Zhang, H. Zhao, J. Li, J. Lin, J. Xia, K. Jin, K. Wang, K. Yang, L. Bing, L. Lei, L. Su, Le. Wang, Lu. Wang, N. Wang, Q. Ren, Q. Yang, R. Li, S. Bai, S. Du, S. Li, S. Lin, S. Nie, S. Wang, S. Zhang, S. Z. Wang, T. Ge, Ta. Q. Fang, Ti. Q. Fang, W. Fang, W. Li, W. Zhang, X. Chen (possible past Google (United States) affiliation), X. Li, X. Tang, X. Wang, X. Xu, X. Zhang, X. Q. Wang, X. Y. Wang, Y. Deng, Y. Gao, Y. Hu, Y. Li, Y. Sui, Y. Wang, Y. Xiao, Y. Zhang, Y. Zhou, Z. Chen (possible past Google (United States) affiliation), Z. Cheng, Z. Feng, Z. Liang, Z. Liu, Z. Zhang
Abstract

General-purpose language models can reason and synthesize knowledge, but complex work also requires sustained interaction with files, information sources, and executable code, together with state maintenance, failure recovery, and verifiable delivery. We call this \emph{working capability}: sustained, verifiable progress toward a real-world objective. Apodex 1.1 develops this capability along two complementary dimensions. \emph{Environment Scaling} expands the diversity and verifiability of exec...

📄 The Laws of Context Allocation: Causal Measurement and Closed-Loop Orchestration in Generative Search
🗓️ Published: 8/24/2026
🔗 http://arxiv.org/abs/2608.23252v1
👥 Authors: Peiyang Liu, Xi Wang (possible past Tsinghua University affiliation), Di Liang, Wei Ye (possible past Meta (United States) affiliation)
Abstract

As Retrieval-Augmented Generation (RAG) shifts toward diverse portfolio generation, it is stymied by two critical bottlenecks: flawed measurement of evidence utilization, and suboptimal context budget allocation. We resolve both sequentially. To resolve measurement, we expose a pervasive ``diagnostic illusion'': standard relevance proxies fail catastrophically on hard negatives. We replace them with an efficient causal leave-one-out probe that accurately isolates generative reliance and formal...

*Notable papers are those with at least two authors from a "big" AI/ML lab.