📄 Notable* Recent AI/ML arXiv Papers

Last updated just now...

📄 Implicit Q-learning-bootstrapped ant colony optimization for maritime moving-target observation scheduling with agile satellites
🗓️ Published: 8/25/2026
🔗 http://arxiv.org/abs/2608.24471v1
👥 Authors: He Wang (possible past Stanford University affiliation), Junyu Wu, Yeye Liu, Yifan Zhou, Jie Zhang, Hui Li (possible past Baidu (China) affiliation), Yanjie Song, Liang Li
Abstract

Maritime moving-target observation scheduling with agile Earth observation satellites is a dynamic, sequence-dependent combinatorial optimization problem. Sea-surface targets move continuously, causing feasible observation windows to vary with target motion and satellite orbital geometry. The scheduler must jointly determine task selection, satellite assignment, observation-window selection, and observation ordering under time-window, attitude-maneuvering, onboard-resource, and cloud-affected av...

📄 Reinforcement Learning-Guided Evolutionary Policy Optimization for Preference-Adjustable Heterogeneous Agile Earth Observation Satellite Scheduling
🗓️ Published: 8/25/2026
🔗 http://arxiv.org/abs/2608.24470v1
👥 Authors: He Wang (possible past Stanford University affiliation), Junyu Wu, Hui Li (possible past Baidu (China) affiliation), Yanjie Song, Witold Pedrycz, Liang Li
Abstract

Heterogeneous agile Earth observation satellite (AEOS) scheduling requires task selection, satellite assignment, and observation sequencing under satellite-dependent visibility windows, attitude maneuvering requirements, energy consumption, and onboard storage constraints. Since satellites differ in orbital access, maneuvering capability, and payload resources, the same task may have different feasible windows, transition costs, and resource-consumption patterns on different platforms, which inc...

📄 Not All Tokens Are Equal: Region-Aware Consistency Repair of Backdoors in MLLMs
🗓️ Published: 8/25/2026
🔗 http://arxiv.org/abs/2608.24354v1
👥 Authors: Jiali Wei, Ming Fan, Mingkun Zhang, Haoyu Wang (possible past Tencent (China) affiliation), Jun Sun, Guoheng Sun, Xiaoning Ren, Haijun Wang, Ting Liu (possible past Google (United States) affiliation)
Abstract

MLLMs are increasingly deployed in user-facing applications, yet they inherit backdoor risks from the pipelines used to construct them: triggers may reside in images, texts, or both. Existing model-level backdoor removal methods, largely designed for conventional classifiers, show limited effectiveness on MLLMs, while MLLM-specific defenses mainly operate at inference time, filtering suspicious inputs without removing the backdoor embedded in the model. To address this gap and eliminate latent b...

📄 Contrastive Branch Policy Optimization
🗓️ Published: 8/25/2026
🔗 http://arxiv.org/abs/2608.24300v1
👥 Authors: Ying Wang (possible past Tsinghua University affiliation), Changlin Qiu, Bang Lin, Linbo Jin, Wen Jiang, Zhe Sun (possible past Tsinghua University affiliation), Jingli Yang
Abstract

Reinforcement learning with verifiable rewards (RLVR) enables language models to learn multi-turn interaction with external tools, yet its sparse outcome rewards provide no signal for identifying which intermediate decisions are responsible for success. Branch sampling induces local comparisons among alternative continuations, but existing methods tend to conflate two distinct problems: allocating a fixed rollout budget and translating branch outcomes into token-level credit. We introduce Contra...

📄 RePolicy: Reinforcement Learning for Safety-Policy Invocation in Agent Safeguards
🗓️ Published: 8/25/2026
🔗 http://arxiv.org/abs/2608.24275v1
👥 Authors: Houcheng Jiang, Boxuan Zhang, Qiyong Zhong, Junfeng Fang, Xiang Wang (possible past Tencent (China) affiliation), Xiangnan He (possible past National University Of Singapore affiliation)
Abstract

Safeguarding language model agents requires assessing complete execution trajectories under context-dependent safety policies. Existing policy-aware safeguards mainly rely on prompting or supervised fine-tuning, limiting their ability to adapt to unseen trajectories and changing policy contexts. We propose RePolicy, an agent safeguard that learns safety-policy invocation through reinforcement learning. Given an agent trajectory and a dynamic policy library, RePolicy invokes the applicable policy...

📄 Tlow: Flow-based Item Tokenizer for Recommendation
🗓️ Published: 8/25/2026
🔗 http://arxiv.org/abs/2608.24176v1
👥 Authors: Nian Li, Chonggang Song (possible past Tencent (China) affiliation), Jingtao Ding, Lingling Yi (possible past Tencent (China) affiliation), Yong Li (possible past Tsinghua University affiliation), Qingmin Liao
Abstract

Item tokenizer encodes semantic embeddings into token IDs to replace the randomly assigned item IDs used in traditional recommendation models, fundamentally addressing the problems of excessive parameters and cold starts. However, the most common tokenizer, RQ-VAE, suffers from low decoding efficiency due to the inherent dependencies among its codebooks. Meanwhile, efficient independent tokenizers such as optimized product quantization (OPQ) still struggle with dimensional correlations and distr...

📄 Task-Adaptive Rubrics for GUI Reward Modeling
🗓️ Published: 8/25/2026
🔗 http://arxiv.org/abs/2608.24174v1
👥 Authors: Tao Xiong, Xavier Hu, Wenkai Wang, Qinzhuo Wu, Changqiao Wu, Pengzhi Gao, Wei Liu (possible past Tsinghua University affiliation), Jian Luan, Shengyu Zhang (possible past Tencent (China) affiliation)
Abstract

Recent studies on GUI agents have increasingly focused on outcome reward modeling, which assigns outcome rewards by judging whether an executed trajectory satisfies the success criteria implied by the user instruction. Existing GUI reward verifiers, however, often under-specify how these criteria should be constructed for each task instance. Whether using generic rubric structures or implicit model reasoning, their judging criteria are not sufficiently task-adaptive: they can transfer checks acr...

📄 OmniJudge or OmniBias? Diagnosing Multimodal Judges through Balanced, Decoupled Lenses
🗓️ Published: 8/25/2026
🔗 http://arxiv.org/abs/2608.24160v1
👥 Authors: Guangzheng Hu, Ziyue Jiang, Weixu Qiao, Lixin Zhang, Jianye Kang, Yuru Wu, Rong Bao, Niantong Li, Wei Wang (possible past University Of Oxford affiliation), Ziyi Cheng, Xinfa Zhu, Hangrui Hu, Ting He, Bing Zhao, Lin Qu, Hu Wei, Jin Xu (possible past Tencent (China) affiliation)
Abstract

Multimodal understanding models that can jointly judge text-to-image (T2I), text-to-video (T2V) and text-to-speech (TTS) generation are increasingly used as "OmniJudges" for evaluation and automatic annotation. How reliably they understand what they score remains unclear, since existing benchmarks and training data tend to overemphasize positive examples and to conflate distinct failure modes, so a judge may score well without recognizing failures while its capability gaps stay hidden. Motivated...

📄 TransPhy: Visual In-Context Learning for Physically Grounded Image Editing
🗓️ Published: 8/25/2026
🔗 http://arxiv.org/abs/2608.24119v1
👥 Authors: Siyi Xie, Xuanke Shi, Jinsheng Quan, Haoran Tang (possible past University Of California, Berkeley affiliation), Zukai Chen, Lei Yang (possible past Google (United States) affiliation), Quan Wang (possible past Google (United States) affiliation)
Abstract

Visual demonstrations provide a natural interface for specifying image transformations that are difficult to describe exhaustively with text. However, existing visual in-context learning (VICL) methods primarily focus on appearance-level relation transfer and provide limited support for physically grounded transformations, whose outcomes depend on material properties, geometry, object interactions, and environmental conditions. Given a source--target exemplar pair and a query image, physically g...

📄 SAGE: From Direct Answering to Evidence-Grounded Inference for Chinese Ancient Document Understanding
🗓️ Published: 8/25/2026
🔗 http://arxiv.org/abs/2608.24011v1
👥 Authors: Yuchuan Wu, Xuan Luo (possible past University Of Washington affiliation), Yinglian Zhu, Meng Fang (possible past Tencent (China) affiliation), Xiangyang Xue, Bin Li
Abstract

Chinese ancient document understanding demands complex visual, linguistic, and historical reasoning. Current Large Vision-Language Models (LVLMs) typically rely on an opaque, single-pass generation paradigm, often producing overconfident and weakly grounded responses. To address this, we propose SAGE, an evidence-grounded multi-agent framework that reformulates Chinese ancient document understanding as evidence-grounded inference rather than direct answer generation. SAGE coordinates specialized...

📄 Incorporating Cognitive Load and Knowledge Transfer for Multi-Domain Knowledge Tracing
🗓️ Published: 8/25/2026
🔗 http://arxiv.org/abs/2608.24005v1
👥 Authors: Haotian Zhang (possible past Stanford University affiliation), Shucun Wang, Jinze Wu, Liang Ding, Shuochen Liu, Zhenya Huang, Jing Sha, Shijin Wang, Qi Liu (possible past Tencent (China) affiliation)
Abstract

Knowledge Tracing (KT) aims to assess students' dynamic knowledge states from their learning histories. While most existing KT methods focus on single-domain learning with notable success, real-world learning scenarios often involve multiple domains simultaneously, introducing two critical factors: 1) Cognitive load, arising from managing learning across domains in both temporal and knowledge dimensions. 2) Knowledge transfer, where knowledge states in one domain influence related states both wi...

📄 ConvergeFlow: Language Flow with Provable Convergence to Token Embeddings
🗓️ Published: 8/24/2026
🔗 http://arxiv.org/abs/2608.23551v1
👥 Authors: Na Li (possible past Tencent (China) affiliation), Yuchen Jiao, Changxiao Cai, Gen Li (possible past University Of Edinburgh affiliation)
Abstract

Recent advances in continuous diffusion and flow-based language models (LMs) have achieved performance competitive with discrete LMs. However, existing continuous frameworks still rely on decoders supervised with cross entropy (CE) because the flow trajectories are not guaranteed to terminate at valid token embeddings. Motivated by this limitation, we introduce \textbf{ConvergeFlow}, an embedding-space flow-based LM, which constrains the data predictor to the convex hull of token embeddings and ...

📄 StrategyBench: Evaluating Explicit Strategy Induction in Large Language Models
🗓️ Published: 8/24/2026
🔗 http://arxiv.org/abs/2608.23475v1
👥 Authors: Jinghan Tan, Yuanzheng Wang (possible past Baidu (China) affiliation), Lu Chen, Zijun Chen (possible past Google (United States) affiliation), Yuqian Wang, Maosong Sun (possible past Tsinghua University affiliation)
Abstract

As large language models are increasingly used in data-scarce and evolving task scenarios, few-shot in-context learning (ICL) has become a key paradigm for task adaptation. However, direct ICL often uses a small set of examples without explicitly abstracting task rules, making it sensitive to example construction. In contrast, human learners often reduce such sensitivity by first summarizing task rules from examples and then applying them to new instances. To evaluate this ability, we propose St...

📄 Apodex 1.1: Scaling Agentic Intelligence for Complex Work
🗓️ Published: 8/24/2026
🔗 http://arxiv.org/abs/2608.23283v1
👥 Authors: Apodex Team, B. An, B. Li, B. Wang, B. Zhang, B. L. Wang, C. Feng, C. Wei, C. Xue, C. Zhang, D. Ng, D. Ye, E. Min, F. Chen, F. Liu, F. Yang, F. Ye, H. Xu, H. Yang, H. Ye, H. Zhang, H. Zhao, J. Li, J. Lin, J. Xia, K. Jin, K. Wang, K. Yang, L. Bing, L. Lei, L. Su, Le. Wang, Lu. Wang, N. Wang, Q. Ren, Q. Yang, R. Li, S. Bai, S. Du, S. Li, S. Lin, S. Nie, S. Wang, S. Zhang, S. Z. Wang, Ta. Q. Fang, Ti. Q. Fang, W. Fang, W. Li, W. Zhang, X. Chen (possible past Google (United States) affiliation), X. Li, X. Tang, X. Wang, X. Xu, X. Zhang, X. Q. Wang, X. Y. Wang, Y. Deng, Y. Gao, Y. Hu, Y. Li, Y. Sui, Y. Wang, Y. Xiao, Y. Zhang, Z. Chen (possible past Google (United States) affiliation), Z. Cheng, Z. Feng, Z. Liang, Z. Zhang
Abstract

General-purpose language models can reason and synthesize knowledge, but complex work also requires sustained interaction with files, information sources, and executable code, together with state maintenance, failure recovery, and verifiable delivery. We call this \emph{working capability}: sustained, verifiable progress toward a real-world objective. Apodex 1.1 develops this capability along two complementary dimensions. \emph{Environment Scaling} expands the diversity and verifiability of exec...

📄 Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data
🗓️ Published: 8/24/2026
🔗 http://arxiv.org/abs/2608.23256v1
👥 Authors: Yinhao Tang, Youqing Fang, Yanan Sun (possible past Tencent (China) affiliation), Jiangning Liu, Ziyi Wang, Xun Zhao (possible past Tencent (China) affiliation), Weiming Zhang, Bin Liu, Kuikun Liu, Wenwei Zhang, Kai Chen (possible past Shanghai Jiao Tong University affiliation)
Abstract

Recent work proposes next-chunk reasoning RL for leveraging no-CoT data---corpora such as worked solutions and textbook derivations that contain reasoning-rich content but lack explicit chain-of-thought annotations. The method trains a model to generate implicit reasoning traces and rewards them by their ability to predict the next chunk of text. While promising, existing evaluations primarily compare against conventional SFT baselines, leaving open whether the gains come from the RL formulation...

📄 NeuralParker: A Reinforcement Learning Planner for Irregular Parking Environments
🗓️ Published: 8/25/2026
🔗 http://arxiv.org/abs/2608.24485v1
👥 Authors: Zihan Wang (possible past Tsinghua University affiliation), Bai Huang, Yang Guan, Xiao Li, Haoyu Xu, Naizheng Wang, Shengbo Eben Li (possible past Tsinghua University affiliation)
Abstract

Automated parking commonly assumes marked slots and short approach maneuvers. Delivery and service vehicles, however, may need to reach an operator-specified pose in an irregular bounded environment from a distant start. Existing learning-based parking planners often rely on local observations, which can restrict long-range route reasoning. To address this problem, we present NeuralParker, a reinforcement learning-based hybrid planner for arbitrary-pose parking. NeuralParker encodes full-environ...

📄 The Laws of Context Allocation: Causal Measurement and Closed-Loop Orchestration in Generative Search
🗓️ Published: 8/24/2026
🔗 http://arxiv.org/abs/2608.23252v1
👥 Authors: Peiyang Liu, Xi Wang (possible past Tsinghua University affiliation), Di Liang, Wei Ye (possible past Meta (United States) affiliation)
Abstract

As Retrieval-Augmented Generation (RAG) shifts toward diverse portfolio generation, it is stymied by two critical bottlenecks: flawed measurement of evidence utilization, and suboptimal context budget allocation. We resolve both sequentially. To resolve measurement, we expose a pervasive ``diagnostic illusion'': standard relevance proxies fail catastrophically on hard negatives. We replace them with an efficient causal leave-one-out probe that accurately isolates generative reliance and formal...

📄 Learning Generalizable Behaviors for Terminal Agents
🗓️ Published: 8/23/2026
🔗 http://arxiv.org/abs/2608.22631v1
👥 Authors: Yihang Yao, Bo Pang, Xuan Phi Nguyen, Ding Zhao (possible past Google (United States) affiliation), Shafiq Joty, Semih Yavuz (possible past Google (United States) affiliation)
Abstract

Terminal agents are a compelling application of large language models (LLMs), with the potential to integrate deeply into users' daily workflows. Reinforcement learning (RL) is a key technique for improving their capabilities, making scalable training environments a central challenge. Since public real-user interaction data are scarce, synthetic environments provide a practical alternative, but often suffer from domain gaps and limited fidelity, leading to poor generalization. Existing work main...

📄 Scale-invariant Optimal Sampling for Rare-events Data with Sparse Models
🗓️ Published: 8/23/2026
🔗 http://arxiv.org/abs/2608.22597v1
👥 Authors: Jing Wang (possible past Google (United States) affiliation), Haiying Wang, Qiang Zhang (possible past Tsinghua University affiliation), Hao Helen Zhang
Abstract

Subsampling is effective in tackling computational challenges for massive data with rare events. Overly aggressive subsampling may adversely affect estimation efficiency, and optimal subsampling is essential to mitigate the information loss. However, existing optimal subsampling probabilities depend on data scales, and some scaling transformations may result in inefficient subsamples. This problem is more significant when there are inactive features, because their influence on the subsampling pr...

*Notable papers are those with at least two authors from a "big" AI/ML lab.