📄 Notable* Recent AI/ML arXiv Papers

Last updated just now...

📄 OpenForgeRL: Train Harness-native Agents in Any Environment
🗓️ Published: 7/23/2026
🔗 http://arxiv.org/abs/2607.21557v1
👥 Authors: Xiao Yu, Baolin Peng, Ruize Xu, Hao Zou, Qianhui Wu, Hao Cheng (possible past Tencent (China) affiliation), Wenlin Yao, Nikhil Singh, Zhou Yu, Jianfeng Gao (possible past Microsoft (United States) affiliation)
Abstract

Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to external systems. While powerful, these complex harnesses also make agents hard to train end-to-end with open infrastructure, whose SFT/RL stacks cannot natively express stateful, multi-process harness inference. To address this, we present OpenForgeRL, an open-source framework for training harness-based agents end-to-end in diverse environments. ...

📄 MIRROR: Learning from the Other View for Multi-Modal Reasoning
🗓️ Published: 7/23/2026
🔗 http://arxiv.org/abs/2607.21552v1
👥 Authors: Wen Ye, Yuxiao Qu, Aviral Kumar (possible past University Of California, Berkeley affiliation), Xuezhe Ma (possible past Carnegie Mellon University affiliation)
Abstract

Unlike large language models (LLMs) that exhibit strong reasoning capabilities, vision-language models (VLMs) struggle with visual reasoning, even on geometry problems that admit equivalent text, diagram, and combined diagram+text views. We show that these views often elicit different behaviors: a model may solve a problem from text but fail on the corresponding diagram, or succeed visually while failing textually. This inconsistency suggests that different views expose complementary reasoning p...

📄 PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning
🗓️ Published: 7/23/2026
🔗 http://arxiv.org/abs/2607.21419v1
👥 Authors: Yipeng Shi, Zhipeng Ma, Yue Wang, Qitai Tan, Yang Li (possible past Google (United States) affiliation), Peng Chen (possible past Tencent (China) affiliation), Zhengzhou Zhu
Abstract

In long-horizon LLM agent reinforcement learning, weak policies often repeat similar failures, producing uninformative rollout trajectories and limiting effective policy optimization. Existing skill-centric methods improve exploration by optimizing, filtering, or internalizing reusable skills. However, they remain centered on the skills themselves rather than being designed as adaptive training-time support for the evolving policy. To address this, we propose a policy-centric training paradigm t...

📄 Hilbert Operator for Progressive Encoding (HOPE): A Mathematical Framework for Deconstructing Learned Representations in Deep Networks
🗓️ Published: 7/23/2026
🔗 http://arxiv.org/abs/2607.21366v1
👥 Authors: Hossein Mobahi (possible past Massachusetts Institute Of Technology affiliation), Peter L. Bartlett (possible past University Of California, Berkeley affiliation)
Abstract

Deep neural networks encode complex representations, but deconstructing this internal knowledge remains a challenge. Given the link between learning and compression, network compression offers a promising lens to analyze this knowledge. However, standard compression heuristics often suffer from scale symmetries and architectural biases. To resolve these, we introduce Hilbert Operator for Progressive Encoding (HOPE), a mathematical framework to gradually deconstruct the representations in trained...

📄 SPORD: A Simulation-Propose-then-OR-Dispose Approach for Supply Chain Planning
🗓️ Published: 7/23/2026
🔗 http://arxiv.org/abs/2607.21354v1
👥 Authors: Jiayin He, Yutong Pan, Sen Yang (possible past Tencent (China) affiliation), Ningxuan Kang, Yongzhi Qi, Jianshen Zhang, Wei Qi (possible past Baidu (China) affiliation), Zuo-Jun Max Shen
Abstract

For years, supply chain planning at e-commerce firms has operated as a collection of isolated projects. Each planning task from static network planning to dynamic warehouse assortment planning requires analysts to spend weeks building models from scratch, calibrating and persuading executives to act on outputs they cannot verify. Three barriers drive this: bespoke models proliferate because standardization is difficult (operational fragmentation); once unified, the combinatorial scale of million...

📄 SciExplore: Evaluating Autonomous Agents from Scientific Navigation to Information Integration
🗓️ Published: 7/23/2026
🔗 http://arxiv.org/abs/2607.20926v1
👥 Authors: Yinhao Tang, Youqing Fang, Yanan Sun (possible past Tencent (China) affiliation), Wenran Liu, Weiming Zhang, Bin Liu, Kuikun Liu, Wenwei Zhang, Kai Chen (possible past Shanghai Jiao Tong University affiliation)
Abstract

Scientific research involves complex information-seeking and reasoning workflows across heterogeneous sources. However, existing benchmarks primarily emphasize general-domain retrieval or static scientific question answering, and therefore fail to assess key capabilities required in realistic scientific research workflows. We introduce SciExplore, a benchmark designed to evaluate scientific information-seeking and reasoning capabilities of LLMs and agents. SciExplore comprises four task types co...

📄 Probabilistic Residual Learning for Online Recommendations
🗓️ Published: 7/23/2026
🔗 http://arxiv.org/abs/2607.20863v1
👥 Authors: Wenyuan Wang, Yusong Zhao, Zihao Xu, Hengyi Wang, Qi Xu, Zhigang Hua, Yan Xie, Yi Wang, Zihao Zhao (possible past Tsinghua University affiliation), Bo Long, Chengzhi Mao, Shuang Yang, Hengguan Huang, Hao Wang (possible past Tsinghua University affiliation)
Abstract

Modern recommender systems are typically based on deep learning (DL) models, where a dense encoder learns representations of users and items. As a result, these systems often suffer from the black-box nature and computational complexity of the underlying models, making it difficult to systematically enhance their recommendation capabilities. To address this problem, we propose Probabilistic Residual Learning (PRL), a causal Bayesian recommendation model that models the residual between ground-tr...

📄 RealVDeblur: One-Step Diffusion for Generalizable Real-World Video Deblurring
🗓️ Published: 7/22/2026
🔗 http://arxiv.org/abs/2607.20628v1
👥 Authors: Renbiao Jin, Mingxin Yang, Yutian Chen (possible past Deepmind (United Kingdom) affiliation), Junhao Zhuang, Xin Cai, Mulin Yu, Linning Xu, Wenxian Yu, Danping Zou, Shi Guo, Tianfan Xue (possible past Massachusetts Institute Of Technology affiliation)
Abstract

Real-world video deblurring remains challenging due to diverse motion patterns, complex degradations, and the scarcity of realistic training data, yet robust restoration is critical for downstream pipelines such as mobile imaging and 3D reconstruction. This work presents \textbf{RealVDeblur}, an efficient generative framework designed to improve in-the-wild robustness under diverse real capture conditions. First, a large-scale, physically grounded blur synthesis pipeline is constructed from scen...

📄 Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning
🗓️ Published: 7/22/2026
🔗 http://arxiv.org/abs/2607.20166v1
👥 Authors: Siqian Tong, Xuan Li (possible past Baidu (China) affiliation), Chaozhuo Li, Baolong Bi, Yiwei Wang (possible past Google (United States) affiliation), Yujun Cai, Shenghua Liu, Chengpeng Hao
Abstract

Large Audio Language models (LALMs) have made rapid progress on acoustic understanding, yet they still struggle with fine-grained audio reasoning (e.g., recognizing event order, repetitions and duration). Existing post-training methods heavily rely on expensive external labels or provide only coarse semantic signals. To bridge this gap, we introduce Audio-Zero, the first label-free self-evolution framework in the field of LALMs that improves fine-grained auditory perception and reasoning. Audio-...

📄 SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
🗓️ Published: 7/22/2026
🔗 http://arxiv.org/abs/2607.20145v1
👥 Authors: Dongfang Li, Xiaodong Luo, Ruoyu Sun, Xuhui Chen, Linyuan Qiu, Jian Meng, Zhengxuan Lu, Yiting Wang, Yucheng Xie, Tao Guo, Tianxiang Fang, Jing Li (possible past Tencent (China) affiliation), Sihang Chen, Shihao Hong, Chang Liu, Weihua Dai, Zirong Zeng, Ziwei Zhu, Zhuohan Wang, Zhengjun Yue, Igor Vasilyev, Min Liu, Weijian Sun, Xin Chen (possible past Tencent (China) affiliation), Yingmeng Gao, Jinhua Zhou, Taolue Chen, Chenwei Wu, Dong Zhang (possible past Nvidia (United States) affiliation), Wenlong Jin, Jinmin Xiang, Barkova Maria, Ushakov Anton, Xianfei Jin, Tian Ding, Zhihang Lin, Qian Chen (possible past Shanghai Jiao Tong University affiliation), Linxin Yang, Mingzhe Yang, Bingwei Zhang, Hongzhang Yang, Fangxue Zhang, Shijun Qin, Jie Yu, Cuihua Hu, Tolstykh Vasiliy, Nosov Ivan, Abdullin Amir, Zhichen Zhou, Xin Zhang (possible past Google (United States) affiliation), Zhixiong Ning, Xutong Zhao, Junjie Huang, Jiajun Liu, Weiyan Kong, Zheng Zhang, Wenhan Luo (possible past Tencent (China) affiliation), Lin Hu, Yangbo Guo, Li Zeng, Shihao Zeng, Baotian Hu, Min Zhang (possible past Tsinghua University affiliation), Haizhou Li, Zhiquan Luo
Abstract

Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. While most large-scale LLM training systems are built around GPU-based clusters, this report presents an end-to-end optimization practice on the Ascend NPU SuperPOD. Using the DeepSeek-V4 model family as the target workload, we develop a hi...

📄 Local Causal Structure Learning in the Presence of Latent Variables and Selection Bias
🗓️ Published: 7/22/2026
🔗 http://arxiv.org/abs/2607.19866v1
👥 Authors: Zheng Li, Hao Zhang (possible past Tencent (China) affiliation), Ruxin Wang, Ruichu Cai, Kun Zhang (possible past Google (United States) affiliation), Feng Xie
Abstract

Discovering the direct causes and effects of a target variable from observational data is a fundamental problem in causal discovery, with broad applications in domains such as gene regulatory analysis and biomedical research. Existing causal discovery methods either learn a global causal structure, which incurs substantial computational cost, or assume the absence of latent variables and selection bias, assumptions that are often violated in real-world settings. Motivated by these challenges, we...

📄 DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations
🗓️ Published: 7/22/2026
🔗 http://arxiv.org/abs/2607.19865v1
👥 Authors: Jiazhen Jiang, Boxi Cao, Lingyong Yan, Yaojie Lu, Hongyu Lin, Shuaiqiang Wang (possible past Baidu (China) affiliation), Dawei Yin (possible past Baidu (China) affiliation), Xianpei Han (possible past Tencent (China) affiliation), Le Sun
Abstract

As autonomous agents rapidly evolve, their ability to reliably manipulate ubiquitous digital documents has become critical for enabling general-purpose AI assistants and automating complex workspace workflows. In this paper, we introduce DocOps, a deterministically verifiable evaluation framework underpinned by a hierarchical taxonomy that deconstructs document operations inspired by real-world practices into atomic dimensions and escalating workflow complexities. Based on DocOps, we systematica...

📄 Hypothesis-and-Refinement Learning of Organic Structures from Multimodal Spectroscopic Data
🗓️ Published: 7/22/2026
🔗 http://arxiv.org/abs/2607.19816v1
👥 Authors: Chengchun Liu, Zhiyuan Yan, Li Yuan (possible past National University Of Singapore affiliation), Hao Li (possible past Tsinghua University affiliation), Boxuan Zhao, Yonghong Tian (possible past Peking University affiliation), Bartosz A. Grzybowski, Fanyang Mo
Abstract

Determining molecular structures from spectroscopic data remains fundamentally challenging because the inverse problem is intrinsically underdetermined: individual spectra are sparse, low-dimensional, and encode only partial structural evidence relative to the vast space of possible molecules. We address this challenge by formulating automated structure elucidation as a scalable hypothesis-refinement paradigm that tightly integrates spectral evidence with large-scale molecular priors. To supply ...

📄 Convergence-Latency-Aware Adaptive Modulation and Resource Allocation in RIS-Assisted Wireless Federated Learning
🗓️ Published: 7/22/2026
🔗 http://arxiv.org/abs/2607.19759v1
👥 Authors: Liwei Wang (possible past Tencent (China) affiliation), Wen Chen, Jun Li, Qingqing Wu, Ming Ding (possible past Tsinghua University affiliation), Xusheng Zhu, Qiong Wu
Abstract

Federated learning (FL) over wireless networks suffers from significant training latency and degraded convergence due to unreliable wireless transmission, especially under blocked propagation environments. Although reconfigurable intelligent surfaces (RISs) can improve communication reliability, existing wireless FL studies rarely characterize the trade-off between learning convergence and communication delay under modulation-dependent transmission errors. In this paper, we consider a wireless F...

*Notable papers are those with at least two authors from a "big" AI/ML lab.