📄 Notable* Recent AI/ML arXiv Papers

Last updated just now...

📄 ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning
🗓️ Published: 9/1/2026
🔗 http://arxiv.org/abs/2609.01058v1
👥 Authors: Fanrui Zhang, Ruixue Ding, Qiang Zhang (possible past Tsinghua University affiliation), Xi Chen (possible past University Of California, Berkeley affiliation), Boli Chen, Shihang Wang, Qiuchen Wang, Hongmin Zhan, Jinxin Bian, Li Xingchao, Peijin Zheng, Hao Cheng (possible past Tencent (China) affiliation), Pengjun Xie, Kaipeng Zhang (possible past Tencent (China) affiliation), Jiawei Liu, Zheng-Jun Zha
Abstract

Training open-ended agents via reinforcement learning (RL) is hindered by the lack of verifiable gold answers and scalable rubrics. Moreover, even near the model's capability boundary, long-horizon open-ended agentic tasks often yield brittle and unstable rewards, resulting in weak or noisy rollout contrast that obscures fine-grained optimization signals for group-based policy learning. To address these challenges, we propose ARISE-RL, a novel full-cycle self-evolution framework that couples a t...

📄 User Representation via Cross Multi-source Behavior Pre-training for Mobile Games
🗓️ Published: 9/1/2026
🔗 http://arxiv.org/abs/2609.01057v1
👥 Authors: Chengqi Yang, Yiran Qiao, Feng Liu, Xingyu Lou, Zijun Zhou, Xiaoyun Mo, Changwang Zhang (possible past Tencent (China) affiliation), Jiayuan Xu, Jun Wang (possible past Tencent (China) affiliation), Xiang Ao
Abstract

User representation pre-training has become a fundamental paradigm for alleviating data sparsity in downstream personalization tasks. However, existing studies predominantly focus on single-app or app-level behaviors, overlooking the inherently cross-source and multi-granular nature of user activities on mobile devices. At the device level, user intent emerges from complex interactions among heterogeneous behavior sources and hierarchical action structures, posing challenges that cannot be addre...

📄 Figures as Programs: Recursive Generation of Editable Scientific Figures
🗓️ Published: 9/1/2026
🔗 http://arxiv.org/abs/2609.01006v1
👥 Authors: Yepeng Liu, Dasen Dai, Chengzhi Liu, Yiren Song, Hai Ci, Yu Zhang (possible past Google (United States) affiliation), Qi Zhang (possible past Tencent (China) affiliation), Mike Zheng Shou (possible past National University Of Singapore affiliation), Xin Eric Wang, Yuheng Bu
Abstract

Scientific methodology figures are essential for communicating complex methods clearly, yet creating them remains labor-intensive and typically requires multiple rounds of refinement. Recent image-generation models can synthesize visually appealing raster figures, but producing a human-satisfactory result in a single generation step remains difficult. Moreover, precise edits to raster figures are challenging for both humans and models. We formulate scientific figure generation as recursive SVG p...

📄 Polished but Unresolved: Identifying Late-Stage Pressure States in Long-Horizon Tool-Use Agents
🗓️ Published: 9/1/2026
🔗 http://arxiv.org/abs/2609.00823v1
👥 Authors: Haoyang Chen, Yi Liu (possible past Google (United States) affiliation), Jianzhi Shao, Xiaozhou Xu, Zhe Sun (possible past Tsinghua University affiliation), Wei Hu
Abstract

Long-horizon tool-use agents need not only to search and plan, but also to decide when to finalize. We study late-stage pressure states, in which an agent is biased toward submitting a final answer that appears complete and polished while key constraints remain unresolved. We first train a linear probe to show that this pressure state is identifiable from the agent's hidden states. Then, we use activation interventions along this pressure direction and find that shifting the hidden states change...

📄 ContextPipe: Database-Inspired Context Assembly for Long-Horizon Agents
🗓️ Published: 9/1/2026
🔗 http://arxiv.org/abs/2609.00749v1
👥 Authors: Peng Xu (possible past Google (United States) affiliation), Zuyu Zhang, Yuze Sun, Feng Tian, Long Wang, Chen Zhang (possible past Peking University affiliation)
Abstract

Long-horizon large language model (LLM) agents require context assembly: the runtime must decide what to include in each prompt, in what order, and when to compact history under a hard context-window budget and a byte-sensitive prompt cache. In production agentic systems, this logic is scattered across prompt builders, ad hoc compaction routines, cache-break workarounds, and per-provider shims. We argue that context assembly is structurally isomorphic to query execution in a relational database:...

📄 ChatDev 2.0: A No-Code Multi-Agent Platform for Developing Everything
🗓️ Published: 9/1/2026
🔗 http://arxiv.org/abs/2609.00714v1
👥 Authors: Yufan Dang, Shu Yao, Bowen Lai, Chenting Xu, Ruijie Shi, Wai-Shing Leung, Huatao Li, Chen Qian (possible past Shanghai Jiao Tong University affiliation), Zhiyuan Liu (possible past Tsinghua University affiliation)
Abstract

Large language model (LLM)-based multi-agent systems (MAS) have shown strong potential for solving complex tasks, yet their development forces a tradeoff: code frameworks are expressive but engineering-intensive, while no-code builders simplify authoring but constrain agent interactions to author-defined workflows. We present ChatDev 2.0: DevAll (hereafter DevAll), a no-code platform for building, executing, and inspecting heterogeneous MAS that delivers both high expressiveness and ease of use....

📄 EM^2Mem: Event-Centric Multimodal Memory for Large Language Models
🗓️ Published: 9/1/2026
🔗 http://arxiv.org/abs/2609.00551v1
👥 Authors: Yijun Chen, Yaqi Zheng, Yanya Li, Boyi Xiao, Buqiang Xu, Shuofei Qiao, Jizhan Fang, Xinle Deng, Yunzhi Yao, Xuehai Wang, Liuxin Zhang, Hui Li (possible past Baidu (China) affiliation), Huajun Chen (possible past Alibaba Group (China) affiliation), Shumin Deng (possible past Alibaba Group (China) affiliation)
Abstract

Multimodal memory offers a scalable interface for long-video question answering, but existing methods often retrieve captions, frames, transcripts, summaries, or graph facts as isolated fragments. Although searchable, such fragments are not generation-ready: language models must reconstruct cross-modal and temporal alignments at inference time, when context is limited and attribution is difficult. We propose EM^2Mem, an event-centric multimodal memory framework that binds heterogeneous evidence ...

📄 ISO-RAG: Isoperimetric Noise Control for Retrieval-Augmented Generation
🗓️ Published: 9/1/2026
🔗 http://arxiv.org/abs/2609.00513v1
👥 Authors: Siyuan Zhang, Hanchen Wang (possible past University Of Cambridge affiliation), Dong Wen, Ying Zhang (possible past Tencent (China) affiliation), Wenjie Zhang
Abstract

Retrieval-Augmented Generation (RAG) mitigates large language models (LLMs) hallucinations, yet conventional dense retrieval struggles with the complex reasoning paths of multi-hop question answering (QA). Graph-based RAG captures multi-step relationships but suffers from severe semantic drift and high online latency due to noisy global graph traversals. Thus, we propose ISO-RAG (ISOperimetric Retrieval-Augmented Generation), a geometry-aware RAG framework. By projecting the underlying knowledge...

📄 EvoFlint: An Evolutionary Atlas of Multi-Turn LLM Vulnerabilities
🗓️ Published: 8/31/2026
🔗 http://arxiv.org/abs/2609.00487v1
👥 Authors: Feitong Qiao, Liren Peng, Shiming Ren, Aishwarya Jadhav, Arghavan Bahadorinejad, Marinette Chen, Muhan Zhang (possible past Meta (United States) affiliation), Abdulaziz Suria, Gennevi Lu, Anish Das Sarma (possible past Stanford University affiliation)
Abstract

Frontier language models that refuse harmful single-turn prompts often comply when the same intent is reached gradually over many turns, making multi-turn attacks one of the least understood failure modes of large language models. Most automated red-teaming methods treat this as a generation problem: produce attacks that break the model. We argue it is better framed as a search problem: discover, organize, and iteratively refine a diverse archive of attack strategies, producing a structured map ...

📄 EGT-KG: Evidence-Grounded Typed KG Retrieval for Practical Scientific QA with Small Language Models
🗓️ Published: 8/31/2026
🔗 http://arxiv.org/abs/2609.00479v1
👥 Authors: Muran Yu, Jiechao Gao, Yuandong Pan, Barney H. Miao, Andrew C. Lesh, Kincho H. Law (possible past Stanford University affiliation), Jie Wang (possible past Tsinghua University affiliation), Michael D. Lepech
Abstract

For emerging scientific research domains, local Small Language Models (SLMs) are becoming more attractive, as they offer stronger privacy control and more stable deployment pipelines than Large Language Models. However, in practice, scientific question-answering on SLMs often operates under inevitable constraints: small literature collections, fragmented evidence, limited context window and reasoning abilities. We propose the Evidence-Grounded Typed Knowledge Graph (EGT-KG), a retrieval framewor...

📄 Distributed Implicit Harm: A Compositional Safety Blind Spot in MLLM-Based Video Moderation
🗓️ Published: 8/31/2026
🔗 http://arxiv.org/abs/2609.00206v1
👥 Authors: Ruotong Wang, Zihao Zhu, Siwei Lyu (possible past University Of Washington affiliation), Xin Tao (possible past Tencent (China) affiliation), Baoyuan Wu (possible past Tencent (China) affiliation)
Abstract

Despite their growing use in video moderation, multimodal large language models (MLLMs) exhibit a compositional safety blind spot: videos composed of seemingly benign components can convey harmful meaning when interpreted as a whole. We refer to this phenomenon as Distributed Implicit Harm (DIH), where harm arises from relations among components distributed along a decomposition axis of the video, rather than from any single explicit cue. Among many possible axes, we study two representative cas...

📄 ReDeck: Step-Level Render-Grounded Refinement for Document-to-Slide Generation
🗓️ Published: 8/31/2026
🔗 http://arxiv.org/abs/2609.00194v1
👥 Authors: Muzhao Tian, Zezi Zeng, Yifan Yang (possible past Tencent (China) affiliation), Xin Gao, Yan Li (possible past Tencent (China) affiliation), Zisu Huang, Xiaohua Wang, Changze Lv, Mingxi Cheng, Bei Liu, Kai Qiu, Qi Dai, Dong Chen, Yue Dong, Xiaoqing Zheng, Ji Li, Chong Luo (possible past Google (United States) affiliation)
Abstract

Document-to-slide generation is challenging because slides are dense editable artifacts that require both faithful content selection and precise spatial layout. Recent slide agents adopt iterative reflection, but typically follow a monolithic "one version, one feedback" loop: a slide or deck is rewritten, rendered afterward, and critiqued only at the turn boundary. This delayed feedback makes local failures such as overflow, overlap, clipping, and off-canvas placement difficult to attribute and ...

📄 IMPACT: Attention Is the Interaction Map for Scalable Interaction-Aware World Model Training
🗓️ Published: 8/31/2026
🔗 http://arxiv.org/abs/2609.00161v1
👥 Authors: Rongze Tang, Jianjie Fang, Zhaolu Wang, Ziyou Wang, Xvyuan Liu, Haisheng Su, Xin Zhang (possible past Google (United States) affiliation), Wei Wu (possible past Tencent (China) affiliation), Chen Gao, Yong Li (possible past Tsinghua University affiliation), Zhibo Chen (possible past Tencent (China) affiliation)
Abstract

World models have made remarkable progress in action-conditioned future prediction for embodied agents, yet still struggle to model physically plausible interactions. Existing approaches address this limitation by constraining the generation process with external representations encoding motion, geometry, or semantics. Obtaining these spatiotemporally dense representations typically requires auxiliary estimators or manual annotations, limiting training scalability. We instead revisit the trainin...

📄 Dense Process Supervision for Search Agents via Fact Utility Estimation
🗓️ Published: 9/1/2026
🔗 http://arxiv.org/abs/2609.00833v1
👥 Authors: Rongzhi Zhu, Xiangyu Liu, Yi Liu (possible past Google (United States) affiliation), Shuo Zhang (possible past National University Of Defense Technology affiliation), Ruirui Zhang (possible past Tencent (China) affiliation), Rui Wu (possible past Google (United States) affiliation), Tao Jiang (possible past Alibaba Group (China) affiliation), Zequn Sun, Wenhao Xu, Wei Hu
Abstract

Reinforcement learning (RL) for search agents typically relies on outcome rewards. However, it often fails to achieve effective credit assignment, due to the unclear value of intermediate steps. It is hard to separate their contributions from the final result. In this paper, we propose a dense process supervision method based on fact utility estimation, which models the reasoning process as the accumulation of discrete evidence facts. We first extract structured facts from raw observations and o...

📄 HarmoCore: Functional Latent Diffusion for Sparse Reconstruction of Oscillatory Wave Fields
🗓️ Published: 9/1/2026
🔗 http://arxiv.org/abs/2609.00679v1
👥 Authors: Lihao Chen, Xinyu Zhang (possible past Baidu (China) affiliation), Panqi Chen, Lei Cheng, Ting Zhang (possible past Meta (United States) affiliation), Jianlong Li, Shikai Fang
Abstract

Reconstructing oscillatory wave fields from scattered sensors is a severely underdetermined inverse problem. Beyond the challenges of general physical-field reconstruction, wave responses are complex-valued, frequency-sensitive, and highly oscillatory, while costly simulation and sensing often leave only extreme-sparse observations. Existing low-rank, operator, and diffusion approaches are largely designed for real-valued, smoother fields; dense pixel-space diffusion is particularly inefficient ...

📄 Group Adaptive Clipping Policy Optimization
🗓️ Published: 8/31/2026
🔗 http://arxiv.org/abs/2609.00444v1
👥 Authors: Sheng Jia, Xiao Wang (possible past Google (United States) affiliation), Shiva Prasad Kasiviswanathan, Rein Houthooft (possible past University Of California, Berkeley affiliation)
Abstract

Group relative policy optimization for reinforcement learning with verifiable rewards (RLVR) typically uses a fixed importance-sampling (IS) ratio clipping boundary across all rollouts. We identify a key limitation: rare correct rollouts on harder problems and abundant correct rollouts on easier problems are clipped at comparable rates, despite contributing very different learning signals. Rollouts with low group success exhibit larger IS ratios and carry stronger gradient signal for exploration...

📄 Synthetic Worlds for Temporal Evaluation and Knowledge Updating in LLMs
🗓️ Published: 8/31/2026
🔗 http://arxiv.org/abs/2609.00184v1
👥 Authors: Jonathan Zheng, Zirui Shao, Alan Ritter (possible past Carnegie Mellon University affiliation), Wei Xu (possible past Tencent (China) affiliation)
Abstract

Large language models (LLMs) rely on static pretraining corpora, causing their knowledge to become outdated over time. Existing approaches for evaluating knowledge edits either suffer from rapid contamination or rely on counterfactual edits that conflict with rigid existing knowledge. In this work, we propose a synthetic, simulation-driven framework for studying knowledge insertion in LLMs. We introduce {\sc ParallelEvents}, a benchmark of fictional yet realistic future worlds that generates coh...

📄 Sycophantic Agreement Transfers with Neutral Data via Contrastive Preference Optimization
🗓️ Published: 8/31/2026
🔗 http://arxiv.org/abs/2608.31079v1
👥 Authors: Camila Blank, Zhuofan Ying, Christopher Potts (possible past Tencent (China) affiliation), Peter Hase, Jing Huang (possible past Meta (United States) affiliation)
Abstract

Sycophantic agreement refers to a behavior in which language models excessively affirm the user, often at the cost of factual accuracy. Although sycophantic agreement is a well-known failure of model alignment, there is limited understanding of how it emerges from model training. In this work, we demonstrate that sycophantic agreement can emerge as an unintended consequence of widely used contrastive preference optimization objectives. Using the OLMo 3 post-training pipeline, we show that, for v...

📄 Language-Informed Flow Matching for Trend-Guided Structure-Based 3D Molecular Generation
🗓️ Published: 8/31/2026
🔗 http://arxiv.org/abs/2608.31009v1
👥 Authors: Tianyu Gao (possible past Tsinghua University affiliation), Zhikai Su, Jiashu Li, Wenjun Gao, Zichuan Ying, Zhe Zhao (possible past Tencent (China) affiliation), Fei Zhang, Ye Wei
Abstract

Structure-based drug design (SBDD) requires ligands that satisfy both 3D target affinity and 1D chemical validity. Existing controllable generation methods often rely on task-specific fine-tuning or externally imposed sampling-time guidance, adding cost and potentially conflicting with evolving 3D geometric constraints. We propose LiFT, a language-informed cross-modal framework built on Flow Matching for trend-guided 3D molecular generation across both de novo design and scaffold hopping. LiFT u...

📄 A Human-in-the-Loop Autonomous Agent for Industry Time Series Forecasting
🗓️ Published: 8/31/2026
🔗 http://arxiv.org/abs/2608.30976v1
👥 Authors: Xiaoyu Tao, Mingyue Cheng, Ze Guo, Bokai Pan, Qi Liu (possible past Tencent (China) affiliation), Shijin Wang, Enhong Chen (possible past Baidu (China) affiliation)
Abstract

Real-world time-series forecasting is rarely a one-shot model invocation: practitioners must formulate tasks, connect data and models, incorporate domain expertise, assess prediction plausibility, and communicate uncertainty. Specialized forecasting models provide strong numerical predictions but usually operate in fixed pipelines, while general-purpose large language model (LLM) agents often lack forecasting-specific checks, constraints, and stopping rules. We present CastClaw, a human-in-the-l...

📄 Safin-1: Safety from Within through Memory-Native State Evolution
🗓️ Published: 8/31/2026
🔗 http://arxiv.org/abs/2609.00092v1
👥 Authors: Ming Zhang (possible past Peking University affiliation), Kaisen Yang, Shu Yu, Ermo Hua, Zhekai Chen, Cheng Jin, Jingnan Zheng, Yi Zhang (possible past Google (United States) affiliation), Zhongtian Ma, Jiawei Zhou, Sirui Chen, Qiaosheng Zhang, Xiang Wang (possible past Tencent (China) affiliation), Ning Ding (possible past Tsinghua University affiliation), Xia Hu, Bowen Zhou, Youbang Sun, Chaochao Lu
Abstract

Long-horizon complex tasks require foundation models to accumulate information, maintain internal states, and adapt over extended interactions. Safety should be an intrinsic property of the model itself, rather than a behavioral constraint relying solely on external safeguards or post-hoc alignment such as supervised fine-tuning. This motivates Safety from Within, where safety-relevant capabilities are represented and invoked through the model's native computation. We present Safin-1, a family o...

📄 E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation
🗓️ Published: 8/31/2026
🔗 http://arxiv.org/abs/2608.30730v1
👥 Authors: Wei Fan (possible past Tencent (China) affiliation), Xinjie Shen, Xudong Guo, Jianhong Tu, Yang Su, Yinger Zhang, Lianghao Deng, Fengyu Wang, Baohua Dong, Yangqiu Song (possible past Tsinghua University affiliation), Dayiheng Liu
Abstract

Long-horizon agentic tasks go beyond chaining short tasks over more interaction turns. Their evolving dynamic environments and long-range dependencies require Large Language Models (LLMs) to continually explore, learn from experience, and adapt their policies over thousands of steps. We introduce E-Commerce Bench, the first open-source benchmark that integrates multi-round counterpart negotiation and dynamic events into a year-long business operation. Over a 365-day year, an LLM agent concurrent...

📄 Collapsibility of Performance Metrics in Clinical Predictive AI
🗓️ Published: 8/31/2026
🔗 http://arxiv.org/abs/2608.30568v1
👥 Authors: João Matos, Ben Van Calster, Richard D. Riley, Paula Dhiman (possible past University Of Oxford affiliation), Gary S. Collins (possible past University Of Oxford affiliation)
Abstract

Background: Population level assessments of predictive artificial intelligence (AI) can conceal performance disparities across subgroups. Fairness evaluations commonly rely on performance analyses across subgroups. However, some performance metrics are non-collapsible, meaning that the overall population performance value does not equal the weighted average of subgroup specific values. Objective: To examine the collapsibility properties of commonly reported performance metrics in predictive AI...

*Notable papers are those with at least two authors from a "big" AI/ML lab.