📄 Notable* Recent AI/ML arXiv Papers

Last updated just now...

📄 Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening
🗓️ Published: 9/16/2026
🔗 http://arxiv.org/abs/2609.18708v1
👥 Authors: Yizhuo Li, Jianhao Yan, Yun Luo, Zhi Wang, Futing Wang, Rong-Xi Tan, Kanghui Tian, Ganqu Cui (possible past Tsinghua University affiliation), Ning Ding (possible past Tsinghua University affiliation), Peilin Zhao (possible past Tencent (China) affiliation), Yafu Li, Yu Cheng (possible past National University Of Singapore affiliation)
Abstract

In reinforcement learning for large language models, Proximal Policy Optimization (PPO) commonly uses a critic to estimate state values and reduce the variance of policy updates. However, we uncover a systematic failure mode in PPO critics, which we call Value Flattening: state values, estimated from multiple Monte Carlo continuations, change sharply across intermediate states while critic predictions remain comparatively flat. We further observe this phenomenon in a controlled FrozenLake enviro...

📄 ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models
🗓️ Published: 9/16/2026
🔗 http://arxiv.org/abs/2609.18487v1
👥 Authors: Shijie Lian, Bin Yu, Zhaolong Shen, Xiaopeng Lin, Yichao Du, Zhirui Zhang (possible past Tencent (China) affiliation), Laurence T. Yang, Kai Chen (possible past Shanghai Jiao Tong University affiliation)
Abstract

Action tokenizers play a central role in autoregressive vision-language-action (VLA) models, determining both the targets for policy training and the executable commands recovered from predicted tokens. Their fidelity is commonly evaluated using pointwise reconstruction metrics such as mean squared error (MSE), yet small individual errors do not fully characterize how faithfully action adjustments across demonstrations are preserved. After compression, similar actions may still cluster around a ...

📄 Disentangling Long-Term Memory via Latent Neuro-Symbolic Reasoning
🗓️ Published: 9/16/2026
🔗 http://arxiv.org/abs/2609.18461v1
👥 Authors: Cai Ke, Xinghao Chen, Xiaoyu Shen, Keyu Chen, Siyu An, Junnan Dong, Ruifeng Xu, Ruizhi Qiao (possible past Tencent (China) affiliation), Xing Sun (possible past Tencent (China) affiliation)
Abstract

Personalized agents are required to reason over long-term history interactions to infer both explicit preferences and implicit behavioral evidence. While early flat retrieval methods score memory fragments independently and neglect the distributed information, current structured memory frameworks rely on query-agnostic static graphs that fail to capture the context-dependent relations. Crucially, raw textual memories are inherently entangled and noisy, making fine-grained personalization and cro...

📄 BENCHCOMPASS: From Scores to Signals for Training and Harness Decisions in Payment-Domain LLMs
🗓️ Published: 9/16/2026
🔗 http://arxiv.org/abs/2609.18270v1
👥 Authors: Sijie Dong, Wei Ren, Xuanwei Hu, Jiawei Luo, Zifan Wang, Xiaoyun Feng, Hui Cai, Lyuxin Xue, Peng Lu, Jianshe Li, Xin Zhang (possible past Google (United States) affiliation), Wei Wu (possible past Tencent (China) affiliation)
Abstract

Payment operations are a critical financial infrastructure, but the value of large language models in this domain remains unclear because payment rules change quickly, evidence is fragmented, and decisions depend on transaction state, participant role, region, and payment rail. Existing benchmarks do not isolate whether failures come from missing payment-rule knowledge, poor use of supplied evidence, or brittleness under imperfect harness inputs. We introduce BENCHCOMPASS, a payment-domain bench...

📄 WFM: Wiki Foundation Model for Complex Agentic Reasoning
🗓️ Published: 9/16/2026
🔗 http://arxiv.org/abs/2609.18182v1
👥 Authors: Junnan Dong, Linhao Luo, Senlei Zhang, Gong Chen, Taian Guo, Yifei Yu, Rong Tao, Tao Guo, Qian-Wen Zhang, Siyu An, Ruizhi Qiao (possible past Tencent (China) affiliation), Xing Sun (possible past Tencent (China) affiliation)
Abstract

Real-world agents fundamentally require persistent non-parametric knowledge for dynamic reasoning, i.e., long-term memory and retrieval-augmented generation. While graphs have shown reliable advantages in providing structured evidence, the sparse graph representations naturally restrict machine readability and semantic density required for complex agentic workflows. Driven by this limitation, the entire industry is witnessing a paradigm shift from traditional sparse graphs to LLM Wiki, an agent-...

📄 Agora: Git as Shared Memory for Collective AutoResearch
🗓️ Published: 9/16/2026
🔗 http://arxiv.org/abs/2609.18094v1
👥 Authors: Yifan Zhang, Yunheng Zou, Shaokun Zhang, Jian Hu, Hao Zhang (possible past Tencent (China) affiliation), Binfeng Xu, Jan Kautz (possible past Nvidia (United States) affiliation), Yi Dong
Abstract

Autonomous research loops such as AutoResearch show that one coding agent can improve a training setup unattended. Run several of them and each session starts from scratch, so more agents tend to mean more duplicated search rather than more discovery. Agora is a shared memory for such agents: research is recorded as an append-only directed acyclic graph (DAG) stored in Git, so that every claim is a commit anyone can check out and rerun. Each result, insight, hypothesis, verification, and report ...

📄 Verifiable Social Reasoning for LLM Assistants
🗓️ Published: 9/15/2026
🔗 http://arxiv.org/abs/2609.17496v1
👥 Authors: Amir Taubenfeld, Zorik Gekhman, Avigail Grinstein-Dabush, Itay Laish, Ariel Goldstein, Marian Croak, Avinatan Hassidim (possible past Google (United States) affiliation), Yossi Matias (possible past Google (United States) affiliation), Amir Feder (possible past Technion – Israel Institute Of Technology affiliation)
Abstract

LLM assistants are widely used for daily social advice, yet evaluating their social reasoning in such consultation settings remains challenging since (i) it requires setups where the assistant learns about social situations from subjective user narratives, and (ii) social properties, such as others' intentions, typically lack verifiable ground truth. To address these challenges, we introduce Fuse, a multi-agent simulation framework for studying user-mediated social reasoning. In Fuse, a target a...

📄 LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence
🗓️ Published: 9/15/2026
🔗 http://arxiv.org/abs/2609.17488v1
👥 Authors: Xingxuan Zhang, Gang Ren (possible past Google (United States) affiliation), Hao Yuan, Hao Zou, Hongze Tan, Hui Wang, Jianhao Song, Jiansheng Li, Jiayao Zhang, Jinghan Zhang, Kaifang Li, Lang Mo, Li Mao, Mingchao Hao, Nuo Xu, Rui Ding, Ruiji Zhang, Shuyang Li, Siyu Mei, Tianyang Zhang, Weiyang Mu, Yancheng Dong, Yongxian Wei, Yuan Xue (possible past Google (United States) affiliation), Yuanrui Wang, Yue He, Zijia Yang, Ziyun Li, Dongzhe Li, Fuqiang Wang, Jiandong Liu, Jiawei Chen (possible past Tencent (China) affiliation), Jiaxin Du, Kaijie Cheng, Kehan Li, Lei Sun, Linjun Zhou, Ningbo Dai, Qi Wang (possible past Tsinghua University affiliation), Renzhe Xu, Shaoxing Du, Shumeng Yang, Wang Lu, Wenjing Chu, Xiannan Huang, Xiaoyu Lin, Xing Ai, Xinyan Han, Xuanyue Li, Xuanyue Su, Xukun Zhang, Yan Lu, Yaxin Zhang, Yi Qin, Yifei Huang, Yihan Xu, Yongle Lv, Yuanyuan Jiang, Yushan Han, Peng Cui (possible past Tsinghua University affiliation)
Abstract

We introduce LimiX-2, a new model in the LimiX family, developed through model and data scaling guided by our previously established scaling laws. LimiX-2 adopts the Contextual Mechanism Networks (CMNs) paradigm and is pretrained with Context-Conditional Masked Modeling (CCMM). CMNs shifts the organizing principle of in-context learning from target-centric prediction to mechanism-oriented joint modeling. Rather than centering the network on the $p(y \mid x, D_{\mathrm{context}})$ objective of co...

📄 From Transient Prompts to Persistent Control: Scientific Poster Generation via Recursive Semantic-Geometric Contracts
🗓️ Published: 9/15/2026
🔗 http://arxiv.org/abs/2609.17326v1
👥 Authors: Runze Li, Yukun Zhao, Can Xu (possible past Google (United States) affiliation), Yucheng Shen, Shuaiqiang Wang (possible past Baidu (China) affiliation), Jianmin Wu, Lingyong Yan, Dawei Yin (possible past Baidu (China) affiliation)
Abstract

Scientific poster generation distills a multimodal paper into a single-page visual artifact, forcing strict trade-offs between informational coverage and readability under a fixed spatial budget. Existing methods pass plans as transient prompts and validate individual stages in isolation. This strategy causes requirements to drift across content and layout modules, and previous checks to be silently invalidated. We introduce PosterVisor, a control framework that shifts poster generation from tra...

📄 Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations
🗓️ Published: 9/16/2026
🔗 http://arxiv.org/abs/2609.19101v1
👥 Authors: Leon Bergen, Usha Bhalla, Andrew Lee, Barak Widawsky, Linas Nasvytis, Connor Watts, Siddharth Boppana, Sidharth Baskaran, Dron Hazra, Michael Byun, Atticus Geiger (possible past Stanford University affiliation), Owen Lewis, Matthew Kowal, Vasudev Shyam, Thomas Fel, Thomas Mcgrath (possible past Google (United States) affiliation), Ekdeep Singh Lubana, Jack Merullo
Abstract

As models scale, reward hacking becomes more frequent, more sophisticated, and more consequential. Does it leave a telltale signature in model representations? This work analyzes how reward hacking is represented internally in frontier open source LLMs, and how those representations can be used to understand and discover the range of hacking behaviors a model displays. In particular, we find that simple difference of means vectors coherently represent reward hacking in Kimi K3, GLM 5.2, and Qwen...

📄 Physics-based prediction, uncertainty quantification and decision-making for IN718 crystallographic texture intensity across LPBF defocus regimes
🗓️ Published: 9/16/2026
🔗 http://arxiv.org/abs/2609.18863v1
👥 Authors: Yisheng Lu, John Riris, Jie Song (possible past Eth Zurich affiliation), Yao Fu, Jie Chen (possible past Tencent (China) affiliation)
Abstract

Reliable prediction of crystallographic texture in laser powder bed fusion is critical for linking process conditions with anisotropic response and for qualification. However, black-box models may fail under shift and cannot distinguish weak data support from loss of physical validity. This study develops a two-stage physics-based model for <001> || BD (build direction) texture in Inconel 718. Stage 1 maps process variables to melting mode and melt pool geometry. Stage 2 predicts texture by comb...

📄 LIGE-GR: A Smooth Leap from Ranking to Generative Recommendation in the LLM Era
🗓️ Published: 9/16/2026
🔗 http://arxiv.org/abs/2609.18148v1
👥 Authors: Venkat Srinivas, Chenzhang He, Sam Woodmansee, Shawn Lian, Wenjie Hu, Renjie Jiang, Ziheng Huang, Xinyuan Zhang, Zhihao Zheng, Zhuoran Yu (possible past Google (United States) affiliation), Rui Li (possible past Google (United States) affiliation), Lei Yuan, Ziwei Li, Jimmy Jia, Mert Terzihan, Ekrem Kocaguneli, Yiming Liao, Zhichen Zhao, Yue Yin, Yue Weng, Wanlin Ma, Xufeng Cai, Weimiao Wu, Yezhou Huang, Du Zhang, Yukun Ding, Aaron Johnston, Yueming Wang, Zhaojie Gong, Yuting Zhang, Serena Li, Adithya Ganesh, Boying Liu, Haichuan Yang, Xialu Li, Matt Ma, Qunshu Zhang, John Joshua Miller, Praveen Rathinavelu, Cheng Huang, Aadhar Sachdeva, Josh Karns, Andres Aaron Gutierrez, Neil Agarwal, Gustas Pladis, Vladimir Batygin, Gopal Ray, Aditya Priyadarshi, Shantanu Patil, Zhe Wang (possible past Deepmind (United Kingdom) affiliation), Penny Pan, Yiping Han, Arun Singh, Guangdeng Liao, Bi Xue, Xinyao Hu, Yang Song (possible past Stanford University affiliation), Yisong Song, Meihong Wang, Haotian Wu, Deepak Agarwal, Ji Liu (possible past Tencent (China) affiliation)
Abstract

The remarkable success of large language models (LLMs) has provided important inspiration for the next generation of recommender systems. Structurally, recommendation and language generation share a similarity: both aim to produce an ordered sequence that optimizes the user's experience. However, how to precisely absorb the essence of the LLM paradigm into mature industrial recommender systems remains an open problem. There are two challenges. First, it is unclear how to incorporate sequence-l...

📄 Token Latency Fairness: Performance Isolation for Multi-Tenant LLM Serving
🗓️ Published: 9/16/2026
🔗 http://arxiv.org/abs/2609.18112v1
👥 Authors: Dev Bali, Soujanya Ponnapalli, Yichuan Wang, Natacha Crooks, Scott Shenker (possible past University Of California, Berkeley affiliation), Matei Zaharia (possible past University Of California, Berkeley affiliation)
Abstract

LLM serving is typically offered as a shared, multi-tenant service, where high-demand workloads from one client can cause latency SLO violations for others. Existing solutions for performance isolation equalize client throughput in the long run, for example through queueing and batching fairness. However, these approaches do not provide latency isolation guarantees; as a result, well-behaved clients can still experience significant degradation to their token-level latencies. In this paper, we ...

📄 Locating Hidden Failures Makes Long-Horizon Agents More Reliable
🗓️ Published: 9/15/2026
🔗 http://arxiv.org/abs/2609.17930v1
👥 Authors: Salman Rahman, Yubin Kim, Mihir Parmar, A. Ali Heydari, Genglin Liu, Simon A. Lee, Weizhi Zhang, Arian Hosseini, Ahmed A. Metwally, Yuzhe Yang, Baharan Mirzasoleiman (possible past Eth Zurich affiliation), Xin Liu, Pavel Izmailov, Saadia Gabriel, Mark Malhotra, Shwetak Patel, Daniel Mcduff (possible past Google (United States) affiliation), Hamid Palangi
Abstract

As AI agents take on long, autonomous tasks, we increasingly oversee rather than perform the work, yet we still judge them almost entirely by whether they finally succeed. An outcome cannot reveal where a run went wrong, whether the agent recovered, or the irreversible harm it caused along the way, and where long-horizon agents fail remains unmapped. We study $2518$ agent trajectories across software engineering, computer use, and science, close to real deployment, and classify $6967$ mistakes i...

*Notable papers are those with at least two authors from a "big" AI/ML lab.