📄 Notable* Recent AI/ML arXiv Papers

Last updated just now...

📄 ConvergeFlow: Language Flow with Provable Convergence to Token Embeddings
🗓️ Published: 8/24/2026
🔗 http://arxiv.org/abs/2608.23551v1
👥 Authors: Na Li (possible past Tencent (China) affiliation), Yuchen Jiao, Changxiao Cai, Gen Li (possible past University Of Edinburgh affiliation)
Abstract

Recent advances in continuous diffusion and flow-based language models (LMs) have achieved performance competitive with discrete LMs. However, existing continuous frameworks still rely on decoders supervised with cross entropy (CE) because the flow trajectories are not guaranteed to terminate at valid token embeddings. Motivated by this limitation, we introduce \textbf{ConvergeFlow}, an embedding-space flow-based LM, which constrains the data predictor to the convex hull of token embeddings and ...

📄 StrategyBench: Evaluating Explicit Strategy Induction in Large Language Models
🗓️ Published: 8/24/2026
🔗 http://arxiv.org/abs/2608.23475v1
👥 Authors: Jinghan Tan, Yuanzheng Wang (possible past Baidu (China) affiliation), Lu Chen, Zijun Chen (possible past Google (United States) affiliation), Yuqian Wang, Maosong Sun (possible past Tsinghua University affiliation)
Abstract

As large language models are increasingly used in data-scarce and evolving task scenarios, few-shot in-context learning (ICL) has become a key paradigm for task adaptation. However, direct ICL often uses a small set of examples without explicitly abstracting task rules, making it sensitive to example construction. In contrast, human learners often reduce such sensitivity by first summarizing task rules from examples and then applying them to new instances. To evaluate this ability, we propose St...

📄 Apodex 1.1: Scaling Agentic Intelligence for Complex Work
🗓️ Published: 8/24/2026
🔗 http://arxiv.org/abs/2608.23283v1
👥 Authors: Apodex Team, B. An, B. Li, B. Wang, B. Zhang, B. L. Wang, C. Feng, C. Wei, C. Xue, C. Zhang, D. Ng, D. Ye, E. Min, F. Chen, F. Liu, F. Yang, F. Ye, H. Xu, H. Yang, H. Ye, H. Zhang, H. Zhao, J. Li, J. Lin, J. Xia, K. Jin, K. Wang, K. Yang, L. Bing, L. Lei, L. Su, Le. Wang, Lu. Wang, N. Wang, Q. Ren, Q. Yang, R. Li, S. Bai, S. Du, S. Li, S. Lin, S. Nie, S. Wang, S. Zhang, S. Z. Wang, Ta. Q. Fang, Ti. Q. Fang, W. Fang, W. Li, W. Zhang, X. Chen (possible past Google (United States) affiliation), X. Li, X. Tang, X. Wang, X. Xu, X. Zhang, X. Q. Wang, X. Y. Wang, Y. Deng, Y. Gao, Y. Hu, Y. Li, Y. Sui, Y. Wang, Y. Xiao, Y. Zhang, Z. Chen (possible past Google (United States) affiliation), Z. Cheng, Z. Feng, Z. Liang, Z. Zhang
Abstract

General-purpose language models can reason and synthesize knowledge, but complex work also requires sustained interaction with files, information sources, and executable code, together with state maintenance, failure recovery, and verifiable delivery. We call this \emph{working capability}: sustained, verifiable progress toward a real-world objective. Apodex 1.1 develops this capability along two complementary dimensions. \emph{Environment Scaling} expands the diversity and verifiability of exec...

📄 Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data
🗓️ Published: 8/24/2026
🔗 http://arxiv.org/abs/2608.23256v1
👥 Authors: Yinhao Tang, Youqing Fang, Yanan Sun (possible past Tencent (China) affiliation), Jiangning Liu, Ziyi Wang, Xun Zhao (possible past Tencent (China) affiliation), Weiming Zhang, Bin Liu, Kuikun Liu, Wenwei Zhang, Kai Chen (possible past Shanghai Jiao Tong University affiliation)
Abstract

Recent work proposes next-chunk reasoning RL for leveraging no-CoT data---corpora such as worked solutions and textbook derivations that contain reasoning-rich content but lack explicit chain-of-thought annotations. The method trains a model to generate implicit reasoning traces and rewards them by their ability to predict the next chunk of text. While promising, existing evaluations primarily compare against conventional SFT baselines, leaving open whether the gains come from the RL formulation...

📄 LLM-based Agents for Forecasting and Prediction: Methods, Training, Evaluation, and Applications
🗓️ Published: 8/24/2026
🔗 http://arxiv.org/abs/2608.23058v1
👥 Authors: Xiaogang Xu, Jiaqi Tang, Jianmin Chen (possible past Google (United States) affiliation), Yingying Yan, Zhenchao Tang, Xiangxin Zhou, Xiaobin Hu (possible past Tencent (China) affiliation), Wei Wei (possible past Google (United States) affiliation), Jinfeng Wu, Qifeng Chen, Lu Zhou, Jiafei Wu, Zhe Liu, Jianwei Yin, Weimin Zheng
Abstract

Large language models (LLMs) now support forecasting systems that combine language-based reasoning with temporal data, evidence retrieval, external tools, and iterative prediction. We investigate LLM-based forecasting agents, meaning systems in which a language model contributes to a scored prediction about a future or currently unobserved target. We organize architectures into three groups. Standalone LLM workflows operate on encoded time series or event context. Tool- and retrieval-augmented a...

📄 What Proves You Wrong: Benchmarking Language Models on Falsifiable Research Ideation
🗓️ Published: 8/24/2026
🔗 http://arxiv.org/abs/2608.22948v1
👥 Authors: Ziyue Wang, Aomufei Yuan, Yiran Yao, Linli Yao, Hongyao Zuo, Ziwen Gong, Yuanxin Liu, Shicheng Li, Yishuo Cai, Tong Yang (possible past Peking University affiliation), Xu Sun (possible past Peking University affiliation), Xiaohui Li, Haoli Bai
Abstract

Large language models are increasingly used to propose research ideas, yet the prevailing ways of judging such ideas supply no shared decision rule: free-form judging sways with style and position, and scoring against a later paper rewards recovery of one realized trajectory. We introduce a benchmark that carries a proposal from Literature to Test: the Lit2Test benchmark centers on a six-field contract organized around a falsifying outcome, so that every proposal precommits the observation that ...

📄 SEAM: Shot Entity-Attribute Memory for Consistent Short-Drama Generation at Scale
🗓️ Published: 8/24/2026
🔗 http://arxiv.org/abs/2608.22725v1
👥 Authors: Jiaqi Liu, Maolin Ran, Xiaoyang Lu, Jian Wang (possible past Baidu (China) affiliation), Weiwen Liu, Jianghao Lin, Yong Yu (possible past Shanghai Jiao Tong University affiliation), Weinan Zhang (possible past Shanghai Jiao Tong University affiliation)
Abstract

Short-drama generation has grown into a large, industrialized pipeline, and as it scales from isolated shots to the episode level, visual continuity has become a critical bottleneck. Current agent frameworks generate each shot in isolation, so context drifts across shots and props, character posture, and blocking turn inconsistent. Once assembled, these small discrepancies amplify into severe visual breaks. We present SEAM (Shot Entity-Attribute Memory), a training-free, model-agnostic memory gr...

📄 The Laws of Context Allocation: Causal Measurement and Closed-Loop Orchestration in Generative Search
🗓️ Published: 8/24/2026
🔗 http://arxiv.org/abs/2608.23252v1
👥 Authors: Peiyang Liu, Xi Wang (possible past Tsinghua University affiliation), Di Liang, Wei Ye (possible past Meta (United States) affiliation)
Abstract

As Retrieval-Augmented Generation (RAG) shifts toward diverse portfolio generation, it is stymied by two critical bottlenecks: flawed measurement of evidence utilization, and suboptimal context budget allocation. We resolve both sequentially. To resolve measurement, we expose a pervasive ``diagnostic illusion'': standard relevance proxies fail catastrophically on hard negatives. We replace them with an efficient causal leave-one-out probe that accurately isolates generative reliance and formal...

📄 Learning Generalizable Behaviors for Terminal Agents
🗓️ Published: 8/23/2026
🔗 http://arxiv.org/abs/2608.22631v1
👥 Authors: Yihang Yao, Bo Pang, Xuan Phi Nguyen, Ding Zhao (possible past Google (United States) affiliation), Shafiq Joty, Semih Yavuz (possible past Google (United States) affiliation)
Abstract

Terminal agents are a compelling application of large language models (LLMs), with the potential to integrate deeply into users' daily workflows. Reinforcement learning (RL) is a key technique for improving their capabilities, making scalable training environments a central challenge. Since public real-user interaction data are scarce, synthetic environments provide a practical alternative, but often suffer from domain gaps and limited fidelity, leading to poor generalization. Existing work main...

📄 Scale-invariant Optimal Sampling for Rare-events Data with Sparse Models
🗓️ Published: 8/23/2026
🔗 http://arxiv.org/abs/2608.22597v1
👥 Authors: Jing Wang (possible past Google (United States) affiliation), Haiying Wang, Qiang Zhang (possible past Tsinghua University affiliation), Hao Helen Zhang
Abstract

Subsampling is effective in tackling computational challenges for massive data with rare events. Overly aggressive subsampling may adversely affect estimation efficiency, and optimal subsampling is essential to mitigate the information loss. However, existing optimal subsampling probabilities depend on data scales, and some scaling transformations may result in inefficient subsamples. This problem is more significant when there are inactive features, because their influence on the subsampling pr...

📄 MCP-Universe RL: A Framework for Training MCP Tool-Use Agents via Reinforcement Learning
🗓️ Published: 8/23/2026
🔗 http://arxiv.org/abs/2608.22167v1
👥 Authors: Ziyang Luo, Yan Yang (possible past Google (United States) affiliation), Xiangru Jian, Ziji Shi, Xiaoqiang Lin, Jun Hao Liew, Silvio Savarese (possible past Stanford University affiliation), Junnan Li
Abstract

Reinforcement learning (RL) has become an effective way to improve the tool-use ability of large language models (LLMs), but most existing RL frameworks stop at the policy update. For every new domain, the user is left with two hard systems problems: standing up an isolated environment for each of hundreds of concurrent trajectories and connecting it to training, and scheduling the rollout so that the GPU stays busy across long, multi-turn episodes that spend much of their time stalled on slow t...

📄 Decoupled Physical Modeling and Execution for Physics Reasoning
🗓️ Published: 8/22/2026
🔗 http://arxiv.org/abs/2608.22126v1
👥 Authors: Ye Zhang (possible past Google (United States) affiliation), Xuehang Guo, Rui Pan, Pengfei Yu (possible past Tsinghua University affiliation), Denghui Zhang, Manling Li, Qingyun Wang
Abstract

Physics reasoning requires constructing a consistent model of the underlying physical system rather than relying solely on symbolic or formula-based manipulation. Although large language models have shown strong ability in solving math and coding problems, they still struggle with physics problems, as these problems entangle the physical modeling process with mathematical calculations. Humans approach physics by first building a representation of the system before performing calculations. Inspir...

*Notable papers are those with at least two authors from a "big" AI/ML lab.