📄 Notable* Recent AI/ML arXiv Papers

Last updated just now...

📄 CogEvol: Towards Efficient and Reliable Learning Environment Generation
🗓️ Published: 8/31/2026
🔗 http://arxiv.org/abs/2608.30968v1
👥 Authors: Shangqing Tu, Daniel Zhang-Li, Yucheng Wang, Shiyu Gan, Yanpeng Wang, Huiqiang Rong, Mofei Chen, Shen Yang, Yini Chen, Yinuo Duan, Haoxuan Li, Binglin Liu, Ye He, Danqi Zheng, Zhanxin Hao, Yuxuan Wu, Mengting Tao, Yuqiu Liu, Jifan Yu, Juanzi Li, Bin Xu, Lei Hou (possible past Tsinghua University affiliation), Huiqin Liu, Yu Zhang (possible past Google (United States) affiliation)
Abstract

We present CogEvol, a family of models trained specifically for Learning Environment Generation: turning a course brief into a finished learning artifact (structured-JSON slides or self-contained interactive HTML pages) in a single pass. Across 220k production requests, CogEvol completes a slide in a median of 17 seconds and an interactive page in 59, replacing minutes-long multi-turn agent scaffolding. Reliability is enforced rather than hoped for: a production-grounded data pipeline turns real...

📄 LOCI: A Locator-Critic with Refinement Loop
🗓️ Published: 8/31/2026
🔗 http://arxiv.org/abs/2608.30959v1
👥 Authors: Walid Bousselham, Mathilde Caron, Arsha Nagrani (possible past University Of Oxford affiliation), Cordelia Schmid (possible past Google (United States) affiliation)
Abstract

Vision-Language Models (VLMs) still struggle on tasks requiring complex visual understanding. We argue that the core issue is not high-level reasoning, but instead failing to locate critical details in the image. Due to this shortcoming, VLMs generate often plausible but incorrect reasoning based on flawed perceptual grounding. To address this, we propose Locator-Critic (LOCI), a training-free framework that decouples visual search from evidence verification. LOCI employs a Locator agent to prop...

📄 CAER: Causal Action Effect Reweighting for World Model Training
🗓️ Published: 8/31/2026
🔗 http://arxiv.org/abs/2608.30897v1
👥 Authors: Jianjie Fang, Xvyuan Liu, Ziyou Wang, Rongze Tang, Zhaolu Wang, Zhuohang Li, Xin Zhang (possible past Google (United States) affiliation), Haisheng Su, Chen Gao, Wei Wu (possible past Tencent (China) affiliation), Xinlei Chen (possible past Tsinghua University affiliation), Yong Li (possible past Tsinghua University affiliation)
Abstract

World models are becoming core infrastructure for embodied intelligence, with action-conditioned video generation providing controllable predictions of how scenes evolve after agent interventions. Yet existing models are commonly trained with space-time-uniform mean squared error, allowing abundant background tokens to dominate the gradient while sparse interaction dynamics remain under-optimized; such uniform fitting rewards reconstructing appearance rather than learning how actions change the ...

📄 Collapsibility of Performance Metrics in Clinical Predictive AI
🗓️ Published: 8/31/2026
🔗 http://arxiv.org/abs/2608.30568v1
👥 Authors: João Matos, Ben Van Calster, Richard D. Riley, Paula Dhiman (possible past University Of Oxford affiliation), Gary S. Collins (possible past University Of Oxford affiliation)
Abstract

Background: Population level assessments of predictive artificial intelligence (AI) can conceal performance disparities across subgroups. Fairness evaluations commonly rely on performance analyses across subgroups. However, some performance metrics are non-collapsible, meaning that the overall population performance value does not equal the weighted average of subgroup specific values. Objective: To examine the collapsibility properties of commonly reported performance metrics in predictive AI...

📄 TuringLLM: Efficiently Scaling Foundation Models Toward Physical AI
🗓️ Published: 8/31/2026
🔗 http://arxiv.org/abs/2608.30567v1
👥 Authors: Yuheng Zhang, Yizhao Wang, Da Zhu, Hua Zhou, Yue He, Jiahui Hu, Shaman Tang, Hanlin Chen, Yuhua Wei, Anhua Liu, Shuang Su (possible past Baidu (China) affiliation), Rui Xin, Mingyuan Wang, Minghao Li, Haojie Yang, Siqi Liu (possible past University Of Oxford affiliation), Jianlei Zheng, Weichao Huang, Qiman Wu (possible past Baidu (China) affiliation), Hang Zhang (possible past Amazon (United States) affiliation), Honggou Yang, Xianming Liu (possible past Meta (United States) affiliation)
Abstract

We present Turing-20B-A2B, a 20B-parameter Mixture-of-Experts language model that activates approximately 2B parameters per token, designed for long-context and latency-sensitive physical AI applications. The model adopts Quantile Routing in a dynamic top-k configuration, enabling token-adaptive expert allocation while maintaining balanced expert utilization and a controlled average compute budget. During deployment, we further apply capacity-constrained routing to prompt prefill for more regula...

📄 SPARK: Skeleton-Guided Reasoning Synthesis from Large-Scale Scientific Literature
🗓️ Published: 8/31/2026
🔗 http://arxiv.org/abs/2608.30214v1
👥 Authors: Yu Li (possible past Tencent (China) affiliation), Wei Li (possible past Peking University affiliation), Xin Gao, Mengyuan Sun, Xiaoyang Wang, Qizhi Pei, Lijun Wu
Abstract

Scientific reasoning remains challenging for open-source models, largely due to the lack of high-quality scientific reasoning data. Existing datasets are often dominated by factual recall or formulaic problem solving, with limited emphasis on mechanism understanding, evidence-grounded reasoning, and hypothesis evaluation. To address this, we introduce SPARK (Scientific Paper Abstracted Reasoning sKeleton), a paper-oriented synthesis framework built on Sci-Base, a large-scale corpus of research p...

📄 Reachability-Based Capability Confinement for LLM Agents under Indirect Prompt Injection
🗓️ Published: 8/30/2026
🔗 http://arxiv.org/abs/2608.30041v1
👥 Authors: Wujie Xiong, Rabimba Karanjai, Yang Lu (possible past Meta (United States) affiliation), Weidong Shi, Lei Xu (possible past Tsinghua University affiliation)
Abstract

Large language model agents place outputs from external skills into their execution context, allowing attacker-controlled data to influence later privileged actions. Existing defenses mainly classify untrusted content or authorize proposed operations. They do not directly address how an agent's future authority should change once untrusted data enters its state. We present SkillGuard, a harness-level enforcement layer that treats this event as contamination and restricts future capabilities to d...

📄 IndicDetect: Evaluating Cross-Lingual LLM-Generated Text Detection for Hindi, Telugu, and Tamil
🗓️ Published: 8/30/2026
🔗 http://arxiv.org/abs/2608.29919v1
👥 Authors: Bhaskar Ganesh Devalla, Junchao Wu, Nilesh Dokuparthi, Greeshma Yaluru, Tatiana Muniz Rodriguez, Lidia S. Chao (possible past Tencent (China) affiliation), Derek F. Wong (possible past Tencent (China) affiliation)
Abstract

The rapid proliferation of LLMs has further heightened the need to develop dependable AI-generated text detection, especially beyond English. Nevertheless, current benchmarks pay little attention to Indic languages and test detectors in idealized settings that do not represent the real world. We present a generalized benchmark for AI-generated text detection in Hindi, Telugu, and Tamil, which we call IndicDetect, designed to assess the robustness of detectors under realistic distribution shifts....

📄 FRAMEWORKERS: A Dynamic Multi-Agent Framework for AI-Generated Video Production
🗓️ Published: 8/30/2026
🔗 http://arxiv.org/abs/2608.29814v1
👥 Authors: Zhendong Li, Lei Sun, Letian Shi, Deheng Zhang, Ruibo Ming, Mengshun Hu, Dannong Xu, Jian Wang (possible past Baidu (China) affiliation), Danda Paudel, Luc Van Gool (possible past Google (United States) affiliation), Jinjin Gu
Abstract

Modern video generators excel at synthesizing individual clips, but complete video production requires coordinating a long sequence of interdependent creative steps, including scripting, storyboarding, generation, and editing. It further demands persistent asset management and dynamic task orchestration as intermediate outputs, dependencies, and execution states evolve over time. Existing automated systems typically rely on rigid pipelines that are difficult to adapt to diverse inputs and changi...

📄 Higher-Dimensional Rotary Position Embedding
🗓️ Published: 8/30/2026
🔗 http://arxiv.org/abs/2608.29715v1
👥 Authors: Yixing Li, Ruobing Xie (possible past Tencent (China) affiliation), Yudong Zhang, Yushi Bai, Samm Sun, Yu Cheng (possible past National University Of Singapore affiliation)
Abstract

Transformers rely on position embedding mechanisms in long context modeling in most cases. Rotary Position Embedding (RoPE) embeds positional information with independent 2D rotations, forming relative position terms in self-attention. However, its pairwise, block-based, and decoupled structure limits deep mixing and robustness across channels. We propose HD-RoPE, which extends RoPE from independent 2D rotations to higher-dimensional rotations and introduces a Paley-I orthogonal basis to obtain ...

📄 Sycophantic Agreement Transfers with Neutral Data via Contrastive Preference Optimization
🗓️ Published: 8/31/2026
🔗 http://arxiv.org/abs/2608.31079v1
👥 Authors: Camila Blank, Zhuofan Ying, Christopher Potts (possible past Tencent (China) affiliation), Peter Hase, Jing Huang (possible past Meta (United States) affiliation)
Abstract

Sycophantic agreement refers to a behavior in which language models excessively affirm the user, often at the cost of factual accuracy. Although sycophantic agreement is a well-known failure of model alignment, there is limited understanding of how it emerges from model training. In this work, we demonstrate that sycophantic agreement can emerge as an unintended consequence of widely used contrastive preference optimization objectives. Using the OLMo 3 post-training pipeline, we show that, for v...

📄 Language-Informed Flow Matching for Trend-Guided Structure-Based 3D Molecular Generation
🗓️ Published: 8/31/2026
🔗 http://arxiv.org/abs/2608.31009v1
👥 Authors: Tianyu Gao (possible past Tsinghua University affiliation), Zhikai Su, Jiashu Li, Wenjun Gao, Zichuan Ying, Zhe Zhao (possible past Tencent (China) affiliation), Fei Zhang, Ye Wei
Abstract

Structure-based drug design (SBDD) requires ligands that satisfy both 3D target affinity and 1D chemical validity. Existing controllable generation methods often rely on task-specific fine-tuning or externally imposed sampling-time guidance, adding cost and potentially conflicting with evolving 3D geometric constraints. We propose LiFT, a language-informed cross-modal framework built on Flow Matching for trend-guided 3D molecular generation across both de novo design and scaffold hopping. LiFT u...

📄 A Human-in-the-Loop Autonomous Agent for Industry Time Series Forecasting
🗓️ Published: 8/31/2026
🔗 http://arxiv.org/abs/2608.30976v1
👥 Authors: Xiaoyu Tao, Mingyue Cheng, Ze Guo, Bokai Pan, Qi Liu (possible past Tencent (China) affiliation), Shijin Wang, Enhong Chen (possible past Baidu (China) affiliation)
Abstract

Real-world time-series forecasting is rarely a one-shot model invocation: practitioners must formulate tasks, connect data and models, incorporate domain expertise, assess prediction plausibility, and communicate uncertainty. Specialized forecasting models provide strong numerical predictions but usually operate in fixed pipelines, while general-purpose large language model (LLM) agents often lack forecasting-specific checks, constraints, and stopping rules. We present CastClaw, a human-in-the-l...

📄 E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation
🗓️ Published: 8/31/2026
🔗 http://arxiv.org/abs/2608.30730v1
👥 Authors: Wei Fan (possible past Tencent (China) affiliation), Xinjie Shen, Xudong Guo, Jianhong Tu, Yang Su, Yinger Zhang, Lianghao Deng, Fengyu Wang, Baohua Dong, Yangqiu Song (possible past Tsinghua University affiliation), Dayiheng Liu
Abstract

Long-horizon agentic tasks go beyond chaining short tasks over more interaction turns. Their evolving dynamic environments and long-range dependencies require Large Language Models (LLMs) to continually explore, learn from experience, and adapt their policies over thousands of steps. We introduce E-Commerce Bench, the first open-source benchmark that integrates multi-round counterpart negotiation and dynamic events into a year-long business operation. Over a 365-day year, an LLM agent concurrent...

📄 Compact and Infinite-Order Error Analysis for Null-Space SVD Estimation
🗓️ Published: 8/31/2026
🔗 http://arxiv.org/abs/2608.30374v1
👥 Authors: Xin Li (possible past Google (United States) affiliation), Jonathan Cohen (possible past Nvidia (United States) affiliation), Rami Puzis
Abstract

We study null-space estimation from a noisy matrix. For a simple left null space, we first derive an exact compact expression for the error of the smallest left singular vector. We then give an all-order series for the SVD vector and projector, followed by compact and consistently truncated series forms for the fixed-realization empirical risk and conditional population generalization risk. The recursion extends to a multiple-dimensional null space by following the complete invariant subspace. T...

📄 Certified Safety Radii in Forecast-Error Space for Wasserstein Distributionally Robust Small Signal Stability-Constrained AC Optimal Power Flow via Lifted Spectrahedral Containment
🗓️ Published: 8/31/2026
🔗 http://arxiv.org/abs/2608.30201v1
👥 Authors: Ziqi Zhang (possible past Tencent (China) affiliation), Xi Chen (possible past University Of California, Berkeley affiliation)
Abstract

Directly robustifying small-signal stability in AC optimal power flow is challenging since the stability boundary in the original uncertainty space is implicit, highly nonconvex, and changes with the operating decision. This paper exploits an alternative geometry. For a fixed model-specific stability certificate admitting suitable physical lifts, the small-signal stability requirement becomes an affine positive semidefinite constraint in the lifted variables, thereby defining a convex certified ...

*Notable papers are those with at least two authors from a "big" AI/ML lab.