πŸ“„ Notable* Recent AI/ML arXiv Papers

Last updated just now...

πŸ“„ HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark
πŸ—“οΈ Published: 8/13/2026
πŸ”— http://arxiv.org/abs/2608.13555v1
πŸ‘₯ Authors: Dairu Liu, Zekun Qi, Jiayu Zeng, Ruixi Yu, Yu Guan, Yintianrun Zhang, Xuchuan Chen, Sikai Liang, Zekai Li, Chenghuai Lin, Xinqiang Yu, Wenyao Zhang, He Wang (possible past Stanford University affiliation), Li Yi (possible past Stanford University affiliation)
Abstract

Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with what people perceive in videos. Kinematic errors average per-frame pose differences but miss the physical artifacts that matter most, particularly unstable support and incorrect contacts such as foot skating and mistimed touch-downs. Meanwhile, widely used test suites are small and lack the diversity needed to stress contact-rich, long-horizon behaviors. We introduce HumanTracker to...

πŸ“„ AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (v1.1)
πŸ—“οΈ Published: 8/13/2026
πŸ”— http://arxiv.org/abs/2608.13492v1
πŸ‘₯ Authors: Alayaworld Team, Kaipeng Zhang (possible past Tencent (China) affiliation), Chuanhao Li, Yifan Zhan, Yongtao Ge, Yuanyang Yin, Jiaming Tan, Kang He, Liaoyuan Fan, Mingliang Zhai, Ruicong Liu, Xiaojie Xu, Xuangeng Chu, Zhen Li (possible past Google (United States) affiliation), Zhengyuan Lin, Zhixiang Wang, Zian Meng, Zihui Gao
Abstract

This report presents an improved version of AlayaWorld. While the backbone architecture, chunk-wise autoregressive generation scheme, and training data remain unchanged from the previous release, we substantially revise how conditioning signals are represented and integrated into the model. The new design is guided by a simple principle: conditioning signals should match the generated content as closely as possible in both latent representation and temporal structure. To this end, we make two ma...

πŸ“„ Synthetic Persona Pretraining: Alignment from Token Zero
πŸ—“οΈ Published: 8/13/2026
πŸ”— http://arxiv.org/abs/2608.13482v1
πŸ‘₯ Authors: Julian Minder, Viktor Moskvoretskii, Raghav Singhal, Difan Jiao, Andy Arditi, Shaobo Cui, Yiderigun Borjigin, Kartik Bali, Stefan Krsteski, Harsh Raj, Huu Nguyen, Jannik Brinkmann, Ashton Anderson (possible past Stanford University affiliation), Roland Aydin, Robert West (possible past Stanford University affiliation)
Abstract

As language-model-based AI is increasingly deployed in autonomous settings, aligning its goals and values with those of humans becomes critical. Today, alignment, and the assistant identity itself, are typically introduced only after pretraining, once behavioral priors are already established. This can make values a thin overlay, rather than deeply rooted, and facilitate subsequent misalignment. Pursuing a different paradigm, we introduce Synthetic Persona Pretraining (SPP), which installs the d...

πŸ“„ Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
πŸ—“οΈ Published: 8/13/2026
πŸ”— http://arxiv.org/abs/2608.13417v1
πŸ‘₯ Authors: Yiwei Li, Wanli Yang, Hexiang Tan, Xiangzhou Huang, Zhengyu Chen, Ziran Li, Borun Chen, Shanglin Lei, Huaisheng Zhu, Hao Tian (possible past Baidu (China) affiliation), Fei Sun (possible past Meta (United States) affiliation), Xunliang Cai, Jingang Wang
Abstract

Autonomous agents are increasingly capable of improving models, systems, and other technical artifacts through long-horizon experimentation. To understand the current state of this capability, however, evaluation must go beyond final scores, which neither reveal where progress is gained or lost nor indicate whether accumulated experience improves later decisions. We therefore present a systematic evaluation of seven frontier models on 36 long-horizon tasks based on a new framework that uses rule...

πŸ“„ Better Decomposition, Free Aggregation: A Synthesizer-Folding Framework for Multilingual Multi-Hop Question Answering
πŸ—“οΈ Published: 8/13/2026
πŸ”— http://arxiv.org/abs/2608.13160v1
πŸ‘₯ Authors: Yilin Wang (possible past Google (United States) affiliation), Yuchun Fan, Weidong Bao (possible past National University Of Defense Technology affiliation), Zili Wei, Shi Feng, Tong Xiao, Zhengtao Yu, Jingbo Zhu
Abstract

Multilingual retrieval-augmented generation (mRAG) equips large language models with access to globally distributed external knowledge for complex multilingual question answering. Recent approaches either translate retrieved documents into English or the query language to bridge the cross-lingual semantic gap, or decompose a complex query into sub-questions and aggregate the intermediate reasoning process. However, both lines of work suffer from two limitations. First, one-size-fits-all translat...

πŸ“„ Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing
πŸ—“οΈ Published: 8/13/2026
πŸ”— http://arxiv.org/abs/2608.13156v1
πŸ‘₯ Authors: Sheng Ren, Yadong Wang, Naiqiang Tan, Jiangang Kong, Jun Fang, Rui Liu, Jun Wang (possible past Tencent (China) affiliation), Kai Chen (possible past Shanghai Jiao Tong University affiliation), Lipeng Liang, Xiang Chen (possible past Tencent (China) affiliation)
Abstract

Pre-norm is the standard normalization placement in modern Transformers because it facilitates joint optimization of full-depth models. We ask whether this preference persists when depth is introduced through a curriculum. In curriculum depth growth, each appended block receives the boundary representation produced by a trained prefix, making normalization placement relevant to forward conditioning. We therefore test whether placement and training curriculum interact. In a controlled distillatio...

πŸ“„ EEG-PRIME: Prototype-Aligned Representation Learning with Multi-Level Conditioning for EEG Decoding
πŸ—“οΈ Published: 8/13/2026
πŸ”— http://arxiv.org/abs/2608.13072v1
πŸ‘₯ Authors: Shuailei Zhang, Muyun Jiang, Wei Zhang (possible past Tsinghua University affiliation), Jinbo Chen, Zhiwei Guo, Yong Li (possible past Tsinghua University affiliation), Yi Ding, Cuntai Guan
Abstract

Electroencephalography (EEG) decoding models often generalize poorly across datasets and subjects due to domain shifts in acquisition protocols and individual neurophysiology. We propose EEG-PRIME, a two-stage EEG foundation model for cross-dataset multi-task decoding. EEG-PRIME combines masked pretraining with prototype-aligned instruction tuning to enable instruction-aware and subject-invariant decoding across diverse BCI paradigms. During pretraining, an EEG encoder learns transferable repres...

πŸ“„ From Local Mismatch to Global Impact: Optimizing Cache Reuse Policy for Efficient Diffusion
πŸ—“οΈ Published: 8/13/2026
πŸ”— http://arxiv.org/abs/2608.13043v1
πŸ‘₯ Authors: Xichen Ye, Yifan Wu (possible past Carnegie Mellon University affiliation), Zhikang Xie, Xiangyu Yue (possible past University Of California, Berkeley affiliation), Cheng Jin, Weizhong Zhang
Abstract

Diffusion models have achieved dominant performance in visual generation but suffer from substantial inference overhead. While cache-based acceleration has emerged as a promising solution, existing policies rely on local similarity heuristics, which we identify as being significantly misaligned with final generation quality. This discrepancy stems from the non-uniform propagation and accumulation of errors along the denoising trajectory. To address this, we propose Global-Impact Cache (GCache). ...

πŸ“„ FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving
πŸ—“οΈ Published: 8/13/2026
πŸ”— http://arxiv.org/abs/2608.12932v1
πŸ‘₯ Authors: Zekai Li, Yihao Liang, Hongfei Zhang, Jian Chen (possible past Baidu (China) affiliation), Yesheng Liang, Zhijian Liu (possible past Massachusetts Institute Of Technology affiliation)
Abstract

Vision-Language-Action (VLA) models promise to bring end-to-end reasoning to autonomous driving, but their computational cost remains far too high for real-time control. The core challenge is structural: VLA inference is not a single bottleneck but a cascade of four. Visual encoding wastes compute on overlapping video frames; language-model prefill recomputes context that could be carried over from the previous timestep; reasoning tokens are generated serially despite low entropy; and flow-match...

πŸ“„ Discovering Efficient and Explainable Communication Topologies for LLM-based Multi-Agent Systems via Causal Inference
πŸ—“οΈ Published: 8/13/2026
πŸ”— http://arxiv.org/abs/2608.12921v1
πŸ‘₯ Authors: Junzhi Li, Peng He (possible past Tencent (China) affiliation), Qirui Ji, Wei Wang (possible past University Of Oxford affiliation), Lixiang Liu, Chuxiong Sun
Abstract

The performance of large language model (LLM)-based multi-agent systems (MAS) largely depends on effective communication topologies. Existing topology generation methods, however, typically learn communication topologies through black-box optimization driven solely by task-level rewards. While effective, such optimization provides little insight into why particular communication edges are selected, making it difficult to identify the critical communication subgraphs responsible for successful co...

πŸ“„ Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories
πŸ—“οΈ Published: 8/13/2026
πŸ”— http://arxiv.org/abs/2608.12847v1
πŸ‘₯ Authors: Yifei Li, Heng Wang, Lingling Zhang (possible past Google (United States) affiliation), Muye Huang, Xinyu Zhang (possible past Baidu (China) affiliation), Jiashuai Liu, Hang Yan, Rongman Xu
Abstract

Retrieval can identify a past trajectory that may matter, yet it does not specify how an acting agent should use that trajectory after users, entities, constraints, or environment state have changed. We identify this post-retrieval reuse step as a distinct bottleneck for long-horizon trajectory memory and formulate an evaluation framework that holds candidate retrieval, target state, model, decoding, and tool budget fixed while varying the support delivered to the agent. We instantiate the frame...

πŸ“„ General Probabilities of Causation with Causal Knowledge
πŸ—“οΈ Published: 8/12/2026
πŸ”— http://arxiv.org/abs/2608.12657v1
πŸ‘₯ Authors: Xin Shu, Zhen Lei (possible past Beijing Academy Of Artificial Intelligence affiliation), Ang Li (possible past Google (United States) affiliation)
Abstract

Probabilities of causation (PoCs) characterize individual causal responses that cannot be directly observed and therefore generally require partial identification. Tian and Pearl first derived theoretically sharp bounds for binary PoCs, including the probability of necessity (PN), the probability of sufficiency (PS), and the probability of necessity and sufficiency (PNS). Mueller et al. subsequently tightened the bounds for binary PNS by incorporating causal information encoded in covariates and...

πŸ“„ DiG-bench: Discovery in Games
πŸ—“οΈ Published: 8/12/2026
πŸ”— http://arxiv.org/abs/2608.12593v1
πŸ‘₯ Authors: Ruairidh M. Battleday, Kai Sandbrink, Jimi Cullen-Drohan, Zihan Yan (possible past Google (United States) affiliation), Timothy Muller, Clare Maguire, Ales Kubicek, Fraser Greenlee-Scott, Sukrit Sumant, Tri Dao, JΓΌrgen Schmidhuber, Michal Valko, Joshua Tenenbaum, Thomas L. Griffiths (possible past University Of California, Berkeley affiliation), Zeb Kurth-Nelson, James C. R. Whittington (possible past University Of Oxford affiliation)
Abstract

Discovery---formulating novel generalizations---is a central part of the scientific process. Despite its importance, there is a gap in the current AI benchmark landscape, with few benchmarks directly probing the capacity for discovering new knowledge with experimentation in controlled environments where the objective is unknown. To address this gap, we release a new benchmark: DiG-bench (Discovery in Games). DiG-bench consists of a set of 70 independent games. Each game is encoded as a short str...

πŸ“„ Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams
πŸ—“οΈ Published: 8/12/2026
πŸ”— http://arxiv.org/abs/2608.12262v1
πŸ‘₯ Authors: Weihao Bo, Shan Zhang, Yanpeng Sun, Jie Liu (possible past Tencent (China) affiliation), Yongke Yao, Jinhao Du, Wei He (possible past Baidu (China) affiliation), Kai Zou, Zechao Li, Jingdong Wang (possible past Baidu (China) affiliation)
Abstract

Multimodal Large Language Models (MLLMs) have been growing the capability for scientific writing and collaboration. For example, OpenAI Prism is a free workspace for scientific writing and collaboration. One important feature in Prism is turning scientific diagrams directly into LaTeX TikZ code. In this paper, we build a benchmark, Diagram-MMU, a multi-modal benchmark designed to assess MLLMs' ability for scientific diagram parsing and understanding. Diagram-MMU features 3.7k curated diagrams an...

πŸ“„ One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL
πŸ—“οΈ Published: 8/12/2026
πŸ”— http://arxiv.org/abs/2608.12253v1
πŸ‘₯ Authors: Simon Yu, Nicholas Tomlin, Marwa Abdulhai, Ximing Lu, Derek Chong, Abe Hou, Dilara Soylu, Sergey Levine (possible past University Of Washington affiliation), Christopher D. Manning (possible past Stanford University affiliation), Weiyan Shi
Abstract

Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior. We show that this approach systematically fails to generalize, and trace the failure to simulator collapse: because the simulator LLM is mode-collapsed, an LLM policy trained against it overfits to narrow strategies that exploit the simulator's dominant mode, and such a policy transfers poorly to unseen simulators and real users. We formalize this collapse theo...

πŸ“„ Intern-S2-Preview: Scientific Agentic Foundation Model
πŸ—“οΈ Published: 8/13/2026
πŸ”— http://arxiv.org/abs/2608.13505v1
πŸ‘₯ Authors: Lei Bai, Jiaqi Cao, Chiyu Chen, Guanzhou Chen, Kai Chen (possible past Shanghai Jiao Tong University affiliation), Guangran Cheng, Erfei Cui, Xuanlang Dai, Shengyuan Ding, Shangheng Du, Yanhui Duan, Yue Fan, Youqing Fang, Quan Gan, Yuanyuan Gao, Jiaye Ge, Lixin Gu, Yuzhe Gu, Qipeng Guo, Junjun He, Xin Hong, Ming Hu, Zhouqi Hua, Haian Huang, Junhao Huang, Zixian Huang, Minxi Jin, Lingkai Kong, Alexander Lam, Zehao Li, Zonglin Li, Tianhao Liang, Dahua Lin, Junyao Lin, Tianyang Lin, Zhouhan Lin, Jiangning Liu, Jin Liu, Kuikun Liu, Wenran Liu, Yifei Liu, Yuhong Liu, Yuhong Liu, Zhoumianze Liu, Ziyan Liu, Ziyu Liu, Haijun Lv (possible past Baidu (China) affiliation), Han Lv, Chengqi Lyu, Le Ma, Ningsheng Ma, Zerun Ma, Haoyang Peng, Runyu Peng, Jifei Shan, Zixin Shang, Kou Shi, Xiang Shi, Qisheng Su, Xuerui Su, Hao Sun, Xiao Sun, Yanan Sun (possible past Tencent (China) affiliation), Yu Sun (possible past Baidu (China) affiliation), Huanze Tang, Yinghao Tang, Wenhui Tian, Zhongbo Tian, Bingli Wang, Haomin Wang, Jiarui Wang, Jingzhi Wang, Rui Wang (possible past Tencent (China) affiliation), Xiquan Wang, Yi Wang, Zhecan Wang, Ziyi Wang, Zun Wang, Rubin Wei, Lianyi Wu, Wen Wu, Yue Wu, Yuhan Wu, Zhenyu Wu, Zijian Wu, Shuhao Xing, Jun Xu (possible past Google (United States) affiliation), Xingle Xu, Xuenan Xu, Xiangchao Yan, Ziang Yan, Bowen Yang, Danni Yang, Lin Yang, Zhiqi Yang, Qian Yao, Haochen Ye, Peng Ye, Jinhui Yin, Jiashuo Yu, Dingbo Yuan, Fei Yuan, Yuhang Zang, Bo Zhang (possible past Tencent (China) affiliation), Chao Zhang, Chen Zhang (possible past Peking University affiliation), Hongjie Zhang, Junming Zhang, Wenlong Zhang, Wenwei Zhang, Yiming Zhang, Zhuo Zhang, Ziyang Zhang, Haiteng Zhao, Penghao Zhao, Yibo Zhao, Zhonghan Zhao, Zhihang Zhong, Bowen Zhou, Peiheng Zhou, Xin Zhou (possible past Stanford University affiliation), Xinyu Zhou, Yunhua Zhou, Dongsheng Zhu, Yicheng Zou
Abstract

Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and environments, and sustain progress across long task horizons. We present Intern-S2-Preview, a series of scientific agentic foundation models designed to support multimodal scientific understanding, reasoning, generation, and long-horizon tasks. The training pipeline begins with scientific multimodal pre-training over rendered scientific do...

πŸ“„ Doubly Robust Estimation of Causal Effect on CVR with Targeted Regularization
πŸ—“οΈ Published: 8/13/2026
πŸ”— http://arxiv.org/abs/2608.13461v1
πŸ‘₯ Authors: Jiayi Dan, Bo Li (possible past Tencent (China) affiliation), Lu Deng, Yong Wang (possible past Baidu (China) affiliation)
Abstract

Post-click conversion rate (CVR) is a key metric in various scenarios including e-commerce and advertising, reflecting the efficiency and user experience in the second stage of the conversion process. Estimating the causal effect on CVR is therefore of great practical importance. However, directly applying existing causal inference methods to clicked samples introduces sample selection bias and increased variance due to the exclusion of non-click data. Recent studies on CVR prediction introduce ...

πŸ“„ Diagnosing JEPA World Models with Action-Conditioned Predictive Consistency
πŸ—“οΈ Published: 8/13/2026
πŸ”— http://arxiv.org/abs/2608.12939v1
πŸ‘₯ Authors: Guo An, Zijing Wu, Honghua Dong, Yuhao Yan, Zixuan Gui, Haochong Chen, Shanzhao Ruan, Xiang Wang (possible past Tencent (China) affiliation), Yurong Ling, Qi Tian (possible past Huawei Technologies (China) affiliation)
Abstract

Joint-embedding predictive architectures (JEPAs) learn world models that predict in a compact latent space rather than in pixels, reducing the pressure to model nuisance appearance. Yet this provides no guarantee against visual perturbations: they can still alter the encoded representation and affect subsequent action-conditioned predictions. Bisimulation captures this requirement precisely: two observations should be treated as the same state only when their action-conditioned consequences agre...

πŸ“„ From Visual Widgets to UI Code: Efficient Tool-Grounded Generation
πŸ—“οΈ Published: 8/12/2026
πŸ”— http://arxiv.org/abs/2608.12611v1
πŸ‘₯ Authors: Houston H. Zhang, Tao Zhang (possible past Nvidia (United States) affiliation), Li Gu, Linfeng Ye, Yuanhao Yu, Xinxin Zuo, Yang Wang (possible past Baidu (China) affiliation), Zhixiang Chi
Abstract

Existing screenshot-to-code systems face a trade-off between flexibility and controllability. Direct multimodal generation can hallucinate visible details, whereas structured pipelines reduce such errors through component-wise decomposition, predefined templates, and customized intermediate representations. These structures, however, introduce additional generative orchestration and restrict outputs to designs covered by the representation. We investigate whether selective tool grounding can imp...

πŸ“„ MARCH: Scaling Recurrent Memory with Content-Routed State Anchors
πŸ—“οΈ Published: 8/12/2026
πŸ”— http://arxiv.org/abs/2608.12435v1
πŸ‘₯ Authors: Ming Zhang (possible past Peking University affiliation), Kaisen Yang, Shu Yu, Ermo Hua, Ning Ding (possible past Tsinghua University affiliation), Xia Hu, Bowen Zhou, Chaochao Lu, Youbang Sun
Abstract

Transformers owe much of their strong long-context retrieval capability to a token-level memory that grows with context length. This flexibility, however, incurs a quadratic computation complexity during training and a key--value cache that grows linearly during autoregressive inference. Recurrent alternatives offer efficient decoding by compressing the entire history into a fixed-size state, but often underperform on recall-intensive tasks since earlier associations usually get overwritten by s...

πŸ“„ Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
πŸ—“οΈ Published: 8/12/2026
πŸ”— http://arxiv.org/abs/2608.12036v1
πŸ‘₯ Authors: Mengru Wang, Junfeng Fang, Shuofei Qiao, Zhenqian Xu, Haoming Xu, Haoxiong Wang, Shumin Deng (possible past Alibaba Group (China) affiliation), Linyi Yang, Zhixiang Cui, Xin Xu, Yunzhi Yao, Buqiang Xu, Fei Shen, Haozhe Luo, Yunxiang Wei, Ningyu Zhang (possible past Tencent (China) affiliation), Julian Mcauley, Tat Seng Chua, Huajun Chen (possible past Alibaba Group (China) affiliation)
Abstract

AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood. As AI development becomes faster and increasingly automated, mechanistic exploration remains largely manual, widening the gap between what models can do and our ability to understand and control them. To bridge this gap, we introduce Mechanist, an agentic system that uses AI as a scientific instrument for the autonomous discover...

*Notable papers are those with at least two authors from a "big" AI/ML lab.