📄 Notable* Recent AI/ML arXiv Papers

Last updated just now...

📄 Verifiable Social Reasoning for LLM Assistants
🗓️ Published: 9/15/2026
🔗 http://arxiv.org/abs/2609.17496v1
👥 Authors: Amir Taubenfeld, Zorik Gekhman, Avigail Grinstein-Dabush, Itay Laish, Ariel Goldstein, Marian Croak, Avinatan Hassidim (possible past Google (United States) affiliation), Yossi Matias (possible past Google (United States) affiliation), Amir Feder (possible past Technion – Israel Institute Of Technology affiliation)
Abstract

LLM assistants are widely used for daily social advice, yet evaluating their social reasoning in such consultation settings remains challenging since (i) it requires setups where the assistant learns about social situations from subjective user narratives, and (ii) social properties, such as others' intentions, typically lack verifiable ground truth. To address these challenges, we introduce Fuse, a multi-agent simulation framework for studying user-mediated social reasoning. In Fuse, a target a...

📄 LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence
🗓️ Published: 9/15/2026
🔗 http://arxiv.org/abs/2609.17488v1
👥 Authors: Xingxuan Zhang, Gang Ren (possible past Google (United States) affiliation), Hao Yuan, Hao Zou, Hongze Tan, Hui Wang, Jianhao Song, Jiansheng Li, Jiayao Zhang, Jinghan Zhang, Kaifang Li, Lang Mo, Li Mao, Mingchao Hao, Nuo Xu, Rui Ding, Ruiji Zhang, Shuyang Li, Siyu Mei, Tianyang Zhang, Weiyang Mu, Yancheng Dong, Yongxian Wei, Yuan Xue (possible past Google (United States) affiliation), Yuanrui Wang, Yue He, Zijia Yang, Ziyun Li, Dongzhe Li, Fuqiang Wang, Jiandong Liu, Jiawei Chen (possible past Tencent (China) affiliation), Jiaxin Du, Kaijie Cheng, Kehan Li, Lei Sun, Linjun Zhou, Ningbo Dai, Qi Wang (possible past Tsinghua University affiliation), Renzhe Xu, Shaoxing Du, Shumeng Yang, Wang Lu, Wenjing Chu, Xiannan Huang, Xiaoyu Lin, Xing Ai, Xinyan Han, Xuanyue Li, Xuanyue Su, Xukun Zhang, Yan Lu, Yaxin Zhang, Yi Qin, Yifei Huang, Yihan Xu, Yongle Lv, Yuanyuan Jiang, Yushan Han, Peng Cui (possible past Tsinghua University affiliation)
Abstract

We introduce LimiX-2, a new model in the LimiX family, developed through model and data scaling guided by our previously established scaling laws. LimiX-2 adopts the Contextual Mechanism Networks (CMNs) paradigm and is pretrained with Context-Conditional Masked Modeling (CCMM). CMNs shifts the organizing principle of in-context learning from target-centric prediction to mechanism-oriented joint modeling. Rather than centering the network on the $p(y \mid x, D_{\mathrm{context}})$ objective of co...

📄 From Transient Prompts to Persistent Control: Scientific Poster Generation via Recursive Semantic-Geometric Contracts
🗓️ Published: 9/15/2026
🔗 http://arxiv.org/abs/2609.17326v1
👥 Authors: Runze Li, Yukun Zhao, Can Xu (possible past Google (United States) affiliation), Yucheng Shen, Shuaiqiang Wang (possible past Baidu (China) affiliation), Jianmin Wu, Lingyong Yan, Dawei Yin (possible past Baidu (China) affiliation)
Abstract

Scientific poster generation distills a multimodal paper into a single-page visual artifact, forcing strict trade-offs between informational coverage and readability under a fixed spatial budget. Existing methods pass plans as transient prompts and validate individual stages in isolation. This strategy causes requirements to drift across content and layout modules, and previous checks to be silently invalidated. We introduce PosterVisor, a control framework that shifts poster generation from tra...

📄 FluxVLA Engine: A One-Stop VLA Engineering Platform for Embodied Intelligence
🗓️ Published: 9/15/2026
🔗 http://arxiv.org/abs/2609.17210v1
👥 Authors: Yinhao Li, Weixin Mao, Zihan Lan, Jikun Rong, Qirui Hu, Yiming Zhang, Weipeng Deng, Bowen Shen, Minzhao Zhu, Yiming Mao, Yan Yang (possible past Google (United States) affiliation), Chenguang Cui, Hongyuan Chen, Xu Huang, Zheyi Zhao, Pinxi Shen, Bozhen He, Zhen Fu, Yifan Wang (possible past Stanford University affiliation), Zexin Zhang, Ang Gao, Haoyu Chen, Chengqi Shi, Hua Chen
Abstract

Vision-language-action (VLA) models, world-action models (WAMs), and offline reinforcement learning methods are rapidly expanding the design space of embodied policies, yet turning these algorithms into reliable robot systems remains constrained by fragmented data formats, training stacks, evaluation protocols, inference runtimes, and embodiment-specific interfaces. We present $\mathrm{FluxVLA}$ Engine, an open, configuration-driven platform that turns heterogeneous embodied-policy components in...

📄 Sparse MLLM Anchors, Dense Adaptation: Breaking the Self-Referential Loop in Wild Test-Time Adaptation
🗓️ Published: 9/15/2026
🔗 http://arxiv.org/abs/2609.17040v1
👥 Authors: Zhenbin Wang, Lei Zhang, Lituan Wang, Yan Wang (possible past Tencent (China) affiliation), Zhao Zhang, Wei Huang (possible past Google (United States) affiliation)
Abstract

Wild test-time adaptation (WTTA) updates a source model online under small test batches, concurrent distribution shifts, and time-varying class imbalance. Most WTTA methods derive their adaptation signals, including predictive uncertainty, sample reliability, and local feature geometry, from the model being adapted. When the source model is unreliable under shift, these signals can reinforce its own errors, forming a self-referential loop. We introduce MASA (Multimodal-LLM-Anchored Semantic Adap...

📄 Beyond Token-Local Imitation: Reward-Compatible Temporal Credit Assignment for On-Policy Distillation
🗓️ Published: 9/15/2026
🔗 http://arxiv.org/abs/2609.16937v1
👥 Authors: Shiqi Liu, Zeyu He, Letian Tao, Guojian Zhan, Jiaxin Gao, Feihong Zhang, Jingliang Duan (possible past Tsinghua University affiliation), Wei Xiong, Kehua Sheng, Bo Zhang (possible past Tencent (China) affiliation), Yang Guan, Shengbo Eben Li (possible past Tsinghua University affiliation)
Abstract

On-policy distillation (OPD) has emerged as an effective approach for large language model post-training, yet existing objectives face a trade-off between objective fidelity and optimization stability. Token-level OPD provides stable but local supervision, whereas sequence-level OPD captures future credit at the cost of horizon-dependent variance. We establish a unified temporal-credit view of these formulations, showing that practical token-level OPD can be interpreted as a temporal approximati...

📄 StackTok: Accelerating VLMs Inference with Budget-Adaptive Visual Token Selection
🗓️ Published: 9/15/2026
🔗 http://arxiv.org/abs/2609.16841v1
👥 Authors: Zhenbin Wang, Lei Zhang, Lituan Wang, Wei Huang (possible past Google (United States) affiliation), Yan Wang (possible past Tencent (China) affiliation), Zhenwei Zhang
Abstract

Increasing image resolution produces ever-longer visual-token sequences in vision-language models (VLMs), substantially raising their inference cost. To reduce this overhead without retraining, existing methods select compact token subsets that prioritize query relevance, visual coverage, or a fixed trade-off between them. The appropriate balance, however, varies across queries and token budgets: localized questions favor relevance, whereas holistic questions demand broader visual coverage. We i...

📄 Layers, Sinks, and Scaling: Adaptive Evidence Selection for Multimodal Large Language Models
🗓️ Published: 9/15/2026
🔗 http://arxiv.org/abs/2609.16795v1
👥 Authors: Zhenbin Wang, Lei Zhang, Lituan Wang, Wei Huang (possible past Google (United States) affiliation), Yan Wang (possible past Tencent (China) affiliation), Zhenwei Zhang
Abstract

Multimodal large language models (MLLMs) can answer knowledge-intensive visual questions by combining visual evidence from images with facts retrieved from external sources. However, MLLMs may overlook relevant evidence in both modalities, attending weakly to the textual sentences or visual regions needed for the correct answer. Recent efforts address this by highlighting retrieved text and marking visual regions before generation, but apply a fixed, one-shot policy that cannot adapt to three so...

📄 Turn-level Multiscale Density Ratio Estimation for LLM Agents
🗓️ Published: 9/15/2026
🔗 http://arxiv.org/abs/2609.16760v1
👥 Authors: Zishuo Zhao, Kai Chen (possible past Shanghai Jiao Tong University affiliation), Ao Li, Yuan Liu (possible past Google (United States) affiliation)
Abstract

With the rapid development of Large language model (LLM), agent systems enhanced by LLMs show huge potential in being able to deal with complex tasks, especially involving multi-step thinking or interaction with tools. For applying LLM techniques with a well-designed agent paradigm, post-training of LLM in multiple agent scenarios is necessary to achieve better performance. Among the variable post-training techniques, alignment methods such as PPO, DPO, DIL, and GRPO become popular because many ...

📄 World Models for Embodied Intelligence: From Plausible to Controllable to Actionable
🗓️ Published: 9/15/2026
🔗 http://arxiv.org/abs/2609.16697v1
👥 Authors: Nanjie Yao, Hao Wang (possible past Tsinghua University affiliation), Chong Cheng, Zhikang Chen, Wenzhe Li, Jiafei Lyu, Li Shen (possible past Tencent (China) affiliation), Peilin Zhao (possible past Tencent (China) affiliation), Zongqing Lu, Gao Huang (possible past Tsinghua University affiliation), Steven Hoi, Dacheng Tao, Deheng Ye (possible past Tencent (China) affiliation)
Abstract

World models connect perception and decision-making in embodied intelligence by maintaining hidden state, anticipating consequences, comparing interventions, and adapting when execution departs from expectations. Although progress is often measured by visual fidelity, their value lies in improving behavior. Before reaching for a cup, a person anticipates its weight and resistance to grasping, shaping the hand before contact. Such anticipation is coarse and rarely pictorial, yet it guides action....

📄 AI for Games in the Foundation Model Era
🗓️ Published: 9/15/2026
🔗 http://arxiv.org/abs/2609.16679v1
👥 Authors: Meng Luo, Yanlin Li, Hao Li (possible past Tsinghua University affiliation), Hongzhan Lin, Pengfei Zhou, Tianjie Ju, Ran Zhang, Yeying Jin, Mong-Li Lee, Wynne Hsu (possible past National University Of Singapore affiliation)
Abstract

Foundation models, alongside advances in learned game-world models, are reshaping AI across the game lifecycle. Beyond playing games, recent systems model players and game dynamics, support design and development, adapt player-facing experiences at runtime, and evaluate resulting artifacts. Yet these directions have evolved largely separately, obscuring which capabilities transfer across settings and which remain tied to particular games, engines, interfaces, or player populations. We organize t...

📄 A Vision-Language Foundation Model for Precise and Comprehensive Brain Tumor Diagnosis from Preoperative Multimodal Data
🗓️ Published: 9/15/2026
🔗 http://arxiv.org/abs/2609.16597v1
👥 Authors: Yinong Wang, Jianwen Chen, Zhou Chen, Shuwen Kuang, Haoning Jiang, Yanzhao Shi, Huichun Yuan, Yan-Ran, Wang, Bing Wang, Lei Wu, Bin Tang, Li Meng, Baihua Luo, Bin Zhou, Wei Ding, Weiming Zhong, Wei Hou, Yuanbing Chen, Zhiping Wan, Wei Wang (possible past University Of Oxford affiliation), Zhenkun Xiao, Wenwu Wan, Allen He, Yuyin Zhou, Longbo Zhang, Feifei Wang, Zhixiong Liu, Michael Iv, Xuan Gong, Liangqiong Qu (possible past Stanford University affiliation)
Abstract

Background Non-invasive presurgical diagnosis of brain tumor types from Magnetic Resonance Imaging (MRI) is essential but challenging due to overlapping imaging features across tumor types, inter-observer variability, and the extensive training required for expertise. We aimed to develop an MRI-based Artificial Intelligence (AI) model for automatic and reliable brain tumor classification with diagnostic uncertainty quantification and radiology reports generation. Methods We developed BrainVLM ...

📄 What Does Layer-Importance Reveal About Transformers and State-Space Models?
🗓️ Published: 9/15/2026
🔗 http://arxiv.org/abs/2609.16537v1
👥 Authors: Istabrak Abbes, Nizar Islah, Irina Rish (possible past Deepmind (United Kingdom) affiliation), Sarath Chandar (possible past Mila - Quebec Artificial Intelligence Institute affiliation)
Abstract

Transformers and state-space models (SSMs) are the two dominant families of sequence models, and a central open question is how far the analytical knowledge built for transformers transfers to SSMs. We address this through the lens of layer importance which underpins compression, selective fine-tuning, and interpretability across both families. We decompose layer importance into two distinct notions. \emph{Necessity} captures how much the pretrained model depends on a layer's existing contributi...

📄 Atria Dawn: The Dawn of Agentic Superintelligence
🗓️ Published: 9/14/2026
🔗 http://arxiv.org/abs/2609.15818v1
👥 Authors: Honglin Guo, Tao Gui, Yicheng Chen, Guanting Dong, Qiming Ge, Yuyang Hu, Zixian Huang, Jiajie Jin, Alexander Lam, Yining Li, Jiahang Lin, Yanjiang Liu, Xinyu Lu, Haijun Lv (possible past Baidu (China) affiliation), Junlin Shang, Qisheng Su, Guoqiang Wang, Rui Wang (possible past Tencent (China) affiliation), Zhecan Wang, Hao Xiang, Xinchen Xie, Shuhao Xing, Xiaoyu Xing, Wanghan Xu, Xinyu Yang, Yajie Yang, Chengfeng Zhao, Haoran Zhao, Ruojun Zhou, Yunhua Zhou, Yicheng Zou, Kun Cai, Qiye Cai, Xinmeng Che, Haodong Chen, Jiabei Chen, Jiahao Chen, Jiayi Chen, Yujia Chen, Lizhi Cui, Youheng Dai, Xin Deng, Yi Dong, Shihan Dou, Chenya Gu, Xu Guo, Ding Han, Feiyang Hao, Haotan He, Jie Hou, Binze Hu, Zijian Hu, Junhao Huang, Huicheng Jiang, Jiazhen Jiang, Shufan Jiang, Jiahao Kuang, Bowen Lai, Bo Li (possible past Tencent (China) affiliation), Jiaqiang Li, Peng Li (possible past Tsinghua University affiliation), Qilong Li, Zhuoqun Li, Jiaxiang Liu, Shuainan Liu, Tong Liu, Yi Liu (possible past Google (United States) affiliation), Zhonghang Lu, Jianwen Luo, Yanyi Luo, Huijie Lv, Ningsheng Ma, Zerun Ma, Houcheng Min, Chengjun Pan, Qiyuan Peng, Xiaoxuan Peng, Jianmin Qian, Jiantao Qiu, Wanying Ren, Huayu Sha, Jifei Shan, Zixin Shang, Bing Shao, Zhuohui Sheng, Jiayang Shi, Yang Shu, Aierpanjiang Simayi, Sirui Song, Yuxiao Song, Zhe Sun (possible past Tsinghua University affiliation), Zhichao Sun, Wenzhe Tan, Wenhui Tian, Zhongbo Tian, Hanchen Wang (possible past University Of Cambridge affiliation), Pengbo Wang, Rui Wang (possible past Tencent (China) affiliation), Yiding Wang, Yuhui Wang, Zhiheng Xi, Caijun Xu, Chao Xu, Yongfeng Xu, Xiaolei Yang, Zhixiong Yang, Qian Yao, Shihong Yi, Yuankai Ying, Jia Yu, Dingbo Yuan, Hao Yuan, Junjie Yuan, Bo Zhang (possible past Tencent (China) affiliation), Caixian Zhang, Qiuyinzhe Zhang, Jiyuan Zhao, Penghao Zhao, Ying Zhao (possible past Stanford University affiliation), Pujun Zheng, Xiaoxue Zhong, Xiaohao Zhou, Xinyu Zhou, Dongsheng Zhu, Guanru Zhu, Yulun Zhu, Yaojie Lu, Tao Ji, Hongyu Lin, Yutao Zhu, Pengfei Cao, Guoxiu He, Xianpei Han (possible past Tencent (China) affiliation), Ben He, Zhicheng Dou, Kang Liu, Qi Zhang (possible past Tencent (China) affiliation), Le Sun, Jun Zhao, Ji-Rong Wen, Xuanjing Huang, Yu-Gang Jiang, Bowen Zhou
Abstract

As AI agents become participants in the development of their successors, they reshape both the production of intelligence and the role of human researchers. We introduce Atria Dawn Preview, a foundation agentic language model designed for scientific research and engineering workflows, with the goal of expanding the frontier of agent productivity in the real world. This model is trained via a Verifiable Experience Pipeline that connects tool-mediated interactions to executable environments and ex...

📄 Navigating Sparse Evidence: Agentic Visual RAG via Explicit Context Selection and Consolidation
🗓️ Published: 9/14/2026
🔗 http://arxiv.org/abs/2609.15800v1
👥 Authors: Yucheng Shen, Lingyong Yan, Jiulong Wu, Shuaiqiang Wang (possible past Baidu (China) affiliation), Jianmin Wu, Dawei Yin (possible past Baidu (China) affiliation), Min Cao
Abstract

Visual Retrieval-Augmented Generation (VRAG) empowers models to navigate and answer queries about visually rich documents by retrieving relevant page images as visual evidence and reasoning over their content. However, effectively utilizing this visual evidence is usually impeded by two main challenges. First, answer-relevant evidence is sparse and may be concentrated in a small region of one page or dispersed across multiple pages. Second, existing agentic methods often generate answers based o...

📄 Bench2Dex: Benchmarking Visuo-Tactile Bimanual Dexterous Manipulation Across Dexterous Hands
🗓️ Published: 9/14/2026
🔗 http://arxiv.org/abs/2609.15726v1
👥 Authors: Zhenjie Yang, Yideng Zhang, Dongjie Zhang, Chenyu Jiang, Xianshuai Liu, Yufeng Li, Zuhao Ge, Xingyu Jiao, Zheng Zhang, Kaiyu He, He Wang (possible past Stanford University affiliation), Yuwen Zhong, Yi Deng, Muyun Jiang, Xianliang Huang, Haisheng Su, Donghang Zhang, Jian Zhang (possible past Tencent (China) affiliation), Xue Yang, Hongyang Li, Zuxuan Wu, Yu-Gang Jiang, Xiaosong Jia, Junchi Yan (possible past Shanghai Jiao Tong University affiliation)
Abstract

Tactile sensing provides contact information that can be difficult to infer from vision alone, but tactile hardware for dexterous hands has not converged to a common design. Dexterous hands differ in finger structure, contact surfaces, and sensor layouts, while simulated tactile signals still differ from measurements produced by physical sensors. These factors make it difficult to study visuo-tactile manipulation across diverse dexterous hands within a consistent experimental setting. We present...

📄 OptiPrime: Optimizing Private Inference through Protocol-Hardware Co-design
🗓️ Published: 9/15/2026
🔗 http://arxiv.org/abs/2609.16898v1
👥 Authors: Jiangrui Yu, Ye Yu, Si Chen, Chenqi Lin, Wenxuan Zeng, Junfeng Fan, Mingyu Gao (possible past Tsinghua University affiliation), Meng Li (possible past Meta (United States) affiliation)
Abstract

Private deep neural network (DNN) inference based on hybrid homomorphic encryption (HE) and multi-party computation (MPC) can protect user data with a formal guarantee, but at the cost of significant latency overhead due to HE. Customized HE accelerators have been proposed and have achieved orders-of-magnitude speedup for individual HE operations. However, when directly applying a commercial HE accelerator to state-of-the-art HE-MPC frameworks, we observe only limited end-to-end performance gain...

📄 The Neverwhere Visual Parkour Benchmark Suite
🗓️ Published: 9/14/2026
🔗 http://arxiv.org/abs/2609.16443v1
👥 Authors: Ziyu Chen, Henghui Bao, Haoran Chang, Alan Yu, Ran Choi, Kai Mcclennen, Gio Huh, Kevin Yang, Ri-Zhao Qiu, Yajvan Ravan, John J. Leonard, Xiaolong Wang (possible past Carnegie Mellon University affiliation), Phillip Isola (possible past University Of California, Berkeley affiliation), Ge Yang, Yue Wang
Abstract

State-of-the-art visual locomotion controllers are increasingly capable at handling complex visual environments, making evaluating their real-world performance before deployment increasingly difficult. This work intends to narrow this train/evaluation gap by developing a collection of hyper-photo-realistic, closed-loop evaluation environments - The Neverwhere Benchmark Suite - comprised of over sixty 3D Gaussian Splatting reconstructions of urban indoor and outdoor scenes. Our goal is to encoura...

📄 LLM Inference in a Flash!
🗓️ Published: 9/14/2026
🔗 http://arxiv.org/abs/2609.16161v1
👥 Authors: Sebastian Zhao, Minseo Kim, Coleman Hooper, Luca Manolache, Michael W. Mahoney (possible past Stanford University affiliation), Yakun Sophia Shao (possible past University Of California, Berkeley affiliation), Kurt Keutzer (possible past University Of California, Berkeley affiliation), Amir Gholami
Abstract

Large Language Models (LLMs) have shown impressive capabilities across a range of natural language processing tasks, and LLM inference has emerged as a critical workload for enabling downstream applications. The demands of serving LLM inference are becoming increasingly challenging as requests shift toward longer sequences and heavier inference, driven by retrieval-augmented generation, inference-time compute scaling, and long-context applications. Additionally, these challenges are compounded b...

📄 MoveBench: A Benchmark for Global-Scale Wildlife Movement Forecasting
🗓️ Published: 9/14/2026
🔗 http://arxiv.org/abs/2609.15780v1
👥 Authors: Justin Kay, Shir Bar, Ellen O. Aikens, Martin Becker, Francesca Cagnacci, Juliet Cohen, Scott W. Forrest, Jessica Kendall-Bar, Madeleine Lucas, Macon Overcast, Meredith S. Palmer, Will Rogers, Nicholas J. Russo, Christian Rutz, Larissa T. Beumer, Michael Brown (possible past Google (United States) affiliation), Ying-Chi Chan, Sarah C. Davidson, Diego Ellis Soto, Anne G. Hertel, Roland Kays, Benjamin Koger, Guram Mikaberidze, Thomas Mueller, Ruth Oliver, Thorsten Papenbrock, Robert Patchett, Jared A. Stabach, Dane Taylor, Scott W. Yanco, Sara Beery (possible past Microsoft (United States) affiliation)
Abstract

Understanding and predicting wildlife movement is critical for ecology and conservation. While trajectory forecasting has advanced for human and vehicle movement, wildlife trajectories present distinct challenges: they are unconstrained in space, highly stochastic, and influenced by environmental conditions. We introduce MoveBench, the first large-scale benchmark for probabilistic wildlife movement forecasting, containing 2.6M GPS locations from 800+ individuals across 110 species in 127 countri...

📄 Circuit-MLLM: Topological Logic-Guided Latent-Space Visual Reasoning for Circuit Schematic Understanding
🗓️ Published: 9/14/2026
🔗 http://arxiv.org/abs/2609.15668v1
👥 Authors: Jinyuan Deng, Yuqi Jiang, Wenjing Huang, Xin Li (possible past Google (United States) affiliation), Qi Sun (possible past Google (United States) affiliation), Cheng Zhuo
Abstract

Through pre-training on extensive text and image datasets, current multi-modal large language models (MLLMs) achieve strong performance on general tasks. However, circuit schematics present a unique challenge for MLLMs due to their dense component layouts and distinct topological logic, demanding fine-grained structural parsing to extract the electrical semantics. To address this, we propose Circuit-MLLM, a multimodal reasoning framework that reformulates circuit topology analysis as a process o...

📄 The Token Before the Value Is the Key: How Hybrid Architectures Organize Induction Circuits
🗓️ Published: 9/14/2026
🔗 http://arxiv.org/abs/2609.15545v2
👥 Authors: Ke Cheng, Xin Xu, Yixiao Chen, Lei Xin, Jianbo Zhao, Fanhu Zeng, Yue Liu, Jun Zhang (possible past Tencent (China) affiliation), Jie Jiang (possible past Tencent (China) affiliation)
Abstract

Hybrid language models can improve capability as well as efficiency, raising the question of how architectural complementarity becomes learned computation. We examine the established induction roles of Carrying predecessor information, Matching a source by content, and Copying its value. How are these position-sensitive and content-based computations allocated across heterogeneous layers? We introduce layer-type-agnostic paired probes that track Carrying and Matching through a common block-updat...

*Notable papers are those with at least two authors from a "big" AI/ML lab.