📄 Notable* Recent AI/ML arXiv Papers

Last updated just now...

📄 STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization
🗓️ Published: 9/29/2026
🔗 http://arxiv.org/abs/2609.38169v1
👥 Authors: Bingchen Yao, Haobo Xu, Haokun Lin, Yichen Wu, Ziyu Guo, Renrui Zhang, Zhichao Lu (possible past Google (United States) affiliation), Zhenan Sun, Ying Wei (possible past Tencent (China) affiliation)
Abstract

Linear attention replaces growing KV caches with fixed-size recurrent states, yet these persistent states can become a substantial memory bottleneck under concurrent serving. Directly quantizing recurrent states to low precision often leads to severe accuracy degradation, as quantization errors propagate through successive state updates. We discover that the impact of these errors depends on two complementary dimensions: temporally, errors in long-lived memory can persist across many decoding st...

📄 LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization
🗓️ Published: 9/29/2026
🔗 http://arxiv.org/abs/2609.38166v1
👥 Authors: Yi Pan, Haocheng Xi, Kan Zhu, Xingyang Li, Yibo Wu, Mayank Mishra, Hongtao Zhang, William X. Zheng, Baris Kasikci, Song Han (possible past Stanford University affiliation), Kurt Keutzer (possible past University Of California, Berkeley affiliation), Rishabh Iyer, Ion Stoica (possible past University Of California, Berkeley affiliation)
Abstract

Recent LLMs increasingly adopt hybrid designs that replace standard attention with linear attention, such as Gated DeltaNet (GDN) and Kimi Delta Attention (KDA). Although they compress the context into a fixed-size recurrent state and substantially reduce the cost of long-context processing, repeatedly reading and updating that state remains a major inference bottleneck. Quantization offers a natural way to reduce this cost, but can significantly degrade model quality, due to the accumulation of...

📄 Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning
🗓️ Published: 9/29/2026
🔗 http://arxiv.org/abs/2609.38147v1
👥 Authors: Paras Dahal, Anton Bakhtin (possible past Google (United States) affiliation), Taco Cohen, Zhengxing Chen, Carole-Jean Wu (possible past Meta (United States) affiliation), Rob Fergus (possible past Meta (United States) affiliation), Scott Yih, Gabriel Synnaeve (possible past Meta (United States) affiliation), Ruslan Salakhutdinov (possible past University Of Toronto affiliation), Sanjeev Arora, Jason Weston (possible past Stanford University affiliation), Anirudh Goyal
Abstract

As agents take on longer and more complex problems, controlling the execution becomes a task in its own right. Each step in the run brings new control choices, like which partial work to build on, whether to start fresh, or when to stop. We introduce agentic meta-reasoning, an inference-time harness that makes these choices an explicit and structured reasoning process. Workers carry out the task-level computation, while a controller consolidates what the run has established, explores next option...

📄 Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE
🗓️ Published: 9/29/2026
🔗 http://arxiv.org/abs/2609.38140v1
👥 Authors: Yu Xu (possible past Tencent (China) affiliation), Yuxin Zhang, Xiao Yang (possible past Tencent (China) affiliation), Haotian Yang, Yizhi Wang, Xinwei Huang, Minxuan Lin, Angtian Wang, Chongyang Ma, Fan Tang
Abstract

Mixture-of-Experts (MoE), popularized by large language models, is a promising paradigm for scaling visual generative models. However, conventional token-wise MoE routes tokens independently within a homogeneous expert pool and regularizes expert usage toward uniformity, making it poorly matched to video data that is spatiotemporally redundant and semantically long-tailed. We show that existing visual MoEs fall into a uniformity trap: semantically under-organized routing, compounded by uniform e...

📄 Does Local Video Understanding Transfer Across Encounters? The EgoGears Benchmark
🗓️ Published: 9/29/2026
🔗 http://arxiv.org/abs/2609.37938v1
👥 Authors: Yuedong Tan, Lei Qi, Yu Liu, Di Wen, Ruiping Liu, Xiaoye Wang, Yufan Chen, Junwei Zheng, Chengzhi Wu, Chen Zhang (possible past Peking University affiliation), Zhihang Chen, Haiwen Sun, Zongwei Wu, Radu Timofte (possible past Eth Zurich affiliation), Danda Pani Paudel, Kunyu Peng
Abstract

Embodied systems must make knowledge acquired during one encounter usable in another despite changes in viewpoint, motion, and illumination. Yet aggregate cross-video accuracy conflates failures of local perception with failures to preserve observation identity, establish correspondence, and compose evidence, obscuring whether local video understanding actually transfers. We introduce EgoGears, a complementary single- and multi-video benchmark designed to diagnose this transition. It contains 56...

📄 Rollout-Marginal Distillation for Long-Horizon Autoregressive Video Generation
🗓️ Published: 9/29/2026
🔗 http://arxiv.org/abs/2609.37925v1
👥 Authors: Chenjian Gao, Zhihao Hu, Jianqi Ma, Jun Zhang (possible past Tencent (China) affiliation), Weidong Zhang, Tianfan Xue (possible past Massachusetts Institute Of Technology affiliation)
Abstract

Autoregressive (AR) video diffusion enables low-latency, streamable video generation, but prediction errors often accumulate over long rollouts. Training the generator on its own rollouts exposes it to these imperfect histories. However, existing video-level distribution matching distillation (DMD) scores the whole rollout jointly. Because a chunk is evaluated together with its past and future, its correction can favor matching artifacts in the surrounding context merely to preserve temporal con...

📄 Pixels to Keys: Exploring Spatial and Motion Cues in Gameplay Inverse Dynamics
🗓️ Published: 9/29/2026
🔗 http://arxiv.org/abs/2609.37907v1
👥 Authors: Abhishek Pillai, Ekta Prashnani, Joohwan Kim (possible past Nvidia (United States) affiliation), Iuri Frosio (possible past Nvidia (United States) affiliation)
Abstract

Video games offer scalable environments for studying perception and control in embodied agents.Abundant online gameplay videos could supply demonstrations, but they rarely include player inputs for training. Inverse Dynamics Models (IDMs) have thus been proposed to infer inputs from frames. Large (up to 1B parameters) IDMs trained on $\sim$1K-2K gameplay hours demonstrate feasibility and cross-environment generalization at this scale, but researchers do not clarify what the key components are to...

📄 ExceptionDrive: A Planning-Oriented Counterfactual Corner-Case Benchmark for Autonomous Driving
🗓️ Published: 9/29/2026
🔗 http://arxiv.org/abs/2609.37871v1
👥 Authors: Ziyi Luo, Zhe Sun (possible past Tsinghua University affiliation), Yehao Lu, Lei Zhou (possible past Apple (United States) affiliation), Lisheng Wu, Xuewei Li, Zequn Qin, Xi Li
Abstract

Average performance on routine driving benchmarks does not establish planner reliability under rare, safety-critical hazards. We proposed ExceptionDrive, a counterfactual planning benchmark that uses VLM-assisted screening, localized multi-view editing, and quality auditing to insert hazards into real nuScenes scenes while preserving their context. Its 21 tasks span six safety families and define hazard or conflict regions, local safety constraints, and acceptable responses. Because hazard inser...

📄 Context Language Models
🗓️ Published: 9/29/2026
🔗 http://arxiv.org/abs/2609.37725v1
👥 Authors: Rulin Shao, Shannon Zejiang Shen, Junjie Oscar Yin, Yuetai Li, Minheng Wang, Hamish Ivison, Radha Poovendran, Nathan Lambert (possible past University Of California, Berkeley affiliation), Teng Xiao, Mike Lewis (possible past Meta (United States) affiliation), Wen-Tau Yih (possible past Microsoft (United States) affiliation), Luke Zettlemoyer (possible past University Of Washington affiliation), Pang Wei Koh
Abstract

We introduce Context Language Models (CLMs), language models that natively manage their own context. We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file. This allows the model to learn what is most important to maintain in context, and naturally extends to multi-agent systems where multiple agent contexts coexist as files. Building CLMs zero-shot with existing models outperforms SOTA context management strategies across a variety o...

📄 PolyOCR-Venus: Unified OCR Foundation Models for Text-Centric Visual Intelligence
🗓️ Published: 9/29/2026
🔗 http://arxiv.org/abs/2609.37712v1
👥 Authors: Guangjian Team, Kaili Huang, Yongshuo Zhang, Bingtao Fu, Changjiang Jiang, Chenfan Qu, Chenfeng Zhang, Fangming Cui, Gaoyang Zhang, Jiangwei Xie, Jianshu Li (possible past National University Of Singapore affiliation), Jing Huang (possible past Meta (United States) affiliation), Jingwen Bai, Mingqi Fang, Tao Fang, Weihong Zhang, Wenbo Du, Xiongfei Bai, Xuekang Zhu, Yinan Xia, Zhenming Wang, Jian Liu, Jingjing Liu (possible past Microsoft (United States) affiliation), Xiang Qi, Weiqiang Wang
Abstract

Optical Character Recognition (OCR) is evolving from plain-text transcription toward general visual intelligence, requiring models to recognize, localize, and reason over textual information in complex visual environments. However, existing OCR systems often excel at only some tasks and struggle to balance recognition, parsing, and reasoning across scenarios. In this report, we present PolyOCR, a family of unified OCR foundation models of varying scales. PolyOCR combines a shared instruction-fol...

📄 GARDiff: Graph-Aligned Residual Diffusion for Probabilistic Multivariate Time-Series Forecasting
🗓️ Published: 9/29/2026
🔗 http://arxiv.org/abs/2609.37694v1
👥 Authors: Rui Han, Min Yang (possible past Baidu (China) affiliation), Xu Zhang (possible past Tencent (China) affiliation), Xinghao Yang, Wei Liu (possible past Tsinghua University affiliation), Yongshun Gong
Abstract

Diffusion models have recently shown strong potential for probabilistic multivariate time-series forecasting by modeling complex conditional distributions. Recent decoupled diffusion frameworks further separate forecasting into deterministic prediction and stochastic residual generation, making it natural to derive dependency graphs from deterministic representations and use them to guide residual diffusion. However, we show that this direct structural transfer is unreliable. Although determinis...

📄 KUPAS MASTER: Distilling the Tacit Expertise of Master Practitioners into Agent-Ready Experience Corpora
🗓️ Published: 9/29/2026
🔗 http://arxiv.org/abs/2609.37673v1
👥 Authors: Changmian Wang, Yuchao Ma, Xuchao Lu, Chen Zhang (possible past Peking University affiliation), Ping Sun, Jiazheng Wang, Shan Wang, Xuanwen Chen, Yihe Sun, Ziyu Lu, Jianqiang Huang, Hongzhi Li, Ziqing Xia, Kaihua Tang, Xian-Sheng Hua (possible past Microsoft (United States) affiliation), Qinghua Zheng
Abstract

Experienced professionals know more than just facts and conclusions. They know which cues matter, why a judgment is reasonable, and which action to take. Routine work records often leave out this tacit knowledge, making it difficult for Large Language Model (LLM) agents to use professional experience effectively. We introduce KUPAS MASTER, an experience engineering platform built around nine-layer cognitive corpus construction. It turns heterogeneous work records and practitioner interviews into...

📄 EnterpriseBench: Benchmarking LLM Agents on Enterprise-Level Strategic Reasoning and Decision-Making
🗓️ Published: 9/29/2026
🔗 http://arxiv.org/abs/2609.37658v1
👥 Authors: Min Yang (possible past Baidu (China) affiliation), Yichen Pan, Jinghua Piao, Dandan Song, Yongshun Gong, Yong Li (possible past Tsinghua University affiliation)
Abstract

LLM agents are increasingly expected to support enterprise workflows, where tasks often involve missing information, uncertainty, feedback, and long-term trade-offs. However, existing enterprise and financial benchmarks mainly test static capabilities such as information extraction, numerical calculation, domain knowledge, and financial QA, leaving interactive and long-horizon decision-making underexplored. To bridge this gap, we introduce EnterpriseBench, a benchmark that evaluates LLM agents a...

📄 RLTL;DR: Self-improvement by Internalizing Self-generated Feedback
🗓️ Published: 9/29/2026
🔗 http://arxiv.org/abs/2609.37633v1
👥 Authors: Michael Kirchhof, Eleonora Gualdoni, Andrew Szot, Khashayar Gatmiry, Aryo Lotfi, Abbas Kazerouni (possible past Stanford University affiliation), Omar Attia, Sanjoy Chowdhury, Alexander Toshev (possible past Google (United States) affiliation)
Abstract

The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This becomes problematic in the realms of self-improvement, where tasks are so difficult that the agent has a low or even no chance of success, and where there are no teacher models or example solutions to distill from. In this paper, we introduce RLTL;DR. After each failed attempt, we show the policy the verifier outputs and le...

📄 ReLMem: Learning Recurrent Memory for Longitudinal EHR Modeling
🗓️ Published: 9/29/2026
🔗 http://arxiv.org/abs/2609.37587v1
👥 Authors: Zijie Meng, Xiwei Dai, Yingying Zhang, Jian Wu (possible past Tencent (China) affiliation), Xian Wu (possible past Tencent (China) affiliation), Zuozhu Liu
Abstract

Longitudinal electronic health record (EHR) modeling requires integrating new visits with an expanding patient history. Yet the continual accumulation of clinical information imposes increasing computational and memory costs on large language models (LLMs) when they process and retain complete patient histories. A practical alternative is visit-wise recurrent compression, which incorporates each incoming visit into a compact, continually updated patient memory. However, under a fixed memory budg...

📄 TReVS: Integrating Textual Relevance and Visual Saliency for Efficient Vision-Language Model Token Pruning
🗓️ Published: 9/29/2026
🔗 http://arxiv.org/abs/2609.37581v1
👥 Authors: Jing Wang (possible past Google (United States) affiliation), Zhiping Wu, Dongdong Ren, Youfang Han, Wei Zhao (possible past Tencent (China) affiliation), Wenbin Li
Abstract

Vision-Language Models (VLMs) excel at visual understanding and reasoning but often incur substantial inference costs due to the large number of visual tokens. Recent visual token pruning methods increasingly follow a two-stage paradigm: they first remove visually redundant tokens after the vision encoder and then discard tokens irrelevant to the textual query within the Large Language Model (LLM). However, since the first stage typically relies solely on vision-encoder saliency, it may prematur...

📄 Learning to Retrieve Missing Evidence for Long-Term Memory QA
🗓️ Published: 9/29/2026
🔗 http://arxiv.org/abs/2609.37443v1
👥 Authors: Yi-Xuan Deng, Yi Zhang (possible past Google (United States) affiliation), Wei Liu (possible past Tsinghua University affiliation), Chao Xue, Shuojin Yang
Abstract

Long-term memory enables language models to use past interactions in future conversations. However, evidence needed to answer a question may be scattered across distant turns, while the question itself omits clues needed to locate it. Retrieved facts can reveal these clues, motivating retrieval decisions conditioned on evidence already found. We introduce MERA (Missing-Evidence Retrieval Augmentation), which separates globally searchable memory from a question-specific evidence state. Verified e...

📄 Transolver-$σ$: Joint Spectral-Physical Subspace Modeling for Neural PDE Solving
🗓️ Published: 9/29/2026
🔗 http://arxiv.org/abs/2609.37279v1
👥 Authors: Haonan Shangguan, Hang Zhou (possible past Baidu (China) affiliation), Haixu Wu, Yuezhou Ma, Jianmin Wang (possible past Tsinghua University affiliation), Mingsheng Long (possible past Tsinghua University affiliation)
Abstract

Neural solvers offer efficient surrogates for numerical simulation of partial differential equations (PDEs). For time-dependent problems, strong one-step accuracy does not necessarily translate into reliable autoregressive rollout. We observe that a solver based only on physical-state modeling can achieve lower one-step error, whereas its spectral-only counterpart can become more accurate at later rollout steps. Motivated by this observation, we present Transolver-$σ$, a neural PDE solver based ...

📄 UniAfford: Token-Routed Multitask Learning for Generalizable 2D-3D Affordance Perception
🗓️ Published: 9/29/2026
🔗 http://arxiv.org/abs/2609.37264v1
👥 Authors: Yuhao Liu (possible past Baidu (China) affiliation), Yiming Zhong, Hanqing Wang, Shaocheng Yan, Yuhang Zhang, Wenzhou Lyu, Ziyang Ding, Wei Zhang (possible past Tsinghua University affiliation), Xue Zhao, Jin Pan, Yuexin Ma, Xinge Zhu
Abstract

Affordance perception aims to localize actionable regions supporting embodied interaction, yet 2D and 3D affordance grounding have evolved as separate problems, with different task definitions, supervision formats, datasets, and evaluation protocols. This fragmentation limits the learning of transferable object-affordance semantics across visual and geometric spaces. We propose Token Router for Tasks, a multitask training paradigm for MLLM-based systems that routes contextual hidden states to ta...

📄 V-JEPA Policy: Building Effective World-Action Models on Predictive Visual Latents
🗓️ Published: 9/29/2026
🔗 http://arxiv.org/abs/2609.37250v1
👥 Authors: Yang Zhang (possible past Tsinghua University affiliation), Jiangyuan Zhao, Chenyou Fan, Jiayu Hu, Xiu Yuan, Chenjia Bai, Xiu Li (possible past Tsinghua University affiliation)
Abstract

World-action models (WAMs) couple future visual-state prediction with action generation. By adapting video generators or image-editing models pretrained at scale, a prominent line of recent WAMs inherits both predictive knowledge and the models in which it was learned. We ask whether a predictive visual latent space induced by large-scale predictive pretraining can instead provide a sufficient foundation for effective WAM learning without inheriting a complete pretrained visual generative model....

📄 Interpolated Policy Distillation: A Controllable Continuum Between Off-Policy and On-Policy Distillation
🗓️ Published: 9/29/2026
🔗 http://arxiv.org/abs/2609.37170v1
👥 Authors: Youxu Shi, Yifan Sun (possible past Baidu (China) affiliation), Dacheng Yin, Haomiao Tang, Guangting Wang, Fengyun Rao (possible past Tencent (China) affiliation), Jing Lyu, Dong Liu
Abstract

Off-policy and on-policy distillation have traditionally been formulated as separate paradigms, each favoring a different property of distillation trajectories. Teacher-generated (off-policy) traces are typically high-quality but lie far from the student's distribution, whereas student-generated (on-policy) rollouts are more learnable but often contain erroneous reasoning. We view these paradigms as the endpoints of a policy continuum and posit that a more effective rollout policy may lie in bet...

📄 Lucid Dreaming for World Models: Learning to Doubt Imagination and Decide by Trust
🗓️ Published: 9/29/2026
🔗 http://arxiv.org/abs/2609.37156v1
👥 Authors: Ziqi Wen, Ting Xu, Lianyu Wang, Xian Lin, Yanda Meng, Huazhu Fu (possible past Inception Institute Of Artificial Intelligence affiliation), Meng Wang (possible past Google (United States) affiliation), Ching-Yu Cheng
Abstract

World models enable agents to learn and plan in imagination, but predictions beyond their experience can become unreliable and mislead decisions. Existing uncertainty estimates derived from predictions can remain overconfident on unfamiliar state-action pairs. We propose the Lucid World Model (LucidWM), which learns doubt from experience and propagates trust through imagination. By integrating Subjective Logic into categorical latent transitions, LucidWM distinguishes predicted outcomes from the...

📄 VACE: Validation-Gated Alternating Co-Evolution of Agent Models and Harnesses
🗓️ Published: 9/29/2026
🔗 http://arxiv.org/abs/2609.37105v1
👥 Authors: Jiexing Qi, Yu He (possible past Google (United States) affiliation), Jun Liu (possible past Tencent (China) affiliation), Qichen Huang, Shaohua Hu, Zhan Dang, Guohua Chen, Rui Yang, Wen Jiang, Yang Liu (possible past Tsinghua University affiliation), Tao Lyu, Fangming Li
Abstract

Language model agents can be improved by updating their model weights or refining the harness that guides task execution. These components are coupled: weight updates change how the model uses the harness, while harness updates change the trajectories used for training. We propose VACE, Validation-Gated Alternating CoEvolution, which alternates agentic reinforcement learning with trajectory-driven harness refinement. After each RL stage, VACE reuses the collected trajectories to propose a harnes...

📄 MatToolBench: Benchmarking Multimodal Agents in Real-World Materials Science Workflows
🗓️ Published: 9/29/2026
🔗 http://arxiv.org/abs/2609.37053v1
👥 Authors: Mei Wu, Rui Xie, Runyu Zhang, Yuqiang Li, Tianfan Fu, Bo Chen (possible past Tencent (China) affiliation), Kai Yu (possible past Baidu (China) affiliation), Xin Chen (possible past Tencent (China) affiliation), Lu Chen
Abstract

Multimodal GUI agents have achieved impressive results on general software benchmarks, yet their ability to operate professional scientific software remains largely unexplored. In materials science, sparse domain-specific web data, specialized interfaces, and tacit workflow conventions create blind spots that general-purpose pretraining cannot readily bridge. We present MatToolBench, the first real-environment benchmark for evaluating multimodal GUI agents on professional materials science softw...

📄 NowcastDiT: Diffusion Transformers are Effective Precipitation Nowcasters
🗓️ Published: 9/29/2026
🔗 http://arxiv.org/abs/2609.37038v1
👥 Authors: Haoran Xu, Xingzhuo Guo, Yuchen Zhang (possible past University Of California, Berkeley affiliation), Jincheng Zhong, Jianmin Wang (possible past Tsinghua University affiliation), Mingsheng Long (possible past Tsinghua University affiliation)
Abstract

Precipitation nowcasting demands accurate short-term forecasts under strong spatiotemporal variability. Diffusion models are well suited to modeling complex precipitation distributions, yet existing approaches often introduce increasingly specialized designs, leaving the capability of a standard diffusion architecture underexplored. We show that a standard Diffusion Transformer already provides a simple and scalable foundation for precipitation nowcasting, with domain-specific requirements accom...

📄 Explore Broadly, Reason Sharply: Push Small Models toward the Frontier via Sampling
🗓️ Published: 9/29/2026
🔗 http://arxiv.org/abs/2609.38104v1
👥 Authors: Panagiotis Theodoropoulos, Nan Jiang (possible past Stanford University affiliation), Xintong Duan, Ali Hasan, Yuriy Nevmyvaka, Evangelos A. Theodorou, Wei Deng (possible past Apple (United States) affiliation)
Abstract

Power-sharpened sampling is an inference-time alternative to reinforcement-learning (RL) post-training for enhancing reasoning in large language models (LLMs). High-probability sequences are amplified under the base model without parameter updates or external rewards, avoiding the costly optimization and jagged generalization of RL. However, this approach faces a fundamental exploration--exploitation trade-off, as % strong sharpening restricts exploration, trapping samplers in plausible but inco...

📄 Graph-Conditioned On-Policy Agent Distillation from Off-the-Shelf Teachers
🗓️ Published: 9/29/2026
🔗 http://arxiv.org/abs/2609.37522v1
👥 Authors: Xiaohan Yi, Wen Luo, Yani Huang, Junfeng Zhan, Asher Qin, Peilin Zhao (possible past Tencent (China) affiliation), Xi Xiao (possible past Tsinghua University affiliation)
Abstract

On-policy distillation (OPD) trains compact language agents with teacher feedback on student-generated trajectories. In multi-turn tasks, compounding errors can move students beyond the teacher's effective supervision. We introduce Graph-Conditioned On-Policy Agent Distillation (GC-OPD), which enriches an off-the-shelf teacher's scoring context with execution evidence. A graph indexes repeated teacher executions by shared states while preserving complete successful and failed histories. After ea...

📄 From Dissonance to Orchestration: Teacher Intervention in On-Policy Distillation
🗓️ Published: 9/29/2026
🔗 http://arxiv.org/abs/2609.37510v1
👥 Authors: Yuhao Wang, Ruiyang Ren (possible past Baidu (China) affiliation), Yinan Zhang, Ruiqing Zhang (possible past Baidu (China) affiliation), Jing Liu (possible past Baidu (China) affiliation), Chunyan Miao
Abstract

On-policy distillation (OPD) trains a student on its own reasoning trajectories using feedback from a stronger teacher. Teacher interventions can improve these trajectories, but also change the distribution on which the student learns. Our controlled studies show that rollout quality alone is an incomplete criterion for allocating teacher guidance. Deeper intervention yields diminishing gains in rollout accuracy while increasing off-policy load. In a training probe with a restricted rollout hori...

*Notable papers are those with at least two authors from a "big" AI/ML lab.