πŸ“„ Notable* Recent AI/ML arXiv Papers

Last updated just now...

πŸ“„ Turbo Harness: Instance-Adaptive Harness Optimization
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.40330v1
πŸ‘₯ Authors: Tunyu Zhang, Hao Wang (possible past Tsinghua University affiliation), Kai Xu (possible past National University Of Defense Technology affiliation), Dimitris N. Metaxas
Abstract

Automating the search for effective harnesses is an important step toward enabling agents to recursively self-improve. Existing harness optimizations typically produce a single global harness that is applied uniformly across task instances. However, a harness that works well on average may not be optimal for every instance. We introduce Turbo Harness, a framework that can adapt a globally optimized harness to each instance by reusing information generated during the original optimization process...

πŸ“„ WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.40325v1
πŸ‘₯ Authors: Ziyan Jiang, Jingbo Yang, Jiabao Ji, Yujian Liu, Qiucheng Wu, Tommi Jaakkola, Yang Zhang (possible past Tsinghua University affiliation), Shiyu Chang (possible past Tencent (China) affiliation)
Abstract

As interactive 3D worlds are increasingly used to study intelligent behavior, it becomes important to develop efficient pipelines for identifying anomalies in these simulated environments, such as floating objects, traversable walls, or objects inconsistent with the surrounding scene. Multimodal AI systems, including vision-language models (VLMs) and vision-language-action models (VLAs), have shown potential for automating this task. However, 3D world auditing is complex, requiring the close cou...

πŸ“„ Cogentic: Multi-Agent Orchestration for Automated Proof Discovery
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.40324v1
πŸ‘₯ Authors: Yang Cai, Vineet Gupta (possible past Google (United States) affiliation), Yanchen Jiang, Christopher Liaw, Aranyak Mehta (possible past Google (United States) affiliation), Grigoris Velegkas, Di Wang
Abstract

We present Cogentic, a multi-agent harness for automated proof discovery on open research problems. While frontier language models can generate strong mathematical ideas in a single shot, single-shot generation is often insufficient for open problems that require exploring multiple competing conjectures, overcoming subtle technical obstructions, and retaining intermediate progress over a long horizon. Cogentic addresses these challenges through an iterative prove--verify loop in which an orchest...

πŸ“„ PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.40285v1
πŸ‘₯ Authors: Yinghui He, Yapei Chang, Khushi Bhardwaj, Daniele Molinari, Tugrul Konuk, Jan Kautz (possible past Nvidia (United States) affiliation), Ali Hatamizadeh (possible past Nvidia (United States) affiliation)
Abstract

On-policy distillation (OPD) is a promising approach for training language agents, providing dense teacher supervision on student-generated trajectories. However, in multi-turn interaction, an incorrect action changes the states the student encounters later, so errors compound across turns. In preliminary experiments across three Qwen3 models (8B to 235B), we find that more than half of the failed rollouts contain a pivotal mistake, an action that moves the agent farther from completing the task...

πŸ“„ cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.40284v1
πŸ‘₯ Authors: Pranjal Aggarwal, Lawrence Keunho Jang, Sean Welleck, Daniel Fried, Ruslan Salakhutdinov (possible past University Of Toronto affiliation), Jing Yu Koh (possible past Google (United States) affiliation)
Abstract

Computer use agents (CUAs), which use graphical user interfaces (GUIs) to complete tasks on a computer, have recently surpassed human performance on many standard benchmarks, including difficult long-horizon tasks. Their capabilities are undoubtedly impressive, however, a key barrier to the widespread adoption and deployment of CUAs remains their speed and cost. Progress towards faster yet capable CUAs requires reliable evaluation of their speed, but many CUA benchmarks currently face a reproduc...

πŸ“„ Belief-Aware Multi-Agent Path Finding under Map Uncertainty
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.40269v1
πŸ‘₯ Authors: Viraj Parimi, Shao-Hung Chan, Han Zhang (possible past Tsinghua University affiliation), Jingkai Chen, Brian Williams (possible past Google (United States) affiliation)
Abstract

Multi-Agent Path Finding (MAPF) aims to find collision-free paths for multiple agents in a shared environment. Classical MAPF assumes that all static obstacles are known in advance, but real-world environments can change unexpectedly due to fallen objects, spills, or other local disturbances. When such changes are spatially correlated, an observation can inform traversability estimates beyond the observed location. Prior approaches address uncertainty in traversability through contingent plans o...

πŸ“„ Learning Skills from Historical Action Trajectories: Action Experience Dictionary for World Action Models
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.40219v1
πŸ‘₯ Authors: Qi Lyu, Jiahua Dong, Hao Shen, Xudong Wang, Hongyuan Yu, Baichen Liu, Henghui Ding, Zhi Han, Nicu Sebe, Ivan Laptev, Fahad Shahbaz Khan (possible past Inception Institute Of Artificial Intelligence affiliation), Salman Khan (possible past Inception Institute Of Artificial Intelligence affiliation)
Abstract

World Action Models (WAMs) couple visual dynamics prediction with action generation, yet they do not explicitly support the reuse of action experience across manipulation tasks. Furthermore, existing WAMs struggle to capture underlying cross-task semantic relationships that could guide target action prediction, as redundant background elements interfere with the extraction of key visual information. To address these challenges, we develop a novel Action Experience Dictionary (AED) that encodes h...

πŸ“„ MemLife: Curating and Reasoning over Long-Term Egocentric Video Memories
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.40195v1
πŸ‘₯ Authors: Guangzhi Xiong, Xinyuan Zhang, Xiao Yang (possible past Tencent (China) affiliation), Hyokun Yun, Kai Zhang, Shiun-Zu Kuo, Hyeonjeong Ha, Xilun Chen, Kai Sun, Lucas Liang, Guangqiang Dong, Ejaz Ahmed, Ahmed A Aly, Anuj Kumar (possible past Meta (United States) affiliation), Raffay Hamid, Aidong Zhang, Xin Luna Dong (possible past University Of Washington affiliation)
Abstract

Long-term egocentric video enables personalized AI assistants to reason about daily life. However, as video histories grow to hundreds of hours spanning months or years, reprocessing raw clips for every query becomes computationally prohibitive. Memory systems offer a scalable alternative by compacting videos into text representations, but often fail on practical benchmarks: either the memory does not preserve key evidence, or the retriever fails to locate relevant entries due to retrieval compe...

πŸ“„ Learning from Research: Toward Lifelong Agent Harness Evolution
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.40169v1
πŸ‘₯ Authors: Jingbo Yang, Kwei-Herng Lai, Xiaowen Wang, Yaar Harari, Evgeniy Gabrilovich (possible past Google (United States) affiliation), Shiyu Chang (possible past Tencent (China) affiliation)
Abstract

Language agents are expected to solve increasingly complex tasks, creating a growing need for continual improvement. One promising approach is to evolve the agent harness, the software that governs tool use, memory management, and task execution, while keeping the underlying language model fixed. Recent methods automate this process by using a meta coding agent to modify the harness based on execution feedback. However, relying on that agent's existing knowledge and observed failures can restric...

πŸ“„ Game-Guided Skill Discovery through Self-Play for Playable Agent Control
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.40137v1
πŸ‘₯ Authors: Seungeun Rho, Jeonghwan Kim, Xue Bin Peng (possible past University Of California, Berkeley affiliation), Sehoon Ha (possible past Google (United States) affiliation)
Abstract

We present Game-Guided Skill Discovery (GGSD), a framework that uses self-play in games to discover motor skills that are directly playable by humans. Playable skills provide a compact abstraction for controlling embodied agents through a small set of learned behaviors rather than low-level actions. To be effective, these skills should be semantically distinct, interpretable, and expressive; properties that existing unsupervised skill-discovery methods often fail to achieve simultaneously. GGSD ...

πŸ“„ Tactile Curiosity Drives Robot Interaction
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.40134v1
πŸ‘₯ Authors: Klemens Iten, Alexander Proshkin, Bhavya Sukhija, Stelian Coros, Andreas Krause (possible past Eth Zurich affiliation), Pieter Abbeel (possible past University Of California, Berkeley affiliation), Carmelo Sferrazza
Abstract

Mastering robot manipulation skills via reinforcement learning (RL) remains largely sample-inefficient. The most common RL algorithms rely on random action sampling to discover new strategies, resulting in agents that allocate most of their training budget to motions in free space, away from the contacts from which manipulation skills emerge. Existing intrinsic motivation methods based on model disagreement or epistemic uncertainty improve on isotropic noise, but they can also reward uncertainty...

πŸ“„ LongEmo: Towards Emotion Understanding and Reasoning in Long Videos
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.40079v1
πŸ‘₯ Authors: Shuo Zhang (possible past National University Of Defense Technology affiliation), Yifan Zhou, Han Wang (possible past Peking University affiliation), Jinsong Zhang, Jingyu Li, Hongbing Li, Zhejun Zhang, Chengyi Zhao, Yuquan Hao, Yitong Liu, Jiyin Li, Ruiqi Tang, Zixuan Lin, Yi Luo, Xurui Zhang, Ronghao Chen, Huacan Wang, Lei Li (possible past Carnegie Mellon University affiliation)
Abstract

While recent Multimodal Large Language Models (MLLMs) have shown promise in affective computing, their reasoning capabilities are largely confined to short video clips with limited interactions. However, real-world emotions are not merely isolated instantaneous reactions but dynamic and cumulative processes deeply shaped by past experiences and ongoing events. To bridge this gap, we introduce LongEmoBench, a benchmark dedicated to emotion understanding and reasoning in long videos. It assesses p...

πŸ“„ CoVisco: Codec-Native Vision Encoder with Native Token Compression for Unified Image-Video Understanding
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.39924v1
πŸ‘₯ Authors: Yulong Liu, Xiaotian Han, Junyuan Shang, Yuchen Ding (possible past Baidu (China) affiliation), Zhenyu Zhang, Shuohuan Wang (possible past Baidu (China) affiliation), Guibo Zhu, Sirui Han, Dianhai Yu (possible past Baidu (China) affiliation)
Abstract

Vision-language models face a fundamental scaling bottleneck: the number of visual tokens grows with both temporal duration and spatial resolution, making long-video understanding expensive for the vision encoder and the language model. Existing methods often compress visual tokens after dense encoding, creating a mismatch between the representation used during training and the compact interface required at deployment. We present CoVisco, a codec-native vision encoder with native token compressi...

πŸ“„ OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.39903v1
πŸ‘₯ Authors: Dingyuan Dai, Heli Qi, Lei Liu, Yinxi Li, Baiding Chen, Zijun Dou, Qingcheng Zeng, Qi Kang, Oliver Sun, Eric Wang, Bo Zhou, Haixin Wang, Yufan Du, Shi Bo, Ruihan Lin, Mengqi Yuan, Dunjie Lu, Steven Dillmann, Yiming Shi, Tina Su, Amy Xin, Minghao Liu, Xi Wang (possible past Tsinghua University affiliation), Xu Huang, Ge Zhang, Pengyu Nie, Zhen Yang (possible past Tsinghua University affiliation), Jie Tang (possible past Tsinghua University affiliation), Juanzi Li, Weihao Xuan, Tianyu Liu
Abstract

Scientific software presents a demanding test for computer-using agents based on visual language models (VLMs): completing a research workflow requires interpreting specialized interfaces, manipulating scientific objects, and producing verifiable results. We thus introduce OSWorld-Science, a benchmark and evaluation environment that combines scientifically meaningful tasks, artifact-based evaluation, and an efficient agent harness for studying computer use in the scientific domain. The benchmark...

πŸ“„ GFD-OPD: Guidance-Folded On-Policy Distillation of Diffusion Models Across Scales
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.39692v1
πŸ‘₯ Authors: Zhenxing Zhang, Jiayan Teng, Wenxu Wu, Zhuoyi Yang, Jiazheng Xu, Wendi Zheng, Jie Tang (possible past Tsinghua University affiliation), Dan Guo, Meng Wang (possible past Google (United States) affiliation)
Abstract

On-policy distillation (OPD) has demonstrated two important capabilities in language models: compressing large teachers into smaller students and merging expert models into a single model. Existing diffusion OPD, however, mostly focus on the latter, with teachers and students sharing the same backbone and scale. We investigate large-to-small diffusion opd from large teachers to a small student and find that the standard recipe fails. To find the underlying cause, we propose Fixed-State KL, an ef...

πŸ“„ ChronoGraph: Functional 4D Scene Graphs with Vision-Language Models for Interaction Understanding and Grounded Planning
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.39665v1
πŸ‘₯ Authors: Chenyangguang Zhang, Malgorzata Gwiazda, Guanlong Jiao, Yuanchen Ju, Federico Tombari (possible past Google (United States) affiliation), Koushil Sreenath, Marc Pollefeys (possible past Google (United States) affiliation), Sunghwan Hong
Abstract

Embodied agents must determine where to act, anticipate the resulting scene changes, and interpret observed outcomes to guide subsequent actions. This requires connecting 4D interaction understanding, which explains how past actions changed the scene, with spatially grounded planning, which determines how and where to act toward a goal and anticipates the resulting scene changes. We introduce ChronoGraph, a functional 4D scene graph that links actions on affordance parts to semantic and geometri...

πŸ“„ GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.39601v1
πŸ‘₯ Authors: Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen (possible past Google (United States) affiliation), Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei (possible past Google (United States) affiliation), Ruihai Wu, Hang Zhang (possible past Amazon (United States) affiliation), Yixiao Ge (possible past Tencent (China) affiliation), Shuchang Zhou, Shilong Liu, Xianming Liu (possible past Meta (United States) affiliation), Ping Luo (possible past Shanghai Artificial Intelligence Laboratory affiliation), Shiyu Huang
Abstract

Precise grounding matters. It specifies which object is the target and where that object is, even in clutter and for tiny objects, and it has to be fast enough for closed-loop control. Yet vision-language-action (VLA) and world-action models (WAMs) take perception from general-purpose vision-language and video-generation backbones, which still fail in these settings. We introduce GroundingPI, a 4B grounding foundation model that generates points and boxes as quantized coordinates in a shared voc...

πŸ“„ GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.39600v1
πŸ‘₯ Authors: Qize Yu, Lianrui Fan, Bowen Ping, Xini Ding, Zetian Song, Junbo Niu, Kaixuan Wang, Tianxing Chen, Yue Chen (possible past Google (United States) affiliation), Minghua He, Yuran Wang, Jie Huang, Haojun Zhang, Min Chen, Hao Li (possible past Tsinghua University affiliation), Wenxuan Song, Ruihai Wu, Xianming Liu (possible past Meta (United States) affiliation), Shilong Liu, Shuchang Zhou, Ping Luo (possible past Shanghai Artificial Intelligence Laboratory affiliation), Shiyu Huang
Abstract

Autoregressive (AR) grounding models serialize spatial predictions, introducing sequential latency and imposing a causal order on output tokens. We view grounding as visual evidence extraction: objects, locations, and spatial relations are jointly constrained by the image and query, yet their dependencies do not imply an intrinsic left-to-right generation order. This distinction makes bidirectional diffusion a natural fit, allowing spatial hypotheses to emerge in parallel and be jointly refined ...

πŸ“„ Self-Spec Verifiable Code Generation
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.39568v1
πŸ‘₯ Authors: Jiaru Qian, Yihong Dong, Yongmin Li, Hao Zhu (possible past Tsinghua University affiliation), Bin Gu, Ge Li (possible past Peking University affiliation)
Abstract

Large language models (LLMs) may generate unreliable code on corner cases missed by testing, while formal verification can provide machine-checkable guarantees. Recently, researchers have proposed several benchmarks to evaluate the capabilities of LLMs in generating formally verifiable code, where LLMs need to formulate formal specifications, generate the corresponding code, and verify its correctness. However, existing benchmarks have two key limitations: (I) They primarily evaluate specificati...

πŸ“„ OmniReasoning: Pushing the Limits of Audio-Visual Joint Reasoning
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.39490v1
πŸ‘₯ Authors: Junming Lin, Yuxuan Wang (possible past Google (United States) affiliation), Zhenxin Lei, Yuxin Liu, Ruixun Liu, Yinsong Yan, Ling Wang (possible past University Of Oxford affiliation), Minghao Han, Yunfei Chu, Shun Lei, Xueyao Zhang, Qize Yang, Jin Xu (possible past Tencent (China) affiliation), Yiwu Zhong
Abstract

Recent advances have enabled unified omni-modal models in understanding audio, vision, and language. However, existing benchmarks, training data, and learning methods largely treat the modalities independently, leaving the capability of audio-visual joint reasoning poorly evaluated and insufficiently elicited. We address this gap with a benchmark, data engine, and learning method. First, we introduce OmniReasoningBench, a benchmark where both audio and visual evidence are indispensable. It compr...

πŸ“„ Fyan: A Human--AI Harness with Semantic Auditing for Document-Level Formalization
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.39228v1
πŸ‘₯ Authors: Wei Zhao (possible past Tencent (China) affiliation), Yangshuo Zou, Chengxiang Ding, Yifan Wu (possible past Carnegie Mellon University affiliation), Xuchuan Wang, Zimu Mao, Lei Zhang, Tao Luo
Abstract

We present FYAN, a human--AI harness for document-level mathematical formalization. Rather than treating theorems in isolation, FYAN coordinates an end-to-end workflow spanning specification, proof planning, logical review, Lean proof construction, knowledge curation, and validation, with support for independent supervision and human guidance. A central component is evidence-grounded semantic auditing, which assesses whether formal statements faithfully preserve their informal specifications. A ...

πŸ“„ RefCon: Iterative Refinement and Contrastive Memory Extraction for Context-Evolving Agent
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.39143v1
πŸ‘₯ Authors: Ubaidillah Ariq Prathama, Bo Liu (possible past Meta (United States) affiliation), Yeo Boon Hong, Yu-Xuan Huang, Yangkai Ding, Tao Yu (possible past University Of Washington affiliation)
Abstract

Long-horizon agent interactions generate useful but noisy experience, and retraining models to absorb it is expensive. Context-evolving agents therefore need memory extraction methods that improve with more test-time compute without relying on gold labels. We propose RefCon, which combines sequential self-refinement with parallel self-contrast to extract higher-quality memories without gold labels. Evaluated on AppWorld and BFCL-V3 across multiple context-evolving agent frameworks, RefCon delive...

πŸ“„ Trustworthy Runtime Error Healing in Real-World Repositories: A Benchmark and Guardrail
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.39086v1
πŸ‘₯ Authors: Gou Tan, Pengfei Chen, Zhensu Sun, Jieke Shi, Junkai Chen, Ting Zhang (possible past Meta (United States) affiliation), Weifeng Sun, Junda He, Shuai Liang, Chuanfu Zhang, Lwin Khin Shar, David Lo (possible past Stanford University affiliation)
Abstract

Runtime error healing lets a crashed program continue by generating code that repairs its live runtime state. Recent work shows that LLMs can generate such healing code, but it is evaluated only on small competition programs, and executing LLM-generated code inside a live process raises safety concerns that remain unaddressed. In this paper, we take LLM-based runtime healing toward practical use in real-world repositories. We first build HealBench, a benchmark of 265 runtime errors from 18 real-...

πŸ“„ Can Agents Trust Their Skills? Uncovering Unsafe Chains of Trust in Skill-Based LLM Agents
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.39065v1
πŸ‘₯ Authors: Yan Wang (possible past Tencent (China) affiliation), Zhihao Zhang, Ke Chen (possible past Tencent (China) affiliation), Kai Chen (possible past Shanghai Jiao Tong University affiliation), Yaqin Zhang, Duohe Ma, Jun Dai, Xiaoyan Sun
Abstract

LLM agents increasingly rely on installable skills, which are packages of instructions, code, and resources that equip them with task-specific capabilities and, once installed, can be automatically invoked across subsequent user tasks. This creates a chain of trust in which users delegate authority to agents, while agent frameworks admit skill-provided content into the agents' context with insufficient validation, allowing malicious skills to influence agent behavior under that delegated authori...

πŸ“„ Structure-aware Reinforcement Learning for Protein Directed Evolution
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.39048v1
πŸ‘₯ Authors: Zikun Nie, Suyuan Zhao, Yizhen Luo (possible past Tsinghua University affiliation), Siqi Fan, Zaiqing Nie (possible past Tsinghua University affiliation)
Abstract

Protein optimization remains a longstanding goal in life sciences. Existing machine learning-assisted directed evolution (MLDE) methods primarily rely on sequence-only features, overlooking the critical spatial constraints and co-evolutionary interactions encoded in protein structures. However, directly integrating structural information remains challenging due to the scarcity of reliable mutant structures. To address these issues, we propose StructEvo, a novel structure-aware reinforcement lear...

πŸ“„ Looped Diffusion Transformer
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.40305v1
πŸ‘₯ Authors: Yong Xien Chng, Tianyi Chen, Wenwen Tong, Haiwen Diao, Zhongang Cai, Lei Yang (possible past Google (United States) affiliation), Ziwei Liu, Lewei Lu, Dahua Lin, Gao Huang (possible past Tsinghua University affiliation)
Abstract

Improving text-to-image models has traditionally relied on increasing model size or the number of denoising steps. In this work, we explore an alternative way to scale computation by repeatedly running shared Transformer blocks within each denoising step, effectively increasing computational depth while keeping the parameter count fixed. This looped computation enables iterative refinement of internal representations without explicit reasoning tokens. However, naive looping fails to consistently...

πŸ“„ How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.40295v1
πŸ‘₯ Authors: Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting (possible past Google (United States) affiliation), Mohit Iyyer (possible past Google (United States) affiliation), Max Spero, Bradley Emi
Abstract

Web text makes up the majority of pretraining data and is increasingly AI-generated. After applying FineWeb quality filtering, we find that 27.5% of tokens from June 2026 web data are labeled as AI-generated by Pangram, rising to 31.1% by August. Unlike synthetic data or model-collapse setups, this *wild* AI text comes from many models, is written for human readers, and arrives unlabeled in pretraining corpora. How does AI text in the wild affect language model pretraining? To answer this questi...

πŸ“„ Beyond Accuracy: Prefix-Invariant Realizations of Low-Precision Fast Matrix Multiplication
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.39816v1
πŸ‘₯ Authors: Shuxiao Xie, Shuyang Xie, Yuan Cao (possible past Google (United States) affiliation), Dezhi Ran, Wei Yang (possible past Tencent (China) affiliation), Tao Xie
Abstract

Fast matrix multiplication saves multiplications through exact cancellation, but rounding sums that mix token rows can leave contributions from later tokens in earlier language model outputs. This threatens prefix invariance, which multiple-choice likelihood scoring relies on: a scored likelihood must depend only on its allowed prefix. On Qwen2.5-14B-Instruct, two fast FP8 realizations repaired to ordinary-looking accuracy still change the answers chosen by likelihood on 5.83% and 10.00% of 240 ...

πŸ“„ A library for differentiable signal processing and machine learning on the sphere
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.39737v1
πŸ‘₯ Authors: Thorsten Kurth (possible past Nvidia (United States) affiliation), Max Rietmann (possible past Nvidia (United States) affiliation), Mauro Bisson, Andrea Paris, Alberto Carpentieri, Jean Kossaifi (possible past Nvidia (United States) affiliation), Anima Anandkumar (possible past Nvidia (United States) affiliation), Christian Hundt, Boris Bonev
Abstract

The two-dimensional sphere embedded in three-dimensional Euclidean space S2, plays a central role in a variety of scientific and engineering domains, including geophysics, planetary science, geodesy, atmospheric physics, quantum chemistry, cosmology, and virtual reality, among many others. As machine learning increasingly permeates these fields, the demand grows for robust tools that process and model functions on the sphere, while respecting the inherent topological and symmetry properties of t...

πŸ“„ KilometerVision: A New Frontier for Large-Scale Spatial Intelligence in VLMs
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.39588v1
πŸ‘₯ Authors: Aravindh Mahendran (possible past Google (United States) affiliation), Michael King, Matthew Koichi Grimes, Antoine Yang, Tyler Zhu (possible past University Of California, Berkeley affiliation), Joseph Heyward, Tengda Han, Shiry Ginosar, Chen Sun (possible past Google (United States) affiliation), Dima Damen, Simon Osindero (possible past Google (United States) affiliation), Noah Snavely (possible past Google (United States) affiliation), Simon Lynen (possible past Eth Zurich affiliation), JoΓ£o Carreira (possible past University Of California, Berkeley affiliation), Viorica PΔƒtrΔƒucean
Abstract

We push the frontier of large-scale spatial intelligence in Vision-Language Models (VLMs) and introduce the first benchmark that probes geographical layout understanding from real-world videos, spanning up to 1km distances. Inspired by the cognitive science literature, we evaluate models against the hierarchical stages of human spatial awareness: anchoring via landmarks, connecting them through routes, and integrating these into global mental maps. Extensive experiments reveal a fundamental dive...

πŸ“„ HAPMoE: Heterogeneity-Aware Automatic Parallelism Planning for Mixture-of-Experts Models Training
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.39350v1
πŸ‘₯ Authors: Mengyuan Fan, Peizhuang Cong, Zixiao Huang, Si Xu, Tong Qiao, Yanghao Li (possible past Meta (United States) affiliation), Jing Yang, Tong Yang (possible past Peking University affiliation), Quanlu Zhang, Yu Wang (possible past Tsinghua University affiliation)
Abstract

As model sizes continue to scale, distributed training has become inevitable. Automatic parallelization techniques can derive efficient training parallelism strategies at low cost while achieving superior performance. The difficulty of this problem is jointly determined by the complexity of the model and the underlying compute cluster. Meanwhile, mixture-of-experts (MoE) models are increasingly emerging as the dominant architecture and the rapid evolution of accelerator hardware has made cluster...

πŸ“„ ElectrolyteFM: Unifying Electrolyte Property Prediction through Cross-Property Knowledge Learning
πŸ—“οΈ Published: 9/30/2026
πŸ”— http://arxiv.org/abs/2609.39340v1
πŸ‘₯ Authors: Jiaxin Yu, Shuo Wang (possible past Nvidia (United States) affiliation), Peng Wang (possible past Peking University affiliation), Yongcai Wang, Deying Li
Abstract

Electrolyte formulation design requires balancing multiple physicochemical properties, yet existing models often focus on a limited subset. Learning each property in isolation can overlook transferable chemical information, whereas indiscriminate sharing can introduce cross-property interference. Our directed transfer analysis shows that jointly learning two property prediction tasks can improve or degrade prediction relative to separate training, with asymmetric transfer effects between the tas...

*Notable papers are those with at least two authors from a "big" AI/ML lab.