๐Ÿ“„ Notable* Recent AI/ML arXiv Papers

Last updated just now...

๐Ÿ“„ Long-WAM: Scaling the Context of World-Action Models
๐Ÿ—“๏ธ Published: 10/7/2026
๐Ÿ”— http://arxiv.org/abs/2610.10528v1
๐Ÿ‘ฅ Authors: Wei Huang (possible past Google (United States) affiliation), Bohan Zhang, Chenzhi Liu, Isabella Liu, Shuai Yang, Weian Mao, Luozhou Wang, Yicheng Xiao, Weifeng Lin, Qixin Hu, Bryan Chu, Sifei Liu (possible past Nvidia (United States) affiliation), Linxi Fan, Xiaojuan Qi (possible past University Of Oxford affiliation), Song Han (possible past Stanford University affiliation), Yukang Chen
Abstract

Real-time robot control demands enough visual history to infer motion and task progress, but processing that history can delay action. We present Long-WAM, a model-system framework for scaling the context of causal world-action models under real-time control constraints. Our central finding is that access to history is not the same as using it: longer histories pay off far more when the video foundation is pretrained autoregressively (AR). We first learn causal prediction from robot and egocentr...

๐Ÿ“„ RoboJEPA: Scaling Robotic Latent World Models
๐Ÿ—“๏ธ Published: 10/7/2026
๐Ÿ”— http://arxiv.org/abs/2610.10515v1
๐Ÿ‘ฅ Authors: Artem Zholus, Nicolas Beltran-Velez, Jianhao Yuan, Sarath Chandar (possible past Mila - Quebec Artificial Intelligence Institute affiliation), Tushar Nagarajan, Daniel Severo, Koustuv Sinha, Michal Drozdzal, Adriana Romero Soriano, Jeannette Bohg (possible past Stanford University affiliation), Nicolas Ballas, Mahmoud Assran
Abstract

Latent world models have shown a remarkable ability to predict future states and to plan in the real world. In practice, however, we lack a principled way to estimate how their capabilities scale with model size, data, and compute, an open problem that slows progress in the field. In this work we present RoboJEPA, a world model based on the Joint Embedding Predictive Architecture (JEPA) and trained on a large-scale dataset spanning 12 robotic embodiments. We show that RoboJEPA's imagination erro...

๐Ÿ“„ Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models
๐Ÿ—“๏ธ Published: 10/7/2026
๐Ÿ”— http://arxiv.org/abs/2610.10478v1
๐Ÿ‘ฅ Authors: Tan Yu (possible past Baidu (China) affiliation), Alexander Bukharin, Khushi Bhardwaj, Jennifer Williams, Zirui Liu, Jonathan Lingjie Li, Soumye Singhal, Joseph Jennings, Sanjeev Satheesh (possible past Baidu (China) affiliation), Yash Jain, Ashish Vaswani (possible past Google (United States) affiliation), Venkat Krishna Srinivasan, Matthew Papakipos, Hyunwoo Kim, Jian Zhang (possible past Tencent (China) affiliation), Oleksii Kuchaiev (possible past Nvidia (United States) affiliation), Markus Kliegl, Mostofa Patwary (possible past Nvidia (United States) affiliation), Mohammad Shoeybi (possible past Nvidia (United States) affiliation), Bryan Catanzaro (possible past University Of California, Berkeley affiliation), Jonathan Cohen (possible past Nvidia (United States) affiliation), Jiantao Jiao
Abstract

How can we predict which base checkpoint is worth an expensive round of agentic post-training? End-to-end pass@$K$ tests whether successful behavior already appears in a base model's distribution, but it is a poor fit for agentic coding: many base checkpoints cannot reliably produce the well-formed tool invocation required to complete a task end-to-end. Single-shot or short-horizon tasks avoid these tool-calling failures by collapsing a multi-step interaction into a fixed prompt and a single pat...

๐Ÿ“„ MIRA: A Musical Intent Refinement Agent for Aligning Text-to-Music Generation with User Intent
๐Ÿ—“๏ธ Published: 10/7/2026
๐Ÿ”— http://arxiv.org/abs/2610.10355v1
๐Ÿ‘ฅ Authors: Zekai Liu, Zhilin Wang, Xuzheng He, Yu Cheng (possible past National University Of Singapore affiliation), Yang Yang (possible past Tencent (China) affiliation)
Abstract

Text-to-music systems produce increasingly convincing audio, yet evaluation reveals little about whether the result matches user intent. A global text-audio relevance score can overlook the implicit intent in underspecified prompts and mask failures in specific requirements, such as instrumentation, structure, rhythm, or mood progression. To bridge this gap, we formulate text-to-music intent alignment as satisfying a per-request rubric of independently verifiable items covering both a request's ...

๐Ÿ“„ Fault-tolerant foundation models
๐Ÿ—“๏ธ Published: 10/7/2026
๐Ÿ”— http://arxiv.org/abs/2610.10311v1
๐Ÿ‘ฅ Authors: Trevor Mccourt (possible past Google (United States) affiliation), Ila R. Fiete, Isaac L. Chuang (possible past Massachusetts Institute Of Technology affiliation)
Abstract

Emerging computer hardware often trades reliability for energy efficiency; here we show that large-language models (LLMs) can be trained to tolerate this unreliability, and that rather than degrading, their error resilience actually increases as they grow. Modified neural scaling laws inferred from 40,000 GPU-hours of training runs on simulated faulty digital hardware quantify this trend and suggest that models learn to compute within "good" error-correcting codes, whose relative overhead remain...

๐Ÿ“„ ExperienceIndex: Artifact-Grounded Memory
๐Ÿ—“๏ธ Published: 10/7/2026
๐Ÿ”— http://arxiv.org/abs/2610.10091v1
๐Ÿ‘ฅ Authors: Peter Baile Chen, Geoffrey X. Yu, Xinming Liu, Samuel Madden (possible past Massachusetts Institute Of Technology affiliation), Dan Roth, Jacob Andreas (possible past University Of California, Berkeley affiliation), Doug Downey (possible past Allen Institute For Artificial Intelligence affiliation), Michael Cafarella (possible past University Of Washington affiliation)
Abstract

Knowledge-intensive tasks require answering many questions by reasoning about a shared corpus of artifacts (e.g., court cases, or scientific literature). As humans interact with these corpora, they naturally accumulate experiential knowledge about artifacts, enabling them to quickly identify the complete set of relevant artifacts for each new task. However, existing AI agents lack appropriate memory solutions to build or reuse such artifact-grounded experience, leading to lower answer quality an...

๐Ÿ“„ From Pixel to Coding: Evaluating the Figure Reproduction Capabilities of MLLMs
๐Ÿ—“๏ธ Published: 10/7/2026
๐Ÿ”— http://arxiv.org/abs/2610.10066v1
๐Ÿ‘ฅ Authors: Zijian Chen, Zhengyu Chen, Bohan Liang, Lirong Deng, Yushuo Zheng, Yanwei Jiang, Qi Jia, Kaiwei Zhang, Wenjun Zhang (possible past Shanghai Jiao Tong University affiliation), Guangtao Zhai (possible past Shanghai Jiao Tong University affiliation)
Abstract

Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in both visual understanding and code generation. However, existing benchmarks typically evaluate these two modalities in isolation, lacking a dedicated assessment of their unification, i.e., how a model can perceive complex visual structures and synthesize them into precise, executable code. Moreover, current visual code generation benchmarks often rely on simplified layouts within single programming environments...

๐Ÿ“„ Learning to Accumulate Knowledge with Mutual Information
๐Ÿ—“๏ธ Published: 10/7/2026
๐Ÿ”— http://arxiv.org/abs/2610.10042v1
๐Ÿ‘ฅ Authors: Yuyang Zhao, Lizi Liao, Leyang Shen, Xiaoyan Zhao, Yang Zhang (possible past Tsinghua University affiliation), Fuli Feng (possible past National University Of Singapore affiliation), Xiangnan He (possible past National University Of Singapore affiliation)
Abstract

Large language model (LLM) agents can improve their performance by reusing knowledge distilled from past interactions. However, curating new experiences into a knowledge bank that becomes more useful as it grows remains challenging. Effective knowledge accumulation should limit redundant overlap among entries and ensure that new knowledge contributes beyond what the bank already provides. Yet training a curator with Group Relative Policy Optimization (GRPO) on standalone task success can reinfor...

๐Ÿ“„ UltraText Bench: A Comprehensive Bilingual Benchmark for Evaluating Visual Text Rendering in Image Generation
๐Ÿ—“๏ธ Published: 10/7/2026
๐Ÿ”— http://arxiv.org/abs/2610.09823v1
๐Ÿ‘ฅ Authors: Deyuan Liu, Yihao Hu, Jingxuan Zhang, Xingying Li, Jun Xie (possible past Tencent (China) affiliation), Jiacheng Liu, Jungang Li, Yu Huang (possible past Tencent (China) affiliation), Xuanyi Liu, Yue Ding, Zecheng Wang, Lei Zhao, Mingda Wang, Zhenglin Cheng, Peng Sun (possible past Tencent (China) affiliation), Tao Lin
Abstract

Dense visual text requires image generators to reproduce long strings across multiple regions with correct placement and legibility. As short-string rendering improves, evaluation must test sustained performance across more demanding scenes. We introduce UltraText Bench, a bilingual benchmark for prompt-only generation of dense visual text. It contains 432 prompts spanning 24 real-world scene categories and three difficulty levels, split equally between English and Chinese. Each human-reviewed p...

๐Ÿ“„ Artificial intelligence pathways from weather to climate
๐Ÿ—“๏ธ Published: 10/7/2026
๐Ÿ”— http://arxiv.org/abs/2610.09770v1
๐Ÿ‘ฅ Authors: Tom Beucler, J. David Neelin, Hui Su (possible past Tencent (China) affiliation), Shivanshi Asthana, Chris Bretherton, Will Chapman, Costa Christopoulos, Spencer K. Clark, Aditya Grover (possible past University Of California, Berkeley affiliation), Ignacio Lopez-Gomez, Tapio Schneider, Adam Subel, Oliver Watt-Meyer
Abstract

Deep learning has made rapid advances in weather forecasting: autoregressive models trained on atmospheric reanalyses now rival dynamical models across nowcasting, medium-range, and subseasonal-to-seasonal lead times, producing well-calibrated ensemble forecasts at reduced cost. We review these advances and consider their extension to climate horizons, where the challenge shifts from initial-condition skill to producing reliable statistical responses under altered forcings. AI-powered climate pr...

๐Ÿ“„ Dual- versus Single-Suggestion AI Support for Radiographic Interpretation in Residents: Randomized Multireader Study
๐Ÿ—“๏ธ Published: 10/7/2026
๐Ÿ”— http://arxiv.org/abs/2610.09589v1
๐Ÿ‘ฅ Authors: Lin Wu, Zhe Xu, Hongyi Wang, Feifei Zhou, Wei Deng (possible past Apple (United States) affiliation), Chunlong Zhang, Yuting Zhu, Kaixiao Chen, Xiao Liang, Chen Yang (possible past Tencent (China) affiliation), Yeyuan Chen, Hao Chen, Fuqing Zhou
Abstract

Purpose: To compare dual- and single-suggestion AI support for radiographic interpretation by residents, particularly when the shared AI suggestion was incorrect. Materials and Methods: This prospective, multicenter, randomized three-arm reader study was conducted at three hospitals in China from July to September 2026 (ChiCTR2600129243). After specialty stratification, 132 residents with fewer than 3 years of clinical experience were randomized 1:1:1 to GPT-5.4 alone (group A), GPT-5.4 plus K...

๐Ÿ“„ Correspondences as Decisions: JevNexus for Decision-Centric Schema Matching
๐Ÿ—“๏ธ Published: 10/7/2026
๐Ÿ”— http://arxiv.org/abs/2610.09487v1
๐Ÿ‘ฅ Authors: Runze Li, Hanchen Wang (possible past University Of Cambridge affiliation), Ying Zhang (possible past Tencent (China) affiliation), Wenjie Zhang
Abstract

Schema matching increasingly uses generative language models to rerank retrieved column candidates, although the underlying task is a bounded correspondence decision. We present JevNexus, which combines typed pairwise decisions with schema/instance evidence and invokes listwise refinement only when the evidence disagrees and the fused margin is small. The evaluation covers 561 cases from six benchmark families. JevNexus obtains dataset-macro MRR and Hits@1 of 0.930 and 0.909, compared with 0.926...

๐Ÿ“„ GeoPrior-Mamba: Structured Process Priors with Mamba for Fine-Resolution XCO2 Reconstruction
๐Ÿ—“๏ธ Published: 10/7/2026
๐Ÿ”— http://arxiv.org/abs/2610.09456v1
๐Ÿ‘ฅ Authors: Zhao Meng, Yinan Cai, Siru Zhong, Juepeng Zheng (possible past Tsinghua University affiliation), Haohuan Fu (possible past Tsinghua University affiliation)
Abstract

Reconstructing fine-resolution column-averaged dry-air CO2 (XCO2) fields from sparse satellite observations requires models to infer spatial structure that is only weakly constrained by direct measurements. Existing learning-based methods typically treat environmental covariates as ordinary numerical inputs and must therefore learn heterogeneous source-sink relationships largely from sparse supervision. We introduce GeoPrior-Mamba, a multi-directional Mamba framework augmented with offline langu...

๐Ÿ“„ RSI-Forge: From Research Papers to Environments for Recursive Self-Improvement
๐Ÿ—“๏ธ Published: 10/7/2026
๐Ÿ”— http://arxiv.org/abs/2610.09426v1
๐Ÿ‘ฅ Authors: Renxiong Wang, Darvin Yi (possible past Stanford University affiliation), Abril Herrlein, Anas Mahmoud, Advait Gosai, Lisiman Hua, Mohammadhossein Rezaei, Xingang Guo, Anisha Gunjal, Utkarsh Tyagi, David J. Lee, Minglai Yang, Haris Riaz, Chenguang Wang (possible past Amazon (United States) affiliation), Huaxiu Yao, Daniel Yue Zhang, Aakash Sabharwal, Tong Zhao, Yunzhong He
Abstract

Environments are the foundation of recursive self-improvement: they provide the problems agents work on and the feedback used to evaluate progress. Yet constructing challenging research environments with reliable evaluation still depends on domain experts, limiting their scale and disciplinary coverage. We introduce RSI-Forge, a multi-agent pipeline that turns published papers into executable environments for self-improvement. Three agents coordinate construction, reproduction, and review to pro...

๐Ÿ“„ Shared Geometry As A Rosetta Stone: Cross-Modal Alignment Without Paired Data
๐Ÿ—“๏ธ Published: 10/7/2026
๐Ÿ”— http://arxiv.org/abs/2610.09411v1
๐Ÿ‘ฅ Authors: Dominik Schnaus, Thomas Dagรจs, Daniel Cremers, Xi Wang (possible past Tsinghua University affiliation), Phillip Isola (possible past University Of California, Berkeley affiliation)
Abstract

Multimodal representations enable zero-shot classification and retrieval, but aligning independently trained models usually requires large amounts of paired data. Yet, the Platonic Representation Hypothesis suggests that models trained on different modalities may converge spontaneously toward a shared representation geometry. But then, do we even need paired examples for cross-modal alignment? Remarkably, we show that paired examples are unnecessary for coarse cross-modal alignment. Our simple W...

๐Ÿ“„ Conditional Accuracy Profiles: Diagnosing LLM Judges across Deployment Conditions
๐Ÿ—“๏ธ Published: 10/6/2026
๐Ÿ”— http://arxiv.org/abs/2610.09229v1
๐Ÿ‘ฅ Authors: Wenqi Li (possible past Nvidia (United States) affiliation), Bin Liu, Mindi Ruan, Chuanbo Hu, Minglei Yin, Xin Li (possible past Google (United States) affiliation)
Abstract

LLM-as-judge is now a standard tool for scalable evaluation, but judge performance is still often summarized by a single accuracy number. This aggregate view hides the deployment conditions under which a judge succeeds or fails. We introduce \textbf{Conditional Accuracy Profiling} (CAP), a post-hoc diagnostic framework that decomposes pairwise LLM-judge accuracy into eight conditions organized into content sensitivity, robustness, and rationale quality. CAP is benchmark-agnostic: it can be appli...

๐Ÿ“„ RLDISCOVER: LLM-driven co-evolution of reinforcement learning algorithms
๐Ÿ—“๏ธ Published: 10/6/2026
๐Ÿ”— http://arxiv.org/abs/2610.09218v1
๐Ÿ‘ฅ Authors: Haoran Li, Zengle Ge, Xiaomin Yuan, Yui Lo, Songlin Zhou, Jiahua Ying, Haoxin Li, Qianhui Liu, Yuanhang Liu, Jiaqun Liu, Guokai Chen, Mingju Chen, Ruinan Wang, Annan Li, Jianmin Wu, Dawei Yin (possible past Baidu (China) affiliation), Dou Shen (possible past Baidu (China) affiliation)
Abstract

LLM-guided program evolution has enabled discoveries in mathematics and computational optimization, raising the prospect of reinforcement learning (RL) algorithms that self-evolve to improve how agents learn. However, realizing this prospect faces two obstacles. Joint search over coupled algorithmic components is difficult to scale: simultaneous changes can disrupt learning, while isolated changes overlook their dependencies. Evaluating candidate algorithms also requires costly training, with fi...

๐Ÿ“„ A Deterministic Evidence Layer for Vision-Language Autism Screening from Naturalistic Home Video
๐Ÿ—“๏ธ Published: 10/6/2026
๐Ÿ”— http://arxiv.org/abs/2610.09217v1
๐Ÿ‘ฅ Authors: Wenqi Li (possible past Nvidia (United States) affiliation), Mindi Ruan, Chuanbo Hu, Shuo Wang (possible past Nvidia (United States) affiliation), Xin Li (possible past Google (United States) affiliation)
Abstract

Autism spectrum disorder (ASD) is diagnosed through specialist observation of a child's social behavior, and access to that expertise is the bottleneck for early identification. Vision-language models (VLMs) describe a child's behavior from video well; the verdict drawn from the description is unstable: at temperature~0, across eight pipeline configurations on one backbone, 16--37\% of clips change their predicted label between repeated runs, and the cause lies in the serving stack. We keep the ...

๐Ÿ“„ AGAR: a reinforcement learning substrate for LLM program evolution
๐Ÿ—“๏ธ Published: 10/6/2026
๐Ÿ”— http://arxiv.org/abs/2610.09215v1
๐Ÿ‘ฅ Authors: Haoran Li, Zengle Ge, Xiaomin Yuan, Yui Lo, Haoxin Li, Songlin Zhou, Qianhui Liu, Jiahua Ying, Yuanhang Liu, Mingju Chen, Annan Li, Jianmin Wu, Dawei Yin (possible past Baidu (China) affiliation), Dou Shen (possible past Baidu (China) affiliation)
Abstract

Given a task and an evaluator, a language model can rewrite a candidate program while a search loop decides which rewrites survive, offering a practical route to algorithm discovery. But that loop is governed by five constants set by hand: which parent to select, how hard to mutate, how to keep diversity, what to remember, and a scalar score that never says which part of the program earned it. Reinforcement learning already has an estimator for each. The obstacle is that program evolution is not...

๐Ÿ“„ Frozen Models, Evolving Expertise: Model-Agnostic Learning from Deployment Experience for Multimodal Medical AI
๐Ÿ—“๏ธ Published: 10/6/2026
๐Ÿ”— http://arxiv.org/abs/2610.09146v1
๐Ÿ‘ฅ Authors: Yexiao He, Yucheng Tang, Pengfei Guo, Yufan He (possible past Nvidia (United States) affiliation), Andriy Myronenko (possible past Nvidia (United States) affiliation), Can Zhao, Ang Li (possible past Google (United States) affiliation), Daguang Xu (possible past Nvidia (United States) affiliation), Dong Yang (possible past Nvidia (United States) affiliation)
Abstract

Large language models (LLMs) and vision-language models (VLMs) are usually frozen after deployment, so they do not learn from the cases they solve. This is especially concerning in medicine, where new clinical evidence, updated guidelines, and new therapies can change established practice. Fine-tuning can update the model, but it requires access to model weights and additional training. Parameter-free methods avoid training, but they may overfit a fixed validation set, lack reliable domain knowl...

๐Ÿ“„ COPC: Coupled Off-Policy Correction for Asynchronous LLM Reinforcement Learning
๐Ÿ—“๏ธ Published: 10/7/2026
๐Ÿ”— http://arxiv.org/abs/2610.09597v1
๐Ÿ‘ฅ Authors: Zicheng Hu, Zhijian Zhou, Xuan Zhang (possible past Meta (United States) affiliation), Yuchen Liu, Cheng Chen (possible past Google (United States) affiliation), Yuan Li (possible past Google (United States) affiliation), Qi Gu, Yan Feng, Hongyan Hao, Chao Qu
Abstract

Asynchronous RL accelerates large language model post-training by decoupling rollout generation from optimization, but trains on stale trajectories. Existing methods primarily correct token-level policy mismatch through importance-ratio control in the actor objective. We show that this \emph{policy-side correction} alone is insufficient: advantage estimates also inherit mismatch from behavior-policy continuations, which we term \emph{advantage staleness}. We derive exact bias and variance decomp...

*Notable papers are those with at least two authors from a "big" AI/ML lab.