๐Ÿ“„ Notable* Recent AI/ML arXiv Papers

Last updated just now...

๐Ÿ“„ Rethinking On-Policy Distillation of Large Language Models II: One Training Example
๐Ÿ—“๏ธ Published: 9/3/2026
๐Ÿ”— http://arxiv.org/abs/2609.04172v1
๐Ÿ‘ฅ Authors: Zixuan Fu, Bingxiang He, Yuxin Zuo, Haohuan Huang, Jinqian Zhang, Ruhang Xiao, Cheng Qian, Qinyu Luo, Huan-Ang Gao, Yudong Wang, Zhiyuan Liu (possible past Tsinghua University affiliation), Ning Ding (possible past Tsinghua University affiliation), Chaojun Xiao
Abstract

On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. Existing work has mainly studied its algorithmic behavior, leaving the role of training data unclear. We examine this role at the data-minimal limit by training on a single query. One-shot OPD keeps improving for hundreds of steps and recovers most of full-data OPD's gain across task domains and model families. We explain this result through the states visited during training and t...

๐Ÿ“„ A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms
๐Ÿ—“๏ธ Published: 9/3/2026
๐Ÿ”— http://arxiv.org/abs/2609.04170v1
๐Ÿ‘ฅ Authors: Davide Paglieri, Logan Cross, Tim Genewein, Joel Z. Leibo (possible past Google (United States) affiliation), Nenad Tomasev, Alexander Sasha Vezhnevets (possible past Google (United States) affiliation)
Abstract

Multi-agent AI science ecosystems rely on agents possessing tools that allow them to communicate, coordinate, and build on each other's work. Yet this shared infrastructure can also introduce vulnerabilities by creating a substrate for the contagious spread of unintended and undesirable behaviors. We report a case study on a research collective of 100 autonomous LLM agents tasked with proving formal mathematical conjectures. Within the swarm, cheating spontaneously emerged and was later challeng...

๐Ÿ“„ Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments
๐Ÿ—“๏ธ Published: 9/3/2026
๐Ÿ”— http://arxiv.org/abs/2609.04148v1
๐Ÿ‘ฅ Authors: Jie Wu, Zhenru Zhang, Beichen Zhang, Xuwu Wang, Yuhui Su, Mouxiang Chen, Peng Wang (possible past Peking University affiliation), Zhihai Wang, Que Shen, Hao Zhou, An Yang, Fei Huang, Yujiu Yang (possible past Tsinghua University affiliation), Dayiheng Liu
Abstract

As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while realistic, executable environments remain scarce. However, environments are what agent post-training actually requires: each can be re-queried into many verifiable tasks and provides execution feedback, whereas a trajectory is a single frozen demonstration. Rather than generating environments from scratch, we observe that the tool-execution history in existing trajectories exposes the structure an...

๐Ÿ“„ Efficient Test-Time Adaptation through Human-AI Interaction
๐Ÿ—“๏ธ Published: 9/3/2026
๐Ÿ”— http://arxiv.org/abs/2609.04141v1
๐Ÿ‘ฅ Authors: Zora Zhiruo Wang, Apurva Gandhi, Rulin Shao, Aspen Chen, Jonas Mueller, Zhiqi Liang, Jett Chen, Michael Ryan (possible past Google (United States) affiliation), Qianou Ma, Luxi He, Zhoujun Cheng, Andre He, Seungone Kim, Jiayi Geng, Mingqian Zheng, Weiwei Sun, Zheyuan Zhang, Xinran Zhao, Yike Wang, Abe Hou, Liwei Jiang, Pang Wei Koh, Diyi Yang (possible past Stanford University affiliation), Graham Neubig (possible past Carnegie Mellon University affiliation), Daniel Fried
Abstract

AI agents are trained on population-scale data to encode broad capabilities spanning those of many practitioners. Yet the artifacts they produce rarely meet the personal bar professionals need to stake their reputation on. On realistic, open-ended tasks where success criteria are heterogeneous and insufficiently documented, individual expertise lives precisely in the elevation and departure from the average. In practice, iterative human-agent interaction surfaces criteria that users cannot fully...

๐Ÿ“„ Xiaomi-TabLDM: A Tabular Foundation Model Technical Report
๐Ÿ—“๏ธ Published: 9/3/2026
๐Ÿ”— http://arxiv.org/abs/2609.03880v1
๐Ÿ‘ฅ Authors: Xiaomi-Tabldm Team, :, Penghui Wang, Wei Liu (possible past Tsinghua University affiliation), Hong Wang, Chengyue Huang, Yuxi Sun, Zirui Wang, Hongming Huang, Quan Wang (possible past Google (United States) affiliation), Chunxiao Liu, Erli Meng, Bin Wang
Abstract

We introduce Xiaomi-TabLDM, a tabular large data foundation model for classification and regression via in-context learning, which delivers superior prediction accuracy without requiring task-specific fine-tuning. Pretrained exclusively on synthetic data generated from structural causal models (SCMs), our model enables more flexible context utilization and more efficient capacity scaling. i) A new performance standard. Strong regression performance across benchmarks: Xiaomi-TabLDM ranks 1st on...

๐Ÿ“„ Bioinfoysis Technical Report
๐Ÿ—“๏ธ Published: 9/3/2026
๐Ÿ”— http://arxiv.org/abs/2609.03871v1
๐Ÿ‘ฅ Authors: Qingyang Shao, Xin Zhang (possible past Google (United States) affiliation), Zhouyang Yuan, Xianying Chen, Yujia Xiang, Zihao Yang, Tong Ye, Yangqi Zhang, Jiakang Xu, Xiaoqing Yan, Xuan Luo (possible past University Of Washington affiliation), Keyi Li, Enci Fan, Kai Kang, Zhuohan Liu, Xingyu Jin, Chunran Teng, Tao Li (possible past Baidu (China) affiliation), Xinyu Lv, Minghui Wang, Wenfeng Li, Yidan Gao, Siyu Liu, Mingrui Luo, Zhu Liang, Guanren Qiao, Zhiping Xu
Abstract

Large language model agents have shown promise in bioinformatics, but most existing systems focus primarily on producing final answers, treating planning, tool use, and code execution as transient interactions. This design is poorly suited to long-horizon bioinformatics tasks, where conclusions must remain connected to the data, computations, and intermediate evidence that support them. We introduce \textbf{Bioinfoysis}, a multi-agent harness that represents each request as a persistent, artifac...

๐Ÿ“„ GazeFS: Target-Centered Gaze-Trajectory Forecasting and Stabilization from Gaze-Head History
๐Ÿ—“๏ธ Published: 9/3/2026
๐Ÿ”— http://arxiv.org/abs/2609.03868v1
๐Ÿ‘ฅ Authors: Yaozheng Xia, Zaiping Zhu, Bo Pang, Minghao Xie, Hui Li (possible past Baidu (China) affiliation), Shaorong Wang, Sheng Li (possible past Google (United States) affiliation)
Abstract

Target-centered gaze interaction requires more than suppressing frame-to-frame fluctuations: target acquisition produces task-aligned changes in gaze-head dynamics, while a gaze trace may retain a persistent target-relative residual direction. We formulate gaze correction as online target-centered gaze-trajectory forecasting and stabilization and introduce GazeFS, which maps a variable-length gaze-head history to the next target-center direction and a short-horizon Search/Focus estimate without ...

๐Ÿ“„ LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes
๐Ÿ—“๏ธ Published: 9/3/2026
๐Ÿ”— http://arxiv.org/abs/2609.03796v1
๐Ÿ‘ฅ Authors: Chuyan Chen, Haoxing Chen, Kun Chen, Zhenglin Cheng, Long Cui, Ruishan Fang, Zhangxuan Gu, Zhicheng Huang, Zhenzhong Lan (possible past Google (United States) affiliation), Yuanting Lei, Haoquan Li, Jianguo Li, Rongchuan Li, Sidu Li, Tao Lin, Deyuan Liu, Jiacheng Liu, Lin Liu, Yuxuan Lou, Zhisheng Lu, Yuxin Ma, Shuheng Shen, Peng Sun (possible past Tencent (China) affiliation), Chaoyang Wang, Hongjun Wang, Xiaomei Wang, Yongxin Wang, Chengzhang Wu, Hongru Wu, Jun Xie (possible past Tencent (China) affiliation)
Abstract

We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch with a frozen vision-language understanding module built on the LLaDA2.0-Mini diffusion language model backbone. Instead of relying heavily on paired image-text data from the beginning, we first build a strong visual generative prior through image-only pre-training and mid-training. The generation pipeline comprises 220M samples, 98 of which are real images. For efficient and scalable o...

๐Ÿ“„ DNative-Twin: Decision Graphs and Digital Twins for Reconstructable Agentic Decisions
๐Ÿ—“๏ธ Published: 9/3/2026
๐Ÿ”— http://arxiv.org/abs/2609.03787v1
๐Ÿ‘ฅ Authors: Junjie Pang, Zhenzhen Xie, Haoke Han, Ying He, Jing Wang (possible past Google (United States) affiliation), Gang Liu (possible past Tencent (China) affiliation)
Abstract

AI agents increasingly gather evidence, invoke tools, apply constraints, and produce decisions that people or software may commit to action. A final output alone cannot show which evidence, tool state, rule, authorization, or action path produced it. We present DNative-Twin, a graph-native digital twin that records a committed agentic decision as a typed trajectory and re-executes its decision mechanism under declared conditions. The graph links the state observed by the agent, the path it follo...

๐Ÿ“„ CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning
๐Ÿ—“๏ธ Published: 9/3/2026
๐Ÿ”— http://arxiv.org/abs/2609.03526v1
๐Ÿ‘ฅ Authors: Bo Zeng, Linfeng Gao, Peiqin Lin, Yu Zhao (possible past Tencent (China) affiliation), Mingyan Zeng, Yu Tong (possible past University Of California, Berkeley affiliation), Xintong Wang, Linlong Xu, Longyue Wang (possible past Tencent (China) affiliation), Weihua Luo, Qinggang Zhang, Jinsong Su
Abstract

Multimodal language models achieve near-ceiling scores on food recognition benchmarks, yet it remains unclear whether this success reflects genuine cultural understanding or mere visual matching. To probe this distinction, we introduce CulturalMenuBench, a benchmark of 4,870 items in 10 languages across 18 regions; its 10 tasks pair final-dish and step-by-step cooking images with ingredients, procedural text, and regional labels, spanning basic recognition to process-grounded cultural attributio...

๐Ÿ“„ BRIDGE: An Open-Source Humanoid Platform via Morphology-Control Co-Design for Physical AI
๐Ÿ—“๏ธ Published: 9/3/2026
๐Ÿ”— http://arxiv.org/abs/2609.03497v1
๐Ÿ‘ฅ Authors: Jianren Wang, Letian Qian, Zikai Wang, Weiwei Wu, Junjie Zong, Abhinav Gupta (possible past Google (United States) affiliation), Deepak Pathak (possible past University Of California, Berkeley affiliation)
Abstract

Developing humanoid robots capable of leveraging human behavioral data is essential for general-purpose embodiment, yet conventional development remains bottlenecked by a decoupled paradigm that isolates hardware design from whole-body control. This approach leads to suboptimal systems that compromise human-like fluidity and agility. To bridge this gap, we introduce a data-driven morphology-control co-design framework that optimizes humanoid morphology for human-like movement. To quantify morpho...

๐Ÿ“„ Post-Training Language Models for Gold-Medal Performance in Coding Competitions
๐Ÿ—“๏ธ Published: 9/2/2026
๐Ÿ”— http://arxiv.org/abs/2609.02849v1
๐Ÿ‘ฅ Authors: Aleksander Ficek, Sean Narenthiran, Mehrzad Samadi (possible past Nvidia (United States) affiliation), Somshubra Majumdar (possible past Nvidia (United States) affiliation), Boris Ginsburg (possible past Nvidia (United States) affiliation)
Abstract

Competitive programming has become a key test of large language model reasoning, with international competitions such as IOI and ICPC representing its most challenging settings. We present an end-to-end specialization pipeline combining large-scale problem curation, synthetic reasoning traces, supervised fine-tuning (SFT), and reinforcement learning (RL). Using 22,000 curated problems, we train Nemotron-3-Nano-CC (30B-A3B) with SFT and RL and Nemotron-3-Ultra-CC (550B-A55B) with SFT alone. We fu...

๐Ÿ“„ Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems
๐Ÿ—“๏ธ Published: 9/2/2026
๐Ÿ”— http://arxiv.org/abs/2609.02750v1
๐Ÿ‘ฅ Authors: Yihang Chen, Yuxiang Chen, Yuxuan Huang, Meng Fang (possible past Tencent (China) affiliation), Weilin Luo, Jun Wang (possible past Tencent (China) affiliation)
Abstract

Multi-agent LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improve through textual reflection. Despite strong empirical results, these systems lack a unified account of coordination, memory improvement, and the role of external verification. We model orchestrator-worker interaction as a bilevel coordination game: under bounded coupling, the workers' local-update game is an approximate potential game whose equilibrium slack is controlled by decompositi...

๐Ÿ“„ Structure and Implementation of New Practical English Textbooks Driven by Artificial Intelligence
๐Ÿ—“๏ธ Published: 9/2/2026
๐Ÿ”— http://arxiv.org/abs/2609.02981v1
๐Ÿ‘ฅ Authors: Ya Wang (possible past Peking University affiliation), Lei Zhang, Xueguang Yang, Bo Chen (possible past Tencent (China) affiliation)
Abstract

Artificial intelligence is changing the form of applied English materials from fixed paper sequences to adaptive learning systems that can diagnose learners, recommend tasks, and provide formative feedback. This paper studies the structure and application of a new practical English textbook driven by artificial intelligence. A five-layer architecture is proposed: knowledge mapping, learner profiling, task generation, feedback orchestration, and teacher-side governance. A prototype was tested on ...

๐Ÿ“„ ProbeMatchDTI: Probe-Driven Multi-Scale Biochemical Pattern Matching for Drug-Target Interaction Prediction
๐Ÿ—“๏ธ Published: 9/2/2026
๐Ÿ”— http://arxiv.org/abs/2609.02549v1
๐Ÿ‘ฅ Authors: Quan Hao, Mengyue Fan, Zifan Dong, Youru Li, Jianduo Zhao, Lechuan Xu, Hao Zhang (possible past Tencent (China) affiliation), Fei Xia (possible past Stanford University affiliation), Jigang Wang, Chong Qiu, Liguo Zhang
Abstract

Drug-target interaction (DTI) prediction is an important task in AI-driven drug discovery. Although recent biochemical representation learning methods have improved DTI prediction, their passive feature aggregation tends to favor dominant molecular patterns while suppressing weak yet binding-relevant signals, such as functional groups and residue-context patterns, limiting the modeling of multi-scale biochemical correspondences. To address this issue, we propose ProbeMatchDTI, a pattern-probe-dr...

๐Ÿ“„ Fine-Grained Anomaly Perception in Wild UGC-Enhanced Images: A Comprehensive Dataset and Difference-Fusion Framework
๐Ÿ—“๏ธ Published: 9/2/2026
๐Ÿ”— http://arxiv.org/abs/2609.02529v1
๐Ÿ‘ฅ Authors: Yan Zhong, Gefei Chen, Qiufang Ma, Zhen Wang, Zhiwei Fan, Lei Shi (possible past Baidu (China) affiliation), Tingting Jiang (possible past Tsinghua University affiliation)
Abstract

Image enhancement and restoration have become standard back-end operations on short-video and social media platforms to boost UGC visual experience. Yet these processes inevitably introduce visual anomalies--especially in faces, texts, and textures--that directly undermine perceptual fidelity and viewer trust. While existing IQA methods perform well on classic distortions, they target holistic quality assessment and fail to capture the specific, localized anomalies caused by enhancement algorith...

๐Ÿ“„ Before the Script, Set the Stage: How Worldview Simulation Amplifies Psychologically Grounded Persuasion in Multi-Turn Jailbreaking
๐Ÿ—“๏ธ Published: 9/2/2026
๐Ÿ”— http://arxiv.org/abs/2609.02414v1
๐Ÿ‘ฅ Authors: Siyu Chen, Haoran Wang, Xiaojian Li, Yao Huang, Yinpeng Dong (possible past Tsinghua University affiliation), Wei Xu (possible past Tencent (China) affiliation)
Abstract

Multi-turn jailbreak attacks demonstrate that harmful intent can be distributed across dialogue, yet existing methods obscure what conversational mechanisms drive vulnerability. We introduce BLUEPRINT, a safety-evaluation framework separating a factorized social-influence strategy space from WORLDVIEWSIM, a cross-turn situational context module. Monte Carlo Tree Search optimizes turn-level combinations of 18 theory-grounded influence factors across a four-turn trajectory. Across six frontier mod...

๐Ÿ“„ Two-Stage Reinforcement Learning for Sound and Adversarial Test Generation in Code LLMs
๐Ÿ—“๏ธ Published: 9/3/2026
๐Ÿ”— http://arxiv.org/abs/2609.03955v1
๐Ÿ‘ฅ Authors: Jiacheng Xu, Wentao Zhang (possible past Mila - Quebec Artificial Intelligence Institute affiliation), Zhiyi Lyu, Fuxiang Zhang, Chaojie Wang, Yang Liu (possible past Tsinghua University affiliation), Bo An
Abstract

Reinforcement learning (RL) has substantially advanced code generation with large language models (LLMs) through executable feedback. The feedback for coding problems mainly comes from specific test cases, where high-quality test cases are often scarce since they should be both sound and discriminative. We thus turn to study the auto-generation of test cases using the learned model. We find this is naturally an adversarial RL problem: the model is expected to generate effective test cases as cou...

๐Ÿ“„ WeatherNext 3: Increasing resolution and performance of global weather models with raw observations
๐Ÿ—“๏ธ Published: 9/3/2026
๐Ÿ”— http://arxiv.org/abs/2609.03582v1
๐Ÿ‘ฅ Authors: Stephan Rasp (possible past Google (United States) affiliation), Boris Babenko (possible past Google (United States) affiliation), Dominic Masters, Andrew El-Kadi, Samier Merchant, Guy Shalev (possible past Google (United States) affiliation), Ilan Price, Fred Zyda, Remi Lam, Sasha Shysheya, Matthew Willson (possible past Deepmind (United Kingdom) affiliation), Stratis Markou, Shreya Agrawal (possible past Google (United States) affiliation), Suhani Vora (possible past Google (United States) affiliation), Mohammed Alewi Hassen, Sunny Mak, Tom R. Andersson, Megan Bela, Akib Uddin, Nofar Peled Levi, Ben Gaiarin, Ferran Alet (possible past Deepmind (United Kingdom) affiliation), Aaron Bell, Peter Battaglia (possible past Massachusetts Institute Of Technology affiliation), Alvaro Sanchez-Gonzalez
Abstract

State-of-the-art AI weather models have shown impressive medium-range forecast skill and computational efficiency, but suffer two key shortcomings: their forecasts have lower spatial and temporal resolution than the best physics-based models and they are exclusively initialized with and trained on analysis data. As a result, they cannot directly make use of observations, and any biases in the analysis are inherited by the forecast. WeatherNext 3 addresses these shortcomings and establishes a new...

๐Ÿ“„ SimpleDesign: A Joint Model for Protein Sequence and Structure Codesign
๐Ÿ—“๏ธ Published: 9/3/2026
๐Ÿ”— http://arxiv.org/abs/2609.03377v1
๐Ÿ‘ฅ Authors: Jiarui Lu, Yuyang Wang, Yizhe Zhang, Jiatao Gu (possible past Meta (United States) affiliation), Navdeep Jaitly (possible past University Of Toronto affiliation), Joshua M. Susskind, Miguel รngel Bautista
Abstract

Proteins are fundamental to biological processes, with their function determined by the complex interplay between the amino acid sequence and the three-dimensional structure. Developing generative models capable of understanding this intrinsically multi-modal relationship is crucial for fields like drug discovery and protein engineering. Existing models often rely on a multi-stage training process where autoencoders that tokenize data into latent representations are trained in a first stage. Sec...

๐Ÿ“„ Kernel Reboot: Breaking the Boundaries of Neural Tangent Kernels for Neural Fields
๐Ÿ—“๏ธ Published: 9/2/2026
๐Ÿ”— http://arxiv.org/abs/2609.03117v1
๐Ÿ‘ฅ Authors: Amir Mallak, Alaa Maalouf, Lior Wolf, Daniela Rus (possible past Massachusetts Institute Of Technology affiliation), Dan Rosenbaum (possible past Deepmind (United Kingdom) affiliation)
Abstract

Neural fields (NFs) map continuous coordinates to signals such as color or density, but fast high-quality reconstruction from sparse observations remains difficult. Classical Neural Tangent Kernel (NTK) regression gives closed-form fits, yet it is fundamentally linear and cannot accumulate reusable task priors. We develop three algorithms that address these gaps. NTK-KIP learns a distilled support set of coordinates (and optional labels) so that a finite NTK can inpaint large missing regions fro...

๐Ÿ“„ Tail-Likelihood Reinforcement Learning
๐Ÿ—“๏ธ Published: 9/2/2026
๐Ÿ”— http://arxiv.org/abs/2609.02987v1
๐Ÿ‘ฅ Authors: Shrinivas Ramasubramanian, Daman Arora, Fahim Tajwar, Guanning Zeng, Qingyang Wu, Zhongzhu Zhou, Chenfeng Xu (possible past University Of California, Berkeley affiliation), Haiwen Feng, Yuda Song, Aarti Singh (possible past Carnegie Mellon University affiliation), Ruslan Salakhutdinov (possible past University Of Toronto affiliation), J. Andrew Bagnell (possible past Carnegie Mellon University affiliation), Jeff Schneider, Andrea Zanette
Abstract

Reinforcement learning typically optimizes average reward. For generative policies, the average can hide an important distinction: two policies can achieve the same mean reward while having very different chances of producing a rare but high-reward rollout. This matters as sampling increases during training and inference, since its benefit depends on retaining probability mass on high-reward outcomes. We propose to optimize this coverage directly. Rather than considering only expected reward, we...

๐Ÿ“„ oHC: Orthogonal Hyper-Connections on SO(4) via Quaternions
๐Ÿ—“๏ธ Published: 9/2/2026
๐Ÿ”— http://arxiv.org/abs/2609.02672v1
๐Ÿ‘ฅ Authors: Haoqiang Guo, Xuyi Chen, Bo Ke, Yishu Lei, Ziyang Xu, Shikun Feng (possible past Baidu (China) affiliation), Ximen, Wenhan Luo (possible past Tencent (China) affiliation)
Abstract

Hyper-Connections (HC) replace the single residual stream of a Transformer with $n$ parallel ones, mixing them at every layer with a learned $n \times n$ residual matrix. Leaving that matrix unconstrained places no limit on the factor by which the mixing step rescales the residual streams, and that factor compounds across layers, which destabilizes training. Manifold-constrained Hyper-Connections (mHC) address this by restricting the matrix to the doubly stochastic matrices. That caps the factor...

๐Ÿ“„ Equation Recast for Canonical Operator Learning Across Parametric PDEs
๐Ÿ—“๏ธ Published: 9/2/2026
๐Ÿ”— http://arxiv.org/abs/2609.02982v1
๐Ÿ‘ฅ Authors: Qiyun Cheng, Valentin Duruisseaux, Cesar F. Clauser, Md Hossain Sahadath, Huihua Yang, Shaowu Pan, Nathaniel Ferraro, Anima Anandkumar (possible past Nvidia (United States) affiliation), Wei Ji (possible past Tencent (China) affiliation), Cristina Rea
Abstract

Learning solution operators across broad parameter ranges can require substantial coverage of both input functions and physical parameters, particularly for purely data-driven parametric models. In addition, the resulting models may fail silently outside the training distribution. We introduce equation recast, which reformulates parametric operator learning as the learning of a single canonical operator. Parameter-induced operator variations are derived analytically from the governing equation a...

*Notable papers are those with at least two authors from a "big" AI/ML lab.