📄 Notable* Recent AI/ML arXiv Papers

Last updated just now...

📄 4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes
🗓️ Published: 10/2/2026
🔗 http://arxiv.org/abs/2610.03715v1
👥 Authors: Ruihong Shen, Žiga Kovačič, Peter Kulits, Xingrui Wang, Zizhang Li, Joshua B. Tenenbaum (possible past Massachusetts Institute Of Technology affiliation), Alan Yuille (possible past Google (United States) affiliation), Jieneng Chen, Jiajun Wu (possible past Massachusetts Institute Of Technology affiliation)
Abstract

We introduce 4DCodeBench, a benchmark for 4D inverse graphics through code generation, in which agents reconstruct dynamic scenes from video as executable graphics programs. To accomplish this, agents must translate visual observations into compact representations of scene structure and dynamics, by implementing abstractions such as physical simulations to reproduce complex behavior. To evaluate this capability, we curate a set of real-world videos and construct synthetic scenes spanning diverse...

📄 EyeRobot 2.0: Active Gaze for Precise Manipulation without Wrist Cameras
🗓️ Published: 10/2/2026
🔗 http://arxiv.org/abs/2610.03710v1
👥 Authors: Kush Hari, Justin Kerr, Nidhya Shivakumar, Samarth Mahapatra, Carmelo Sferrazza, Jiahui Lei, Jitendra Malik (possible past University Of California, Berkeley affiliation), C. Karen Liu (possible past Stanford University affiliation), Ken Goldberg (possible past University Of California, Berkeley affiliation), Angjoo Kanazawa (possible past University Of California, Berkeley affiliation)
Abstract

Inspired by human vision, we introduce a framework using active gaze to enable fine-grained bimanual manipulation with only a single stereo camera. EyeRobot 2.0 physically attends to a 3D fixation point in the scene by swiveling two eye viewpoints to center their gaze on it. The resulting images are processed foveally by allocating more visual tokens to the image centers, focusing computation on task-relevant features. Such Active Visual Fixation (AVF) requires carefully coordinated gaze during ...

📄 LoGo: Local-Global Rewards for Consistent Long-Horizon Video Generation
🗓️ Published: 10/2/2026
🔗 http://arxiv.org/abs/2610.03636v1
👥 Authors: Ziqi Ma, Shreya Sharma, Mohamed El Banani, Katja Schwarz, Chongjie Ye, Chao-Yuan Wu, Li Fei-Fei (possible past Stanford University affiliation), Ben Mildenhall (possible past University Of California, Berkeley affiliation), Georgia Gkioxari (possible past University Of California, Berkeley affiliation), Justin Johnson (possible past Stanford University affiliation), Gowthami Somepalli
Abstract

Camera-controlled video models are rapidly advancing toward long generation horizons and complex camera control. A key failure mode is 3D inconsistency: as the camera moves, objects lose permanence and scene structures shift. Existing post-training techniques, which assign a single scalar reward to the entire generation, are poorly suited to correcting these inconsistencies over long horizons. We introduce LoGo, which blends global and spatially localized rewards for camera-controlled video mode...

📄 MobiAgent: Dual-Loop Recursive Policy Self-Improvement for Long-Horizon Mobile Manipulation
🗓️ Published: 10/2/2026
🔗 http://arxiv.org/abs/2610.03476v1
👥 Authors: Chenzhi Liu, Yue Zhang, Jiehong Lin, Jianan Wang (possible past Deepmind (United Kingdom) affiliation), Bo Wang (possible past Tencent (China) affiliation), Zhongrui Wang, Xiaojuan Qi (possible past University Of Oxford affiliation)
Abstract

Long-horizon mobile manipulation presents significant challenges due to compounding execution errors and capacity interference between locomotion and arm control. While recent Vision-Language-Action models excel at short-horizon tasks, they lack the hierarchical reasoning required for multi-stage objectives. Furthermore, existing hierarchical agents suffer from rigid sub-task mapping, inflexible replanning, and a lack of continuous learning. To address these limitations, we introduce MobiAgent, ...

📄 Multimodal reasoning for broadly neutralizing antibody discovery from label-free human B cell repertoires across virus families
🗓️ Published: 10/2/2026
🔗 http://arxiv.org/abs/2610.03160v1
👥 Authors: Hantao Lou, Jianqing Zheng, Can Yue, Meihan Zhang, Yuanchao Bao, Yu Chen (possible past Meta (United States) affiliation), Mengting Huang, Yupeng Yang, Qianyu Pan, Nana Fu, Yansong Shi, Hongli Li, Yangyang Chai, Ruyi Chen, Wansheng Li, Zhu Liang, Rongmei Yao, Yuanhan Mo, Lei Wang (possible past Baidu (China) affiliation), Chunmei Wang, Yun Quan, Qiong Zhang, Xiangxi Wang, Xuetao Cao
Abstract

Discovering broadly neutralizing antibodies (bnAbs) from human natural immune repertoires remains a fundamental challenge in immunology, hindered by: the extreme rarity of bnAb, incomplete understanding of their cellular origins across pathogens, and the inability of existing computational tools to generalize across emerging viral threats. Here we present ImmuneAgent, a closed-loop AI system that integrates multimodal reasoning with continual meta-learning and wet-lab feedback to overcome these ...

📄 ULTRADISCOVERY: Abductive Exploration in an Interconnected, Epistemically Open Universe
🗓️ Published: 10/2/2026
🔗 http://arxiv.org/abs/2610.03092v1
👥 Authors: Weihan Li, Tianshi Zheng, Yangqiu Song (possible past Tsinghua University affiliation), Ginny Y. Wong, Simon See (possible past Nvidia (United States) affiliation)
Abstract

Scientific discovery often begins when scattered clues call for a new way of describing the world. Such abductive exploration can require constructing the representation in which an explanation is stated, when the world is epistemically open, and composing evidence scattered across contexts, when it is structurally interconnected. Existing benchmarks rarely separate these two demands or control them independently. We introduce ULTRADISCOVERY, an interactive world of five domains in which an agen...

📄 Verifiable, Articulable, and Tacit Components of Preference
🗓️ Published: 10/2/2026
🔗 http://arxiv.org/abs/2610.03025v1
👥 Authors: Alexander Spangher, Sheldon Huang, Andreas Haupt, Noah D. Goodman (possible past Stanford University affiliation), Diyi Yang (possible past Stanford University affiliation), Daniel E. Ho, Sanmi Koyejo
Abstract

What makes a short story gripping; a news article newsworthy; or a math proof elegant? These constructs resist articulation or verification; their meaning is at least partially tacit. However, modern AI models are improved primarily via articulated constitutions, rubrics and verifiers (i.e. in RLAIF and RLVR); tacit components of preferences are typically understudied. We introduce a large, labeled preference dataset CreativePreferences, containing 2.8M texts labeled by 317M human preference jud...

📄 HASTE: Evolving Agent Harnesses Against Emerging Attacks Using Sparse Evidence
🗓️ Published: 10/2/2026
🔗 http://arxiv.org/abs/2610.02920v1
👥 Authors: Xiqiao Xiong, Moxin Li, Zhixin Ma, Ouxiang Li, Wenjie Wang, Fuli Feng (possible past National University Of Singapore affiliation), Xiangnan He (possible past National University Of Singapore affiliation)
Abstract

Agent harnesses play a critical role in defenses by enforcing safety constraints to prevent unsafe actions. However, rapidly emerging attacks outpace manual harness adaptation, motivating automated harness evolution. Yet the signals available for harness evolution are often sparse, such as brief descriptions or a few attack examples in threat reports and preprints. To address this limitation, we introduce HASTE, a multi-agent framework that evolves agent harnesses from sparse threat evidence thr...

📄 AMBER: Multi-View Adaptive Budget Allocation for Listwise Vision-Language Reranking
🗓️ Published: 10/2/2026
🔗 http://arxiv.org/abs/2610.02831v1
👥 Authors: Wenteng Chen, Jiachen Zhu, Rong Shan, Tianyi Xu, Yuxiang Chen, Congmin Zheng, Teng Wang, Junjie Wu, Weiwen Liu, Changwang Zhang (possible past Tencent (China) affiliation), Weinan Zhang (possible past Shanghai Jiao Tong University affiliation), Jun Wang (possible past Tencent (China) affiliation), Jianghao Lin
Abstract

Vision-language models (VLMs) are powerful listwise rerankers for multimodal retrieval, but high inference costs restrict them to evaluating small local candidate views. Existing multi-call strategies rely on fixed schedules, wasting expensive VLM calls on uninformative candidate pairs and easy queries. To address this, we propose Adaptive Multi-view Budgeted Elo Reranking (AMBER), an online, budgeted multi-view reranking framework that dynamically optimizes global resource allocation. AMBER tre...

📄 MLCommons Jailbreak Benchmark v1.0
🗓️ Published: 10/2/2026
🔗 http://arxiv.org/abs/2610.02827v1
👥 Authors: Carsten Maple, Cagatay Yucel, Isaac Holeman, Chris Knotz, Peter Mattson (possible past Google (United States) affiliation), James Goel, Jonathan Petit, Sean Mcgregor, James Ezick, Abhishek Kumar (possible past Google (United States) affiliation), Alicia Parrish, Murali Emani, Kashyap Iyer, Faiza Khan Khattak, Washington Mbonu, Daniel Machlab, Eileen Long, Shaona Ghosh, Jibin Varghese, Roman Lutz, Andrew Gruen, Bennett Hillenbrand, Prabal Gupta, Mohammed Serrhini, Dhivya Nagasubramanian, Aakash Gupta, Jun, Lu, Kurt Bollacker, Chang Liu, Jonathan Petit, Cong Chen, Jean-Philippe Monteuuis, Brent Miller, Apurv Verma, Roman Eng, Armstrong Foundjem, Mohammed Serrhini
Abstract

Modern AI systems are designed to refuse hazardous requests. A jailbreak is a prompt crafted to bypass those safeguards and elicit outputs that the system would normally refuse to provide. The MLCommons Jailbreak Benchmark v1.0 provides an end-to-end methodology for evaluating the robustness of large language models to single-turn, text-based jailbreak attacks. It combines criteria-driven system and attack selection, paired baseline and adversarial evaluation, human annotation, automated evaluat...

📄 Improving Atomic-Fact Recall via Focused Views in Unstructured Knowledge Editing
🗓️ Published: 10/2/2026
🔗 http://arxiv.org/abs/2610.02772v1
👥 Authors: Ding Wu, Ye Zhang (possible past Google (United States) affiliation), Haoyu Wang (possible past Tencent (China) affiliation), Tianci Liu
Abstract

Large language models (LLMs) increasingly serve as general-purpose interfaces to factual knowledge, but their parameters do not automatically reflect information that changes after pretraining. Knowledge editing (KE) provides a targeted alternative to costly retraining by modifying selected knowledge and preserving unrelated knowledge and general capabilities. Conventional KE uses structured factual triples, whereas unstructured KE (UKE) uses free-form passages containing multiple facts. Nonethe...

📄 Prospective Hindsight: Self-Calibrating Reinforcement Learning via Prediction-Reality Gaps
🗓️ Published: 10/2/2026
🔗 http://arxiv.org/abs/2610.02740v1
👥 Authors: Jiaxin Zhang, Xiangyu Peng, Qinglin Chen, Yu Li (possible past Tencent (China) affiliation), Hiroaki Hayashi, Chien-Sheng Wu (possible past Salesforce (United States) affiliation)
Abstract

Reinforcement learning for long-horizon agents relies on purely retrospective training signals: credit is assigned only after observing environmental consequences, leaving the agent's belief at action time invisible to the gradient. We introduce Prospective Hindsight (PH), a self-calibrating training principle that augments any retrospective base method with a signal derived from the gap between the agent's prospective prediction (before feedback) and the retrospective evaluation (after feedback...

📄 LEAP: Learning Efficient Action Proposals For LLM Agents
🗓️ Published: 10/2/2026
🔗 http://arxiv.org/abs/2610.02670v1
👥 Authors: Zhen Xu (possible past Google (United States) affiliation), Qizheng Zhang, Gerry Wan, Shang Zhu, Ce Zhang (possible past Eth Zurich affiliation)
Abstract

LLM agents are known to be slow in rollouts. An agent completes a task one step at a time. At each step, it reasons and then chooses an action to execute. The next step and action cannot start until the previous one has finished. Speculative decoding accelerates the rollouts at the reason phase by drafting and verifying the inference tokens. Recent works have also started to apply similar ideas at the action phase. These works use off-the-shelf models, usually large, to draft action proposals fo...

📄 Large Language Continuous Diffusion Models
🗓️ Published: 10/2/2026
🔗 http://arxiv.org/abs/2610.02665v1
👥 Authors: Zhihan Yang, Wei Guo, Jean-Marie Lemercier, Simon Welker, Yonggan Fu, Mohammad Mahdi Kamani, Sajad Norouzi, Julius Berner, Tomas Geffner, Karsten Kreis (possible past Nvidia (United States) affiliation), Yongxin Chen, Molei Tao, John Thickstun, Pavlo Molchanov (possible past Nvidia (United States) affiliation), Ante Jukić, Arash Vahdat (possible past Nvidia (United States) affiliation), Morteza Mardani (possible past Nvidia (United States) affiliation)
Abstract

Despite the success of discrete diffusion language models (dLMs) for fast parallel decoding, their non-smooth, high-dimensional space hinders trajectory steering for reasoning and inference acceleration. To overcome this, we present Sigma, the first large-scale (3B/8B) continuous dLM built on steerable, low-dimensional ODE/SDE latent trajectories. Trained blockwise via likelihood optimization, Sigma jointly denoises Gaussian-corrupted token embeddings while learning an optimal embedding geometry...

📄 Generalization Properties of Score-matching Diffusion Models for Intrinsically Low-dimensional Data
🗓️ Published: 10/2/2026
🔗 http://arxiv.org/abs/2610.02663v1
👥 Authors: Saptarshi Chakraborty, Quentin Berthet (possible past University Of Cambridge affiliation), Peter L. Bartlett (possible past University Of California, Berkeley affiliation)
Abstract

Despite the remarkable empirical success of flow-matching models, their statistical generalization guarantees remain underdeveloped. Existing analyses often impose restrictive assumptions on the estimated velocity field and yield convergence rates that fail to reflect the intrinsic low-dimensional structure common in real data, such as natural images and molecular geometries. In this work, we study the statistical generalization of flow-matching models for learning an unknown distribution $P_{\m...

📄 Mitigating Social Sycophancy via Pluralistic Preference Optimization
🗓️ Published: 10/1/2026
🔗 http://arxiv.org/abs/2610.02568v1
👥 Authors: Stephane Hatgis-Kessell, Myra Cheng (possible past Deepmind (United Kingdom) affiliation), Xiaoxuan Hou, Qian Hu, Rahul Gupta (possible past Google (United States) affiliation), Natasha Jaques (possible past University Of California, Berkeley affiliation), Emma Brunskill
Abstract

Personal advice, including relationship advice, now ranks among the most common uses of generative AI. But language models (LMs) exhibit sycophancy: they affirm users much more often than humans do, which can make people overconfident and less willing to repair their relationships after a conflict. Prior work on mitigating sycophancy has focused on factual settings where a response can be checked against a ground truth answer, while mitigations for social sycophancy (e.g., personal advice, where...

📄 Normal-Form Correlation in Markov Games
🗓️ Published: 10/2/2026
🔗 http://arxiv.org/abs/2610.03621v1
👥 Authors: Ioannis Anagnostides, Constantinos Daskalakis (possible past University Of California, Berkeley affiliation), Gabriele Farina, Noah Golowich, Tuomas Sandholm, Brian Hu Zhang (possible past Stanford University affiliation)
Abstract

There has been a surge of recent work on correlated equilibrium concepts in Markov games. However, existing results focus on concepts weaker than normal-form correlated equilibria (NFCEs), leaving open the more challenging question of computing such equilibria, which goes back to the seminal work of Papadimitriou and Roughgarden (JACM'08). Here, we establish the first efficient algorithm for NFCEs in finite-horizon Markov games with a fixed number of players $n$. In particular, with $S$ states, ...

📄 HyperThink: Text-to-Parameter Hypernetworks for Efficient Reasoning
🗓️ Published: 10/2/2026
🔗 http://arxiv.org/abs/2610.03039v1
👥 Authors: Donggyun Kim, Jack Lu, Chanwoo Kim (possible past Google (United States) affiliation), Mengye Ren (possible past University Of Toronto affiliation), Seunghoon Hong
Abstract

Long-form thinking traces can substantially improve the multi-step reasoning performance of large language models (LLMs), but they introduce high inference-time overhead, with latency dominated by sequential decoding. We propose HyperThink, a text-to-parameter approach that amortizes this reasoning computation into a single query-conditioned parameter update: a lightweight hypernetwork reads the question and predicts updates to a small subset of the base LLM's parameters, while a vector-quantize...

📄 Learning from Evolving Errors: Adaptive Iterative Repair for On-Policy Distillation
🗓️ Published: 10/2/2026
🔗 http://arxiv.org/abs/2610.02700v1
👥 Authors: Rui Li (possible past Google (United States) affiliation), Liyang He, Zheng Zhang, Zhenya Huang, Linbo Zhu, Qi Liu (possible past Tencent (China) affiliation)
Abstract

On-policy self-distillation (OPSD) supplies dense token-level feedback on trajectories sampled from the student's own policy, a richer training signal than the outcome-level rewards of reinforcement learning. This feedback comes from a teacher conditioned on a full reference solution unavailable to the student. The reference solution specifies the target but not how to move from the student's current error toward it, creating a solution-conditioned shortcut risk. We introduce AIR-OPD, an adaptiv...

*Notable papers are those with at least two authors from a "big" AI/ML lab.