📄 Notable* Recent AI/ML arXiv Papers

Last updated just now...

📄 CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents
🗓️ Published: 7/21/2026
🔗 http://arxiv.org/abs/2607.19338v1
👥 Authors: Qijia He, Jiayi Cheng, Chenqian Le, Rui Wang (possible past Tencent (China) affiliation), Xunmei Liu, Yixian Chen, Jie Mei, Zhihao Wang, Xupeng Chen, Yuhuan Chen, Tao Wang (possible past Stanford University affiliation)
Abstract

Coding agents increasingly operate in executable environments where a failed attempt produces actionable feedback rather than merely an incorrect answer. Existing cost-aware systems typically treat such failures as cascade decisions: try a cheap model first, then escalate hard cases to a stronger and more expensive model. In coding, however, execution feedback can also make further cheap-model recovery worthwhile, raising a budgeted deployment question: when should an agent spend more cheap comp...

📄 Provable diffusion-based posterior sampling for linear inverse problems via DDIM
🗓️ Published: 7/21/2026
🔗 http://arxiv.org/abs/2607.19333v1
👥 Authors: Yuchen Jiao, Na Li (possible past Tencent (China) affiliation), Changxiao Cai, Yuxin Chen, Gen Li (possible past University Of Edinburgh affiliation)
Abstract

Diffusion-based methods have achieved remarkable empirical success in solving inverse problems. However, many existing posterior samplers either lack rigorous theoretical guarantees or incur substantial computational overhead. We propose a simple and efficient algorithm, called \pddim, for solving linear inverse problems with diffusion priors via a DDIM-type sampler. Our method requires only lightweight, coordinate-wise modifications to the standard DDIM update, while explicitly incorporating th...

📄 ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
🗓️ Published: 7/21/2026
🔗 http://arxiv.org/abs/2607.19191v1
👥 Authors: Fan Jiang (possible past Shanghai Jiao Tong University affiliation), Zhaoxu Sun, Mengchao Wang, Ziyu Zhu, Chiyu Wang, Yunpeng Zhang, Wenlin Liu, Yun Wang, Xue Zheng, Rui Sun, Junfeng Ni, Hongyu Pan, Zhongxu Sun, Fei Yu, Zengye Ge, Mengmeng Du, Nianfei Fan, Mingchao Sun, Yu Liu, Yongchang, Yanqing Zhu, Jiahang Wang, Ning Ying, Yuze Xuan, Di Yang, Zhicheng Liu, Zhe Gao, Tingbing Xu, Jiacheng Sui, Wenjin Yang, Junnan Lai, Shufeng Liu, Yuan Liu (possible past Google (United States) affiliation), Zheng Zhou (possible past Tencent (China) affiliation), Yingliang Peng, Dawei Cao, Kaifeng Sheng, Yuxiang Cai, Fei Lu, Mu Xu, Ning Guo
Abstract

We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-source data infrastructure spanning AAA games, simulation engines, and internet videos to learn controllable world dynamics. WorldExplorer performs agent-driven collection guided by training feedback, while a unified pipeline applies 14 deterministic quality checks, VLM-based assessment, and synchronized action and text annotation. We progressively distill a ...

📄 DAIS: Dependency-Aware Intermediate QA Supervision for Complex Reasoning
🗓️ Published: 7/21/2026
🔗 http://arxiv.org/abs/2607.19088v1
👥 Authors: Yu Wang (possible past Tsinghua University affiliation), Ming Fan, Xicheng Zhang, Zhiyong Li, Zhihu Wang, Caiyue Xu, Dahai Hu, Ting Liu (possible past Google (United States) affiliation)
Abstract

Chain-of-thought (CoT) supervision exposes intermediate rationales, but flat rationale targets usually optimize a single reasoning sequence and provide limited supervision on how local conclusions should support later decisions. We introduce Dependency-Aware Intermediate QA Supervision (DAIS), a training-time framework that converts filtered teacher rationales into stage-level QA records. Each intermediate record predicts a local answer conditioned on the previous states needed for that decision...

📄 Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
🗓️ Published: 7/21/2026
🔗 http://arxiv.org/abs/2607.19064v1
👥 Authors: Xinjie Zhang, Peng Zhang, Shicheng Zheng, Jinghao Guo, Zhaoyang Jia, Yifei Shen, Xun Guo, Yuxuan Luo (possible past Baidu (China) affiliation), Jiahao Li, Wenxuan Xie, Fanyi Pu, Xiaoyi Zhang (possible past Apple (United States) affiliation), Kaichen Zhang, Zongyu Guo, Tianci Bi, Dongnan Gui, Zhening Liu, Zimo Wen, Zihan Zheng, Senqiao Yang, Xiao Li, Jinglu Wang, Bin Li, Yan Lu
Abstract

Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. The stack is built from two co-designed components: Mage-VAE, a lightweight high-fidelity latent tokenizer, and a Native-Resolution Multimodal Diffusion Transformer trained with rectified flow matching. Mage-VAE uses one-step diffusion-style encoding and decoding w...

📄 FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling
🗓️ Published: 7/21/2026
🔗 http://arxiv.org/abs/2607.19038v1
👥 Authors: Jialong Zuo, Haotong Zuo, Shiwei Zhang, Xiang Wang (possible past Tencent (China) affiliation), Chen Li (possible past Tencent (China) affiliation), Nong Sang, Changxin Gao, Xiang Bai
Abstract

Translating novels into films poses a grand challenge for generative artificial intelligence, requiring conversion of abstract literary prose into long-form, multi-scene visual narratives. While current video generation models excel at short, single-scene clips within narrow temporal and spatial contexts, novel-to-film generation operates in a more complex regime, demanding long-duration content across diverse scenes with dynamically evolving entity states. To address this, we formalize novel-to...

📄 AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents
🗓️ Published: 7/21/2026
🔗 http://arxiv.org/abs/2607.18754v1
👥 Authors: Kunlun Zhu, Xuyan Ye, Zhiguang Han, Yuchen Zhao, Bingxuan Li, Weijia Zhang, Muxin Tian, Xiangru Tang (possible past University Of Cambridge affiliation), Pan Lu (possible past Baidu (China) affiliation), James Zou, Jiaxuan You, Heng Ji
Abstract

LLM agent failures are difficult to debug because the step where an error surfaces is often not the one that caused it. Existing observability tools replay execution traces but provide little support for identifying the root cause or translating diagnosis into recovery. We present AgentDebugX, an open-source debugging framework that organizes debugging as a closed loop of Detect, Attribute, Recover, and Rerun. At its core, DeepDebug performs multi-turn root-cause diagnosis through global traject...

📄 Attributes Should Come from Images, Not Class Names: Distribution-Conditioned Attribute Selection for Vision-Language Models
🗓️ Published: 7/21/2026
🔗 http://arxiv.org/abs/2607.18695v1
👥 Authors: Gautam Rajendrakumar Gare, Jia Shi, Zhiqiu Lin, Deepak Pathak (possible past University Of California, Berkeley affiliation), John Galeotti, Deva Ramanan (possible past Carnegie Mellon University affiliation)
Abstract

A popular route to interpretable zero-shot classification asks a large language model (LLM) to describe each class name and prompts CLIP with the resulting descriptors. We show that these descriptors carry little visual evidence of their own: removing the class name from the prompt collapses ImageNet accuracy from 59.5% to 15.5%. The diagnosis is that the descriptors are conditioned on the label rather than on the images, so they describe the concept in general and mislead exactly when the data ...

📄 Simple Domain Generalization for Strong Pixel-Level Image Tampering Detection in Modern VLMs
🗓️ Published: 7/20/2026
🔗 http://arxiv.org/abs/2607.18230v1
👥 Authors: Yi Tang, Xinyi Shang, Jiacheng Cui, Sondos Mahmoud Bsharat, Jiacheng Liu, Xiaohan Zhao, Tran Dinh Tien, Ahmed Elhagry, Salwa K. Al Khatib, Tianjun Yao, Yonina C. Eldar, Jing-Hao Xue, Hao Li (possible past Tsinghua University affiliation), Salman Khan (possible past Inception Institute Of Artificial Intelligence affiliation), Zhiqiang Shen
Abstract

Modern vision-language models (VLMs) have significantly improved image generation and editing capabilities, making pixel-level image tampering detection increasingly important yet challenging under cross-model and out-of-distribution shifts. This work studies domain generalization for pixel-level image tampering detection in modern VLMs like ChatGPT, Gemini, Qwen-Image, etc., aiming to learn tampering localization models that remain robust across diverse VLM-generated manipulation distributions....

📄 GigaPath-Flash and GigaTIME-Flash: Efficient Pathology Foundation Models for Whole-Slide and Tumor Microenvironment Analysis
🗓️ Published: 7/20/2026
🔗 http://arxiv.org/abs/2607.18218v1
👥 Authors: Naoto Usuyama, Jeya Maria Jose Valanarasu, Sicong Yao, Hanwen Xu, Jaspreet Bagga, Guanghui Qin, Robert E. Kramer, Cliff Wong, Soohee Lee, Hao Qiu, Theodore Zhengde Zhao, Racheli Ben Shimol, Angela Crabtree, Kevin Matlock, Eduardo Alejandro Lozano Garcia, Naiteek Sangani, Alberto Santamaria-Pang, Jason Entenmann, Alexandra Q. Bartlett, Bill J. Wright, Bernard A. Fox, Brian Piening, Sheng Zhang, Sheng Wang (possible past Tencent (China) affiliation), Tristan Naumann, Carlo Bifulco, Hoifung Poon (possible past Microsoft (United States) affiliation)
Abstract

Foundation models have emerged as a driving force in computational pathology, with the potential to transform cancer diagnosis, prognosis, and treatment selection by learning transferable representations from large-scale histopathology data. A growing landscape of pathology foundation models now spans diverse data sources, architectures, and downstream applications. However, most pretrained models operate only at the image-tile level, use restrictive licenses, and remain computationally expensiv...

📄 AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report
🗓️ Published: 7/20/2026
🔗 http://arxiv.org/abs/2607.18367v1
👥 Authors: Alayaworld Team, Kaipeng Zhang (possible past Tencent (China) affiliation), Chuanhao Li, Yifan Zhan, Yongtao Ge, Yuanyang Yin, Jiaming Tan, Kang He, Liaoyuan Fan, Mingliang Zhai, Ruicong Liu, Xiaojie Xu, Xuangeng Chu, Zhen Li (possible past Google (United States) affiliation), Zhengyuan Lin, Zhixiang Wang, Zian Meng, Zihui Gao
Abstract

Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user inputs instantly. It enable us to create customized, explorable, and continuously evolving virtual world from text, an image, or video. Realizing this vision requires four tightly coupled capabilities: interaction, persistent spatiotemporal consistency, stable long-horizon generation, and ef...

📄 O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning
🗓️ Published: 7/20/2026
🔗 http://arxiv.org/abs/2607.18142v1
👥 Authors: Mei Yuan, Qi Long, Qifeng Wu, Zhenyang Li, Yizhou Zhao, Lei Wang (possible past Baidu (China) affiliation), Yang Liu (possible past Tsinghua University affiliation), Min Xu
Abstract

Industrial Video Anomaly Detection (IVAD) aims to identify anomalous objects and events in an industrial process, which is crucial for modern manufacturing and quality control systems. Existing VLM-based anomaly reasoning methods are capable of detecting open-ended anomalies in general domains. However, their performance declines in industrial settings characterized by intricate object transformations, strict physics, and procedural constraints. To tackle the complexity of such interaction-inten...

📄 CaT-GS: Efficient 3DGS Rendering for Large Scale Scenes via Inter-frame Caching and Tile Scheduling
🗓️ Published: 7/20/2026
🔗 http://arxiv.org/abs/2607.17842v1
👥 Authors: Tingjia Zhang, Bo Chen (possible past Tencent (China) affiliation), Shengzhong Liu, Fan Wu, Guihai Chen (possible past Shanghai Jiao Tong University affiliation)
Abstract

Recent breakthroughs in 3D Gaussian Splatting (3DGS) have advanced neural rendering with high fidelity and speed. However, its performance degrades significantly in large-scale scenes due to the computational burden of tile-based rasterization. Existing optimization efforts either require costly scene re-training or focus on narrow aspects of the pipeline, overlooking critical inefficiencies in real-world deployments. Through a comprehensive analysis, we identify three primary sources of redunda...

📄 Visual Semantic Decoding of Electrocorticography from Video Stimuli using End-to-End Deep Learning
🗓️ Published: 7/21/2026
🔗 http://arxiv.org/abs/2607.18923v1
👥 Authors: Stella Ho, Joel Villalobos, Joseph West, Jingyang Liu (possible past Google (United States) affiliation), Weijie Qi, Haruhiko Kishima, Ryohei Fukuma, Takufumi Yanagisawa, Sam E. John, David B. Grayden (possible past Google (United States) affiliation)
Abstract

ECoG-based visual semantic decoding enables inference of semantic interpretation of visual perception from complex, noisy brain activity. This study examines the feasibility of visual semantic decoding using an end-to-end deep learning framework using electrocorticography (ECoG). Specifically, the decoding task is to predict visual categories from video stimuli using time-series neural inputs. A previously collected ECoG dataset from participants ($n=17$) with drug-resistant epilepsy is used for...

📄 The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric
🗓️ Published: 7/20/2026
🔗 http://arxiv.org/abs/2607.18237v1
👥 Authors: Sheng-Yu Wang, Yotam Nitzan, Aaron Hertzmann, Jun-Yan Zhu (possible past University Of California, Berkeley affiliation), Eli Shechtman, Alexei A. Efros (possible past University Of California, Berkeley affiliation), Richard Zhang
Abstract

Human visual similarity judgments are context-dependent. For example, two images may be similar in shape but distinct in color. Existing perceptual similarity metrics, however, collapse these nuances into a single scalar value, offering no mechanism to condition on specific aspects. To bridge this gap, we introduce a large-scale dataset of human similarity judgments over image triplets, where each triplet is annotated across multiple, free-form semantic aspects of similarity. Benchmarking a broa...

📄 Patch Policy: Efficient Embodied Control via Dense Visual Representations
🗓️ Published: 7/20/2026
🔗 http://arxiv.org/abs/2607.18236v1
👥 Authors: Gaoyue Zhou, Zichen Jeff Cui, Ada Langford, Bowen Tan, Yann Lecun (possible past Meta (United States) affiliation), Lerrel Pinto (possible past Carnegie Mellon University affiliation)
Abstract

Pretrained dense visual features from Vision Transformers (ViTs) are powerful yet have been underutilized in robot learning. Modern robot policies either compress each observation into a single global token, or rely on visual backbones trained from scratch, sacrificing both fine-grained spatial detail and the benefits of large-scale visual pre-training. While there exist policies that do operate on dense patch features like large vision-language-action models (VLAs), they tend to be heavy and sl...

📄 Three-Body Scattering for Generative Modeling
🗓️ Published: 7/20/2026
🔗 http://arxiv.org/abs/2607.18198v1
👥 Authors: Peng Sun (possible past Tencent (China) affiliation), Zhenglin Cheng, Deyuan Liu, Jun Xie (possible past Tencent (China) affiliation), Xinyi Shang, Tao Lin
Abstract

Modern generative models typically rely on an adversarial critic, a prescribed noise-to-data path, or an autoregressive factorization. Instead, we show that a proper distributional energy can induce sample-level motion and provide direct regression supervision for a one-step generator. Three-Body Scattering Modeling (TBSM) for generation turns the energy distance into a constant-size per-projectile interaction: each projectile is attracted toward one real source and repelled from one independent...

*Notable papers are those with at least two authors from a "big" AI/ML lab.