📄 Notable* Recent AI/ML arXiv Papers

Last updated just now...

📄 Simple Domain Generalization for Strong Pixel-Level Image Tampering Detection in Modern VLMs
🗓️ Published: 7/20/2026
🔗 http://arxiv.org/abs/2607.18230v1
👥 Authors: Yi Tang, Xinyi Shang, Jiacheng Cui, Sondos Mahmoud Bsharat, Jiacheng Liu, Xiaohan Zhao, Tran Dinh Tien, Ahmed Elhagry, Salwa K. Al Khatib, Tianjun Yao, Yonina C. Eldar, Jing-Hao Xue, Hao Li (possible past Tsinghua University affiliation), Salman Khan (possible past Inception Institute Of Artificial Intelligence affiliation), Zhiqiang Shen
Abstract

Modern vision-language models (VLMs) have significantly improved image generation and editing capabilities, making pixel-level image tampering detection increasingly important yet challenging under cross-model and out-of-distribution shifts. This work studies domain generalization for pixel-level image tampering detection in modern VLMs like ChatGPT, Gemini, Qwen-Image, etc., aiming to learn tampering localization models that remain robust across diverse VLM-generated manipulation distributions....

📄 GigaPath-Flash and GigaTIME-Flash: Efficient Pathology Foundation Models for Whole-Slide and Tumor Microenvironment Analysis
🗓️ Published: 7/20/2026
🔗 http://arxiv.org/abs/2607.18218v1
👥 Authors: Naoto Usuyama, Jeya Maria Jose Valanarasu, Sicong Yao, Hanwen Xu, Jaspreet Bagga, Guanghui Qin, Robert E. Kramer, Cliff Wong, Soohee Lee, Hao Qiu, Theodore Zhengde Zhao, Racheli Ben Shimol, Angela Crabtree, Kevin Matlock, Eduardo Alejandro Lozano Garcia, Naiteek Sangani, Alberto Santamaria-Pang, Jason Entenmann, Alexandra Q. Bartlett, Bill J. Wright, Bernard A. Fox, Brian Piening, Sheng Zhang, Sheng Wang (possible past Tencent (China) affiliation), Tristan Naumann, Carlo Bifulco, Hoifung Poon (possible past Microsoft (United States) affiliation)
Abstract

Foundation models have emerged as a driving force in computational pathology, with the potential to transform cancer diagnosis, prognosis, and treatment selection by learning transferable representations from large-scale histopathology data. A growing landscape of pathology foundation models now spans diverse data sources, architectures, and downstream applications. However, most pretrained models operate only at the image-tile level, use restrictive licenses, and remain computationally expensiv...

📄 O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning
🗓️ Published: 7/20/2026
🔗 http://arxiv.org/abs/2607.18142v1
👥 Authors: Mei Yuan, Qi Long, Qifeng Wu, Zhenyang Li, Yizhou Zhao, Lei Wang (possible past Baidu (China) affiliation), Yang Liu (possible past Tsinghua University affiliation), Min Xu
Abstract

Industrial Video Anomaly Detection (IVAD) aims to identify anomalous objects and events in an industrial process, which is crucial for modern manufacturing and quality control systems. Existing VLM-based anomaly reasoning methods are capable of detecting open-ended anomalies in general domains. However, their performance declines in industrial settings characterized by intricate object transformations, strict physics, and procedural constraints. To tackle the complexity of such interaction-inten...

📄 CaT-GS: Efficient 3DGS Rendering for Large Scale Scenes via Inter-frame Caching and Tile Scheduling
🗓️ Published: 7/20/2026
🔗 http://arxiv.org/abs/2607.17842v1
👥 Authors: Tingjia Zhang, Bo Chen (possible past Tencent (China) affiliation), Shengzhong Liu, Fan Wu, Guihai Chen (possible past Shanghai Jiao Tong University affiliation)
Abstract

Recent breakthroughs in 3D Gaussian Splatting (3DGS) have advanced neural rendering with high fidelity and speed. However, its performance degrades significantly in large-scale scenes due to the computational burden of tile-based rasterization. Existing optimization efforts either require costly scene re-training or focus on narrow aspects of the pipeline, overlooking critical inefficiencies in real-world deployments. Through a comprehensive analysis, we identify three primary sources of redunda...

📄 ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video
🗓️ Published: 7/20/2026
🔗 http://arxiv.org/abs/2607.17790v1
👥 Authors: Xiaozhong Lyu, Gen Li (possible past University Of Edinburgh affiliation), Zhiyin Qian, Xucong Zhang (possible past Eth Zurich affiliation), Marc Pollefeys (possible past Google (United States) affiliation), Siyu Tang (possible past Eth Zurich affiliation)
Abstract

Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the continuous interaction between a human viewer and the surrounding environment. A holistic and efficient multimodal model capable of reconstructing this 4D representation is therefore highly desirable. However, existing approaches often rely on auxiliary inputs such as pre-computed camera trajectories, treat scene perception and human ego-motion modeling as separate problems despite their str...

📄 Thinking in Video: Can Video Generators Really Reason About the Real World?
🗓️ Published: 7/20/2026
🔗 http://arxiv.org/abs/2607.17523v1
👥 Authors: Yongheng Zhang, Guang Yang, Ruihan Hou, Qiguang Chen, Ziang Liu, Xiaolong Liu, Manman Zhang, Yanchao Hao, Zheng Wei, Hao Wu (possible past Tencent (China) affiliation), Libo Qin, Peishan Dai, Yinghui Li, Di Yin, Xing Sun (possible past Tencent (China) affiliation)
Abstract

Recent advances in world models and video generation have given rise to an emerging reasoning paradigm that leverages video generative models to simulate, predict, and reason about real-world dynamics. We redefine this paradigm as Thinking in Video, where video is not merely an output artifact but a medium for constructing, extending, and verifying causal thought. However, this promise remains unverified: convincing rollouts may reflect memorized appearances rather than causal understanding, whi...

📄 Pailitao-MMSearch: Building Native E-Commerce Multimodal Search Foundation
🗓️ Published: 7/20/2026
🔗 http://arxiv.org/abs/2607.17499v1
👥 Authors: Xiaohan Ye, Xu Chen (possible past Tencent (China) affiliation), Zihan Gong, Jian Ding (possible past Baidu (China) affiliation), Lianyu Du, Baicheng Chen, Yunmeng Shu, Jingqian Zhao, Zhixiang Zhao, Shuaiqi Jia, Chong Ma, Shuwen Xiao, Xiangheng Kong, Yuan Gao (possible past Tencent (China) affiliation), Jun Song, Jinsong Lan, Xiaoyong Zhu, Bo Zheng
Abstract

The evolution of e-commerce has fundamentally transformed how users search for products, shifting from simple text-based keyword queries to complex multimodal interactions that seamlessly combine product images, natural language descriptions, and mixed-intent instructions. However, existing approaches face a critical dilemma: single-modal specialist models, deployed independently for text retrieval, visual search, and voice recognition, operate in isolation and cannot handle cross-modal queries,...

📄 Toward Anthropomorphic Dialogue: A Closed-Loop Framework for Human-Like Chat Generation, Evaluation, and Preference Alignment
🗓️ Published: 7/19/2026
🔗 http://arxiv.org/abs/2607.17191v1
👥 Authors: Wentao Liu, Siyu Song, Xi Chen (possible past University Of California, Berkeley affiliation), Youjia Li, Xiaokun Wang, Min Ji, Ji Wang (possible past Tencent (China) affiliation)
Abstract

Human-like private chat requires more than fluent response generation: a system must preserve persona, relationship, memory, bounded knowledge, medium-specific timing, and a coherent multi-turn arc. We present AnthroDial, a closed-loop framework that formulates anthropomorphic dialogue as a joint problem of system architecture, executable evaluation, and diagnostic alignment. It combines (1) a role-conditioned scheduled dialogue runtime with persona and scenario cards, long-term memory, virtual ...

📄 Fourier Geometric Wind Power Forecasting with Numerical Weather Prediction
🗓️ Published: 7/19/2026
🔗 http://arxiv.org/abs/2607.17095v1
👥 Authors: Shiyuan Piao, Fan Zehui, Yang Liu (possible past Tsinghua University affiliation), Hong Cheng, Juepeng Zheng (possible past Tsinghua University affiliation), Jie Zhou (possible past Tsinghua University affiliation), Fugee Tsung
Abstract

Accurate short-term wind power forecasting is essential for grid stability and operational planning, yet remains challenging due to the complex interactions between atmospheric conditions and turbine dynamics. However, existing methods fail to effectively incorporate weather forecasting with wind turbine data (i.e., SCADA), leading to suboptimal solutions. To address this, we introduce a multimodal framework that integrates historical point-based SCADA data with grid-based Numerical Weather Pred...

📄 EvoGUI: An Evolution-Aware Benchmark for GUI State-Transition Understanding
🗓️ Published: 7/19/2026
🔗 http://arxiv.org/abs/2607.17050v1
👥 Authors: Yaohan Yang, Minglei Shi, Borui Zhang, Jie Zhou (possible past Tsinghua University affiliation), Jiwen Lu (possible past Tsinghua University affiliation)
Abstract

GUI agents must reason about how actions transform interface states, but end-to-end success rates entangle this ability with perception, grounding, planning, and recovery. We introduce EvoGUI, a diagnostic framework that converts normalized GUI trajectories into three complementary visual question answering probes: temporal ordering, inverse action/value prediction, and contrastive one-step successor discrimination. Their labels are derived from trajectory order and logged actions, requiring no ...

📄 Training Continuous Chain of Thought Models: A Tale of Two Regimes
🗓️ Published: 7/18/2026
🔗 http://arxiv.org/abs/2607.16972v1
👥 Authors: Varun Yerram, He He (possible past Stanford University affiliation), Eunsol Choi (possible past Google (United States) affiliation)
Abstract

Continuous Chain-of-Thought methods replace verbose reasoning traces with a short sequence of dense latent representations. Earlier continuous CoT methods indirectly supervise the latent representations such that its final state match that of verbose reasoning traces, requiring autoregressive, slow generation during training. We introduce C-MTP, a simpler, faster direct supervision approach that models each latent as an average of the embeddings in the CoT traces to be compressed. Our approach o...

📄 Environment-free Synthetic Data Generation for API-Calling Agents
🗓️ Published: 7/18/2026
🔗 http://arxiv.org/abs/2607.16900v1
👥 Authors: Seanie Lee, Sanjoy Chowdhury, Chao Jiang, Cheng-Yu Hsieh (possible past University Of Washington affiliation), Ting-Yao Hu, Alexander T Toshev, Oncel Tuzel (possible past Apple (United States) affiliation), Raviteja Vemulapalli (possible past Google (United States) affiliation)
Abstract

Training API-calling large language model (LLM) agents demands massive amounts of high-quality trajectories. However, collecting such data at scale typically requires fully implemented environments with executable APIs and realistic, pre-populated backend databases, creating a major bottleneck for scalability. To overcome this, we propose an environment-free synthetic data generation approach that leverages LLMs as on-the-fly digital world models. Given only API specifications, our method genera...

📄 The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric
🗓️ Published: 7/20/2026
🔗 http://arxiv.org/abs/2607.18237v1
👥 Authors: Sheng-Yu Wang, Yotam Nitzan, Aaron Hertzmann, Jun-Yan Zhu (possible past University Of California, Berkeley affiliation), Eli Shechtman, Alexei A. Efros (possible past University Of California, Berkeley affiliation), Richard Zhang
Abstract

Human visual similarity judgments are context-dependent. For example, two images may be similar in shape but distinct in color. Existing perceptual similarity metrics, however, collapse these nuances into a single scalar value, offering no mechanism to condition on specific aspects. To bridge this gap, we introduce a large-scale dataset of human similarity judgments over image triplets, where each triplet is annotated across multiple, free-form semantic aspects of similarity. Benchmarking a broa...

📄 Patch Policy: Efficient Embodied Control via Dense Visual Representations
🗓️ Published: 7/20/2026
🔗 http://arxiv.org/abs/2607.18236v1
👥 Authors: Gaoyue Zhou, Zichen Jeff Cui, Ada Langford, Bowen Tan, Yann Lecun (possible past Meta (United States) affiliation), Lerrel Pinto (possible past Carnegie Mellon University affiliation)
Abstract

Pretrained dense visual features from Vision Transformers (ViTs) are powerful yet have been underutilized in robot learning. Modern robot policies either compress each observation into a single global token, or rely on visual backbones trained from scratch, sacrificing both fine-grained spatial detail and the benefits of large-scale visual pre-training. While there exist policies that do operate on dense patch features like large vision-language-action models (VLAs), they tend to be heavy and sl...

📄 Three-Body Scattering for Generative Modeling
🗓️ Published: 7/20/2026
🔗 http://arxiv.org/abs/2607.18198v1
👥 Authors: Peng Sun (possible past Tencent (China) affiliation), Zhenglin Cheng, Deyuan Liu, Jun Xie (possible past Tencent (China) affiliation), Xinyi Shang, Tao Lin
Abstract

Modern generative models typically rely on an adversarial critic, a prescribed noise-to-data path, or an autoregressive factorization. Instead, we show that a proper distributional energy can induce sample-level motion and provide direct regression supervision for a one-step generator. Three-Body Scattering Modeling (TBSM) for generation turns the energy distance into a constant-size per-projectile interaction: each projectile is attracted toward one real source and repelled from one independent...

*Notable papers are those with at least two authors from a "big" AI/ML lab.