📄 Notable* Recent AI/ML arXiv Papers

Last updated just now...

📄 Video, Ergo Genero: Unifying Video Tasks via Spatiotemporal Analogy
🗓️ Published: 9/27/2026
🔗 http://arxiv.org/abs/2609.33935v1
👥 Authors: Chia-Hsiang Kao, Belinda Zeng, Bharath Hariharan (possible past University Of California, Berkeley affiliation), Menglin Jia (possible past Google (United States) affiliation)
Abstract

Adapting video models to new tasks typically requires dedicated data curation and fine-tuning. While visual analogy provides a training-free alternative by specifying tasks in-context, it remains restricted to the image domain. To explore whether analogy-based methods can unify diverse video tasks and generalize to out-of-distribution scenarios, we introduce ViGeo, a framework that extends visual in-context learning to the video domain via spatiotemporal canvas completion. Evaluated on a diverse...

📄 Robot-GST: geometry-aware spatial-temporal robot policy representation and evaluation
🗓️ Published: 9/27/2026
🔗 http://arxiv.org/abs/2609.33872v1
👥 Authors: Sichao Liu, Zekun Wang, Lixuan Tang, Yiming Li (possible past Tsinghua University affiliation), Xiaohan Wang (possible past Baidu (China) affiliation), Hanzhi Zhang, Daqiang Guo, Peng Zhou (possible past Tencent (China) affiliation), Lihui Wang
Abstract

Robotic manipulation policies are advancing rapidly with increasing reliance on vision-language models for end-to-end decision making. However, reliable deployment remains challenging because many policies lack explicit mechanisms for predicting task outcomes and evaluating whether generated actions will achieve desired final states, causing execution errors to accumulate during long-horizon manipulation. We present Robot-GST, a geometry-aware spatio-temporal behaviour representation and evaluat...

📄 Achieve What You Imagined: Learning to Align Actions with Visual Plans
🗓️ Published: 9/27/2026
🔗 http://arxiv.org/abs/2609.33832v1
👥 Authors: Yuheng Qiao, Ziran Wei, Xiaohan Wang (possible past Baidu (China) affiliation), Daqiang Guo, Yichen Luo, Zhibo Pang, Peng Zhou (possible past Tencent (China) affiliation), Sichao Liu
Abstract

World-action models can jointly predict future visual observations and robot actions. However, discrepancies may exist between their visual predictions and the consequences implied by generated actions. We observe that WAMs can often generate visually plausible task-completion outcomes before producing action sequences that reliably achieve them. Consequently, we treat the WAM-generated visual prediction as a goal-conditioned visual proposal rather than a directly executable plan. We use a froze...

📄 Diffusion Reward Models
🗓️ Published: 9/27/2026
🔗 http://arxiv.org/abs/2609.33803v1
👥 Authors: Xiangyang Wang, Bingxiang He, Zeyuan Liu, Jiaze Wangziqing Qiao, Yuxin Zuo, Huan-Ang Gao, Cheng Qian, Wenbin Zhang, Ran Li, Youbang Sun, Ning Ding (possible past Tsinghua University affiliation), Yuanchun Shi, Zhiyuan Liu (possible past Tsinghua University affiliation), Chaojun Xiao, Chun Yu
Abstract

Reward models underpin the alignment of large language models, yet the dominant designs reduce each prompt--response pair to a point estimate or to a distribution from a fixed parametric family. This is at odds with human preference, which is inherently multimodal: the same response can be reasonably judged in many ways, and no single family covers all of them. To better fit this structure, we introduce DRM, a Diffusion Reward Model that recasts reward modeling as conditional density estimation ...

📄 Supervision Recovery for Time Series Anomaly Detection via Context-Anchored Pairing
🗓️ Published: 9/27/2026
🔗 http://arxiv.org/abs/2609.33610v1
👥 Authors: Yifei Gao, Tian Lan, Yimeng Lu, Xuming An, Meng Wang (possible past Google (United States) affiliation), Wenjun He, Yijie Li, Chen Zhang (possible past Peking University affiliation)
Abstract

Time series anomaly detection (TSAD) remains challenging not only because anomaly labels are scarce, but also because temporal anomalies are highly context-dependent. Existing methods often rely on unsupervised objectives or surrogate abnormal patterns, providing limited supervision for context-dependent normal--anomalous distinctions. We propose Context-Anchored Pair Supervision (CAPS), a supervision-recovery framework for TSAD. CAPS views ideal anomaly supervision as a matched comparison betwe...

📄 PPG-LM: A Photoplethysmography-Language Model with Multi-Level Clinical Alignment
🗓️ Published: 9/27/2026
🔗 http://arxiv.org/abs/2609.33516v1
👥 Authors: Xiaoda Wang, Minxiao Wang, Maxwell A Xu, Patrick Langer, Kaiqiao Han, Defu Cao, Xiao Luo, Yuzhe Yang, Yan Liu (possible past Tencent (China) affiliation), Xiao Hu, Yizhou Sun, Wei Wang (possible past University Of Oxford affiliation), Carl Yang
Abstract

Photoplethysmography (PPG) is widely recorded by clinical monitors and consumer wearables, providing a scalable source of continuous physiological information. These recordings offer an opportunity for physiological assessment at scale, but realizing this potential requires models to learn from both signal-derived physiological supervision and broader clinical context captured in electronic health records (EHRs). This involves aligning information spanning local observations, care events, and en...

📄 Just Let Linear States Forget the Distant Past: Prefix Caching via Suffix Replay for Hybrid LLMs
🗓️ Published: 9/27/2026
🔗 http://arxiv.org/abs/2609.33477v1
👥 Authors: Yirui Liu, Ruoling Qi, Xuaner Wu, Yuxin Jin, Jian Chen (possible past Baidu (China) affiliation), Penghang Liu, Yafei Huang, Jiawei Shao, Xuelong Li (possible past Tencent (China) affiliation)
Abstract

Hybrid LLMs interleave full-attention layers with linear-attention layers to reduce long-context inference cost, but this structure complicates prefix caching. Full-attention KV caches are token-addressable, whereas linear-attention layers maintain recurrent states that cannot be rolled back to arbitrary prefix boundaries. Existing systems materialize recurrent-state checkpoints, restricting prefix reuse to checkpoint-aligned positions. We present SuffixReplay, the first prefix caching system ...

📄 TTRSD: Test-Time Reinforcement Learning with Self-Distillation for Vision-Language Models
🗓️ Published: 9/27/2026
🔗 http://arxiv.org/abs/2609.33414v1
👥 Authors: Shuning Wang, Zhiheng Wu, Xun Zhou, Chongyang Cui, Chen Jia, Bowen Liu, Chuanjie Li, Xiang Chen (possible past Tencent (China) affiliation), Yi Yang (possible past Baidu (China) affiliation), Yumeng Zhang, Wenjie Huang
Abstract

Test-time reinforcement learning enables vision-language models (VLMs) to adapt using unlabeled inputs. However, repeated sampling under fixed visual conditions can reinforce shared perceptual errors, while sequence-level rewards fail to isolate visual perception the foundational bottleneck that anchors multimodal reasoning risking the degradation of pre-trained reasoning capabilities. We propose TTRSD, a test-time reinforcement learning framework combining multi-view answer-level self-distillat...

📄 CoViST: Visual Token Compression via Composable States
🗓️ Published: 9/27/2026
🔗 http://arxiv.org/abs/2609.33397v1
👥 Authors: Qi Zhang (possible past Tencent (China) affiliation), Xiandong Meng, Ronggang Wang, Siwei Ma (possible past Peking University affiliation)
Abstract

Visual token compression lowers the inference cost of vision--language models by representing images with fewer tokens. However, most existing methods compress visual tokens to a reduced set, leaving the amount of visual evidence represented by each token and its original spatial context implicit. Therefore, the compressed representation does not explicitly encode how much visual information each representative carries or where it lies in the original image. This limitation arises even after a s...

📄 ZeroGAR: Benchmarking the Adversarial Robustness of Zero-Shot Graph Models
🗓️ Published: 9/27/2026
🔗 http://arxiv.org/abs/2609.33314v1
👥 Authors: Zhongjian Zhang, Xiao Wang (possible past Google (United States) affiliation), Busheng Zhang, Bo Yan, Xingtong Yu, Yue Gao (possible past Tsinghua University affiliation), Jia Li (possible past Google (United States) affiliation), Chuan Shi
Abstract

Zero-shot graph models (ZGMs), which learn transferable knowledge from source graphs and directly apply to unseen target graphs without any adaptation, have achieved promising performance and attracted considerable attention. Despite their proliferation, existing ZGMs are predominantly evaluated on clean graphs, while existing graph robustness benchmarks mainly focus on supervised settings, leaving a fundamental question largely unexplored: How robust are ZGMs when their unseen target graphs are...

📄 TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces
🗓️ Published: 9/27/2026
🔗 http://arxiv.org/abs/2609.33295v1
👥 Authors: Dehai Min, Daoan Zhang, Yiming Zeng, Huayi Zhang, Ziyi Chen, Yan Zhang, Qinbo Bai, Mengyuan Chao, Jing Ning, Qiyue Hua, Huiyi Chen, Hanrong Zhang, Henry Peng Zou, Jie Yang (possible past Shanghai Jiao Tong University affiliation), Wei Xu (possible past Tencent (China) affiliation), Philip S. Yu (possible past Tsinghua University affiliation)
Abstract

An agent can complete a task while exhibiting undesirable behavior during execution. Developers need tests for the specific behaviors encountered in deployment, beyond fixed benchmark suites. We present TraceDance, an agent system that constructs targeted benchmarks from deployment traces for user-specified undesirable behaviors. For efficient construction, Anchor-and-Confirm combines programmable retrieval with candidate-level confirmation by a Flash large language model (LLM), while the Anchor...

📄 Next Thoughts Are Distributions: Generative Autoregressive Reasoning in the Latent Space
🗓️ Published: 9/27/2026
🔗 http://arxiv.org/abs/2609.33271v1
👥 Authors: Yang Li (possible past Google (United States) affiliation), Yi Wang, Shiyuan Huang, Yang Liu (possible past Tsinghua University affiliation), Hao Wang (possible past Tsinghua University affiliation), Chengzhi Mao
Abstract

Reasoning problems often admit multiple valid ways to proceed. Continuous reasoning promises to move computation beyond language tokens into a more compact latent space, but representing several plausible ways to think next remains difficult. We introduce Autoregressive Thought Flow (ATF), which models the next continuous thought as a multimodal distribution. A causal autoregressive model performs the reasoning computation, while a lightweight diffusion head generates a plausible next thought fr...

📄 VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis
🗓️ Published: 9/27/2026
🔗 http://arxiv.org/abs/2609.33253v1
👥 Authors: Kangjie Chen, Xiangyu Li, Dongbin Zhang, Chaoda Zheng, Shijia Chen, Jinhao Deng, Hongbin Lin, Choo Sin Wai, Minqi Wang, Minghao Yang, Dake Zhong, Guorui Song, Yu Zhang (possible past Google (United States) affiliation), Xianming Liu (possible past Meta (United States) affiliation), Boyang Wang
Abstract

We present VGGT-Diff, a geometry-routed multi-view diffusion model for sparse-view novel view synthesis. Existing novel view synthesis (NVS) methods face a fundamental trade-off: reconstruction-based approaches preserve observed geometry but struggle to synthesize unseen regions, while diffusion-based methods provide strong generative priors yet rely on implicit source-to-query correspondence. VGGT-Diff bridges these regimes by routing visual geometry latents from VGGT-Ω into a pretrained video ...

📄 Inspire: Benchmarking Scientific Literature Search for Open Research Problems
🗓️ Published: 9/27/2026
🔗 http://arxiv.org/abs/2609.33233v1
👥 Authors: Jianrong Ding, Zhengyan Shi, Jianyuan Zhong, Kai Qiu, Qi Dai, Yifan Yang (possible past Tencent (China) affiliation), Chong Luo (possible past Google (United States) affiliation), Qiang Xu
Abstract

Scientific literature search often begins with an open research problem rather than a known target paper or a fixed candidate set. We introduce INSPIRE, a benchmark for evaluating agents that search prior literature to make progress on solution-redacted research problems. Each instance pairs a research brief with a target-specific cutoff three months before a later paper and evaluates ranked outputs against graded cited antecedents from that paper's realized research lineage. Search proceeds ove...

📄 When Do Models Admit They Are Wrong? Failure Disclosure Is Unstable Under Reinforcement Learning
🗓️ Published: 9/27/2026
🔗 http://arxiv.org/abs/2609.33220v1
👥 Authors: Steven Y. Feng (possible past Carnegie Mellon University affiliation), Noah D. Goodman (possible past Stanford University affiliation), Michael C. Frank, Evan Hubinger, Paul C. Bogdan, Andrew Lampinen
Abstract

Outcome-based reinforcement learning can produce models with similar task performance but very different ways of communicating about their mistakes. We study failure disclosure: whether a model admits that an attempted solution failed rather than staying silent or presenting it as successful. Across repeated outcome-only GRPO training runs, failure disclosure varies far more than task accuracy. The pattern extends to a second reasoning task and stabilized PPO, persists at 7B, and also appears in...

📄 SeOPD: Self-Evolving LLMs via Online Policy Distillation from Self-Generated Chain-of-Thought
🗓️ Published: 9/27/2026
🔗 http://arxiv.org/abs/2609.33181v1
👥 Authors: Xiaoshu Chen, Xiangyu Wong, Sihang Zhou (possible past National University Of Defense Technology affiliation), Ke Liang, Xinwang Liu (possible past National University Of Defense Technology affiliation)
Abstract

Recent advances in online policy self-distillation (OPSD) have demonstrated that large language models (LLMs) can improve their capabilities by leveraging external privileged information (PI), such as manual annotations or feedback from external environments. However, obtaining accurate annotations and constructing sophisticated environments often require substantial human effort and computation, limiting the scalability of OPSD. While a few recent studies have explored self-improvement without ...

📄 Binding Multiple Modalities via Multimodal Wasserstein Barycenter
🗓️ Published: 9/27/2026
🔗 http://arxiv.org/abs/2609.33800v1
👥 Authors: Xiaole Tang, Jiayi Xu, Xiang Gu, Yan Yang (possible past Google (United States) affiliation), Jian Sun (possible past Microsoft (United States) affiliation)
Abstract

Multimodal learning beyond two modalities commonly leverages a specific modality (e.g., text) to bind other modalities. However, how to establish a more balanced representation space that approximates shared semantics while respecting the holistic geometry of $n$-modal data remains challenging. In this work, we present BaryBind, which aims to transport the specific modality towards the Wasserstein barycenter (WB) optimized across all modalities and introduces a volumetric alignment objective to ...

📄 Positions Are Not Facts: The Mismatch Between KV Caches and Memory
🗓️ Published: 9/27/2026
🔗 http://arxiv.org/abs/2609.33759v1
👥 Authors: Changhai Zhou, Yuhua Zhou, Shiyang Zhang, Jun Gao (possible past Nvidia (United States) affiliation), Zhen Li (possible past Google (United States) affiliation), Hua Wu (possible past Baidu (China) affiliation), Hanchao Yu, Haifeng Wang (possible past Google (United States) affiliation)
Abstract

When a fact changes, how should a language model update the history stored in its key-value (KV) cache? Hiding the old record is cheap, but it may still contain needed details or answer questions about the past. We compare hiding whole records, hiding only replaced values, and deleting old text and recomputing the cache. In a controlled quantity task, masking makes all eight models prefer the new value more strongly, yet six lose complete answers through unit errors or failure to stop; keeping t...

📄 DuoOPD: Learning from Joint Teacher-Student Outcomes for Multi-Task On-Policy Distillation
🗓️ Published: 9/27/2026
🔗 http://arxiv.org/abs/2609.33711v1
👥 Authors: Ao Yu, Weibo Gao, Heng Zhou, Linan Yue, Rui Li (possible past Google (United States) affiliation), Suyi Liu, Yu Yan, Yizhong Zhang, Qi Liu (possible past Tencent (China) affiliation)
Abstract

On-policy distillation (OPD) trains a student on its own responses with token-level feedback from a stronger teacher, yet the teacher can fail on questions the student already answers correctly, and how often each model succeeds varies across tasks. OPD ignores these outcomes and, on average, pushes down even the student's correct responses; gating feedback by student correctness fixes the direction but uses the teacher in the same way whether or not it succeeded. We introduce DuoOPD, in which t...

📄 SafeMol: Dual-Modality Safety Alignment for Molecular Multimodal Models
🗓️ Published: 9/27/2026
🔗 http://arxiv.org/abs/2609.33640v1
👥 Authors: Xinmiao Wang, Ruijie Wang, Menghui Wang, Jiawei Chen (possible past Tencent (China) affiliation), Haoyue Deng, Ran Zhang, Xingxuan Zhang, Xiao Wang (possible past Google (United States) affiliation)
Abstract

Molecular multimodal models support diverse understanding and generation tasks but may introduce safety vulnerabilities when handling hazardous molecules. In this work, We reveal substantial jailbreak vulnerabilities under both text-only and graph-conditioned settings. Our analysis further shows that safety robustness must hold across input modalities while balancing safety, over-refusal, and utility. To address these challenges, we construct SafeMolBench, a molecular multimodal safety-alignment...

📄 You Only Edit Once: Incentivizing In-Context Capability of LLMs via Local Demonstration Refinement
🗓️ Published: 9/27/2026
🔗 http://arxiv.org/abs/2609.33609v1
👥 Authors: Jiarong Wen, Qi Wang (possible past Tsinghua University affiliation), Yun Qu, Yixiu Mao, Heming Zou, Haoang Chi, Lizhou Cai, Yiqin Lv, Kaiyu Zhang, Yuhang Jiang, Xiangyang Ji (possible past Tsinghua University affiliation)
Abstract

In-context learning (ICL) is crucial for boosting the inference performance of large language models (LLMs). However, the effectiveness of ICL in LLMs is greatly influenced by the choice of demonstration sets. Exhaustive searches over these sets are combinatorial, and existing selectors often rely on relevance or likelihood proxies to implicitly assess ICL quality. Making repeated queries to the target LLM with these strategies can incur substantial costs. This work simplifies selection by frami...

📄 Hierarchical Response Preservation for Continual Adaptation of Zero-Shot Graph-Text Models
🗓️ Published: 9/27/2026
🔗 http://arxiv.org/abs/2609.33607v1
👥 Authors: Haopeng Zhang, Yuhan Wang (possible past Tencent (China) affiliation), Yubing Su, Yingxin Chen, Xiao Wang (possible past Google (United States) affiliation), Ruijie Wang, Jianxin Li
Abstract

Pretrained graph-text models align graph representations with textual semantics, enabling recognition of unseen classes and transfer across graph domains. However, as graph data and classes continually arrive, models should learn from new supervision while retaining their zero-shot transfer capabilities and historical task knowledge. Two challenges arise: (i) new classes can overturn historical predictions despite preserved distinctions among historical classes, and (ii) overly strict response p...

*Notable papers are those with at least two authors from a "big" AI/ML lab.