πŸ“„ Notable* Recent AI/ML arXiv Papers

Last updated just now...

πŸ“„ MindTopo: Can Foundation Models Reason in Topological Space?
πŸ—“οΈ Published: 9/10/2026
πŸ”— http://arxiv.org/abs/2609.11900v1
πŸ‘₯ Authors: Yunfei Ge, Anbang Liu, Qineng Wang, Johnalbert Garnica, Jianwen Lyu, Zihan Wang (possible past Tsinghua University affiliation), Reuben Tan, Jianfeng Gao (possible past Microsoft (United States) affiliation), Ruohan Zhang, Yining Hong, Jiajun Wu (possible past Massachusetts Institute Of Technology affiliation), Manling Li
Abstract

Spatial reasoning depends not only on metric properties such as distance, angle, and shape, but also on topological relations that remain invariant under continuous deformation. Cognitive science identifies these relations as foundational to spatial understanding, yet foundation-model evaluations largely focus on metric or viewpoint-dependent relations. We introduce MindTopo, a benchmark of topological intuition across five properties grounded in cognitive science and formal topology: continuity...

πŸ“„ The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement
πŸ—“οΈ Published: 9/10/2026
πŸ”— http://arxiv.org/abs/2609.11873v1
πŸ‘₯ Authors: Yi Duan, Ying Liu, Zirui Tang, Haodong Chen, Jun Zhou, Yumou Liu, Bangrui Xu, Yukai Wu, Sidi Chen, Yuhan Zhou, Haoyu Wang (possible past Tencent (China) affiliation), Xiaoyou Yu, Shaokun Han, Xuzhou Zhu, Le Zhou, Bolin Lu, Wei Zhou, Jiachen Liu (possible past Baidu (China) affiliation), Nuozhou Fang, Jiaxin Tian, Ruoyu Chen, Yuxuan Li, Kai Zuo, Kaiyan Zhang, Jiantao Qiu, Conghui He (possible past Tsinghua University affiliation), Guoliang Li (possible past Tsinghua University affiliation), Bowen Zhou, Zhiyuan Liu (possible past Tsinghua University affiliation), Zhoufutu Wen, Jihua Kang, Xuanhe Zhou, Fan Wu
Abstract

Recursive self-improvement (RSI) enables AI systems to turn experience and feedback into persistent changes that improve both their capabilities and the process of future improvement. We first use the Headroom-Closed Index (HCI) to reveal the problems of existing LLMs, then introduce the RSI concept and its development roadmap: from improvement-execution autonomy, improvement-strategy autonomy, experience-acquisition autonomy, and environment-adaptation autonomy, to recursive meta-improvement. N...

πŸ“„ ActSafeGuard: Differentiable and Training-Aligned Constraint Enforcement for Flow-Matching Policies
πŸ—“οΈ Published: 9/10/2026
πŸ”— http://arxiv.org/abs/2609.11697v1
πŸ‘₯ Authors: Jianming Ma, Rongjun Jin, Xiaxi Si, Yang Zhang (possible past Tsinghua University affiliation), Yiheng Li, Yue Gao (possible past Tsinghua University affiliation)
Abstract

Vision-Language-Action (VLA) and World-Action Models (WAMs) have demonstrated strong capabilities in general-purpose robotic manipulation, yet their generated actions may violate hard physical constraints and therefore be unsafe or infeasible for deployment. Existing safety approaches either optimize statistical safety objectives without deterministic per-step guarantees or correct unsafe actions only during inference, creating a mismatch between policy training and execution. We introduce ActSa...

πŸ“„ Learning Interaction between Image and Layout Priors for Joint Image-Layout Generation in Design Templates
πŸ—“οΈ Published: 9/10/2026
πŸ”— http://arxiv.org/abs/2609.11519v1
πŸ‘₯ Authors: Shirong Yang, Bo Yang (possible past Tencent (China) affiliation), Ying Cao (possible past Baidu (China) affiliation)
Abstract

In this paper, we address the problem of graphic design template creation, which generates a background image and a layout of foreground elements over the background to form a harmonious composition from an input text. Prior work on graphic design generation mostly adopts a sequential paradigm, where design elements are generated sequentially. We argue that such a sequential scheme falls short of faithfully capturing the dependency between the background and layout (and thus the joint image-la...

πŸ“„ RouteRepair: Instance-Level Failure Diagnosis and Targeted Repair in LLM-Based Automated Heuristic Design for Routing Optimization
πŸ—“οΈ Published: 9/10/2026
πŸ”— http://arxiv.org/abs/2609.11452v1
πŸ‘₯ Authors: Binghao Ji, Di Huang (possible past Google (United States) affiliation), Jiahui Fang, Zhiyuan Liu (possible past Tsinghua University affiliation)
Abstract

Efficient routing optimization is essential to freight transportation, urban logistics, and shared mobility, where high-quality heuristics are often required under limited computational budgets. Recent large language model (LLM)-based automated heuristic design methods can generate effective routing rules, but aggregate evaluation may mask recurrent failures on particular instance structures. To address this limitation, this study develops RouteRepair, which diagnoses parent-specific weaknesses ...

πŸ“„ SWRouter: Similarity-Contractive Window Routing for Multi-Turn Large Language Model Conversations
πŸ—“οΈ Published: 9/10/2026
πŸ”— http://arxiv.org/abs/2609.11414v1
πŸ‘₯ Authors: Yu Wang (possible past Tsinghua University affiliation), Yuchen Li, Rui Kong, Xinran Chen, Jiamin Chen, Hengyi Cai, Shuaiqiang Wang (possible past Baidu (China) affiliation), Jiashu Zhao, Yulun Zhang, Zhonghao Lyu, Haoyi Xiong (possible past Baidu (China) affiliation), Linghe Kong, Jimmy Xiangji Huang (possible past Tencent (China) affiliation), Dawei Yin (possible past Baidu (China) affiliation)
Abstract

Large language models exhibit complementary strengths, motivating routing methods that dispatch each query to the most suitable model. Although existing routers are effective in single-turn settings, they do not directly transfer to multi-turn dialogue, where routing performance critically depends on how historical context is segmented, retained, and incorporated into the current prompt. This introduces two fundamental challenges: preventing information loss and information confusion during cont...

πŸ“„ X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation
πŸ—“οΈ Published: 9/10/2026
πŸ”— http://arxiv.org/abs/2609.11412v1
πŸ‘₯ Authors: Haojun Zhang, Yi Zou, Min Chen, Qize Yu, Lianrui Fan, Xini Ding, Hao Li (possible past Tsinghua University affiliation), Shuchang Zhou, Xianming Liu (possible past Meta (United States) affiliation), Shiyu Huang
Abstract

Reducing audio-encoder depth lowers the inference cost of speech large language models, but removing complete blocks perturbs the embeddings consumed by the decoder and can cause deletion and premature end-of-sequence errors. We introduce X-AuT, a progressive framework that selects layer combinations through short behavioral probes and restores the pruned model through representation alignment, cross-scale distillation, scheduled student-policy supervision, and LoRA finetuning. The language-mode...

πŸ“„ Your Model Already Knows Don't Teach It, Learn to Ask It: Soft Prompting for Few-Shot Adaptation of Vision-Language Models
πŸ—“οΈ Published: 9/10/2026
πŸ”— http://arxiv.org/abs/2609.11310v1
πŸ‘₯ Authors: Gautam Rajendrakumar Gare, Siyi Li, Hewei Wang, Cesar Daniel Hernandez, Wei Zhao (possible past Tencent (China) affiliation), Wolfgang M. Pauli, John Galeotti, Deva Ramanan (possible past Carnegie Mellon University affiliation)
Abstract

We address few-shot object detection with vision-language models (VLMs) in out-of-domain settings such as aerial, industrial, and medical imagery, using only ten annotated images for supervision. Existing adaptation methods are discrete prompt optimization and LoRA fine-tuning. We revisit a third option: soft prompting, where a small number of continuous prompt tokens are optimized while the pretrained backbone remains frozen. We identify two key design choices. First, placing prompt tokens at...

πŸ“„ Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents
πŸ—“οΈ Published: 9/10/2026
πŸ”— http://arxiv.org/abs/2609.11243v1
πŸ‘₯ Authors: Jiaqiang Li, Yajie Yang, Zhiheng Xi, Jiadong Chen, Enyu Zhou, Senjie Jin, Yang Nan, Jiazheng Zhang, Han Wang (possible past Peking University affiliation), Yanxin Li, Dingwei Zhu, Bicheng Deng, Yuhui Wang, Xiang Zheng, Qi Zhang (possible past Tencent (China) affiliation), Lei Bai, Xingjun Ma, Tao Gui
Abstract

Autonomous research agents are increasingly expected to search the literature, analyze experimental evidence, and generate scientific hypotheses. These capabilities require multi-step evidence grounded reasoning that progressively acquires, integrates, and verifies evidence before reaching a conclusion. Existing multimodal benchmarks, however, largely evaluate final-answer accuracy, leaving open whether predictions are actually supported by traceable scientific evidence. We introduce Sci-MMR, a ...

πŸ“„ NovGauge: A Fine-Grained Benchmark for Diagnosing LLMs' Capability in Paper Novelty Assessment
πŸ—“οΈ Published: 9/10/2026
πŸ”— http://arxiv.org/abs/2609.11234v1
πŸ‘₯ Authors: Guoqiang Zhang, Kexin Tan, Ming Zhang (possible past Peking University affiliation), Li Ju, Wenqing Jing, Zhonghan Yue, Jiayi Chen, Shiqiang Wu, Shaofan Liu, Yue Zhang, Yuankai Ying, Yang Shi, Tao Gui, Qi Zhang (possible past Tencent (China) affiliation), Xuanjing Huang
Abstract

Large language models (LLMs) are increasingly used in peer review at major AI conferences, yet novelty remains a persistent weak point. Existing benchmarks assess novelty as a single holistic score, making it difficult to diagnose which dimension a model misjudges or whether its evidence is faithful. We present NovGauge, a human-anchored benchmark for fine-grained novelty assessment diagnosis. The benchmark contains 619 paper pairs and 50 multi-paper sets, drawn from two expert sources: ICLR rev...

πŸ“„ Autonomous Chemical Mechanistic Discovery through Agentic Reasoning and Validation
πŸ—“οΈ Published: 9/10/2026
πŸ”— http://arxiv.org/abs/2609.11147v1
πŸ‘₯ Authors: Dong Li, Sixuan Mi, Zihao Ye, Huan Xiong, Tao Xu (possible past Meta (United States) affiliation), Tong Zhu (possible past Nvidia (United States) affiliation), Aijia Zhang, Junqi Gao, Kaiyan Zhang, Shijie Wang, Bowen Zhou, Yuqiang Li, Biqing Qi
Abstract

Unraveling reaction mechanisms is central to modern chemistry, yet automating these investigations remains challenging because computational workflows still rely heavily on expert intervention. Here we introduce ARCHE, an autonomous agentic system that integrates a general-purpose reasoning model, a domain-specialized computational chemistry model, and a structured tool registry to transform mechanistic inquiry into a scalable, self-validating process. ARCHE interprets scientific questions, gene...

πŸ“„ Fork Where the Model Changes Its Mind: Belief-Shift Branching for Tree-Structured Reinforcement Learning
πŸ—“οΈ Published: 9/10/2026
πŸ”— http://arxiv.org/abs/2609.11061v1
πŸ‘₯ Authors: Bin Lei, Yu Li (possible past Tencent (China) affiliation), Prafulla Kumar Choubey, Jiaxin Zhang, Becky Xiangyu Peng, Qinyuan Ye, Kartik Narayan, Caiwen Ding, Silvio Savarese (possible past Stanford University affiliation), Chien-Sheng Wu (possible past Salesforce (United States) affiliation)
Abstract

Tree-structured rollouts give critic-free reinforcement learning with verifiable rewards (RLVR) step-level credit: fork a chain at an intermediate point, and sibling outcome differences estimate step value. Each fork adds sampling cost, so realistic budgets typically allow only a few forks per chain. A fork placed where the outcome is already largely settled yields siblings that mostly agree and provide almost no credit signal; hence, for a given tree size, where forks are placed largely determi...

πŸ“„ BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure
πŸ—“οΈ Published: 9/10/2026
πŸ”— http://arxiv.org/abs/2609.11028v1
πŸ‘₯ Authors: Shenghan Zheng, Zonglin Di, Yimin Liu, Kyoung Whan Choe, Jiankai Sun, Heguang Lin, Penghao Jiang, Yifeng He, Xiao Cheng, Jicheng Wang, Wenbo Chen, Alex Yates, Yinzhe Zhao, Bingran You, Yuan Gao (possible past Tencent (China) affiliation), Ayush Munot, Shubham Gaur, Zhe Ye, Hao Wang (possible past Tsinghua University affiliation), Xiangyi Li, Dawn Song (possible past University Of California, Berkeley affiliation), Christophe Hauser
Abstract

LM-agent benchmarks increasingly function as interactive evaluation infrastructure. Agents observe state, call tools, modify workspaces, submit artifacts, and receive rewards from outcome procedures. This interactivity makes evaluations vulnerable to reward hacking: an agent improves its measured score by exploiting the reward-relevant trajectory instead of solving the intended task. Existing defenses rely largely on task-specific patches, prompt instructions, or post-hoc detectors. They d...

πŸ“„ Demystifying the Privacy-Utility Trade-off in LLM Interactions
πŸ—“οΈ Published: 9/10/2026
πŸ”— http://arxiv.org/abs/2609.10992v1
πŸ‘₯ Authors: Zhenhua Liu (possible past Tencent (China) affiliation), Zhanxu Xie, Junjie Yu, Tong Zhu (possible past Nvidia (United States) affiliation), Lijun Li, Wenliang Chen
Abstract

The integration of Large Language Models into daily tasks relies on context-rich instructions, inevitably exposing sensitive user information. Current privacy-preserving methods typically employ context-agnostic static rules, causing severe utility degradation. However, the specific mechanisms governing how sanitization impacts downstream performance remain largely underexplored. To address this, we conduct a systematic analysis to deconstruct the privacy-utility trade-off, uncovering three unde...

πŸ“„ Tapes Together Strong: The Co-evolution of Computation and Cooperation
πŸ—“οΈ Published: 9/9/2026
πŸ”— http://arxiv.org/abs/2609.10817v1
πŸ‘₯ Authors: Kunal Jha, Francesco Cicala, Blaise AgΓΌera Y Arcas (possible past Google (United States) affiliation), Blake Aaron Richards, Natasha Jaques (possible past University Of California, Berkeley affiliation), Max Kleiman-Weiner, Eyvind Niklasson (possible past Google (United States) affiliation)
Abstract

How does cooperation evolve in complex agentic systems? Prior work in evolutionary game theory studies why individuals are incentivized to cooperate by isolating social interactions from the physical costs of behavior, while artificial life models traditionally study emergent self-replication without formalizing the dilemma between acquiring resources and preserving the shared energy needed to reproduce. In contrast, we introduce Autopoietic Game Theory, a computational model where social intera...

πŸ“„ Show-Harness: Just a VLM Agent Can Play Robots
πŸ—“οΈ Published: 9/9/2026
πŸ”— http://arxiv.org/abs/2609.10522v1
πŸ‘₯ Authors: Yanzhe Chen, Zechen Bai, Zhijun Cao, Wenzheng Zeng, Kevin Qinghong Lin (possible past National University Of Singapore affiliation), Yiqi Lin, Guoqiang Liang, Kevin Yuchen Ma, Qiming Huang, Mike Zheng Shou (possible past National University Of Singapore affiliation)
Abstract

Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains challenging. We present Show-Harness, an Embodied Harness that enables VLMs to "play" robots through a compact semantic interface linking intent to action. Show-Harness exposes discrete semantic action units that VLMs can naturally reason over, while embodiment-specific interpreters deterministically ground them into local robot actions, keeping the VL...

πŸ“„ Beyond One-Size-Fits-All: Sample-Adaptive Strategy Routing for Vision Token Pruning in MLLMs
πŸ—“οΈ Published: 9/9/2026
πŸ”— http://arxiv.org/abs/2609.10346v1
πŸ‘₯ Authors: Haiji Liang, Pengfei Zhou, Zhenglin Wan, Wei Wang (possible past University Of Oxford affiliation), Yang You (possible past University Of California, Berkeley affiliation), Wangbo Zhao
Abstract

Multimodal large language models (MLLMs) process hundreds or thousands of visual tokens per image, incurring prohibitive inference costs. While existing vision token pruning methods mitigate this overhead, they implicitly assume that a single fixed pruning strategy can be applied uniformly across all inputs. Our analysis further reveals that ranking pruning methods by average benchmark accuracy conceals substantial sample-wise complementarity: although the average-best strategy excels overall, a...

πŸ“„ GANDR: Claim Auditing for Verifiable Legal Answer Generation
πŸ—“οΈ Published: 9/9/2026
πŸ”— http://arxiv.org/abs/2609.10293v1
πŸ‘₯ Authors: Chen Qian (possible past Shanghai Jiao Tong University affiliation), Yimeng Wang, Yu Chen (possible past Meta (United States) affiliation), Lingfei Wu (possible past Tencent (China) affiliation), Andreas Stathopoulos
Abstract

In high-stakes domains such as legal practice, a language-model answer is only useful to the extent that a reader can verify each claim against the source the system cites. Current grounded-generation pipelines score the answer as a whole, so a correct conclusion can rest on fabricated or loosely matched citations and still score well. Closing this gap requires both a system built for per-claim verification and an evaluation that measures it. We introduce GANDR (Grounded ANswer DRafter), a two-a...

πŸ“„ A Trust-Network-Based Federated Learning Framework for Multi-Center Aging Clock Prediction
πŸ—“οΈ Published: 9/9/2026
πŸ”— http://arxiv.org/abs/2609.10108v1
πŸ‘₯ Authors: Chunxu Zhang, Bo Li (possible past Tencent (China) affiliation), Wenliang Wang, Yang Liu (possible past Tsinghua University affiliation), Di Jiang, Yuan Huang, Yo-Ichi Nabeshima, Akinori Yamamura, Bo Yang (possible past Tencent (China) affiliation), Qiang Yang
Abstract

Aging clocks quantify biological aging and help characterize individual health status. What protein interactions are important for accurate aging clocks, and are they zeroth-order or higher-order? Addressing these questions requires learning from large molecular datasets distributed across medical centers, where privacy constraints prevent centralized data sharing. Federated learning offers a natural solution but faces four challenges in this setting: limited local sample sizes, sparse and direc...

πŸ“„ Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data
πŸ—“οΈ Published: 9/10/2026
πŸ”— http://arxiv.org/abs/2609.11917v1
πŸ‘₯ Authors: Atindra Jha, Margaret Li, Jure Leskovec (possible past Stanford University affiliation), Percy Liang (possible past Stanford University affiliation), Luke Zettlemoyer (possible past University Of Washington affiliation)
Abstract

As the supply of human-written text is exhausted, it has become standard practice to repeat language model training data. Prior work has studied data repetition for densely activated Transformers, but the effects of data repetition remains largely unexplored for recently dominant sparse architectures such as Mixture-of-Experts (MoE), despite their increased compute efficiency. We vary data repetition rates across single- and multi-domain data mixes, and across MoE settings, including expert coun...

πŸ“„ CoRA-NAS: Coarse Ranking and Anchor-Residual Refinement for Neural Architecture Search
πŸ—“οΈ Published: 9/10/2026
πŸ”— http://arxiv.org/abs/2609.11884v1
πŸ‘₯ Authors: Yifan Yang (possible past Tencent (China) affiliation), Zhaoyan Wang, Zheng Gao, Xiaoyu Li (possible past Tencent (China) affiliation), Jiaojiao Jiang
Abstract

Zero-cost proxies rank architectures cheaply, but their reliability varies across search spaces. We introduce CoRA-NAS (COarse Ranking + Anchor-residual), a two-stage framework combining a static ranking prior with low-cost learning-curve refinement. CoRA-Rank aggregates capacity and structure-at-initialization proxies through an equal-weight log-rank consensus and a target-free consensus gate. CoRA-Refine samples anchors across this prior, extrapolates their early validation curves, and propaga...

πŸ“„ Reflex-Informed Neuromuscular Reinforcement Learning for Muscle-Driven Locomotion
πŸ—“οΈ Published: 9/10/2026
πŸ”— http://arxiv.org/abs/2609.11733v1
πŸ‘₯ Authors: Jian Zhou (possible past Tencent (China) affiliation), Xingyu Zhang, Rui Ma, Yu Cao (possible past University Of California, Berkeley affiliation), Shane Xie, Zhi-Qiang Zhang
Abstract

Muscle-driven locomotion provides a physically grounded approach to generating realistic human movement. However, achieving both physiological plausibility and adaptability to changes in musculoskeletal capacity and external disturbances remains a fundamental challenge. To address this limitation, we propose a Reflex-Informed Neuromuscular Reinforcement Learning framework for muscle-driven locomotion. Within this framework, a fixed phase-dependent reflex controller serves as the underlying neuro...

πŸ“„ Coherent Floquet quantum reservoirs for molecular property prediction
πŸ—“οΈ Published: 9/10/2026
πŸ”— http://arxiv.org/abs/2609.11071v1
πŸ‘₯ Authors: Luofei Wang, Da Zhang, Congren Wang, Yiming Li (possible past Tsinghua University affiliation), Yuxiao Yang, Xuan Zhang (possible past Meta (United States) affiliation), Xuefeng Cui, Zhang-Qi Yin
Abstract

Quantum reservoir computing (QRC) uses quantum dynamics to represent input histories for prediction through a trained classical readout. Discrete time crystals (DTCs) exhibit robust subharmonic responses under periodic driving, and previous work has used their dynamics to construct DTC-QRC. Here we construct a DTC-based reservoir architecture to predict molecular properties from structural and dynamical observations. Coherent Floquet evolution processes local molecular graph events and surface-h...

πŸ“„ From Connectivity to Rewards: Dense Reward Learning with Directed State Graphs
πŸ—“οΈ Published: 9/9/2026
πŸ”— http://arxiv.org/abs/2609.10781v1
πŸ‘₯ Authors: Shuyuan Zhang (possible past Google (United States) affiliation), Zihan Wang (possible past Tsinghua University affiliation), Xiao-Wen Chang, Doina Precup (possible past Deepmind (United Kingdom) affiliation)
Abstract

The integration of graphs with Goal-Conditioned Hierarchical Reinforcement Learning (GCHRL) has received increasing attention, as graphs naturally encode task hierarchies for effective subgoal sampling. However, existing methods often overlook intrinsic connectivity information, failing to fully leverage the underlying topology for efficient learning. Most graph-based GCHRL methods use the graph as a stochastic sampling tool rather than as an environmental model that encodes connectivity and sta...

*Notable papers are those with at least two authors from a "big" AI/ML lab.