πŸ“„ Notable* Recent AI/ML arXiv Papers

Last updated just now...

πŸ“„ Show-Harness: Just a VLM Agent Can Play Robots
πŸ—“οΈ Published: 9/9/2026
πŸ”— http://arxiv.org/abs/2609.10522v1
πŸ‘₯ Authors: Yanzhe Chen, Zechen Bai, Zhijun Cao, Wenzheng Zeng, Kevin Qinghong Lin (possible past National University Of Singapore affiliation), Yiqi Lin, Guoqiang Liang, Kevin Yuchen Ma, Qiming Huang, Mike Zheng Shou (possible past National University Of Singapore affiliation)
Abstract

Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains challenging. We present Show-Harness, an Embodied Harness that enables VLMs to "play" robots through a compact semantic interface linking intent to action. Show-Harness exposes discrete semantic action units that VLMs can naturally reason over, while embodiment-specific interpreters deterministically ground them into local robot actions, keeping the VL...

πŸ“„ Beyond One-Size-Fits-All: Sample-Adaptive Strategy Routing for Vision Token Pruning in MLLMs
πŸ—“οΈ Published: 9/9/2026
πŸ”— http://arxiv.org/abs/2609.10346v1
πŸ‘₯ Authors: Haiji Liang, Pengfei Zhou, Zhenglin Wan, Wei Wang (possible past University Of Oxford affiliation), Yang You (possible past University Of California, Berkeley affiliation), Wangbo Zhao
Abstract

Multimodal large language models (MLLMs) process hundreds or thousands of visual tokens per image, incurring prohibitive inference costs. While existing vision token pruning methods mitigate this overhead, they implicitly assume that a single fixed pruning strategy can be applied uniformly across all inputs. Our analysis further reveals that ranking pruning methods by average benchmark accuracy conceals substantial sample-wise complementarity: although the average-best strategy excels overall, a...

πŸ“„ GANDR: Claim Auditing for Verifiable Legal Answer Generation
πŸ—“οΈ Published: 9/9/2026
πŸ”— http://arxiv.org/abs/2609.10293v1
πŸ‘₯ Authors: Chen Qian (possible past Shanghai Jiao Tong University affiliation), Yimeng Wang, Yu Chen (possible past Meta (United States) affiliation), Lingfei Wu (possible past Tencent (China) affiliation), Andreas Stathopoulos
Abstract

In high-stakes domains such as legal practice, a language-model answer is only useful to the extent that a reader can verify each claim against the source the system cites. Current grounded-generation pipelines score the answer as a whole, so a correct conclusion can rest on fabricated or loosely matched citations and still score well. Closing this gap requires both a system built for per-claim verification and an evaluation that measures it. We introduce GANDR (Grounded ANswer DRafter), a two-a...

πŸ“„ A Trust-Network-Based Federated Learning Framework for Multi-Center Aging Clock Prediction
πŸ—“οΈ Published: 9/9/2026
πŸ”— http://arxiv.org/abs/2609.10108v1
πŸ‘₯ Authors: Chunxu Zhang, Bo Li (possible past Tencent (China) affiliation), Wenliang Wang, Yang Liu (possible past Tsinghua University affiliation), Di Jiang, Yuan Huang, Yo-Ichi Nabeshima, Akinori Yamamura, Bo Yang (possible past Tencent (China) affiliation), Qiang Yang
Abstract

Aging clocks quantify biological aging and help characterize individual health status. What protein interactions are important for accurate aging clocks, and are they zeroth-order or higher-order? Addressing these questions requires learning from large molecular datasets distributed across medical centers, where privacy constraints prevent centralized data sharing. Federated learning offers a natural solution but faces four challenges in this setting: limited local sample sizes, sparse and direc...

πŸ“„ OntologyAligner: Ontology-Aligned Retrieval and Hierarchy-Guided Large Language Model Reranking for Biomedical Ontology Normalization
πŸ—“οΈ Published: 9/9/2026
πŸ”— http://arxiv.org/abs/2609.10055v1
πŸ‘₯ Authors: Jie Song (possible past Eth Zurich affiliation), Zhichuan Xu, Ziyu Lu, Meng Xiao, Cheng Bi, Yuxin Zhang, Xin Zheng, Xiaoran Li, Qiongfang Cao, Hao Yang (possible past Tencent (China) affiliation), Bairong Shen
Abstract

Biomedical ontology normalization maps free-text expressions to standardized concepts, enabling consistent integration and analysis of biomedical data. This task remains challenging because lexical variation and subtle distinctions among hierarchically related concepts can obscure concept boundaries. We present OntologyAligner, a three-stage framework that combines ontology-aligned retrieval, large language model candidate reranking, and selective hierarchy-guided refinement. We also construct P...

πŸ“„ Structural Process Supervision for Latent Chain-of-Thought Reasoning
πŸ—“οΈ Published: 9/9/2026
πŸ”— http://arxiv.org/abs/2609.09928v1
πŸ‘₯ Authors: Yiqi Li, Xu Chen (possible past Tencent (China) affiliation), Chen Ju, Jiangchao Yao, Zhaoyang Li, Jinsong Lan, Xiaoyong Zhu, Bo Zheng, Yu Wang (possible past Tsinghua University affiliation)
Abstract

Latent reasoning approaches enhance token-level efficiency and robustness by replacing verbose, explicit chain-of-thought (CoT) tokens with compact continuous-space embeddings. However, existing methods lack direct process supervision over these latent embeddings, which often leads to representation collapse and uneven information distribution. To address this, we propose Prototype-Mediated Process Supervision (PMPS), which introduces learnable reasoning prototypes as semantic anchors to provide...

πŸ“„ Kernel-Complexity Edge Sanitization for Training-Free Defense against Structural Graph Attacks
πŸ—“οΈ Published: 9/9/2026
πŸ”— http://arxiv.org/abs/2609.09698v1
πŸ‘₯ Authors: Yaning Jia, Shenyang Deng, Yaoqing Yang (possible past University Of California, Berkeley affiliation), Chiyu Ma, Wenxuan Xu, Soroush Vosoughi (possible past Google (United States) affiliation)
Abstract

Graph Neural Networks (GNNs) have achieved remarkable success across diverse applications, yet they remain highly vulnerable to adversarial attacks that maliciously perturb graph structure. Existing defenses often lack rigorous theoretical grounding, rely on attack-specific heuristics, or require costly retraining procedures such as adversarial training. To address these limitations, we propose Kernel-Complexity Edge Sanitization (KCES), a training-free and model-agnostic framework for defending...

πŸ“„ CityPlanner: A Sandbox Agent for Executable Urban Planning
πŸ—“οΈ Published: 9/9/2026
πŸ”— http://arxiv.org/abs/2609.09578v1
πŸ‘₯ Authors: Wentao Zhang (possible past Mila - Quebec Artificial Intelligence Institute affiliation), Jingyuan Wang, Zetong Zhou, Yifan Yang (possible past Tencent (China) affiliation), Wenrui Wang
Abstract

Urban planning is a real-world spatial optimization problem that requires selecting feasible actions from large candidate spaces under practical objectives such as cost and service quality. Existing optimization and reinforcement learning methods are effective for fixed formulations, but often depend on task-specific representations and constraint handling. We propose \emph{CityPlanner}, a sandbox-agent framework for executable urban planning. CityPlanner introduces \emph{UrbanSandbox}, a unifie...

πŸ“„ ExecCritic: Learn to Test, Test to Improve for Coding Agents
πŸ—“οΈ Published: 9/8/2026
πŸ”— http://arxiv.org/abs/2609.09133v1
πŸ‘₯ Authors: Leitian Tao, Baolin Peng, Haorui Wang, Hang Wang, Hao Cheng (possible past Tencent (China) affiliation), Wenlin Yao, Qianhui Wu, Tao Ge, Sharon Li, Jianfeng Gao (possible past Microsoft (United States) affiliation)
Abstract

Execution feedback can guide coding agents toward correct repository repairs, but only when the tests capture the behavior requested by the issue. Agent-generated tests can encode incomplete or incorrect behavioral targets; when the same trajectory writes both the patch and the test, their errors can agree and create false confidence. We introduce ExecCritic, combining a test--verify--revise scaffold with a role-specific reinforcement learning recipe for training agents within it. The scaffold s...

πŸ“„ PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving
πŸ—“οΈ Published: 9/8/2026
πŸ”— http://arxiv.org/abs/2609.08965v1
πŸ‘₯ Authors: Yuan Gao (possible past Tencent (China) affiliation), Sebastian MΓΌller, Mattia Piccinini, Marc Kaufeld, Yuchen Zhang (possible past University Of California, Berkeley affiliation), Finn Rasmus SchΓ€fer, Qunying Song, Johannes Betz
Abstract

Ensuring the safety of autonomous driving is a critical challenge. Scenario-based testing is a systematic process used to validate Autonomous Driving Systems (ADSs), but it remains a fragmented modular pipeline in which scenario generation, retrieval, modification, ADS execution, and results analysis are performed by separate tools with little interaction. Large Language Model (LLM) agents have shown promise across ADS sub-systems such as perception, planning, and control. However, no prior work...

πŸ“„ Kairos: A Dataset for Fine-Grained Video-Language Modeling over Space, Time, and Dynamics
πŸ—“οΈ Published: 9/8/2026
πŸ”— http://arxiv.org/abs/2609.08755v1
πŸ‘₯ Authors: Ruibo Ming, Lei Sun, Deheng Zhang, He Zhang, Jialu Li, Jian Wang (possible past Baidu (China) affiliation), Zhendong Li, Mengshun Hu, Danda Pani Paudel, Luc Van Gool (possible past Google (United States) affiliation), Jinjin Gu
Abstract

Many emerging video language modeling tasks require systems to move beyond clip-level abstraction and model visual content as it unfolds over extended time horizons. However, most existing video datasets rely on coarse or sparsely aligned supervision, which compresses temporal variation and limits the ability of models to learn reusable representations of continuous visual dynamics. We introduce Kairos, a video dataset for video-language modeling with time-resolved annotations. Kairos consists o...

πŸ“„ Field-level prediction of mid-plane stress tensor fields in concrete target penetration: a cross-velocity graph neural operator surrogate
πŸ—“οΈ Published: 9/9/2026
πŸ”— http://arxiv.org/abs/2609.10032v1
πŸ‘₯ Authors: Wenpu Du, Peng Zhou (possible past Tencent (China) affiliation), Yunlong Xia, Sinuo Xin, Congcong Zhang, Boyang Zhang, Yi Zhang (possible past Google (United States) affiliation), Wenzheng Xu
Abstract

Although the impact resistance of concrete has been studied extensively, a framework linking mesoscale heterogeneity to full-field stress-tensor prediction has been lacking. Data were generated with a full-scale aggregate-resolved LS-DYNA model (projectile diameter 45 mm, mass 2.13 kg, target diameter 500 mm x thickness 200 mm, mesh 10 mm), verified against published penetration experiments (Frew 2006, Hanchak 1992, Forrestal 1996) by configuration similarity. The dataset contains six-component ...

πŸ“„ TEFM: Token-Efficient Faithful Modeling for Structured Data
πŸ—“οΈ Published: 9/9/2026
πŸ”— http://arxiv.org/abs/2609.09552v1
πŸ‘₯ Authors: Zhichao Hou, Lingdao Sha, Xueyu Mao, Yang Liu (possible past Tsinghua University affiliation), Peijie Qiu, Rui Song (possible past Peking University affiliation)
Abstract

In this paper, we solve two fundamental obstacles in applying LLMs to critical domains: token efficiency and faithfulness. To address both constraints jointly, we present TEFM (Token-Efficient Faithful Modeling), a framework designed for structured data analysis in critical domains. TEFM achieves token efficiency by compressing lengthy structured observations into compact Behavioral Code tokens, dramatically reducing token consumption with minimal information loss. Moreover, TEFM enables faithfu...

πŸ“„ Building the Harness Automatically: Self-Play in Code Distills a Text Harness for Black-Box Optimization
πŸ—“οΈ Published: 9/8/2026
πŸ”— http://arxiv.org/abs/2609.09468v1
πŸ‘₯ Authors: Yi Wu (possible past University Of California, Berkeley affiliation), Zheng Ren, Zhiyu Hu, Haochen Wang, Daryl Chang, Li Wei (possible past Google (United States) affiliation), Ting Wang, Zhen Li (possible past Google (United States) affiliation), Pooja Gupta, Nitin Jindal (possible past Google (United States) affiliation), Lukasz Heldt (possible past Google (United States) affiliation)
Abstract

Can an agent learn a numerical search strategy through executable practice and then transfer that strategy as text? We study low-budget black-box optimization, where unaided language models remain well below strong classical optimizers. During development, an agent repeatedly writes and evaluates optimizer programs. It then distills the resulting program and practice record once into a 197-word primary Harness A, which is frozen before evaluation. Harness A reduces Gemini Flash regret by 48\% in...

*Notable papers are those with at least two authors from a "big" AI/ML lab.