πŸ“„ Notable* Recent AI/ML arXiv Papers

Last updated just now...

πŸ“„ BrickBench: Evaluating Agentic Brick Design
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.12452v1
πŸ‘₯ Authors: Peter Kulits, Yiqing Xu, R. Kenny Jones, Cordelia Schmid (possible past Google (United States) affiliation), Jiajun Wu (possible past Massachusetts Institute Of Technology affiliation)
Abstract

We propose BrickBench, a benchmark for agentic text-conditioned LEGO-set design. Given a prompt, an agent is tasked with producing an assembly that not only satisfies semantic and design criteria, but that can also be physically built. To do so, it must select parts from a discrete library and reason jointly about local and global constraints. We score validity, alignment, and design across three settings that vary in scale and part availability. We provide BrickAgent, an environment for coding ...

πŸ“„ Predicting Alignment Generalization with Value Representations
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.12410v1
πŸ‘₯ Authors: Andy Liu (possible past Nvidia (United States) affiliation), Mehar Bhatia, Karolina Stanczak, Mona Diab (possible past Carnegie Mellon University affiliation), Vered Shwartz, Daniel Fried
Abstract

LLM developers post-train their models to exhibit prosocial values and behavioral traits, which are enumerated in an alignment target. However, while recent post-training developments have yielded models that score highly on alignment evaluations, training models on sets of narrow behaviors still influences their behavior across unseen contexts and environments in unexpected ways. In this paper, we establish the task of alignment generalization prediction, i.e., predicting how fine-tuning a mode...

πŸ“„ Looking Inside LLMs: Small-World Connectivity as a Signature of Reasoning Performance
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.12304v1
πŸ‘₯ Authors: Zheng Huang, Sansheng Cao, Enpei Zhang, Weikang Qiu, Elynn Chen, Xiang Zhang, Yaoqing Yang (possible past University Of California, Berkeley affiliation), Rex Ying (possible past Stanford University affiliation), Dawei Zhou, Yujun Yan
Abstract

Understanding large language model (LLM) reasoning requires looking beyond behavioral performance to examine how reasoning ability is reflected in internal organization. Inspired by neuroscience findings linking higher intelligence to stronger small-world organization in functional brain networks, we investigate small-world connectivity as a structural signature of LLM reasoning. We construct functional graphs from attention-head activation similarities and find that a higher small-world index (...

πŸ“„ ReSI: Recursive Safety Improvement toward Resistant and Resilient AI
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.12233v1
πŸ‘₯ Authors: Jingnan Zheng, Dongcheng Zhang, Yi Zhang (possible past Google (United States) affiliation), Ming Zhang (possible past Peking University affiliation), Qiaosheng Zhang, Youbang Sun, An Zhang, Xiangnan He (possible past National University Of Singapore affiliation), Tat-Seng Chua, Xia Hu, Bowen Zhou, Chaochao Lu, Xiang Wang (possible past Tencent (China) affiliation)
Abstract

Recursive self-improvement, the participation of AI systems in improving their own capabilities, is beginning to move from theoretical prospect to practice, posing both challenges and opportunities for safety alignment. Models evolve through frequent updates, and their safety alignment requires continual adaptation to each new checkpoint. Meanwhile, with evolving red-teaming methods exposing new vulnerabilities, safety improvement for each checkpoint needs to mitigate exposed vulnerabilities and...

πŸ“„ PulseBound: Future-Beat State Forecasting Under an Explicit Information Boundary
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.12010v1
πŸ‘₯ Authors: Chenyang Xu, Donglin Xie, Xi Xiang, Xiaoyu Li (possible past Tencent (China) affiliation), Yufan Lu, Jiqiun Gao, Yi Zhao, Xin-Yi Li, Guangpu Zhu, Zijian Wang, Xiwen Yang, Dezhen Wang, Lin Chen, Shenda Hong (possible past Peking University affiliation), Leilei Li
Abstract

Predictive representation learning from photoplethysmography (PPG) can violate causal information access even with causal attention, as normalization, nonlocal transforms, or companion views may depend on withheld samples. We introduce PulseBound, a PPG representation learner combining physiologically structured future-beat prediction with an explicit stored-window information boundary. A content-independent cutoff separates the visible prefix from the prediction target. Prefix-only normalizatio...

πŸ“„ MetaOPD: Meta-Learned Token Weighting for On-Policy Distillation
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.11989v1
πŸ‘₯ Authors: Zipeng Wang, Xinpeng Dong, Yuefan Wang, Pingchen Lu, Xian Wei, Kun Kuang, Fei Wu (possible past Google (United States) affiliation), Zhongxiang Dai, Min Zhang (possible past Tsinghua University affiliation)
Abstract

On-policy distillation (OPD) trains a student on its own generated responses using token-level teacher supervision. However, uniform weighting overlooks differences in token learning value, while existing weighting methods rely on predefined mappings from prediction signals to token weights. These mappings are not learned from the effectiveness of the resulting student updates, limiting their ability to adapt to evolving learning needs. In this paper, we propose MetaOPD, a bilevel optimization f...

πŸ“„ Beyond Visual Enhancement: Adaptive Multi-Context Steering to Mitigate LVLM Hallucinations
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.11907v1
πŸ‘₯ Authors: Shuran Ma, Jiale Li, Yuxin Dong, Shan Zheng, Qingyun Jiang, Xiang Chen (possible past Tencent (China) affiliation), Qi Zhu, Deyi Ji, Yifan Yang (possible past Tencent (China) affiliation), Jianfeng Pan, Yu Tian, Xue Yang
Abstract

Hallucination remains a significant challenge in Large Vision-Language Models (LVLMs). Existing training-free methods generally mitigate hallucinations through contrastive decoding or visual enhancement, often increasing the relative influence of visual evidence during generation. This raises a fundamental question: Can LVLMs dynamically regulate the contributions of different context sources to suppress hallucinations? In this work, we investigate and quantify how LVLMs coordinate multiple cont...

πŸ“„ How Is Automated Research Evaluated? A Survey of Benchmarks and Evaluation Practices
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.11877v1
πŸ‘₯ Authors: Liulei Zhang, Dejing Zhou, Chuyue Huang, Guanhua Chen, Yutong Yao, Lidia S. Chao (possible past Tencent (China) affiliation), Chi Man Vong, Derek F. Wong (possible past Tencent (China) affiliation)
Abstract

Automated research systems support literature synthesis, ideation, experiments, writing, and peer review, but their evaluation is dispersed across tasks, benchmarks, and studies that are difficult to compare directly. We review this literature from the perspective of evaluation design and evidence, covering six targets: literature synthesis, research ideation, executable workflows, scholarly writing and communication, automatic peer review, and end-to-end research. We compare task construction, ...

πŸ“„ From Suppression to Repair: Mitigating Object Hallucination in Large Vision-Language Models via Localized Distribution Alignment
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.11826v1
πŸ‘₯ Authors: Chen Zhao (possible past Stanford University affiliation), Xingping Dong, Jiachun Shi, Liang Peng, Chong Wang (possible past Google (United States) affiliation), Zhen Lei (possible past Beijing Academy Of Artificial Intelligence affiliation), Ran He, Bo Du
Abstract

Object hallucination remains a major obstacle for large vision-language models (LVLMs) to generate reliable content. An intuitive mitigation strategy is to suppress hallucination-related components in hidden representations. However, these components may also contain useful information, and suppressing them can weaken the model's multimodal capabilities. In this paper, we propose ResOT, a training-free method that repairs representations at inference time through localized distribution alignment...

πŸ“„ Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.11794v1
πŸ‘₯ Authors: Haoyu Zhao, Zhengxu Yu, Zhiyuan He, Meng Fang (possible past Tencent (China) affiliation), Rasul Tutunov, Haitham Bou-Ammar, Weilin Luo, Jun Wang (possible past Tencent (China) affiliation)
Abstract

Learning to act in unfamiliar environments requires agents to infer how the world works and revise that understanding as new evidence arrives. Yet limited observations can support multiple world models that explain past interactions but predict different outcomes in unseen states. We introduce Memento 3, building on the Memento series to enable frozen LLM agents to continually learn explicit world models through external memory. The agent maintains a natural-language rulebook as persistent seman...

πŸ“„ Timer-M1: A Multivariate Time Series Foundation Model via Learning Primitives
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.11734v1
πŸ‘₯ Authors: Haoran Zhang, Haixuan Liu, Xingjian Su, Yong Liu, Zhi Chen, Yuxuan Wang (possible past Google (United States) affiliation), Jianmin Wang (possible past Tsinghua University affiliation), Mingsheng Long (possible past Tsinghua University affiliation)
Abstract

We introduce Timer-M1, a pretrained multivariate time series foundation model that learns with primitives for zero-shot forecasting. Across domains, time series share elementary temporal and relational patterns, termed primitives, yet differ in how these primitives manifest and evolve across different contexts. Despite progress in zero-shot and task-general forecasting, existing foundation models may still struggle to generalize to complex real-world scenarios. To this end, we develop a primitiv...

πŸ“„ HI3D 3.0 (Twinkle3D): Object-specific 3D Asset Generation with High Resolution
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.11685v1
πŸ‘₯ Authors: Ziying Li, Shengchu Zhao, Huiang He, Yiyang Chen, Jianwen Huang, Bailin Li, Changhao Li, Jianhui Li, Jie Li, Ruiyang Liu, Yibo Luo, Tengjiao Sun, Pei Tang, Shiwen Wang, Jiaqi Wu, Kang Wu, Kaiqiao Yang, Zherui Yang, Hu Zhang (possible past Google (United States) affiliation), Xuezhi Zhao, Xinhe Zheng, Yukun Li (possible past Baidu (China) affiliation), Heliang Zheng, Rongfei Jia
Abstract

Image-to-3D generation has become increasingly capable of producing objects that closely resemble the input image, and an outstanding challenge is to reproduce the depicted object itself, including the specific geometry that defines it. Inscriptions, brand marks, and repeated structures are frequently distorted or lost, despite being critical to object identity. We present Hi3D 3.0, an image-to-3D generation system targeting object-specific fidelity, with Twinkle3D as its geometry model for gene...

πŸ“„ Harness Evolution Hits a Ceiling: When Weight Training Should Begin
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.11655v1
πŸ‘₯ Authors: Yuan Tian, Bing Hu (possible past Meta (United States) affiliation), Hao Wang (possible past Tsinghua University affiliation), Binghang Lu, Fang Wu
Abstract

Improving a long-horizon LLM agent means evolving the harness around a frozen model or training its weights. We let a self-evolving harness make the system stronger first, then cross seed and evolved harnesses with base and trained weights to learn which gains the trained model keeps and which still need the runtime. We show that the right lever can be read off the agent's failure composition: labelling failed trajectories by the first signal that fires separates process failures (blocked calls,...

πŸ“„ Sera: Semantic Representation Aggregation for Reliable and Interpretable Battery Health Forecasting
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.11567v1
πŸ‘₯ Authors: Jiawei Li, Fang Liu (possible past Massachusetts Institute Of Technology affiliation), Wei Zhang (possible past Tsinghua University affiliation), Zuming Liu, Man-Fai Ng, Zhi Wei Seh
Abstract

Battery state of health (SoH) forecasting is important for battery management, but remains challenging due to nonlinear degradation and heterogeneity across batteries. Existing data-driven approaches primarily use temporal models to learn from numerical battery time series, and higher-level degradation characteristics are often not explicitly represented. These characteristics, however, can provide degradation guidance to support reliable forecasting and make the influence of degradation more in...

πŸ“„ ReTeach: Building a Self-Teacher through Multi-Round Reflection and Retry
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.11529v1
πŸ‘₯ Authors: Yafeng Tang, Hao Li (possible past Tsinghua University affiliation), Hongsheng Yu (possible past Tencent (China) affiliation), Qiang Fu (possible past Tencent (China) affiliation)
Abstract

Self-distillation can improve reasoning without a separately trained, more capable teacher, but its effectiveness depends on how the self-teacher gains an advantage over the student. Conditioning the teacher on reference answers or solutions can provide such an advantage, but this information may be unavailable. Reflection offers a way to derive explicit error diagnoses and revision guidance from self-generated attempts, yet existing reflection-based methods often combine it with reference infor...

πŸ“„ Compactness and Consistency: A Conjoint Framework for Deep Graph Clustering
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.11506v1
πŸ‘₯ Authors: Wei Ju, Siyu Yi, Kangjie Zheng, Yifan Wang (possible past Stanford University affiliation), Ziyue Qiao, Li Shen (possible past Tencent (China) affiliation), Yongdao Zhou, Xiaochun Cao, Jiancheng Lv
Abstract

Graph clustering is a fundamental task in data analysis, aiming at grouping nodes with similar characteristics in the graph into clusters. This problem has been widely explored using graph neural networks (GNNs) due to their ability to leverage node attributes and graph topology for effective cluster assignments. However, representations learned through GNNs typically struggle to capture global relationships between nodes via local message-passing mechanisms. Moreover, the redundancy and noise i...

πŸ“„ BridgeGuard: Explicit Safety Drift for Diffusion-based Autonomous Driving
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.11483v1
πŸ‘₯ Authors: Zhenjun Qiu, Jianing Huang, Dongang Liu, Baiyu Du, Yixun Niu, Hao Yang (possible past Tencent (China) affiliation), Xinyu Huang (possible past Baidu (China) affiliation), Chuan Hu, Shu Liu (possible past Tencent (China) affiliation)
Abstract

Diffusion-based driving planners capture diverse behaviors but can generate unsafe trajectories under distribution shift. We propose BridgeGuard, a safety-constrained diffusion planning method that progressively strengthens a constraint term during denoising to drive intermediate trajectories toward a scene-dependent safety domain. Corrections operate in a low-dimensional curve space, promoting geometric coherence. A learned module, DistanceFieldNet, predicts a time-dependent distance field from...

πŸ“„ Zatom-2: Multitask Pretraining on Atomistic Data for Generative Modeling across Domains
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.11454v1
πŸ‘₯ Authors: Miruna Cretu, Alex Abrudan, Antonia Panescu, Tynan Perez, Rishabh Anand, N. Benjamin Erichson, Michael W. Mahoney (possible past Stanford University affiliation), Samuel Blau, Joseph Jacobson, Rafael GΓ³mez-Bombarelli, Rex Ying (possible past Stanford University affiliation), Tuomas Knowles, Pietro LiΓ², Alex Morehead
Abstract

Unified atomistic modeling has the potential to accelerate discovery in chemistry, materials science, and biology by bridging data-rich chemical domains and data-scarce biological contexts. However, existing generative approaches to atomistic modeling remain highly specialized to scientific disciplines (chemistry vs. biology) or do not leverage both high-volume organic (molecule) and inorganic (material) data for general-purpose pretraining. To this end, we introduce Zatom-2, an atomistic genera...

πŸ“„ Cognition-Oriented Emotion Tracing from Causes to Consequences in Real-World Social Scenes
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.11410v1
πŸ‘₯ Authors: Hao Li (possible past Tsinghua University affiliation), Jinye Zhang, Bobo Li, Mong-Li Lee, Wynne Hsu (possible past National University Of Singapore affiliation), Zheng Wang, Hao Fei, Min Zhang (possible past Tsinghua University affiliation)
Abstract

Affective computing has progressed from categorical emotion recognition to open-ended affective analysis with large multimodal models. Yet affective science describes emotion as an unfolding process shaped by appraisal, regulation, and social interpretation, which remains underexplored computationally. We propose TRACE, a cognition-oriented framework that formalizes an affective episode through three interrelated stages: Condition, Affect, and Effect, integrating observable cues with cognitive f...

πŸ“„ Environmental Feedback Modeling Matters: Rethinking Feedback Treatment in Agentic Hindsight Self-Distillation
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.11384v1
πŸ‘₯ Authors: Hangxi Guo, Fengyuan Liu, Yue Wang, Yuhua Qi, Haoyi Xiong (possible past Baidu (China) affiliation), Fei Sun (possible past Meta (United States) affiliation), Mengnan Du
Abstract

Reinforcement learning is commonly used to train language agents in interactive environments, but cannot be directly applied when rewards are unavailable. Recent methods use environmental feedback as privileged context for hindsight self-distillation, but our analysis suggests that simply conditioning the teacher on feedback is insufficient, motivating us to rethink how environmental feedback is used in agentic self-distillation. Given that environmental feedback contains rich supervision for mo...

πŸ“„ SynCo: Data Synthesis Co-Training for Self-Evolving LLMs via Multi-Agent Reinforcement Learning
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.11345v1
πŸ‘₯ Authors: Wei Yang (possible past Tencent (China) affiliation), Shawn Li, Yuehan Qin, Yawei Wang, Mingxi Wang, Shixuan Li, Tiankai Yang, Jiate Li, Jesse Thomason (possible past University Of Washington affiliation), Xuezhe Ma (possible past Carnegie Mellon University affiliation), Yue Zhao
Abstract

Self-evolving LLM agents promise to improve autonomously through continual interaction and learning, reducing their dependence on manually curated supervision. Realizing this promise requires not only updating the agent, but also evolving its training experience as its capabilities change. However, most existing pipelines rely on static datasets or separately updated synthesis models, causing previously useful tasks to become trivial while overly difficult tasks remain uninformative. This growin...

πŸ“„ ReCast: Attribution-Oriented Step Representation Learning for LLM-Based Agent Systems
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.11334v1
πŸ‘₯ Authors: Weilin Jin, Mingyu Wang, Taiyu Zhu, Ziqi Zhou, Wenbo Li, Haoyang Huang, Nan Duan, Yifan Wu (possible past Carnegie Mellon University affiliation), Ying Li (possible past Meta (United States) affiliation), Zhonghai Wu
Abstract

In LLM-based agent systems, failures can originate from early steps whose effects propagate through subsequent interactions, making their origins difficult to identify. To trace such failures back to their origin, failure attribution has been formulated as the task of identifying the earliest step responsible for the failure. Recent methods leverage LLM internal signals for failure attribution, typically using hidden states as step representations. We therefore conduct an empirical study to eval...

πŸ“„ Mine Odyssey: Benchmarking Spatial Agentic Intelligence in the Wild
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.11328v1
πŸ‘₯ Authors: Yuxuan Cao, Junlong Li, Hao Li (possible past Tsinghua University affiliation), Junxian He (possible past Carnegie Mellon University affiliation)
Abstract

Advances in foundation models are driving efforts to introduce agents to assist people in the physical world. Such agents require agentic spatial intelligence: exploring unfamiliar environments, updating spatial understanding through interaction, and adapting actions based on feedback to sustain progress toward a sequence of goals. Existing benchmarks cover only a limited range of spatial layouts, scales, and traversal requirements. We introduce Mine Odyssey, a benchmark for evaluating agentic s...

πŸ“„ MetaEncoder: Exploring the Limit of Bi-Encoders for Multimodal System One Decision Making with Natural Language Interface
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.11316v1
πŸ‘₯ Authors: Jianpeng Cheng, Guangyu Sun (possible past Peking University affiliation), Aashu Singh, Benyu Zhang, Haixing Dai, Hossein Mansour, Jiangfan Zhang, Shlok Kumar Mishra, Wei Sun (possible past Google (United States) affiliation), Xuanming Cui, Yanli Liu, Qi Guo, Max Xiangjun Fan, Jun Xiao
Abstract

System One models output constrained decisions and probability distributions rather than free-form text generation. While prevailing paradigms rely on structured schema objects to encode state, intent, and candidate choices, we revisit a fully natural language-based System One interface. In this framework, both the user request and each candidate option are expressed in natural language, supported by multimodal (image and video) auxiliary inputs. We introduce MetaEncoder, which fine-tunes a pre-...

πŸ“„ A Balanced Data Diet: Addressing the Exploration Bottleneck in Mega-Scale RL for Robot Control
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.12465v1
πŸ‘₯ Authors: Octi Zhang, Mateo Guaman Castro, Patrick Yin, Ignacio Dagnino, Abhishek Gupta (possible past University Of California, Berkeley affiliation), Rosario Scalise, Byron Boots (possible past Carnegie Mellon University affiliation)
Abstract

General-purpose robots must perform a wide range of tasks from agile locomotion to dexterous manipulation. While sim-to-real reinforcement learning (RL) has proven to be a useful tool for this goal, current RL pipelines depend on engineering-heavy, per-task structural priors such as shaped rewards and demonstrations. Recent work has shown that diverse simulator resets, combined with massively parallel simulation, can alleviate much of this engineering burden on several manipulation problems. How...

πŸ“„ VioLA: Learning Generalist Humanoid Control Policies from Human Data
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.12435v1
πŸ‘₯ Authors: Mert Albaba, Jens Beißwenger, Anna Manasyan, Daniel Marta, Michael J. Black, Wieland Brendel, Andreas Krause (possible past Eth Zurich affiliation), Georg Martius, Martin Riedmiller (possible past Google (United States) affiliation)
Abstract

Teaching a humanoid to follow instructions with its whole body runs into two obstacles. Its action space is large and tightly coupled: legs, arms, and fingers must move together while the robot keeps its balance, which makes joint-level actions hard to learn. And humanoid demonstrations are scarce, so current humanoid generalist policies do not follow new instructions out of the box and are fine-tuned on teleoperated demonstrations of each task before deployment. Human demonstrations exist in fa...

πŸ“„ Learning Kilometer-Scale Weather Prediction with Global-Regional Alignment
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.12401v1
πŸ‘₯ Authors: Guowen Li, Yang Liu (possible past Tsinghua University affiliation), Yujie Wang, Qiuyan Sun, Haoyuan Liang, Juepeng Zheng (possible past Tsinghua University affiliation), Hong Cheng, Haohuan Fu (possible past Tsinghua University affiliation)
Abstract

Kilometer-scale regional weather forecasting is essential for local weather warnings and weather-sensitive decisions. Existing data-driven approaches often rely on numerical forecasts for large-scale guidance or require additional training of global forecasting components. Pretrained global weather models offer an efficient source of large-scale forecasts, motivating their reuse to guide high-resolution regional prediction. However, this coupling requires aligning global and regional representat...

πŸ“„ Perception Test 2026: Challenge Summary and Extension to City-scale Audio-Visual Reasoning
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.12081v1
πŸ‘₯ Authors: Fedor Kitashov, JoΓ£o Carreira (possible past University Of California, Berkeley affiliation), Shiry Ginosar, Dima Damen, Andrew Zisserman (possible past University Of Oxford affiliation), Viorica PΔƒtrΔƒucean
Abstract

Continuing the Perception Test challenge series, we organised the fourth edition as a workshop at the European Conference on Computer Vision (ECCV) 2026 in MalmΓΆ, Sweden. This edition focused on spatial intelligence and featured four different tracks: unified multiple-choice videoQA and grounded videoQA from the original Perception Test benchmark, alongside two new tracks based on city-scale walking-tour videos (KilometerAudio and KilometerVision). In this report, we describe the new benchmarks ...

πŸ“„ TACROSS: An Efficient and Low-Cost Scalable Human Touch System Across Heterogeneous Tactile Sensors for Dexterous Robot Learning
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.11945v1
πŸ‘₯ Authors: Bo Chen (possible past Tencent (China) affiliation), Huanzhang Hu, Junyang Ma, Bo Yue, Fangdi Yu, Haijier Chen, Xianxin Lai, Shuyu Pan, Zhen Yang (possible past Tsinghua University affiliation), Xiaoquan Sun, Wenze Cui, Zhongliang Jiang, Shaopeng Liu, Jiayu Chen
Abstract

Collecting tactile demonstrations on robots is costly and slow, motivating the use of lower-cost human tactile gloves for scalable data collection. However, human capacitive/piezoresistive gloves and robotic tactile sensors differ fundamentally in transduction principle, sensor layout, spatial resolution, and dynamic response, making alignment of raw sensor channels ill-posed. To address this problem, we present TACROSS, a scalable system for learning from human touch and transferring it to robo...

πŸ“„ Beyond Action Entropy: Quotient-Space Exploration for Genome-Scale Metabolic Model Repair
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.11627v1
πŸ‘₯ Authors: Xuan Gong, Hanbo Huang, Wenbin Dai, Jing Wang (possible past Google (United States) affiliation), Lei Bai, Xiang Xiao, Weishu Zhao, Shiyu Liang (possible past Google (United States) affiliation)
Abstract

Repairing scientific models from functional observations differs fundamentally from supervised prediction: feedback may certify a solution without revealing which structural correction is responsible. We study this setting for genome-scale metabolic model (GEM) repair, where multiple reaction edits can explain the same phenotypes and many apparently distinct edits correspond to the same biological mechanism. This many-to-one structure creates a hidden failure mode for conventional exploration: d...

πŸ“„ Smoothing the Top-k Exposure Boundary for Sparse Mixture-of-Experts
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.11575v1
πŸ‘₯ Authors: Yunkai Chai, Tong Zhu (possible past Nvidia (United States) affiliation), Xiaoye Qu, Xuyang Hu, Guanjie Chen, Qipeng Guo, Yu Cheng (possible past National University Of Singapore affiliation)
Abstract

Sparse Mixture-of-Experts models scale parameter capacity efficiently while maintaining a fixed compute budget per token. However, traditional training paradigms enforce a static choice of top-$k$ experts, which converts a continuous routing distribution into a rigid step function. This constraint introduces a brittle boundary where highly competitive experts are arbitrarily separated into full-supervision and zero-feedback zones based on minor score fluctuations. To address this issue, we propo...

πŸ“„ WAM-Cache: Staleness-Bounded KV Reuse for Efficient World Action Models
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.11401v1
πŸ‘₯ Authors: Kai Ding, Yang He, Ruijie Quan (possible past Baidu (China) affiliation), Yi Yang (possible past Baidu (China) affiliation)
Abstract

World Action Models (WAMs) enable generalist robot manipulation by conditioning an action expert on representations from a pretrained video Diffusion Transformer (DiT). In closed-loop control, the video DiT runs at every chunk to encode the current observation into layerwise key-value (KV) pairs that the action expert queries. This prefill dominates the per-chunk computational cost, yet existing training-free accelerations leave it fully dense. We present WAM-Cache, a training-free framework tha...

πŸ“„ Being-M0.7: A Latent World-Action Model for Humanoid Robots
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.11283v1
πŸ‘₯ Authors: Junpeng Yue, Boyuan Li, Yuxuan Wang (possible past Google (United States) affiliation), Zepeng Wang, Yuhui Fu, Feiyang Xie, Yu Zhang (possible past Google (United States) affiliation), Jing Zhang (possible past University Of Washington affiliation), Xianqi Zhang, Weibo Li, Xiaofei Zheng, Yuming Fang, Jiangxing Wang, Zongqing Lu
Abstract

Humanoid loco-manipulation requires coordinated locomotion and manipulation informed by future scene evolution and whole-body motion, yet learning these capabilities is constrained by scarce robot demonstrations. Human video and motion datasets offer scalable supervision, but many contain only video or motion rather than paired video-motion data. Moreover, human motion does not directly specify executable robot actions. We present Being-M0.7, a latent world-action model that transfers visual-mot...

πŸ“„ PageWeaver: KV-Guided Query Unions for Sparse Attention
πŸ—“οΈ Published: 10/8/2026
πŸ”— http://arxiv.org/abs/2610.11201v1
πŸ‘₯ Authors: Zhiyuan Li (possible past Peking University affiliation), Zihan Li, Zefang Yuan, Lei Wang (possible past Baidu (China) affiliation), Hao Wang (possible past Tsinghua University affiliation)
Abstract

Dynamic sparse attention limits the KV pages selected by each query, but a small support does not necessarily yield efficient GPU work. Query unions share page loads and populate Tensor Core tiles; their cost depends on which queries are grouped together. We present PageWeaver, an execution design that uses selected-page affinity to assemble query groups while preserving each query's original support and complete output ownership. A bounded GPU search produces query IDs, and an ID-aware two-CTA ...

*Notable papers are those with at least two authors from a "big" AI/ML lab.