πŸ“„ Notable* Recent AI/ML arXiv Papers

Last updated just now...

πŸ“„ SPADE: Self-Play in Adaptive Synthetic Executable Environments
πŸ—“οΈ Published: 8/19/2026
πŸ”— http://arxiv.org/abs/2608.19197v1
πŸ‘₯ Authors: Bo Liu (possible past Meta (United States) affiliation), Simon Yu, Yiding Jiang (possible past Google (United States) affiliation), Ao Qu, Andrew Zhao, Zichen Liu, Junsu Kim, Zijian Zhou, Seungone Kim, Tongzheng Ren, Mickel Liu, Hanfei Yu, Zhaorun Chen, Weiyan Shi, Paul Pu Liang, Luke Zettlemoyer (possible past University Of Washington affiliation), Yejin Choi (possible past Allen Institute For Artificial Intelligence affiliation), Natasha Jaques (possible past University Of California, Berkeley affiliation)
Abstract

Continuous self-improvement requires an ever-expanding pool of self-generated, diverse, adaptive goals. For language agents, existing training environment pools (hand-curated, statically synthesized, or frozen-verifier) keep the goal distribution fixed as the learner scales. We introduce SPADE (Self-Play in Adaptive Synthetic Executable Environments), a self-play RL framework in which a single LLM plays two roles: an Environment Designer that writes complete, long-horizon training environments a...

πŸ“„ ADEPT: Accelerating Dexterity via Pre-Training and Post-Training using Reinforcement Learning
πŸ—“οΈ Published: 8/19/2026
πŸ”— http://arxiv.org/abs/2608.19182v1
πŸ‘₯ Authors: Jayjun Lee, Jessica Yin, Asif Rana, Nicholas Blauch, Sam Mady, Mohak Bhardwaj, Nima Fazeli, Nathan Ratliff (possible past Nvidia (United States) affiliation), Karl Van Wyk (possible past Nvidia (United States) affiliation), Ankur Handa (possible past Nvidia (United States) affiliation)
Abstract

We introduce Accelerating Dexterity via Pre-Training (ADEPT), a large-scale reinforcement learning (RL) framework for learning sim-to-real transferable dexterity across high degree-of-freedom (DoF) robot embodiments that can solve long-horizon tasks directly from raw visuo-tactile perception. ADEPT pretrains a dexterous policy on a generic object reposing task, then post-trains downstream policies with this pretrained behavior as a prior. ADEPT enables learning new behaviors that are otherwise d...

πŸ“„ Interpretable AI predicts a 2026 summer dry anomaly in central China
πŸ—“οΈ Published: 8/19/2026
πŸ”— http://arxiv.org/abs/2608.19163v1
πŸ‘₯ Authors: Anran Wang, Wen Shi, Yong Luo (possible past Tsinghua University affiliation), Jianbin Huang (possible past Tsinghua University affiliation), Lijuan Chen, Junhu Zhao, Weixin Jin, Huihui Yuan
Abstract

Seasonal precipitation anomalies are largely regulated by atmospheric circulation, which dynamical models predict with greater reliability than precipitation itself. Here, we employ a deep learning model that translates dynamical circulation predictions into precipitation estimates. Predictions initialized from March to May consistently indicate a dry anomaly over central China in summer 2026. Retrospective evaluations revealed higher predictive skill in the analogue years, which also tended to ...

πŸ“„ What is Missing from AI Post-Training AI: An Empirical Analysis
πŸ—“οΈ Published: 8/19/2026
πŸ”— http://arxiv.org/abs/2608.19072v1
πŸ‘₯ Authors: Joy Jia Yin Lim, Xin Huang (possible past Baidu (China) affiliation), Hao Peng (possible past Tsinghua University affiliation), Yaxi Lu, Xin Cong, Zhong Zhang, Maosong Sun (possible past Tsinghua University affiliation), Yankai Lin (possible past Tsinghua University affiliation)
Abstract

Large language model (LLM) agents can now post-train an LLM end-to-end. They can write code, launch training, evaluate checkpoints, and improve downstream performance, raising the prospect of AI-for-AI. We argue that this picture conflates two distinct capabilities: execution-level capability, iterating within a selected training strategy; and strategy-level capability, revising the high-level judgment as experimental evidence accumulates. Analyzing a large corpus of publicly released post-train...

πŸ“„ Harness Continual Learning: Continual Adaptation Beyond Model Parameters
πŸ—“οΈ Published: 8/19/2026
πŸ”— http://arxiv.org/abs/2608.19013v1
πŸ‘₯ Authors: Borui Kang, Jinrui Gu, Junhan Lv, Wenbin Li, Lei Wang (possible past Baidu (China) affiliation), Yang Gao (possible past Tencent (China) affiliation)
Abstract

Continual learning has largely been model-centric, treating model parameters as the state that changes with sequential experience. Modern agents can also adapt through a harness of prompts, memories, tools, skills, and routing rules. Because these contents jointly shape later execution, a harness update can disrupt previously reliable behavior even when the model is frozen. This raises a new question: how can an agent continually improve its state outside the model while retaining behavior acqui...

πŸ“„ MedUAG: Unified Understanding and Generation for Medical Multimodal Models
πŸ—“οΈ Published: 8/19/2026
πŸ”— http://arxiv.org/abs/2608.18937v1
πŸ‘₯ Authors: Zijie Meng, Yuncheng Zhang, Hualiang Wang, Yitian Tang, Xiaotang Gai, Chen Shen (possible past Tencent (China) affiliation), Songtao Jiang, Shaosheng Cao, Jian Wu (possible past Tencent (China) affiliation), Xian Wu (possible past Tencent (China) affiliation), Zuozhu Liu
Abstract

Recent Multimodal Large Language Models (MLLMs) are rapidly evolving into unified understanding and generation (UAG) frameworks. However, extending these unified paradigms to the medical domain is hindered by: the absence of comprehensive training and evaluation benchmarks, and the lack of broadly validated unified medical model. To address these gaps, we present a comprehensive foundation for medical UAG. First, we construct MedUAGCorpus, the largest unified medical understanding and generation...

πŸ“„ SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents
πŸ—“οΈ Published: 8/19/2026
πŸ”— http://arxiv.org/abs/2608.18852v1
πŸ‘₯ Authors: Qingyao Li, Wenxiang Jiao, Shuai Shao, Kangning Zhang, Yuan Lu, Yi Guo, Weiwen Liu, Weinan Zhang (possible past Shanghai Jiao Tong University affiliation), Yong Yu (possible past Shanghai Jiao Tong University affiliation)
Abstract

Agent frameworks increasingly package procedural knowledge as skills: instruction files an agent reads on demand, while public libraries now hold thousands of them. Which skill to read has thus become a decision the policy itself makes in the middle of an episode, yet no existing signal trains it. We show that the default remedy, outcome-rewarded RL over the candidate slate, cannot teach it, for a structural reason we identify and name selector credit starvation: under a broadcast, sequence-leve...

πŸ“„ MorphoGP: A Nonparametric Framework for Predicting Equilibrium Beach Profiles Under Tidal Influence
πŸ—“οΈ Published: 8/19/2026
πŸ”— http://arxiv.org/abs/2608.18558v1
πŸ‘₯ Authors: Xi Wu (possible past Google (United States) affiliation), Yanqing Wei, Hang Yin, Pengze Li, Hongshuai Qi, Xi Chen (possible past University Of California, Berkeley affiliation)
Abstract

The prediction of equilibrium beach profiles under tidal influence is of fundamental importance for sustainable coastal development, informing shoreline protection strategies and managing coastal ecosystems under changing environmental conditions. However, it remains challenging due to the highly nonlinear interactions among wave, tide, and sedimentary processes. Traditional empirical and numerical models often exhibit limited adaptability across diverse coastal environments, with especially pro...

πŸ“„ DART-SD: Diamond-topology Aware Retrieval and Tuning for Self-Distillation of Multi-Turn Tool-Calling Agents
πŸ—“οΈ Published: 8/19/2026
πŸ”— http://arxiv.org/abs/2608.18524v1
πŸ‘₯ Authors: Hangrui Xu, Jiarui Wang, Yang Yang (possible past Tencent (China) affiliation), Chuanbo Zhu, Fangda Chen, Ziqi Wu, Jingming Cai, Yan Song (possible past Tencent (China) affiliation)
Abstract

Equipping Large Language Models (LLMs) with multi-turn tool-calling capabilities is essential for building autonomous agents. However, progress is fundamentally limited by the reliance on full-length trajectory imitation. For tasks involving multiple order-independent sub-goals, the optimal solution space forms a vast combinatorial diamond lattice. Forcing this rich topology into monolithic trajectories causes a severe topological collapse, indiscriminately penalizing valid alternative explorati...

πŸ“„ Coupled-cluster molecular properties across the main group that extrapolate beyond training size
πŸ—“οΈ Published: 8/18/2026
πŸ”— http://arxiv.org/abs/2608.18346v1
πŸ‘₯ Authors: Wenhao He, Xu Chen (possible past Tencent (China) affiliation), Noah Song, Haowei Xu, Tim S. Hindges, Bohan Li (possible past Google (United States) affiliation), Zihan Lin, Yu Yao, Avetik R. Harutyunyan, Fang Liu (possible past Massachusetts Institute Of Technology affiliation), Yao Wang, Hao Tang, Ju Li (possible past Massachusetts Institute Of Technology affiliation)
Abstract

Coupled-cluster theory defines the accuracy standard for molecular electronic-structure properties but scales too steeply for routine application, whereas density-functional theory is affordable yet systematically biased. We resolve this trade-off with a single equivariant network, MEHnet-MG, that predicts an effective one-electron Hamiltonian from one inexpensive B3LYP/def2-SVP calculation and derives a broad suite of properties from it (energy, optical gap, dipole, quadrupole, polarizability, ...

πŸ“„ From Inference to Adaptation: A Unified Optimal Transport View of Vision Language Model
πŸ—“οΈ Published: 8/18/2026
πŸ”— http://arxiv.org/abs/2608.18339v1
πŸ‘₯ Authors: Qi Yu (possible past Tencent (China) affiliation), Zhichen Zeng, Katherine Tieu, Xiyuan Yang, Ruizhong Qiu, Yuchen Yan, Lihui Liu, Yanjun Zhao, Lingjie Chen, Jingrui He, Hanghang Tong (possible past Ibm (United States) affiliation)
Abstract

Vision-language models (VLMs) have demonstrated remarkable zero-shot capabilities yet remain sensitive to real-world distribution shifts during inference. Although significant efforts are devoted to adapting VLMs at test time, they rely heavily on noisy pseudo-labels predicted directly from raw embedding similarities during inference, which are unreliable under distribution shift and mislead the adaptation. To avoid noise amplification, existing works craft coarse-grained surrogate objectives du...

πŸ“„ FedCoRe: Target-Adaptive Completion for Missing Modalities in Healthcare Federated Learning
πŸ—“οΈ Published: 8/18/2026
πŸ”— http://arxiv.org/abs/2608.18311v1
πŸ‘₯ Authors: Holger R. Roth (possible past Nvidia (United States) affiliation), Ziyue Xu (possible past Nvidia (United States) affiliation), Peter Cnudde
Abstract

Federated multimodal models often assume every site has every modality, although hospitals differ in access to EHRs, chest radiographs, and ECGs. We study this setting on a MIMIC-derived respiratory deterioration task with simulated FL clients and introduce FedCoRe (Federated Cross-Modal Representation Completion). FedCoRe learns representation- or logit-space corrections rather than generating synthetic ECGs or CXR images. When a client observes a modality that may be missing at deployment, it ...

πŸ“„ GigaBrain-WBC-0.5: A Behavior World Model for Robust Whole-Body Control with Environment Interaction
πŸ—“οΈ Published: 8/18/2026
πŸ”— http://arxiv.org/abs/2608.18234v1
πŸ‘₯ Authors: Ziyang Cheng, Tianshu Tang, Jinxin Lan, Xinze Chen, Yuhan Gong, Zhichao Liu, Changzhong Wu, Yahao Mao, Zongyan Deng, Mingxuan Ma, Huasen Xi, Yilong Liu, Yutong Wu, Xiaofeng Wang, Yang Wang (possible past Baidu (China) affiliation), Yun Ye, Guan Huang, Xiaojie Jin (possible past National University Of Singapore affiliation), Zheng Zhu, Jiwen Lu (possible past Tsinghua University affiliation)
Abstract

Whole-body motion tracking policies turn a humanoid into a robust control interface: the teleoperator---or an upstream model---only supplies a coarse movement intent, while the low-level policy keeps the robot balanced and physically feasible. Existing trackers deliver this interface only on flat ground: trained in empty scenes, they never learn how contact with terrain and objects reshapes their dynamics, and they attempt to teach the policy to balance under any command by continually enlarging...

πŸ“„ On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification
πŸ—“οΈ Published: 8/18/2026
πŸ”— http://arxiv.org/abs/2608.18066v1
πŸ‘₯ Authors: Qinyuan Ye, Yu Li (possible past Tencent (China) affiliation), Yada Pruksachatkun, Jiaxin Zhang, Chien-Sheng Wu (possible past Salesforce (United States) affiliation)
Abstract

Memory-based self-improving agents--those that learn from an online stream of tasks and improve over time by maintaining a textual memory bank--have shown great promise in recent literature. However, the reliability aspects of these methods have been critically overlooked. In this work, we conduct a comprehensive re-evaluation of two memory-based methods, broadening the scope of evaluation along two axes: (1) including multiple runs to quantify variance, and (2) randomly shuffling the tasks to i...

πŸ“„ Against Political Polarization: A Unified Framework for Tracing Evolving Political Ideologies on Social Media
πŸ—“οΈ Published: 8/18/2026
πŸ”— http://arxiv.org/abs/2608.17987v1
πŸ‘₯ Authors: Yijie Xu, Chao Wang (possible past Google (United States) affiliation), Hui Xiong (possible past Baidu (China) affiliation)
Abstract

The rapid growth of social media has greatly influenced political discourse, highlighting the need to understand individual political ideologies and their temporal dynamics. This task faces challenges such as data scarcity, abundant non-political content, costly and bias-prone manual annotation, and difficulty in modeling future ideological inclinations. To address these issues, we propose TSN4PI, a unified framework for tracking the evolution of political ideologies on social media. It includes...

πŸ“„ Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation
πŸ—“οΈ Published: 8/18/2026
πŸ”— http://arxiv.org/abs/2608.17941v1
πŸ‘₯ Authors: Zhizhao Liu, Zhiliang Tian (possible past Baidu (China) affiliation), Xi Wang (possible past Tsinghua University affiliation), Zhihua Wen, Yihang Xiong, Zhiquan Lai, Dongsheng Li
Abstract

Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models but relies on costly rollout exploration. Assigning the same exploration budget to samples with different difficulty levels is inefficient: easy samples may receive redundant rollouts, whereas difficult but learnable samples may receive too little exploration. Existing adaptive schedulers address this mismatch through curriculum-based sample selection or non-uniform rollout allocati...

πŸ“„ StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
πŸ—“οΈ Published: 8/18/2026
πŸ”— http://arxiv.org/abs/2608.17800v1
πŸ‘₯ Authors: Liya Zhu, Xin Ma, Tao Liu (possible past Baidu (China) affiliation), Haodong Wang, Ge Zhang, Jingzhe Ding, Qingshui Gu, Yongjie Zhong, Jinxiang Meng, Yuan Gao (possible past Tencent (China) affiliation), Yunqiu Zhou, Hao Zhu (possible past Tsinghua University affiliation), Jifeng He, Yongzhi Liao, Xinyi Zhang, Chaoxin Li, Yi Zhu, Xi Lin, Duju Zeng, Xiang Gao, Wen Zhang, Yunyang Wang, Duo Wang, Huan Zhou, Zuo Wang, Jin Chen, Kaiyuan Zhang, Chuqian Yu, Tianhao Yu, Longxiang Liu, Jianbo Xue, Huimin Che, Jiahao Wang, Yujia Qin, Jiaheng Liu, Shen Yan, Xiaolong Chang, Wenhao Huang
Abstract

Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on researcher-selected tasks, leaving uncertain whether such progress extends to the work that real-world users actually demand from AI systems. We introduce \textbf{StartupBench}, an E2E agent benchmark grounded in market-validated AI startup products. Rather than defining tasks from pre-defined assumptions about useful ag...

πŸ“„ Evaluating the Diversity of AI-Generated Content with Diversity Profiles
πŸ—“οΈ Published: 8/18/2026
πŸ”— http://arxiv.org/abs/2608.17731v1
πŸ‘₯ Authors: Xiuyuan Hu, Xuege Hou, Guoqing Liu, Yang Zhao (possible past Google (United States) affiliation), Jieran Li, Dongbiao Sun, JosΓ© Miguel HernΓ‘ndez-Lobato (possible past University Of Cambridge affiliation), Hao Zhang (possible past Tencent (China) affiliation), Xue Liu
Abstract

Diversity is a fundamental criterion for evaluating generative artificial intelligence (AI) systems, yet its measurement remains inherently ambiguous. Existing approaches typically represent generated samples in an embedding space, compute pairwise distances or similarities, and aggregate them into a single scalar score. Such scalar summaries are convenient, but they often encode different inductive biases and may yield contradictory rankings of the same sample sets. In this paper, we argue that...

πŸ“„ Benchmarking Automated Security Patch Backporting: How Far Are We?
πŸ—“οΈ Published: 8/18/2026
πŸ”— http://arxiv.org/abs/2608.17671v1
πŸ‘₯ Authors: Jincheng Yang, Yulong Fu, Chengwei Liu, Lyuye Zhang, Fangyuan Zhang, Bingyang Ren, Yang Liu (possible past Tsinghua University affiliation), Hui Li (possible past Baidu (China) affiliation)
Abstract

Automated security patch backporting is critical for mitigating N-day vulnerabilities. Recent tools report success rates above 80% on their respective datasets. However, these evaluations are often confined to homogeneous environments, such as one repository or specific project versions. Consequently, it remains unclear how well these tools generalize beyond their originally targeted scenarios. We present Porting Benchmark, a curated dataset of 1,234 security patch backporting cases spanning cro...

πŸ“„ Learning Canonical Register Automata over Ordered Data Domains
πŸ—“οΈ Published: 8/19/2026
πŸ”— http://arxiv.org/abs/2608.18765v1
πŸ‘₯ Authors: Yong Li (possible past Tsinghua University affiliation), Qiyi Tang (possible past Tencent (China) affiliation), Di-De Yen
Abstract

Register automata are finite automata equipped with memory that recognize data languages over infinite alphabets. In this work, we investigate active learning algorithms for deterministic register automata (DRAs) over ordered data domains--covering both dense domains, such as the rationals, and non-dense domains such as the integers. We show that the active learning problem for DRAs over both dense and non-dense ordered domains can be treated within a single unified framework. More specifically,...

πŸ“„ Role-Conditioned Sub-Token Routing for Efficient Vision-Language-Action Policies
πŸ—“οΈ Published: 8/19/2026
πŸ”— http://arxiv.org/abs/2608.18410v1
πŸ‘₯ Authors: Wei Jiang (possible past Apple (United States) affiliation), Wei Wang (possible past University Of Oxford affiliation)
Abstract

Vision-Language-Action (VLA) models process long multimodal token sequences, making inference expensive in both memory and computation. Existing efficiency methods mainly reduce visual tokens, but aggressive token pruning becomes fragile because removing a token discards its entire representation. Sub-token compression provides a complementary alternative by retaining more tokens while reducing their value width. However, directly applying sub-token compression to VLA policies is less effective ...

πŸ“„ Debate Training Reduces Reward Hacking in RLAIF
πŸ—“οΈ Published: 8/18/2026
πŸ”— http://arxiv.org/abs/2608.17776v1
πŸ‘₯ Authors: Zachary Kenton, Lili Janzer, Rory Greig, Tian Huey Teh, Kirill Tyshchuk, Jonah Brown-Cohen, Harri Edwards (possible past Openai (United States) affiliation), Senthooran Rajamanoharan, Noah Y. Siegel, Natasha Jaques (possible past University Of California, Berkeley affiliation), Rohin Shah
Abstract

We demonstrate that RL finetuning an LLM using debate, a two-player adversarial game between a generator and a critic adjudicated by a weaker LLM judge, reduces reward hacking compared to a reinforcement learning from AI feedback (RLAIF) baseline. Reward hacking is a central obstacle in RLAIF: as training progresses, the policy learns to exploit systematic errors in its AI judge, degrading task performance, a problem that worsens precisely when the judge is weaker than the policy, the setting mo...

*Notable papers are those with at least two authors from a "big" AI/ML lab.