πŸ“„ Notable* Recent AI/ML arXiv Papers

Last updated just now...

πŸ“„ CoEM: Empowering Long-Context Reasoning with Commit-on-Evidence Memory
πŸ—“οΈ Published: 9/29/2026
πŸ”— http://arxiv.org/abs/2609.36935v1
πŸ‘₯ Authors: Jingguang Li, Yebo Wu, Zuyi Guo, Kailang Ma, Xianjie Dai, Han Zheng, Benwang Chen, Li Li (possible past Google (United States) affiliation), Can Rong (possible past Peking University affiliation), Heye Huang
Abstract

Long-context reasoning is essential for complex and long-horizon tasks, yet the performance of large language models (LLMs) degrades as context length increases. Recent approaches address this by processing input chunk by chunk while maintaining a bounded textual memory in model context. However, premature information compression can discard critical details essential for subsequent reasoning. In this paper, we introduce Commit-on-Evidence Memory (CoEM), which learns when to convert source evide...

πŸ“„ PrecogUI: Proactive GUI Agents via Pre-cognitive Simulation and Experience Retrieval
πŸ—“οΈ Published: 9/29/2026
πŸ”— http://arxiv.org/abs/2609.36923v1
πŸ‘₯ Authors: Bin Kang (possible past Tencent (China) affiliation), Jiarui Ouyang, Li Jiang (possible past Tencent (China) affiliation), Bin Chen, Zhuotao Tian
Abstract

Existing reactive Graphical User Interface (GUI) agents often fail in long-horizon, dynamic scenarios, where unexpected disturbances trigger attention-diverting and cascading failures. To address this, we propose PrecogUI, a pre-cognitive architecture that shifts the paradigm from reactive execution to proactive decision-making. Specifically, we design a Proactive Experience Pool (PEP), which caches recurring anomaly and success patterns as "state-action-result" tuples in a dual-memory repositor...

πŸ“„ WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents
πŸ—“οΈ Published: 9/29/2026
πŸ”— http://arxiv.org/abs/2609.36887v1
πŸ‘₯ Authors: Bo Mao, Hang He, Linting Wang, Lizhi Lin, Maosen Zhou, Guanming Liu, Jinxiu Liu, Tianyu Huai, Chaoyun Zhang (possible past University Of Edinburgh affiliation), Bingxuan Li, Kepeng Lei, Guanting Dong, Zhou Shao, Rui Zheng, Hang Yan, Jie Zhou (possible past Tsinghua University affiliation), Chengcheng Wan, Tao Gui, Liang He, Xipeng Qiu
Abstract

Recent efforts to scale tool-use post-training have largely centered on the synthesis of executable environments, which constitute only one component of a broader agentic interaction system comprising the environment, task, agent harness, and evaluator. Scaling environments in isolation, however, does not guarantee commensurate gains in model performance, because reliable learning signals depend on coherent interactions among all components of the agentic interaction system. To address this prob...

πŸ“„ When Upstream Messages Override Correct Answers: A Controlled Study of Multi-Agent LLM Collaboration
πŸ—“οΈ Published: 9/29/2026
πŸ”— http://arxiv.org/abs/2609.36855v1
πŸ‘₯ Authors: Yaxin Gong, Gangyi Zhang, Chongming Gao, Leyang Shen, Chenxiao Fan, Jiakai Wang, Dong Wang (possible past Tsinghua University affiliation), Yang Liu (possible past Tsinghua University affiliation), Wenjie Wang, Xiangnan He (possible past National University Of Singapore affiliation)
Abstract

Multi-agent LLM systems rely on message passing among specialized agents to accomplish complex tasks. However, an upstream agent may provide useful information or an incorrect answer that causes a downstream agent to override a correct answer supported by its own evidence. Prior work has not clearly separated the benefits of communication from the damage caused by incorrect messages. We study this problem with controlled experiments across five benchmarks and five receivers, keeping the downstre...

πŸ“„ CAD-Native Transformer Operators for AI-Aided Engineering
πŸ—“οΈ Published: 9/29/2026
πŸ”— http://arxiv.org/abs/2609.36806v1
πŸ‘₯ Authors: Daniel Leibovici, Nikola Borislavov Kovachki, Dawon Ahn, Ruben Ohana, Ira J. S. Shokar, Abouzar Ghasemi, Semih Akkurt, Rishikesh Ranade, Neil Ashton, Jan Kautz (possible past Nvidia (United States) affiliation), Jean Kossaifi (possible past Nvidia (United States) affiliation)
Abstract

Modern engineering systems, from automobiles to aircraft, are designed by using precise, continuous parametric computer-aided design (CAD) models. Evaluating design changes through numerical simulation requires meshing the continuous geometry, a computationally expensive and often brittle process that can require manual intervention and replaces the continuous representation with a discrete approximation. Most neural surrogates accelerate the simulation, but inherit this representation gap by re...

πŸ“„ Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving
πŸ—“οΈ Published: 9/29/2026
πŸ”— http://arxiv.org/abs/2609.36750v1
πŸ‘₯ Authors: Yiming Wang, Yikang Liu, Qingyuan Tian, Xingyu Chen (possible past Tencent (China) affiliation), Zhuosheng Zhang, Zhaopeng Tu (possible past Tencent (China) affiliation), Rui Wang (possible past Tencent (China) affiliation)
Abstract

Self-rewarding reinforcement learning (RL) enables large language models (LLMs) to self-evolve without human labels. Existing ensemble-based methods construct reward references from rollout groups and assign rewards accordingly. However, a response's reward representation also depends on its randomly sampled group context, i.e., the other responses in its group. Using only one group-context realization may miss desired reward signals and provide unreliable guidance for policy optimization. To ad...

πŸ“„ SIPO: Unifying Reinforcement Learning with On-Policy Self-Distillation
πŸ—“οΈ Published: 9/29/2026
πŸ”— http://arxiv.org/abs/2609.36742v1
πŸ‘₯ Authors: Zhenrui Yue, Huimin Zeng, Yueqi Wang, Yaokun Liu, Fengran Mo, Jinghan Zhang, Mung Yao Jia, Gyuseok Lee, Yang Zhang (possible past Tsinghua University affiliation), Na Wei, Dong Wang (possible past Tsinghua University affiliation)
Abstract

Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for improving large language models (LLMs) on various tasks, yet its sparse outcome rewards lack token-level credit assignment for intermediate steps. To address this, on-policy self-distillation (OPSD) leverages a self-teacher with privileged context to provide additional dense learning signals. However, because the self-teacher is often overconfident and imposes excessive penalties on long reasoning trajectori...

πŸ“„ RankBuffer: Efficient Ranking-Based Rewards for Open-Ended Generation
πŸ—“οΈ Published: 9/29/2026
πŸ”— http://arxiv.org/abs/2609.36652v1
πŸ‘₯ Authors: Zixuan Yang, Yiqun Chen, Qi Liu (possible past Tencent (China) affiliation), Wei Yang (possible past Tencent (China) affiliation), Erhan Zhang, Liyi Chen, Qimeng Wang, Yan Gao, Jiaxin Mao
Abstract

Open-ended generation lacks canonical answers, making pointwise rewards difficult to calibrate for group-based reinforcement learning. Directly ranking same-query rollouts provides a more suitable relative reward signal, but existing ranking-based reward methods can incur substantial judging cost. We introduce RankBuffer, which maintains an ordered, query-specific buffer of previously judged responses as a reusable quality scale. Each rollout is first inserted into an anchor interval through an ...

πŸ“„ SAKI: Maximal-Coupling-Routed Teacher Supervision for On-Policy Distillation
πŸ—“οΈ Published: 9/29/2026
πŸ”— http://arxiv.org/abs/2609.36601v1
πŸ‘₯ Authors: Miteto Wei, Xiaohan Wang (possible past Baidu (China) affiliation), Zehao Chen, Jiajun Chai, Sichao Liu, Li Wang (possible past Tesla (United States) affiliation), Haoyuan Xu, Zhaoyu Hu, Wei Lin, Guojun Yin
Abstract

On-policy distillation (OPD) reduces train-test state mismatch by training a student on its own generated trajectories, but weak students may visit teacher-misaligned prefixes where supervision is less representative. We introduce SAKI (Supervision Allocation with KL-constrained Interpolation), which combines a KL-constrained teacher-guided rollout with maximal coupling and reuses realized accept/correction events to route token-level supervision. Accepted positions retain sampled-token reverse-...

πŸ“„ Second-Moment Stochastic Approximation Methods
πŸ—“οΈ Published: 9/29/2026
πŸ”— http://arxiv.org/abs/2609.36600v1
πŸ‘₯ Authors: Tao Jiang (possible past Alibaba Group (China) affiliation), Lin Xiao (possible past Microsoft (United States) affiliation)
Abstract

Classical stochastic approximation methods rely on estimators of the first moment (mean) of a random regression function. We study methods that employ estimators of both the first and the second moments, which include modern deep-learning optimizers such as Adam and Muon as special cases. We derive second-moment stochastic approximation methods through the lens of optimal preconditioning for solving matrix equations, and develop a two-stage framework for their convergence analysis. The first sta...

πŸ“„ Long-Term Memory-Guided Enhancement for Target Perception in Audio-Language Models
πŸ—“οΈ Published: 9/29/2026
πŸ”— http://arxiv.org/abs/2609.36577v1
πŸ‘₯ Authors: Zhenhong Zhou, Xuanyue Zhao, Youji Liu, Yuanhe Zhang, Xiaoyu Ma (possible past Google (United States) affiliation), Lianyu Hu, Yang Liu (possible past Tsinghua University affiliation)
Abstract

Audio large language models (ALLMs) can reason about the content of audio recordings to perform complex tasks. However, these capabilities usually collapse in real-world environments when background noise and competing sources mix the target sound. Inspired by long-term memory in human listening, we propose Long-Term Memory-Guided Audio Enhancement (LTM-AE) to improve selective target perception by refining the audio representations of ALLMs without training. LTM-AE extracts representations in h...

πŸ“„ LIBERO-MAX: Do Robot Policies Adapt When the World Changes?
πŸ—“οΈ Published: 9/29/2026
πŸ”— http://arxiv.org/abs/2609.36518v1
πŸ‘₯ Authors: Yunbei Zhang, Zijian Jin, Yuanzhe Liu, Janet Wang, Xilun Zhang, Yuyou Zhang, Zhenyu Zhang, Daoan Zhang, Shuaicheng Niu, Gen Li (possible past University Of Edinburgh affiliation), Jianfei Yang, Jihun Hamm, Ismini Lourentzou, Weirui Ye, Bo Liu (possible past Meta (United States) affiliation), Peter Stone, Marco Pavone (possible past Stanford University affiliation)
Abstract

Robots must often continue a task after a target moves, the viewpoint shifts, or an obstacle appears, even though their earlier observations and committed actions reflect the previous scene. Many simulation robustness benchmarks fix external conditions at reset, leaving this temporal challenge underexamined. We introduce LIBERO-MAX, a benchmark of 8,000 paired cases spanning eight types of changes to geometry, observations, appearance, clutter, and paths. Each pair compares task execution with a...

πŸ“„ Staircase Policy: Streaming Inference for World-Action Models with Large Action Chunks
πŸ—“οΈ Published: 9/29/2026
πŸ”— http://arxiv.org/abs/2609.36471v1
πŸ‘₯ Authors: Guoheng Sun, Chen Chen (possible past Tencent (China) affiliation), Jin Wang, Ang Li (possible past Google (United States) affiliation), Teresa Lv
Abstract

World-Action Models (WAMs) improve robotic manipulation by conditioning action generation on predicted future observations, but future prediction adds further inference overhead to already expensive iterative action generation. Action chunking can amortize this cost over multiple actions, yet performance degrades over long execution horizons because later actions remain conditioned on stale observations. We introduce STAIRCASE POLICY, a streaming inference and training framework that turns a flo...

πŸ“„ Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression
πŸ—“οΈ Published: 9/28/2026
πŸ”— http://arxiv.org/abs/2609.36322v1
πŸ‘₯ Authors: Xingyu Zhu, Pu, Yi, Ziheng Cheng, Ang Lv, Jing Liu (possible past Baidu (China) affiliation), Lexing Ying (possible past Stanford University affiliation), Yiyuan Ma, Xin Dong (possible past Tsinghua University affiliation)
Abstract

Chunked KV-cache compression reduces the memory and attention costs of long-context inference by compressing windows of consecutive tokens into fewer cache entries at a fixed stride. Such compression also introduces a new positional coordinate: a token's phase, or its position relative to compression-window boundaries. We uncover a systematic asymmetry in models using such compression: the same information can be easy to retrieve at one phase and difficult at another. We call this periodic varia...

πŸ“„ CheatBench: Measuring Reward Gaming in AI Agents
πŸ—“οΈ Published: 9/28/2026
πŸ”— http://arxiv.org/abs/2609.36308v1
πŸ‘₯ Authors: Long Phan, Stephen K. Yang, Jason J. Lim, Mantas Mazeika, Wenyu Zhang (possible past Tencent (China) affiliation), Zheyuan Liu, Richard Ren, Jingxiang Meng, Yaoteng Tan, Weiliang Zhao, Addison Wu, Matei Anghel, Dan Hendrycks (possible past University Of California, Berkeley affiliation)
Abstract

Reinforcement learning has helped AI agents solve increasingly difficult tasks, but high rewards do not always reflect the work users intended. In recent incidents and controlled evaluations across the AI industry, agents trained to maximize reward have accessed unauthorized information, attempted to evade monitoring systems, and even breached sandbox protections to attack external systems. As agents become more capable, this behavior could pose increasingly serious risks. To measure this proble...

πŸ“„ PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents
πŸ—“οΈ Published: 9/28/2026
πŸ”— http://arxiv.org/abs/2609.36199v1
πŸ‘₯ Authors: Vighnesh Subramaniam, Boris Katz, Brian Cheung (possible past University Of California, Berkeley affiliation), Chun-Liang Li, Tomas Pfister (possible past University Of Oxford affiliation), Yale Song
Abstract

Diffusion models can produce striking images and videos, but they still struggle with the compositional details that make a generation faithful to a prompt, such as object counts, attribute binding, spatial relations, and temporally grounded actions. A common way to improve prompt satisfaction is to spend more compute at test time through Best-of-N sampling, but final-sample selection is fixed. Best-of-N can only choose among completed outputs and cannot repair a promising trajectory before it f...

πŸ“„ Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning
πŸ—“οΈ Published: 9/28/2026
πŸ”— http://arxiv.org/abs/2609.36178v1
πŸ‘₯ Authors: Dongwon Jung, Hemanth Neelgund Ramesh, Yifan Wang (possible past Stanford University affiliation), Xiaomin Li, Yuexing Hao, Yu Hu, Muhao Chen, Varun Chandrasekaran, Andrzej Banburski-Fahey, Jaron Lanier (possible past Microsoft (United States) affiliation)
Abstract

Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment of trajectory-level advantages to all policy tokens fails to distinguish consequential decisions from less relevant ones, obscuring which intermediate decisions contributed to success. We introduce ProVer, a framework that targets potentially pivotal decisions for fine-grained credit assignment in agentic reinforcement learning. Given a rollout group...

πŸ“„ Draft in Parallel, Condition Through Depth: Adjacent Causal Injection for Speculative Decoding
πŸ—“οΈ Published: 9/28/2026
πŸ”— http://arxiv.org/abs/2609.36173v1
πŸ‘₯ Authors: Haohui Zhang, Keyu Chen, Haocheng Sun, Weibo Gu, Ruizhi Qiao (possible past Tencent (China) affiliation), Xing Sun (possible past Tencent (China) affiliation), Bo Jiang
Abstract

Parallel speculative drafting generates multiple candidates in one backbone pass, but independent token selection can produce inconsistent continuations that shorten the accepted prefix. Existing methods mostly leave conditional decoding to a lightweight module after the backbone, which limits the flow of predecessor information to successors. Our analysis of DFlash shows that early positions already form recoverable predictions in shallow layers, and that accurate adjacent predecessors help suc...

πŸ“„ Reasoning with Neural Cellular Automata
πŸ—“οΈ Published: 9/28/2026
πŸ”— http://arxiv.org/abs/2609.36126v1
πŸ‘₯ Authors: Mayalen Etcheverry, Pietro Miotti, Aidan Sirbu, Konstantin SchΓΌrholt, Mariia Drozdova, Arna Ghosh, Blaise AgΓΌera Y Arcas (possible past Google (United States) affiliation), James Manyika, Blake Richards, Eyvind Niklasson (possible past Google (United States) affiliation)
Abstract

Modern AI architectures used to solve visual reasoning tasks typically rely heavily on global connectivity and synchronization. As biological systems demonstrate, though, sophisticated computation can be performed in a more decentralized fashion. In this work, we test the reasoning capabilities of Neural Cellular Automata (NCAs), networks of recurrent cells that use strictly local connectivity and asynchronous updates. NCAs have been extensively studied in artificial life experiments, but it is ...

πŸ“„ Unifying Distributional Training for One-Step Visual Generation
πŸ—“οΈ Published: 9/28/2026
πŸ”— http://arxiv.org/abs/2609.35763v1
πŸ‘₯ Authors: Chi Zhang (possible past Peking University affiliation), Haoyang Shi, Yueyi Liu, Ruichuan An, Junkang Zhou, Chang Li, Xiuyuan Lu, Yichi Zhang, Bo Wang (possible past Tencent (China) affiliation), Yuhang Wu, Sen Cui, Miao Liu
Abstract

\emph{Distributional training} provides collective supervision for one-step visual generation by matching real and generated features in frozen representation spaces. We introduce \emph{a unified theoretical framework} that separates distribution modeling from matching discrepancy and connects global objectives to pointwise feature updates through Wasserstein gradient flow. Under this framework, FD-Loss and Gaussian-kernel Drifting are recovered through Gaussian optimal transport and kernel-dens...

πŸ“„ Improving Test-Time Scaling with Adaptive Looped Transformers
πŸ—“οΈ Published: 9/28/2026
πŸ”— http://arxiv.org/abs/2609.35748v1
πŸ‘₯ Authors: Yichen You, Tianyu Fu, Aosong Feng, Xingtai Lv, Xuefei Ning, Ning Ding (possible past Tsinghua University affiliation), Yu Wang (possible past Tsinghua University affiliation)
Abstract

Looped transformers have demonstrated promising parameter efficiency by reusing layers for latent computation. Prior studies compare looped and non-looped models at matched parameters or per-token FLOPs. However, to the best of our knowledge, whether looping improves test-time scaling as outputs grow longer remains underexplored. Through post-training looped transformers, we study the accuracy-compute slope, measured as the accuracy gain per doubling of test-time decoding FLOPs. We find that exi...

πŸ“„ Rethinking Circuit Evaluation: Do Circuits Explain Model Errors?
πŸ—“οΈ Published: 9/28/2026
πŸ”— http://arxiv.org/abs/2609.35686v1
πŸ‘₯ Authors: Li Zhang (possible past University Of Oxford affiliation), Chuqin Geng, Mark Zhang, Chen Yang (possible past Tencent (China) affiliation), Luke Zhang, Haolin Ye, Xujie Si
Abstract

Mechanistic interpretability (MI) aims to explain a model's behaviour through analyzing its internal computations; circuit-based explanations aim to isolate these computations with compact subnetworks validated by ablating the rest of the model. We show that circuits validated this way may fail to recover the underlying mechanism of the model's behaviour by closely reproducing its successful decisions while failing to account for most of its errors. Such explanations should account for the model...

πŸ“„ One Proposal for Every Margin: Zero-Shot Amortized Sequential Importance Sampling for Binary Matrices
πŸ—“οΈ Published: 9/28/2026
πŸ”— http://arxiv.org/abs/2609.35514v1
πŸ‘₯ Authors: Ruishuo Chen, Weijia Li (possible past Tsinghua University affiliation), Xun Wang, Yu Chen (possible past Meta (United States) affiliation), Leheng Cai, Longbo Huang
Abstract

In ecology, psychometrics, and the analysis of social and financial networks, binary matrices are often analyzed conditional on their observed row and column sums, which restricts the problem to a finite sample space of matrices with the same margins. Two fundamental problems are to count this space and to sample uniformly from it. Sequential importance sampling (SIS) addresses both with independent weighted samples and an unbiased count estimator, but its efficiency depends critically on the pr...

πŸ“„ The Hidden Ratio in Adam: Stable Structure, Compression, and Sign Dynamics
πŸ—“οΈ Published: 9/28/2026
πŸ”— http://arxiv.org/abs/2609.35392v1
πŸ‘₯ Authors: Yihe Zhou, Tongtian Zhu, Yingxiao Huo, Satya Prakash Dash, Can Wang (possible past Tsinghua University affiliation), Samuel Kaski, Mingfei Sun (possible past Tencent (China) affiliation)
Abstract

Adam is the default optimizer for training modern deep neural networks, yet its adaptive behavior remains poorly understood due to the complex interaction between its first- and second-moment exponential moving averages (EMAs). We study Adam in the tied-$Ξ²$ regime, where the two EMA decay rates are equal, and show that its adaptive dynamics can be expressed through a transformed ratio with approximately scale-stable behavior. Empirically, this transformed ratio exhibits a stable, heavy-tailed di...

πŸ“„ Scaffold Then Internalize: Representation Injection for Diffusion Transformers
πŸ—“οΈ Published: 9/28/2026
πŸ”— http://arxiv.org/abs/2609.35292v1
πŸ‘₯ Authors: Han Fu, Jiacheng Chen, Baoquan Zhao, Weidong Chen (possible past Tencent (China) affiliation), Wei Liu (possible past Tsinghua University affiliation), Qing Li, Xudong Mao
Abstract

Recent representation alignment (REPA) methods accelerate diffusion transformer training by aligning projections of the transformer's hidden states with representations from pretrained visual encoders. In this work, we explore a reverse and complementary direction to REPA: rather than projecting diffusion representations into the encoder's space, we inject encoder representations into the diffusion transformer, allowing them to actively participate in the denoising process. To this end, we intro...

πŸ“„ Sol-H3: Recursive Self-Improvement for MiniMax-H3 Inference Acceleration on Sol-Engine across Cloud and Edge
πŸ—“οΈ Published: 9/28/2026
πŸ”— http://arxiv.org/abs/2609.35110v1
πŸ‘₯ Authors: Yitong Li, Jincheng Yu, Junsong Chen, Haopeng Li, Shuchen Xue, Haozhe Liu, Ping Luo (possible past Shanghai Artificial Intelligence Laboratory affiliation), Song Han (possible past Stanford University affiliation), Enze Xie
Abstract

Video diffusion models are rapidly scaling and exhibiting enhanced generation capabilities. Among these recent advancements, MiniMax-H3 stands out as a highly capable, production-level open-source model. However, its 33-billion parameters and multi-step iterative denoising process introduce substantial computational overhead. Consequently, their practical production is hindered by generation latency in the cloud deployment like NVIDIA-GB200, alongside strict memory limits that pose further chall...

πŸ“„ EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
πŸ—“οΈ Published: 9/28/2026
πŸ”— http://arxiv.org/abs/2609.35047v1
πŸ‘₯ Authors: Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum (possible past Massachusetts Institute Of Technology affiliation), Adrian Weller (possible past University Of Cambridge affiliation), Zenna Tavares, Tom Silver, Kevin Ellis
Abstract

A robot should be able to learn through experiments how unfamiliar objects behave and interact, then plan with that knowledge. It need not start from scratch: physics engines supply knowledge of motion and contact, but can omit entire mechanisms, such as glue curing, water heating, or wind. We present EMPIRIC, an agent that learns a residual world model: a physics engine extended with code for the missing mechanisms. The learned programs can introduce new forces, constraints, and hidden state, a...

πŸ“„ Beyond Verbalized Confidence: Calibrating Reasoners with Differentiable Readouts
πŸ—“οΈ Published: 9/28/2026
πŸ”— http://arxiv.org/abs/2609.34857v1
πŸ‘₯ Authors: Chenxiao Fan, Chongming Gao, Gangyi Zhang, Leyang Shen, Yaxin Gong, Jiamin Wang, Jiakai Wang, Dong Wang (possible past Tsinghua University affiliation), Yang Liu (possible past Tsinghua University affiliation), Fuli Feng (possible past National University Of Singapore affiliation), Xiangnan He (possible past National University Of Singapore affiliation)
Abstract

Reinforcement learning with verifiable rewards (RLVR) trains reasoning models to produce correct answers, but does not ensure that their stated confidence is calibrated. The resulting models are systematically overconfident. Recent methods train calibration inside the RLVR loop by having the model state a numerical confidence alongside its answer, but they all obtain the confidence by sampling it as text. This choice imposes two costs: a sampled confidence introduces variance and in practice col...

πŸ“„ MW-Nowcast: Six-hour ensemble nowcasting of extreme precipitation
πŸ—“οΈ Published: 9/28/2026
πŸ”— http://arxiv.org/abs/2609.34836v1
πŸ‘₯ Authors: Ning Wang, Zuliang Fang, Weixin Jin, Zhongjian Lv, Shuang Qin, Pengcheng Zhao, Siqi Xiang, Jiang Bian (possible past Baidu (China) affiliation), Haoyi Xiong (possible past Baidu (China) affiliation), Nan Guan, Bin Zhang, Liangjie Zhang, Denvy Deng, Qi Zhang (possible past Tencent (China) affiliation), Matt Corey, Jitu Keshri, Sridhar Iyer, Hongyu Sun, Kit Thambiratnam, Jonathan Weyn, Richard E. Turner (possible past University Of Cambridge affiliation), Haiyu Dong
Abstract

Extending reliable nowcasting of extreme precipitation could provide critical additional time for warnings and emergency response during high-impact events such as flash floods. Radar-based generative machine-learning models have enabled skilful hyperlocal precipitation nowcasting, but accurate prediction of intense precipitation remains confined to the first few hours. Because storm-scale structure is predictable for longer than individual cells, a natural strategy is to predict that structure ...

*Notable papers are those with at least two authors from a "big" AI/ML lab.