📄 Notable* Recent AI/ML arXiv Papers

Last updated just now...

📄 Agentic Commerce World: An Auditable and Verifiable Environment for Vibe Commerce
🗓️ Published: 8/3/2026
🔗 http://arxiv.org/abs/2608.02441v1
👥 Authors: Shicheng Fan, Mingdai Yang, Duohao Wang, Canyu Chen, Yongfeng Zhang, Hua Wei (possible past Google (United States) affiliation), Manling Li, Julian Mcauley, Kun Zhang (possible past Google (United States) affiliation), Philip S. Yu (possible past Tsinghua University affiliation), Kejing Yu, Zhiwei Liu
Abstract

In vibe coding, people describe software in natural language and delegate implementation to AI agents. By analogy, vibe commerce allows people to express buying or selling goals in natural language and delegate the corresponding tasks to agents. Commerce, however, requires independently controlled Buyer and Merchant agents to interact in a shared market while preserving their private objectives and distinct authority. We introduce Agentic Commerce World (ACWorld), an environment for evaluating s...

📄 Mamba with Hierarchical Memory: Solving Representation Bottleneck in Long Sequence Modeling
🗓️ Published: 8/3/2026
🔗 http://arxiv.org/abs/2608.02347v1
👥 Authors: Qinwen Wang, Jieping Luo, Aoxiang Qin, Ruoyu Zhao, Jianxiong Tang, Wei Zhang (possible past Tsinghua University affiliation), Zhichao Lu (possible past Google (United States) affiliation), Luziwei Leng
Abstract

Recurrent linear attention models (RLAs) such as Mamba offer efficient linear-time sequence modeling as an alternative to Transformers, yet their fixed-capacity recurrent states limit long-sequence modeling. Drawing inspiration from hierarchical human memory, we propose Hierarchical Memory Mamba (HMM) to address this limitation. Building upon a pre-trained Mamba backbone, HMM integrates a lightweight working memory that extracts slow paragraph-level semantics (PLS) from the fast sensory memory e...

📄 BRiG-AFA: Bellman Risk-to-Go Learning for Non-Myopic Active Feature Acquisition
🗓️ Published: 8/3/2026
🔗 http://arxiv.org/abs/2608.02305v1
👥 Authors: Jiaorong Feng, Qian Li (possible past National University Of Defense Technology affiliation), Ying Li (possible past Meta (United States) affiliation)
Abstract

Active feature acquisition (AFA) asks which unobserved feature to measure next for each test instance under a budget. Greedy rules are easy to train but can overlook context features whose value is realized only through later acquisitions, while reinforcement-learning and generative approaches introduce difficult optimization or conditional-density estimation. We introduce \method, a deployable, supervised alternative that learns a separate candidate-conditioned risk-to-go function for every rem...

📄 SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation
🗓️ Published: 8/3/2026
🔗 http://arxiv.org/abs/2608.02287v1
👥 Authors: Zelin Tan, Yiqun Zhang, Hao Li (possible past Tsinghua University affiliation), Zhiyao Cui, Hejia Geng, Shao Zhang, Hangfan Zhang, Yang Chen (possible past Tencent (China) affiliation), Xiaosong Wang (possible past Nvidia (United States) affiliation), Lilong Wang, Zhenfei Yin, Shuyue Hu, Chen Zhang (possible past Peking University affiliation), Lei Bai
Abstract

Agent skills have become an important mechanism for equipping language-model agents with reusable procedural knowledge. However, providing skills alone does not guarantee that current models can effectively identify, apply, and coordinate them. To improve skill-use capabilities, we introduce SKT, a verified data synthesis pipeline that constructs skill-grounded tasks and executable trajectories from large collections of agent skills. SKT selects suitable single-skill and multi-skill configuratio...

📄 Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories
🗓️ Published: 8/3/2026
🔗 http://arxiv.org/abs/2608.02276v1
👥 Authors: Shuai Shao, Kangning Zhang, Qingyao Li, Shijian Wang, Hao Wang (possible past Tsinghua University affiliation), Wenxiang Jiao, Yuan Lu, Yi Guo, Weiwen Liu, Weinan Zhang (possible past Shanghai Jiao Tong University affiliation)
Abstract

Agents built around large language models continually accumulate interaction trajectories during deployment, yet their behavior typically remains fixed. Beyond updating model weights, these trajectories can improve the agent harness that constructs context, mediates tools, validates actions, and recovers execution. We introduce Harness-R1, the first method, to our knowledge, that makes failure-conditioned, lifecycle-wide editing of an existing executable runtime a learned capability. It post-tra...

📄 PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs
🗓️ Published: 8/3/2026
🔗 http://arxiv.org/abs/2608.02218v1
👥 Authors: Haojie Hu, Chenhao Dang, Yaojia Liu, Hengrui Kang, Conghui He (possible past Tsinghua University affiliation), Weijia Li (possible past Tsinghua University affiliation)
Abstract

Scientific poster construction compresses a long multimodal paper into a readable, editable canvas. Existing systems hide request-level failures by scoring only completed outputs; direct image generation is not element-editable, while coding-agent workflows are costly. PosterMELD is a template-conditioned multi-agent pipeline: capacity-aware slots guide writing before rendering, and deterministic gates plus vision-language model (VLM) review route failures to bounded repair. Each accepted reques...

📄 MEGRAG: Multi-Granular Evidence Graphs for Answer-Aware Multi-Hop RAG
🗓️ Published: 8/3/2026
🔗 http://arxiv.org/abs/2608.02195v1
👥 Authors: Weidong Bao (possible past National University Of Defense Technology affiliation), Yingying Sun, Jun Yang (possible past Tsinghua University affiliation), Yilin Wang (possible past Google (United States) affiliation), Zili Wei, Yubin Bao, Fangling Leng, Minghe Yu, Tiancheng Zhang, Ge Yu
Abstract

Multi-hop question answering is a fundamental challenge in retrieval-augmented generation (RAG), because deriving an answer requires integrating dispersed evidence. Iterative RAG (iRAG) is widely used for this challenge, but existing methods have two limitations. First, most methods still support each reasoning step with single-granularity evidence, making it difficult to balance information density and contextual noise. Second, existing methods often answer the original question only after aggr...

📄 From Profiling to Synthesis: Benchmarking Implicit Behavioral Alignment in Personalized LLM Agents
🗓️ Published: 8/3/2026
🔗 http://arxiv.org/abs/2608.02171v1
👥 Authors: Jiajia Song, Bobo Li, Haiwen Yi, Zibo Ji, Meishan Zhang, Hao Fei, Min Zhang (possible past Tsinghua University affiliation), Mong-Li Lee, Wynne Hsu (possible past National University Of Singapore affiliation)
Abstract

Large Language Models have enabled increasingly capable autonomous agents, yet personalization remains critical for making such agents practically useful. Recent benchmarks have begun evaluating personalization in agents, but they largely rely on static preference snapshots, fixed interaction logs, or question answering over predefined user profiles. Such designs fail to capture the complexity of evolving user preferences and neglect preference-conditioned task execution-a discrepancy we term as...

📄 UniqueSplat: View-conditioned 3D Gaussian Splatting for Generalizable 3D Reconstruction
🗓️ Published: 8/3/2026
🔗 http://arxiv.org/abs/2608.02145v1
👥 Authors: Haixu Song, Xiaoke Yang, Shengjun Zhang, Jiwen Lu (possible past Tsinghua University affiliation), Yueqi Duan (possible past Stanford University affiliation)
Abstract

In this paper, we propose UniqueSplat, a view-conditioned feed-forward 3D Gaussian Splatting model to reconstruct customized 3D radiance fields for each view query. Existing feed-forward methods such as pixelSplat and MVSplat aim to generate fixed Gaussians across all views of each scene by minimizing the error between rendered views and ground-truth images. However, such fixed Gaussians generally render images from all views and lack the ability to adapt to specific viewpoints, as they do not i...

📄 Self-Improving Large Language Models via Progressive Experience Evolution
🗓️ Published: 8/3/2026
🔗 http://arxiv.org/abs/2608.02139v1
👥 Authors: Shijie Ren, Xiting Wang, Meng Li (possible past Meta (United States) affiliation), Yujie Guo, Yunhang Yao, Ziheng Peng, Xunlong Wang, Yuetan Chen, Haoyang Zhou, Yunlong Liang, Fandong Meng (possible past Tencent (China) affiliation)
Abstract

Large language models (LLMs) capable of self-improvement require not only effective policy optimization, but also a principled mechanism for transforming transient interaction experience into persistent model capabilities. Existing self-improvement paradigms remain fragmented: test-time methods can explicitly extract experience but cannot internalize it into model parameters, whereas training-time optimization methods can update model parameters but lack an explicit mechanism for accumulating tr...

📄 IACM-RL: Intent-Aware Context Management and Reinforcement Learning for Complex Tool Invocation under Dynamic Intent Fluctuations
🗓️ Published: 8/3/2026
🔗 http://arxiv.org/abs/2608.02110v1
👥 Authors: Dingwei Zhu, Jiahan Li, Chengjun Pan, Yunxian Yang, Yunbin Zhao, Yunke Zhang, Zhonghang Lu, Zhuohui Sheng, Chenhao Huang, Jiahang Lin, Yajie Yang, Junlin Shang, Shichun Liu, Yuhui Wang, Honglin Guo, Junjie Ye, Xin Guo, Jiazheng Zhang, Ming Zhang (possible past Peking University affiliation), Shihan Dou, Zhiheng Xi, Tao Gui, Qi Zhang (possible past Tencent (China) affiliation), Xipeng Qiu, Xuanjing Huang
Abstract

Executing long-horizon tool invocations in real-world environments is severely challenged by dynamic user intent noise. Existing methods attempt robustness via implicit history scanning or text compression, yet predominantly assume perfect instructions in simplistic scenarios. Inevitably, under fluctuating contexts, obsolete constraints dilute model attention, triggering catastrophic intent deviation and infinite API loops. To resolve this, we propose IACM-RL, a comprehensive framework for robus...

📄 DeGS: A Scalable 3DGS Architecture via Decoupled Workload Parsing and Reorganization
🗓️ Published: 8/3/2026
🔗 http://arxiv.org/abs/2608.02099v1
👥 Authors: Minnan Pei, Gang Li (possible past Tsinghua University affiliation), Zeyu Zhu, Siting Wang, Junwen Si, Zhuoran Song, Yu Feng (possible past University Of California, Berkeley affiliation), Fangxin Liu, Xiaoyao Liang, Jian Cheng
Abstract

3D Gaussian Splatting (3DGS) has emerged as a leading technique for real-time novel view synthesis, yet existing 3DGS accelerators suffer from poor architectural scalability: increasing the number of PEs leads to marginal performance improvement during rendering. We identify that the root cause is the tightly coupled ``checking-while-blending'' dataflow, which exacerbates PE underutilization caused by spatial redundancy from irregular Gaussian coverage and temporal redundancy from asynchronous p...

📄 Fetch-then-Explore: Decoupling Selection from Extraction over a Persistent Workspace for Search Agents
🗓️ Published: 8/3/2026
🔗 http://arxiv.org/abs/2608.02097v1
👥 Authors: Qi Liu (possible past Tencent (China) affiliation), Yiqun Chen, Zidan Chen, Yan Gao, Yi Wu (possible past University Of California, Berkeley affiliation), Yao Hu, Jiaxin Mao, Fengbin Zhu, Tat-Seng Chua
Abstract

Search agents now answer questions that take dozens of searches to settle, yet how such an agent reads a page has drawn far less attention than how it finds one. Nearly all of them use one of two document interfaces, and both tie a page to the moment it is opened. \emph{Visit-and-read} injects a reading of the page into the message history at fetch time, fixing that reading before the agent knows which fact it will need. Stateful \emph{browsing} instead extracts on demand from the page in hand, ...

📄 Recompute or Reuse? Diagnosing and Mitigating Textual Shortcuts in VLM Self-Reflection
🗓️ Published: 8/3/2026
🔗 http://arxiv.org/abs/2608.01930v1
👥 Authors: Wenxiao Fan, Jingling Fu, Fang Li, Luohang Liu, Yu He (possible past Google (United States) affiliation), Lichen Ma, Zhiyang Yu, Weishan Bi, Junshi Huang, Yan Li (possible past Tencent (China) affiliation), Gu Simiu, Kan Li
Abstract

Vision-language models (VLMs) are expected to revise their reasoning when visual evidence changes. Failures to do so are often attributed to insufficient visual attention or contextual inertia, leaving unclear what models reuse instead of recomputing from the current image. We show that evidence-bearing reasoning in a prior chain of thought (CoT) can form a textual shortcut that competes behaviorally with visual recomputation. Across 16 VLMs, a matched counterfactual analysis identifies evidence...

📄 DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents
🗓️ Published: 8/3/2026
🔗 http://arxiv.org/abs/2608.01827v1
👥 Authors: Huanyao Zhang, Jiepeng Zhou, Runhao Zhao, Yanzhe Shan, Jiaoyang Chen, Bowen Zhou, Bo Li (possible past Tencent (China) affiliation), Fang Wang (possible past Tencent (China) affiliation), Jialong Wu, Zhengwei Tao, Lang Mei, Xiaohan Yu, Liyan Liu, Chong Chen, Wentao Zhang (possible past Mila - Quebec Artificial Intelligence Institute affiliation)
Abstract

Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametric knowledge limits their ability to address knowledge-intensive and dynamically evolving open-world problems. To move beyond this limitation, multimodal deep search has emerged as a key direction for open-world information access, evolving from single-turn factual retrieval toward long-horizon, multi-turn search guided by visual evidence. However, existing methods typically confin...

📄 CoNav-UAV: Cooperative Dual-Altitude Aerial Navigation via Stackelberg Learning
🗓️ Published: 8/3/2026
🔗 http://arxiv.org/abs/2608.01802v1
👥 Authors: Junru Song, Wenhao Zhang, Yang Yang (possible past Tencent (China) affiliation), Xuekai Qiu, Feifei Wang, Weien Zhou, Tingsong Jiang, Ying Wen, Yang Li (possible past Google (United States) affiliation), Wen Yao
Abstract

Target-oriented vision-and-language navigation (VLN) on aerial platforms is attracting growing attention for missions such as disaster rescue, infrastructure inspection, and security patrol. In this task, an unmanned aerial vehicle (UAV) needs to locate targets given only a concise description of their appearance and surroundings. This requires global exploration and grounding as well as collision-free close-range approach, two interleaved processes difficult to reconcile within a single agent. ...

📄 PICopilot: An LLM-based Agentic Framework for Assisting Photonic Integrated Circuit Design via Script Generation
🗓️ Published: 8/3/2026
🔗 http://arxiv.org/abs/2608.01791v1
👥 Authors: Xiaohan Jiang, Zeyu Li (possible past Peking University affiliation), Wei Zhang (possible past Tsinghua University affiliation), Jiang Xu
Abstract

The rapid development of photonic integrated circuits (PICs) is shifting the design flow from traditional graphical user interface (GUI)-based methods to script-based methods for higher flexibility, portability, and maintainability. However, script-based design introduces new challenges, requiring designers to possess additional proficiency in tool application programming interfaces (APIs) and programming. It also demands greater effort and time because it is inherently less intuitive and more c...

📄 ProtoAct: Turning Wet-Lab Protocols into Embodied Robotic Actions
🗓️ Published: 8/3/2026
🔗 http://arxiv.org/abs/2608.01690v1
👥 Authors: Zhe Liu, Jiaming Gu, Zhaohui Du, Zhe Wang (possible past Deepmind (United Kingdom) affiliation), Huanbo Jin, Quan Lu, Qi Wang (possible past Tsinghua University affiliation), Ting Xiao, Minting Pan, Dongzhan Zhou
Abstract

Biological wet-lab protocols are written for trained researchers and often leave routine operations, state-dependent conditions, and contextual parameters implicit, making them difficult to translate into robot-executable actions. We present ProtoAct, a structured protocol-grounding framework that converts free-form biological procedures into state-aware, embodiment-ready action sequences. ProtoAct uses ProtoRAG to retrieve manually annotated examples for context-sensitive parsing, employs Refin...

📄 GABench: A Comprehensive Benchmark for Evaluating LLM Agents on Graph Analysis Tasks
🗓️ Published: 8/3/2026
🔗 http://arxiv.org/abs/2608.01684v1
👥 Authors: Jiarui Tan, Zhongjian Zhang, Yabo Guo, Jiawei Liu, Yujie Xing, Muhan Zhang (possible past Meta (United States) affiliation), Cheng Yang (possible past Tsinghua University affiliation), Chuan Shi
Abstract

Large language model (LLM) agents are increasingly capable of planning, using tools, and interacting with external environments. They are typically supported by harnesses, which manage state and coordinate multi-step execution. Graph analysis provides a promising setting for evaluating their agentic capabilities, because it requires agents to access data and execute operations in a graph environment. However, existing graph benchmarks for LLMs provide limited coverage of graph tasks and graph ty...

📄 RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States
🗓️ Published: 8/3/2026
🔗 http://arxiv.org/abs/2608.02508v1
👥 Authors: Yi Yang (possible past Baidu (China) affiliation), Zhennan Chen, Yihong Zhuang, Tiehan Fan, Yinan Chen, Jian Li (possible past Tencent (China) affiliation), Jian Yang, Ying Tai (possible past Tencent (China) affiliation)
Abstract

Learning-based memory systems for self-evolving LLM agents face two tightly coupled challenges. First, trajectory-indexed utilities grow with the interaction history, thereby dispersing limited feedback over an ever-expanding state space. Second, because trajectory-level rewards are jointly assigned to co-retrieved memories, irrelevant experiences may receive misleading utility updates and consequently enter the memory-reward trap. To address these challenges, we introduce Reduced-Order Memory R...

📄 Advancing Relevance Measurement with Vision-Language Models for Web-Scale Search
🗓️ Published: 8/3/2026
🔗 http://arxiv.org/abs/2608.02446v1
👥 Authors: Han Wang (possible past Peking University affiliation), Alex Whitworth, Pak Ming Cheung, Zhenjie Zhang, Krishna Kamath, Xi Chen (possible past University Of California, Berkeley affiliation), Roberto Konow, Kurchi Subhra Hazra
Abstract

Relevance evaluation plays a crucial role in personalized search systems, serving as a guardrail alongside user engagement metrics to ensure that search results align with user queries and intent. While human annotation is the traditional method for relevance evaluation, its high cost and long turnaround time limit its scalability. In this work, we present a VLM-based automated relevance evaluation pipeline deployed within Pinterest Search for online A/B experiments. We rigorously validate the a...

📄 Qwen-CUA: Native Computer Use for (almost) Everything
🗓️ Published: 8/3/2026
🔗 http://arxiv.org/abs/2608.02352v1
👥 Authors: Dunjie Lu, Shuai Bai, Tianyi Bai, Sicheng Fan, Chang Gao, Jian Guan, Feng Hu, Mianqiu Huang, Xingyang Huang, Yizhen Jiang, Yuheng Jing, Dehui Kong, Ning Li, Dayiheng Liu, Shixuan Liu, Zheng Liu, Que Shen, Bowen Wang, Junli Wang, Chencan Wu, Rui Xie, Tianbao Xie, Zhihui Xie, Haiyang Xu, An Yang, Tao Yu (possible past University Of Washington affiliation), Wenzhen Yuan (possible past Massachusetts Institute Of Technology affiliation), Xi Zhang, Zhenru Zhang, Mingkang Zhu, Zhaoqing Zhu, Yizhong Cao, Kai Dang, Binyuan Hui, Kaixin Li, Junyang Lin, Haiquan Wang, Zekun Wang, Yiheng Xu, Fan Yan, Mengqi Yuan, Danyang Zhang, Jiajun Zhang, Zhipeng Zhang, Fan Zhou, Fan Zhou
Abstract

Native computer use offers a general interface for agents to operate almost any software available to people, but requires long-horizon state tracking, large-scale interactive experience, and learning from sparse yet verifiable outcomes. We introduce Qwen-CUA, a native computer-use agent with a 397B-A17B Qwen mixture-of-experts backbone. It observes only screenshots and acts through keyboard and mouse events, without DOM trees, accessibility metadata, or task-specific APIs. Its scaffold maintain...

📄 Start Classifying: Categorical Critics for LLM Reinforcement Learning
🗓️ Published: 8/3/2026
🔗 http://arxiv.org/abs/2608.02181v1
👥 Authors: Zhijian Zhou, Long Li, Xuan Zhang (possible past Meta (United States) affiliation), Zongkai Liu, Yulei Qin, Ke Li (possible past University Of California, Berkeley affiliation), Xing Sun (possible past Tencent (China) affiliation), Xiaoyu Tan, Chao Qu, Yuan Qi
Abstract

Proximal Policy Optimization (PPO) for large language models typically trains its critic by mean-squared-error (MSE) regression on scalar value targets. Although scalar MSE is statistically valid for estimating the conditional expected return, sparse binary rewards in reinforcement learning with verifiable rewards (RLVR) make critic optimization and calibration especially consequential: small value errors directly distort the scalar advantages used by PPO. We study whether a classification-based...

📄 One QK Channel, Many Sources: Guarding Low-Precision Attention Collapse
🗓️ Published: 8/3/2026
🔗 http://arxiv.org/abs/2608.02091v1
👥 Authors: Shuxiao Xie, Shuyang Xie, Yuan Cao (possible past Google (United States) affiliation), Dezhi Ran, Wei Yang (possible past Tencent (China) affiliation), Tao Xie
Abstract

A bfloat16 transformer can train normally for many steps and then collapse abruptly. Distinct low-precision errors can trigger the same failure, leaving unclear whether each source needs its own repair or one shared route can be blocked. We isolate a reproduced GPT-2-class collapse to the streaming-softmax accumulator, where fp32 accumulation repairs it, and use the fault as an assay for moving controlled errors across sources. Errors placed outside attention still drive the same query-key (QK) ...

📄 HarnessCompass: Guiding Automatic Harness Evolution toward Generalizable and Effective Agent Harnesses
🗓️ Published: 8/3/2026
🔗 http://arxiv.org/abs/2608.01918v1
👥 Authors: Luan Zhang, Ruochen Zhou, Dandan Song, Zhengyu Chen, Yuhang Tian, Jun Yang (possible past Tsinghua University affiliation), Huipeng Ma, Chenhao Li, Guangyuan Feng, Xudong Li, Yizhou Jin, Yan Xu (possible past Peking University affiliation)
Abstract

Harness design plays a critical role in agent performance by shaping how large language models (LLMs) perceive, reason over, and act within executable environments. Recent work has proposed automatic harness evolution, which iteratively improves the harness from agent--environment interactions. However, existing methods often overfit to the evolution tasks, rely exclusively on trajectory-derived signals, and optimize harness components jointly, causing interference across components. We propose ...

📄 Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning
🗓️ Published: 8/3/2026
🔗 http://arxiv.org/abs/2608.01743v1
👥 Authors: Li Wang (possible past Tesla (United States) affiliation), Xiaodong Lu, Xiaohan Wang (possible past Baidu (China) affiliation), Jiajun Chai, Wei Lin, Tianhao Peng, Guojun Yin
Abstract

Reinforcement learning (RL) has become a central paradigm for large language model (LLM) post-training, but optimization toward new objectives can degrade capabilities already present in the base model. KL regularization is widely used to mitigate such forgetting by constraining policy drift toward a reference model. However, standard full-policy KL regularization constrains the entire response distribution and may unnecessarily restrict exploration and target-task learning. This raises a natura...

📄 LieStoNet: Learning Lie Symmetries from Spatiotemporal Data for Stochastic Dynamical Systems
🗓️ Published: 8/3/2026
🔗 http://arxiv.org/abs/2608.01582v1
👥 Authors: Shida Liu, Abhishek Gupta (possible past University Of California, Berkeley affiliation), Sumit Sinha, L. Mahadevan (possible past Google (United States) affiliation)
Abstract

Symmetry is central to modern machine learning and physics: invariances and equivariances improve sample efficiency, robustness, and out-of-distribution generalization, while symmetry principles guide scientific modeling. Yet for stochastic dynamical systems the relevant continuous symmetries are rarely known, and symmetry discovery for SDEs has remained essentially unexplored. We introduce \textit{LieStoNet}, an end-to-end, \emph{template-free} framework for discovering Lie-point symmetries of ...

📄 Conformalized Large Language Models under Configuration Shift
🗓️ Published: 8/2/2026
🔗 http://arxiv.org/abs/2608.01460v1
👥 Authors: Yuqicheng Zhu, Jialin Yu, Lin Li, Gengyuan Zhang, Zhen Yang (possible past Tsinghua University affiliation), Steffen Staab, Puneet Dokania, Philip Torr (possible past University Of Oxford affiliation), Jie Tang (possible past Tsinghua University affiliation), Evgeny Kharlamov
Abstract

Conformal prediction (CP) is a distribution-free framework for uncertainty quantification that has recently been adapted to large language models (LLMs), providing prediction sets with finite-sample coverage guarantees under exchangeability. Yet for LLMs, nonconformity scores are often induced by an inference pipeline, not just a fixed model, making them depend not only on the data distribution but also on configurable factors such as the prompt template, decoding parameters, and deployment sett...

*Notable papers are those with at least two authors from a "big" AI/ML lab.