📄 Notable* Recent AI/ML arXiv Papers

Last updated just now...

📄 Post-Training Language Models for Gold-Medal Performance in Coding Competitions
🗓️ Published: 9/2/2026
🔗 http://arxiv.org/abs/2609.02849v1
👥 Authors: Aleksander Ficek, Sean Narenthiran, Mehrzad Samadi (possible past Nvidia (United States) affiliation), Somshubra Majumdar (possible past Nvidia (United States) affiliation), Boris Ginsburg (possible past Nvidia (United States) affiliation)
Abstract

Competitive programming has become a key test of large language model reasoning, with international competitions such as IOI and ICPC representing its most challenging settings. We present an end-to-end specialization pipeline combining large-scale problem curation, synthetic reasoning traces, supervised fine-tuning (SFT), and reinforcement learning (RL). Using 22,000 curated problems, we train Nemotron-3-Nano-CC (30B-A3B) with SFT and RL and Nemotron-3-Ultra-CC (550B-A55B) with SFT alone. We fu...

📄 Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems
🗓️ Published: 9/2/2026
🔗 http://arxiv.org/abs/2609.02750v1
👥 Authors: Yihang Chen, Yuxiang Chen, Yuxuan Huang, Meng Fang (possible past Tencent (China) affiliation), Weilin Luo, Jun Wang (possible past Tencent (China) affiliation)
Abstract

Multi-agent LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improve through textual reflection. Despite strong empirical results, these systems lack a unified account of coordination, memory improvement, and the role of external verification. We model orchestrator-worker interaction as a bilevel coordination game: under bounded coupling, the workers' local-update game is an approximate potential game whose equilibrium slack is controlled by decompositi...

📄 ProbeMatchDTI: Probe-Driven Multi-Scale Biochemical Pattern Matching for Drug-Target Interaction Prediction
🗓️ Published: 9/2/2026
🔗 http://arxiv.org/abs/2609.02549v1
👥 Authors: Quan Hao, Mengyue Fan, Zifan Dong, Youru Li, Jianduo Zhao, Lechuan Xu, Hao Zhang (possible past Tencent (China) affiliation), Fei Xia (possible past Stanford University affiliation), Jigang Wang, Chong Qiu, Liguo Zhang
Abstract

Drug-target interaction (DTI) prediction is an important task in AI-driven drug discovery. Although recent biochemical representation learning methods have improved DTI prediction, their passive feature aggregation tends to favor dominant molecular patterns while suppressing weak yet binding-relevant signals, such as functional groups and residue-context patterns, limiting the modeling of multi-scale biochemical correspondences. To address this issue, we propose ProbeMatchDTI, a pattern-probe-dr...

📄 Fine-Grained Anomaly Perception in Wild UGC-Enhanced Images: A Comprehensive Dataset and Difference-Fusion Framework
🗓️ Published: 9/2/2026
🔗 http://arxiv.org/abs/2609.02529v1
👥 Authors: Yan Zhong, Gefei Chen, Qiufang Ma, Zhen Wang, Zhiwei Fan, Lei Shi (possible past Baidu (China) affiliation), Tingting Jiang (possible past Tsinghua University affiliation)
Abstract

Image enhancement and restoration have become standard back-end operations on short-video and social media platforms to boost UGC visual experience. Yet these processes inevitably introduce visual anomalies--especially in faces, texts, and textures--that directly undermine perceptual fidelity and viewer trust. While existing IQA methods perform well on classic distortions, they target holistic quality assessment and fail to capture the specific, localized anomalies caused by enhancement algorith...

📄 Before the Script, Set the Stage: How Worldview Simulation Amplifies Psychologically Grounded Persuasion in Multi-Turn Jailbreaking
🗓️ Published: 9/2/2026
🔗 http://arxiv.org/abs/2609.02414v1
👥 Authors: Siyu Chen, Haoran Wang, Xiaojian Li, Yao Huang, Yinpeng Dong (possible past Tsinghua University affiliation), Wei Xu (possible past Tencent (China) affiliation)
Abstract

Multi-turn jailbreak attacks demonstrate that harmful intent can be distributed across dialogue, yet existing methods obscure what conversational mechanisms drive vulnerability. We introduce BLUEPRINT, a safety-evaluation framework separating a factorized social-influence strategy space from WORLDVIEWSIM, a cross-turn situational context module. Monte Carlo Tree Search optimizes turn-level combinations of 18 theory-grounded influence factors across a four-turn trajectory. Across six frontier mod...

📄 RouteGraph-Mona: Confusion-Aware Routing Fine-Tuning for Mineral Image Classification
🗓️ Published: 9/2/2026
🔗 http://arxiv.org/abs/2609.02282v1
👥 Authors: Jierui Li, Zhiyuan Qi, Hao Zhu (possible past Tsinghua University affiliation), Yufan Liu, Jixian Liu, Shaojie Jiang, Jianda Wang, Yaqi Liu, Xiaotong Li, Wei Wang (possible past University Of Oxford affiliation)
Abstract

Mineral image classification is important for geological exploration and resource development, but it remains challenging due to substantial intra-class variations in appearance and high inter-class visual similarity. Multi-cognitive Visual Adapter (Mona) is a vision-oriented parameter-efficient adapter that adapts pre-trained visual models by tuning only a few parameters. However, Mona statically aggregates responses from multiple scales, limiting its ability to accommodate sample-specific scal...

📄 PGPO: Potential-Guided Policy Optimization for Multi-Turn Agentic Tasks
🗓️ Published: 9/2/2026
🔗 http://arxiv.org/abs/2609.02236v1
👥 Authors: Yuyao Zheng, Haipeng Sun, Junwei Bao, Lemao Liu (possible past Tencent (China) affiliation), Hongfei Jiang, Yang Song (possible past Stanford University affiliation), Dejing Dou (possible past Baidu (China) affiliation)
Abstract

Group-based reinforcement learning (RL) has become an effective paradigm for LLM post-training, but in multi-turn agentic tasks with sparse terminal rewards, it often provides coarse credit for intermediate actions. To obtain more fine-grained credit assignment, recent work such as GiGPO introduces step-level advantages for intermediate actions. However, these step-level signals still rely on the final outcome of each individual trajectory. As a result, actions within failed trajectories can rem...

📄 SkillGLoW: Procedural-Family Skill Consolidation for Self-Improving Agents on Long-Horizon Task Streams
🗓️ Published: 9/2/2026
🔗 http://arxiv.org/abs/2609.02217v1
👥 Authors: Ao Yan, Xin Zhang (possible past Google (United States) affiliation), Jiawei Du, Joey Tianyi Zhou (possible past Tencent (China) affiliation)
Abstract

LLM agents increasingly self-improve by writing and reusing textual skills, kept either as one global document or as a flat pool of per-task entries, though most of the evidence comes from domains with structurally similar tasks. On long-horizon workloads where each task demands a different solution, the two forms fail in opposite ways: the document collapses into generic discipline, while the pool inflates and its entries stay bound to the instance that wrote them. We argue the missing unit of ...

📄 OmegaUse-SOP: SOP Engineering for Professional Computer Use from Human Demonstrations
🗓️ Published: 9/2/2026
🔗 http://arxiv.org/abs/2609.02149v1
👥 Authors: Yixiong Xiao (possible past Tsinghua University affiliation), Lang An, Hucheng Yang, Pinxue Ma, Yongquan Chen, Jingjia Cao, Yusai Zhao, Ting Wang, Ting Liu (possible past Google (United States) affiliation), Siqi Bao (possible past Baidu (China) affiliation), Jingbo Zhou (possible past Baidu (China) affiliation), Hua Wu (possible past Baidu (China) affiliation)
Abstract

Large language models (LLMs) are increasingly evolving from conversational assistants into agents capable of operating external digital environments. Graphical user interface (GUI) agents play an important role in this transition, as many real-world workflows remain accessible only through user-facing software interfaces. However, despite recent progress on general computer-use benchmarks, domain-specific professional standard operating procedures (SOPs) remain challenging for GUI agents because...

📄 READY or Not: Reliable Enterprise Agent Deployment
🗓️ Published: 9/2/2026
🔗 http://arxiv.org/abs/2609.02095v1
👥 Authors: Veronica Chatrath, Bryan Zhu, Jingxuan Fan, George Pu, Soham Dinesh Tiwari, Soham Dan, Ryan Young, Yuan, Li, Yuang Yao, Apaar Shanker, Minglai Yang, Daniel Yue Zhang, Yunzhong He, Ying Liu, Chenguang Wang (possible past Amazon (United States) affiliation), Zhijun Yin, Yuan Xue (possible past Google (United States) affiliation)
Abstract

An AI agent can perform well on benchmarks and still be unsuitable for deployment. Existing AI-agent benchmarks measure whether an agent can complete realistic professional work, whereas enterprise deployment asks a different question: whether an agent can meet a required reliability level, under acceptable human oversight, and at tolerable cost. We introduce Reliable Enterprise Agent Deployment (READY), a framework for qualifying AI agents for deployment on enterprise workflows. READY preserves...

📄 MineTRACE: An Evidence-Grounded Interactive Reasoning System for Mineral Prospectivity
🗓️ Published: 9/2/2026
🔗 http://arxiv.org/abs/2609.02060v1
👥 Authors: Yiran Zhang, Jinwen Liu, Daniel Su, Yisu Chen, Qiang Sun (possible past University Of Toronto affiliation), Chris Gonzalez, Eun-Jung Holden, Marco Fiorentini, Wei Liu (possible past Tsinghua University affiliation), Yihao Ding
Abstract

Mineral exploration requires integrating heterogeneous geochemical, geophysical, and geological evidence, yet existing prospectivity systems often provide only opaque scores or heatmaps. We present MineTRACE, a web-based system for evidence-grounded exploration of eight commodities: Cu, Au, Ni, W, Sn, Co, Ta, and Mn. Users can explore prospectivity maps, query locations or regions, inspect supporting evidence, and interact through natural language. A transparent expert tree, informed by geologic...

📄 Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement
🗓️ Published: 9/1/2026
🔗 http://arxiv.org/abs/2609.01481v1
👥 Authors: Haoyang Yan, Min-Le Su, Hangfan Zhang, Zhanhao Li, Chen Zhang (possible past Peking University affiliation), Shao Zhang, Yang Chen (possible past Tencent (China) affiliation), Lei Bai, Shuyue Hu
Abstract

This paper studies autonomous software development, in which LLM-based coding agents transform high-level requirements into complete, functional, and usable software systems without human intervention. We introduce Harness-of-Harness (HoH), a framework that enables coding agents to continually improve software during autonomous development. HoH operates on existing coding-agent harnesses, and organizes their executions into iterative planning-coding-testing loops. To sustain improvement across l...

📄 EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents
🗓️ Published: 9/1/2026
🔗 http://arxiv.org/abs/2609.01281v1
👥 Authors: Wei Wang (possible past University Of Oxford affiliation), Wenqiao Zhang, Yutong Lin, Yuqian Yuan, Tianwei Lin (possible past Baidu (China) affiliation), Jinhao Mao, Zhenxuan Fan, Mingjian Gao, Yang Dai, Wentong Li, Zheqi Lv, Zheng Dong, Yingjie Niu, Jiaqi Zhu, Jun Xiao, Chao Li (possible past Baidu (China) affiliation), Yueting Zhuang
Abstract

Vision-language-action (VLA) models map visual observations and language instructions directly to robot actions, but long-horizon tasks require more than action prediction. An agent must coordinate perception, planning, execution, progress verification, and recovery as the physical state evolves. An action prediction or a model-generated skill decision does not, by itself, guarantee that the proposed operation is valid in the current state or that its outcome will be verified. We propose Embodie...

📄 oHC: Orthogonal Hyper-Connections on SO(4) via Quaternions
🗓️ Published: 9/2/2026
🔗 http://arxiv.org/abs/2609.02672v1
👥 Authors: Haoqiang Guo, Xuyi Chen, Bo Ke, Yishu Lei, Ziyang Xu, Shikun Feng (possible past Baidu (China) affiliation), Ximen, Wenhan Luo (possible past Tencent (China) affiliation)
Abstract

Hyper-Connections (HC) replace the single residual stream of a Transformer with $n$ parallel ones, mixing them at every layer with a learned $n \times n$ residual matrix. Leaving that matrix unconstrained places no limit on the factor by which the mixing step rescales the residual streams, and that factor compounds across layers, which destabilizes training. Manifold-constrained Hyper-Connections (mHC) address this by restricting the matrix to the doubly stochastic matrices. That caps the factor...

📄 Humanoid Safe Stop via Learned Stoppability Value
🗓️ Published: 9/2/2026
🔗 http://arxiv.org/abs/2609.02358v1
👥 Authors: Junfeng Long, Pieter Abbeel (possible past University Of California, Berkeley affiliation), Koushil Sreenath, Roberto Horowitz, Guanya Shi, C. Karen Liu (possible past Stanford University affiliation)
Abstract

Humanoid robots responding to emergency stop commands typically execute a fixed maneuver, without reasoning about whether a safe stop is actually feasible from the current state. We cast emergency stopping as a reach-avoid problem and propose Safe-Stop, a task-agnostic framework that pairs a learned stop policy with learned stoppability estimators. The estimators are complementary: a stop-probability estimator supervised by the actual outcomes of the fixed stop policy, and a reach-avoidance esti...

📄 TC-Next: Zero-Shot Multimodal Cyclone Forecasting
🗓️ Published: 9/2/2026
🔗 http://arxiv.org/abs/2609.02085v1
👥 Authors: Zhe Wang (possible past Deepmind (United Kingdom) affiliation), Sijie Chen, Yiming Luo, Daehyun Kim (possible past Samsung (South Korea) affiliation), Chien-Yi Chang
Abstract

We present TropicalCycloneNext (TC-Next), a multimodal deep learning model that forecasts tropical cyclone track and intensity at $6$-$24$ h leads by leveraging a foundation model's forecast fields of atmospheric kinematic and thermodynamic fields and GridSat infrared satellite imagery. Trained only on GraphCast forecasts over the Western Pacific (WP), yet reliant only on generic atmospheric variables, TC-Next on GraphCast lowers track error by $15$-$44\%$ and intensity error by a factor of $3$-...

📄 Train What You Deploy: Closing the MLP Reachability Gap in Low-Rank Clone Distillation
🗓️ Published: 9/2/2026
🔗 http://arxiv.org/abs/2609.02006v1
👥 Authors: Wenhui Chen, Zhifeng Li (possible past Tencent (China) affiliation), Jie Zhou (possible past Tsinghua University affiliation), Navan Preet Singh, Madalina Ciobanu, Chenghua Wang, Qingqing Mao, Ritankar Das
Abstract

A compressed student has two shapes that need not agree: the weight it deploys at inference and the weight family its training can reach. We show that a state-of-the-art weight-inheritance distiller, Low-Rank Clone (LRC), deploys a full-width student MLP but ties training to a teacher-induced slice, leaving 62.5-81.4% of each deployed matrix's independent linear degrees of freedom unreachable-paid for at inference, never trainable. Our principle is one line: train what you deploy. From the ident...

📄 From Language to Behavior: Scaling Sequence Transformers for Industrial Recommendation Ranking with Rec-Native Designs
🗓️ Published: 9/1/2026
🔗 http://arxiv.org/abs/2609.01240v1
👥 Authors: Jie Chen (possible past Tencent (China) affiliation), Xiangqian Yu, Yanchao Lian, Tan Lu, Run Yang, Zhengchun Shang, Xing Wang (possible past Tencent (China) affiliation), Cheng Chen (possible past Google (United States) affiliation), Ke Hu (possible past Google (United States) affiliation), Qiang Li, Tianjiu Yin, Xiaobing Liu (possible past Google (United States) affiliation)
Abstract

Scaling Transformers has driven large gains in language modeling, but transplanting this to behavior-sequence modeling in production ranking is challenging: recommendation differs in signal quality, where behavior sequences are noisy, temporally irregular, and sparsely supervised, and in computation asymmetry, where each request scores many candidates against one shared user history under tight latency budgets. We propose ReST, a recommendation-native Transformer scaling framework. For signal qu...

*Notable papers are those with at least two authors from a "big" AI/ML lab.