📄 Notable* Recent AI/ML arXiv Papers

Last updated just now...

📄 Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning
🗓️ Published: 7/22/2026
🔗 http://arxiv.org/abs/2607.20166v1
👥 Authors: Siqian Tong, Xuan Li (possible past Baidu (China) affiliation), Chaozhuo Li, Baolong Bi, Yiwei Wang (possible past Google (United States) affiliation), Yujun Cai, Shenghua Liu, Chengpeng Hao
Abstract

Large Audio Language models (LALMs) have made rapid progress on acoustic understanding, yet they still struggle with fine-grained audio reasoning (e.g., recognizing event order, repetitions and duration). Existing post-training methods heavily rely on expensive external labels or provide only coarse semantic signals. To bridge this gap, we introduce Audio-Zero, the first label-free self-evolution framework in the field of LALMs that improves fine-grained auditory perception and reasoning. Audio-...

📄 SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
🗓️ Published: 7/22/2026
🔗 http://arxiv.org/abs/2607.20145v1
👥 Authors: Dongfang Li, Xiaodong Luo, Ruoyu Sun, Xuhui Chen, Linyuan Qiu, Jian Meng, Zhengxuan Lu, Yiting Wang, Yucheng Xie, Tao Guo, Tianxiang Fang, Jing Li (possible past Tencent (China) affiliation), Sihang Chen, Shihao Hong, Chang Liu, Weihua Dai, Zirong Zeng, Ziwei Zhu, Zhuohan Wang, Zhengjun Yue, Igor Vasilyev, Min Liu, Weijian Sun, Xin Chen (possible past Tencent (China) affiliation), Yingmeng Gao, Jinhua Zhou, Taolue Chen, Chenwei Wu, Dong Zhang (possible past Nvidia (United States) affiliation), Wenlong Jin, Jinmin Xiang, Barkova Maria, Ushakov Anton, Xianfei Jin, Tian Ding, Zhihang Lin, Qian Chen (possible past Shanghai Jiao Tong University affiliation), Linxin Yang, Mingzhe Yang, Bingwei Zhang, Hongzhang Yang, Fangxue Zhang, Shijun Qin, Jie Yu, Cuihua Hu, Tolstykh Vasiliy, Nosov Ivan, Abdullin Amir, Zhichen Zhou, Xin Zhang (possible past Google (United States) affiliation), Zhixiong Ning, Xutong Zhao, Junjie Huang, Jiajun Liu, Weiyan Kong, Zheng Zhang, Wenhan Luo (possible past Tencent (China) affiliation), Lin Hu, Yangbo Guo, Li Zeng, Shihao Zeng, Baotian Hu, Min Zhang (possible past Tsinghua University affiliation), Haizhou Li, Zhiquan Luo
Abstract

Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. While most large-scale LLM training systems are built around GPU-based clusters, this report presents an end-to-end optimization practice on the Ascend NPU SuperPOD. Using the DeepSeek-V4 model family as the target workload, we develop a hi...

📄 Overview of FinMMEval 2026 Task 2: Multilingual Financial Short-Answer Question Answering
🗓️ Published: 7/22/2026
🔗 http://arxiv.org/abs/2607.19867v1
👥 Authors: Zhuohan Xie, Xueqing Peng, Georgi Georgiev, Dimitar Dimitrov, Yuyang Dai, Rania Elbadry, Vanshikaa Jani, Lingfei Qian, Fan Zhang, Jimin Huang, Jiahui Geng, Yankai Chen, Ye Yuan (possible past Carnegie Mellon University affiliation), Haolun Wu, Yuxia Wang, Ivan Koychev, Veselin Stoyanov (possible past Meta (United States) affiliation), Mingzi Song, Yu Chen (possible past Meta (United States) affiliation), Xue Liu, Preslav Nakov (possible past Tencent (China) affiliation)
Abstract

FinMMEval 2026 Task 2 evaluates short-answer financial question answering over multilingual evidence. Each final-test item pairs an English question with financial statements and news in English, Chinese, Japanese, Spanish, and Greek. Participating systems submit one concise answer per item in JSONL format. The final-test set contains 256 items, split evenly between easy and expert tiers; each tier contains four question templates instantiated over 32 company-report groups. Gold answers were wit...

📄 DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations
🗓️ Published: 7/22/2026
🔗 http://arxiv.org/abs/2607.19865v1
👥 Authors: Jiazhen Jiang, Boxi Cao, Lingyong Yan, Yaojie Lu, Hongyu Lin, Shuaiqiang Wang (possible past Baidu (China) affiliation), Dawei Yin (possible past Baidu (China) affiliation), Xianpei Han (possible past Tencent (China) affiliation), Le Sun
Abstract

As autonomous agents rapidly evolve, their ability to reliably manipulate ubiquitous digital documents has become critical for enabling general-purpose AI assistants and automating complex workspace workflows. In this paper, we introduce DocOps, a deterministically verifiable evaluation framework underpinned by a hierarchical taxonomy that deconstructs document operations inspired by real-world practices into atomic dimensions and escalating workflow complexities. Based on DocOps, we systematica...

📄 Overview of FinMMEval 2026 Task 1: Multilingual Financial Multiple-Choice Question Answering
🗓️ Published: 7/22/2026
🔗 http://arxiv.org/abs/2607.19856v1
👥 Authors: Zhuohan Xie, Yuyang Dai, Rania Elbadry, Vanshikaa Jani, Georgi Georgiev, Dimitar Dimitrov, Fan Zhang, Xueqing Peng, Lingfei Qian, Jimin Huang, Jiahui Geng, Yankai Chen, Ye Yuan (possible past Carnegie Mellon University affiliation), Haolun Wu, Yuxia Wang, Ivan Koychev, Veselin Stoyanov (possible past Meta (United States) affiliation), Mingzi Song, Yu Chen (possible past Meta (United States) affiliation), Xue Liu, Preslav Nakov (possible past Tencent (China) affiliation)
Abstract

FinMMEval 2026 Task 1 evaluates multilingual financial multiple-choice question answering in English, Chinese, Arabic, and Hindi. The task tests whether systems can select the correct answer to finance questions involving domain terminology, numerical interpretation, and conceptual financial reasoning across languages and scripts. The final-test set contains 800 questions, with 200 questions per language; gold answers were withheld during submission, and each language was ranked independently by...

📄 Beyond Fail-to-Pass: Iterative Hardening of Co-Generated Bug Reproduction Tests and Fixes
🗓️ Published: 7/22/2026
🔗 http://arxiv.org/abs/2607.19843v1
👥 Authors: Yuhao Tan, Zhibang Yang, Fangkai Yang, Yuan Yao (possible past Tsinghua University affiliation), Yu Kang, Lu Wang (possible past University Of Washington affiliation), Pu Zhao, Xin Zhang (possible past Google (United States) affiliation), Xiaoxing Ma, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang
Abstract

Large language models (LLMs) have made automated program repair (APR) increasingly practical for real-world bugs, but repairing directly from bug reports remains underconstrained. Bug reproduction tests (BRTs) help close this gap by turning a bug report into an executable, bug-specific signal that can guide repair and validate candidate patches. Existing work has therefore studied BRT generation as a core subproblem in APR and mainly evaluates a generated BRT using the fail-to-pass (F->P) criter...

📄 Convergence-Latency-Aware Adaptive Modulation and Resource Allocation in RIS-Assisted Wireless Federated Learning
🗓️ Published: 7/22/2026
🔗 http://arxiv.org/abs/2607.19759v1
👥 Authors: Liwei Wang (possible past Tencent (China) affiliation), Wen Chen, Jun Li, Qingqing Wu, Ming Ding (possible past Tsinghua University affiliation), Xusheng Zhu, Qiong Wu
Abstract

Federated learning (FL) over wireless networks suffers from significant training latency and degraded convergence due to unreliable wireless transmission, especially under blocked propagation environments. Although reconfigurable intelligent surfaces (RISs) can improve communication reliability, existing wireless FL studies rarely characterize the trade-off between learning convergence and communication delay under modulation-dependent transmission errors. In this paper, we consider a wireless F...

📄 CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents
🗓️ Published: 7/21/2026
🔗 http://arxiv.org/abs/2607.19338v1
👥 Authors: Qijia He, Jiayi Cheng, Chenqian Le, Rui Wang (possible past Tencent (China) affiliation), Xunmei Liu, Yixian Chen, Jie Mei, Zhihao Wang, Xupeng Chen, Yuhuan Chen, Tao Wang (possible past Stanford University affiliation)
Abstract

Coding agents increasingly operate in executable environments where a failed attempt produces actionable feedback rather than merely an incorrect answer. Existing cost-aware systems typically treat such failures as cascade decisions: try a cheap model first, then escalate hard cases to a stronger and more expensive model. In coding, however, execution feedback can also make further cheap-model recovery worthwhile, raising a budgeted deployment question: when should an agent spend more cheap comp...

📄 Provable diffusion-based posterior sampling for linear inverse problems via DDIM
🗓️ Published: 7/21/2026
🔗 http://arxiv.org/abs/2607.19333v1
👥 Authors: Yuchen Jiao, Na Li (possible past Tencent (China) affiliation), Changxiao Cai, Yuxin Chen, Gen Li (possible past University Of Edinburgh affiliation)
Abstract

Diffusion-based methods have achieved remarkable empirical success in solving inverse problems. However, many existing posterior samplers either lack rigorous theoretical guarantees or incur substantial computational overhead. We propose a simple and efficient algorithm, called \pddim, for solving linear inverse problems with diffusion priors via a DDIM-type sampler. Our method requires only lightweight, coordinate-wise modifications to the standard DDIM update, while explicitly incorporating th...

📄 ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
🗓️ Published: 7/21/2026
🔗 http://arxiv.org/abs/2607.19191v1
👥 Authors: Fan Jiang (possible past Shanghai Jiao Tong University affiliation), Zhaoxu Sun, Mengchao Wang, Ziyu Zhu, Chiyu Wang, Yunpeng Zhang, Wenlin Liu, Yun Wang, Xue Zheng, Rui Sun, Junfeng Ni, Hongyu Pan, Zhongxu Sun, Fei Yu, Zengye Ge, Mengmeng Du, Nianfei Fan, Mingchao Sun, Yu Liu, Yongchang, Yanqing Zhu, Jiahang Wang, Ning Ying, Yuze Xuan, Di Yang, Zhicheng Liu, Zhe Gao, Tingbing Xu, Jiacheng Sui, Wenjin Yang, Junnan Lai, Shufeng Liu, Yuan Liu (possible past Google (United States) affiliation), Zheng Zhou (possible past Tencent (China) affiliation), Yingliang Peng, Dawei Cao, Kaifeng Sheng, Yuxiang Cai, Fei Lu, Mu Xu, Ning Guo
Abstract

We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-source data infrastructure spanning AAA games, simulation engines, and internet videos to learn controllable world dynamics. WorldExplorer performs agent-driven collection guided by training feedback, while a unified pipeline applies 14 deterministic quality checks, VLM-based assessment, and synchronized action and text annotation. We progressively distill a ...

📄 DAIS: Dependency-Aware Intermediate QA Supervision for Complex Reasoning
🗓️ Published: 7/21/2026
🔗 http://arxiv.org/abs/2607.19088v1
👥 Authors: Yu Wang (possible past Tsinghua University affiliation), Ming Fan, Xicheng Zhang, Zhiyong Li, Zhihu Wang, Caiyue Xu, Dahai Hu, Ting Liu (possible past Google (United States) affiliation)
Abstract

Chain-of-thought (CoT) supervision exposes intermediate rationales, but flat rationale targets usually optimize a single reasoning sequence and provide limited supervision on how local conclusions should support later decisions. We introduce Dependency-Aware Intermediate QA Supervision (DAIS), a training-time framework that converts filtered teacher rationales into stage-level QA records. Each intermediate record predicts a local answer conditioned on the previous states needed for that decision...

📄 Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
🗓️ Published: 7/21/2026
🔗 http://arxiv.org/abs/2607.19064v2
👥 Authors: Xinjie Zhang, Peng Zhang, Shicheng Zheng, Jinghao Guo, Zhaoyang Jia, Yifei Shen, Xun Guo, Yuxuan Luo (possible past Baidu (China) affiliation), Jiahao Li, Wenxuan Xie, Fanyi Pu, Xiaoyi Zhang (possible past Apple (United States) affiliation), Kaichen Zhang, Zongyu Guo, Tianci Bi, Dongnan Gui, Zhening Liu, Zimo Wen, Zihan Zheng, Senqiao Yang, Xiao Li, Jinglu Wang, Bin Li, Yan Lu
Abstract

Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. The stack is built from two co-designed components: Mage-VAE, a lightweight high-fidelity latent tokenizer, and a Native-Resolution Multimodal Diffusion Transformer trained with rectified flow matching. Mage-VAE uses one-step diffusion-style encoding and decoding w...

📄 FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling
🗓️ Published: 7/21/2026
🔗 http://arxiv.org/abs/2607.19038v1
👥 Authors: Jialong Zuo, Haotong Zuo, Shiwei Zhang, Xiang Wang (possible past Tencent (China) affiliation), Chen Li (possible past Tencent (China) affiliation), Nong Sang, Changxin Gao, Xiang Bai
Abstract

Translating novels into films poses a grand challenge for generative artificial intelligence, requiring conversion of abstract literary prose into long-form, multi-scene visual narratives. While current video generation models excel at short, single-scene clips within narrow temporal and spatial contexts, novel-to-film generation operates in a more complex regime, demanding long-duration content across diverse scenes with dynamically evolving entity states. To address this, we formalize novel-to...

📄 Local Causal Structure Learning in the Presence of Latent Variables and Selection Bias
🗓️ Published: 7/22/2026
🔗 http://arxiv.org/abs/2607.19866v1
👥 Authors: Zheng Li, Hao Zhang (possible past Tencent (China) affiliation), Ruxin Wang, Ruichu Cai, Kun Zhang (possible past Google (United States) affiliation), Feng Xie
Abstract

Discovering the direct causes and effects of a target variable from observational data is a fundamental problem in causal discovery, with broad applications in domains such as gene regulatory analysis and biomedical research. Existing causal discovery methods either learn a global causal structure, which incurs substantial computational cost, or assume the absence of latent variables and selection bias, assumptions that are often violated in real-world settings. Motivated by these challenges, we...

📄 Hypothesis-and-Refinement Learning of Organic Structures from Multimodal Spectroscopic Data
🗓️ Published: 7/22/2026
🔗 http://arxiv.org/abs/2607.19816v1
👥 Authors: Chengchun Liu, Zhiyuan Yan, Li Yuan (possible past National University Of Singapore affiliation), Hao Li (possible past Tsinghua University affiliation), Boxuan Zhao, Yonghong Tian (possible past Peking University affiliation), Bartosz A. Grzybowski, Fanyang Mo
Abstract

Determining molecular structures from spectroscopic data remains fundamentally challenging because the inverse problem is intrinsically underdetermined: individual spectra are sparse, low-dimensional, and encode only partial structural evidence relative to the vast space of possible molecules. We address this challenge by formulating automated structure elucidation as a scalable hypothesis-refinement paradigm that tightly integrates spectral evidence with large-scale molecular priors. To supply ...

📄 Visual Semantic Decoding of Electrocorticography from Video Stimuli using End-to-End Deep Learning
🗓️ Published: 7/21/2026
🔗 http://arxiv.org/abs/2607.18923v1
👥 Authors: Stella Ho, Joel Villalobos, Joseph West, Jingyang Liu (possible past Google (United States) affiliation), Weijie Qi, Haruhiko Kishima, Ryohei Fukuma, Takufumi Yanagisawa, Sam E. John, David B. Grayden (possible past Google (United States) affiliation)
Abstract

ECoG-based visual semantic decoding enables inference of semantic interpretation of visual perception from complex, noisy brain activity. This study examines the feasibility of visual semantic decoding using an end-to-end deep learning framework using electrocorticography (ECoG). Specifically, the decoding task is to predict visual categories from video stimuli using time-series neural inputs. A previously collected ECoG dataset from participants ($n=17$) with drug-resistant epilepsy is used for...

*Notable papers are those with at least two authors from a "big" AI/ML lab.