πŸ“„ Notable* Recent AI/ML arXiv Papers

Last updated just now...

πŸ“„ ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding
πŸ—“οΈ Published: 7/27/2026
πŸ”— http://arxiv.org/abs/2607.24743v1
πŸ‘₯ Authors: Hangjie Yuan, Yichen Qian, Zhiwei Tang, Xianzhe Xu, Lirong Wu, Sicheng Yang, Jinwang Wang, Pengju Wang, Zhitao Zeng, Yizeng Han, Yan Xing, Shengxuan Luo, Tao Feng, Qing Xie, Weigen Yao, Yi Yang (possible past Baidu (China) affiliation), Zuozhu Liu, Jiasheng Tang, Shaocheng Wang, Jitao Wang, Jiahong Dong, Weihua Chen, Feng Xu, Fan Wang (possible past Baidu (China) affiliation)
Abstract

Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment. In this paper, we introduce ClinFusion, a vision-centric MLLM designed for holistic medical un...

πŸ“„ EchoBridge: Long-Tail-Aware ECG-Echocardiography Text Alignment for Echocardiography-Derived Cardiac Findings
πŸ—“οΈ Published: 7/27/2026
πŸ”— http://arxiv.org/abs/2607.24553v1
πŸ‘₯ Authors: Xiaocheng Fang, Jieyi Cai, Guangkun Nie, Haoyu Wang (possible past Tencent (China) affiliation), Jiarui Jin, Yujie Xiao, Bo Liu (possible past Meta (United States) affiliation), Chenyang He, Qinghao Zhao, Gaofeng Cheng, Hongyan Li, Shenda Hong (possible past Peking University affiliation)
Abstract

Standardized echocardiography conclusions provide meaningful supervision for learning ECG representations of echocardiography-derived cardiac findings. Global ECG--text alignment may entangle modality-specific factors, while long-tailed finding distributions provide sparse positive supervision for low-prevalence conditions. We propose EchoBridge with Complementary Shared--Private Projection (CSPP) and Adaptive Prototype Boundary Calibration (APBC). CSPP maps each modality into shared and auxilia...

πŸ“„ DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes
πŸ—“οΈ Published: 7/27/2026
πŸ”— http://arxiv.org/abs/2607.24516v1
πŸ‘₯ Authors: Jiahao Xie, Zhongbin Guo, Qianle Wang, Ruiqi Lu, Dongling Xiao (possible past Baidu (China) affiliation), Wanxuan Sun, Cheng Yang (possible past Tsinghua University affiliation)
Abstract

While data curation for Vision Language Models (VLMs) is increasingly active, public practice for constructing pretraining mixtures remains largely heuristic: practitioners stack datasets that pass quality filters, set cross-domain ratios by intuition, and lack a principled, attributable criterion for admitting new data, while frontier recipes remain undisclosed. We formulate data construction as a systematic mixture-optimization problem and turn it into a reproducible engineering discipline by ...

πŸ“„ Are Prompt Optimizers Blind? Cross-Modal Visual Feedback for Automatic Prompt Optimization
πŸ—“οΈ Published: 7/27/2026
πŸ”— http://arxiv.org/abs/2607.24354v1
πŸ‘₯ Authors: Haoyue Liu, Xiaoyu Ma (possible past Google (United States) affiliation), Ye Chen, Yuexian Zou (possible past Peking University affiliation), Xiaoying Tang
Abstract

Automatic prompt optimization (APO) has been widely adopted to adapt vision-language models (VLMs) to downstream tasks without weight updates, yielding promising results. However, on multimodal tasks, the effectiveness of APO is fundamentally bottlenecked by a blind feedback channel: the optimizer reads the question, the prediction, and the gold answer, but never the input image on which the model failed, and therefore cannot diagnose visually grounded errors. As a remedy, we introduce Cross-Mod...

πŸ“„ FilmBench: A Film-Grade Benchmark for Cinematic Video Generation
πŸ—“οΈ Published: 7/27/2026
πŸ”— http://arxiv.org/abs/2607.24241v1
πŸ‘₯ Authors: Shengyi Wang, Niantong Li, Guangzheng Hu, Hong Qi, Fei Ding, Weixu Qiao, Jinlin Wang, Xiaotong Lv, Peng Han, Zimeng Li, Fanshu Ding, Yushu Wang, Han Wu, Jingjing Chen, Chongxiao Wang, Yanhao Wu, Chenglong Huang, Xiaoqian Zhu, Jie Tian (possible past Shanghai Jiao Tong University affiliation), Hua Li, Jingjing Fan, Mingshuang Tang, Zhong Li (possible past Tencent (China) affiliation), Hengxia Qiang, Weibin Chen, Jinyang Zhen, Bing Zhao, Lin Qu, Jing Li (possible past Tencent (China) affiliation), Hu Wei
Abstract

Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary (overall visual quality, coarse text alignment and temporal smoothness) rather than the professional Cinematic Language criteria by which films are actually made and judged, so they as...

πŸ“„ StanceFlip: A Comprehensive Multi-Dimensional Benchmark for Multimodal Conversational Stance Flipping Forecasting
πŸ—“οΈ Published: 7/27/2026
πŸ”— http://arxiv.org/abs/2607.24191v1
πŸ‘₯ Authors: Heyan Chai, Xin Li (possible past Google (United States) affiliation), Wenjie Wang, Jianyang Qin, Chaoyang Li, Lu Wang (possible past University Of Washington affiliation), Hao Chen, Qing Liao
Abstract

Conversational stance detection has shifted from static text analysis to dynamic multimodal modeling. However, existing benchmarks exhibit three key limitations: failure to capture the dynamic evolution of beliefs, particularly during stance reversals; difficulty in disentangling affective states from logical reasoning; and neglect of the critical role of multimodal cues in resolving pragmatic ambiguities such as sarcasm. To address these limitations, we propose StanceFlip, a benchmark designed ...

πŸ“„ Agent-UCT: Upper Confidence Bounds Applied to Trees for Agentic Workflow Optimization with Cost-Awareness
πŸ—“οΈ Published: 7/27/2026
πŸ”— http://arxiv.org/abs/2607.24162v1
πŸ‘₯ Authors: Yang Li (possible past Google (United States) affiliation), Hai Liu, Dian Shao, Yu Wang (possible past Tsinghua University affiliation), Xiyu Chen, Sergey Volkov, Bozhi Wang, Ziyu Sun, Sihang Liu, Ye Luo, Xiaowei Zhang
Abstract

Optimizing agentic workflows, such as retrieval-augmented generation (RAG) pipelines, requires navigating a combinatorial space of discrete component choices under tight evaluation budgets. Existing approaches - heuristic search, black-box optimization, and standard tree search methods - do not explicitly exploit the compositional structure of these workflows, leading to redundant computation and inefficient budget allocation. We introduce Agent-UCT (Agent-based Cost-Aware Upper Confidence Bound...

πŸ“„ Scaling GUI Agents with Visual State Transitions
πŸ—“οΈ Published: 7/27/2026
πŸ”— http://arxiv.org/abs/2607.24112v1
πŸ‘₯ Authors: Xiangyan Liu, Kaixin Li, Haonan Wang, Biao Wu, Meng Fang (possible past Tencent (China) affiliation), Longxu Dou, Chao Du, Michael Qizhe Shieh, Tianyu Pang (possible past Tsinghua University affiliation)
Abstract

We introduce State Transition Pretraining (STP) as a new scaling axis for GUI agents. During the STP stage, we continually pretrain a unified multimodal model on visual state transitions by jointly optimizing inverse dynamics (predicting actions from state changes) and forward dynamics (predicting next states from current states and actions). This optimization equips the model with better action-grounded visual representations and an internal world model of GUI dynamics. When subsequently fine-t...

πŸ“„ Towards High-Level Semantic Intelligence
πŸ—“οΈ Published: 7/27/2026
πŸ”— http://arxiv.org/abs/2607.24082v1
πŸ‘₯ Authors: Xiujie Song, Gefei Yang, Yining You, Jiahui Gan, Qi Jia, Shota Watanabe, Tianxi Wan, Mengyue Wu (possible past Shanghai Jiao Tong University affiliation), Kai Yu (possible past Baidu (China) affiliation)
Abstract

Recent advances in AI have substantially expanded its cognitive and reasoning capabilities. From the perspective of semantic complexity, the development of AI reveals a clear trajectory from simple to complex semantic processing. While early AI systems mainly addressed tasks involving direct and literal semantic perception or expression, contemporary systems are increasingly expected to perform more sophisticated cognitive reasoning, enabling the understanding and generation of High-Level Semant...

πŸ“„ Offline-Online Curriculum RL for Multimodal Reasoning
πŸ—“οΈ Published: 7/26/2026
πŸ”— http://arxiv.org/abs/2607.23700v1
πŸ‘₯ Authors: Wendi Deng, Hang Du, Guoshun Nan, Haokun Tian, Jiaqi Yu, Xinlei Cao, Jaile Li, Jingfeng Chen, Ling Deng, Ting Li, Hao Yang (possible past Tencent (China) affiliation), Jun Liu (possible past Tencent (China) affiliation), Xudong Jiang, Sicong Leng
Abstract

Multimodal large language models exhibit capabilities on reasoning tasks, yet often produce flawed intermediate steps while yielding correct final answers. This behavior undermines interpretability and reliability, suggesting reliance on spurious shortcuts rather than faithful reasoning. Although efforts have explored step-level supervision, distinguishing decisive steps from redundant ones remains challenging. We propose $O^2$-CritiCuRL, a novel curriculum reinforcement learning framework that ...

πŸ“„ Offline-to-Online Creative Optimization with Generative Models and Adaptive Testing
πŸ—“οΈ Published: 7/26/2026
πŸ”— http://arxiv.org/abs/2607.23696v1
πŸ‘₯ Authors: Kevin Lee, Benjamin Letham (possible past Meta (United States) affiliation), Zhiyuan Jerry Lin, Elodie Samson, Eric Onofrey, Poppy Zhang, Shawndra Hill, Eytan Bakshy (possible past Meta (United States) affiliation)
Abstract

Ad creative optimization is increasingly constrained by evaluation rather than generation. Generative models can produce many plausible creatives, but reliable evaluation requires online experiments, in which only a limited slate can be tested. We study how to use data from historical A/B tests to generate and select the candidates in that slate. We developed and deployed a performance-driven offline-to-online workflow that guides creative generation with a predictive model as an inference-time ...

πŸ“„ Hybrid Advantage Estimation with Unified Critic for VLM Agentic Reinforcement Learning
πŸ—“οΈ Published: 7/26/2026
πŸ”— http://arxiv.org/abs/2607.23605v1
πŸ‘₯ Authors: Wenxuan Zhang, Yuhui Wang, Donggang Jia, Xiaoqian Shen, Jian Ding (possible past Baidu (China) affiliation), Ivan Viola, JΓΌrgen Schmidhuber, Mohamed Elhoseiny (possible past Meta (United States) affiliation)
Abstract

Large Vision-Language Models (VLMs) now act as agents in interactive environments, where success requires coherent reasoning and decision-making across turns. Although end-to-end training in agentic environments can improve such multi-turn decision-making abilities, current methods mainly rely on either token-wise optimization over concatenated token trajectories or turn-wise optimization with uniform within-turn credit. In this work, we establish theoretical formulations for the two levels of o...

πŸ“„ Separating Capability from Permission: A Governance Framework for Agentic AI Autonomy Levels
πŸ—“οΈ Published: 7/26/2026
πŸ”— http://arxiv.org/abs/2607.23438v1
πŸ‘₯ Authors: Haining Zheng, Qian Dong, Rodolfo K. Depena, Jonathan D. Bhatia, Feng Xiao (possible past Google (United States) affiliation), Peng Xu (possible past Google (United States) affiliation)
Abstract

As AI systems increasingly exhibit agentic behavior, discussions of autonomy often conflate what systems are technically capable of doing with what they should be permitted to do in practice. This paper introduces a governance framework that explicitly separates Allowed Autonomy Levels (AAL), which define the degree of autonomy an AI agent is authorized to exercise given risk, oversight, and accountability considerations, from Autonomous Capability Levels (ACL), which characterize an agent's inh...

πŸ“„ RareLens: Towards End-to-End Rare Disease Care via Aligning Divergent Large Language Model Reasoning
πŸ—“οΈ Published: 7/25/2026
πŸ”— http://arxiv.org/abs/2607.23290v1
πŸ‘₯ Authors: Xi Chen (possible past University Of California, Berkeley affiliation), Hongru Zhou, Shiyu Feng, Hanyu Zhou, Huahui Yi, Rongsheng Wang, Tiancheng He, Kun Wang, Pingping Liu, Qiankun Li, Sicheng Lin, Huiying Ou, Xiaohong Zheng, Tianying Zang, Zhuohang Wu, Leheng Jiang, Kexin Cao, Wenhan Zhang, Chengyi Li, Zhiyang Wang, Songlin Li, Benyou Wang (possible past Tencent (China) affiliation), Ningbei Yin, Shaoting Zhang (possible past Baidu (China) affiliation), Weili Fu, Jian Li (possible past Tencent (China) affiliation), Kang Li
Abstract

Rare diseases collectively affect an estimated 3.5% to 5.9% of the population, yet more than 70% of patients are misdiagnosed and many endure years of evaluation before a diagnosis is reached, because early presentations are nonspecific and relevant expertise is scarce and unevenly distributed. Artificial intelligence could provide support, but existing systems address isolated stages of care, overwhelmingly diagnosis. They typically depend on the results of downstream investigations, and they t...

πŸ“„ MMOE: Modernizing Diffusion Transformers with Efficient Expert Design
πŸ—“οΈ Published: 7/27/2026
πŸ”— http://arxiv.org/abs/2607.24665v1
πŸ‘₯ Authors: Yanhao Jia, Jiepeng Wang, Haibin Huang, Chi Zhang (possible past Peking University affiliation), Erik Cambria, Xuelong Li (possible past Tencent (China) affiliation)
Abstract

Modern large language models scale successfully by pairing capacity growth with efficiency, keeping per-token and deployment costs under control as capacity grows. AIGC Foundation Models (AFMs), especially diffusion-transformer backbones, have begun to adopt sparse experts, but recent efforts mostly enlarge total parameter counts and sparsity ratios without importing the efficiency mechanisms that made LLM scaling practical, so generation quality is seldom balanced against training and deploymen...

πŸ“„ Kimi K3: Open Frontier Intelligence
πŸ—“οΈ Published: 7/27/2026
πŸ”— http://arxiv.org/abs/2607.24653v1
πŸ‘₯ Authors: Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, M. C., Jianfeng Cai, Xinyuan Cai, Peizhou Cao, Yuxuan Cao, Ziwei Chai, Y. Charles, H. S. Che, Guanduo Chen, Guangyu Chen (possible past Microsoft (United States) affiliation), Guanzheng Chen, Huarong Chen, Jia Chen, Jianlong Chen, Jun Chen, Kexin Chen, Peng Chen (possible past Tencent (China) affiliation), Ruijue Chen, Wentao Chen, Xin Chen (possible past Tencent (China) affiliation), Yang Chen (possible past Tencent (China) affiliation), Yanru Chen, Yifei Chen (possible past Baidu (China) affiliation), Yingjiang Chen, Yuankun Chen, Yujie Chen, Yutian Chen (possible past Deepmind (United Kingdom) affiliation), Zhirong Chen, Dazhi Cheng, Yean Cheng, Jialei Cui, Jingbing Cui, Anqi Dai, Jiaqi Deng, Hao Ding, Rui Ding, Shaofeng Ding, Mengfan Dong, Mengnan Dong, Yuhao Dong, Yuxin Dong, Angang Du, Chenzhuang Du, Dikang Du, Jusen Du, Yulun Du, Yu Fan, Jing Feng, Qiulin Feng, Yichen Feng, Kelin Fu, Qiang Fu (possible past Tencent (China) affiliation), Fuxuan Gao, Hongcheng Gao, Jingyue Gao, Tong Gao, Weijia Gao, Shangyi Geng, Jie Gong, Linhu Gong, Shengao Gong, Xiaochen Gong, Qizheng Gu, Yicheng Gu, Shuhao Guan, Haiqing Guo, Shiqi Guo, Xiang Guo, Zhengyan Guo, Beixi Hao, Wenxin Hao, Xiaoru Hao, Dailan He, Haotian He, Lehan He, Qi He, Weiran He, Xinran He (possible past Meta (United States) affiliation), Xinyi He, Yibo He, Yunjia He, Chao Hong, Tiange Hong, Hao Hu, Jiaxi Hu, Ruikun Hu, Weiming Hu, Yangyang Hu, Zhenxing Hu, Liang Hua, Jinbin Huang, Ke Huang, Ruiyuan Huang, Siying Huang, Weixiao Huang, Yan Huang (possible past Tencent (China) affiliation), Zhengjie Huang (possible past Baidu (China) affiliation), Zhiqi Huang, Yulong Hui, Chaobo Jia, Yutong Jiang, Zhejun Jiang, Zuoyou Jiang, Wenyi Jin, Xinyi Jin, Yu Jing, Huanjun Kong, Guokun Lai (possible past Carnegie Mellon University affiliation), Aidi Li, Cheng Li (possible past Google (United States) affiliation), Chengyuan Li, Cong Li (possible past Google (United States) affiliation), Fang Li, Guanyu Li, Haoyang Li, Jia Li (possible past Google (United States) affiliation), Junxiong Li, Lei Li (possible past Carnegie Mellon University affiliation), Letian Li, Lincan Li, Weihong Li, Wentao Li, Xintong Li, Yang Li (possible past Google (United States) affiliation), Yishen Li, Yiwei Li, Yuxiao Li, Zhaowei Li, Zhaoxi Li, Zheming Li, Zhengxiao Li, Zhiyuan Li (possible past Peking University affiliation), Jiawei Lin, Xiaohan Lin, Yibo Lin (possible past Peking University affiliation), Zichao Lin, Ziyan Lin, Bill Liu, Boxiao Liu, Chuan Liu, Liang Liu (possible past Tencent (China) affiliation), Shaowei Liu, Shudong Liu, Shuran Liu, Tianwei Liu, Weizhou Liu, Yangyang Liu, Yanming Liu, Yibo Liu, Yipeng Liu, Zhengying Liu, Zhiheng Liu, Enzhe Lu, Haoyu Lu, Linqiang Lu, Tingzhan Lu, Zhiyuan Lu, Aotian Luo, G. Luo, Junyu Luo, Yifan Luo, B. Lyu, Wenzhou Lyu, Shaoguang Mao, Yuan Mei, Xin Men, Minqing Ni, Yixuan Niu, Siyuan Pan, Shujun Peng, Zhangyang Qi, Ruoyu Qin, Zechao Qin, Zeyu Qin, Haiquan Qiu, Jianxin Qiu, Jiezhong Qiu (possible past Tsinghua University affiliation), Bowen Qu, Yuhao Qu, Zeyu Shang, Youbo Shao, Han Shen, Jincheng Shi, Juanfeng Shi, Lidong Shi, Shengyuan Shi, Wingchun Siu, Pengwei Song, Xiaoxi Song, Jianlin Su, Yunfeng Su, Zhaochen Su, Lin Sui, Jingsong Sun, Junyao Sun, Shaoning Sun, Shuzhe Sun, Tongyu Sun, Yujun Sun, Yunpeng Tai, Chuning Tang, Heyi Tang, Sirui Tang, Zecheng Tang, Chaoran Tian, Rongpeng Tian, Yu Tian, Wei Tu, Chensi Wang, Chuang Wang, Chunjie Wang, Dinglu Wang, Feng Wang, Hailong Wang (possible past Google (United States) affiliation), Haiming Wang, Hao Wang (possible past Tsinghua University affiliation), Hao Wang (possible past Tsinghua University affiliation), Huaqing Wang, Hui Wang, Jiayi Wang, Jinglong Wang, Jinhong Wang, Jiuzheng Wang, Linian Wang, Shaobo Wang, Shenzhi Wang, Shuyi Wang, Si Wang, Siyuan Wang, Tianfu Wang, Wenjue Wang, Xingran Wang, Xinmei Wang, Xinyuan Wang, Xusheng Wang, Yalin Wang, Yangkun Wang, Yao Wang, Yaoyu Wang, Yejie Wang, Yiqin Wang, Yucheng Wang, Yuzhi Wang, Zhaoji Wang, Zhaowei Wang, Zhengtao Wang, Zhenhao Wang, Zhongsheng Wang, Zifan Wang, Chu Wei, Ming Wei, Shouxin Wei, Zichen Wen, Fan Wu, Haoning Wu, Rucong Wu, Wenhao Wu, Xiaoxue Wu, Yingcong Wu, Yongqi Wu, Yuxin Wu (possible past University Of California, Berkeley affiliation), Zijian Wu, Xinglang Xian, Chenxuan Xiang, Yuye Xiang, Bocheng Xiao, Chenjun Xiao, Xin Xiao, Jin Xie, Xiaotong Xie, Yifeng Xie, Zhe Xie, Bowei Xing, Yiming Xiong, Baosheng Xu, Boyu Xu, Jiale Xu, Jianfan Xu, Jing Xu (possible past Meta (United States) affiliation), Jinjing Xu, L. H. Xu, Qingtao Xu, Shuyao Xu, Suting Xu, Tiantian Xu, Tianxiang Xu, Weixin Xu, Xinran Xu, Yangchuan Xu, Ye Xu, Yueni Xu, Ziyao Xu, Haonan Xue, Junjie Yan, Yaoyao Yan, Fan Yang (possible past Tencent (China) affiliation), Guangyao Yang, Hao Yang (possible past Tencent (China) affiliation), Junwei Yang, Ruoyu Yang, Wenjie Yang, Xiaofei Yang, Xinyu Yang, Yi Yang (possible past Baidu (China) affiliation), Yiling Yang, Ying Yang, Yuchen Yang, Zhen Yang (possible past Tsinghua University affiliation), Zhilin Yang (possible past Stanford University affiliation), Zian Yang, Zuhao Yang, Haotian Yao, Dan Ye, Haoran Ye, Wenjie Ye, Zhanbo Ye, Bohong Yin, Haoxiang Yin, Xietong Yin, Chengzhen Yu, Haozhen Yu, Longhui Yu, Shengnan Yu, Shuying Yu, Tianxiang Yu, Enming Yuan, Mengjie Yuan, Tongtian Yue, Wei Yue, Yang Yue, Dunyuan Zha, Haobing Zhan, B. H. Zhang, Dehao Zhang, Fei Zhang, Hao Zhang (possible past Tencent (China) affiliation), Haoyuan Zhang, Huanyu Zhang, Jiapei Zhang, Jiaxuan Zhang, Jin Zhang, Kaiyi Zhang, Miaozhen Zhang, Puqi Zhang, Qinglei Zhang, Rong Zhang, Rui Zhang, Shaoshuai Zhang, Shiyi Zhang, Xiaobin Zhang, Xiaoyun Zhang (possible past Shanghai Jiao Tong University affiliation), Y. Zhang, Yangkun Zhang, Ye Zhang (possible past Google (United States) affiliation), Yichi Zhang, Yikun Zhang, Yizhi Zhang, Yongting Zhang, Yu Zhang (possible past Google (United States) affiliation), Yutao Zhang, Yutong Zhang, Zheng Zhang, Zijing Zhang, Bin Zhao, Chenguang Zhao, Feifan Zhao, Jinglun Zhao, Jinxiang Zhao, Shuai Zhao, Wenshuo Zhao, Xiangyu Zhao, Xuanle Zhao, Yikai Zhao, Zijia Zhao, Haozhi Zheng, Huabin Zheng, Ruihan Zheng, Shaojie Zheng, Tengyang Zheng, Haofeng Zhong, Lei Zhong, Longguang Zhong, M. Zhou, Qiankang Zhou, Runjie Zhou, Ruozhang Zhou, Xinyu Zhou, Yiqiao Zhou, Zaida Zhou, Jinguo Zhu, Liya Zhu, Xinhao Zhu, Yangjunfeng Zhu, Yuxuan Zhu, Zhen Zhu, Chen Zhuang, Weiyu Zhuang, Xinxing Zu
Abstract

We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve information flow across sequence length and model depth. Together with Stable LatentMoE, which effectively activates 16 of 896 routed experts per token, and refined training and data recipes, these advances yield an approximately 2.5x improvement in o...

πŸ“„ BeyondFusion: Self-Aligned Latent Diffusion for Calibration-Free Infrared Super-Resolution and Infrared-Visible Fusion
πŸ—“οΈ Published: 7/27/2026
πŸ”— http://arxiv.org/abs/2607.24110v1
πŸ‘₯ Authors: Minchong Chen, Xiaoyun Yuan, Minyu Cao, Jianing Zhang, Jun Zhang (possible past Tencent (China) affiliation), Shuyang Liu, Xiaokang Yang (possible past Shanghai Jiao Tong University affiliation)
Abstract

Mobile infrared-visible imaging typically pairs a compact infrared sensor with a high-resolution visible camera for complementary perception. While cross-sensor misalignment caused by different optics, viewpoints, fields of view, and exposure timings hinders practical deployment. In this paper, we propose BeyondFusion, a unified latent diffusion framework for calibration-free visible-guided infrared super-resolution and infrared-visible fusion tasks. The proposed framework supports both task-spe...

πŸ“„ SpecFormer: Mitigating Embedding and Attention Collapse via Spectral-Aware Transformer for Recommendation
πŸ—“οΈ Published: 7/27/2026
πŸ”— http://arxiv.org/abs/2607.24025v1
πŸ‘₯ Authors: Yu Cui, Yi Xu, Jiahao Wang, Hao Zhang (possible past Tencent (China) affiliation), Yu Zhang (possible past Google (United States) affiliation), Xiaoyi Zeng, Can Wang (possible past Tsinghua University affiliation), Jinxin Hu, Jiawei Chen (possible past Tencent (China) affiliation)
Abstract

Transformer architectures have achieved remarkable success across diverse domains; however, directly applying their standard self-attention mechanism to recommendation often yields suboptimal performance, sometimes even trailing behind well-designed simple recommendation models. In this paper, we reveal that this performance bottleneck stems from severe embedding and attention collapse unique to recommendation scenarios. The heterogeneity and long-tail nature of recommendation data lead to a sev...

πŸ“„ SCTA: An Agentic Framework for Stable and Interpretable Target Gene Discovery from Single-Cell RNA Sequencing
πŸ—“οΈ Published: 7/26/2026
πŸ”— http://arxiv.org/abs/2607.23821v1
πŸ‘₯ Authors: Shuyu Chen, Chen Zhu (possible past Baidu (China) affiliation), Ye Zhang (possible past Google (United States) affiliation), Yang Li (possible past Google (United States) affiliation), Qiqi Xie, Haohan Wang
Abstract

Identifying therapeutic target genes from single-cell RNA sequencing (scRNA-seq) data remains a fundamental challenge in translational biology. Unlike bulk assays, scRNA-seq captures heterogeneous cellular states and rare subpopulations, but this same heterogeneity makes target discovery highly sensitive to analytical choices throughout the pipeline, including preprocessing, cell population selection, differential expression analysis, and downstream biological interpretation. As a result, existi...

πŸ“„ MS-GPT: Rethinking MS/MS De Novo Structure Elucidation as Spectrum-Induced Posterior Querying of a Molecule-Language Model
πŸ—“οΈ Published: 7/26/2026
πŸ”— http://arxiv.org/abs/2607.23607v1
πŸ‘₯ Authors: Xin Zhao, Yumin Liu, Zhuo Li, Weichu Zheng, Feng Zhu (possible past Baidu (China) affiliation), Xiaokang Yang (possible past Shanghai Jiao Tong University affiliation), Yaohui Jin, Yanyan Xu
Abstract

Molecular structure elucidation from tandem mass spectra (MS/MS) is a central inverse problem in analytical chemistry. Most existing approaches to MS/MS identification remain tied to reference libraries or predefined candidate sets, whereas de novo methods aim to generate structures directly from spectra. A common de novo route predicts a molecular fingerprint from the spectrum and then decodes structures from it, enabling decoder pretraining on large molecule-only corpora. However, this paradig...

*Notable papers are those with at least two authors from a "big" AI/ML lab.