πŸ“„ Notable* Recent AI/ML arXiv Papers

Last updated just now...

πŸ“„ MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
πŸ—“οΈ Published: 7/28/2026
πŸ”— http://arxiv.org/abs/2607.25948v1
πŸ‘₯ Authors: Mingqiao Ye, Zhaochong An, Zhitong Gao, Xian Liu, FranΓ§ois Fleuret, Chuan Li, Amir Zadeh, Serge Belongie (possible past Google (United States) affiliation), Afshin Dehghan, Jesse Allardice, David Mizrahi, Oğuzhan Fatih Kar, Roman Bachmann, Amir Zamir (possible past University Of California, Berkeley affiliation)
Abstract

Any-to-any models predict any modality from any combination of others within a single network, a formulation used in multimodal vision and vision-language models, and increasingly in scientific domains such as ecology and astronomy. Existing any-to-any models are typically trained from scratch using encoder-decoder or diffusion architectures, impacting their performance and preventing them from using strong pre-trained decoder-only models as a prior. In this work, we investigate decoder-only any...

πŸ“„ Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases
πŸ—“οΈ Published: 7/28/2026
πŸ”— http://arxiv.org/abs/2607.25933v1
πŸ‘₯ Authors: Rui Yang, Weihao Xuan, Yi Lin, Zhuhan Bao, Jonathan Chong Kai Liew, Matthew Yu Heng Wong, NicolΓ‘s Lescano, Nikita R. Paripati, Emily Ling-Lin Pai, Jiarui Liu, Heli Qi, Heng-Jui Chang, Benny Kai Guo Loo, Huitao Li, Kunyu Yu, Yufan Wang, Chuan Hong, Shijian Lu, Douglas Teodoro, Naoto Yokoya, Ross Koppel, Mona Diab (possible past Carnegie Mellon University affiliation), Hua Xu, David W. Bates, Nan Liu, Yifan Peng (possible past Stanford University affiliation)
Abstract

Clinical diagnostic evaluation should not only assess whether models can provide correct diagnoses, but also reflect the realities of clinical practice, including progressive disclosure of multimodal information, dynamic updating of diagnostic hypotheses, and continuous refinement of clinical reasoning. However, existing evaluations of multimodal large language models (MLLMs) typically rely on single-turn or isolated tasks, making it difficult to fully capture the complexity of real-world clinic...

πŸ“„ A Machine-Learning-Based Gas Lift Optimization Workflow for Unconventional Fields
πŸ—“οΈ Published: 7/28/2026
πŸ”— http://arxiv.org/abs/2607.25885v1
πŸ‘₯ Authors: Sha, Miao, Alexandra Vendetti, Logan Smart, Gunta Chomchalerm, Yang Chen (possible past Tencent (China) affiliation), Christopher Frazier, Dustin Haralson, Jeremy Sorenson, Xiao Ma, Huafei Sun, Aaron Shinn, Haining Zheng, Xiao-Hui Wu, Peng Xu (possible past Google (United States) affiliation)
Abstract

In this paper, we present an automated data-driven workflow using Machine Learning (ML) for gas lift optimization in unconventional fields. This workflow integrates a ML model that accurately forecasts the Gas Lift Performance Curve, and a Bayesian Optimization Framework to solve for the optimal gas injection rates under the constraints of facility capacity. The ML model leverages the historical production time series data without requiring downhole gauges or multi-rate well tests. We piloted th...

πŸ“„ HiSkill: Empowering LLM Agents with Hierarchical Skill Graphs
πŸ—“οΈ Published: 7/28/2026
πŸ”— http://arxiv.org/abs/2607.25853v1
πŸ‘₯ Authors: Yu Hao, Jinxuan Cai, Qi Zhang (possible past Tencent (China) affiliation), Yawen Li, Zhiqiang Zhang, Chuan Shi, Cheng Yang (possible past Tsinghua University affiliation)
Abstract

Skills have become an important abstraction for enabling large language model (LLM) agents to reuse past experience in long-horizon interactive tasks. However, existing trajectory-to-skill methods often produce flat collections of high-level textual skills that are stored and retrieved independently, leaving skill relations underutilized and maintaining a gap between high-level skills and executable actions. In this paper, we propose HiSkill, a hierarchical skill graph framework that organizes i...

πŸ“„ Cognivia: A Cognitive Behavioral Therapy Copilot for Evidence-Based Mental Healthcare
πŸ—“οΈ Published: 7/28/2026
πŸ”— http://arxiv.org/abs/2607.25681v1
πŸ‘₯ Authors: Qi Chen (possible past Baidu (China) affiliation), Siria Xiyueyao Luo, Jian Wang (possible past Baidu (China) affiliation), Yuan Shi, Haocong Rao, Xuejiao Zhao
Abstract

Cognitive distortion amplifies negative emotions and contributes to mental health disorders. Cognitive Behavioral Therapy (CBT) is an effective way to address cognitive distortions, but its large-scale application is limited by the shortage of professional therapists. Although large language models (LLMs) have recently been explored for mental health applications, existing methods still suffer from limited domain specificity, overly flattering responses, and the absence of well-defined annotatio...

πŸ“„ DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space
πŸ—“οΈ Published: 7/28/2026
πŸ”— http://arxiv.org/abs/2607.25675v1
πŸ‘₯ Authors: Jiangwang Chen, Zixin Song, Junlin Liu, Shuaiyu Zhou, Haiyan Wu, Haihan Shi, Chenxi Zhou, Hanqing Li, Xiao Yang (possible past Tencent (China) affiliation), Da Zhu, Guanjun Jiang, Hai Wan, Xibin Zhao (possible past Tsinghua University affiliation)
Abstract

Text-space optimization adapts large language models (LLMs) by editing external natural-language artifacts rather than model weights, so the optimized artifacts remain inspectable and the model can be treated as a black box. However, most existing text-space methods keep evaluation fixed. On open-ended tasks, this can become a bottleneck: once the solver improves on the criteria a rubric measures, omitted dimensions remain invisible to the optimization signal. Simply evolving the rubric is also ...

πŸ“„ Visual prompt engineering for video models
πŸ—“οΈ Published: 7/28/2026
πŸ”— http://arxiv.org/abs/2607.25537v1
πŸ‘₯ Authors: Robert Geirhos, Yuxuan Li, ThaddΓ€us Wiedemer, Neha Kalibhat, Zi Wang, Mani Malek, Oyvind Tafjord, Kevin Swersky (possible past Google (United States) affiliation), Been Kim (possible past Google (United States) affiliation), Priyank Jaini
Abstract

In the age of foundation models, a model is only as good as its prompt. For this reason, prompt engineering has become an essential technique for improving language model performance. Since video models are currently becoming foundation models for visual tasks (e.g., visual reasoning), we here ask whether they similarly benefit from visual prompt engineering: automatically modifying the task image to improve model performance. For example, for a visual physics reasoning task ("Where does the bal...

πŸ“„ At-the-Roofline Sparse Tensor Contractions on Vector Processors for Transformer Inference
πŸ—“οΈ Published: 7/28/2026
πŸ”— http://arxiv.org/abs/2607.25504v1
πŸ‘₯ Authors: Bowen Wang, Chi Zhang (possible past Peking University affiliation), Diyou Shen, Renzo Andri, Navaneeth Kunhi Purayil, Luca Benini (possible past Eth Zurich affiliation)
Abstract

Fine-grained weight pruning and activation sparsification have emerged as effective approaches for reducing the compute and memory cost of inference for Transformer models. In the moderate-sparsity regime, Gustavson's dataflow provides a natural execution model for exploiting both activation and weight sparsity on vector processors through metadata-driven indexed accumulation. However, existing RVV architectures lack native support for this pattern, forcing kernels to rely on software index deco...

πŸ“„ CAST: Game Solvers as Turn-Level Teachers for LLM Agents
πŸ—“οΈ Published: 7/28/2026
πŸ”— http://arxiv.org/abs/2607.25308v1
πŸ‘₯ Authors: Yu Wang (possible past Tsinghua University affiliation), Yi-Kai Zhang, Wentao Shi, Ziang Ye, Yuchun Miao, Yueqing Sun, Qi Gu, Xunliang Cai, Lan-Zhe Guo, Han-Jia Ye, Fuli Feng (possible past National University Of Singapore affiliation)
Abstract

Training large language models (LLMs) to act in long-horizon games is a promising step toward generalist decision-making, yet reinforcement learning with verifiable rewards (RLVR) relies on sparse final rewards that reveal little about which decisions determine success. Denser process signals could supply this missing turn-level credit, but existing sources are hard to keep both cheap and accurate. We observe that changes in a game solver's state value reveal whether an action advances the state...

πŸ“„ ScaleResfusion: Residual Rectified Flow based on Residual Vector Field
πŸ—“οΈ Published: 7/28/2026
πŸ”— http://arxiv.org/abs/2607.25275v1
πŸ‘₯ Authors: Zhenning Shi, Chen Xu, Junhao Zhang, Kefei Zhang, Linjie Liu, Zhedong Zheng (possible past Baidu (China) affiliation), Tao Li (possible past Baidu (China) affiliation)
Abstract

Real-world Image Restoration (Real-IR) aims to recover high-quality (HQ) images from complex and unknown degradations. Although recent diffusion-based methods have substantially improved perceptual quality, their current designs leave two key challenges unresolved. Methods that start from Gaussian noise are slow and often less faithful to the degraded input. Residual-based methods usually train from scratch, which makes it hard to exploit modern pre-trained generative priors. In this paper, we p...

πŸ“„ Where Steering Signals Come From: Activation Source Selection in Activation Steering
πŸ—“οΈ Published: 7/28/2026
πŸ”— http://arxiv.org/abs/2607.25270v1
πŸ‘₯ Authors: Jiaran Ye, Lingxu Ran, Zijun Yao, Chenpeng Wang, Yong Jiang (possible past Tsinghua University affiliation), Lei Hou (possible past Tsinghua University affiliation), Juanzi Li, Liangming Pan
Abstract

Activation steering controls language models by adding vectors or features to hidden states at inference time, but the upstream source of these steering signals is often treated as a secondary detail. We study this source choice as activation source selection: the combination of source context and activation readout policy used to collect the hidden states from which a steering signal is built. Holding the downstream intervention fixed, we show across three instruction-tuned models and four stee...

πŸ“„ The User Asks, Platforms Compete: How Agentic Recommendation Markets Take Shape
πŸ—“οΈ Published: 7/28/2026
πŸ”— http://arxiv.org/abs/2607.25253v1
πŸ‘₯ Authors: Deyao Hong, Kehan Zheng, Qian Li (possible past National University Of Defense Technology affiliation), Jun Zhang (possible past Tencent (China) affiliation), Jie Jiang (possible past Tencent (China) affiliation), Hongning Wang
Abstract

Online recommendation has traditionally taken place after a user enters a platform, which determines the candidate pool and the ranking shown to the user. LLM-based user agents enable a different recommendation process: a user specifies a need before choosing a platform, leaving platforms to compete for the user's attention, which we refer to as an agentic recommendation market. In our controlled LLM-based experiments across three product domains, we find this new setting of recommendation creat...

πŸ“„ Observing sycophantic AI validate others reduces its appeal but not its persuasiveness
πŸ—“οΈ Published: 7/28/2026
πŸ”— http://arxiv.org/abs/2607.25166v1
πŸ‘₯ Authors: Meryl Ye, Robert Kraut (possible past Carnegie Mellon University affiliation), Steve Rathje (possible past University Of Cambridge affiliation)
Abstract

AI chatbots can be ``sycophantic,'' or overly agreeable and flattering toward users. Sycophantic AI has been shown to entrench attitudes, yet users frequently fail to recognize it (a phenomenon we call ``sycophancy blindness''). We tested whether increasing users' awareness of sycophancy protects them from its harmful effects. In one preregistered experiment (n = 940), participants received a brief written warning about sycophancy before conversing with a sycophantic chatbot. In a second preregi...

πŸ“„ Learning from 53.6K Real-World Developer Edits of AI-Generated Code
πŸ—“οΈ Published: 7/27/2026
πŸ”— http://arxiv.org/abs/2607.25130v1
πŸ‘₯ Authors: Jenny T. Liang, Mihika Bairathi, Wayne Chi, Ameet Talwalkar (possible past University Of California, Berkeley affiliation), Nishant Subramani (possible past Allen Institute For Artificial Intelligence affiliation), Valerie Chen
Abstract

Imperfections in AI-generated code require that software developers modify the generated code manually, or by re-prompting an AI programming assistant. Manual code edits provide more realistic and granular information on editing behavior than Git commits, which only contain final successful code snippets. Yet, due to a lack of high-quality, realistic code editing data, LLMs are mostly trained on publicly available Git data (e.g., commits). To address this gap, we introduce DECODE (Developer Edit...

πŸ“„ ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding
πŸ—“οΈ Published: 7/27/2026
πŸ”— http://arxiv.org/abs/2607.24743v2
πŸ‘₯ Authors: Hangjie Yuan, Yichen Qian, Zhiwei Tang, Xianzhe Xu, Lirong Wu, Sicheng Yang, Jinwang Wang, Pengju Wang, Zhitao Zeng, Yizeng Han, Yan Xing, Shengxuan Luo, Tao Feng, Qing Xie, Weigen Yao, Yi Yang (possible past Baidu (China) affiliation), Zuozhu Liu, Jiasheng Tang, Shaocheng Wang, Jitao Wang, Jiahong Dong, Weihua Chen, Feng Xu, Fan Wang (possible past Baidu (China) affiliation)
Abstract

Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment. In this paper, we introduce ClinFusion, a vision-centric MLLM designed for holistic medical un...

πŸ“„ MMOE: Modernizing Diffusion Transformers with Efficient Expert Design
πŸ—“οΈ Published: 7/27/2026
πŸ”— http://arxiv.org/abs/2607.24665v1
πŸ‘₯ Authors: Yanhao Jia, Jiepeng Wang, Haibin Huang, Chi Zhang (possible past Peking University affiliation), Erik Cambria, Xuelong Li (possible past Tencent (China) affiliation)
Abstract

Modern large language models scale successfully by pairing capacity growth with efficiency, keeping per-token and deployment costs under control as capacity grows. AIGC Foundation Models (AFMs), especially diffusion-transformer backbones, have begun to adopt sparse experts, but recent efforts mostly enlarge total parameter counts and sparsity ratios without importing the efficiency mechanisms that made LLM scaling practical, so generation quality is seldom balanced against training and deploymen...

πŸ“„ Kimi K3: Open Frontier Intelligence
πŸ—“οΈ Published: 7/27/2026
πŸ”— http://arxiv.org/abs/2607.24653v1
πŸ‘₯ Authors: Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, M. C., Jianfeng Cai, Xinyuan Cai, Peizhou Cao, Yuxuan Cao, Ziwei Chai, Y. Charles, H. S. Che, Guanduo Chen, Guangyu Chen (possible past Microsoft (United States) affiliation), Guanzheng Chen, Huarong Chen, Jia Chen, Jianlong Chen, Jun Chen, Kexin Chen, Peng Chen (possible past Tencent (China) affiliation), Ruijue Chen, Wentao Chen, Xin Chen (possible past Tencent (China) affiliation), Yang Chen (possible past Tencent (China) affiliation), Yanru Chen, Yifei Chen (possible past Baidu (China) affiliation), Yingjiang Chen, Yuankun Chen, Yujie Chen, Yutian Chen (possible past Deepmind (United Kingdom) affiliation), Zhirong Chen, Dazhi Cheng, Yean Cheng, Jialei Cui, Jingbing Cui, Anqi Dai, Jiaqi Deng, Hao Ding, Rui Ding, Shaofeng Ding, Mengfan Dong, Mengnan Dong, Yuhao Dong, Yuxin Dong, Angang Du, Chenzhuang Du, Dikang Du, Jusen Du, Yulun Du, Yu Fan, Jing Feng, Qiulin Feng, Yichen Feng, Kelin Fu, Qiang Fu (possible past Tencent (China) affiliation), Fuxuan Gao, Hongcheng Gao, Jingyue Gao, Tong Gao, Weijia Gao, Shangyi Geng, Jie Gong, Linhu Gong, Shengao Gong, Xiaochen Gong, Qizheng Gu, Yicheng Gu, Shuhao Guan, Haiqing Guo, Shiqi Guo, Xiang Guo, Zhengyan Guo, Beixi Hao, Wenxin Hao, Xiaoru Hao, Dailan He, Haotian He, Lehan He, Qi He, Weiran He, Xinran He (possible past Meta (United States) affiliation), Xinyi He, Yibo He, Yunjia He, Chao Hong, Tiange Hong, Hao Hu, Jiaxi Hu, Ruikun Hu, Weiming Hu, Yangyang Hu, Zhenxing Hu, Liang Hua, Jinbin Huang, Ke Huang, Ruiyuan Huang, Siying Huang, Weixiao Huang, Yan Huang (possible past Tencent (China) affiliation), Zhengjie Huang (possible past Baidu (China) affiliation), Zhiqi Huang, Yulong Hui, Chaobo Jia, Yutong Jiang, Zhejun Jiang, Zuoyou Jiang, Wenyi Jin, Xinyi Jin, Yu Jing, Huanjun Kong, Guokun Lai (possible past Carnegie Mellon University affiliation), Aidi Li, Cheng Li (possible past Google (United States) affiliation), Chengyuan Li, Cong Li (possible past Google (United States) affiliation), Fang Li, Guanyu Li, Haoyang Li, Jia Li (possible past Google (United States) affiliation), Junxiong Li, Lei Li (possible past Carnegie Mellon University affiliation), Letian Li, Lincan Li, Weihong Li, Wentao Li, Xintong Li, Yang Li (possible past Google (United States) affiliation), Yishen Li, Yiwei Li, Yuxiao Li, Zhaowei Li, Zhaoxi Li, Zheming Li, Zhengxiao Li, Zhiyuan Li (possible past Peking University affiliation), Jiawei Lin, Xiaohan Lin, Yibo Lin (possible past Peking University affiliation), Zichao Lin, Ziyan Lin, Bill Liu, Boxiao Liu, Chuan Liu, Liang Liu (possible past Tencent (China) affiliation), Shaowei Liu, Shudong Liu, Shuran Liu, Tianwei Liu, Weizhou Liu, Yangyang Liu, Yanming Liu, Yibo Liu, Yipeng Liu, Zhengying Liu, Zhiheng Liu, Enzhe Lu, Haoyu Lu, Linqiang Lu, Tingzhan Lu, Zhiyuan Lu, Aotian Luo, G. Luo, Junyu Luo, Yifan Luo, B. Lyu, Wenzhou Lyu, Shaoguang Mao, Yuan Mei, Xin Men, Minqing Ni, Yixuan Niu, Siyuan Pan, Shujun Peng, Zhangyang Qi, Ruoyu Qin, Zechao Qin, Zeyu Qin, Haiquan Qiu, Jianxin Qiu, Jiezhong Qiu (possible past Tsinghua University affiliation), Bowen Qu, Yuhao Qu, Zeyu Shang, Youbo Shao, Han Shen, Jincheng Shi, Juanfeng Shi, Lidong Shi, Shengyuan Shi, Wingchun Siu, Pengwei Song, Xiaoxi Song, Jianlin Su, Yunfeng Su, Zhaochen Su, Lin Sui, Jingsong Sun, Junyao Sun, Shaoning Sun, Shuzhe Sun, Tongyu Sun, Yujun Sun, Yunpeng Tai, Chuning Tang, Heyi Tang, Sirui Tang, Zecheng Tang, Chaoran Tian, Rongpeng Tian, Yu Tian, Wei Tu, Chensi Wang, Chuang Wang, Chunjie Wang, Dinglu Wang, Feng Wang, Hailong Wang (possible past Google (United States) affiliation), Haiming Wang, Hao Wang (possible past Tsinghua University affiliation), Hao Wang (possible past Tsinghua University affiliation), Huaqing Wang, Hui Wang, Jiayi Wang, Jinglong Wang, Jinhong Wang, Jiuzheng Wang, Linian Wang, Shaobo Wang, Shenzhi Wang, Shuyi Wang, Si Wang, Siyuan Wang, Tianfu Wang, Wenjue Wang, Xingran Wang, Xinmei Wang, Xinyuan Wang, Xusheng Wang, Yalin Wang, Yangkun Wang, Yao Wang, Yaoyu Wang, Yejie Wang, Yiqin Wang, Yucheng Wang, Yuzhi Wang, Zhaoji Wang, Zhaowei Wang, Zhengtao Wang, Zhenhao Wang, Zhongsheng Wang, Zifan Wang, Chu Wei, Ming Wei, Shouxin Wei, Zichen Wen, Fan Wu, Haoning Wu, Rucong Wu, Wenhao Wu, Xiaoxue Wu, Yingcong Wu, Yongqi Wu, Yuxin Wu (possible past University Of California, Berkeley affiliation), Zijian Wu, Xinglang Xian, Chenxuan Xiang, Yuye Xiang, Bocheng Xiao, Chenjun Xiao, Xin Xiao, Jin Xie, Xiaotong Xie, Yifeng Xie, Zhe Xie, Bowei Xing, Yiming Xiong, Baosheng Xu, Boyu Xu, Jiale Xu, Jianfan Xu, Jing Xu (possible past Meta (United States) affiliation), Jinjing Xu, L. H. Xu, Qingtao Xu, Shuyao Xu, Suting Xu, Tiantian Xu, Tianxiang Xu, Weixin Xu, Xinran Xu, Yangchuan Xu, Ye Xu, Yueni Xu, Ziyao Xu, Haonan Xue, Junjie Yan, Yaoyao Yan, Fan Yang (possible past Tencent (China) affiliation), Guangyao Yang, Hao Yang (possible past Tencent (China) affiliation), Junwei Yang, Ruoyu Yang, Wenjie Yang, Xiaofei Yang, Xinyu Yang, Yi Yang (possible past Baidu (China) affiliation), Yiling Yang, Ying Yang, Yuchen Yang, Zhen Yang (possible past Tsinghua University affiliation), Zhilin Yang (possible past Stanford University affiliation), Zian Yang, Zuhao Yang, Haotian Yao, Dan Ye, Haoran Ye, Wenjie Ye, Zhanbo Ye, Bohong Yin, Haoxiang Yin, Xietong Yin, Chengzhen Yu, Haozhen Yu, Longhui Yu, Shengnan Yu, Shuying Yu, Tianxiang Yu, Enming Yuan, Mengjie Yuan, Tongtian Yue, Wei Yue, Yang Yue, Dunyuan Zha, Haobing Zhan, B. H. Zhang, Dehao Zhang, Fei Zhang, Hao Zhang (possible past Tencent (China) affiliation), Haoyuan Zhang, Huanyu Zhang, Jiapei Zhang, Jiaxuan Zhang, Jin Zhang, Kaiyi Zhang, Miaozhen Zhang, Puqi Zhang, Qinglei Zhang, Rong Zhang, Rui Zhang, Shaoshuai Zhang, Shiyi Zhang, Xiaobin Zhang, Xiaoyun Zhang (possible past Shanghai Jiao Tong University affiliation), Y. Zhang, Yangkun Zhang, Ye Zhang (possible past Google (United States) affiliation), Yichi Zhang, Yikun Zhang, Yizhi Zhang, Yongting Zhang, Yu Zhang (possible past Google (United States) affiliation), Yutao Zhang, Yutong Zhang, Zheng Zhang, Zijing Zhang, Bin Zhao, Chenguang Zhao, Feifan Zhao, Jinglun Zhao, Jinxiang Zhao, Shuai Zhao, Wenshuo Zhao, Xiangyu Zhao, Xuanle Zhao, Yikai Zhao, Zijia Zhao, Haozhi Zheng, Huabin Zheng, Ruihan Zheng, Shaojie Zheng, Tengyang Zheng, Haofeng Zhong, Lei Zhong, Longguang Zhong, M. Zhou, Qiankang Zhou, Runjie Zhou, Ruozhang Zhou, Xinyu Zhou, Yiqiao Zhou, Zaida Zhou, Jinguo Zhu, Liya Zhu, Xinhao Zhu, Yangjunfeng Zhu, Yuxuan Zhu, Zhen Zhu, Chen Zhuang, Weiyu Zhuang, Xinxing Zu
Abstract

We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve information flow across sequence length and model depth. Together with Stable LatentMoE, which effectively activates 16 of 896 routed experts per token, and refined training and data recipes, these advances yield an approximately 2.5x improvement in o...

πŸ“„ EchoBridge: Long-Tail-Aware ECG-Echocardiography Text Alignment for Echocardiography-Derived Cardiac Findings
πŸ—“οΈ Published: 7/27/2026
πŸ”— http://arxiv.org/abs/2607.24553v1
πŸ‘₯ Authors: Xiaocheng Fang, Jieyi Cai, Guangkun Nie, Haoyu Wang (possible past Tencent (China) affiliation), Jiarui Jin, Yujie Xiao, Bo Liu (possible past Meta (United States) affiliation), Chenyang He, Qinghao Zhao, Gaofeng Cheng, Hongyan Li, Shenda Hong (possible past Peking University affiliation)
Abstract

Standardized echocardiography conclusions provide meaningful supervision for learning ECG representations of echocardiography-derived cardiac findings. Global ECG--text alignment may entangle modality-specific factors, while long-tailed finding distributions provide sparse positive supervision for low-prevalence conditions. We propose EchoBridge with Complementary Shared--Private Projection (CSPP) and Adaptive Prototype Boundary Calibration (APBC). CSPP maps each modality into shared and auxilia...

πŸ“„ Agent-UCT: Upper Confidence Bounds Applied to Trees for Agentic Workflow Optimization with Cost-Awareness
πŸ—“οΈ Published: 7/27/2026
πŸ”— http://arxiv.org/abs/2607.24162v1
πŸ‘₯ Authors: Yang Li (possible past Google (United States) affiliation), Hai Liu, Dian Shao, Yu Wang (possible past Tsinghua University affiliation), Xiyu Chen, Sergey Volkov, Bozhi Wang, Ziyu Sun, Sihang Liu, Ye Luo, Xiaowei Zhang
Abstract

Optimizing agentic workflows, such as retrieval-augmented generation (RAG) pipelines, requires navigating a combinatorial space of discrete component choices under tight evaluation budgets. Existing approaches - heuristic search, black-box optimization, and standard tree search methods - do not explicitly exploit the compositional structure of these workflows, leading to redundant computation and inefficient budget allocation. We introduce Agent-UCT (Agent-based Cost-Aware Upper Confidence Bound...

πŸ“„ BeyondFusion: Self-Aligned Latent Diffusion for Calibration-Free Infrared Super-Resolution and Infrared-Visible Fusion
πŸ—“οΈ Published: 7/27/2026
πŸ”— http://arxiv.org/abs/2607.24110v1
πŸ‘₯ Authors: Minchong Chen, Xiaoyun Yuan, Minyu Cao, Jianing Zhang, Jun Zhang (possible past Tencent (China) affiliation), Shuyang Liu, Xiaokang Yang (possible past Shanghai Jiao Tong University affiliation)
Abstract

Mobile infrared-visible imaging typically pairs a compact infrared sensor with a high-resolution visible camera for complementary perception. While cross-sensor misalignment caused by different optics, viewpoints, fields of view, and exposure timings hinders practical deployment. In this paper, we propose BeyondFusion, a unified latent diffusion framework for calibration-free visible-guided infrared super-resolution and infrared-visible fusion tasks. The proposed framework supports both task-spe...

*Notable papers are those with at least two authors from a "big" AI/ML lab.