πŸ“„ Notable* Recent AI/ML arXiv Papers

Last updated just now...

πŸ“„ CMA-OT: Hierarchical Expert Supervision for Dance-to-Music Generation
πŸ—“οΈ Published: 9/11/2026
πŸ”— http://arxiv.org/abs/2609.13118v1
πŸ‘₯ Authors: Jinting Wang, Chenxing Li, Dong Yu (possible past Tencent (China) affiliation), Li Liu (possible past National University Of Defense Technology affiliation)
Abstract

Dance-to-music (D2M) generation aims to synthesize music that is rhythmically and stylistically aligned with dance videos. A key challenge arises from the semantic mismatch between sparse dance cues, such as rhythm and style, and the dense information required for music composition, including structure, instrumentation, and expressive dynamics. Existing methods typically rely on these sparse cues and supervise only the final audio output, resulting in poorly learned music representations and gen...

πŸ“„ MP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant
πŸ—“οΈ Published: 9/11/2026
πŸ”— http://arxiv.org/abs/2609.13076v1
πŸ‘₯ Authors: Yi-Jen Shih, Shih-Yun Shan Kuan, Guan-Ting Lin, Kai-Wei Chang, Siddhant Arora, Shu-Wen Yang, Abdelrahman Mohamed (possible past University Of Toronto affiliation), Shinji Watanabe (possible past Carnegie Mellon University affiliation), Hung-Yi Lee, David Harwath
Abstract

Conversational voice agents have advanced significantly, offering increasingly natural human-machine interactions through both cascaded and end-to-end architectures. However, while recent benchmarks extensively evaluate dyadic interactions and passive audio comprehension, they largely overlook a prevalent real-world scenario: multi-party conversations. Evaluating agents in these settings is fundamentally more challenging than in dyadic interactions due to the exponentially greater conversational...

πŸ“„ How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks
πŸ—“οΈ Published: 9/11/2026
πŸ”— http://arxiv.org/abs/2609.13009v1
πŸ‘₯ Authors: Ali Ansari, Haoran Sun, Andy Zeyi Liu, Mark Jabbour, Yongshan Ding, Steven Girvin, Yu He (possible past Google (United States) affiliation), Sohrab Ismail-Beigi, Aleksander Kubica, Owen D. Miller, Corey O'hern, Vidvuds Ozolins, David Poland, A. Douglas Stone, Frank C. Van Den Bosch, Logan Wright, Navid Akbari, Santanu Antu, Kangle Cai, Andrew Calabrese-Day, Mateo CΓ‘rdenes Wuttig, Meng Cheng, Barry T. Chiang, Ali Ghorashi, Shouzhen Gu, Haoyang Huang, Zhibo Kang, Lukas Kienesberger, Hantian Liu, Charles Lomba, Zhongling Lu, Wenchao Ma, Rohin E. Mcintosh, Evan Mckinney, Ivan Rojkov, Xulei Sun, Yarone Meir Tokayer, Naveen Balaji Umasankar, Mira Varma, Leda Wang, Qimin Wang, Tyler Wang, Haoyu Wei, Jinming Yang, Jinchen Zhao, Sherlock Tingrui Zhao, Qinyuan Zheng, Jay S. Zou, Lucas Baker (possible past Deepmind (United Kingdom) affiliation), Arman Cohan, John Sous
Abstract

Low reported scores on leading physics benchmarks, including those featured in the Artificial Analysis Intelligence Index (2026), suggest that frontier language models still struggle with advanced physics, a demanding test of their scientific reasoning and quantitative problem-solving abilities. Yet this impression does not always align with domain experts' experiences using these models in their work. We revisit these reported findings by evaluating frontier models on six widely used physics be...

πŸ“„ SeqMoE: Toward Full-Load Performance via Predictive and Graph-Compatible MoE Offloading
πŸ—“οΈ Published: 9/11/2026
πŸ”— http://arxiv.org/abs/2609.12978v1
πŸ‘₯ Authors: Zihan Wang (possible past Tsinghua University affiliation), Yuqi Wang, Lei Gong, Cheng Tang, Wenqi Lou, Teng Wang, Chao Wang (possible past Google (United States) affiliation), Xuehai Zhou
Abstract

Mixture-of-Experts (MoE) creates a structural advantage for offloading: only a small fraction of activated experts need to reside in device memory, and if they can be loaded in time for computation, offloading can in principle approach full-load performance, where all model weights reside in device memory. Yet translating MoE's structural advantage into practical offloading gains remains challenging. We propose SeqMoE to bridge this gap. To maximize expert hits, we build predictive memory manage...

πŸ“„ UniPart: Towards Zero-shot Language-Grounded 3D Part Segmentation for Embodied Interaction
πŸ—“οΈ Published: 9/11/2026
πŸ”— http://arxiv.org/abs/2609.12898v1
πŸ‘₯ Authors: Xinqiang Yu, Zekun Qi, Jiawei He (possible past Baidu (China) affiliation), Wenyao Zhang, Xuchuan Chen, Guaocai Yao, Li Yi (possible past Stanford University affiliation), Zhaoxiang Zhang (possible past Beijing Academy Of Artificial Intelligence affiliation), He Wang (possible past Stanford University affiliation)
Abstract

Fine-grained robotic manipulation depends on understanding parts, not only whole objects. Existing 3D foundation models tend to be either generalized but object-aware, or part-aware but limited to closed-set taxonomies, which weakens zero-shot transfer. We study text-conditioned 3D part segmentation, where a free-form phrase selects a functional part on point cloud. We introduce UniPart, a feed-forward cross-modal 3D Transformer that conditions CLIP text embedding. To scale supervision, we build...

πŸ“„ Online Video Agent Harness for Long Video Understanding
πŸ—“οΈ Published: 9/11/2026
πŸ”— http://arxiv.org/abs/2609.12818v1
πŸ‘₯ Authors: Sen Yang (possible past Tencent (China) affiliation), Boqiang Duan, Jing Yang, Weihao Bo, Jie Liu (possible past Tencent (China) affiliation), Boyuan Tong, Ze Feng, Wenkang Zhang, Jingdong Wang (possible past Baidu (China) affiliation), Hua Wu (possible past Baidu (China) affiliation)
Abstract

Long video understanding often behaves like a visual needle-in-a-haystack problem: query-relevant evidence is sparsely distributed across long temporal spans, while packing dense frames into a single VLM context incurs \textit{context rot} and high cost. Existing video agents often rely on query-agnostic offline preprocessing or ad hoc tool sets, which can miss query-specific details and waste computation. In this work, we present VideoXAgent, a purely online video-agent harness for long video u...

πŸ“„ SIMS: Scale-Invariant Merit-Function-Based Scalarization for Multi-Task Learning
πŸ—“οΈ Published: 9/11/2026
πŸ”— http://arxiv.org/abs/2609.12599v1
πŸ‘₯ Authors: Zebin Chen, Fei Xing, Yang Chen (possible past Tencent (China) affiliation), Hua Liu, Andy Hf Chow, Yuhua Qian, Yu Zhang (possible past Google (United States) affiliation)
Abstract

Multi-task learning (MTL) requires navigating unavoidable trade-offs among competing objectives. This paradigm is frequently formulated as multi-objective optimization (MOO), where the scalarization is favored to reduce an MOO problem to a single objective. We empirically find that existing merit-function-based scalarization approaches are sensitive to the relative scales of different objectives in practical MTL, where task losses commonly differ by orders of magnitude. The optimization process ...

πŸ“„ Earth-Agent-Pro: Towards Real-World Full-Chain Earth Observation with Agents
πŸ—“οΈ Published: 9/11/2026
πŸ”— http://arxiv.org/abs/2609.12533v1
πŸ‘₯ Authors: Zhutao Lv, Chenhao Dang, Yi Feng, Yanpei Gong, Xiaolei Wang, Junyan Ye, Conghui He (possible past Tsinghua University affiliation), Weijia Li (possible past Tsinghua University affiliation)
Abstract

Real-world Earth observation (EO) agents must translate high-level scientific questions into executable workflows to acquire observations, prepare data, perform domain computations, and derive conclusions from runtime evidence. Existing EO agents typically start from supplied observations, while benchmarks typically provide prepared inputs or candidate answers, leaving full-chain open-world EO execution largely untested. We present Earth-Agent-Pro, an execution-adaptive Plan-and-Execute framewor...

πŸ“„ EvoRS: On-Policy Self-Evolution of Reward Systems for Open-Ended Reinforcement Learning
πŸ—“οΈ Published: 9/11/2026
πŸ”— http://arxiv.org/abs/2609.12459v1
πŸ‘₯ Authors: Weiyuan Li, Aili Chen, Xintao Wang (possible past Tencent (China) affiliation), Yikai Zhang, Qingqing Dong, Jinghan Xu, Hongru Hou, Wenxuan Zhao, Chengkun Lang, Jun Gao (possible past Nvidia (United States) affiliation), Yuanli Guo, Hongcheng Guo, Yanghua Xiao, Deqing Yang
Abstract

Open-ended reinforcement learning often relies on rubric-based rewards for tasks without directly verifiable answers. Yet the policy and reward system form a dynamic feedback loop: as the policy optimizes the current reward, an initially useful reward system may become unreliable due to reward hacking or reduced response discriminability. The reward system should therefore evolve rather than remain fixed during training. Existing dynamic-rubric methods adapt evaluation criteria, but reward failu...

πŸ“„ Bridging Vision Foundation Model Priors with CLIP for Spatial-aware Few-shot Anomaly Detection in Medical Images
πŸ—“οΈ Published: 9/11/2026
πŸ”— http://arxiv.org/abs/2609.12454v1
πŸ‘₯ Authors: Juzheng Miao, Yuchen Yuan (possible past Baidu (China) affiliation), Cheng Chen (possible past Google (United States) affiliation), Pheng-Ann Heng
Abstract

Vision-Language Models such as CLIP enable effective few-shot medical anomaly detection (AD) via strong image-text semantic alignment. However, their globally contrastive pretraining lacks explicit spatial supervision, limiting precise lesion localization. In contrast, Vision Foundation Models (VFMs) such as DINO learn spatially coherent patch representations via self-distillation and local-to-global consistency, better capturing fine-grained anatomical structures. Leveraging this complementarit...

πŸ“„ Debiasing as a Measurement Intervention: Calibrated Ties and Resolution Loss in LLM-as-a-Judge Evaluation
πŸ—“οΈ Published: 9/11/2026
πŸ”— http://arxiv.org/abs/2609.12439v1
πŸ‘₯ Authors: Liang Zhao (possible past Baidu (China) affiliation), Yong Wang (possible past Baidu (China) affiliation), Jiangzhe Chen
Abstract

LLM-as-a-judge protocols are commonly debiased by instructing judges to ignore presentation cues such as citation formatting, source labels, and evidence-display style. We show that this intervention can suppress bias while damaging the resolution of the measurement instrument. We introduce TraceJudgeBench, a diagnostic benchmark for auditing citation-like artifacts in RAG and agent-workflow evaluation, covering content-equivalent pairs, citation ablations, correctness conflicts, human-validated...

πŸ“„ BlueLM-GUI Technical Report: A Real-Device-Centric Flywheel for Self-Improving Mobile GUI Agents
πŸ—“οΈ Published: 9/11/2026
πŸ”— http://arxiv.org/abs/2609.12394v1
πŸ‘₯ Authors: Tong Ye, Kunyang Han, Guozhi Wang, Longqiang Luo, Zhifeng Ding, Yongxiang Zhang, Xiaolei Shen, Yuxuan Zhang, Zhuping Zhang, Tao Xu (possible past Meta (United States) affiliation), Yue Pan, Yucheng Zhao, Yupei Hu, Yuanjiang Ouyang, Danfeng Shen, Runqi Lin, Hongda Cai, Zhaoxiong Wang, Mengjia Yan, Yingjie Zhong, Chen Zhou, Zeyu Zhang, Xuwen Zhu, Penggang Shi, Mingcheng Luo, Ziyang Wu, Min Jin, Mingfu Shen, Zairong Xu, Fan Zhang, Hao Wang (possible past Tsinghua University affiliation), Liang Liu (possible past Tencent (China) affiliation), Zhulin Xie, Lijun Yao, Xiao Liang, Liangmin Wen, Liqiang Feng, Feilong Wu, Min Hu, Min Chen, Guanjing Xiong, Xiaohu Ruan, Xiaoxin Chen
Abstract

Mobile GUI agents are shifting from multi-module frameworks to native models trained end-to-end, yet industrial deployment faces three persistent gaps. Sandbox training produces a distribution mismatch with production environments; expensive real-device failures remain underutilized; and fixed benchmarks saturate, losing the power to guide iteration. We present BlueLM-GUI, a 35B-A3B mobile GUI agent built as a real-device-centric flywheel that closes these gaps through three principles. Every Sa...

πŸ“„ Toward Robust Personalized Alignment for LLMs: Mitigating Persona Drift in Multi-Turn Dialogue
πŸ—“οΈ Published: 9/11/2026
πŸ”— http://arxiv.org/abs/2609.12373v1
πŸ‘₯ Authors: Youyuan Zhang, Siyuan Li (possible past Tencent (China) affiliation), Fangming Liu, Jing Li (possible past Tencent (China) affiliation)
Abstract

Persona drift remains a central challenge for personalized language models, as user profiles evolve over long interactions rather than remain permanently fixed. Models must therefore revise persistent persona states when preferences genuinely change, while avoiding updates driven by transient, ambiguous, or unresolved observations. We propose CORE, which separates turn-local evidence from persistent persona-state revision and selectively updates grounded user preferences through uncertainty-awar...

πŸ“„ AIM: A Privacy-Aware Interoperable Memory Framework for Multi-Agent Multi-User LLM Systems
πŸ—“οΈ Published: 9/11/2026
πŸ”— http://arxiv.org/abs/2609.12320v1
πŸ‘₯ Authors: Zachary Johnson, Nigel Boachie Kumankumah, Somya Chatterjee, Tejas Sathyamurthi, Min Chen, Xinyi Alice Li, Xiao Wang (possible past Google (United States) affiliation), Emily Morgan Gelchie, Jessica Lin, Sadid A. Hasan, Sulaiman Vesal (possible past Stanford University affiliation)
Abstract

Traditional large language models (LLMs) are scoped to individual user sessions, limiting their knowledge to a single conversation and preventing them from learning user preferences that evolve over time. Existing agentic memory systems address this limitation but generally operate at the individual-user level, restricting the public knowledge that could be shared across users to improve downstream responses. We introduce AIM (Agentic Interoperable Memory), a unified, privacy-aware memory framew...

πŸ“„ Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models
πŸ—“οΈ Published: 9/11/2026
πŸ”— http://arxiv.org/abs/2609.12303v1
πŸ‘₯ Authors: Kalyani Marathe, Artidoro Pagnoni, Tomasz Limisiewicz, Margaret Li, Mike Lewis (possible past Meta (United States) affiliation), Luke Zettlemoyer (possible past University Of Washington affiliation), Srinivasan Iyer
Abstract

Small models are made more capable through distillation from a larger one that shares their tokenization scheme. However, do distilled byte and token models behave similarly in terms of scaling trends as compute and data increases? To enable this comparison, we introduce two variants to efficiently convert token logits to Byte Logits: 1) approximate: Marginalize-It, and 2) exact: End-Of-Token. We then present the first large scale study of overtraining decoder-only dense transformer models varyi...

πŸ“„ Reinforcement Learning over Patient Trajectories for Clinical Reasoning in EHR Foundation Models
πŸ—“οΈ Published: 9/10/2026
πŸ”— http://arxiv.org/abs/2609.12277v1
πŸ‘₯ Authors: Yuxin Xiao, Sheng Zhang, Chandan Singh, Tristan Naumann, Hoifung Poon (possible past Microsoft (United States) affiliation), Jianfeng Gao (possible past Microsoft (United States) affiliation), Xiaodong Liu
Abstract

Electronic health record (EHR) foundation models trained on longitudinal patient trajectories have demonstrated strong performance across diverse clinical prediction tasks. However, their clinical reasoning capabilities remain constrained by next-token prediction on limited and incomplete EHR data. To address this, we propose a reinforcement learning (RL) fine-tuning framework that treats EHR foundation models as generative policies over patient trajectories. We formulate common clinical predict...

πŸ“„ Recommendation Retrievers Need Verifiers: Universal Generative Reranking for Sequential Recommendations
πŸ—“οΈ Published: 9/10/2026
πŸ”— http://arxiv.org/abs/2609.12270v1
πŸ‘₯ Authors: Benyu Zhang, Qiang Zhang (possible past Tsinghua University affiliation), Rui Li (possible past Google (United States) affiliation), Qunshu Zhang, Devansh Tandon, Neeraj Bhatia
Abstract

First-stage recommenders in multi-stage systems produce a ranked candidate list from which a limited prefix is forwarded to downstream rankers. Because each forwarded item must be processed by more expensive ranking stages, this shortlist cannot be arbitrarily large. The first-stage objective is therefore high coverage of relevant items within the forwarded prefix, commonly measured by Recall@$k$. A relevant item may be available deeper in the retrieved list but absent from the shorter prefix th...

πŸ“„ Reality Is the Final Verifier: On Two Key Gaps in Agentic Software Engineering
πŸ—“οΈ Published: 9/10/2026
πŸ”— http://arxiv.org/abs/2609.12039v1
πŸ‘₯ Authors: Alexander Krentsel, Shubham Agarwal, Mert Cemri, Shu Liu (possible past Tencent (China) affiliation), Sidharth Sankhe, Ziming Mao, Matei Zaharia (possible past University Of California, Berkeley affiliation), Ion Stoica (possible past University Of California, Berkeley affiliation)
Abstract

Software development follows an implementation-verification loop in which developers or agents iteratively revise an implementation until an evaluator, such as a test suite, accepts it. The evaluator checks the implementation against a set of requirements under a model of the deployment environment. Yet even a formal proof that the implementation satisfies the requirements under the model cannot guarantee acceptable behavior after deployment. Requirements only approximate stakeholder intent, and...

πŸ“„ MindTopo: Can Foundation Models Reason in Topological Space?
πŸ—“οΈ Published: 9/10/2026
πŸ”— http://arxiv.org/abs/2609.11900v1
πŸ‘₯ Authors: Yunfei Ge, Anbang Liu, Qineng Wang, Johnalbert Garnica, Jianwen Lyu, Zihan Wang (possible past Tsinghua University affiliation), Reuben Tan, Jianfeng Gao (possible past Microsoft (United States) affiliation), Ruohan Zhang, Yining Hong, Jiajun Wu (possible past Massachusetts Institute Of Technology affiliation), Manling Li
Abstract

Spatial reasoning depends not only on metric properties such as distance, angle, and shape, but also on topological relations that remain invariant under continuous deformation. Cognitive science identifies these relations as foundational to spatial understanding, yet foundation-model evaluations largely focus on metric or viewpoint-dependent relations. We introduce MindTopo, a benchmark of topological intuition across five properties grounded in cognitive science and formal topology: continuity...

πŸ“„ The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement
πŸ—“οΈ Published: 9/10/2026
πŸ”— http://arxiv.org/abs/2609.11873v1
πŸ‘₯ Authors: Yi Duan, Ying Liu, Zirui Tang, Haodong Chen, Jun Zhou, Yumou Liu, Bangrui Xu, Yukai Wu, Sidi Chen, Yuhan Zhou, Haoyu Wang (possible past Tencent (China) affiliation), Xiaoyou Yu, Shaokun Han, Xuzhou Zhu, Le Zhou, Bolin Lu, Wei Zhou, Jiachen Liu (possible past Baidu (China) affiliation), Nuozhou Fang, Jiaxin Tian, Ruoyu Chen, Yuxuan Li, Kai Zuo, Kaiyan Zhang, Jiantao Qiu, Conghui He (possible past Tsinghua University affiliation), Guoliang Li (possible past Tsinghua University affiliation), Bowen Zhou, Zhiyuan Liu (possible past Tsinghua University affiliation), Zhoufutu Wen, Jihua Kang, Xuanhe Zhou, Fan Wu
Abstract

Recursive self-improvement (RSI) enables AI systems to turn experience and feedback into persistent changes that improve both their capabilities and the process of future improvement. We first use the Headroom-Closed Index (HCI) to reveal the problems of existing LLMs, then introduce the RSI concept and its development roadmap: from improvement-execution autonomy, improvement-strategy autonomy, experience-acquisition autonomy, and environment-adaptation autonomy, to recursive meta-improvement. N...

πŸ“„ ActSafeGuard: Differentiable and Training-Aligned Constraint Enforcement for Flow-Matching Policies
πŸ—“οΈ Published: 9/10/2026
πŸ”— http://arxiv.org/abs/2609.11697v1
πŸ‘₯ Authors: Jianming Ma, Rongjun Jin, Xiaxi Si, Yang Zhang (possible past Tsinghua University affiliation), Yiheng Li, Yue Gao (possible past Tsinghua University affiliation)
Abstract

Vision-Language-Action (VLA) and World-Action Models (WAMs) have demonstrated strong capabilities in general-purpose robotic manipulation, yet their generated actions may violate hard physical constraints and therefore be unsafe or infeasible for deployment. Existing safety approaches either optimize statistical safety objectives without deterministic per-step guarantees or correct unsafe actions only during inference, creating a mismatch between policy training and execution. We introduce ActSa...

πŸ“„ RiPPLE: Cross-Space Performance Prediction from Early Training for Neural Architecture Search
πŸ—“οΈ Published: 9/11/2026
πŸ”— http://arxiv.org/abs/2609.12418v1
πŸ‘₯ Authors: Yifan Yang (possible past Tencent (China) affiliation), Zhaoyan Wang, Zheng Gao, Xiaoyu Li (possible past Tencent (China) affiliation), Jiaojiao Jiang
Abstract

Neural architecture search (NAS) evaluates candidate networks, but fully training enough architectures to rank an entire space is expensive. Zero-cost proxies score architectures at initialization, yet their ranking quality varies across search spaces. Learned predictors reduce evaluation cost but typically require fully trained labels or partial-training features for individual candidates. We introduce $\textbf{RiPPLE}$, $\underline{\textbf{R}}$anking v$\underline{\textbf{i}}$a $\underline{\tex...

πŸ“„ Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data
πŸ—“οΈ Published: 9/10/2026
πŸ”— http://arxiv.org/abs/2609.11917v1
πŸ‘₯ Authors: Atindra Jha, Margaret Li, Jure Leskovec (possible past Stanford University affiliation), Percy Liang (possible past Stanford University affiliation), Luke Zettlemoyer (possible past University Of Washington affiliation)
Abstract

As the supply of human-written text is exhausted, it has become standard practice to repeat language model training data. Prior work has studied data repetition for densely activated Transformers, but the effects of data repetition remains largely unexplored for recently dominant sparse architectures such as Mixture-of-Experts (MoE), despite their increased compute efficiency. We vary data repetition rates across single- and multi-domain data mixes, and across MoE settings, including expert coun...

πŸ“„ CoRA-NAS: Coarse Ranking and Anchor-Residual Refinement for Neural Architecture Search
πŸ—“οΈ Published: 9/10/2026
πŸ”— http://arxiv.org/abs/2609.11884v1
πŸ‘₯ Authors: Yifan Yang (possible past Tencent (China) affiliation), Zhaoyan Wang, Zheng Gao, Xiaoyu Li (possible past Tencent (China) affiliation), Jiaojiao Jiang
Abstract

Zero-cost proxies rank architectures cheaply, but their reliability varies across search spaces. We introduce CoRA-NAS (COarse Ranking + Anchor-residual), a two-stage framework combining a static ranking prior with low-cost learning-curve refinement. CoRA-Rank aggregates capacity and structure-at-initialization proxies through an equal-weight log-rank consensus and a target-free consensus gate. CoRA-Refine samples anchors across this prior, extrapolates their early validation curves, and propaga...

πŸ“„ Reflex-Informed Neuromuscular Reinforcement Learning for Muscle-Driven Locomotion
πŸ—“οΈ Published: 9/10/2026
πŸ”— http://arxiv.org/abs/2609.11733v1
πŸ‘₯ Authors: Jian Zhou (possible past Tencent (China) affiliation), Xingyu Zhang, Rui Ma, Yu Cao (possible past University Of California, Berkeley affiliation), Shane Xie, Zhi-Qiang Zhang
Abstract

Muscle-driven locomotion provides a physically grounded approach to generating realistic human movement. However, achieving both physiological plausibility and adaptability to changes in musculoskeletal capacity and external disturbances remains a fundamental challenge. To address this limitation, we propose a Reflex-Informed Neuromuscular Reinforcement Learning framework for muscle-driven locomotion. Within this framework, a fixed phase-dependent reflex controller serves as the underlying neuro...

πŸ“„ Your Model Already Knows Don't Teach It, Learn to Ask It: Soft Prompting for Few-Shot Adaptation of Vision-Language Models
πŸ—“οΈ Published: 9/10/2026
πŸ”— http://arxiv.org/abs/2609.11310v1
πŸ‘₯ Authors: Gautam Rajendrakumar Gare, Siyi Li, Hewei Wang, Cesar Daniel Hernandez, Wei Zhao (possible past Tencent (China) affiliation), Wolfgang M. Pauli, John Galeotti, Deva Ramanan (possible past Carnegie Mellon University affiliation)
Abstract

We address few-shot object detection with vision-language models (VLMs) in out-of-domain settings such as aerial, industrial, and medical imagery, using only ten annotated images for supervision. Existing adaptation methods are discrete prompt optimization and LoRA fine-tuning. We revisit a third option: soft prompting, where a small number of continuous prompt tokens are optimized while the pretrained backbone remains frozen. We identify two key design choices. First, placing prompt tokens at...

*Notable papers are those with at least two authors from a "big" AI/ML lab.