| M3-Former: Multimodal Transformer with Mixture-of-Experts for Long-Term Vessel Trajectory PredictionarXiv:2609.10559 | Wenzhe Jin, Haina Tang | 11 Sept 2026 | cs.LG | — | 89 |
| RiVaT-Fuse: Reliability-Calibrated Variational Tensor Fusion for Multimodal Prediction under Modality UncertaintyarXiv:2609.10798 | Yingfan Xu, Tieming Liu, Ye Liang +1 | 11 Sept 2026 | cs.LG | — | 89 |
| CoRA-NAS: Coarse Ranking and Anchor-Residual Refinement for Neural Architecture SearcharXiv:2609.11884 | Yifan Yang, Zhaoyan Wang, Zheng Gao +2 | 11 Sept 2026 | cs.LG | — | 89 |
| HuRo: Robotizing Human Videos for Scalable VLA PretrainingarXiv:2609.10706 | Jinho Jeong, Se June Joo, Jaehyun Kang +2 | 11 Sept 2026 | cs.RO | — | 89 |
| Meta-Learning for Data-Efficient Plant Growth Estimation via Vision Transformers and Fuzzy ClusteringarXiv:2609.10749 | Sheikh Hasan Elahi, Rusith Chamara Hathurusinghe Dewage, Habib Ullah +2 | 11 Sept 2026 | cs.CV | — | 89 |
| How Much Velocity Does Off-Ball Space Value Need? A Broadcast-Viewport BenchmarkarXiv:2609.10801 | Seongjin Choi | 11 Sept 2026 | cs.CV | — | 89 |
| Scale-Aware 3D Deep Learning for Robust Brain Metastasis Detection in Multimodal MRIarXiv:2609.10825 | Sylvain Jaume, Hongming Wang, Simon K. Warfield | 11 Sept 2026 | eess.IV | — | 89 |
| Are We Really Doing Few-Shot Learning? A Critical Examination of Pre-Training AssumptionsarXiv:2609.10851 | Alejandro Galan-Cuenca, Marcelo Saval-Calvo, Antonio Javier Gallego | 11 Sept 2026 | cs.CV | — | 89 |
| Symmetry-aware super-resolution of crystal orientation maps via invariant latent-space learningarXiv:2609.10898 | Umang Garg, Warren Zamudio, McLean P. Echlin +2 | 11 Sept 2026 | cs.CV | — | 89 |
| New Evidence, Same Choice: Testing Physical Experiment Selection in Vision Language ModelsarXiv:2609.11022 | Sourajit Saha, Shubhashis Roy Dipta, Nobin Sarwar +2 | 11 Sept 2026 | cs.CV | — | 89 |
| Meta-Learning for Classifier Selection in Image Datasets: A Feature-Driven Framework for Accuracy PredictionarXiv:2609.11041 | Zahra Nabizadeh_Shahre_Babak, Farzaneh Koohestani, Nader Karimi +2 | 11 Sept 2026 | cs.CV | — | 89 |
| TailProp: content-adaptive light- and heavy-tailed propagation for visionarXiv:2609.11081 | Jiahao Kong, Zihan Li | 11 Sept 2026 | cs.CV | — | 89 |
| Improving Faint Object Detection for Space Situational Awareness with Variational AutoencodersarXiv:2609.11269 | Angela Cratere, Luca Ghilardi, Vishnu Reddy +2 | 11 Sept 2026 | cs.CV | — | 89 |
| Your Model Already Knows Don't Teach It, Learn to Ask It: Soft Prompting for Few-Shot Adaptation of Vision-Language ModelsarXiv:2609.11310 | Gautam Rajendrakumar Gare, Siyi Li, Hewei Wang +2 | 11 Sept 2026 | cs.CV | — | 89 |
| Hologram Representation via Quadratic Phase Gaussian SplattingarXiv:2609.11434 | Haolong Wang, Yicheng Zhan, Kaan Ak\c{s}it +1 | 11 Sept 2026 | cs.GR | — | 89 |
| Breaking the Central Bias: Spatially Partitioned Experts for Coordinate-Based NeuroevolutionarXiv:2609.11518 | Romain Claret, Arthur Gygax, Michael O'Neill +2 | 11 Sept 2026 | cs.NE | — | 89 |
| Vidu S2: Real-Time Interactive, Editable, and Spatial Video GenerationarXiv:2609.11638 | Jintao Zhang, Kai Jiang, Jintao Chen +2 | 11 Sept 2026 | cs.CV | — | 89 |
| Multimodal Taxonomic Conditioning for Generative Plankton ImageryarXiv:2609.11673 | Daniela Ivanova, Ozgu Goksu, Nicolas Pugeault | 11 Sept 2026 | cs.CV | — | 89 |
| Logit Refiner: Improving Visual Autoregressive Models via Intra-Scale Dependency ModelingarXiv:2609.11804 | Meimingwei Li, Stefan Andreas Baumann, Felix Krause +1 | 11 Sept 2026 | cs.CV | — | 89 |
| 3D Point Splatting for mmWave Radar Novel View SynthesisarXiv:2609.11894 | Adnan Armouti, Yixuan Gao, Rajalakshmi Nandakumar | 11 Sept 2026 | cs.CV | — | 89 |
| Optimizing Three Critical Factors for Practical and Effective OOD Detection Fine-TuningarXiv:2308.01030 | Hyunjun Choi, JaeHo Chung, Hawook Jeong | 11 Sept 2026 | cs.LG | — | 89 |
| CertDW: Towards Certified Dataset Ownership Verification via Conformal CalibrationarXiv:2506.13160 | Ting Qiao, Yiming Li, Jianbin Li +2 | 11 Sept 2026 | cs.LG | — | 89 |
| Sublinear Variational Optimization of Gaussian Mixture Models with Millions to Billions of ParametersarXiv:2501.12299 | Sebastian Salwig, Till Kahlke, Florian Hirschberger +2 | 11 Sept 2026 | stat.ML | — | 89 |
| Divergence-Based Similarity Function for Multi-View Contrastive LearningarXiv:2507.06560 | Jaehyoung Jeon, Cheolsu Lim, Myungjoo Kang | 11 Sept 2026 | cs.CV | — | 89 |
| Federated Learning for Surgical Vision in Appendicitis Classification: Results of the FedSurg EndoVis 2024 ChallengearXiv:2510.04772 | Max Kirchner, Hanna Hoffmann, Alexander C. Jenke +2 | 11 Sept 2026 | cs.CV | — | 89 |
| Discriminative Span as a Predictor of Synthetic Data Utility via Classifier ReconstructionarXiv:2605.09697 | Radhika Amar Desai, Modigari Narendra | 11 Sept 2026 | cs.CV | — | 89 |
| ProsMAE: Multi-Source MAE Pretraining for ISUP Grade ClassificationarXiv:2607.08162 | Anna Jung, Kyeonghun Kim, Youngung Han +2 | 11 Sept 2026 | cs.CV | — | 89 |
| What to Preserve, Where to Adapt: A Depth-Wise Analysis of Forgetting in Continual Gynecological Image SegmentationarXiv:2608.13660 | Amal Saqib, Tausifa Jan Saleem, Numan Saeed +1 | 11 Sept 2026 | cs.CV | — | 89 |
| GameWAM: A World Action Model for Video GamesarXiv:2608.26200 | Yuncheng Guo, Zhanqiu Zhang, Yiwen Guo +1 | 11 Sept 2026 | cs.AI | — | 89 |
| Motus2: A Self-Evolving General World Model for Dexterous ManipulationarXiv:2608.30237 | Hongzhe Bi, Zihao Zhou, Yihang Tang +2 | 11 Sept 2026 | cs.RO | — | 89 |
| Representation learning of human cortical folding to reveal long lasting neurodevelopmental signaturesarXiv:2609.05438 | Julien Laval, Robin Guiavarch, Antoine Dufournet +2 | 11 Sept 2026 | q-bio.QM | — | 89 |
| Reason Through the Latent! Making Latent Visual Reasoning NecessaryarXiv:2609.06746 | Suhyeong Park, Junha Jung, Jaewoo Kang | 11 Sept 2026 | cs.AI | — | 89 |
| OmniHallu: Unified Hallucination Detection for Cross-Modal Comprehension and Generation in Multimodal Large Language ModelsarXiv:2609.11244 | Jianjiang Yang, Peihang Li, Shanqing Xu +2 | 11 Sept 2026 | cs.CL | — | 89 |
| MultiHuSE: A Multimodal Dataset for Humour Styles and EmotionsarXiv:2609.11322 | Mary Ogbuka Kenneth, Foaad Khosmood, Abbas Edalat | 11 Sept 2026 | cs.CL | — | 89 |
| BodyCam-VQA: Enhanced Body-Worn Camera Video Captioning via Multimodal Reasoning and Probe Question GenerationarXiv:2609.10815 | Karish Gupta, Matthew Alex, Alex Li +2 | 11 Sept 2026 | cs.CV | — | 89 |
| MindTopo: Can Foundation Models Reason in Topological Space?arXiv:2609.11900 | Yunfei Ge, Anbang Liu, Qineng Wang +2 | 11 Sept 2026 | cs.AI | — | 89 |
| Do Vision-Language Models Understand Visual Persuasiveness? A Diagnosis via Visual Persuasive FactorsarXiv:2511.17036 | Gyuwon Park, Hyounghun Kim | 11 Sept 2026 | cs.CL | — | 89 |
| Characterizing Text Branch Sensitivity in Medical Vision-Language Segmentation via Evidence DecouplingarXiv:2609.02663 | Ziquan Liu, Zhewei Zhu, Xuyang Shi | 11 Sept 2026 | cs.CV | — | 89 |
| AgenticGen: Reward-Guided Agentic Video Generation for AdvertisingarXiv:2609.09187 | Xingyuan Bu, Chengru Song, Hao Zhou +2 | 11 Sept 2026 | cs.CV | — | 89 |
| Reliability-Aware Hybrid-K Ensemble Selection for Cervical Cytology Classification: Integrating Discrimination, Calibration, and Selective PredictionarXiv:2609.09189 | Nisreen Albzour, Sarah S. Lam | 11 Sept 2026 | eess.IV | — | 89 |