| HiPerViT: A Hierarchical Perceiver-Vision Transformer Architecture for Multi-Scale Texture RecognitionarXiv:2609.10917 | Jo\~ao Pedro C. A. de S\'a, Odemir Martinez Bruno | 11 Sept 2026 | cs.CV | — | 89 |
| CamPilot: A Multi-Agent Cinematic Assistant for Camera-Controlled Movie GenerationarXiv:2609.10943 | Yang Wu, Stefano Petrangeli, Ishita Dasgupta +1 | 11 Sept 2026 | cs.CV | — | 89 |
| Toward Interpretable Multimodal Fusion: Heat Conduction Modeling for Hyperspectral and LiDAR Joint ClassificationarXiv:2609.11040 | Kan Wei, Jiahui Cui, Jing Yao +2 | 11 Sept 2026 | cs.CV | — | 89 |
| Beyond Benchmarks: Using VLMs to Reveal Systematic Classification Failures Under Real World ConditionsarXiv:2609.11126 | Dieuwertje Alblas, Alma M. Liezenga, Jan Erik van Woerden +2 | 11 Sept 2026 | cs.CV | — | 89 |
| ReconPlusGen: Injecting Reconstruction Prior into Multi-view 3D Generation through Noise Inversion and ModulationarXiv:2609.11129 | Jiarui Liu, Heng Li, Weiyu Li +2 | 11 Sept 2026 | cs.CV | — | 89 |
| LAION-Mobile: Evaluating Deepfake Detectors On One Million Smartphone PhotosarXiv:2609.11134 | Achim von Stryk, Janis Keuper | 11 Sept 2026 | cs.CV | — | 89 |
| UniH$^3$: Unifying Hierarchical Homogeneity and Heterogeneity for All-in-One Medical Image RestorationarXiv:2609.11156 | Zhiwen Yang, Jiayin Li, Chengyu Liu +2 | 11 Sept 2026 | cs.CV | — | 79 |
| Beyond Visual Quality: Evaluating Physical Consistency under Ego-Motion with EgoGenEvalarXiv:2609.11172 | Yilin Long, Chenming Zhu, Zitang Gou +2 | 11 Sept 2026 | cs.CV | — | 89 |
| A Multi-View and Confusion-Guided Ensemble Framework for Robust Synthetic Image AttributionarXiv:2609.11188 | Zuomin Qu | 11 Sept 2026 | cs.CV | — | 89 |
| CEM-TUDASR: Computationally efficient multi-modality transformer based unsupervised domain adaptive super-resolution approacharXiv:2609.11201 | Anjali Sarvaiya, Jay Kadel, Kishor Upla +1 | 11 Sept 2026 | cs.CV | — | 89 |
| Tri-DehazeGS: Scene--Medium Decoupled Gaussian Splatting with Transmittance-Aware OptimizationarXiv:2609.11223 | Kui Jiang, Yang Gu, Jiacheng Liu +2 | 11 Sept 2026 | cs.CV | — | 89 |
| When is Test-Time Adaptation Identifiable From Unlabeled Evidence?arXiv:2609.11235 | Kartik Jhawar, Lipo Wang | 11 Sept 2026 | cs.CV | — | 89 |
| HALDETECT at ImageEval 2026 Shared Tasks: Answer-First Contrastive Grounding with QLoRAarXiv:2609.11236 | Syed Mohaiminul Hoque, Md Sakhawat Hossain | 11 Sept 2026 | cs.CV | — | 89 |
| SCINTILLA-SNN: A Spiking Multi-Scale Selective Aggregation Network for Perineural Invasion PredictionarXiv:2609.11237 | Youngung Han, Yului Jeong, Kyeonghun Kim +2 | 11 Sept 2026 | cs.CV | — | 89 |
| Fast and Accurate Monomodal 3D High Resolution Deep Registration of Drosophila Larval Brain VolumesarXiv:2609.11240 | Daniel Reisenb\"uchler, Yousef Sadegheih, Michael Dittrich +2 | 11 Sept 2026 | cs.CV | — | 89 |
| From Evaluation to Enhancement: Benchmarking and Improving Think-with-Video Reasoning for Video Generative ModelsarXiv:2609.11242 | Meng Luo, Yicheng Liu, Jiahao Wang +2 | 11 Sept 2026 | cs.CV | — | 89 |
| Uncertainty DMD: Restoring Diversity in Few-Step Autoregressive Video DistillationarXiv:2609.11265 | Zixuan Duan, Xunzhi Xiang, Yabo Chen +2 | 11 Sept 2026 | cs.CV | — | 89 |
| Order-Aware 2.5D Multiple Instance Learning for Preoperative MRI-Based Perineural Invasion Risk Assessment in Intrahepatic CholangiocarcinomaarXiv:2609.11271 | Hyunsu Go, Youngung Han, Kyeonghun Kim +2 | 11 Sept 2026 | cs.CV | — | 89 |
| SAMV-DUSt3R: Instance-Centric 3D Scene Decoupling from Sparse Multi-ViewsarXiv:2609.11279 | Langxu Zhao, Zuan Gu, Yingdan Zhang +2 | 11 Sept 2026 | cs.CV | — | 89 |
| GRIPNet: Gaussian Radial Intensity Prior Guided Architecture for Pulmonary Nodule Detection in CTarXiv:2609.11312 | Haojie Yang, Ran Su | 11 Sept 2026 | cs.CV | — | 89 |
| Mi-Ripple: Restoring Images Degraded by Iterative AI EditingarXiv:2609.11317 | Jiayin Chen, Yicheng Xu, Muting Wang | 11 Sept 2026 | cs.CV | — | 79 |
| Predictive Multi-Landmark OCT Tracking for Increased Motion RobustnessarXiv:2609.11330 | Konrad Reuter, Suresh Guttikonda, Chaitali Uday Karekar +2 | 11 Sept 2026 | cs.CV | — | 89 |
| R4Tun: LLM-guided adaptive segmental tunnel lining segmentation in point cloudsarXiv:2609.11360 | Xinghui Tao, Zehao Ye, Guangming Wang +2 | 11 Sept 2026 | cs.CV | — | 89 |
| Vision Transformer-Based Multi-Level Feature Fusion for Multi-Label Sewer Defect ClassificationarXiv:2609.11375 | Xu Fang, Zhuoran Wang, Qing Li +2 | 11 Sept 2026 | cs.CV | — | 89 |
| Brain-PACE: A Deep Siamese MRI Framework for Modelling Longitudinal Brain AccelerationarXiv:2609.11378 | Samuel Maddox (School of Computing Sciences, University of East Anglia), Jacob Newman (School of Computing Sciences +2 | 11 Sept 2026 | cs.CV | — | 89 |
| DINO-Med: A Unified Patch-Based Adaptation Framework for Multi-Modal Medical Image Analysis Applied to Liver Fibrosis StagingarXiv:2609.11380 | Boya Wang, Ruizhe Li, Chao Chen +1 | 11 Sept 2026 | cs.CV | — | 89 |
| Multi-Modal Controlled Coherent Motion GenerationarXiv:2609.11439 | Yifei Liu, Qiong Cao, Hongwei Yi +2 | 11 Sept 2026 | cs.CV | — | 89 |
| BruNet: A Cross-Domain Transfer Framework for Bruise SegmentationarXiv:2609.11463 | Qiming Wang, Richard J. Motley, Ebube E. Obi +2 | 11 Sept 2026 | cs.CV | — | 89 |
| BridgeMatch: Conditional Transport Bridges in Matching Matrix Space for 3D Deformable RegistrationarXiv:2609.11472 | Qianliang Wu, Haobo Jiang, Guangwei Gao +2 | 11 Sept 2026 | cs.CV | — | 89 |
| Pre- and Post-Treatment Brain Metastases Segmentation Using nnU-Net with Post-Processing for BraTS 2026arXiv:2609.11477 | Haobin Liu, Xin Wang | 11 Sept 2026 | cs.CV | — | 89 |
| FreeFlow: A Bias-free Hierarchical Transformer for Optical Flow EstimationarXiv:2609.11486 | Vladislav Bargatin, Alexander Yakovenko, Khaled Abud +1 | 11 Sept 2026 | cs.CV | — | 80 |
| Recursive Code World Models: Building Complex Worlds through Recursive Scene ProgramsarXiv:2609.11499 | Zhiqi Li, Yuxuan Liao, Bo Zhu | 11 Sept 2026 | cs.CV | — | 81 |
| UBone3D: Physics-Rectified Conditional Flow Matching for Anatomical 3D Shape Completion from UltrasoundarXiv:2609.11506 | Weiying Chen, Yuchong Gao, Siyuan Li +2 | 11 Sept 2026 | cs.CV | — | 89 |
| Harnessing Intrinsic Subject-Aware Attention for Controllable Multi-Subject Video GenerationarXiv:2609.11507 | Niange Yu, Ye Tian, Biaolong Chen +2 | 11 Sept 2026 | cs.CV | — | 89 |
| Prototype Matters: Modality-unified Prototype Self-distillation for Unsupervised Visible-infrared Person Re-identificationarXiv:2609.11514 | Menglin Wang, Xiaojin Gong | 11 Sept 2026 | cs.CV | — | 89 |
| LoopVAE: Recurrent Depth Across Scales for Visual TokenizationarXiv:2609.11516 | Zhiying Lu | 11 Sept 2026 | cs.CV | — | 89 |
| Learning Interaction between Image and Layout Priors for Joint Image-Layout Generation in Design TemplatesarXiv:2609.11519 | Shirong Yang, Bo Yang, Ying Cao | 11 Sept 2026 | cs.CV | — | 89 |
| World in World: Explore the World with World ModelsarXiv:2609.11548 | Chenxi Song, Yanming Yang, Chi Zhang | 11 Sept 2026 | cs.CV | — | 81 |
| A Comparative Evaluation of Pre-trained Convolutional Neural Networks for Melanoma DetectionarXiv:2609.11550 | Wagner Moreno Schmitz, Marco Antonio de Castro Barbosa, Thiago Magalh\~aes Amaral +2 | 11 Sept 2026 | cs.CV | — | 89 |
| Learn the Solid, Not the File: Canonical Inputs for Neural Networks on CAD Boundary RepresentationsarXiv:2609.11573 | Heinrich Jiang, Hager Yasser Mohamed, Alexander Hitt +2 | 11 Sept 2026 | cs.CV | — | 89 |