| Forward-Free LLM Depth Pruning via Weight RedundancyarXiv:2609.09883cs.LG | Vincent-Daniel Yun, Woosang Lim | — | 16 Sept 2026 | — | — | — | — |
| GGUF-Metadata Prediction of Single-Sequence llama.cpp Throughput Across Three SystemsarXiv:2609.14864cs.AI | Bo Su, Chuhong Xu, Ruiyang Xu +2 | — | 16 Sept 2026 | — | — | — | — |
| Validating Hybrid-State Cache Recovery for GLM-5.3-Flash with vLLM and LMCachearXiv:2609.15030cs.DC | Frank Li | — | 16 Sept 2026 | — | — | — | — |
| Turkish MMLU Pro: Traceable Option Augmentation and Its Validity Limits in Turkish Multiple-Choice EvaluationarXiv:2609.15467cs.CL | M. Ali Bayram | — | 16 Sept 2026 | — | — | — | — |
| Is INT8 Portable? A Cross-Platform Measurement Study of Quantized Inference on Embedded and Automotive AcceleratorsarXiv:2609.16085cs.AR | Yuyeong Shin | — | 16 Sept 2026 | — | — | — | — |
| The AI-Enabled Scientific FrontierarXiv:2609.16258cs.AI | Emma Fu, Gabriel Manso, Neil Thompson | — | 16 Sept 2026 | — | — | — | — |
| Partition-Aware Scheduling for Mobile Heterogeneous Inference Co-ExecutionarXiv:2609.14213cs.DC | Leana Golubchik, Marco Paolieri, Zhuojin Li | — | 15 Sept 2026 | — | — | — | — |
| Learning Metastable DynamicsarXiv:2609.14712cs.PF | Mahmoud Salamati, Nikhil Singh, Rupak Majumdar +1 | — | 15 Sept 2026 | — | — | — | — |
| HELENA for 5G NR LEO NTN Channel Estimation: A Comparative EvaluationarXiv:2609.14735eess.SP | Johann Marquez-Barja, Miguel Camelo Botero, Nina Slamnik-Krije\v{s}torac | — | 15 Sept 2026 | — | — | — | — |
| Shared KV Caching for Replicated 27B Inference: Correctness Failures and Performance BoundariesarXiv:2609.15021cs.DC | Frank Li | — | 15 Sept 2026 | — | — | — | — |
| Accelerating Transfer-Learning-Based Autotuning with Predictive LLVM IR Performance RankingarXiv:2609.15807cs.PF | Akash Dutta, Ali Jannesari, Md Arafat Hossain +2 | — | 15 Sept 2026 | — | — | — | — |
| Studying quantization trade-offs for efficient inference deployment in machine translationarXiv:2607.29397cs.CL | Douglas Orr, Jim Zhao, Koen Oostermeijer +2 | — | 15 Sept 2026 | — | — | — | — |
| One Simple Trick for Improving the Performance of Energy-Limited Local Inference and TrainingarXiv:2609.11936cs.PF | Dan Alistarh, Erik Schultheis, Maximilian Kleinegger | — | 14 Sept 2026 | — | — | — | — |
| The Battery Price of edge AI: A study of the Environmental Impact of LLM Inference on Mobile DevicesarXiv:2609.11940cs.PF | Tristan Coignion, \'Edouard Gu\'egain | — | 14 Sept 2026 | — | — | — | — |
| Efficient Vision-Language-Action Management and Serving for Robot FactoriesarXiv:2609.12075cs.DC | Basel Fakhri, Christina Giannoula, Dionysios Adamopoulos +1 | — | 14 Sept 2026 | — | — | — | — |
| HoliBench: A Cross-Platform Benchmarking and Deployment Toolkit for Foundation Models in CPS-IoT ApplicationsarXiv:2609.12412cs.DC | Inesh Chakrabarti, Mani Srivastava, Pragya Sharma +1 | — | 14 Sept 2026 | — | — | — | — |
| 4D Parallelism Unlocks Exascale Bayesian Neural Networks for High-Fidelity Atmospheric ModelingarXiv:2609.12815physics.ao-ph | Andreas Herten, Anni Moisala, Arvid Weyrauch +2 | — | 14 Sept 2026 | — | — | — | — |
| Dissecting GPU Utilization for LLM Inference on Nvidia HopperarXiv:2609.12923cs.PF | Dejan Kostic, Gerald Q. Maguire Jr., Marco Chiesa +1 | — | 14 Sept 2026 | — | — | — | — |
| UltraQuant: 4-bit KV Caching for Context-Heavy AgentsarXiv:2606.20474cs.LG | Aditi Ghai Rana, Ashish Sirasao, Bowen Bao +2 | — | 14 Sept 2026 | — | — | — | — |
| Dynamic Expert Quantization for Scalable Mixture-of-Experts InferencearXiv:2511.15015cs.PF | Dawei Xiang, Kexin Chu, Wei Zhang +2 | — | 14 Sept 2026 | — | — | — | — |
| Input Resolution Matters: Real-Time Object Detection LatencyarXiv:2609.12920cs.CV | Fumio Machida, Laura Carnevali, Qingyang Zhang | — | 14 Sept 2026 | — | — | — | — |
| Elastoformer: Enabling Dynamic Adaptivity via Elastic Model TransformationarXiv:2609.10018cs.CV | Sudaksh Kalra, Dolly Sapra | — | 11 Sept 2026 | — | — | — | — |
| FP8 is All You Need (Part 1): Debunking Hardware FP64 as the HPC Holy Grail (Sep 3rd version)arXiv:2606.06510cs.AR | Satoshi Matsuoka | — | 11 Sept 2026 | — | — | — | — |
| FP8 is All You Need (Part 2): Full-FP64 3-D FFT on FP8-Generation Tensor CoresThe Integer-Epilogue Wall and the Minimal Hardware That Would Remove ItarXiv:2606.23698cs.MS | Satoshi Matsuoka | — | 11 Sept 2026 | — | — | — | — |