Skip to content
AI Atlas
Research

Papers

Publications linked to models, labs and benchmarks. Authors, venues and abstracts come from arXiv and publisher pages; model links come from model cards citing the paper.

24 papers

Papers
TitleAuthorsOrganizationPublishedIntroduces Models (and artifacts) whose model card or documentation cites this paper — inbound described_by relations.DatasetsBenchmarksCode
Forward-Free LLM Depth Pruning via Weight RedundancyarXiv:2609.09883cs.LGVincent-Daniel Yun, Woosang Lim16 Sept 2026
GGUF-Metadata Prediction of Single-Sequence llama.cpp Throughput Across Three SystemsarXiv:2609.14864cs.AIBo Su, Chuhong Xu, Ruiyang Xu +216 Sept 2026
Validating Hybrid-State Cache Recovery for GLM-5.3-Flash with vLLM and LMCachearXiv:2609.15030cs.DCFrank Li16 Sept 2026
Turkish MMLU Pro: Traceable Option Augmentation and Its Validity Limits in Turkish Multiple-Choice EvaluationarXiv:2609.15467cs.CLM. Ali Bayram16 Sept 2026
Is INT8 Portable? A Cross-Platform Measurement Study of Quantized Inference on Embedded and Automotive AcceleratorsarXiv:2609.16085cs.ARYuyeong Shin16 Sept 2026
The AI-Enabled Scientific FrontierarXiv:2609.16258cs.AIEmma Fu, Gabriel Manso, Neil Thompson16 Sept 2026
Partition-Aware Scheduling for Mobile Heterogeneous Inference Co-ExecutionarXiv:2609.14213cs.DCLeana Golubchik, Marco Paolieri, Zhuojin Li15 Sept 2026
Learning Metastable DynamicsarXiv:2609.14712cs.PFMahmoud Salamati, Nikhil Singh, Rupak Majumdar +115 Sept 2026
HELENA for 5G NR LEO NTN Channel Estimation: A Comparative EvaluationarXiv:2609.14735eess.SPJohann Marquez-Barja, Miguel Camelo Botero, Nina Slamnik-Krije\v{s}torac15 Sept 2026
Shared KV Caching for Replicated 27B Inference: Correctness Failures and Performance BoundariesarXiv:2609.15021cs.DCFrank Li15 Sept 2026
Accelerating Transfer-Learning-Based Autotuning with Predictive LLVM IR Performance RankingarXiv:2609.15807cs.PFAkash Dutta, Ali Jannesari, Md Arafat Hossain +215 Sept 2026
Studying quantization trade-offs for efficient inference deployment in machine translationarXiv:2607.29397cs.CLDouglas Orr, Jim Zhao, Koen Oostermeijer +215 Sept 2026
One Simple Trick for Improving the Performance of Energy-Limited Local Inference and TrainingarXiv:2609.11936cs.PFDan Alistarh, Erik Schultheis, Maximilian Kleinegger14 Sept 2026
The Battery Price of edge AI: A study of the Environmental Impact of LLM Inference on Mobile DevicesarXiv:2609.11940cs.PFTristan Coignion, \'Edouard Gu\'egain14 Sept 2026
Efficient Vision-Language-Action Management and Serving for Robot FactoriesarXiv:2609.12075cs.DCBasel Fakhri, Christina Giannoula, Dionysios Adamopoulos +114 Sept 2026
HoliBench: A Cross-Platform Benchmarking and Deployment Toolkit for Foundation Models in CPS-IoT ApplicationsarXiv:2609.12412cs.DCInesh Chakrabarti, Mani Srivastava, Pragya Sharma +114 Sept 2026
4D Parallelism Unlocks Exascale Bayesian Neural Networks for High-Fidelity Atmospheric ModelingarXiv:2609.12815physics.ao-phAndreas Herten, Anni Moisala, Arvid Weyrauch +214 Sept 2026
Dissecting GPU Utilization for LLM Inference on Nvidia HopperarXiv:2609.12923cs.PFDejan Kostic, Gerald Q. Maguire Jr., Marco Chiesa +114 Sept 2026
UltraQuant: 4-bit KV Caching for Context-Heavy AgentsarXiv:2606.20474cs.LGAditi Ghai Rana, Ashish Sirasao, Bowen Bao +214 Sept 2026
Dynamic Expert Quantization for Scalable Mixture-of-Experts InferencearXiv:2511.15015cs.PFDawei Xiang, Kexin Chu, Wei Zhang +214 Sept 2026
Input Resolution Matters: Real-Time Object Detection LatencyarXiv:2609.12920cs.CVFumio Machida, Laura Carnevali, Qingyang Zhang14 Sept 2026
Elastoformer: Enabling Dynamic Adaptivity via Elastic Model TransformationarXiv:2609.10018cs.CVSudaksh Kalra, Dolly Sapra11 Sept 2026
FP8 is All You Need (Part 1): Debunking Hardware FP64 as the HPC Holy Grail (Sep 3rd version)arXiv:2606.06510cs.ARSatoshi Matsuoka11 Sept 2026
FP8 is All You Need (Part 2): Full-FP64 3-D FFT on FP8-Generation Tensor CoresThe Integer-Epilogue Wall and the Minimal Hardware That Would Remove ItarXiv:2606.23698cs.MSSatoshi Matsuoka11 Sept 2026

24 results

Author lists and categories are copied from the paper's own metadata (arXiv, publisher). Linked models, datasets, benchmarks and code come from stated relations only; a dash means no source stated one.