Frameworks
Frameworks & runtimes
Libraries, inference engines and agent frameworks. Versions and stars are read from PyPI, GitHub and release pages — never estimated.
66 total
| Framework | Latest version | Release date | License | Stars | Language | Organization | Quality | Compare |
|---|---|---|---|---|---|---|---|---|
| text-generation-inferenceLarge Language Model Text Generation Inference | 3.3.7 | 19 Dec 2025 | Apache-2.0 | 10,886 | — | Hugging Face | ||
| onnxruntimeONNX Runtime: cross-platform, high performance ML inferencing and training accelerator | ONNX Runtime v1.30.0 | 10 Sept 2026 | MIT | 21,827 | — | Microsoft | ||
| TensorRT-LLMTensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate | 1.3.0rc26 | 9 Sept 2026 | — | 14,597 | — | NVIDIA | ||
| sglangSGLang is a high-performance serving framework for large language models and multimodal models. | 0.5.19 | 5 Sept 2026 | Apache-2.0 | 35,824 | — | SGLang project | ||
| ollamaGet up and running with Kimi-K2.6, GLM-5.2, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models. | 0.34.0 | 10 Sept 2026 | MIT | 180,690 | — | Ollama | ||
| mlx-lmRun LLMs with MLX | 0.31.3 | 22 Apr 2026 | MIT | 6,976 | — | Apple ML Explore (MLX) | ||
| mlxMLX: An array framework for Apple silicon | 0.32.2 | 25 Aug 2026 | MIT | 28,386 | — | Apple ML Explore (MLX) | ||
| llama-cpp-pythonPython bindings for llama.cpp | 0.3.35-cu123 | 17 Aug 2026 | MIT | 10,615 | — | — | ||
| llama.cppLLM inference in C/C++ | b10917 | 11 Sept 2026 | MIT | 127,879 | — | ggml.ai | ||
| vllmA high-throughput and memory-efficient inference and serving engine for LLMs | 0.29.0 | 10 Sept 2026 | Apache-2.0 | 91,519 | — | vLLM project | ||
| nanoGPTThe simplest, fastest repository for training/finetuning medium-sized GPTs. | — | — | MIT | 62,999 | — | — | ||
| SpeechA scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech) | NVIDIA NeMo Speech 3.0 | 7 Aug 2026 | Apache-2.0 | 18,424 | — | NVIDIA | ||
| flash-attentionFast and memory-efficient exact attention | fa4-v4.0.0.beta30 | 9 Sept 2026 | BSD-3-Clause | 24,890 | — | Dao AI Lab | ||
| tritonDevelopment repository for the Triton language and compiler | gfx950-tutorial-v2.1 | 31 Aug 2026 | MIT | 20,132 | — | OpenAI | ||
| BentoMLThe easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more! | 1.4.39 | 7 May 2026 | Apache-2.0 | 8,834 | — | BentoML | ||
| kubeflowMachine Learning Toolkit for Kubernetes | 1.10.0 | 25 Mar 2025 | Apache-2.0 | 15,860 | — | Kubeflow | ||
| rayRay is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads. | Ray-2.58.0 | 23 Aug 2026 | Apache-2.0 | 43,782 | — | Anyscale | ||
| pytorch-lightningPretrain, finetune ANY AI model of ANY size on 1 or 10,000+ GPUs with zero code changes. | Lightning 2.6.6 | 10 Sept 2026 | Apache-2.0 | 31,336 | — | Lightning AI | ||
| axolotlGo ahead and axolotl questions | 0.19.0 | 10 Sept 2026 | Apache-2.0 | 12,461 | — | Axolotl AI | ||
| unslothLocal UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, DeepSeek-V4, MiniMax-H3, Gemma 4, FLUX and more. | Large Performance Gains + Fixes | 9 Sept 2026 | Apache-2.0 | 76,034 | — | Unsloth | ||
| trlTrain transformer language models with reinforcement learning. | 1.13.0 | 10 Sept 2026 | Apache-2.0 | 19,286 | — | Hugging Face | ||
| peft🤗 PEFT: State-of-the-art Parameter-Efficient Fine-Tuning. | 0.20.0 | 28 Jul 2026 | Apache-2.0 | 21,657 | — | Hugging Face | ||
| diffusers🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch. | Diffusers 0.40.0: New pipelines, tensor-parallel support, improved CLI, and more | 20 Aug 2026 | Apache-2.0 | 34,497 | — | Hugging Face | ||
| jaxComposable transformations of Python+NumPy programs: differentiate, vectorize, JIT to GPU/TPU, and more | JAX v0.11.1 | 17 Aug 2026 | Apache-2.0 | 36,295 | — | |||
| transformers🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training. | Release 5.17.0 | 10 Sept 2026 | Apache-2.0 | 165,129 | — | Hugging Face | ||
| pytorchTensors and Dynamic neural networks in Python with strong GPU acceleration | viable/strict/1789158535: [ROCm][CD] Add gfx1250 (MI450) to the nightly wheel arch list (#196610) | 11 Sept 2026 | — | 102,930 | — | PyTorch Foundation |
Stars are a GitHub metric observed at crawl time (see each framework's Sources tab for the snapshot). Versions come from PyPI or the project's release feed.