Ollama
Score: โOne-command local prototyping and personal agents.
VRAM
โ
Tok/s
โ
Weighted scoring for Ollama, vLLM, SGLang, TensorRT-LLM, MLX, ExLlamaV2, KTransformers, and more. Pick by GPU, model size, quantization, batch size, and priority.
Select a preset to auto-tune weights, batch size, and recommendation.
Recommended engine
โ
Score: โ
VRAM needed
โ
Est. tok/s
โ
Select your setup to see a recommendation.
Model weights: โ
KV cache: โ
Total estimate: โ
GPU VRAM: โ
One-command local prototyping and personal agents.
VRAM
โ
Tok/s
โ
Maximum hardware compatibility and edge deployment.
VRAM
โ
Tok/s
โ
High-throughput multi-user serving.
VRAM
โ
Tok/s
โ
Structured generation, batch programming, and fast multi-turn.
VRAM
โ
Tok/s
โ
Maximum NVIDIA throughput with pre-compiled engines.
VRAM
โ
Tok/s
โ
Native Apple Silicon performance with unified memory.
VRAM
โ
Tok/s
โ
Fast single-GPU EXL2/GPTQ serving with OpenAI-compatible API.
VRAM
โ
Tok/s
โ
Maximum 4-bit throughput on a single NVIDIA GPU.
VRAM
โ
Tok/s
โ
Reference implementation and newest models.
VRAM
โ
Tok/s
โ
Portable single-file model distribution.
VRAM
โ
Tok/s
โ
Storytelling/roleplay with built-in web UI.
VRAM
โ
Tok/s
โ
Production serving of Hugging Face models with sharding.
VRAM
โ
Tok/s
โ
Very large MoE and dense models with CPU+GPU offloading.
VRAM
โ
Tok/s
โ
Cross-platform deployment and optimized transformer models.
VRAM
โ
Tok/s
โ
| Engine | Formats | Platforms | OpenAI API | Batching | Multi-GPU | Structured | License |
|---|---|---|---|---|---|---|---|
| Ollama | GGUF | Macos, Linux, Windows | โ | Basic | โ | โ | MIT |
| llama.cpp | GGUF | Macos, Linux, Windows | โ | Basic | โ | โ | MIT |
| vLLM | Safetensors, PyTorch, AWQ, GPTQ, GGUF | Linux, Windows | โ | Continuous | โ | โ | Apache 2.0 |
| SGLang | Safetensors, PyTorch, AWQ, GPTQ, FP8 | Linux | โ | Continuous | โ | โ | Apache 2.0 |
| TensorRT-LLM | TensorRT-LLM engines | Linux, Windows | โ | Continuous | โ | โ | Apache 2.0 |
| MLX (Apple Silicon) | Safetensors, MLX | Macos | โ | Basic | โ | โ | MIT |
| TabbyAPI | EXL2, GPTQ | Linux, Windows | โ | Continuous | โ | โ | AGPL |
| ExLlamaV2 | EXL2, GPTQ | Linux, Windows | โ | Continuous | โ | โ | MIT |
| Hugging Face Transformers | Safetensors, PyTorch | Macos, Linux, Windows | โ | Basic | โ | โ | Apache 2.0 |
| llamafile | GGUF | Macos, Linux, Windows | โ | Basic | โ | โ | Apache 2.0 |
| KoboldCpp | GGUF | Windows, Linux, Macos | โ | Basic | โ | โ | AGPL |
| Text Generation Inference (TGI) | Safetensors, PyTorch | Linux | โ | Continuous | โ | โ | Apache 2.0/Hugging Face |
| KTransformers | GGUF, Safetensors | Linux, Windows | โ | Basic | โ | โ | MIT |
| ONNX Runtime | ONNX | Macos, Linux, Windows | โ | Basic | โ | โ | MIT |
For 4-bit quantized dense models, TabbyAPI / ExLlamaV2 or TensorRT-LLM are usually fastest. For native FP16/BF16 models, vLLM and SGLang extract the highest throughput thanks to continuous batching.
Use vLLM or SGLang for open-model serving with continuous batching and OpenAI-compatible endpoints. If you only serve NVIDIA and can pre-compile engines, TensorRT-LLM can squeeze out extra throughput. TGI is also a proven production option.
Not fully loaded in FP16. Use KTransformers or llama.cpp with CPU/NVMe offloading, or quantize to Q4 GGUF. Expect lower tok/s because expert layers move through system memory.
MLX is purpose-built for M-series chips and gives the best tokens-per-watt. Ollama and llama.cpp are also excellent if you already have GGUF models.
Only if you serve multiple concurrent users. Single-user chat or local agents rarely benefit; one of the simpler GGUF engines is often easier to set up.
No. The calculator uses directional multipliers based on memory bandwidth, quantization, and engine efficiency. Real numbers depend on drivers, prompt length, KV cache, and PCIe/NVLink topology.
Scores and tok/s estimates are directional, based on public benchmarks, release notes, and typical user reports. Real-world performance depends on driver version, quantization recipe, context length, batching, CPU overhead, and PCIe/NVLink topology.
๐ Get AI automation insights daily
15:00 MST. One-click unsubscribe.