GPUFits

LLM VRAM Calculator

Calculate exact VRAM requirements for any LLM: weights, KV cache, and overhead by quantization level and context length. Free, data-sourced, no signup.

Weights
2.0 GB
KV cache
0.5 GB
Overhead
1.5 GB
Total
3.9 GB

Many runtimes default to a smaller context than the model maximum (e.g. Ollama defaults to 2048 tokens). Actual usage depends on your configuration.