Synced from Hugging Face — 2026-09-14
How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need?
Exact on-disk weight sizes below are pulled directly from Hugging Face. Total VRAM figures also include an estimated KV cache and runtime overhead — see the breakdown per quantization, and the calculator below to adjust context length.
By quantization
Meta-Llama-3.1-70B-Instruct-IQ1_M
15.60 GB exact, on disk
~18.54 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)
Meta-Llama-3.1-70B-Instruct-IQ2_M
22.46 GB exact, on disk
~26.08 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)
Meta-Llama-3.1-70B-Instruct-IQ2_S
20.71 GB exact, on disk
~24.16 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)
Meta-Llama-3.1-70B-Instruct-IQ2_XS
19.69 GB exact, on disk
~23.03 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)
Meta-Llama-3.1-70B-Instruct-IQ2_XXS
17.79 GB exact, on disk
~20.94 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)
Meta-Llama-3.1-70B-Instruct-IQ3_M
29.74 GB exact, on disk
~34.09 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)
Meta-Llama-3.1-70B-Instruct-IQ3_XS
27.29 GB exact, on disk
~31.39 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)
Meta-Llama-3.1-70B-Instruct-IQ4_XS
35.30 GB exact, on disk
~40.20 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)
Meta-Llama-3.1-70B-Instruct-Q2_K
24.56 GB exact, on disk
~28.39 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)
Meta-Llama-3.1-70B-Instruct-Q2_K_L
25.52 GB exact, on disk
~29.45 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)
Meta-Llama-3.1-70B-Instruct-Q3_K_L
34.59 GB exact, on disk
~39.42 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)
Meta-Llama-3.1-70B-Instruct-Q3_K_M
31.91 GB exact, on disk
~36.48 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)
Meta-Llama-3.1-70B-Instruct-Q3_K_S
28.79 GB exact, on disk
~33.04 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)
Meta-Llama-3.1-70B-Instruct-Q3_K_XL
35.45 GB exact, on disk
~40.37 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)
Meta-Llama-3.1-70B-Instruct-Q4_K_L
40.33 GB exact, on disk
~45.74 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)
Meta-Llama-3.1-70B-Instruct-Q4_K_M
39.60 GB exact, on disk
~44.94 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)
Meta-Llama-3.1-70B-Instruct-Q4_K_S
37.58 GB exact, on disk
~42.71 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)
Meta-Llama-3.1-70B-Instruct-Q5_K_L/Meta-Llama-3.1-70B-Instruct-Q5_K_L-00001-of-00002
37.26 GB exact, on disk
~42.36 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)
Meta-Llama-3.1-70B-Instruct-Q5_K_L/Meta-Llama-3.1-70B-Instruct-Q5_K_L-00002-of-00002
9.87 GB exact, on disk
~12.23 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)
Meta-Llama-3.1-70B-Instruct-Q5_K_M/Meta-Llama-3.1-70B-Instruct-Q5_K_M-00001-of-00002
37.14 GB exact, on disk
~42.23 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)
Meta-Llama-3.1-70B-Instruct-Q5_K_M/Meta-Llama-3.1-70B-Instruct-Q5_K_M-00002-of-00002
9.38 GB exact, on disk
~11.69 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)
Meta-Llama-3.1-70B-Instruct-Q5_K_S
45.32 GB exact, on disk
~51.23 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)
Meta-Llama-3.1-70B-Instruct-Q6_K/Meta-Llama-3.1-70B-Instruct-Q6_K-00001-of-00002
37.13 GB exact, on disk
~42.22 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)
Meta-Llama-3.1-70B-Instruct-Q6_K/Meta-Llama-3.1-70B-Instruct-Q6_K-00002-of-00002
16.79 GB exact, on disk
~19.84 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)
Meta-Llama-3.1-70B-Instruct-Q6_K_L/Meta-Llama-3.1-70B-Instruct-Q6_K_L-00001-of-00002
37.18 GB exact, on disk
~42.27 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)
Meta-Llama-3.1-70B-Instruct-Q6_K_L/Meta-Llama-3.1-70B-Instruct-Q6_K_L-00002-of-00002
17.20 GB exact, on disk
~20.29 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)
Meta-Llama-3.1-70B-Instruct-Q8_0/Meta-Llama-3.1-70B-Instruct-Q8_0-00001-of-00002
37.07 GB exact, on disk
~42.15 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)
Meta-Llama-3.1-70B-Instruct-Q8_0/Meta-Llama-3.1-70B-Instruct-Q8_0-00002-of-00002
32.75 GB exact, on disk
~37.40 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)
VRAM calculator
Estimate total VRAM for a given model size, quantization, and context length. Estimates only — see the note at the bottom.
- Model weights
- —
- KV cache
- —
- Runtime overhead (10%)
- —
- Total VRAM
- —
Fits on:
This won't fit on local hardware
Based on the estimate above, this configuration exceeds every local GPU profile listed. Renting cloud GPU capacity is the practical option for a workload this size.
RunPod
On-demand GPU pods, billed by the minute.
Lambda Labs
Reserved and on-demand cloud GPU instances.
Disclosure: We may earn a commission from cloud providers if you spin up an instance through these links, at no extra cost to you. This never affects the VRAM math above — see /methodology.
Live changes
- No changes yet — adjust a control above.
Generic estimate: weight size
is bits-per-weight × parameter count; KV cache uses a reference
architecture for the selected size bucket (7B/13B/70B: real published
models; 120B: extrapolated, no real model exists at that exact
size). Model-specific (HF-synced):
weight size is the exact .gguf file size fetched from
Hugging Face — no estimation. KV cache is still estimated:
this schema doesn't carry layer count, head count, or head dimension,
so the calculator parses an approximate parameter count from the
model's title (e.g. "8B") and reuses the nearest size bucket's
reference architecture for the KV math, same caveats as the generic
mode. If no parameter count can be parsed from the title, it falls
back to the 7B architecture and says so next to the quantization
dropdown. Batch size is fixed at 1 in both modes. Treat all figures
here as a starting estimate, not a guarantee. Full formula and known
limitations: /methodology.
FAQ
How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-IQ1_M?
The Meta-Llama-3.1-70B-Instruct-IQ1_M quantization is 15.60 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 18.54 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.
How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-IQ2_M?
The Meta-Llama-3.1-70B-Instruct-IQ2_M quantization is 22.46 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 26.08 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.
How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-IQ2_S?
The Meta-Llama-3.1-70B-Instruct-IQ2_S quantization is 20.71 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 24.16 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.
How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-IQ2_XS?
The Meta-Llama-3.1-70B-Instruct-IQ2_XS quantization is 19.69 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 23.03 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.
How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-IQ2_XXS?
The Meta-Llama-3.1-70B-Instruct-IQ2_XXS quantization is 17.79 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 20.94 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.
How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-IQ3_M?
The Meta-Llama-3.1-70B-Instruct-IQ3_M quantization is 29.74 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 34.09 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.
How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-IQ3_XS?
The Meta-Llama-3.1-70B-Instruct-IQ3_XS quantization is 27.29 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 31.39 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.
How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-IQ4_XS?
The Meta-Llama-3.1-70B-Instruct-IQ4_XS quantization is 35.30 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 40.20 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.
How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q2_K?
The Meta-Llama-3.1-70B-Instruct-Q2_K quantization is 24.56 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 28.39 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.
How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q2_K_L?
The Meta-Llama-3.1-70B-Instruct-Q2_K_L quantization is 25.52 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 29.45 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.
How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q3_K_L?
The Meta-Llama-3.1-70B-Instruct-Q3_K_L quantization is 34.59 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 39.42 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.
How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q3_K_M?
The Meta-Llama-3.1-70B-Instruct-Q3_K_M quantization is 31.91 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 36.48 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.
How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q3_K_S?
The Meta-Llama-3.1-70B-Instruct-Q3_K_S quantization is 28.79 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 33.04 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.
How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q3_K_XL?
The Meta-Llama-3.1-70B-Instruct-Q3_K_XL quantization is 35.45 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 40.37 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.
How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q4_K_L?
The Meta-Llama-3.1-70B-Instruct-Q4_K_L quantization is 40.33 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 45.74 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.
How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q4_K_M?
The Meta-Llama-3.1-70B-Instruct-Q4_K_M quantization is 39.60 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 44.94 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.
How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q4_K_S?
The Meta-Llama-3.1-70B-Instruct-Q4_K_S quantization is 37.58 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 42.71 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.
How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q5_K_L/Meta-Llama-3.1-70B-Instruct-Q5_K_L-00001-of-00002?
The Meta-Llama-3.1-70B-Instruct-Q5_K_L/Meta-Llama-3.1-70B-Instruct-Q5_K_L-00001-of-00002 quantization is 37.26 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 42.36 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.
How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q5_K_L/Meta-Llama-3.1-70B-Instruct-Q5_K_L-00002-of-00002?
The Meta-Llama-3.1-70B-Instruct-Q5_K_L/Meta-Llama-3.1-70B-Instruct-Q5_K_L-00002-of-00002 quantization is 9.87 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 12.23 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.
How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q5_K_M/Meta-Llama-3.1-70B-Instruct-Q5_K_M-00001-of-00002?
The Meta-Llama-3.1-70B-Instruct-Q5_K_M/Meta-Llama-3.1-70B-Instruct-Q5_K_M-00001-of-00002 quantization is 37.14 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 42.23 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.
How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q5_K_M/Meta-Llama-3.1-70B-Instruct-Q5_K_M-00002-of-00002?
The Meta-Llama-3.1-70B-Instruct-Q5_K_M/Meta-Llama-3.1-70B-Instruct-Q5_K_M-00002-of-00002 quantization is 9.38 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 11.69 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.
How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q5_K_S?
The Meta-Llama-3.1-70B-Instruct-Q5_K_S quantization is 45.32 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 51.23 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.
How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q6_K/Meta-Llama-3.1-70B-Instruct-Q6_K-00001-of-00002?
The Meta-Llama-3.1-70B-Instruct-Q6_K/Meta-Llama-3.1-70B-Instruct-Q6_K-00001-of-00002 quantization is 37.13 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 42.22 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.
How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q6_K/Meta-Llama-3.1-70B-Instruct-Q6_K-00002-of-00002?
The Meta-Llama-3.1-70B-Instruct-Q6_K/Meta-Llama-3.1-70B-Instruct-Q6_K-00002-of-00002 quantization is 16.79 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 19.84 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.
How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q6_K_L/Meta-Llama-3.1-70B-Instruct-Q6_K_L-00001-of-00002?
The Meta-Llama-3.1-70B-Instruct-Q6_K_L/Meta-Llama-3.1-70B-Instruct-Q6_K_L-00001-of-00002 quantization is 37.18 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 42.27 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.
How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q6_K_L/Meta-Llama-3.1-70B-Instruct-Q6_K_L-00002-of-00002?
The Meta-Llama-3.1-70B-Instruct-Q6_K_L/Meta-Llama-3.1-70B-Instruct-Q6_K_L-00002-of-00002 quantization is 17.20 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 20.29 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.
How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q8_0/Meta-Llama-3.1-70B-Instruct-Q8_0-00001-of-00002?
The Meta-Llama-3.1-70B-Instruct-Q8_0/Meta-Llama-3.1-70B-Instruct-Q8_0-00001-of-00002 quantization is 37.07 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 42.15 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.
How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q8_0/Meta-Llama-3.1-70B-Instruct-Q8_0-00002-of-00002?
The Meta-Llama-3.1-70B-Instruct-Q8_0/Meta-Llama-3.1-70B-Instruct-Q8_0-00002-of-00002 quantization is 32.75 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 37.40 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.