LocalNodeOps

Synced from Hugging Face — 2026-09-14

How much VRAM does gemma-2-27b-it-GGUF need?

Exact on-disk weight sizes below are pulled directly from Hugging Face. Total VRAM figures also include an estimated KV cache and runtime overhead — see the breakdown per quantization, and the calculator below to adjust context length.

By quantization

gemma-2-27b-it-IQ2_M

8.75 GB exact, on disk

~13.06 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

gemma-2-27b-it-IQ2_S

8.06 GB exact, on disk

~12.30 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

gemma-2-27b-it-IQ2_XS

7.82 GB exact, on disk

~12.04 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

gemma-2-27b-it-IQ3_M

11.60 GB exact, on disk

~16.20 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

gemma-2-27b-it-IQ3_XS

10.76 GB exact, on disk

~15.27 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

gemma-2-27b-it-IQ3_XXS

10.01 GB exact, on disk

~14.45 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

gemma-2-27b-it-IQ4_XS

13.80 GB exact, on disk

~18.62 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

gemma-2-27b-it-Q2_K

9.73 GB exact, on disk

~14.14 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

gemma-2-27b-it-Q2_K_L

10.00 GB exact, on disk

~14.44 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

gemma-2-27b-it-Q3_K_L

13.52 GB exact, on disk

~18.31 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

gemma-2-27b-it-Q3_K_M

12.50 GB exact, on disk

~17.19 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

gemma-2-27b-it-Q3_K_S

11.33 GB exact, on disk

~15.90 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

gemma-2-27b-it-Q3_K_XL

13.79 GB exact, on disk

~18.61 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

gemma-2-27b-it-Q4_K_L

15.77 GB exact, on disk

~20.78 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

gemma-2-27b-it-Q4_K_M

15.50 GB exact, on disk

~20.49 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

gemma-2-27b-it-Q4_K_S

14.66 GB exact, on disk

~19.56 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

gemma-2-27b-it-Q5_K_L

18.34 GB exact, on disk

~23.61 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

gemma-2-27b-it-Q5_K_M

18.08 GB exact, on disk

~23.33 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

gemma-2-27b-it-Q5_K_S

17.59 GB exact, on disk

~22.79 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

gemma-2-27b-it-Q6_K

20.81 GB exact, on disk

~26.33 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

gemma-2-27b-it-Q6_K_L

21.08 GB exact, on disk

~26.63 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

gemma-2-27b-it-Q8_0

26.95 GB exact, on disk

~33.08 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

gemma-2-27b-it-Q8_0_L

27.98 GB exact, on disk

~34.22 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

gemma-2-27b-it-f32.gguf/gemma-2-27b-it-f32-00001-of-00003

36.89 GB exact, on disk

~44.02 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

gemma-2-27b-it-f32.gguf/gemma-2-27b-it-f32-00002-of-00003

37.13 GB exact, on disk

~44.28 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

gemma-2-27b-it-f32.gguf/gemma-2-27b-it-f32-00003-of-00003

27.42 GB exact, on disk

~33.60 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

gemma-2-27b-it-f32/gemma-2-27b-it-f32-00001-of-00003

36.89 GB exact, on disk

~44.02 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

gemma-2-27b-it-f32/gemma-2-27b-it-f32-00002-of-00003

37.13 GB exact, on disk

~44.28 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

gemma-2-27b-it-f32/gemma-2-27b-it-f32-00003-of-00003

27.42 GB exact, on disk

~33.60 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

VRAM calculator

Estimate total VRAM for a given model size, quantization, and context length. Estimates only — see the note at the bottom.

7B13B70B120B
2K4K8K16K32K64K128K
Model weights
KV cache
Runtime overhead (10%)
Total VRAM

Fits on:

    Live changes

    • No changes yet — adjust a control above.

    Generic estimate: weight size is bits-per-weight × parameter count; KV cache uses a reference architecture for the selected size bucket (7B/13B/70B: real published models; 120B: extrapolated, no real model exists at that exact size). Model-specific (HF-synced): weight size is the exact .gguf file size fetched from Hugging Face — no estimation. KV cache is still estimated: this schema doesn't carry layer count, head count, or head dimension, so the calculator parses an approximate parameter count from the model's title (e.g. "8B") and reuses the nearest size bucket's reference architecture for the KV math, same caveats as the generic mode. If no parameter count can be parsed from the title, it falls back to the 7B architecture and says so next to the quantization dropdown. Batch size is fixed at 1 in both modes. Treat all figures here as a starting estimate, not a guarantee. Full formula and known limitations: /methodology.

    FAQ

    How much VRAM does gemma-2-27b-it-GGUF need at gemma-2-27b-it-IQ2_M?

    The gemma-2-27b-it-IQ2_M quantization is 8.75 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 13B reference architecture for the KV cache estimate, total estimated VRAM is approximately 13.06 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does gemma-2-27b-it-GGUF need at gemma-2-27b-it-IQ2_S?

    The gemma-2-27b-it-IQ2_S quantization is 8.06 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 13B reference architecture for the KV cache estimate, total estimated VRAM is approximately 12.30 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does gemma-2-27b-it-GGUF need at gemma-2-27b-it-IQ2_XS?

    The gemma-2-27b-it-IQ2_XS quantization is 7.82 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 13B reference architecture for the KV cache estimate, total estimated VRAM is approximately 12.04 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does gemma-2-27b-it-GGUF need at gemma-2-27b-it-IQ3_M?

    The gemma-2-27b-it-IQ3_M quantization is 11.60 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 13B reference architecture for the KV cache estimate, total estimated VRAM is approximately 16.20 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does gemma-2-27b-it-GGUF need at gemma-2-27b-it-IQ3_XS?

    The gemma-2-27b-it-IQ3_XS quantization is 10.76 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 13B reference architecture for the KV cache estimate, total estimated VRAM is approximately 15.27 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does gemma-2-27b-it-GGUF need at gemma-2-27b-it-IQ3_XXS?

    The gemma-2-27b-it-IQ3_XXS quantization is 10.01 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 13B reference architecture for the KV cache estimate, total estimated VRAM is approximately 14.45 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does gemma-2-27b-it-GGUF need at gemma-2-27b-it-IQ4_XS?

    The gemma-2-27b-it-IQ4_XS quantization is 13.80 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 13B reference architecture for the KV cache estimate, total estimated VRAM is approximately 18.62 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does gemma-2-27b-it-GGUF need at gemma-2-27b-it-Q2_K?

    The gemma-2-27b-it-Q2_K quantization is 9.73 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 13B reference architecture for the KV cache estimate, total estimated VRAM is approximately 14.14 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does gemma-2-27b-it-GGUF need at gemma-2-27b-it-Q2_K_L?

    The gemma-2-27b-it-Q2_K_L quantization is 10.00 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 13B reference architecture for the KV cache estimate, total estimated VRAM is approximately 14.44 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does gemma-2-27b-it-GGUF need at gemma-2-27b-it-Q3_K_L?

    The gemma-2-27b-it-Q3_K_L quantization is 13.52 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 13B reference architecture for the KV cache estimate, total estimated VRAM is approximately 18.31 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does gemma-2-27b-it-GGUF need at gemma-2-27b-it-Q3_K_M?

    The gemma-2-27b-it-Q3_K_M quantization is 12.50 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 13B reference architecture for the KV cache estimate, total estimated VRAM is approximately 17.19 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does gemma-2-27b-it-GGUF need at gemma-2-27b-it-Q3_K_S?

    The gemma-2-27b-it-Q3_K_S quantization is 11.33 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 13B reference architecture for the KV cache estimate, total estimated VRAM is approximately 15.90 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does gemma-2-27b-it-GGUF need at gemma-2-27b-it-Q3_K_XL?

    The gemma-2-27b-it-Q3_K_XL quantization is 13.79 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 13B reference architecture for the KV cache estimate, total estimated VRAM is approximately 18.61 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does gemma-2-27b-it-GGUF need at gemma-2-27b-it-Q4_K_L?

    The gemma-2-27b-it-Q4_K_L quantization is 15.77 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 13B reference architecture for the KV cache estimate, total estimated VRAM is approximately 20.78 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does gemma-2-27b-it-GGUF need at gemma-2-27b-it-Q4_K_M?

    The gemma-2-27b-it-Q4_K_M quantization is 15.50 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 13B reference architecture for the KV cache estimate, total estimated VRAM is approximately 20.49 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does gemma-2-27b-it-GGUF need at gemma-2-27b-it-Q4_K_S?

    The gemma-2-27b-it-Q4_K_S quantization is 14.66 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 13B reference architecture for the KV cache estimate, total estimated VRAM is approximately 19.56 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does gemma-2-27b-it-GGUF need at gemma-2-27b-it-Q5_K_L?

    The gemma-2-27b-it-Q5_K_L quantization is 18.34 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 13B reference architecture for the KV cache estimate, total estimated VRAM is approximately 23.61 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does gemma-2-27b-it-GGUF need at gemma-2-27b-it-Q5_K_M?

    The gemma-2-27b-it-Q5_K_M quantization is 18.08 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 13B reference architecture for the KV cache estimate, total estimated VRAM is approximately 23.33 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does gemma-2-27b-it-GGUF need at gemma-2-27b-it-Q5_K_S?

    The gemma-2-27b-it-Q5_K_S quantization is 17.59 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 13B reference architecture for the KV cache estimate, total estimated VRAM is approximately 22.79 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does gemma-2-27b-it-GGUF need at gemma-2-27b-it-Q6_K?

    The gemma-2-27b-it-Q6_K quantization is 20.81 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 13B reference architecture for the KV cache estimate, total estimated VRAM is approximately 26.33 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does gemma-2-27b-it-GGUF need at gemma-2-27b-it-Q6_K_L?

    The gemma-2-27b-it-Q6_K_L quantization is 21.08 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 13B reference architecture for the KV cache estimate, total estimated VRAM is approximately 26.63 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does gemma-2-27b-it-GGUF need at gemma-2-27b-it-Q8_0?

    The gemma-2-27b-it-Q8_0 quantization is 26.95 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 13B reference architecture for the KV cache estimate, total estimated VRAM is approximately 33.08 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does gemma-2-27b-it-GGUF need at gemma-2-27b-it-Q8_0_L?

    The gemma-2-27b-it-Q8_0_L quantization is 27.98 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 13B reference architecture for the KV cache estimate, total estimated VRAM is approximately 34.22 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does gemma-2-27b-it-GGUF need at gemma-2-27b-it-f32.gguf/gemma-2-27b-it-f32-00001-of-00003?

    The gemma-2-27b-it-f32.gguf/gemma-2-27b-it-f32-00001-of-00003 quantization is 36.89 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 13B reference architecture for the KV cache estimate, total estimated VRAM is approximately 44.02 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does gemma-2-27b-it-GGUF need at gemma-2-27b-it-f32.gguf/gemma-2-27b-it-f32-00002-of-00003?

    The gemma-2-27b-it-f32.gguf/gemma-2-27b-it-f32-00002-of-00003 quantization is 37.13 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 13B reference architecture for the KV cache estimate, total estimated VRAM is approximately 44.28 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does gemma-2-27b-it-GGUF need at gemma-2-27b-it-f32.gguf/gemma-2-27b-it-f32-00003-of-00003?

    The gemma-2-27b-it-f32.gguf/gemma-2-27b-it-f32-00003-of-00003 quantization is 27.42 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 13B reference architecture for the KV cache estimate, total estimated VRAM is approximately 33.60 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does gemma-2-27b-it-GGUF need at gemma-2-27b-it-f32/gemma-2-27b-it-f32-00001-of-00003?

    The gemma-2-27b-it-f32/gemma-2-27b-it-f32-00001-of-00003 quantization is 36.89 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 13B reference architecture for the KV cache estimate, total estimated VRAM is approximately 44.02 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does gemma-2-27b-it-GGUF need at gemma-2-27b-it-f32/gemma-2-27b-it-f32-00002-of-00003?

    The gemma-2-27b-it-f32/gemma-2-27b-it-f32-00002-of-00003 quantization is 37.13 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 13B reference architecture for the KV cache estimate, total estimated VRAM is approximately 44.28 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does gemma-2-27b-it-GGUF need at gemma-2-27b-it-f32/gemma-2-27b-it-f32-00003-of-00003?

    The gemma-2-27b-it-f32/gemma-2-27b-it-f32-00003-of-00003 quantization is 27.42 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 13B reference architecture for the KV cache estimate, total estimated VRAM is approximately 33.60 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.