Quantization
AWQ 4-bit
Faster inference than GGUF at similar bit-width, at a small extra quality cost.
+0.44 Perplexity delta vs FP16
Community-reported Full benchmark writeup pending — placeholder entry for homepage layout.
Quantization
Faster inference than GGUF at similar bit-width, at a small extra quality cost.
Full benchmark writeup pending — placeholder entry for homepage layout.