GPU
RTX 4090 24GB
24GB VRAM handles 70B models at Q4 quantization with room for an 8K context window.
18.2 Tokens/sec (Llama 3 70B, Q4)
Verified Full benchmark writeup pending — placeholder entry for homepage layout.
GPU
24GB VRAM handles 70B models at Q4 quantization with room for an 8K context window.
Full benchmark writeup pending — placeholder entry for homepage layout.