Llama 3 8B in float16 uses 16 GB for weights on an A100-80GB with 90% utilization and 2 GB overhead. KV cache costs 128 KB per token. If you switch to AWQ 4-bit (4.5 GB weights), how many additional concurrent 4K-context sequences can you serve compared to float16?
- About 22 additional sequences, because the 11.5 GB freed from weight compression translates to roughly 22 more 4K sequences at 0.5 GB of KV cache each
- About 10 additional sequences, because the freed VRAM is mostly consumed by the increased dequantization buffer overhead that AWQ requires during inference
- About 90 additional sequences, because AWQ reduces both the model weights and the KV cache memory proportionally, doubling the effective capacity
- About 50 additional sequences, because 4-bit quantization halves all memory requirements uniformly including KV cache, activations, and CUDA context overhead
Why
With float16 weights at 16 GB, available KV cache memory is 80 times 0.9 minus 16 minus 2, which equals 54 GB, supporting 54 divided by 0.5 equals 108 concurrent 4K sequences. With AWQ 4-bit weights at 4.5 GB, available memory becomes 80 times 0.9 minus 4.5 minus 2, which equals 65.5 GB, supporting 65.5 divided by 0.5 equals 131 sequences. The difference is 131 minus 108, approximately 22 additional sequences. The 10-sequence estimate incorrectly assumes large dequantization overhead, but AWQ's kernel fuses dequantization with the matrix multiply, adding negligible memory. The 90-sequence estimate incorrectly claims KV cache is also reduced, but KV cache stores attention states in float16 regardless of weight quantization. The 50-sequence estimate incorrectly assumes quantization halves all memory, but only model weights are quantized. This calculation demonstrates why weight quantization is valuable beyond the throughput gains from smaller matrix multiplies: the freed VRAM goes directly to KV cache, enabling more concurrent sequences. The 21% capacity increase from 108 to 131 sequences translates directly to higher throughput since decode throughput scales linearly with batch size.