can-i-run-this-llm

Can I run this LLM?

Can a NVIDIA H100 SXM (80GB) run Laguna S 2.1 (MoE)?

Laguna · openmdw-1.1 · estimated for a 80.0 GB machine at 4K context.

Yes — it runs great.

Best fit: Laguna S 2.1 (MoE) at Q3_K_M · 54.0 GB · 221–331 tok/s on a NVIDIA H100 SXM (80GB).

Every quantization on this machine

QuantSizeSpeedFit
Q3_K_M54.0 GB221–331 tok/sRuns great
Q4_K_M73.1 GB198–296 tok/sWon't fit
Q5_K_M87.9 GB183–274 tok/sWon't fit
Q6_K97.9 GB174–261 tok/sWon't fit
Q8_0125.0 GB154–231 tok/sWon't fit
FP16235.2 GB105–158 tok/sWon't fit

Ran Laguna S 2.1 (MoE) on a NVIDIA H100 SXM (80GB)? Add your real numbers →