can-i-run-this-llm

Can I run this LLM?

Can a NVIDIA GH200 Grace Hopper (144GB HBM + 480GB) run Qwen3.8 Flash Next?

Qwen · See model card · estimated for a 144.0 GB machine at 4K context.

Yes, but it's tight.

Best fit: Qwen3.8 Flash Next at Q8_0 · 188.2 GB · 12–18 tok/s on a NVIDIA GH200 Grace Hopper (144GB HBM + 480GB).

Every quantization on this machine

QuantSizeSpeedFit
Q8_0188.2 GB12–18 tok/sTight / slow
FP16354.0 GB4–5 tok/sTight / slow

Ran Qwen3.8 Flash Next on a NVIDIA GH200 Grace Hopper (144GB HBM + 480GB)? Add your real numbers →