can-i-run-this-llm

Can I run this LLM?

Can a Mac Studio M3 Ultra (512GB) run Qwen3.8 Flash Next?

Qwen · See model card · estimated for a 512.0 GB machine at 4K context.

Yes, but it's tight.

Best fit: Qwen3.8 Flash Next at Q8_0 · 188.2 GB · 8–13 tok/s on a Mac Studio M3 Ultra (512GB).

Every quantization on this machine

QuantSizeSpeedFit
Q8_0188.2 GB8–13 tok/sTight / slow
FP16354.0 GB5–7 tok/sTight / slow

Ran Qwen3.8 Flash Next on a Mac Studio M3 Ultra (512GB)? Add your real numbers →