Oh yes for the FP8, you will need 500GB ish. 4bit around 250GB - offloading MoE ...

vFunct · 2025-07-22T23:29:17 1753226957

Do we know if the full model is FP8 or FP16/BF16? The hugging face page says BF16: https://huggingface.co/Qwen/Qwen3-Coder-480B-A35B-Instruct

So likely it needs 2x the memory.

danielhanchen · 2025-07-22T23:47:20 1753228040

I think it's BF16 trained then quantized to FP8, but unsure fully - I was also trying to find out if they used FP8 for training natively!

jychang · 2025-07-22T23:55:16 1753228516

Qwen uses 16bit, Kimi and Deepseek uses FP8.

danielhanchen · 2025-07-23T03:07:51 1753240071

Oh ok cool thanks!