NVIDIA Optimizes MoE Training for Biological Foundation Models
NVIDIA uses Transformer Engine and MXFP8 to increase MoE training throughput by up to 2.21x on B200 GPUs for biologic ...
News on open-weight models you can run locally
NVIDIA uses Transformer Engine and MXFP8 to increase MoE training throughput by up to 2.21x on B200 GPUs for biologic ...