Skip to main content
AMD MI355X CDNA4 GPU benchmarking AI training performance in MLPerf v6.0, showcasing competitive results with high-speed data

Editorial illustration for AMD's MI355X CDNA4 GPU Shows Competitive Training Times in MLPerf v6.0

AMD's MI355X CDNA4 GPU Shows Competitive Training Times...

Updated: 3 min read

AMD just matched Nvidia’s top chip. In a head-to-head sprint, the company's new MI355X accelerator fine-tuned a Llama 2 70B model in the same time as a Nvidia B200 GPU, using an eight-accelerator setup. It repeated that on a Llama 3.1 8B pre-training run.

The June 2024 MLPerf tests also saw AMD scale its MI325X chip across eight nodes to generate images from text. This round delivered a crucial debut: the first appearance of AMD’s Primus training framework and its MXFP4 data format running natively on the company's own hardware.

AMD’s MLPerf Training v6.0 submission demonstrates continued progress across both hardware and software. On the hardware side, the CDNA4-generation MI355X delivers competitive time-to-train results against NVIDIA B200 on both single-node LLM benchmarks — Llama 2 70B LoRA fine-tuning and Llama 3.1 8B pretraining — at an iso-GPU count of 8, while the MI325X powers an 8-node Flux.1 Schnell text-to-image submission.

The significance is in the execution. AMD ran its MXFP4 recipe directly on the MI355X's FP4 silicon, bypassing the drag of software emulation. Its Primus framework managed both benchmark types.

That signals a more cohesive software stack. Nvidia, of course, still won more overall benchmarks. But on this specific eight-GPU battleground, AMD’s result is clear.

The training performance gap for core AI workloads has tightened. Now, watch where the fight moves next.

Common Questions Answered

How did AMD's MI355X GPU perform against Nvidia's B200 in the MLPerf v6.0 benchmarks?

AMD's MI355X matched Nvidia's B200 GPU in head-to-head performance, successfully fine-tuning a Llama 2 70B model in the same time using an eight-accelerator setup. The MI355X repeated this competitive result on a Llama 3.1 8B pre-training run, demonstrating that the training performance gap for core AI workloads has significantly tightened between the two companies.

What is MXFP4 and how does AMD's implementation provide an advantage?

MXFP4 is AMD's new data format that debuted in the MLPerf v6.0 tests. AMD ran MXFP4 directly on the MI355X's FP4 silicon, bypassing the performance drag of software emulation, which gives it a significant efficiency advantage over traditional approaches.

What role does AMD's Primus training framework play in these benchmark results?

AMD's Primus framework is a new training framework that managed both benchmark types in the MLPerf v6.0 tests, signaling a more cohesive software stack. This unified framework demonstrates AMD's ability to handle diverse AI workloads efficiently across its accelerators.

What was the significance of AMD's MI325X performance in the MLPerf v6.0 image generation tests?

AMD scaled its MI325X chip across eight nodes to generate images from text in the June 2024 MLPerf tests, showcasing its capability for multi-node deployment. This demonstrated AMD's competitive positioning not just in language model training but also in generative image tasks.

LIVE00:14DeepSeek Boosts Agent, Coding Performance in Open-Source V4-Flash Model