Justy: Hey, have you seen AMD's latest MLPerf Training 6.0 results? Cody: Yeah, I glanced at it. They seem to be making some bold claims about their Instinct GPUs. What caught your eye? Justy: Well, apparently they've made a 3.5X generational leap on Llama 2-70B. That's impressive. Cody: Let's not get too excited. We need to dive into the details. What specifically did they improve on? Justy: They mentioned a production-ready MXFP4 (FP4) training recipe across both LLM benchmarks. That's a big deal. Cody: Agreed. But we should also consider how this compares to NVIDIA's performance. Are they really ahead or just catching up? Justy: From what I read, their performance on core LLM workloads is competitive with NVIDIA B200. That's a good sign. Cody: Okay, that's interesting. But what about multi-node training? That's crucial for large-scale deployments. Justy: AMD's first multi-node submission with FLUX.1 is a big step forward. It shows they're serious about scaling. Cody: Alright, I think they're making progress. But let's not forget that MLPerf is just one benchmark. How does this translate to real-world applications? Justy: Fair point. But for AI training infrastructure, this kind of performance and scalability are essential. It changes the game for companies that need to train large models quickly and efficiently. Cody: Agreed. I think AMD's platform readiness, with their ROCm software and growing ecosystem, is what's really noteworthy here.