Laura: Okay, Harper, this is basically your favorite kind of fight. Big anti-lock-in claim, little benchmark detail, and me going, yeah but what if the product story is still real. Harper: It is extremely my thing. My first read is that the article is selling a category dream more than a demonstrated result. ZML says its LLMD server can run open models across Nvidia, A M D, Google T P U, Apple Metal, and Intel Arc, and sometimes hit max speed or better. That is a huge claim. Laura: Right. Harper: And the evidence in the piece is mostly Morin saying the silos are bad, inference matters more now, and customers want optionality. I buy all of that. What I do NOT think the article establishes is that one server is already extracting near-peak performance across that many backends. Laura: I think that's fair, but you're flattening the useful part. The central argument isn't just speed bragging. It's that inference software has become the control point, and if you can make chip choice less painful, suddenly mixed fleets become a practical buying decision instead of a science project. Harper: Mm-hm. Laura: That matters even before you prove universal best-in-class kernels on every device. If I'm a cloud or enterprise team staring at cost and power, a serving layer that lets me test cheaper or lower-power hardware without rewriting everything is already interesting. Harper: Laura, yes, but then say THAT. Say portability, say operational leverage, say reduced switching cost. Don't say maximum available speed across lots of chips unless you've brought receipts. Support is one thing. Fast is another. Peak fast everywhere is where I start squinting. Laura: You really do hear one inflated clause and turn into the benchmark police. Harper: Because somebody has to. Also the fresh context makes this worse, not better. Their own launch is labeled alpha, and the Docker image calls it a technical preview not intended for production use. So now we're two levels away from "this changes enterprise inference today." Laura: Sure. Harper: And that doesn't kill it. It just means the article is early. Which, fine, TechCrunch is allowed to be early. But technically, the hard problem here is not making one binary start up on five hardware families. It's scheduling, kernels, memory behavior, quantization paths, model-specific weirdness, and not falling apart when the workload changes. Laura: I don't disagree. I just think you're underrating how much value there is in collapsing the evaluation loop. If LLMD gives teams one self-contained server where they can run Llama, Gemma, Qwen, and Mistral across different chips, that saves real time. You don't need the final forever stack for that to be a product win. Harper: Yeah. Laura: And the article's reasoning on market dynamics is pretty solid. Inference is where spending is exploding, Nvidia's still dominant, but buyers are desperate for leverage. So software that weakens the default lock-in gets attention even if it starts as a very good adapter layer. Harper: That part I actually buy. This rhymes with our boring-answer-keeps-winning thing from a while back. The valuable move is often not some magical new model idea. It's a nasty systems layer that makes the economics less stupid. Laura: Exactly. Harper: But I still want numbers. Baselines against v L L M or S G Lang. Same model, same batch shape, same latency target, same chip, same quantization. Otherwise "faster" is just vibes with venture funding behind it. Laura: Okay, tiny side note. "Vibes with venture funding behind it" is such an Exploring Next sentence that I need it framed somewhere. Harper: Eight months of this show and somehow that's what survives. Laura: I mean, that's our real cap table. Bad phrases and grudges. Laura: Back to it. The co-designing silicon line is where I got cautious. That's the kind of sentence that can mean deep compiler and kernel collaboration, which is real, or it can mean founder-theater around roadmap conversations. Harper: Yes. And because the article names all these European chip startups, it kind of gestures at an ecosystem thesis. Which may be true. A neutral serving stack could help weird new accelerators get tried. But again, helping a chip get tried is not the same as making it competitive. Laura: I think the people who should care are exactly those in-between buyers. Not the team that's all-in on Nvidia and happy. The ones with procurement pressure, power constraints, or regional hardware options. For them, even partial de-risking changes the conversation. Harper: I'd put sixty-five thirty-five on LLMD finding real developer traction as an evaluation layer this year. Production default across heterogeneous fleets? Much lower. Maybe thirty percent until they publish serious benchmarks and get past technical preview status. Laura: That's basically where I land, just less grumpy about it. I could go seventy-thirty that it becomes useful before it becomes dominant. Which, honestly, is enough. A lot of infrastructure earns its keep by making the market more contestable, not by winning the whole stack. Harper: And I do appreciate that Morin isn't pretending Nvidia is over. The article is clear on that. This is not "Nvidia is done." It's "there's finally room for software to make other chips less annoying." Smaller claim. Better claim. Laura: Yeah. Also, if they want the Build Next version of this, the concrete thing is just LLMD alpha and the ZML repo. That is at least testable, which I respect. Harper: Same. Run the server, pick one of the supported open models, compare it against your existing v L L M or S G Lang setup, and see if the promise survives contact with your actual workload. That is a much healthier use of a Wednesday than arguing about the phrase hot French startup. Laura: Debatable. Okay, episode six fourteen. You go collect error bars, I'll go look for the user who just wants the chip drama to stop.