Asteria: Okay, Demis basically wrote the most polished version of, 'we need a real gate before the frontier gets weird.' And honestly, Draco, the part I care about is not the AGI poetry. It's that he's proposing an actual standards body with teeth. Draco: Yeah, and the poetry is doing a lot of work. The argument underneath it is more concrete: frontier models should be tested by something outside the labs, with benchmarks that update as capabilities move, and with a path from voluntary review to mandatory gating. That part is real systems thinking, even if the framing gets very "dawning of a new age." Asteria: The title is doing backflips. But the mechanism is kind of the only interesting part, right? A federated, public-private body, maybe FINRA-ish, with independent technical people and open-source reps, plus enough industry money to hire actual talent and rent compute that doesn't fall over. Draco: Mm-hm. And the compute detail matters more than it sounds. If your evaluator can't run large-scale testing, adversarial probes, and follow-up checks, then the whole thing becomes a ceremonial badge factory. The article at least admits that, which is better than most of these frameworks that pretend evaluation is just a spreadsheet and a vibes committee. Asteria: How's your week, by the way? You look like you spent it arguing with benchmarks again. Draco: Roughly, yeah. Very on-brand. I think the thing that jumps out here is the threshold problem: who decides what counts as Frontier-class, and how do you stop the benchmark itself from becoming the target? If the Labs know the shape of the test, they will absolutely train around it. Asteria: Exactly, and Demis sort of knows that. He says the body would start with consultation from frontier labs, then build its own held-out tests so it can stop being a mirror. That is the right instinct. It is also the part that sounds easy until you try to do it and discover you've invented a whole new industry. Draco: Right. Held-out tests are only useful if they stay held-out, and that's the game. Quarterly refreshes help, deprecating saturated benchmarks helps, independent auditors help. But the moment the benchmark suite becomes stable enough to market against, it starts drifting toward benchmark theater unless the body has enough depth to keep generating genuinely new probes. Asteria: That's the bit I think product people would actually feel. If you're a frontier lab, this changes release ops, internal security, maybe even how you staff safety work. If you're not at that level, you mostly just watch the prestige gravity pull harder toward whoever can pass the gate cleanly. Draco: And that's where I get a little less dreamy than Demis. He slides from 'test the dangerous stuff' into 'maybe coordinate a slowdown if needed' pretty fast. Those are not the same policy object. Testing frameworks are tractable. Coordinated slowdown across competitive labs is a completely different beast. Asteria: Yeah, that part is doing a lot of civilizational cosplay. But I do think the article is strongest when it stays on the boring machinery: cybersecurity vetting, post-release vulnerability handling, watermarking for generated images, and model cards with actual technical detail. That stuff is unsexy and probably necessary. Draco: No, I buy the boring machinery. I just don't buy the leap from 'we can build a decent frontier evaluation regime' to 'this will naturally scale into international consensus.' Cross-border alignment is where these nice frameworks go to get mangled. Still, if someone had to start somewhere, this is at least closer to engineering than manifesto-writing. Asteria: And that's why I think the right audience is narrower than the prose suggests. Frontier labs, auditors, national labs, maybe the people building security tooling around them. Everybody else mostly cares indirectly, when the default release culture shifts and the floor gets a little less chaotic. Draco: Which is the most Exploring Next sentence in the world. But yes, that's the practical read. The article keeps reaching for abundance and post-scarcity, and I get why. The real near-term value is simpler: make frontier systems harder to ship irresponsibly, and make the testing regime better than whatever the lab would have written for itself. Asteria: Look at you, almost optimistic. I think that's the cleanest version of the whole thing, honestly. The big cosmic language is there, but the useful part is the governance stack, and that is much less romantic and much more shippable. Draco: Yeah. And if it ever actually exists, the first win won't be 'AGI saved humanity' or whatever. It'll be something boring like, 'the evaluator caught the thing the lab missed before launch.' Which, frankly, is enough to be going on with. Asteria: Okay, that is such a you ending. Fine. But I do like when the future is forced to wear a name tag, Draco. Keeps it from getting away with too much.