Talon: Okay, Wildflower, I think this is either a real cost-performance jump or another case of benchmark bundling dressed up as a model launch… and you already look annoyed. Wildflower: I am a little annoyed. The article’s actual claim is not just “the model is better.” It’s that GPT-5.6 gets more useful work per dollar, across coding, knowledge work, science, browsing, computer use, all of it. And then halfway through, that claim quietly starts absorbing ultra, which is four agents in parallel by default. Those are NOT the same thing. Talon: Right. Wildflower: If your headline win comes from better base weights, great. If it comes from packaging a multi-agent harness into the product, also fine. But say which one is doing the work. You and I have been on this for months now. The harness is the whole loop, not some footnote. Talon: Yeah, but I don’t think that makes the launch fake. If I’m a team trying to ship coding agents, I care whether the thing solves more tasks for less money and less waiting. I do not need the victory to be spiritually pure. That is such an Exploring Next disease you have. Talon: Seriously. The strongest part of the post is pretty plain. Sol hits eighty on the Artificial Analysis Coding Agent Index, two point eight above Fable 5, with less than half the output tokens, less than half the time, and about one-third lower estimated cost. That’s a real product claim, not just “vibes were excellent in our internal demo.” Wildflower: I buy that more than the broad “most polished collaborator yet” stuff. The coding section at least gives you multiple angles. Coding Agent Index, Terminal-Bench two point one, DeepSWE, then the token and latency deltas. That hangs together better technically than the design-judgment section, which is basically, look, it made a nicer slide deck. Talon: Mm-hm. Wildflower: And even there, the slide example is doing a very specific thing. Template following, style transfer from a reference deck, preserving master-slide conventions. Useful, yes. But that’s narrower than “expert-level shareable artifacts from your messy work life.” That sentence is doing a LOT. Talon: I mean… yes. OpenAI discovered marketing copy again. But I still think the practical audience is obvious. If you’re running code agents, browser-heavy research, or document production inside a company, this matters because the article is basically saying the same work can clear with fewer model round trips and fewer tokens. That’s budget and latency, not just ego. Wildflower: Sure. Talon: And Terra might be the sneaky important one. Sol gets the flagship treatment, Luna gets the cheap-and-fast slot, but Terra sounds like the default-buy model. Balanced, close enough on benchmarks, lower cost. That’s the one product teams actually standardize on if they don’t want to think about model routing all day. Wildflower: Talon, that part I actually agree with. The family story is stronger than the “new standard for intelligence” line. Sol is the halo. Terra is probably the workhorse. Luna is there to make abundance credible instead of rhetorical. The article even hints at that when it says the smaller models beat Fable 5 at around one-sixteenth the cost, though I’d want the exact setup before I get too excited. Talon: Oh interesting. Wildflower: Because again, benchmark context matters. We already did this with Grok four point five. Competitive, maybe even strong, is not the same as universally best. And some of these evals are ecosystem-shaped. If Sol is being run in OpenAI’s own codex-style environment, I’m not shocked it looks especially good there. Talon: I don’t even think that’s a gotcha, though. If the environment is part of the shipped experience, then performance in that environment matters. The thing I’d push on is the article sometimes blurs “model can do this” with “Responses API plus Programmatic Tool Calling plus multi-agent beta can do this.” Those are different purchasing decisions. Wildflower: Exactly. Talon: But they’re both valid decisions. This is where your skepticism tips into infrastructure purism. Users buy outcomes. If Programmatic Tool Calling lets the system filter intermediate junk without shoving every tool response back through the model, that’s meaningful. It sounds like the same boring answer that keeps winning: better loop design, less waste. Wildflower: No, I’m with you on the mechanism. That part is actually one of the more believable sections. Writing small programs to coordinate tools, monitor progress, and choose the next action is much cleaner than making the model narrate every microscopic step. We’ve basically been saying since the loop episode that repetition without new information is just churn. This is a way to cut churn. Talon: Right, right. Wildflower: My issue is only attribution. Ultra especially. Four agents in parallel by default, plus optional sixteen-agent comparisons on BrowseComp and SEC-Bench Pro, and then the chart moves up and left on score versus latency. Cool. But that is a systems result. It should be sold as a systems result. Talon: Okay but that’s still a good result. Honestly, I’d put seventy-thirty on ultra becoming a standard paid mode across the frontier labs by fall, just because “faster hard-task completion through parallel workers” is an easy thing to sell. Wildflower: I’d go higher than that, maybe eighty-twenty. Not because it’s profound. Because it’s legible. You can explain it in one sentence without pretending the model woke up enlightened. Wildflower: Also, tiny side note, I cannot believe we’ve been doing this since November and I’m now arguing that the honest part of a frontier launch is the part where they admit it’s four workers and a bill. Talon: That’s growth. That’s emotional growth. Talon: My honest read is pretty simple. The coding and cost claims look strong enough that teams should test this now, especially Sol versus Terra in their own harness. The knowledge-work and design stuff, I’d treat as promising until somebody stress-tests it outside the pretty examples. And ultra is interesting precisely because it’s not magic. Wildflower: Yeah. Credible release, overbroad framing. The argument holds up best where they show repeated gains on agentic coding and workflow evals with token, time, and cost attached. It gets mushier when “good slide hygiene” turns into “professional collaborator.” That gap is where I’d keep my eyebrow raised. Talon: Fair. Episode six twenty-nine, and you only hated about thirty percent of it. I’m calling that a lovely place to stop.