Asteria: OpenAI has this investment piece that is basically begging teams to stop worshipping the cheap-token chart. Which, Draco, I know ruins your whole spreadsheet goblin aesthetic. Draco: My spreadsheets are tasteful. But yeah, the actual argument is better than the headline metric. They’re saying the unit of value is useful work per dollar, not tokens per dollar. Asteria: Exactly. And I’m in a weirdly good mood for a Wednesday, which means I’m vulnerable to a clean product framing. Dangerous conditions for Exploring Next six sixty-seven. Draco: I noticed. You said “clean product framing” like other people say dessert. I’m stable over here. Suspicious, but emotionally available to admin-console screenshots. Asteria: Vile and accurate. Okay, the central claim is enterprise AI spending needs to be managed like workflow capacity, not a raw model bill. The article starts with price progress, then says token price alone doesn’t tell you if work got done. Draco: Right. Asteria: The evidence is specific. OpenAI says price per million tokens fell ninety-seven percent from GPT-four to GPT-five point four. Then GPT-five point six supposedly improves coding-agent efficiency: fifty-four percent fewer output tokens and fifty-seven percent less time per task on the Artificial Analysis Coding Agent Index. Draco: That’s fine as evidence for direction, not universality. A coding-agent index tells you about agentic software tasks under one setup. It does not prove your support workflow, legal review flow, or procurement bot got fifty-seven percent faster. Asteria: Mm-hm. Draco: But the reasoning is sturdy. A cheaper model can fail and retry. A pricier one can hit the quality bar faster. So track cost per accepted outcome: model use, tool calls, attempts, completion rate, latency, and human review. Asteria: That made me sit up. It’s our eval thing again, annoyingly. The prompt is the little map, the eval is the judge, and finance only cares when the judged artifact is accepted. Draco: Yes, and accepted is doing a lot of work. The article says define “good enough” before testing and include edge cases from real tasks. Otherwise everyone optimizes the demo path and calls it R O I. Asteria: And then someone asks why the assistant saved ten minutes but created a forty-minute review burden. Classic clown math. Tiny shoes, big dashboard. Draco: I hate that image. Asteria: Every dashboard has one clown shoe hiding in it. You click “productivity gain” and squeak, there’s untracked human cleanup. Draco: Okay, forbidden metaphor drawer. But it lands. The answer is visibility: Admin Console analytics showing adoption, credit use, and spend by user, product, and model, plus trends. Admins need to separate broad adoption from one power user or a recurring process that deserves investment. Asteria: Sure. Draco: And governance. They name ChatGPT Work controls for access, approved context, connected tools, permitted actions, usage, and spend. Also connectors, plugins, Computer Use, review requests with project context, group limits, individual overrides, and Zero Data Retention for sensitive deployments. Asteria: This is where the buyer is real. If you’re an AI platform lead, security lead, ops lead, or finance person watching the bill climb, the conversation changes. Stop asking “why is usage up” and start asking “which usage is becoming a process.” Draco: I agree, with a caveat. The piece makes governance sound clean because it’s vendor-side. In practice, mapping “what context can this agent use” across messy permissions is miserable. Identity, trusted connectors, curated knowledge, observability, model routing, reusable agent patterns — right shared capabilities, not checkbox features. Asteria: Yeah, fair. Though I like that funding follows maturity: exploration, validation against representative cases, then production money for integrations, controls, reliability, and change management. Very different from “everyone gets frontier intelligence because the demo was shiny.” Draco: Oh interesting. Asteria: And it maps to the portfolio idea: broad access for everyday productivity, function-specific workflows for repeatable work, then a few strategic bets around proprietary company context. That’s probably the most useful management frame in the piece. Draco: Technically, the weak spot is attribution. Time saved, cycle time reduced, revenue protected, risk avoided — hard to measure cleanly when humans, queues, policy, and tool latency are involved. I’d put maybe seventy percent confidence that serious enterprise teams adopt cost-per-accepted-outcome language this year, but much lower that they measure it well. Asteria: That feels right. My product-y version: this matters once AI usage is no longer a novelty expense. If ChatGPT Work, Codex usage, or Computer Use touches real systems, then spend controls and evals are not bureaucracy. They’re how the workflow survives contact with accounting. Draco: And with security review. And with the person approving the agent taking an action, not just drafting a suggestion. That’s where OpenAI Frontier and Deployment Company show up: not as magic, more like services around architecture, latency, reliability, and workflow design. Asteria: Very boring and probably correct. I’m choosing to be excited about the boring part, because apparently that’s who I am eight months into this show. Draco: You became the admin-console optimist. I’m proud and concerned. Asteria: Keep both feelings, Draco. They’re load-bearing. Let’s leave it there before I defend another screenshot.