top of page

Inside OpenAI's Two-Day Launch: What GPT-5.6 and GPT-Live's Numbers Actually Show

Inside OpenAI's Two-Day Launch
Inside OpenAI's Two-Day Launch: What GPT-5.6 and GPT-Live's Numbers Actually Show

On July 9, OpenAI published two releases within about a day of each other. The first was GPT-5.6, a three-model family (Sol, Terra, Luna) pitched around cost efficiency and agentic reasoning. The second was GPT-Live, a new voice architecture built to listen and speak at the same time rather than waiting for a pause before responding. Neither release exists in isolation. Read together, they show a company spending engineering effort on two different problems at once: getting more useful work out of every token, and making spoken conversation feel less scripted.


The benchmark tables that accompany both releases are worth reading past the summary claims. OpenAI's own comparison charts show GPT-5.6 leading on some evaluations and trailing Claude models on others, a detail that matters more for anyone deciding which model to deploy than the marketing language suggests.


A Model Family Built Around Cost Per Token


GPT-5.6 Sol, Terra, and Luna are positioned as durable capability tiers rather than a single model with settings. On Agents' Last Exam, an evaluation of long-running professional workflows across 55 fields, OpenAI reports Sol scoring 53.6 in its launch narrative, though the summary table elsewhere in the same release lists 52.7 percent against Claude Fable 5's 40.5 percent. Either figure puts Sol ahead on this particular eval, and OpenAI frames the gap as roughly 11 to 13 points depending on reasoning effort, achieved at a fraction of Fable 5's estimated cost. Terra and Luna, the cheaper tiers, are reported to outperform Fable 5 at around one-sixteenth the cost on the same measure.


On the Artificial Analysis Intelligence Index, a broader third-party benchmark spanning agentic work, coding, and general reasoning, the picture is closer: Sol scores 58.9 against Fable 5's 59.9, a gap of about one point, while OpenAI says Sol completes the same tasks in 61 percent less time. The efficiency argument, in other words, rests less on outright superiority and more on getting comparable results faster and cheaper. That framing lines up with what enterprise buyers actually optimize for once a model clears a quality bar: total cost of a completed workflow, not a leaderboard position.


For heavier tasks, GPT-5.6 introduces two ways to spend more compute deliberately. A "max" setting extends reasoning time beyond the prior "xhigh" tier, and an "ultra" mode coordinates four agents in parallel by default, trading token volume for speed and accuracy on demanding work. OpenAI's own charts show this parallel-agent approach shifting the score-versus-latency curve favorably across BrowseComp, SEC-Bench Pro, and Terminal-Bench 2.1, though these are self-reported comparisons and the underlying methodology for scaling to 16 agents isn't detailed in the release.


Coding and Computer Use, With a More Mixed Record


GPT-5.6 Sol posts a new high on the Artificial Analysis Coding Agent Index at 80, about 2.8 points above Fable 5's 77.2, while using under half the output tokens and roughly one-third less estimated cost. It also leads on Terminal-Bench 2.1 and DeepSWE, two evaluations of command-line workflows and long-horizon engineering.


That advantage does not hold everywhere. On SWE-Bench Pro, a benchmark testing real-world software engineering tasks, OpenAI's own table shows Claude Mythos 5 at 80.3 percent and Claude Fable 5 at 80 percent, both well ahead of Sol's 64.6 percent. On GDPval-AA v2, a professional-work evaluation scored in Elo terms, Fable 5 edges out Sol, 1,759.6 to 1,747.8. On Toolathlon, a tool-use benchmark, Mythos 5 and Fable 5 both score 61.7 percent against Sol's 58 percent. The pattern suggests GPT-5.6's strength is concentrated in specific coding and browsing tasks rather than across the board, which is a more useful thing for a buyer to know than "our best coding model yet."


Where GPT-5.6 does show a clearer jump is computer use and design judgment. OpenAI says Sol can inspect a rendered interface after generating it, not just produce the underlying code, and catch visual problems before handing work back. On OSWorld 2.0, a computer-use benchmark, Sol scores 62.6 percent, ahead of Claude Opus 4.8's 54.8 percent, while using what OpenAI describes as 85 percent fewer output tokens. The release also highlights GPT-5.6's ability to infer a slide deck's existing design system, including rules embedded in a Slide Master, and apply it consistently to new content, an improvement over GPT-5.5 in the examples shown.


Cyber and Science Gains, Paired With Access Controls


GPT-5.6 shows large jumps on cybersecurity evaluations. On ExploitBench, which measures progress from a known vulnerability to working exploit code, Sol scores 73.5 percent against GPT-5.5's 47.9 percent, though Claude Mythos 5 still leads at 78 percent. On ExploitGym, Sol nearly doubles GPT-5.5's pass rate under a two-hour cap. OpenAI frames these gains as primarily useful for defensive work such as patching and threat modeling, and says the model does not cross its internally defined "Critical" risk threshold in cybersecurity or biology.


Access to the model's more capable defensive cyber features runs through OpenAI's Trusted Access for Cyber program, which requires identity verification, and the company says individual users will need hardware-backed passkeys enabled by September 1 to retain access to its most cyber-capable models. On biology-related evaluations, including GeneBench Pro and LifeSciBench, OpenAI reports Pareto improvements over GPT-5.5. It's worth noting OpenAI's release states that Claude Fable 5 declines to answer most questions on one of these biology evaluations, a claim that comes from OpenAI's own footnote rather than independent testing.


Accelerating Its Own Research Loop


One of the more specific numbers in the release concerns OpenAI's internal use of the model. The company says average daily output tokens per active researcher during GPT-5.6's internal testing period were more than double the peak seen with GPT-5.5, and that the share of internal research compute devoted to coding inference grew roughly 100-fold over six months, with agentic token usage up about 22-fold over the same period. OpenAI built an internal evaluation suite, the RSI Index, to track this directly, and reports Sol scoring 16.2 points higher than GPT-5.5 on a bundle of tasks measuring progress toward recursive self-improvement.


These figures describe OpenAI's own workflows before they describe anything a customer will touch. That sequencing, capability compounding first inside the tightly controlled environment where a lab already operates before it shows up in a shipped product, is a pattern worth watching across this industry, not just this release.


GPT-Live: A Different Bet on Voice


Alongside GPT-5.6, OpenAI released GPT-Live, a voice architecture built for full-duplex interaction, meaning the model processes incoming audio and generates output simultaneously rather than waiting for silence to signal a turn has ended. When a query needs web search or deeper reasoning, GPT-Live delegates to a frontier text model, GPT-5.5 at launch, in the background while keeping the conversation flowing.


In head-to-head human evaluations against Advanced Voice Mode, OpenAI reports GPT-Live-1 preferred 75.7 percent of the time in matched five-to-ten-minute conversations, with GPT-Live-1 mini preferred 69.2 percent of the time. On GPQA, a test of expert-level scientific reasoning, GPT-Live-1 at high reasoning effort scores 84.2 percent against Advanced Voice Mode's 45.3 percent. On BrowseComp, the gap is larger still: 75.2 percent versus 0.7 percent. A third evaluation, testing voice agents on telecom support tasks, used what OpenAI describes as a customized internal user model, worth flagging since it's not an external, reproducible benchmark.


GPT-Live-1 becomes the default voice model for ChatGPT's paid tiers, with GPT-Live-1 mini defaulting for free users, rolling out globally across iOS, Android, and the web. OpenAI says an API version is planned, with a signup form available now for developers who want early notice.


Safety Measures Built for Real Time


Both releases lean on safety infrastructure OpenAI says goes beyond prior generations. For GPT-5.6, the company describes a layered system combining training-time protections with a real-time reasoning monitor that reviews conversations for potential harm, alongside account-level enforcement calibrated to trust and risk. OpenAI says its cyber safeguards for Sol now block roughly ten times more potentially harmful activity than before, a change the company acknowledges creates friction for legitimate defensive work, which is why it offers an option to retry blocked prompts on lower-capability models.


For GPT-Live, the safety work is voice-specific: new audio-native evaluations, red-teaming focused on self-harm, psychosis, emotional reliance, and sexual content, and built-in mechanisms that can steer a live conversation toward a safer response or end it in higher-risk cases. OpenAI says linked parents may be notified in situations involving signs of potential self-harm among teen users, and that the model is restricted to a fixed set of predefined voices to prevent impersonation of real people.


What to Take From the Combination


Taken together, the two releases point to OpenAI running parallel bets rather than a single roadmap. GPT-5.6 competes primarily on cost per completed task, with a benchmark record that favors OpenAI in some domains, coding agent workflows and computer use chief among them, while Claude models retain a lead on others, including SWE-Bench Pro and general professional-work evaluations. GPT-Live competes on a different axis entirely: the subjective experience of talking to a model, an area where the human-preference data shown here is genuinely lopsided in OpenAI's favor.


For teams evaluating either release, the practical next step is running task-specific comparisons rather than relying on the aggregate claims either company publishes. Cost per completed workflow, not a single index score, is what should drive a deployment decision, and this release shows that picture varies considerably depending on which specific task sits in front of the model.

 
 

JOIN THE AI SPECTATOR MAILING LIST

CONTACT

Contacting You About:

Thanks for submitting!

New York, NY           

Db @DavidBorish.com           

  • LinkedIn
  • Instagram
  • Facebook
  • X
Back to top

© 2026 by David Borish IP, LLC, All Rights Reserved

bottom of page