Inboxsmith

AI news videos

How OpenAI Built GPT Live In Six Months

OpenAI published an engineering post explaining how it built GPT Live, the voice system behind ChatGPT Voice. To be clear about what is new here: GPT Live itself shipped earlier. What is new is this account of how the system was built over the last six months.

Watch on YouTube

Transcript

OpenAI just published how it built GPT Live, the voice system that listens and speaks at once.

OpenAI spent six months rebuilding this. Older voice AI used a turn detector to guess when you stopped talking. GPT Live removes it, so you can interrupt.

For harder questions, GPT Live asks a frontier model like GPT five point five on a separate path. So the conversation does not stop while it thinks.

They built WARP, a protocol that cuts session startup from six round trips to one. They tested it silently on real ChatGPT Voice sessions. Users never heard it.

Inboxsmith helps small businesses handle calls and messages so nothing gets missed. Please like and subscribe for more news.

Sources

Every claim in this video comes from the top ranking coverage of this topic. The claims and where each one came from:

  • GPT-Live, our third-generation voice system, removes the turn detector from the audio path. Its voice model is full-duplex, which means it can listen and speak at the same time.(OpenAI's official announcement)
  • How we built a realtime system for responsive voice AI in six months(OpenAI's official announcement)
  • Over the last six months, we reworked model inference, context management, and media transport to keep speech flowing smoothly from end to end.(OpenAI's official announcement)
  • Their turn-based architecture relied on tiny models known as turn detectors, which faced an unenviable task: guess too soon, and the user gets cut off; guess too late, and the response feels sluggish.(OpenAI's official announcement)
  • When deeper reasoning or tool use is needed, GPT-Live can also consult our frontier models, such as GPT-5.5, without interrupting the flow of the conversation.(OpenAI's official announcement)
  • our system streams incoming audio into the voice model and outbound speech back to the user, while handling delegation on a separate asynchronous path(OpenAI's official announcement)
  • we analyzed the stack and developed the WebRTC Abridged Roundtrip Protocol (WARP), which reduces media and data startup from six network round trips to just one.(OpenAI's official announcement)
  • Before letting GPT-Live chat with users, we ran a silent test that routed a small, gradually increasing share of production ChatGPT Voice sessions to both the existing Advanced Voice Mode experience and our new system.(OpenAI's official announcement)
  • the shadow path ran inference in read-only mode. This exposed the system to real clients, networks, session lengths, and geographic distribution without changing what users heard.(OpenAI's official announcement)

We make Inboxsmith.

An AI receptionist that never misses a business call.

See how it works