On July 8, 2026, OpenAI released "GPT-Live", a new model that overhauls the voice experience in ChatGPT worldwide (OpenAI's official announcement). Announced just ahead of the general availability of GPT-5.6 the following day, its headline feature is a full-duplex architecture that lets it "listen and speak at the same time."

Until now, AI voice has been turn-based: "it waits for you to finish talking, then responds." GPT-Live, by contrast, behaves like a human conversation partner — it backchannels, absorbs interruptions, and doesn't cut off the silences while you think. In this article we explain, based on the official announcement, what GPT-Live is, how it differs from the old Advanced Voice Mode, how it works, which plans can use it, and what it still can't do.

Update as of September 25, 2026: since publication, the background model and voice usage limits changed (September 9), the API became generally available (September 10), and voice came to ChatGPT Work (September 23). We kept the launch-day explanation and rewrote the parts that differ today, with dates (sources: OpenAI release notes, ChatGPT Voice help article, API changelog).

FULL-DUPLEX VOICE · 2026

The end of turn-taking

— toward human-like conversation that listens while it speaks

Full-duplex
Listen + speak at once
Backchannels & interruptions OK
Preferred
Clearly preferred over Advanced Voice Mode
Human evals of 5–10 min chats
2 models
Live-1 / Live-1 mini
mini is available even for free
Delegation
Search and deep reasoning go to another model
The background model is updated as new ones ship

1. What is GPT-Live — the "listen while speaking" full-duplex voice

GPT-Live is a new series of voice models that powers voice conversations in ChatGPT. Replacing the previous voice feature (Advanced Voice Mode), it rolled out worldwide to ChatGPT on iOS, Android, and the web starting July 8, 2026.

Its core is a "full-duplex" design. Like a phone call, it keeps generating output (the AI's voice) while it processes input (your voice) at the same time. Because of that —

  • It can offer backchannels like "mhmm" and "yeah" while you're still talking
  • If you interrupt mid-sentence, it actually listens and changes course while it's speaking
  • If you fall silent to think, it doesn't mistake that for "you're done" and cut you off

In short, GPT-Live delivers real-time conversation with no "waiting your turn."

2. How it differs from the old Advanced Voice Mode

The previous AI voice (Advanced Voice Mode) was turn-based (half-duplex). By design, it had these weaknesses.

AspectOld: Advanced Voice ModeNew: GPT-Live
Basic methodTurn-based (half-duplex) — waits for you to finishFull-duplex — speaks while listening at the same time
Detecting the end of a turnBased on silenceDoesn't wait for a clean break — follows pauses, interruptions, and changes in pace as it listens, and decides in the moment
Thinking pausesTends to mistake them for "done" and interruptWaits without cutting in
BackchannelsEssentially noneReturns "mhmm," "yeah," and the like
InterruptionsCan cut off unnaturallyAbsorbs them and adjusts
Decision frequencyPer turnMultiple times per second (speak / listen / pause / interrupt / call a tool)

According to OpenAI, in a comparison where humans evaluated 5-to-10-minute conversations under the same conditions, both GPT-Live-1 and GPT-Live-1 mini were clearly preferred over Advanced Voice Mode. Improvements were also reported on GPQA (expert-level scientific reasoning), BrowseComp (agentic web search), and customer-support-style tasks.

3. How it works — per-second decisions and delegation to a model behind the scenes

GPT-Live generates voice output while continuously processing voice input, and decides "what to do right now" many times per second — speak, keep listening, pause, interrupt, or call a tool. This high-frequency decision-making is what produces the human-like tempo.

Another key part of the design is that "deep processing is delegated to a model behind the scenes." GPT-Live handles the responsiveness of the conversation, and when web search, deep reasoning, or complex work is needed, it hands the task off to a larger model behind the scenes (GPT-5.5 at launch), then returns to the conversation once results come back. It's a mechanism that combines "thinking" with an uninterrupted conversation.

OpenAI wrote in the announcement that "as we release new frontier models, we'll continuously update the model used by GPT-Live," and sure enough, since September 9, 2026, Voice uses GPT-5.6 or GPT-6 Astra (OpenAI release notes). In other words, the name of the background model will keep changing. What's worth remembering is not the name but the structure: one part keeps the conversation going and another does the thinking. How smart the answers are depends on the background model; how easy it is to talk to depends on GPT-Live.

At launch, you chose the depth of reasoning from three levels: Instant (GPT-5.5 Instant) / Medium / High (GPT-5.5 Thinking). These voice-only levels were deprecated on September 9, 2026; you now pick the model and reasoning effort with the same controls as text chat (what's available depends on your plan; OpenAI release notes). The basic idea hasn't changed: favor speed for light chat and deeper thinking for involved consultations.

4. Two models and availability by plan

PlanAt launch (July 8, 2026)As of September 25, 2026 (Voice in Chat, per 24 hours)
FreeGPT-Live-1 miniGPT-Live-1 mini (limited; limits may change)
GoGPT-Live-13 hours with GPT-Live-1 mini
PlusGPT-Live-13 hours with GPT-Live-1
ProGPT-Live-115 hours on the $100 plan, unlimited on the $200 plan (GPT-Live-1)
Business / Enterprise / EduNot availableAvailable depending on workspace settings. Business Standard gets 3 hours and Premium 15 hours; extra usage costs 1.25 credits per minute

The key point is that even the free plan can use the mini version of GPT-Live. Natural full-duplex conversation isn't limited to paid tiers. At launch GPT-Live-1 was the default on Go/Plus/Pro, but since September 9, 2026, Go uses GPT-Live-1 mini, and Plus and Pro no longer switch to mini after hitting a Voice limit (OpenAI release notes). Business, Enterprise, and Edu workspaces, which were excluded at launch, can now use it depending on their settings (OpenAI Help). It's available on ChatGPT for iOS, Android, and ChatGPT.com, rolled out in stages worldwide. The limits get revised, so check the table in OpenAI Help before relying on them.

5. The main advances

🗣️ Backchannels

Responds with "mhmm" while you're talking. It creates the sense of being heard.

⏱️ Quick back-and-forth

It decides many times per second whether to speak or keep listening, so replies keep pace. Neither OpenAI's announcement nor the model page gives a latency figure.

✋ Follows interruptions

Even if you cut in mid-conversation, it listens and course-corrects.

🤫 Doesn't cut off silence

It doesn't misjudge a thinking pause as "you're done" — it waits for you.

6. Current limitations — worth being honest about

Expectations can run ahead of reality, so it's worth being honest about what it can't do. Below, the launch-day limitations are split into those that still apply and those that have been resolved (as of September 25, 2026).

  • 🟡 No video or screen sharing: unchanged since launch — Live doesn't handle video input or screen sharing (OpenAI Help). When you need them, switch to the older Advanced voice in the iOS or Android app.
  • 🟡 Full multilingual support is still to come: at launch, OpenAI wrote that GPT-Live is optimized for the most popular languages and that for certain languages it may have a non-native accent or gaps in fluency.
  • 🟡 Built for one-on-one conversation: it isn't yet optimized for several people speaking at once, and may respond when people are talking to each other (OpenAI Help).
  • 🟢 The API is now available: it was "coming soon" at launch, and became generally available on September 10, 2026 (see chapter 8).
  • 🟡 Staggered rollout: which features you get, and when, depends on your plan, region, workspace settings, and app version.

These are expected to expand in future updates, but it's good to understand that "video and screen sharing right now" is not on the table.

7. Where it shines

  • Hands-free consulting and thinking out loud: organize your thoughts by voice while walking or working. Because your silences aren't cut off, your train of thought is less likely to break.
  • Language conversation practice: with backchannels and natural interruptions, you get practice that's closer to a real conversation.
  • Research and summarizing by voice: deep search is delegated to a model behind the scenes, so you can look things up while you talk.
  • Handing off work by talking: since September 23, 2026, you can speak to ChatGPT Work on web and phone and ask it to create documents, slides, and spreadsheets. If a task is still running when you end the call, it continues in text (requires both Voice and Work access; OpenAI release notes).
  • Accessibility: in situations where typing is hard, a snappy voice interface helps.

OpenAI Help also gives concrete tips for getting more out of it.

  • If you want to think out loud, say so first: at the start, tell it something like "wait until I ask you to respond," and it will hold back. Long pauses or background sounds can still make it respond.
  • If it keeps cutting in, change the setup: use headphones, move somewhere quieter, or turn up your device volume. On iPhone, open Control Center, choose Mic Mode, and turn on Voice Isolation.
  • For dictated text, use Dictation: Voice transcripts aren't verbatim and may not match exactly what was said. When you want to review and edit what you said before sending it, Dictation fits better than Voice.
  • Know how recordings are handled: audio clips are stored with the chat transcript and retained for 30 days. Clips are used for training only if you turn on sharing yourself (sharing isn't possible in Business, Enterprise, or Edu). Transcripts may be used depending on your "Improve the model for everyone" setting.

8. API availability and how it differs from GPT-Realtime

At launch, the GPT-Live API was "coming soon," and developers and businesses could only sign up for notifications. Then, on September 10, 2026, it became generally available in the API (OpenAI API changelog). According to the model page as of September 25, 2026, the model ID is gpt-live-1, it runs on a dedicated Live endpoint (v1/live/sessions), and voice sessions cost $0.05 per minute, billed per second. The backend model and tools are billed separately.

In the API, just as in ChatGPT, you build it as GPT-Live running the conversation plus a backend doing the thinking. You can let OpenAI run the backend model, or connect your own agent or service (official guide). Your application still handles permission checks and anything that touches your own systems.

To clear up the confusion: the API also has voice models separate from GPT-Live — GPT-Realtime-2.1 / GPT-Realtime-2.1 mini (released July 6, 2026), reasoning voice models with tool use for the Realtime API. They're built differently: GPT-Live = a full-duplex conversation layer that listens while speaking and hands the thinking to a backend, while GPT-Realtime = a voice-agent foundation on the Realtime API. OpenAI's voice agents guide compares when to use which.

Summary

  • GPT-Live is a new model that changes ChatGPT's voice from "turn-based" to full-duplex (listening while speaking) (July 8, 2026, worldwide).
  • With backchannels, interruption handling, and tolerance for silence, it achieves a conversational tempo close to a human's. It was clearly preferred over Advanced Voice Mode.
  • GPT-Live handles conversational responsiveness, while deep reasoning and search are delegated to a model behind the scenes (GPT-5.5 at launch; GPT-5.6 or GPT-6 Astra since September 9, 2026).
  • Free and Go use GPT-Live-1 mini, Plus and Pro use GPT-Live-1 (as of September 25, 2026). Business, Enterprise, and Edu can use it depending on settings. Available on ChatGPT for iOS/Android/web.
  • Current limitations: no video or screen sharing, and full multilingual support not yet reached. The API became generally available on September 10, 2026 ($0.05 per minute). GPT-Realtime-2.1 is a separate line of API voice models.

FAQ

Q1. Can I use GPT-Live for free?

Yes. The free plan can use GPT-Live-1 mini with limited access. As of September 25, 2026, Go also uses GPT-Live-1 mini (up to 3 hours a day), while Plus and Pro use the higher-end GPT-Live-1. It's available on ChatGPT for iOS, Android, and web (staged rollout).

Q2. What exactly is "full-duplex"?

It's a method that can listen and speak at the same time, like a phone call. Unlike the old turn-based approach (waiting for you to finish before responding), it can return backchannels while you're talking, absorb interruptions, and wait through your thinking silences without cutting in.

Q3. What's the single biggest difference from the old Advanced Voice Mode?

How it detects the end of a turn. The old approach was silence-based, so it tended to mistake a thinking pause for "done" and interrupt. GPT-Live doesn't wait for a clean break: it follows pauses, interruptions, and changes in pace as it listens, deciding many times per second whether to respond or keep listening (OpenAI's system card), so it's less likely to cut you off and feels more natural. It's also clearly preferred in human evaluations.

Q4. Can it answer difficult questions?

Yes. GPT-Live handles conversational responsiveness, and when web search or deep reasoning is needed, it delegates the processing to a model behind the scenes and brings the results back into the conversation. At launch that model was GPT-5.5; since September 9, 2026, it uses GPT-5.6 or GPT-6 Astra, and you pick the model and reasoning effort with the same controls as text chat.

Q5. Can it do video or screen sharing?

Not as of September 25, 2026. Live doesn't support video or screen sharing, so when you need them, switch to the older Advanced voice in the iOS or Android app (OpenAI Help). At launch, OpenAI said it was working to add these capabilities.

Q6. Can I integrate it into my own app via API?

Yes. The GPT-Live API became generally available on September 10, 2026; the model ID is gpt-live-1, and voice sessions cost $0.05 per minute (the backend model and tools are billed separately). There's also the separate GPT-Realtime-2.1 / mini line, so choose based on whether you prioritize natural conversation or want to build on the existing Realtime API.

Q7. How does it compare to Google's Gemini Live?

Google's Gemini Live lets you share your camera or your phone's screen while you talk (Gemini Live help). As of September 25, 2026, GPT-Live doesn't support video or screen sharing, so if you want to show what you're looking at, Gemini Live has the edge. The strength OpenAI highlights for GPT-Live is natural full-duplex conversation: it gives backchannel cues while listening, handles interruptions mid-sentence, and waits through the pauses while you think. OpenAI hasn't published latency figures, so response speed can't be compared. For more, see the GPT-5.6 vs Gemini comparison.

If your goal is to transcribe an existing recording rather than have a live voice conversation, see how to upload audio files to ChatGPT. Audio attachments, Voice, dictation, and Record have different workflows and limits.

Related articles