Information as of September 16, 2026. An analysis of the launch and published research; the voice sample is embedded below.
Imagine a conversation with an assistant you are showing an app sketch to. You explain what it should do, change your mind a moment later, add a constraint and ask about the cost of building it. You don't want to finish dictating, wait for an answer and start over every single time. You want to keep talking while the other side checks information and does the work. That promise is exactly why Google's latest release invites the question in the title.
Gemini 3.8 Live Extended Thinking gives no grounds for declaring AGI. It does show a direction in which an assistant can become a far more useful collaborator. The distinction matters: the impression of an intelligent presence, a benchmark score and day-to-day reliability are three separate things. Each one deserves a separate look, instead of settling the whole question on the strength of a flashy demo.
Graphic: Google, launch materials from September 15, 2026. Announcement.
What actually shipped?
On September 15 Google released two models: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. The Gemini API changelog describes both as generally available, meaning GA. The first is the default option for fast voice interactions; the second is meant for tasks that need more elaborate reasoning in the background. This is a separate launch from Gemini 3.8 Flash and Flash Cyber, announced on September 2. A shared version number does not mean an identical product or interchangeable test results. Changelog, Flash launch.
The Gemini 3.8 Audio model card indicates that the Live models are built on Gemini 3 Pro. They accept audio, images, video and text, and they produce voice and text. A bump in the version number alone therefore says nothing about whether the model's foundation, the training approach, tool handling or the product wrapper has changed. What is most interesting here is the combination of model capability with continuous conversation. Model card.
That also puts expectations in order. We did not get an announcement of a universal system that will replace any specialist on its own. We got specialized interaction models whose value has to be checked while speaking, listening, using context and carrying out specific instructions.
The conversation doesn't have to wait for the computation
The Extended Thinking documentation describes reasoning and tool calls performed in the background. During that time the model can send short messages instead of leaving the user in silence. The available levels are low, medium and high. Regular Live also supports reasoning interleaved with interaction, but it does not expose the thinkingLevel setting. So the difference should not be reduced to the slogan "one thinks, the other doesn't." Thinking in the Live API.
Picture a sample travel integration: the assistant searches through offers while we add that the return flight has to be in the evening. The value of that architecture would lie in accepting the correction mid-work, keeping the earlier constraints and carrying the search through to the end. That is an example of a possible use, not a report from a test of mine.
The most interesting question, then, is whether the model correctly manages what we have already established. The ability to say "checking now" is easy to demonstrate. What is harder is recognizing a contradiction, withdrawing an outdated plan and not performing an action the user has just canceled. That is how I would judge a conversational partner's practical intelligence.
Voice, vision and tools in one task
Google shows demonstrations of building React components from a sketch and spoken notes, coordinating reservations, and talking about a chessboard in view. It also declares automatic switching across 97 languages. These are vendor materials: they show the intended way of using the product, but they do not give the probability of success for any similar task of your own. Launch materials.
In my view the most important part is being able to refer to a specific object together. Talking about "that button" or "the error in the right-hand corner" can cut the effort needed to describe a problem. For a developer, a teacher or a support agent, that would save a lot of small steps.
At the same time, a camera does not automatically solve the problem of understanding. The assistant can confuse screen elements, read an outdated view or miss a change. In a good product it should be able to say which part of the image it is referring to and ask for a better view. A polished demonstration is where evaluation begins, not where it ends.
Benchmarks: the score is strong
In the Artificial Analysis listing, checked on September 16, Gemini 3.8 Live Extended Thinking in the High setting scores 82.6 points on the Speech to Speech Index. Regular Live reaches 76.0. The Extended Thinking result on the agentic τ-Voice is 68.6%, against 30.1% for regular Live. On Big Bench Audio the stronger variant takes 97.7%. These are concrete pieces of evidence of progress, although each one measures a different slice of capability. Current listing.

Graphic source: Google; measurements: Artificial Analysis. Model settings are part of the comparison. Launch methodology.
On the launch chart for the index, the closest competitors shown are GPT-Live-1 Astra Medium at 81.5 and Grok Voice Think Fast 2.0 High at 81.3. Gemini's lead is therefore far more modest than the headline about first place might suggest. Without an analysis of measurement uncertainty and repeatability, I would not call it a gulf. Official evaluation report.
The practical conclusion is simple: we have a reason to try the model out. We have no reason to derive from a single ranking the claim that it will be the best in every company, every language and every conversation.

Google graphic: τ-Voice results according to Artificial Analysis. The High variant denotes a specific reasoning level, not a separate base model.
What does 82.6 points mean?
The Speech to Speech Index is not an AGI test. The current methodology combines four equally weighted components: voice reasoning, agentic task execution, user preference and Task Success Rate. This is version 2.0 of the index; earlier variants had a different composition. When comparing historical results, you therefore have to check whether you are comparing the same measure. Artificial Analysis methodology.
Big Bench Audio contains a thousand audio questions from four categories: logical inference, navigation, object counting, and truth-and-lie puzzles. A high score tells you about performance on that kind of material. It does not mean command of every field or resistance to every kind of misunderstanding. Big Bench Audio description.
It is tempting to treat a near-maximum score as evidence of universal intelligence. Yet correctly solving a puzzle you heard and walking a customer through a messy complaint call take different skills. In the second case, missing data, changes of mind and the limits of external systems all come into play.
So when I read a ranking, what I look for first is whether the test resembles the work I want to hand to the model. A place on the podium is only the next piece of information after that.
Harder tasks show the limits
Google reports 35.1% on a banking test, called τ-Voice-banking in the announcement and τ³-Banking on the chart. The methodology report describes work with a large, unstructured knowledge base and multi-step tool calls. I am keeping both source labels rather than pretending that all variants of banking benchmarks are interchangeable. Google report.

Google publishes this chart as τ³-Banking while describing the result in the launch text as τ-Voice-banking. The measurements concern a specific test configuration, not serving real bank customers.
For me this is the most important number of the whole launch. A score can be enough to lead a difficult comparison and at the same time sit very far from the reliability expected of an independent worker. Both judgments can be true at once.
It also should not be read as a forecast of the error rate of a future application. Real-world performance depends on the type of cases, the integrations, the data and the option to hand a conversation over to a human. A benchmark reveals limits in a specific environment; it does not hand you a ready risk calculation for any deployment.
Extended Thinking doesn't win everything
Artificial Analysis reports 96.1% for regular Live on conversation dynamics, and 91.9% for Extended Thinking High. Average time to first audio is 1.18 and 1.35 seconds respectively. The regular variant also achieves a higher Task Success Rate in that listing: 93.2% against 89.1%. That is a separate measure from τ-Voice. Results table.
Choosing a model becomes more interesting as a result. If the conversation is short, predictable and calls for a simple data lookup, more elaborate reasoning may bring no benefit. For a diagnosis, or for weighing many dependencies against each other, the extra effort can pay off. I would not pick a model purely by the length of its name.
It is also worth separating the first reaction from the completion of the task. A quick "checking that now" improves the sense of contact, but it says nothing about when we will get a correct result. If I were buying such a service, I would measure both times separately.
EVA-Bench also asks about conversation quality
EVA-Bench separates correctness of execution from user experience. The authors describe, among other things, scoring of task completion, consistency of answers with the available information, and the flow of the conversation. The set covers 213 scenarios across three domains and a study of the effect of disruptions. That is a useful complement to a reasoning ranking on its own. The EVA-Bench paper.

Official Google chart based on a ServiceNow evaluation in the Gemini Enterprise Agent Platform. The "minimal" label comes from that study; the public Gemini API for Extended Thinking documents the levels low, medium and high.
This chart illustrates my main argument well: a pleasant conversation and a correctly completed task are not the same thing. An assistant can be polite, fast and articulate and still carry out the instruction badly. It can also find the right solution while wearing the user out with a long monologue. A good product has to combine both qualities.
What does a developer get?
The model documentation gives an input limit of 131,072 tokens and an output limit of 65,536. Extended Thinking supports Google Search and asynchronous function calling. It does not offer built-in code execution, File Search or structured outputs. You can build your own integrations, but they should not be attributed automatically to the model itself. Extended Thinking specification.
There is also an important protocol change: in Extended Thinking, turnComplete can mean the end of an utterance while the work is still running. The state of the whole task is given by interaction_status, which takes the values IN_PROGRESS and IDLE. The application should keep receiving messages instead of closing the handler after the first response. Live API capability comparison.
From the user's point of view this is an invisible detail. From the product builder's point of view it can decide whether we hear the final result or only a polite announcement. That is why the quality of an assistant also depends on the application code, error handling and confirmation of tool results.
What does such a conversation cost?
The Gemini API price list assigns both Live models the same unit rates: input audio costs $0.005 per minute, output audio $0.018. Text is billed separately: $0.75 per million input tokens and $4.50 per million output tokens, with thinking tokens counted in the output category. Other billable elements can be added on top. Official pricing.
A purely arithmetic example: ten minutes of uploaded audio and five minutes of responses would cost $0.14 for the audio part alone. That is not the price of a complete service or a guarantee of the cost of a fifteen-minute session. What counts is the streams actually processed and any additional usage.
How does the competition compare? Below are standard API rates in USD, without extra services or taxes. The units are deliberately left visible: a minute of session, a minute of audio and a million tokens are not equivalent.
| Model or service | Base voice billing | Additional costs |
|---|---|---|
| Gemini 3.8 Live | $0.005/min input + $0.018/min output | Text, reasoning, tools |
| Gemini 3.8 Live Extended Thinking | Same audio rates as Live | Reasoning usage can be higher |
| GPT-Live 1 | $0.05/min of session time | Backend model and tools billed separately |
| Grok Voice Think Fast 2.0 | $0.08/min of audio | The price list also shows $0.004 for text input |
| GPT-Realtime-2.1 | $32 per million input audio tokens; $64 per million output | Text, images and other usage |
| GPT-Realtime-2.1-mini | $10 per million input audio tokens; $20 per million output | Text, images and other usage |
Rate sources: Google, OpenAI, Grok.
Fifteen minutes of a GPT-Live 1 session comes to $0.75 for the voice layer, before the backend model is billed. Fifteen billed minutes of Grok audio means $1.20. These are examples of how different tariffs behave, not an experiment showing that the models did identical work. For a thousand repetitions of the described scenarios we get $140 for the Gemini audio part, $750 for GPT-Live sessions and $1,200 for Grok audio. Each amount needs the elements left out of the assumptions added back in.

The cost in the Artificial Analysis chart is normalized to input audio time for a specific set of tasks. It covers more than the input rate alone; it is not a universal price list for conversations.
When comparing offers, what would interest me in the end is the cost of a case handled correctly. A cheap model whose output has to be fixed repeatedly can turn out more expensive to operate. A pricier variant, though, does not have to justify its price on simple instructions.
The launch chart for a fixed subset of Big Bench Audio gives $0.84 for regular Live, $3.50 for Extended Thinking High, $4.80 for Grok High and $5.83 for GPT-Live-1 Astra Medium. In that particular sample, Extended Thinking costs about 4.17 times more than Live, despite identical Google unit rates. The result illustrates the effect of usage, configuration and response length. That multiplier must not be carried over directly to any application. Google cost chart.
Availability: a launch is not one switch
Google has started rolling the models out in the Gemini API and AI Studio. The announcement also lists Search Live for the Live variant, and Gemini Live plus selected Workspace experiences for Extended Thinking. Enterprise offerings have a separate schedule that includes a private preview. The availability of a specific feature has to be distinguished from the availability of the model through the API. Launch schedule.
So I would not assume that every user will immediately see an identical feature in their account. I also have not checked how all the variants behave on a Polish account. An official declaration of multilingual support is no substitute for testing Polish surnames, addresses, industry abbreviations or the mixing of Polish with English.
Limits you can't hear in the voice
The model card explicitly lists hallucinations, possible slowdowns and timeouts. It also gives January 2025 as the training knowledge boundary. Access to current information can come from search or tools, but a launch date does not automatically refresh all the knowledge stored in the model. Vendor limitations.
Session continuity is a separate matter. The Live API documentation describes context compression and connection resumption. These solve part of the problem with long conversations, but they are not a promise of flawless memory of everything agreed or of lasting learning about the user. Session management.
For me, memory of constraints would be one of the first tests. I would set a condition at the start of the conversation, change the subject several times and check the result later. Then I would correct an earlier piece of information and check whether the assistant uses the new version. Only trials like that say anything about working together over a longer stretch.
What a three-minute voice sample shows
I ran one acting scene, "The Last Rehearsal," with two live voice models improvising against each other. OpenAI's gpt-realtime-2.1 played Alex in the Cedar voice; Gemini 3.8 Live Extended Thinking played Morgan in the Kore voice with thinking set to HIGH. Twelve turns, 180 seconds of edited audio. Each side received the partner's last line as audio and answered in its own native voice, with no external text-to-speech in between.
The limits matter as much as the recording. The turns alternate, so nothing here measures full-duplex conversation, and network waiting time was cut out, so the video says nothing about real latency. Gemini repeatedly spoke remarks about preparing its answer instead of staying in character, and in the final turn it returned only such remarks plus an IDLE signal, never closing the scene as instructed. One scene cannot separate the influence of role, voice, prompt and model, so it is no proof that one vendor is better.
Open the video file directly (MP4, 3 MB, 3 min).
Video: the full second take, all 12 turns in order, no cherry-picking. Subtitles are API transcripts, not independent ASR.
Why this still isn't proof of AGI
In the paper "Levels of AGI," DeepMind researchers propose judging systems by level of task performance, breadth of competence and autonomy. That perspective is more useful than asking whether a conversation sounds human. A system can be very capable within a particular class of interaction and still need substantial support outside it. Levels of AGI.
In my view, proof of generality would require a much wider set of trials: novel problems, previously unknown rules, transfer of knowledge between domains, and long stretches of operation without a human constantly stepping in. Recognizing the limits of its own knowledge and recovering after a mistake would count too.
The published voice benchmarks settle none of this. Nor do they show that the system picks worthwhile goals on its own, learns an arbitrary skill or reaches human versatility. This is not a charge against the product. The AGI claim is simply much broader than the claim of a very good voice assistant.
We also don't need to settle the question of consciousness to judge how useful a tool is. It is enough to check whether it completes tasks correctly, whether it communicates its limits, and whether a person can effectively correct what it does. Emotional naturalness in conversation should not stand in for those criteria.
The breakthrough may be closer to everyday work
Three possible uses interest me most: technical support with a live image, collaborative design, and learning through conversation. In each of them the user often discovers what they actually need only during the interaction. The assistant should keep up with that process instead of expecting a perfect instruction up front.
Before deploying anything, I would compare regular Live and Extended Thinking on the same set of real cases. I would check the correctness of the final result, the reaction to an interruption, the behavior after a tool error, and the cost. I would also include cases where the right answer is a request for clarification. That is my proposal for assessing the product, not an extra Google benchmark.
I would add one more trial with an ambiguous instruction. The user says "move that to tomorrow," but two meetings came up earlier in the conversation. A good assistant should notice the ambiguity instead of picking an event at random. That small situation separates fluent speech generation from responsible collaboration very cleanly. In daily use, moments like it may matter more than another point scored on a logic task.
Do we have AGI yet? On the basis of this launch, nobody can honestly say so. What we do have is a strong signal that the voice interface is starting to be fit for more complex collaboration. If quality like this holds up outside demonstrations, it may change everyday work faster than another round of arguments about definitions.
The moment that matters most will come when, after a conversation, we see the task done correctly. With no dropped condition, no faked success and no need to start from scratch. That is the test I am waiting for. In the end we judge the result that is left once the talking stops.
Sources
- Google, Gemini 3.8 Live and Live Extended Thinking announcement: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/
- Gemini API changelog: https://ai.google.dev/gemini-api/docs/changelog
- Google, Gemini 3.8 Flash and Flash Cyber announcement: https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/
- Gemini 3.8 Audio model card: https://deepmind.google/models/model-cards/gemini-3-8-audio/
- Thinking in the Live API: https://ai.google.dev/gemini-api/docs/live-api/thinking
- Artificial Analysis, Speech to Speech listing: https://artificialanalysis.ai/speech-to-speech
- Google DeepMind, Gemini 3.8 Live evaluation methodology: https://storage.googleapis.com/deepmind-media/gemini/gemini_3-8_live_model_evaluation.pdf
- Artificial Analysis, Speech to Speech benchmarking methodology: https://artificialanalysis.ai/methodology/speech-to-speech-benchmarking
- Artificial Analysis, speech reasoning benchmarking: https://artificialanalysis.ai/methodology/speech-to-speech-benchmarking
- EVA-Bench paper: https://arxiv.org/abs/2605.13841
- Gemini 3.8 Live Extended Thinking model specification: https://ai.google.dev/gemini-api/docs/models/gemini-3.8-live-extended-thinking
- Live API capability comparison: https://ai.google.dev/gemini-api/docs/live-api/capabilities
- Gemini API pricing: https://ai.google.dev/gemini-api/docs/pricing
- OpenAI API pricing: https://developers.openai.com/api/docs/pricing
- xAI (Grok) pricing: https://docs.x.ai/developers/pricing
- Live API session management: https://ai.google.dev/gemini-api/docs/live-api/session-management
- Levels of AGI paper: https://arxiv.org/abs/2311.02462
