Gemini 3.8 Live models enable real-time voice interaction with visual grounding

Research

Gemini 3.8 Live models enable real-time voice interaction with visual grounding

Google DeepMind launches two new AI models for fluid voice conversations with background task execution and multilingual support.

Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking

Published

September 15, 2026

Reading time

2 minutes

Perspective

Research

Topics

Gemini 3.8 Live · real-time voice AI · k-server conjecture

Google DeepMind has introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two models designed to enhance voice-based interactions through real-time visual context processing, multi-step reasoning, and uninterrupted task execution. These models are available via the Gemini API, Google Workspace, Search, and the Gemini app, with enterprise-grade performance demonstrated on benchmarks such as Artificial Analysis' Speech to Speech Quality Index and ServiceNow’s EVA-Bench.

Real-time visual and language integration in Gemini 3.8 Live

Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking
Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking

Gemini 3.8 Live processes visual inputs in near real-time, enabling context-aware responses during voice conversations. It automatically detects and transitions between 97 languages mid-conversation and executes tool calls in the background while maintaining dialogue flow, as demonstrated in live onboarding and chess-playing scenarios.

Source: Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking · DeepMind Blog

Agent consistency gap revealed in repeated task execution

Your Agent Aced the Task. Will It Do It Again?
Your Agent Aced the Task. Will It Do It Again?

Hugging Face’s analysis shows that a ReAct agent using GPT-4.1 succeeds on 77.4% of task runs on average but only completes all five repetitions successfully in 53.0% of cases. This 24.4-point consistency gap arises from flat decision distributions where minor perturbations alter token selection, even at temperature zero.

Source: Your Agent Aced the Task. Will It Do It Again? · Hugging Face Blog

k-server conjecture proven with tight competitive bound

The k-server conjecture is true
The k-server conjecture is true

Researchers have proven the k-server conjecture, confirming that the Work Function Algorithm achieves optimal k-competitiveness. The result guarantees that any online strategy for routing k servers to serve dynamic requests will never exceed twice the cost of an optimal offline solution, regardless of the number of servers or request patterns.

Source: The k-server conjecture is true · Hacker News

Perplexity deploys GPT-6 Astra for autonomous system operations

Perplexity trusts GPT-6 Astra with end-to-end systems
Perplexity trusts GPT-6 Astra with end-to-end systems

Perplexity uses GPT-6 Astra to autonomously write internal communications, modify software, and monitor production systems, requiring significantly fewer human check-ins than with prior models. This deployment reflects a shift toward sustained, low-intervention agent operation in enterprise environments.

Source: Perplexity trusts GPT-6 Astra with end-to-end systems · OpenAI News

Codex and ChatGPT identify antimicrobial candidates from genomic data

How a researcher uses Codex and ChatGPT to search for new antimicrobial molecules
How a researcher uses Codex and ChatGPT to search for new antimicrobial molecules

César de la Fuente’s lab employs Codex and ChatGPT to scan living and extinct genomes for sequences with antimicrobial potential. The models assist in filtering genomic data to prioritize candidates for experimental validation, accelerating the search for solutions to drug-resistant infections.

Source: How a researcher uses Codex and ChatGPT to search for new antimicrobial molecules · OpenAI News

What to watch next

These developments highlight advances in real-time multimodal interaction, theoretical algorithmic guarantees, autonomous agent reliability, and applied bioinformatics. Each system operates within defined technical boundaries: Gemini models rely on proprietary infrastructure, the k-server proof applies only to deterministic online scenarios, agent consistency remains sensitive to model output distributions, and genomic screening requires experimental validation.

Continue reading

More from COREXA