Gemini Omni
Retour aux articles
10 min de lecture

Guide Gemini 3.1 Flash Live API : agents vocaux, caméra et partage d'écran en temps réel

Créez des agents multimodaux à faible latence avec Gemini 3.1 Flash Live — architecture WebSocket, audio natif, caméra et partage d'écran à 1 FPS, appels d'outils et jetons éphémères.

Gemini Live APIGemini 3.1 Flash LiveReal-TimeVoice AgentsMultimodalDevelopers2026Français

Pourquoi Gemini 3.1 Flash Live compte en août 2026

Google a lancé Gemini 3.1 Flash Live via la Gemini Live API en mars 2026 ; en août, c’est le modèle recommandé pour les agents voix+vision en production. Contrairement aux appels batch generateContent, Live maintient une session WebSocket persistante streamant audio, images vidéo et texte — le modèle entend, voit et parle sans services STT/TTS séparés.

En août 2026 :

  • Modèles 2.5 Flash Live dépréciés — migrer vers gemini-3.1-flash-live-preview.
  • Écosystème mature — Stitch, Ato, Weekend RPG en production.
  • Live API + Gemini 3.7 Flash — Live pour le dialogue ; 3.7 Flash REST pour le code (voir annonce 3.7 Flash).

Architecture: WebSocket sessions

The Live API is stateful — persistent WSS streams audio, video frames (≤1 FPS JPEG/PNG), and text bidirectionally. Setup via BidiGenerateContentSetup, then BidiGenerateContentRealtimeInput.

LayerRole
TransportStateful WSS
InputPCM16 @ 16 kHz, video ≤1 FPS, text
OutputPCM @ 24 kHz, transcripts, tool calls

Deploy server-to-server (proxy + logging) or client-to-server (lower latency; use ephemeral tokens).

Native end-to-end audio

Input: 16-bit PCM LE mono 16 kHz. Output: 24 kHz PCM. Convert browser float32 capture to PCM16. VAD handles turns; push-to-talk uses manual activityStart/activityEnd.

Camera and screen share

Stream JPEG/PNG at 1 FPS max — camera Q&A, screen-share code review (Stitch demo), accessibility. Default TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO bills all frames in a turn; send video only during speech to save cost.

Session limits: 15 min audio-only, 2 min audio+video — extend via session resumption and context compression.

Tool calling

Synchronous function calling + Google Search grounding. 3.1 Flash Live does not support NON_BLOCKING async tools from 2.5.

thinkingLevel and 90+ languages

Use thinkingLevel: minimal (default, lowest latency) through high. Supports 90+ languages with improved tone and pace recognition vs 2.5 Native Audio.

Security

Never ship API keys in clients — mint ephemeral tokens on your backend. For WebRTC scale, use partner integrations (Fishjam, Stream Vision Agents, Voximplant).

Migrate from 2.5 Flash Live

Replace deprecated models with gemini-3.1-flash-live-preview. Switch thinkingBudgetthinkingLevel; use send_realtime_input for mid-session text; handle multi-part server events; review turnCoverage.

3.1 Live vs 3.7 Flash REST

3.1 Live3.7 REST
ProtocolWebSocketHTTP
LatencySub-second audio85ms+ first token
Context128K session2.5M tokens
Best forVoice agents, screen assistCoding agents, RAG

Get started

  1. Google AI StudioStream
  2. pip install google-genai or npm install @google/genai
  3. Capabilities guide