gemini-live-api-dev
Use this skill when building real-time, bidirectional streaming applications with the Gemini Live API. Covers WebSocket-based audio/video/text streaming, voice activity detection (VAD), native audio features, function calling, session management, ephemeral tokens for client-side auth, live translation, and all Live API configuration options. SDKs covered - google-genai (Python), @google/genai (JavaScript/TypeScript).
DeepseekModel
キュレーション済みスキル
品質 優秀 · 90
v1.0.0
取得
https://deepseekmodel.com/api/download.php?id=google-gemini-gemini-skills-skills-gemini-live-api-dev-skill-md&format=skill
ダウンロード .skill
標準形式。system_prompt と model_config を収録し、任意の Agent で利用可能
.skill ファイルの system_prompt フィールドの実際の内容。
name gemini-live-api-dev description Use this skill when building real-time, bidirectional streaming applications with the Gemini Live API. Covers WebSocket-based audio/video/text streaming, voice activity detection (VAD), native audio features, function calling, session management, ephemeral tokens for client-side auth, live translation, and all Live API configuration options. SDKs covered - google-genai (Python), @google/genai (JavaScript/TypeScript). Gemini Live API Development Skill Overview The Live API enables low-latency, real-time voice and video interactions with Gemini over WebSockets. It processes continuous streams of audio, video, or text to deliver immediate, human-like spoken responses. Key capabilities: Bidirectional audio streaming — real-time mic-to-speaker conversations Video streaming — send camera/screen frames alongside audio Text input/output — send and receive text within a live session Audio transcriptions — get text transcripts of both input and output audio Voice Activity Detection (VAD) — automatic interruption handling Native audio — thinking (with configurable thinkingLevel ) Function calling — synchronous tool use Google Search grounding — ground responses in real-time search results Session management — context compression, session resumption, GoAway signals Ephemeral tokens — secure client-side authentication [!NOTE] The Live API currently only supports WebSockets . For WebRTC support or simplified integration, use a partner integration . Models gemini-3.1-flash-live-preview — Optimized for low-latency, real-time dialogue. Native audio output, thinking (via thinkingLevel ). 128k context window. This is the recommended model for all Live API use cases. gemini-3.5-transcribe-live — Real-time streaming speech-to-text with interim hypotheses, finalized transcripts, smart formatting, and Hybrid VAD. gemini-3.5-live-translate-preview — Real-time streaming translation model. [!WARNING] The following Live API models are deprecated and will be shut down. Migrate to gemini-3.1-flash-live-preview . gemini-2.5-flash-native-audio-preview-12-2025 — Migrate to gemini-3.1-flash-live-preview . gemini-live-2.5-flash-preview — Released June 17, 2025. Shutdown: December 9, 2025. gemini-2.0-flash-live-001 — Released April 9, 2025. Shutdown: December 9, 2025. SDKs Python : google-genai — pip install google-genai JavaScript/TypeScript : @google/genai — npm install @google/genai [!WARNING] Legacy SDKs google-generativeai (Python) and @google/generative-ai (JS) are deprecated. Use the new SDKs above. Partner Integrations To streamline real-time audio/video app development, use a third-party integration supporting the Gemini Live API over WebRTC or WebSockets : LiveKit — Use the Gemini Live API with LiveKit Agents. Pipecat by Daily — Create a real-time AI chatbot using Gemini Live and Pipecat. Fishjam by Software Mansion — Create live video and audio streaming applications with Fishjam. Vision Agents by Stream — Build real-time voice and video AI applications with Vision Agents. Voximplant — Connect inbound and outbound calls to Live API with Voximplant. Firebase AI SDK — Get started with the Gemini Live API using Firebase AI Logic. Audio Formats Input : Raw PCM, little-endian, 16-bit, mono. 16kHz native (will resample others). MIME type: audio/pcm;rate=16000 Output : Raw PCM, little-endian, 16-bit, mono. 24kHz sample rate. [!IMPORTANT] Use send_realtime_input / sendRealtimeInput for all real-time user input (audio, video, and text ). send_client_content / sendClientContent is only supported for seeding initial context history (requires setting initial_history_in_client_content in history_config ). Do not use it to send new user messages during the conversation. [!WARNING] Do not use media in sendRealtimeInput . Use the specific keys: audio for audio data, video for images/video frames, and text for text input. Quick Start Authentication Python from google import genai client = genai.Client(api_key= "YOUR_API_KEY" ) JavaScript import { GoogleGenAI } from '@google/genai' ; const ai = new GoogleGenAI ({ apiKey : 'YOUR_API_KEY' }); Connecting to the Live API Python from google.genai import types config = types.LiveConnectConfig( response_modalities=[types.Modality.AUDIO], system_instruction=types.Content( parts=[types.Part(text= "You are a helpful assistant." )] ) ) async with client.aio.live.connect(model= "gemini-3.1-flash-live-preview" , config=config) as session: pass # Session is active JavaScript const session = await ai. live . connect ({ model : 'gemini-3.1-flash-live-preview' , config : { responseModalities : [ 'audio' ], systemInstruction : { parts : [{ text : 'You are a helpful assistant.' }] } }, callbacks : { onopen : () => console . log ( 'Connected' ), onmessage : ( response ) => console . log ( 'Message:' , response), onerror : ( error ) => console . error ( 'Error:' , error), onclose : () => console . log ( 'Closed' ) } }); Sending Text Python await session.send_realtime_input(text= "Hello, how are you?" ) JavaScript session. sendRealtimeInput ({ text : 'Hello, how are you?' }); Sending Audio Python await session.send_realtime_input( audio=types.Blob(data=chunk, mime_type= "audio/pcm;rate=16000" ) ) JavaScript session. sendRealtimeInput ({ audio : { data : chunk. toString ( 'base64' ), mimeType : 'audio/pcm;rate=16000' } }); Sending Video Python # frame: raw JPEG-encoded bytes await session.send_realtime_input( video=types.Blob(data=frame, mime_type= "image/jpeg" ) ) JavaScript session. sendRealtimeInput ({ video : { data : frame. toString ( 'base64' ), mimeType : 'image/jpeg' } }); Receiving Audio and Text [!IMPORTANT] A single server event can contain multiple content parts simultaneously (e.g., audio chunks and transcript). Always process all parts in each event to avoid missing content. Python async for response in session.receive(): content = response.server_content if content: # Audio — process ALL parts in each event if content.model_turn: for part in content.model_turn.parts: if part.inline_data: audio_data = part.inline_data.data # Transcription if content.input_transcription: print ( f"User: {content.input_transcription.text} " ) if content.output_transcription: print ( f"Gemini: {content.output_transcription.text} " ) # Interruption if content.interrupted is True : pass # Stop playback, clear audio queue JavaScript // Inside the onmessage callback const content = response. serverContent ; if (content?. modelTurn ?. parts ) { for ( const part of content. modelTurn . parts ) { if (part. inlineData ) { const audioData = part. inlineData . data ; // Base64 encoded } } } if (content?. inputTranscription ) console . log ( 'User:' , content. inputTranscription . text ); if (content?. outputTranscription ) console . log ( 'Gemini:' , content. outputTranscription . text ); if (content?. interrupted ) { /* Stop playback, clear audio queue */ } Live Translation (Gemini Live Translate) The Live API supports real-time, low-latency streaming translation of speech (audio) across 70+ languages. For full details on options and capabilities, see the Live Translate Guide . Model gemini-3.5-live-translate-preview — The recommended translation model for all Live Translate use cases. Configuration ( TranslationConfig ) To enable translation, specify a TranslationConfig object inside your live session setup: Python SDK : Configure the connection using translation_config on LiveConnectConfig : config = types.LiveConnectConfig( response_modalities=[types.Modality.AUDIO], translation_config=types.TranslationConfig( target_language_code= "es" , # Target language code (e.g. es, fr, pl) echo_target_language= True , ), input_audio_transcription=types.AudioTranscriptionConfig(), output_audio_transcription=types.AudioTranscriptionConfig(), ) Raw WebSockets : Place translationConfig inside generationConfig : { "setup" : { "model" : "models/gemini-3.5-live-translate-preview" , "generationConfig" : { "responseModalities" : [ "AUDIO" ] , "translationConfig" : { "targetLanguageCode" : "es" , "echoTargetLanguage" : true } } } } Live Streaming Transcription (Gemini Live Transcribe) The Live API supports real-time streaming speech-to-text over WebSockets with low-latency interim hypotheses, finalized transcripts, and Hybrid VAD. For full details, see the Live Transcription Guide and Colab Cookbook . Model gemini-3.5-transcribe-live Modes smart : cleans up filler words, resolves inline self-corrections, and structures formatting. verbatim (default): exact word-for-word transcript. Python config = types.LiveConnectConfig( response_modalities=[ "TEXT" ], input_audio_transcription=types.AudioTranscriptionConfig(), ) async with client.aio.live.connect(model= "gemini-3.5-transcribe-live" , config=config) as session: # Stream audio await session.send_realtime_input(audio=types.Blob(data=chunk, mime_type= "audio/pcm;rate=16000" )) # Hybrid VAD: notify turn end on client-detected silence for zero latency await session.send_realtime_input(audio_stream_end= True ) JavaScript const session = await ai. live . connect ({ model : 'gemini-3.5-transcribe-live' , config : { responseModalities : [ 'text' ], inputAudioTranscription : { mode : 'smart' } }, callbacks : { onmessage : ( msg ) => { if (msg. serverContent ?. interimInputTranscription ) { console . log ( 'Interim:' , msg. serverContent . interimInputTranscription . text ); } if (msg. serverContent ?. inputTranscription ) { console . log ( 'Final:' , msg. serverContent . inputTranscription . text ); } } } }); session. sendRealtimeInput ({ audio : { data : chunkBase64, mimeType : 'audio/pcm;rate=16000' } }); session. sendRealtimeInput ({ audioStreamEnd : true }); // Hybrid VAD Raw WebSockets { "setup" : { "model" : "models/gemini-3.5-transcribe-live" , "generationConfig" : { "responseModalities" : [ "TEXT" ] , "speechConfig" : { "voiceConfig" : { } } } , "inputAudioTranscription" : { "mode" : "smart" } } } Limitations Response modality — Only TEXT or AUDIO per session, not both. Native audio models only support audio. Audio-only session — 15 min without compression Audio+video session — 2 min without compression Connection lifetime — ~10 min (use session resumption) Context window — 128k tokens (native audio) / 32k tokens (standard) Async function calling — Not yet supported; function calling is synchronous only. The model will not start responding until you've sent the tool response. Proactive audio — Not yet supported in Gemini 3.1 Flash Live. Remove any configuration for this feature. Affective dialogue — Not yet supported in Gemini 3.1 Flash Live. Remove any configuration for this feature. Code execution — Not supported URL context — Not supported Migrating from Gemini 2.5 Flash Live When migrating from gemini-2.5-flash-native-audio-preview-12-2025 to gemini-3.1-flash-live-preview :
このスキルを起動するキーワード。クリックでコピーできます。
このスキルにはトリガーワードがありません。
ダウンロードした .skill に含まれるフィールド。
| フィールド | 説明 |
|---|---|
| format | フォーマット識別子(skill/v1) |
| skill_id | スキル固有 ID |
| name | スキル名 |
| version | バージョン |
| description | 説明 |
| category | カテゴリ(配列) |
| trigger_words | トリガーワード |
| tags | タグ |
| source | ソース |
| source_url | ソース URL(本ページ) |
| exported_at | エクスポート日時(ダウンロード毎) |
| system_prompt | システムプロンプト本文 |
| model_config | モデル設定:provider / model / temperature / max_tokens / top_p |
| examples | サンプル |
| install_guide | 各プラットフォームの導入説明(Coze / Dify / Claude / カスタム) |