> ## Documentation Index
> Fetch the complete documentation index at: https://visionagents.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Gemini Realtime

<iframe className="w-full aspect-video rounded-xl" src="https://www.youtube.com/embed/8lA6bF2EnvA" title="Gemini Live integration" frameBorder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowFullScreen />

[Google's Gemini](https://ai.google.dev/gemini-api/docs/live) provides native multimodal speech-to-speech over WebSocket with optional video. No separate STT/TTS services required. Requires `google-genai>=2.19.0`.

<Info>
  Vision Agents uses [Stream Video](https://getstream.io/video/) for real-time WebRTC transport by default. [External WebRTC transports](/integrations/introduction-to-integrations#edge-transport) are supported as well. Most AI providers offer free tiers to get started.
</Info>

<Tip>
  Gemini also provides a traditional [LLM](/integrations/llm/gemini) with built-in tools for search, code execution, and RAG, plus streaming [speech-to-text](/integrations/stt/gemini) for custom voice pipelines.
</Tip>

## Installation

```sh theme={null}
uv add "vision-agents[gemini]"
```

Set `GOOGLE_API_KEY` or `GEMINI_API_KEY` in your environment.

## Quick Start

The default Live model is `gemini-3.8-live` (latency-optimized audio-to-audio). Video frames are forwarded when `fps` is set.

```python theme={null}
from vision_agents.core import Agent, User
from vision_agents.plugins import gemini, getstream

agent = Agent(
    edge=getstream.Edge(),
    agent_user=User(name="Assistant", id="agent"),
    instructions="You are a helpful assistant.",
    llm=gemini.Realtime(fps=3),
)
```

For background reasoning and async tools, use Extended Thinking:

```python theme={null}
from vision_agents.plugins.gemini import LIVE_EXTENDED_THINKING_MODEL, Realtime

llm = Realtime(model=LIVE_EXTENDED_THINKING_MODEL)
```

`turn_complete` only ends a streaming chunk. Agent turn-complete events wait for `interaction_status=IDLE` (the deprecated `REQUIRES_ACTION` alias is treated as idle) so thinking and async tool calls can continue.

Use `agent.simple_response(text=...)` for a text instruction, or `await llm.send_client_content(..., turn_complete=True)` to inject structured turns. `turn_complete=True` interrupts ongoing generation.

## Parameters

| Name | Type | Default | Description |
| - | - | - | - |
| `model` | `str` | `"gemini-3.8-live"` | Live model. Use `LIVE_EXTENDED_THINKING_MODEL` (`"gemini-3.8-live-extended-thinking"`) for background reasoning. Pass `"gemini-3.5-live-translate-preview"` for Live Translate, or an older Live model such as `"gemini-3.1-flash-live-preview"`. |
| `fps` | `int` | `1` | Video frames per second forwarded to Gemini |
| `blocking` | `bool` | `False` | Use `BLOCKING` tool execution. Default is `NON_BLOCKING`. `blocking=True` is allowed only on `gemini-3.8-live`; Extended Thinking rejects it. |
| `thinking_level` | `ThinkingLevel` | `None` | Thinking level. Applied automatically as `HIGH` for Extended Thinking when `thinking_config` is omitted. |
| `config` | `LiveConnectConfigDict` | `None` | Optional config dict to customize session behavior |
| `input_audio_pacing` | `AudioInputPacingConfig` | `None` | Buffers irregular upstream PCM and forwards fixed-size chunks (default 20 ms) at a stable wall-clock rate. Auto-enabled with `AudioInputPacingConfig.virtual_microphone()` when `model` is `"gemini-3.5-live-translate-preview"`; pass `None` explicitly to opt out. |
| `api_key` | `str` | `None` | API key (defaults to `GOOGLE_API_KEY` or `GEMINI_API_KEY`) |

## Tools

Register functions on the Realtime LLM. Declarations default to `NON_BLOCKING`.

```python theme={null}
from vision_agents.plugins.gemini import LIVE_EXTENDED_THINKING_MODEL, Realtime

llm = Realtime(model=LIVE_EXTENDED_THINKING_MODEL)

@llm.register_function(description="Get the current weather for a city.")
async def get_weather(city: str) -> dict:
    return {"city": city, "condition": "Sunny", "temperature_celsius": 22}
```

## Voice Activity Detection

Built-in VAD defaults are optimized for low-latency conversations. Video is included in the turn by default (`TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO`). Override via `config`:

```python theme={null}
from google.genai.types import (
    AutomaticActivityDetectionDict,
    EndSensitivity,
    RealtimeInputConfigDict,
    StartSensitivity,
    TurnCoverage,
)

llm = gemini.Realtime(
    config={
        "realtime_input_config": RealtimeInputConfigDict(
            turn_coverage=TurnCoverage.TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO,
            automatic_activity_detection=AutomaticActivityDetectionDict(
                start_of_speech_sensitivity=StartSensitivity.START_SENSITIVITY_HIGH,
                end_of_speech_sensitivity=EndSensitivity.END_SENSITIVITY_HIGH,
                silence_duration_ms=250,
                prefix_padding_ms=50,
            ),
        ),
    },
)
```

| Name | Type | Default | Description |
| - | - | - | - |
| `turn_coverage` | `TurnCoverage` | `TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO` | How audio activity and video are included in a turn |
| `start_of_speech_sensitivity` | `StartSensitivity` | `START_SENSITIVITY_HIGH` | How quickly the model detects the start of speech |
| `end_of_speech_sensitivity` | `EndSensitivity` | `END_SENSITIVITY_HIGH` | How quickly the model detects the end of speech |
| `silence_duration_ms` | `int` | `250` | Milliseconds of silence before the model considers a turn end |
| `prefix_padding_ms` | `int` | `50` | Milliseconds of audio to include before detected speech start |

<Tip>
  Higher sensitivity values make the model react faster to speech starts and stops, which reduces latency but may increase false positives in noisy environments.
</Tip>

## VLM (Vision Language Model)

Use Gemini 3 vision models for multimodal interactions with video frames. The VLM buffers video frames, converts them to JPEG, and sends them alongside text prompts.

```python theme={null}
from vision_agents.core import Agent, User
from vision_agents.plugins import gemini, getstream, deepgram, elevenlabs

agent = Agent(
    edge=getstream.Edge(),
    agent_user=User(name="Vision Agent", id="vision-agent"),
    instructions="Describe what you see in one sentence.",
    llm=gemini.VLM(model="gemini-3-flash-preview"),
    stt=deepgram.STT(),
    tts=elevenlabs.TTS(),
)
```

| Name | Type | Default | Description |
| - | - | - | - |
| `model` | `str` | `"gemini-3-flash-preview"` | Gemini vision model |
| `fps` | `int` | `1` | Video frames per second to capture |
| `frame_buffer_seconds` | `int` | `10` | Seconds of video to buffer for model input |
| `thinking_level` | `ThinkingLevel` | `None` | Thinking level for enhanced reasoning |
| `media_resolution` | `MediaResolution` | `None` | Resolution for multimodal processing |
| `api_key` | `str` | `None` | API key (defaults to `GOOGLE_API_KEY` env var) |

## Next Steps

<CardGroup cols={2}>
  <Card title="Gemini LLM" icon="brain" href="/integrations/llm/gemini">
    LLM with built-in tools and RAG
  </Card>

  <Card title="Build a Video Agent" icon="video" href="/introduction/video-agents">
    Add video processing
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.