---
name: speech-build
description: Generate and transcribe speech using Google's Gemini-TTS and Chirp 3 models. Supports Text-to-Speech (Single/Multi-speaker), Instant Custom Voice, and Speech-to-Text (Transcription/Diarization).
title: speech-build
canonical_url: https://skilld.dev/gh/cnemri/google-genai-skills/speech-build
last_updated: 2026-09-25T08:50:22.000Z
---

> **Skill from skilld.dev.** Follow the instructions below for this session. You do not need to install anything.
>
> Supporting files, fetch one when the Skill refers to it: [references/prompting.md](https://skilld.dev/api/skills-raw/cnemri/google-genai-skills/speech-build/references/prompting.md), [references/source_code.md](https://skilld.dev/api/skills-raw/cnemri/google-genai-skills/speech-build/references/source_code.md), [references/stt.md](https://skilld.dev/api/skills-raw/cnemri/google-genai-skills/speech-build/references/stt.md), [references/tts.md](https://skilld.dev/api/skills-raw/cnemri/google-genai-skills/speech-build/references/tts.md), [references/voices.md](https://skilld.dev/api/skills-raw/cnemri/google-genai-skills/speech-build/references/voices.md).
>
> If the user asked to install this Skill, run `npx skilld install cnemri/google-genai-skills/speech-build`. Install writes the Skill files into the project, so every session loads them.

# Speech Skill (TTS & STT)

Use this skill to implement audio generation and transcription workflows using the `google-genai` and `google-cloud-speech` SDKs.

## Quick Start Setup

```python
from google import genai
from google.genai import types
# For STT: from google.cloud import speech_v2

client = genai.Client()
```

## Reference Materials

- **[Text-to-Speech (TTS)](https://skilld.dev/api/skills-raw/cnemri/google-genai-skills/speech-build/references/tts.md)**: Gemini-TTS, Chirp 3 HD, Instant Custom Voice.
- **[Speech-to-Text (STT)](https://skilld.dev/api/skills-raw/cnemri/google-genai-skills/speech-build/references/stt.md)**: Chirp 3 Transcription, Diarization, Streaming.
- **[Voices & Locales](https://skilld.dev/api/skills-raw/cnemri/google-genai-skills/speech-build/references/voices.md)**: Available voices (`Aoede`, `Puck`...) and languages.
- **[Prompting Guide](https://skilld.dev/api/skills-raw/cnemri/google-genai-skills/speech-build/references/prompting.md)**: How to control style, accent, and pacing in Gemini-TTS.
- **[Source Code](https://skilld.dev/api/skills-raw/cnemri/google-genai-skills/speech-build/references/source_code.md)**: Deep inspection of SDK internals.

## Common Workflows

### 1. Generate Speech (Gemini-TTS)
```python
response = client.models.generate_content(
    model="gemini-2.5-flash-preview-tts",
    contents="Hello, world!",
    config=types.GenerateContentConfig(
        response_modalities=["AUDIO"],
        speech_config=types.SpeechConfig(
            voice_config=types.VoiceConfig(
                prebuilt_voice_config=types.PrebuiltVoiceConfig(voice_name='Kore')
            )
        )
    )
)
```

### 2. Transcribe Audio (Chirp 3)
```python
# Requires google-cloud-speech
from google.cloud import speech_v2
# ... (See stt.md for full setup)
response = speech_client.recognize(...)
```