All skills
cnemri avatar

/speech-build

@ce49e83

Generate and transcribe speech using Google's Gemini-TTS and Chirp 3 models. Supports Text-to-Speech (Single/Multi-speaker), Instant Custom Voice, and Speech-to-Text (Transcription/Diarization).

Use this Skill: https://skilld.dev/gh/cnemri/google-genai-skills/speech-build

This session only. Nothing lands on disk.

referencesstt.md

≈436 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Speech-to-Text (STT) - Chirp 3

Chirp 3 offers state-of-the-art multilingual transcription and speaker diarization.

Transcribe Audio (Synchronous)

For audio < 1 minute.

from google.cloud import speech_v2
from google.cloud.speech_v2.types import cloud_speech

client = speech_v2.SpeechClient(...)

config = cloud_speech.RecognitionConfig(
    auto_decoding_config=cloud_speech.AutoDetectDecodingConfig(),
    model="chirp_3",
    language_codes=["auto"], # Language identification
)

request = cloud_speech.RecognizeRequest(
    recognizer="projects/.../locations/.../recognizers/_",
    config=config,
    content=audio_bytes, # or uri="gs://..."
)

response = client.recognize(request=request)

Batch Recognition

For long audio files.

request = cloud_speech.BatchRecognizeRequest(
    recognizer=recognizer,
    config=config,
    files=[cloud_speech.BatchRecognizeFileMetadata(uri="gs://...")],
    recognition_output_config=cloud_speech.RecognitionOutputConfig(
        gcs_output_config=cloud_speech.GcsOutputConfig(uri="gs://output-bucket")
    ),
)
operation = client.batch_recognize(request=request)
result = operation.result()

Speaker Diarization

Identify different speakers.

config = cloud_speech.RecognitionConfig(
    features=cloud_speech.RecognitionFeatures(
        diarization_config=cloud_speech.SpeakerDiarizationConfig(),
    ),
    model="chirp_3",
    # ...
)

Streaming STT

Real-time transcription.

# Create generator yielding StreamingRecognizeRequest
requests = create_streaming_requests(audio_file)
responses = client.streaming_recognize(requests=requests)
for response in responses:
    print(response.results[0].alternatives[0].transcript)

Source: SKILL.md on GitHub

2 warnings6mo4 checks · Risk SAFE
  • Gen Agent Trust Hub7mo

    The skill provides documentation and examples for using official Google Speech SDKs. It references trusted repositories and follows standard API patterns. No malicious behavior was detected.

  • Socket6mo

    No alerts

  • Snyk7mo

    Risk: MEDIUM · No issues

  • Runlayer7mo

    6/6 files flagged

Signed by skilld at ce49e83. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 months ago.

Dormantupdated 8 months ago

README badge

README badge for cnemri/google-genai-skills/speech-build