All skills
microsoft avatar

/podcast-generation

@a2003b6
by microsoftmicrosoft/skills3.1k stars
351

Generate AI-powered podcast-style audio narratives using Azure OpenAI's GPT Realtime Mini model via WebSocket. Use when building text-to-speech features, audio narrative generation, podcast creation from content, or integrating with Azure OpenAI Realtime API for real audio output. Covers full-stack implementation from React frontend to Python FastAPI backend with WebSocket streaming.

Use this Skill: https://skilld.dev/gh/microsoft/skills/podcast-generation

This session only. Nothing lands on disk.

referencesarchitecture.md

≈1.2k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Architecture Reference

Full-Stack Flow

┌─────────────────────────────────────────────────────────────────┐
│                       FRONTEND (React)                          │
│  AudioNarrative.jsx / PodcastOverview.jsx                       │
│    ↓ POST /api/v1/ai/audio                                     │
│  api.js → aiAPI.generateAudio(sourceType, sourceId, ...)       │
└────────────────────────────┬────────────────────────────────────┘
                             │
┌────────────────────────────▼────────────────────────────────────┐
│                       BACKEND (FastAPI)                         │
│  API: ai.py                                                     │
│    @router.post("/audio") → AudioNarrativeRequest               │
│    @router.get("/audio/{id}/stream") → WAV file                │
│                             │                                   │
│  Service: ai_service.py                                         │
│    generate_audio_narrative()                                   │
│    - Fetch content from DB (tag/bookmark/custom)               │
│    - Build styled prompt                                        │
│    - WebSocket connect to Azure Realtime API                   │
│    - Stream audio chunks + transcript                           │
│    - PCM → WAV conversion                                       │
│    - Save to database                                           │
│                             │                                   │
│  Model: database.py                                             │
│    AudioNarrative table                                         │
└────────────────────────────┬────────────────────────────────────┘
                             │ wss://
┌────────────────────────────▼────────────────────────────────────┐
│              Azure OpenAI Realtime API                          │
│  Model: gpt-realtime-mini                                       │
│  Output: PCM audio (24kHz, 16-bit mono) + transcript           │
└─────────────────────────────────────────────────────────────────┘

Database Schema

class AudioNarrative(Base):
    __tablename__ = "audio_narratives"
    
    id: int                          # Primary key
    source_type: str                 # "tag", "bookmark", "custom"
    source_id: Optional[int]         # Reference to source
    source_name: Optional[str]       # Display name
    title: str                       # Generated title
    script: str                      # Transcript text
    audio_url: Optional[str]         # Stream endpoint
    audio_data: Optional[str]        # Base64 WAV
    duration_seconds: Optional[int]  # Calculated from PCM length
    voice_name: str                  # "alloy", "echo", etc.
    created_at: datetime

Pydantic Schemas

class AudioNarrativeRequest(BaseModel):
    source_type: str          # Required: "tag", "bookmark", "custom"
    source_id: Optional[Union[int, str]] = None
    custom_query: Optional[str] = None
    voice_name: str = "alloy"
    style: str = "podcast"    # "podcast", "summary", "lecture"

class AudioNarrativeResponse(BaseModel):
    id: int
    title: str
    script: str
    audio_url: Optional[str]
    audio_data: Optional[str]  # Base64 WAV
    duration_seconds: Optional[int]
    voice_name: str
    created_at: datetime

Style Instructions

STYLE_INSTRUCTIONS = {
    "podcast": "Speak in a conversational, engaging podcast style with natural transitions. Use phrases like 'Let's dive into...' and 'What's fascinating here is...'",
    "summary": "Speak clearly and informatively, getting straight to the key points in a news anchor style.",
    "lecture": "Speak in an educational, thorough style suitable for learning. Explain concepts clearly like a professor."
}

Source: SKILL.md on GitHub

1 warning16d5 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    This skill provides a robust framework for generating AI-powered audio narratives using Azure OpenAI. It includes security considerations common to LLM-integrated applications, specifically regarding how external data is incorporated into prompts, which warrants standard review during implementation.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: LOW · No issues

  • Runlayer7mo

    5/5 files flagged

  • ZeroLeaks5mo

    2 findings · Score: 80/100

Signed by skilld at a2003b6. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 8 months ago

README badge

README badge for microsoft/skills/podcast-generation