All skills
dpearson2699 avatar

/speech-recognition

@cf3fe87

Transcribe speech to text using Apple's Speech framework. Use when implementing live microphone transcription with AVAudioEngine, recognizing recorded audio files, handling speech and microphone authorization, choosing on-device vs server-backed SFSpeechRecognizer behavior, or adopting SpeechAnalyzer, SpeechTranscriber, DictationTranscriber, AssetInventory, and async result streams on iOS 26+.

Use this Skill: https://skilld.dev/gh/dpearson2699/swift-ios-skills/speech-recognition

This session only. Nothing lands on disk.

referencesspeechanalyzer-patterns.md

≈1.5k tokens on demand. Your agent reads this file only when SKILL.md points to it.

SpeechAnalyzer Patterns

Use this reference when implementing iOS 26+ speech-to-text with SpeechAnalyzer, SpeechTranscriber, DictationTranscriber, SpeechDetector, or AssetInventory.

Contents

Choosing Modules

Prefer SpeechTranscriber for the newer general-purpose on-device model. Before building UI around it, check both device and locale support:

guard SpeechTranscriber.isAvailable,
      let locale = SpeechTranscriber.supportedLocale(equivalentTo: Locale.current)
else {
    // Disable the feature or try DictationTranscriber for compatible devices/locales.
    return
}

Use documented presets only:

  • .transcription for basic accurate transcription.
  • .transcriptionWithAlternatives for editing suggestions.
  • .timeIndexedTranscriptionWithAlternatives for audio-time metadata plus alternatives.
  • .progressiveTranscription for low-latency live UI updates.
  • .timeIndexedProgressiveTranscription for live UI updates with time ranges.

Use DictationTranscriber when SpeechTranscriber is unavailable and the app can accept dictation-model behavior. Add SpeechDetector only with a transcriber module, and only when voice activity detection is worth the risk of dropping speech-like audio.

Preparing Assets

SpeechAnalyzer modules require model assets. The system installs and shares them outside the app bundle, but the app must request installation for the module configuration it plans to use.

let transcriber = SpeechTranscriber(locale: locale, preset: .transcription)

if let request = try await AssetInventory.assetInstallationRequest(
    supporting: [transcriber]
) {
    try await request.downloadAndInstall()
}

For language pickers, use installedLocales, supportedLocales, and AssetInventory.status(forModules:) to distinguish installed, downloadable, and unsupported choices. The app has a limited number of locale reservations; release unused reservations with AssetInventory.release(reservedLocale:).

Transcribing Files

For files, let the analyzer convert the file to a compatible format and finish the session after the file is consumed.

func transcribeFile(at url: URL, locale: Locale) async throws -> AttributedString {
    guard let supportedLocale = SpeechTranscriber.supportedLocale(equivalentTo: locale) else {
        throw SpeechError.unsupportedLocale
    }

    let transcriber = SpeechTranscriber(
        locale: supportedLocale,
        preset: .transcription
    )

    if let request = try await AssetInventory.assetInstallationRequest(
        supporting: [transcriber]
    ) {
        try await request.downloadAndInstall()
    }

    let analyzer = SpeechAnalyzer(modules: [transcriber])
    async let transcript = transcriber.results.reduce(into: AttributedString()) {
        text, result in
        text.append(result.text)
    }

    let file = try AVAudioFile(forReading: url)
    let lastSampleTime = try await analyzer.analyzeSequence(from: file)
    if let lastSampleTime {
        try await analyzer.finalizeAndFinish(through: lastSampleTime)
    } else {
        try analyzer.cancelAndFinishNow()
    }

    return try await transcript
}

Live Audio

For live audio, create an AsyncStream<AnalyzerInput>, convert microphone buffers to the analyzer-compatible format, yield them, and consume results in a separate task.

let transcriber = SpeechTranscriber(
    locale: locale,
    preset: .timeIndexedProgressiveTranscription
)
let analyzer = SpeechAnalyzer(modules: [transcriber])
let audioFormat = await SpeechAnalyzer.bestAvailableAudioFormat(
    compatibleWith: [transcriber]
)
let (inputSequence, inputBuilder) = AsyncStream.makeStream(of: AnalyzerInput.self)

// In the audio-engine tap, convert each AVAudioPCMBuffer to audioFormat first.
inputBuilder.yield(AnalyzerInput(buffer: convertedBuffer))

Use AVAudioConverter or an existing project audio pipeline for the conversion. Do not feed arbitrary input-node formats directly unless they already match a compatible analyzer format.

Handling Results

SpeechTranscriber.Result.text is an AttributedString. Time-indexed presets include audio time range attributes that can drive playback highlighting.

When using progressive presets, volatile results may be replaced by later final results. Keep volatile display state separate so the UI does not duplicate text.

for try await result in transcriber.results {
    if result.isFinal {
        volatileTranscript = AttributedString()
        finalizedTranscript.append(result.text)
    } else {
        volatileTranscript = result.text
    }
}

Finishing Sessions

The analyzer can only analyze one input sequence at a time. Ending your stream does not finish the analyzer session; call a finish or cancel method.

Use:

  • finalizeAndFinish(through:) after analyzeSequence(_:) returns a final sample time.
  • finalizeAndFinishThroughEndOfInput() after autonomous start(inputSequence:).
  • cancelAndFinishNow() for immediate cancellation.

After the session finishes, result streams terminate and most analyzer methods no longer accept new work. Create a new analyzer for a new finished session.

References

Source: SKILL.md on GitHub

No alerts16d5 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    The skill provides markdown documentation and Swift code snippets for implementing speech-to-text functionality using Apple's Speech framework. No security issues or malicious behaviors were detected.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: LOW · No issues

  • Runlayer6mo

    1 file scanned · No issues

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at cf3fe87. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 days ago.

Steadyupdated 3 months ago
  • speech-recognition
  • ios
  • swift
  • avfoundation
  • audio
  • microphone
  • transcription
  • asyncawait
  • authorization

README badge

README badge for dpearson2699/swift-ios-skills/speech-recognition

Transcribe live microphone and pre-recorded audio to text using Apple's Speech framework, supporting both the legacy SFSpeechRecognizer (iOS 10+) and the modern SpeechAnalyzer actor-based API (iOS 26+) with async/await. Covers authorization flows, on-device vs server recognition, partial and final results, and handling audio engine setup with AVAudioEngine.

Generated from the current SKILL.md.

Does this skill cover iOS 26 SpeechAnalyzer as well as the older SFSpeechRecognizer?
Yes. The skill covers both the new SpeechAnalyzer API (iOS 26+) with async/await and AsyncSequence, and the older SFSpeechRecognizer (iOS 10+) with callback-based recognition.
Can I use this skill for live microphone transcription?
Yes. The skill includes patterns for live microphone transcription using AVAudioEngine with SFSpeechAudioBufferRecognitionRequest, plus newer async/await approaches with SpeechAnalyzer.
Does this skill support on-device recognition?
Yes. The skill covers on-device recognition (iOS 13+) via the `requiresOnDeviceRecognition` flag for SFSpeechRecognizer and asset-based setup for SpeechAnalyzer, though not all locales are supported.
What permissions are required for live transcription?
Both speech recognition and microphone permissions are required. The skill shows how to request both and add the necessary Info.plist keys.
Can this skill transcribe pre-recorded audio files?
Yes. The skill includes examples for transcribing audio files using SFSpeechURLRecognitionRequest and SpeechAnalyzer's file-based API.

Generated from the current SKILL.md. These answers refresh after source changes.