All skills
microsoft avatar

/azure-ai-contentunderstanding-py

@e19efc2
by microsoftmicrosoft/skills3.1k stars
351

Azure AI Content Understanding SDK for Python. Use for multimodal content extraction from documents, images, audio, and video. Triggers: "azure-ai-contentunderstanding", "ContentUnderstandingClient", "multimodal analysis", "document extraction", "video analysis", "audio transcription".

Use this Skill: https://skilld.dev/gh/microsoft/skills/azure-ai-contentunderstanding-py

This session only. Nothing lands on disk.

referencesnon-hero-scenarios.md

≈1.3k tokens on demand. Your agent reads this file only when SKILL.md points to it.

azure-ai-contentunderstanding-py non-hero scenarios

These scenarios are intentionally separate from hero flows in SKILL.md. They cover secondary/advanced patterns typically used after the primary end-to-end path is working.

Analyze Image

from azure.ai.contentunderstanding.models import AnalyzeInput

poller = client.begin_analyze(
    analyzer_id="prebuilt-imageSearch",
    inputs=[AnalyzeInput(url="https://example.com/image.jpg")]
)
result = poller.result()
content = result.contents[0]
print(content.markdown)

Analyze Video

from azure.ai.contentunderstanding.models import AnalyzeInput

poller = client.begin_analyze(
    analyzer_id="prebuilt-videoSearch",
    inputs=[AnalyzeInput(url="https://example.com/video.mp4")]
)

result = poller.result()

# Access video content (AudioVisualContent)
content = result.contents[0]

# Get transcript phrases with timing
for phrase in content.transcript_phrases:
    print(f"[{phrase.start_time} - {phrase.end_time}]: {phrase.text}")

# Get key frames (for video)
for frame in content.key_frames:
    print(f"Frame at {frame.time}: {frame.description}")

Analyze Audio

from azure.ai.contentunderstanding.models import AnalyzeInput

poller = client.begin_analyze(
    analyzer_id="prebuilt-audioSearch",
    inputs=[AnalyzeInput(url="https://example.com/audio.mp3")]
)

result = poller.result()

# Access audio transcript
content = result.contents[0]
for phrase in content.transcript_phrases:
    print(f"[{phrase.start_time}] {phrase.text}")

Custom Analyzers

Create custom analyzers with field schemas for specialized extraction:

from azure.ai.contentunderstanding.models import (
    AnalyzeInput,
    ContentAnalyzer,
    ContentFieldDefinition,
    ContentFieldSchema,
)

# Create custom analyzer - returns an LRO poller; wait for provisioning to complete
poller = client.begin_create_analyzer(
    analyzer_id="my-invoice-analyzer",
    resource=ContentAnalyzer(
        description="Custom invoice analyzer",
        base_analyzer_id="prebuilt-documentSearch",
        field_schema=ContentFieldSchema(
            fields={
                "vendor_name": ContentFieldDefinition(type="string"),
                "invoice_total": ContentFieldDefinition(type="number"),
                "line_items": ContentFieldDefinition(
                    type="array",
                    item_definition=ContentFieldDefinition(
                        type="object",
                        properties={
                            "description": ContentFieldDefinition(type="string"),
                            "amount": ContentFieldDefinition(type="number"),
                        },
                    ),
                ),
            }
        ),
    ),
)
poller.result()  # wait until analyzer is ready

# Use custom analyzer
analyze_poller = client.begin_analyze(
    analyzer_id="my-invoice-analyzer",
    inputs=[AnalyzeInput(url="https://example.com/invoice.pdf")]
)

result = analyze_poller.result()

# Access extracted fields from analyzed content
content = result.contents[0]
print(content.fields["vendor_name"].value_string)
print(content.fields["invoice_total"].value_number)

Analyzer Management

# List all analyzers
analyzers = client.list_analyzers()
for analyzer in analyzers:
    print(f"{analyzer.analyzer_id}: {analyzer.description}")

# Get specific analyzer
analyzer = client.get_analyzer("prebuilt-documentSearch")

# Delete custom analyzer
client.delete_analyzer("my-custom-analyzer")

Async Client

import asyncio
import os
from azure.ai.contentunderstanding.aio import ContentUnderstandingClient
from azure.ai.contentunderstanding.models import AnalyzeInput
from azure.identity.aio import DefaultAzureCredential

async def analyze_document():
    endpoint = os.environ["CONTENTUNDERSTANDING_ENDPOINT"]
    async with DefaultAzureCredential() as credential:
        async with ContentUnderstandingClient(
            endpoint=endpoint,
            credential=credential
        ) as client:
            poller = await client.begin_analyze(
                analyzer_id="prebuilt-documentSearch",
                inputs=[AnalyzeInput(url="https://example.com/doc.pdf")]
            )
            result = await poller.result()
            content = result.contents[0]
            return content.markdown

asyncio.run(analyze_document())

Content Types

Class For Provides
DocumentContent PDF, images, Office docs Pages, tables, figures, paragraphs
AudioVisualContent Audio, video files Transcript phrases, timing, key frames

Both derive from AnalysisContent, which provides basic information and a markdown representation.

Model Imports

from azure.ai.contentunderstanding.models import (
    AnalyzeInput,
    AnalyzeResult,
    DocumentContent,
    AudioVisualContent,
)

Client Types

Client Purpose
ContentUnderstandingClient Sync client for all operations
ContentUnderstandingClient (aio) Async client for all operations

Source: SKILL.md on GitHub

2 warnings16d4 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    This skill provides standard integration for the Microsoft Azure AI Content Understanding SDK. It follows security best practices, such as using managed identities via DefaultAzureCredential and context managers for resource cleanup. While the skill processes external content, which is its primary purpose, it does so through official SDK methods.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: MEDIUM · 1 issue

  • Runlayer7mo

    2/2 files flagged

Signed by skilld at e19efc2. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 3 months ago
Other metadata
metadata
{
  "author": "Microsoft",
  "version": "1.0.0",
  "package": "azure-ai-contentunderstanding"
}

README badge

README badge for microsoft/skills/azure-ai-contentunderstanding-py