All skills
microsoft avatar

/microsoft-foundry

@04110d9
by microsoftmicrosoft/skills3.1k stars
351

Build, deploy, evaluate, optimize, fine-tune, and manage Microsoft Foundry agents, models, and resources end to end. USE FOR: foundry, azd ai agent, azd provision/deploy, hosted agent scaffold/develop/run/deploy/troubleshoot, prompt agent create, create agent, update agent, add tool to agent, invoke agent, agent.yaml, agent insights, pull agent insights, evaluate agent, batch eval, continuous eval, continuous monitoring, agent CI/CD, optimize prompt, improve prompt, prompt optimizer, optimize agent instructions, Agent Optimizer scaffold, dataset curation from traces, deploy model, model fine-tuning (SFT/DPO/RFT), Foundry project, RBAC, role assignment, permissions, quota, capacity, region, deployment failure, AI Services, create Foundry resource, knowledge index, customize deployment, onboard, availability, training-data, grader, distillation, large file upload. DO NOT USE FOR: Azure Functions, App Service, general Azure deploy (use azure-deploy), general Azure prep (use azure-prepare).

Use this Skill: https://skilld.dev/gh/microsoft/skills/microsoft-foundry

This session only. Nothing lands on disk.

finetuningreferencesvision-fine-tuning.md

≈1.1k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Vision Fine-Tuning

Fine-tune models with image data to customize visual understanding. Uses the same chat-completions JSONL format as text SFT, but with image content blocks in user messages.

Supported Models

Model Version
gpt-4o 2024-08-06
gpt-4.1 2025-04-14

Image Requirements

Constraint Limit
Max examples with images per training file 50,000
Max images per example 64
Max image file size 10 MB
Supported formats JPEG, PNG, WEBP
Color mode RGB or RGBA
Min examples 10

Important: Images can only appear in user messages, never in assistant responses.

Data Format

Each training example follows the standard SFT messages format. Images are included as image_url content blocks within user messages.

{"messages": [{"role": "system", "content": "You are a helpful AI assistant that describes images."}, {"role": "user", "content": [{"type": "text", "text": "Describe this image."}, {"type": "image_url", "image_url": {"url": "https://example.com/photo.png", "detail": "high"}}]}, {"role": "assistant", "content": "The image shows a cityscape with tall buildings against a blue sky."}]}

Image Sources

Images can be provided in two ways:

1. Public URL:

{"type": "image_url", "image_url": {"url": "https://example.com/image.png"}}

2. Base64 data URI:

{"type": "image_url", "image_url": {"url": "data:image/png;base64,iVBORw0KGgo..."}}

Detail Control

The detail parameter controls image processing fidelity and cost:

Value Behavior Cost
low Downscales to 512×512 pixels Lower
high Full resolution processing Higher
auto Model decides based on image size Default
{"type": "image_url", "image_url": {"url": "https://example.com/image.png", "detail": "low"}}

Use low for tasks where fine visual detail doesn't matter (classification, general description). Use high for tasks needing precise detail (OCR, diagram reading, defect detection).

Content Moderation

Images are screened before training. The following are automatically excluded:

  • Images containing people or faces (face detection only — no identification)
  • CAPTCHAs
  • Content violating Azure usage policies

This screening may add latency to file upload validation.

Best Practices

  • Diverse examples: Vary image content, angles, lighting, and resolution
  • Consistent annotations: Keep assistant response style and detail level uniform
  • Start with detail: low: Cheaper and faster — upgrade to high only if results need it
  • Check for excluded images: After upload, verify the training count matches expectations — some images may be silently skipped due to content moderation
  • Mixed text+image: You can include both text-only and image examples in the same training file

Training Workflow

Vision fine-tuning follows the exact same workflow as text SFT:

  1. Prepare JSONL with image content blocks
  2. Upload training file (validation may take longer due to image screening)
  3. Create fine-tuning job with a supported vision model
  4. Monitor and evaluate as usual
# Upload (image validation may take longer)
train_file = client.files.create(purpose="fine-tune", file=open("vision_train.jsonl", "rb"))
client.files.wait_for_processing(train_file.id)

# Submit — same as text SFT
job = client.fine_tuning.jobs.create(
    model="gpt-4.1-2025-04-14",
    training_file=train_file.id,
    validation_file=val_file.id,
    method={"type": "supervised"}
)

Troubleshooting

Issue Resolution
Images skipped silently Check for people/faces, oversized files, unsupported formats
URL not accessible Ensure URLs are publicly accessible, or use base64 data URIs
Exceeds 10 MB Resize or compress the image
Wrong color mode Convert to RGB or RGBA
Low quality results Try detail: high, add more diverse examples, increase dataset size

Reference

Source: SKILL.md on GitHub

2 warnings3d4 checks · Risk SAFE
  • Gen Agent Trust Hub3d

    This skill provides a comprehensive environment for managing the end-to-end lifecycle of AI agents, models, and infrastructure on Microsoft Foundry. It includes sub-skills for deployment, evaluation, fine-tuning, and troubleshooting. The skill utilizes dynamic code execution and shell command wrappers, which are used within the context of local development and cloud orchestration. All external resources and dependencies originate from trusted organizations and well-known services.

  • Socket3d

    2 alerts: gptSecurity, gptAnomaly

  • Snyk3d

    Risk: LOW · No issues

  • Runlayer7mo

    36/36 files flagged

Signed by skilld at 04110d9. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 20 hours ago.

Activeupdated last week
metadata
{
  "author": "Microsoft",
  "version": "1.2.26"
}

README badge

README badge for microsoft/skills/microsoft-foundry