All skills
google avatar

/gemini-api

@becc4b8
by googlegoogle/skills21k stars
1,698

Use when the user asks about using Gemini in an enterprise environment or explicitly mentions Vertex AI, Google Cloud, or Agent Platform. Guides the usage of the Gemini API on Agent Platform with the Google Gen AI SDK. Covers SDK usage (Python, JS/TS, Go, Java, C#), capabilities like multimodal inputs, tools, media generation, caching, batch prediction, and Live API.

Use this Skill: https://skilld.dev/gh/google/skills/gemini-api

This session only. Nothing lands on disk.

referencesbounding_box.md

≈466 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Bounding Box Detection

Detect and localize objects within images or videos using bounding boxes. The model returns coordinates in the format [y_min, x_min, y_max, x_max], normalized from 0 to 1000.

Implementation (Python)

To ensure structured output, define a BoundingBox class and provide it as the response_schema.

from google import genai
from google.genai.types import (
    GenerateContentConfig,
    Part,
)
from pydantic import BaseModel


# Define the schema for the bounding box
class BoundingBox(BaseModel):
    box_2d: list[int]
    label: str


client = genai.Client()

config = GenerateContentConfig(
    system_instruction="""
    Return bounding boxes as an array with labels.
    Never return masks. Limit to 25 objects.
    """,
    response_mime_type="application/json",
    response_schema=list[BoundingBox],
)

image_uri = "gs://cloud-samples-data/generative-ai/image/socks.jpg"

response = client.models.generate_content(
    model="gemini-3.8-flash",
    contents=[
        Part.from_uri(file_uri=image_uri, mime_type="image/jpeg"),
        "Detect the socks in the image and provide bounding boxes.",
    ],
    config=config,
)

# Access the detected boxes
for bbox in response.parsed:
    print(f"Label: {bbox.label}, Box: {bbox.box_2d}")

Coordinate System

  • Format: [y_min, x_min, y_max, x_max]
  • Normalization: Coordinates are integers from 0 to 1000.
  • Origin: [0, 0] is the top-left corner of the image.

Visualization Helper

To visualize the results, scale the normalized coordinates back to the original image dimensions.

def scale_box(box_2d, width, height):
    y_min, x_min, y_max, x_max = box_2d
    return [
        int(y_min / 1000 * height),
        int(x_min / 1000 * width),
        int(y_max / 1000 * height),
        int(x_max / 1000 * width),
    ]

Source: SKILL.md on GitHub

1 warning9d3 checks · Risk SAFE
  • Gen Agent Trust Hub9d

    This skill provides a detailed integration guide for the Gemini API. It includes security considerations such as prompts designed to override the agent's internal knowledge of model versions, instructional safety filter examples that use provocative content, and the demonstration of dynamic code execution and external tool integration via MCP. These elements are presented as educational samples but should be reviewed when implementing in production.

  • Socket9d

    No alerts

  • Snyk9d

    Risk: MEDIUM · 1 issue

Signed by skilld at becc4b8. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated last week
metadata
{
  "version": "1.0.0",
  "category": "AiAndMachineLearning"
}
compatibility
Requires active Google Cloud credentials and Agent Platform API enabled.

README badge

README badge for google/skills/gemini-api