All skills
oaustegard avatar

/invoking-gemini

@4ac0233

Invokes Google Gemini models for structured outputs, image generation, text-to-speech narration, multi-modal tasks, and Google-specific features. Use when users request Gemini, image generation, Gemini TTS or a synthesized voice, structured JSON output, Google API integration, or cost-effective parallel processing.

Use this Skill: https://skilld.dev/gh/oaustegard/claude-skills/invoking-gemini

This session only. Nothing lands on disk.

CHANGELOG.md

≈2.8k tokens on demand. Your agent reads this file only when SKILL.md points to it.

invoking-gemini - Changelog

2026-09-24

Added — speech generation (Gemini 3.8 Flash TTS, GA 2026-09-23)

  • generate_speech() writes a WAV from text with a prebuilt or designed voice and an optional style; design_voice() creates a stored voice from a description; list_voices() pages the 2,089-voice library. Run through the CF gateway on 2026-09-24, they returned a 7.8 s WAV, a stored voice id with its sample, and 2,092 voices (the library plus this project's designs).
  • SPEECH_MODELS / SPEECH_ALIASES (tts, tts-lite) are kept apart from MODEL_ALIASES, so invoke_gemini() cannot resolve to an audio model.
  • The calls use the Interactions API. On generateContent a "Style:" prefix is spoken and systemInstruction is refused.
  • _rest_request() sends a direct-mode API key in the x-goog-api-key header rather than the URL.

2026-09-03

⚠️ BREAKING — flash alias and DEFAULT_MODEL repointed to gemini-3.8-flash

  • Gemini 3.8 Flash reached GA on 2026-09-02. It is now the registry default and the flash alias. Google also shipped 3.7 Flash on 2026-08-13, which this skill missed; both are added to MODELS.
  • New pinned aliases flash-3.7 and flash-3.6. flash-3.5 and flash-3 are unchanged. Nothing is removed; no 3.x Flash has a shutdown date.
  • Price is the same as 3.6 on Google's current page ($0.75 in / $3.75 out through 2026-12-31, then $1.50 / $7.50), so the swap costs nothing per token. Google says 3.8 "works harder" at higher effort levels, so per-task thinking tokens may go up.

⚠️ BREAKING — pro alias repointed to gemini-3.8-flash; Pro tier off routing

  • Oskar, 2026-09-03: "We should never use 3.1 Pro, its Pareto efficiency is too poor compared to the later Flash models." gemini-3.1-pro-preview costs 2.7× / 3.2× (in / out) what 3.8 Flash does at today's rates and the 3.5+ Flash line already beat it on coding and agentic benchmarks. The ID stays in MODELS for pinned callers; the table marks it DEPRECATED. "Maximum reasoning" is now Flash with thinking_level='high'. A future 3.5 Pro gets the same test before any alias points at it.

Changed — thinking_level='minimal' on 3.7 / 3.8 Flash

  • Both models return HTTP 400 for minimal (verified live 2026-09-03; 3.6, 3.5 and 3.5-lite accept it). invoke_gemini() now downgrades minimal to low on those two model IDs and prints a one-line note to stderr, so callers that pass minimal uniformly (transcription, classification) keep working on the new default instead of burning a non-retriable 400 and returning None. Measured on 3.8: low spent 0 thinking tokens on a one-word reply, the default medium spent 79. On 3.7, low still spent 45–88 on the same prompt, so the downgrade is not free there — budget output generously.
  • For a true no-thinking pass, pin flash-3.6 or lite.

Fixed — stale figures in the model tables

  • gemini-3.1-pro-preview output above 200K is $18.00 / 1M, not $24.00.
  • gemini-3.6-flash pricing carried the flat $1.50 / $7.50 from its launch table; Google's pricing page now lists it on the same introductory schedule as 3.7 and 3.8.
  • 3.5 Pro was still described as "slated for June 2026"; it has not shipped as of 2026-09-03.
  • Deprecation table gains the 3.x preview rows and the 2026-10-02 shutdown of gemini-2.5-flash-image (the nano-banana alias target).
  • Two docstrings still named gemini-3-flash-preview as the default.

2026-07-29

Fixed — nested Pydantic models raised a bare HTTP 400

  • _pydantic_to_schema() returned pydantic's model_json_schema() nearly verbatim, which emits $defs + $ref for every nested model. Gemini's responseSchema does not support $ref/$defs, so any model containing another model (list[Finding], the common case) produced a 400 with no usable body, burned three retries, and returned None. Flat single-level models worked, which is why this went unnoticed. Refs are now inlined and the keywords Gemini rejects are stripped recursively rather than only at the top level (title, default, additionalProperties, discriminator, examples, const).
  • Optional[X] / anyOf is now translated to the non-null branch plus nullable: true instead of being passed through as anyOf, which Gemini also rejects.
  • 4xx responses now raise _NonRetriableAPIError carrying the response body. raise_for_status() discarded it, and Gemini puts the only useful diagnostic there — which schema keyword it refused. 4xx is deterministic, so it no longer burns the retry budget either.
  • invoke_with_structured_output() gained max_output_tokens (default 32768). Thinking tokens count against the output budget, so a small cap truncated the JSON mid-object and surfaced as a pydantic EOF while parsing that reads like a schema error. finishReason=MAX_TOKENS is now detected and reported as truncation.

2026-07-21

Added — audio (and video) input

  • image_path now accepts any supported media file, not just images. Audio input verified working 2026-07-21 (3-beep WAV; gemini-3.6-flash returned the correct count and pitch direction). The param keeps its legacy name for back compat; docstrings now state what it really accepts.
  • Explicit mimeType overrides for audio/video extensions mimetypes guesses wrongly or misses (.m4a/.aac/.flac/.ogg/.opus/.mp3/.wav/.aiff/.mp4/.mov/.webm/ .heic/.heif). A wrong guess previously sent a bad mimeType and produced a confused answer rather than an error.
  • New MediaInputError (subclass of ValueError) for deterministic input problems; retry loops re-raise it immediately instead of burning 3 attempts.
  • 15MB inline cap enforced with an actionable message. Larger files need the Files API, which this client still does not implement.
  • The direct google-generativeai SDK fallback now rejects non-image media with a clear message instead of an opaque PIL UnidentifiedImageError. Audio/video require the CF Gateway path.

Routing note: send audio to gemini-3.6-flash. gemini-3.5-flash-lite was unreliable on the same clip — it reported three beeps then two on identical input, and got the pitch direction wrong both times.

⚠️ BREAKING — lite alias repointed

  • MODEL_ALIASES['lite']: gemini-2.5-flash-lite → gemini-3.5-flash-lite. Output cost goes from $0.40 to $2.50 per 1M (~6x) for any caller using model="lite". Pin gemini-2.5-flash-lite by full ID if you need the old rate, though it is deprecated (below).

  • The Gemini 2.5 text generation is retired from routing: gemini-2.5-flash, gemini-2.5-flash-lite, gemini-2.5-pro. Model IDs stay callable and the stable-flash / stable-pro aliases still resolve, so nothing hard-breaks, but they are no longer recommended targets. Image model nano-banana (gemini-2.5-flash-image) is NOT affected.

  • Added gemini-3.5-flash-lite (GA 2026-07-21) as the cheap/bulk tier.

  • Gemini 3.6 Flash (gemini-3.6-flash) reached GA (2026-07-21). Added it to the model registry and made it the new DEFAULT_MODEL.

  • Repointed the flash alias from gemini-3.5-flash to gemini-3.6-flash. Added flash-3.5 as a stable handle for the prior frontier Flash; flash-3 still pins the older gemini-3-flash-preview.

  • Rationale: 3.6 Flash is ~half Sonnet's cost (in $1.50 / out $7.50 vs ~$3 / ~$15) and improves coding/agentic quality with ~17% fewer output tokens than 3.5 Flash — making it the default for sub-agent delegation.

  • Updated SKILL.md and references/models.md tables (pricing, 64K output cap, benchmark deltas, tone-regression caveat).

  • Not touched: gemini-3.5-flash-lite / gemini-3.5-flash-cyber (shipped same day) are not yet aliased; two helper fns still hardcode a gemini-3-flash-preview default in their signatures (pre-existing).

2026-05-28

  • Nano Banana 2 (gemini-3.1-flash-image-preview) and Nano Banana Pro (gemini-3-pro-image-preview) reached GA on Vertex / Gemini Enterprise.
  • Kept the -preview model IDs: the Gemini Developer API surface this client uses still serves both under -preview (GA IDs without the suffix are Vertex-only and 404 here). Verified against the live image-generation docs.
  • Documented capabilities: 512/1K/2K GA + 4K preview, up to 14 reference images, Search + Image-Search grounding (3.1 Flash), thinking_level control.
  • Noted video-as-input is a Vertex preview only; not available on the Developer API.

All notable changes to the invoking-gemini skill are documented in this file. The format is based on Keep a Changelog.

[0.9.0] - 2026-09-24

Other

  • invoking-gemini 0.9.0: add Gemini 3.8 Flash TTS (speech generation)

[0.8.0] - 2026-09-03

Other

  • invoking-gemini: Gemini 3.8 Flash is the default, add 3.7, retire Pro from routing, handle minimal-thinking 400 (#783)

[0.7.0] - 2026-07-23

Other

  • invoking-gemini: default to gemini-3.6-flash, retire Gemini 2.5 text models (#741)
  • invoking-gemini: Nano Banana 2/Pro GA — keep -preview IDs on Developer API surface (#676)

[0.6.0] - 2026-05-23

Added

  • add Gemini 3.5 Flash + thinking_level, fix stale model docs (#669)

Fixed

  • retry on egress-proxy 503 ('DNS cache overflow') in remembering + invoking-gemini (#580)

Other

  • Remove _MAP.md files, direct agents to tree-sitting for code navigation (#545)

[0.5.0] - 2026-03-31

Added

  • surface Image Generation, add examples (#520)
  • add mapping-features skill for behavioral web app documentation (#432)

Other

  • Regenerate _MAP.md files after @lat: backlink insertion (#504)
  • Lattice v2: bidirectional source-anchored knowledge graph (#503)

[0.3.1] - 2026-03-01

Fixed

  • use camelCase keys for Gemini REST API inline data

[0.3.0] - 2026-03-01

Added

  • add image generation support + fix IMAGE_MODELS registry

[0.3.0] - 2026-03-01

Added

  • generate_image() function for native image generation via Gemini image models
  • image and image-pro model aliases for image generation
  • nano-banana (gemini-2.5-flash-image) to IMAGE_MODELS registry
  • Image generation documentation in SKILL.md with prompt patterns and examples

Fixed

  • IMAGE_MODELS registry now maps display names to actual API model IDs (was mapping names to themselves, causing 404 errors)
  • nano-banana-2 → gemini-3.1-flash-image-preview
  • nano-banana-pro → gemini-3-pro-image-preview
  • nano-banana → gemini-2.5-flash-image

[0.2.0] - 2026-03-01

Added

  • update invoking-gemini model registry to current Gemini lineup

[0.1.0] - 2026-03-01

Added

  • route invoking-gemini through Cloudflare AI Gateway
  • add line numbers, markdown ToC, and other files listing
  • add code maps and CLAUDE.md integration guidance
  • Delete VERSION files, complete migration to frontmatter
  • Migrate all 27 skills from VERSION files to frontmatter

Changed

  • migrate API credential management to project knowledge files

Fixed

  • limit markdown ToC to h1/h2 headings only

Source: SKILL.md on GitHub

1 alert4mo4 checks · Risk SAFE
  • Gen Agent Trust Hub4mo

    The invoking-gemini skill is a professional implementation of a Google Gemini API client with support for Cloudflare AI Gateway. It handles credentials securely via local configuration files and uses trusted infrastructure. No security vulnerabilities or malicious patterns were found.

  • Socket4mo

    No alerts

  • Snyk4mo

    Risk: LOW · No issues

  • Runlayer7mo

    4/10 files flagged

Signed by skilld at 4ac0233. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated last week
metadata
{
  "version": "0.9.0"
}

README badge

README badge for oaustegard/claude-skills/invoking-gemini