All skills
am-will avatar

/gemini-computer-use

@6cd5afa
by am.willam-will/codex-skills1k stars
60

Build and run Gemini 2.5 Computer Use browser-control agents with Playwright. Use when a user wants to automate web browser tasks via the Gemini Computer Use model, needs an agent loop (screenshot → function_call → action → function_response), or asks to integrate safety confirmation for risky UI actions.

Use this Skill: https://skilld.dev/gh/am-will/codex-skills/gemini-computer-use

This session only. Nothing lands on disk.

SKILL.md

≈83 tokens always: the name and description. ≈430 when used: this file. ≈159 more on demand in 1 file.

Gemini Computer Use

Quick start

  1. Source the env file and set your API key:

    cp env.example env.sh
    $EDITOR env.sh
    source env.sh
  2. Create a virtual environment and install dependencies:

    python -m venv .venv
    source .venv/bin/activate
    pip install google-genai playwright
    playwright install chromium
  3. Run the agent script with a prompt:

    python scripts/computer_use_agent.py \
      --prompt "Find the latest blog post title on example.com" \
      --start-url "https://example.com" \
      --turn-limit 6

Browser selection

  • Default: Playwright's bundled Chromium (no env vars required).
  • Choose a channel (Chrome/Edge) with COMPUTER_USE_BROWSER_CHANNEL.
  • Use a custom Chromium-based executable (e.g., Brave) with COMPUTER_USE_BROWSER_EXECUTABLE.

If both are set, COMPUTER_USE_BROWSER_EXECUTABLE takes precedence.

Core workflow (agent loop)

  1. Capture a screenshot and send the user goal + screenshot to the model.
  2. Parse function_call actions in the response.
  3. Execute each action in Playwright.
  4. If a safety_decision is require_confirmation, prompt the user before executing.
  5. Send function_response objects containing the latest URL + screenshot.
  6. Repeat until the model returns only text (no actions) or you hit the turn limit.

Operational guidance

  • Run in a sandboxed browser profile or container.
  • Use --exclude to block risky actions you do not want the model to take.
  • Keep the viewport at 1440x900 unless you have a reason to change it.

Resources

  • Script: scripts/computer_use_agent.py
  • Reference notes: references/google-computer-use.md
  • Env template: env.example

Source: SKILL.md on GitHub

2 warnings13d5 checks · Risk SAFE
  • Gen Agent Trust Hub13d

    This skill implements a browser automation agent using Gemini Computer Use. It is vulnerable to indirect prompt injection from malicious content on websites visited by the agent.

  • Socket13d

    No alerts

  • Snyk13d

    Risk: MEDIUM · 1 issue

  • Runlayer7mo

    4/4 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 6cd5afa. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 months ago.

Steadyupdated 9 months ago
  • Python
  • gemini
  • computer-use
  • playwright
  • browser-automation
  • agent-loop
  • web-automation
  • google-genai

README badge

README badge for am-will/codex-skills/gemini-computer-use

Builds and runs agentic loops for Gemini 2.5 Computer Use, which controls a browser via Playwright to complete web automation tasks. The agent captures screenshots, sends them to the model, executes returned actions (click, type, scroll), and loops until the task completes or a turn limit is reached. Includes safety confirmation prompts for risky UI actions.

Generated from the current SKILL.md.

What model does this skill use?
Gemini 2.5 with Computer Use capabilities. The skill sends screenshots and receives function calls to control the browser.
Does this work with browsers other than Chromium?
The default is Playwright's bundled Chromium, but you can specify Chrome or Edge via COMPUTER_USE_BROWSER_CHANNEL, or any Chromium-based executable (e.g. Brave) via COMPUTER_USE_BROWSER_EXECUTABLE.
How do I prevent the agent from taking risky actions?
Use the --exclude flag to block specific actions you don't want the model to execute, and run the browser in a sandboxed profile or container for isolation.
What happens when the agent encounters a risky action?
If the model marks an action with safety_decision as require_confirmation, the skill prompts you to approve it before executing.
How many turns can the agent run?
You control the turn limit via the --turn-limit flag (default appears to be 6). Each turn is one screenshot, model response, and action execution cycle.

Generated from the current SKILL.md. These answers refresh after source changes.