Topics
- Python
- gemini
- computer-use
- playwright
- browser-automation
- agent-loop
- web-automation
- google-genai
What it does
Builds and runs agentic loops for Gemini 2.5 Computer Use, which controls a browser via Playwright to complete web automation tasks. The agent captures screenshots, sends them to the model, executes returned actions (click, type, scroll), and loops until the task completes or a turn limit is reached. Includes safety confirmation prompts for risky UI actions.
Generated from the current SKILL.md.
Frequently asked
What model does this skill use?
Gemini 2.5 with Computer Use capabilities. The skill sends screenshots and receives function calls to control the browser.
Does this work with browsers other than Chromium?
The default is Playwright's bundled Chromium, but you can specify Chrome or Edge via COMPUTER_USE_BROWSER_CHANNEL, or any Chromium-based executable (e.g. Brave) via COMPUTER_USE_BROWSER_EXECUTABLE.
How do I prevent the agent from taking risky actions?
Use the --exclude flag to block specific actions you don't want the model to execute, and run the browser in a sandboxed profile or container for isolation.
What happens when the agent encounters a risky action?
If the model marks an action with safety_decision as require_confirmation, the skill prompts you to approve it before executing.
How many turns can the agent run?
You control the turn limit via the --turn-limit flag (default appears to be 6). Each turn is one screenshot, model response, and action execution cycle.
Generated from the current SKILL.md. These answers refresh after source changes.