All skills
inference-shell avatar

/agent-browser

@b576e8e

Browser automation for AI agents via inference.sh. Navigate web pages, interact with elements using @e refs, take screenshots, record video. Capabilities: web scraping, form filling, clicking, typing, drag-drop, file upload, JavaScript execution. Use for: web automation, data extraction, testing, agent browsing, research. Triggers: browser, web automation, scrape, navigate, click, fill form, screenshot, browse web, playwright, headless browser, web agent, surf internet, record video

Use this Skill: https://skilld.dev/gh/inference-shell/skills/agent-browser

This session only. Nothing lands on disk.

referencescommands.md

≈1.6k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Command Reference

Complete reference for all agent-browser functions. For quick start, see SKILL.md.

Base Command

All commands follow this pattern:

belt app run agent-browser --function <function> --session <session_id|new> --input '<json>'
  • --function: Function to call (open, snapshot, interact, screenshot, execute, close)
  • --session: Session ID from previous call, or new to start fresh
  • --input: JSON input for the function

Functions

open

Navigate to URL and configure browser. This is the entry point for all sessions.

belt app run agent-browser --function open --session new --input '{
  "url": "https://example.com",
  "width": 1280,
  "height": 720,
  "user_agent": "Mozilla/5.0...",
  "record_video": false,
  "show_cursor": false,
  "proxy_url": null,
  "proxy_username": null,
  "proxy_password": null
}'

Input Fields:

Field Type Default Description
url string required URL to navigate to
width int 1280 Viewport width in pixels
height int 720 Viewport height in pixels
user_agent string null Custom user agent string
record_video bool false Record video (returned on close)
show_cursor bool false Show cursor indicator in screenshots/video
proxy_url string null Proxy server URL
proxy_username string null Proxy auth username
proxy_password string null Proxy auth password

Output:

{
  "session_id": "abc123",
  "url": "https://example.com",
  "title": "Example Domain",
  "elements": [...],
  "elements_text": "@e1 [a] \"More information...\" href=\"...\"\n...",
  "screenshot": "<File>"
}

snapshot

Re-fetch page state with @e refs. Call after navigation or DOM changes.

belt app run agent-browser --function snapshot --session $SESSION_ID --input '{}'

Output: Same as open (url, title, elements, elements_text, screenshot)

interact

Perform actions on the page using @e refs.

belt app run agent-browser --function interact --session $SESSION_ID --input '{
  "action": "click",
  "ref": "@e1"
}'

Input Fields:

Field Type Description
action string Action to perform (see Actions table)
ref string Element ref (e.g., @e1)
text string Text for fill/type/press/select
direction string Scroll direction: up, down, left, right
scroll_amount int Scroll pixels (default 400)
wait_ms int Wait duration in milliseconds
url string URL for goto action
target_ref string Target ref for drag action
file_paths array File paths for upload action

Actions:

Action Required Fields Description
click ref Single click
dblclick ref Double click
fill ref, text Clear input and type text
type text Type text without clearing
press text Press key (Enter, Tab, Escape, etc.)
select ref, text Select dropdown option by label
hover ref Hover over element
check ref Check checkbox
uncheck ref Uncheck checkbox
drag ref, target_ref Drag from ref to target_ref
upload ref, file_paths Upload files to file input
scroll direction Scroll page (optional: scroll_amount)
back - Go back in browser history
wait wait_ms Wait for specified milliseconds
goto url Navigate to different URL

Output:

{
  "success": true,
  "action": "click",
  "message": null,
  "screenshot": "<File>",
  "snapshot": {
    "url": "...",
    "title": "...",
    "elements": [...],
    "elements_text": "..."
  }
}

screenshot

Take a screenshot of the current page.

belt app run agent-browser --function screenshot --session $SESSION_ID --input '{
  "full_page": true
}'

Input Fields:

Field Type Default Description
full_page bool false Capture full scrollable page

Output:

{
  "screenshot": "<File>",
  "width": 1280,
  "height": 720
}

execute

Run JavaScript code on the page.

belt app run agent-browser --function execute --session $SESSION_ID --input '{
  "code": "document.title"
}'

Input Fields:

Field Type Description
code string JavaScript code to execute

Output:

{
  "result": "Example Domain",
  "error": null,
  "screenshot": "<File>"
}

Examples:

# Get page title
'{"code": "document.title"}'

# Count elements
'{"code": "document.querySelectorAll(\"a\").length"}'

# Extract text
'{"code": "document.querySelector(\"h1\").textContent"}'

# Get all links
'{"code": "Array.from(document.querySelectorAll(\"a\")).map(a => a.href)"}'

# Scroll to bottom
'{"code": "window.scrollTo(0, document.body.scrollHeight)"}'

# Get computed style
'{"code": "getComputedStyle(document.body).backgroundColor"}'

close

Close the browser session. Returns video if recording was enabled.

belt app run agent-browser --function close --session $SESSION_ID --input '{}'

Output:

{
  "success": true,
  "video": "<File or null>"
}

Key Combinations

For the press action, use these key names:

Key Name
Enter Enter
Tab Tab
Escape Escape
Backspace Backspace
Delete Delete
Arrow keys ArrowUp, ArrowDown, ArrowLeft, ArrowRight
Modifiers Control, Shift, Alt, Meta

Key combinations:

# Ctrl+A (select all)
'{"action": "press", "text": "Control+a"}'

# Ctrl+C (copy)
'{"action": "press", "text": "Control+c"}'

# Shift+Tab (focus previous)
'{"action": "press", "text": "Shift+Tab"}'

Error Handling

When an action fails, success is false and message contains the error:

{
  "success": false,
  "action": "click",
  "message": "Unknown ref: @e99. Run 'snapshot' to get current elements.",
  "screenshot": "<File>",
  "snapshot": {...}
}

Common errors:

  • Unknown ref: @eN - Ref doesn't exist, re-snapshot needed
  • 'text' required for fill action - Missing required field
  • 'target_ref' required for drag action - Missing drag target
  • Timeout 5000ms exceeded - Element not found or not clickable

Source: SKILL.md on GitHub

1 alert13d5 checks · Risk SAFE
  • Gen Agent Trust Hub13d

    The skill provides powerful browser automation capabilities including JavaScript execution and file uploads. While these are intended features, they introduce risks such as indirect prompt injection from malicious websites and the potential for session data exfiltration via scripts.

  • Socket13d

    2 alerts: gptAnomaly, gptSecurity

  • Snyk13d

    Risk: MEDIUM · 1 issue

  • Runlayer6mo

    7/10 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at b576e8e. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last week.

Activeupdated 5 months ago
What it can do
Runs commands
All 1 allowed tools
Bash(belt *)
  • browser-automation
  • playwright
  • web-scraping
  • form-filling
  • screenshot
  • video-recording
  • javascript-execution
  • inference-sh
  • headless-browser

README badge

README badge for inference-shell/skills/agent-browser

Browser automation using Playwright with inference.sh, controlled via the `belt` CLI. Navigate pages, interact with elements using @e refs, take screenshots, record video, and run JavaScript. Suitable for web scraping, form filling, clicking, typing, drag-drop, file upload, and automated testing workflows.

Generated from the current SKILL.md.

What browser engine does this use?
Playwright. The skill wraps Playwright under the hood with a simplified @e ref system for element interaction.
Do I need to install anything besides the skill?
Yes, you need the belt CLI installed. The skill documentation includes install instructions.
Can I record video of browser sessions?
Yes. Pass record_video: true when opening a session, and the video file is returned when you close the session.
How do I interact with page elements?
Use the @e refs returned by open and snapshot calls. Pass the ref to interact actions like click, fill, drag, upload, or execute JavaScript.
Do element refs stay valid after navigation?
No. Refs are invalidated after navigation, form submission, or dynamic content loading. Always call snapshot after these events to get fresh refs.

Generated from the current SKILL.md. These answers refresh after source changes.