All skills
inference-shell avatar

/agent-browser

@b576e8e

Browser automation for AI agents via inference.sh. Navigate web pages, interact with elements using @e refs, take screenshots, record video. Capabilities: web scraping, form filling, clicking, typing, drag-drop, file upload, JavaScript execution. Use for: web automation, data extraction, testing, agent browsing, research. Triggers: browser, web automation, scrape, navigate, click, fill form, screenshot, browse web, playwright, headless browser, web agent, surf internet, record video

Use this Skill: https://skilld.dev/gh/inference-shell/skills/agent-browser

This session only. Nothing lands on disk.

referencessession-management.md

≈1.3k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Session Management

Browser sessions for state persistence and parallel browsing.

Related: authentication.md for login patterns, SKILL.md for quick start.

Contents

How Sessions Work

Each session maintains an isolated browser context with:

  • Cookies
  • LocalStorage / SessionStorage
  • Browser history
  • Page state
  • Video recording (if enabled)

Sessions persist across function calls, allowing multi-step workflows.

Starting a Session

Use --session new to create a fresh session:

RESULT=$(belt app run agent-browser --function open --session new --input '{
  "url": "https://example.com"
}')
SESSION_ID=$(echo $RESULT | jq -r '.session_id')
echo "Session: $SESSION_ID"

Using Session IDs

All subsequent calls use the session ID:

# Navigate
belt app run agent-browser --function open --session $SESSION_ID --input '{
  "url": "https://example.com/page2"
}'

# Interact
belt app run agent-browser --function interact --session $SESSION_ID --input '{
  "action": "click", "ref": "@e1"
}'

# Screenshot
belt app run agent-browser --function screenshot --session $SESSION_ID --input '{}'

# Close
belt app run agent-browser --function close --session $SESSION_ID --input '{}'

Session State

What Persists

Within a session, these persist across calls:

  • Cookies (login state, preferences)
  • LocalStorage and SessionStorage
  • IndexedDB data
  • Browser history (for back/forward)
  • Current page and DOM state
  • Video recording buffer

What Doesn't Persist

  • Sessions don't persist across server restarts
  • No automatic session recovery
  • Video is only available until close is called

Parallel Sessions

Run multiple independent sessions simultaneously:

#!/bin/bash
# Scrape multiple sites in parallel

# Start sessions
RESULT1=$(belt app run agent-browser --function open --session new --input '{
  "url": "https://site1.com"
}')
SESSION1=$(echo $RESULT1 | jq -r '.session_id')

RESULT2=$(belt app run agent-browser --function open --session new --input '{
  "url": "https://site2.com"
}')
SESSION2=$(echo $RESULT2 | jq -r '.session_id')

# Work with each session independently
belt app run agent-browser --function screenshot --session $SESSION1 --input '{}' &
belt app run agent-browser --function screenshot --session $SESSION2 --input '{}' &
wait

# Clean up both
belt app run agent-browser --function close --session $SESSION1 --input '{}'
belt app run agent-browser --function close --session $SESSION2 --input '{}'

Use Cases for Parallel Sessions

  1. A/B Testing - Compare different pages or user experiences
  2. Multi-site scraping - Gather data from multiple sources
  3. Load testing - Simulate multiple users
  4. Cross-region testing - Use different proxies per session

Session Cleanup

Always close sessions when done:

belt app run agent-browser --function close --session $SESSION_ID --input '{}'

Why close matters:

  • Releases server resources
  • Returns video recording (if enabled)
  • Prevents resource leaks

Error Handling

#!/bin/bash
set -e

cleanup() {
  belt app run agent-browser --function close --session $SESSION_ID --input '{}' 2>/dev/null || true
}
trap cleanup EXIT

SESSION_ID=$(belt app run agent-browser --function open --session new --input '{
  "url": "https://example.com"
}' | jq -r '.session_id')

# ... your automation ...
# cleanup runs automatically on exit

Best Practices

1. Store Session IDs

# Good: Store for reuse
SESSION_ID=$(... | jq -r '.session_id')
belt ... --session $SESSION_ID ...

# Bad: Parse every time
belt ... --session $(... | jq -r '.session_id') ...

2. Close Sessions Promptly

Don't leave sessions open longer than needed. Server resources are limited.

3. Use Meaningful Variable Names

# Good: Clear purpose
LOGIN_SESSION=$(...)
SCRAPE_SESSION=$(...)

# Bad: Generic names
S1=$(...)
S2=$(...)

4. Handle Session Expiry

Sessions may expire after extended inactivity:

# Check if session is still valid
RESULT=$(belt app run agent-browser --function snapshot --session $SESSION_ID --input '{}' 2>&1)
if echo "$RESULT" | grep -q "session not found"; then
  echo "Session expired, starting new one"
  SESSION_ID=$(belt app run agent-browser --function open --session new --input '{
    "url": "https://example.com"
  }' | jq -r '.session_id')
fi

5. One Task Per Session

For clarity, use one session per logical task:

# Good: Separate sessions for separate tasks
LOGIN_SESSION=$(...)  # Handle login
SCRAPE_SESSION=$(...)  # Handle scraping

# Okay for related tasks: One session for a workflow
SESSION=$(...)
# login -> navigate -> extract -> close

Source: SKILL.md on GitHub

1 alert13d5 checks · Risk SAFE
  • Gen Agent Trust Hub13d

    The skill provides powerful browser automation capabilities including JavaScript execution and file uploads. While these are intended features, they introduce risks such as indirect prompt injection from malicious websites and the potential for session data exfiltration via scripts.

  • Socket13d

    2 alerts: gptAnomaly, gptSecurity

  • Snyk13d

    Risk: MEDIUM · 1 issue

  • Runlayer6mo

    7/10 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at b576e8e. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last week.

Activeupdated 5 months ago
What it can do
Runs commands
All 1 allowed tools
Bash(belt *)
  • browser-automation
  • playwright
  • web-scraping
  • form-filling
  • screenshot
  • video-recording
  • javascript-execution
  • inference-sh
  • headless-browser

README badge

README badge for inference-shell/skills/agent-browser

Browser automation using Playwright with inference.sh, controlled via the `belt` CLI. Navigate pages, interact with elements using @e refs, take screenshots, record video, and run JavaScript. Suitable for web scraping, form filling, clicking, typing, drag-drop, file upload, and automated testing workflows.

Generated from the current SKILL.md.

What browser engine does this use?
Playwright. The skill wraps Playwright under the hood with a simplified @e ref system for element interaction.
Do I need to install anything besides the skill?
Yes, you need the belt CLI installed. The skill documentation includes install instructions.
Can I record video of browser sessions?
Yes. Pass record_video: true when opening a session, and the video file is returned when you close the session.
How do I interact with page elements?
Use the @e refs returned by open and snapshot calls. Pass the ref to interact actions like click, fill, drag, upload, or execute JavaScript.
Do element refs stay valid after navigation?
No. Refs are invalidated after navigation, form submission, or dynamic content loading. Always call snapshot after these events to get fresh refs.

Generated from the current SKILL.md. These answers refresh after source changes.