All skills
inference-shell avatar

/agent-browser

@b576e8e

Browser automation for AI agents via inference.sh. Navigate web pages, interact with elements using @e refs, take screenshots, record video. Capabilities: web scraping, form filling, clicking, typing, drag-drop, file upload, JavaScript execution. Use for: web automation, data extraction, testing, agent browsing, research. Triggers: browser, web automation, scrape, navigate, click, fill form, screenshot, browse web, playwright, headless browser, web agent, surf internet, record video

Use this Skill: https://skilld.dev/gh/inference-shell/skills/agent-browser

This session only. Nothing lands on disk.

referencesvideo-recording.md

≈1.8k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Video Recording

Capture browser automation as video for debugging, documentation, or verification.

Related: commands.md for full function reference, SKILL.md for quick start.

Contents

Basic Recording

Enable video recording when opening a session:

# Start with recording enabled
SESSION=$(belt app run agent-browser --function open --session new --input '{
  "url": "https://example.com",
  "record_video": true
}' | jq -r '.session_id')

# Perform actions
belt app run agent-browser --function interact --session $SESSION --input '{
  "action": "click", "ref": "@e1"
}'

belt app run agent-browser --function interact --session $SESSION --input '{
  "action": "fill", "ref": "@e2", "text": "test input"
}'

# Close to get the video
RESULT=$(belt app run agent-browser --function close --session $SESSION --input '{}')
VIDEO=$(echo $RESULT | jq -r '.video')
echo "Video file: $VIDEO"

Cursor Indicator

For demos and documentation, show a visible cursor that follows mouse movements:

SESSION=$(belt app run agent-browser --function open --session new --input '{
  "url": "https://example.com",
  "record_video": true,
  "show_cursor": true
}' | jq -r '.session_id')

The cursor appears as a red dot that:

  • Follows mouse movements in real-time
  • Shows click feedback (shrinks on mousedown)
  • Persists across page navigations
  • Appears in both screenshots and video

This is especially useful for:

  • Tutorial/documentation videos
  • Debugging interaction issues
  • Sharing recordings with non-technical stakeholders

How Recording Works

  1. Start: Pass "record_video": true in the open function
  2. Record: All browser activity is captured throughout the session
  3. Stop: Video is finalized when close is called
  4. Retrieve: Video file is returned in the close response

The video captures:

  • Page loads and navigations
  • Element interactions (clicks, typing)
  • Scrolling and animations
  • Dynamic content changes

Use Cases

Debugging Failed Automation

#!/bin/bash
# Record automation for debugging

SESSION=$(belt app run agent-browser --function open --session new --input '{
  "url": "https://app.example.com",
  "record_video": true
}' | jq -r '.session_id')

# Run automation
RESULT=$(belt app run agent-browser --function interact --session $SESSION --input '{
  "action": "click", "ref": "@e1"
}')

SUCCESS=$(echo $RESULT | jq -r '.success')
if [ "$SUCCESS" != "true" ]; then
  echo "Action failed!"
  echo "Message: $(echo $RESULT | jq -r '.message')"

  # Get video for debugging
  CLOSE_RESULT=$(belt app run agent-browser --function close --session $SESSION --input '{}')
  echo "Debug video: $(echo $CLOSE_RESULT | jq -r '.video')"
  exit 1
fi

belt app run agent-browser --function close --session $SESSION --input '{}'

Documentation Generation

Record workflows for user documentation:

#!/bin/bash
# Record how-to video

SESSION=$(belt app run agent-browser --function open --session new --input '{
  "url": "https://app.example.com/settings",
  "record_video": true,
  "width": 1920,
  "height": 1080
}' | jq -r '.session_id')

# Add pauses for clarity
belt app run agent-browser --function interact --session $SESSION --input '{
  "action": "wait", "wait_ms": 1000
}'

# Step 1: Click settings
belt app run agent-browser --function interact --session $SESSION --input '{
  "action": "click", "ref": "@e5"
}'
belt app run agent-browser --function interact --session $SESSION --input '{
  "action": "wait", "wait_ms": 500
}'

# Step 2: Change setting
belt app run agent-browser --function interact --session $SESSION --input '{
  "action": "click", "ref": "@e10"
}'
belt app run agent-browser --function interact --session $SESSION --input '{
  "action": "wait", "wait_ms": 500
}'

# Step 3: Save
belt app run agent-browser --function interact --session $SESSION --input '{
  "action": "click", "ref": "@e15"
}'
belt app run agent-browser --function interact --session $SESSION --input '{
  "action": "wait", "wait_ms": 1000
}'

# Get the video
RESULT=$(belt app run agent-browser --function close --session $SESSION --input '{}')
echo "Documentation video: $(echo $RESULT | jq -r '.video')"

Test Evidence for CI/CD

#!/bin/bash
# Record E2E test for CI artifacts

TEST_NAME="${1:-e2e-test}"

SESSION=$(belt app run agent-browser --function open --session new --input '{
  "url": "'"$TEST_URL"'",
  "record_video": true
}' | jq -r '.session_id')

# Run test steps
run_test_steps $SESSION
TEST_RESULT=$?

# Always get video
CLOSE_RESULT=$(belt app run agent-browser --function close --session $SESSION --input '{}')
VIDEO=$(echo $CLOSE_RESULT | jq -r '.video')

# Save to artifacts
if [ -n "$CI_ARTIFACTS_DIR" ]; then
  cp "$VIDEO" "$CI_ARTIFACTS_DIR/${TEST_NAME}.webm"
fi

exit $TEST_RESULT

Monitoring and Auditing

#!/bin/bash
# Record automated task for audit trail

TASK_ID=$(date +%Y%m%d-%H%M%S)

SESSION=$(belt app run agent-browser --function open --session new --input '{
  "url": "https://admin.example.com",
  "record_video": true
}' | jq -r '.session_id')

# Perform admin task
# ... automation steps ...

# Save recording
RESULT=$(belt app run agent-browser --function close --session $SESSION --input '{}')
VIDEO=$(echo $RESULT | jq -r '.video')

# Archive for audit
mv "$VIDEO" "/audit/recordings/${TASK_ID}.webm"
echo "Audit recording saved: ${TASK_ID}.webm"

Best Practices

1. Add Strategic Pauses

Pauses make videos easier to follow:

# After significant actions, add a pause
'{"action": "click", "ref": "@e1"}'
'{"action": "wait", "wait_ms": 500}'  # Let viewer see result

2. Use Larger Viewport for Documentation

'{"url": "...", "record_video": true, "width": 1920, "height": 1080}'

3. Handle Errors Gracefully

Always retrieve video even on failure:

cleanup() {
  if [ -n "$SESSION" ]; then
    belt app run agent-browser --function close --session $SESSION --input '{}' 2>/dev/null
  fi
}
trap cleanup EXIT

4. Combine with Screenshots

Use screenshots for key frames, video for flow:

# Record overall flow
'{"record_video": true}'

# Capture key states
belt app run agent-browser --function screenshot --session $SESSION --input '{
  "full_page": true
}'

5. Don't Record Sensitive Sessions

Avoid recording when handling credentials:

if [ "$CONTAINS_SENSITIVE_DATA" = "true" ]; then
  RECORD="false"
else
  RECORD="true"
fi

'{"url": "...", "record_video": '$RECORD'}'

Output Format

  • Format: WebM (VP8/VP9 codec)
  • Compatibility: All modern browsers and video players
  • Quality: Matches viewport size
  • Compression: Efficient for screen content

Limitations

  1. Session-level only - Can't start/stop mid-session
  2. Memory usage - Long sessions consume more memory
  3. File size - Complex pages with animations produce larger files
  4. No audio - Browser audio is not captured
  5. Returned on close - Video only available after session ends

Source: SKILL.md on GitHub

1 alert13d5 checks · Risk SAFE
  • Gen Agent Trust Hub13d

    The skill provides powerful browser automation capabilities including JavaScript execution and file uploads. While these are intended features, they introduce risks such as indirect prompt injection from malicious websites and the potential for session data exfiltration via scripts.

  • Socket13d

    2 alerts: gptAnomaly, gptSecurity

  • Snyk13d

    Risk: MEDIUM · 1 issue

  • Runlayer6mo

    7/10 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at b576e8e. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last week.

Activeupdated 5 months ago
What it can do
Runs commands
All 1 allowed tools
Bash(belt *)
  • browser-automation
  • playwright
  • web-scraping
  • form-filling
  • screenshot
  • video-recording
  • javascript-execution
  • inference-sh
  • headless-browser

README badge

README badge for inference-shell/skills/agent-browser

Browser automation using Playwright with inference.sh, controlled via the `belt` CLI. Navigate pages, interact with elements using @e refs, take screenshots, record video, and run JavaScript. Suitable for web scraping, form filling, clicking, typing, drag-drop, file upload, and automated testing workflows.

Generated from the current SKILL.md.

What browser engine does this use?
Playwright. The skill wraps Playwright under the hood with a simplified @e ref system for element interaction.
Do I need to install anything besides the skill?
Yes, you need the belt CLI installed. The skill documentation includes install instructions.
Can I record video of browser sessions?
Yes. Pass record_video: true when opening a session, and the video file is returned when you close the session.
How do I interact with page elements?
Use the @e refs returned by open and snapshot calls. Pass the ref to interact actions like click, fill, drag, upload, or execute JavaScript.
Do element refs stay valid after navigation?
No. Refs are invalidated after navigation, form submission, or dynamic content loading. Always call snapshot after these events to get fresh refs.

Generated from the current SKILL.md. These answers refresh after source changes.