---
title: "cmux-browser by manaflow-ai · skilld"
canonical_url: "https://skilld.dev/gh/manaflow-ai/cmux"
meta:
  description: "Automates browser interactions within cmux webviews by opening pages, waiting for state changes, snapshotting the DOM, and performing actions like clicks and form fills using… From manaflow-ai/cmux."
  "og:description": "Automates browser interactions within cmux webviews by opening pages, waiting for state changes, snapshotting the DOM, and performing actions like clicks and form fills using… From manaflow-ai/cmux."
  "og:title": "cmux-browser by manaflow-ai"
  "twitter:description": "Automates browser interactions within cmux webviews by opening pages, waiting for state changes, snapshotting the DOM, and performing actions like clicks and form fills using… From manaflow-ai/cmux."
  "twitter:title": "cmux-browser by manaflow-ai"
---

`

[All skills](https://skilld.dev/skills)

[![manaflow-ai avatar](https://skilld.dev/_img/avatar?url=https%3A%2F%2Fgithub.com%2Fmanaflow-ai.png%3Fsize%3D96)](https://skilld.dev/gh/manaflow-ai)

# **/cmux-browser**

[@f97f1e9](https://github.com/manaflow-ai/cmux/commit/f97f1e948071d6adf779cf1c46fb7d508ca9a42a "Your agent reads SKILL.md at commit f97f1e9")

by [manaflow-ai](https://skilld.dev/gh/manaflow-ai)· [manaflow-ai](https://skilld.dev/gh/manaflow-ai)/ [cmux](https://skilld.dev/gh/manaflow-ai/cmux)·28k stars

 2,423

End-user browser automation with cmux. Use when you need to open sites, inspect or interact with browser surfaces, wait for page state, and extract data without stealing focus.

- 13 files
- 36 KB
- Updated 3 weeks ago
- [GitHub](https://github.com/manaflow-ai/cmux/blob/f97f1e948071d6adf779cf1c46fb7d508ca9a42a/skills/cmux-browser/SKILL.md "View SKILL.md on GitHub")
- [1 alert](#third-party-checks "Third-party checks: 1 alert · 5 checks · Risk SAFE")

## SKILL.md

10.7 KB

**≈48** tokens always: the name and description. **≈2.7k** when used: this file. **≈6.1k** more on demand in 9 files.

## Browser Automation with cmux

### Read the CLI contract first

Check the binary that will actually run before giving an exact command:

```
cmux browser --help
cmux --version
```

There are two deliberately different command shapes:

- **Create a browser surface** with `open`, `open-split`, or `new`. These commands may be workspace-scoped and do not need a surface handle.
- **Use an existing surface** with every surface-bound navigation, inspection, interaction, tab, state, or diagnostic command. Pass the handle explicitly with `--surface <handle>` or as the first positional token.

Prefer the flag form in scripts because it makes the target unmissable:

```
SURFACE="surface:7" # use a ref returned by discovery; do not guess an index
cmux browser --surface "$SURFACE" get url
cmux browser --surface "$SURFACE" get-url       # accepted alias
cmux browser --surface "$SURFACE" snapshot --interactive
cmux browser --surface "$SURFACE" snapshot -i   # accepted alias
cmux browser --surface "$SURFACE" url            # accepted alias
cmux browser --surface "$SURFACE" tab list
cmux browser --surface "$SURFACE" click e1 --snapshot-after
```

The positional form is equivalent (`cmux browser "$SURFACE" get url`). `url` and `get-url` are accepted URL aliases, and the short interactive snapshot flag is accepted when a surface is already present; use `get url` and `snapshot --interactive` in new documentation so the target and operation are clear. Surface-bound operations have no unscoped form. The current CLI's explicitly global browser verbs (`open`, `open-split`, `new`, `identify`, `import`, `profile`, `profiles`, `react-grab`, `reactgrab`, `devtools`, `dev-tools`, `focus-mode`, `design-mode`, `zoom`, and `history`) may omit the handle and use caller/workspace routing; do not infer a target from visible focus for any other verb.

### Find an existing browser surface without changing focus

`identify`, `tree`, and list commands are read-only and do not select a workspace, pane, or browser. Do not infer that the visually focused surface is the one the user wants.

First inspect the caller context (useful for the default workspace):

```
cmux identify --json
```

To discover browser surfaces in the caller or another workspace/window, use the all-window tree. It includes parent refs, so a browser in a different workspace can be targeted directly without selecting that workspace:

```
cmux tree --all --json \
  | jq -r '
      .windows[]? as $window
      | $window.workspaces[]? as $workspace
      | $workspace.panes[]? as $pane
      | $pane.surfaces[]?
      | select(.type == "browser")
      | [$window.ref, $workspace.ref, $pane.ref, .ref]
      | @tsv'
```

The filtered output is `window`, `workspace`, `pane`, and `surface` refs. Keep the `surface` ref, then target it explicitly:

```
SURFACE="surface:N" # copied from the filtered tree output
cmux browser --surface "$SURFACE" get url
cmux browser --surface "$SURFACE" snapshot --interactive
```

If the user gives a URL or title instead of a workspace/pane, match that metadata locally and emit only the unique surface ref. This never prints the matched URL or title:

```
MATCH_FIELD="url" # use "title" when matching a page title
MATCH_VALUE="${BROWSER_URL_OR_TITLE:?set BROWSER_URL_OR_TITLE without logging it}"
SURFACE="$(
  cmux tree --all --json |
    jq -r --arg field "$MATCH_FIELD" --arg value "$MATCH_VALUE" '
      [
        .windows[]? as $window
        | $window.workspaces[]? as $workspace
        | $workspace.panes[]? as $pane
        | $pane.surfaces[]?
        | select(.type == "browser")
        | select((if $field == "url" then (.url // "") else (.title // "") end) == $value)
        | .ref
      ] as $matches
      | if ($matches | length) == 1 then $matches[0]
        elif ($matches | length) == 0 then error("no matching browser surface")
        else error("multiple matches; use workspace/pane context")
        end'
)"
if [[ -z "$SURFACE" ]]; then
  printf '%s\n' 'no uniquely matching browser surface; provide workspace/pane context' >&2
  exit 1
fi
cmux browser --surface "$SURFACE" get url
```

For one known workspace, `cmux --json list-pane-surfaces --workspace <workspace>` is a smaller read-only query. Raw tree/list payloads can contain page URLs and titles; filter or redact them before logging or pasting them. Never use a focus/select command merely to discover a surface.

### Core workflow

Open (or create) a surface without stealing focus, capture the returned ref, then use that ref for every existing-surface operation:

```
OPEN_JSON="$(cmux --json browser open https://example.com --focus false)"
SURFACE="$(printf '%s' "$OPEN_JSON" | jq -r '.surface_ref // .surface_id // empty')"
[ -n "$SURFACE" ] || { printf '%s\n' 'browser open did not return a surface ref' >&2; exit 1; }
cmux browser --surface "$SURFACE" get url
cmux browser --surface "$SURFACE" wait --load-state complete --timeout-ms 15000
cmux browser --surface "$SURFACE" snapshot --interactive
cmux browser --surface "$SURFACE" fill e1 "hello"
cmux browser --surface "$SURFACE" click e2 --snapshot-after
cmux browser --surface "$SURFACE" snapshot --interactive
```

After a browser download finishes, inspect the same surface's bounded history without opening the file or consuming a waiter:

```
cmux browser --surface "$SURFACE" download list
cmux browser --surface "$SURFACE" download list --limit 5 --json
```

The JSON records expose the stable `download_id`, filename, actual saved path when known, status (`downloading`, `saved`, or `failed`), byte count when known, and whether a known path still exists. Listing is newest first and repeatable; it remains scoped to the requested surface. Use `download wait` to keep the existing event-wait workflow.

The `open` response contains the new surface ref; in a script, extract it from the JSON response instead of printing the full response. If `get url` is empty or `about:blank`, navigate first instead of waiting on load state. Re-snapshot after navigation, modal open/close, or any major DOM change because refs go stale.

### Wait

```
cmux browser --surface "$SURFACE" wait --selector "#ready" --timeout-ms 10000
cmux browser --surface "$SURFACE" wait --text "Success" --timeout-ms 10000
cmux browser --surface "$SURFACE" wait --url-contains "/dashboard" --timeout-ms 10000
cmux browser --surface "$SURFACE" wait --load-state complete --timeout-ms 15000
cmux browser --surface "$SURFACE" wait --function "document.readyState === 'complete'" --timeout-ms 10000
```

### Viewport sizing (WKWebView)

`cmux browser --surface "$SURFACE" viewport <width> <height>` sets an exact logical viewport from 1 to 4096 CSS pixels. The page is aspect-fitted inside its existing pane, so pane layout and focus stay unchanged, and screenshots use the requested logical dimensions. `viewport reset` returns to native pane sizing.

Close or detach the browser inspector first: its inspector-managed split layout cannot be combined with viewport emulation, and opening or redocking an attached inspector resets emulation to native sizing. Large viewport and page-zoom combinations are bounded; the command returns structured `maximum_page_zoom` details and leaves the viewport unchanged when the combination exceeds WKWebView render limits.

### Limits (WKWebView)

Offline emulation, trace/screencast recording, network route interception/mocking, and low-level raw input injection return `not_supported`; they depend on Chrome/CDP-only APIs. Use `click`, `fill`, `press`, `scroll`, `wait`, and `snapshot` instead.

### Troubleshooting `js_error`

Some complex pages reject the JavaScript behind `snapshot --interactive` and `eval`. Recover by checking whether the page actually navigated, then fall back to raw text or HTML:

```
cmux browser --surface "$SURFACE" get url
cmux browser --surface "$SURFACE" get text body
cmux browser --surface "$SURFACE" get html body
```

If it still fails, navigate to a simpler intermediate page and retry from there. If the CLI and this skill disagree, refresh help (`cmux browser --help`) and refresh the installed skill before continuing; do not invent an implicit target.

### Skill distribution and refresh

The repository copies are the source of truth: `.claude/skills/cmux-browser` and `.agents/skills/cmux-browser` point at `skills/cmux-browser`. Do not edit a mirror by hand. The supported Vercel installer (pinned here to the reviewed `skills` 1.5.23 release) refreshes both global Claude Code and Codex discovery roots and copies the complete skill (including references and templates):

```
# From this checkout while developing the skill:
npx --yes skills@1.5.23 add . --global --yes --skill cmux-browser --agent claude-code codex --copy

# From the published repository after the change is merged:
npx --yes skills@1.5.23 add manaflow-ai/cmux --global --yes --skill cmux-browser --agent claude-code codex --copy
```

Restart an agent session after a refresh if it cached the previous document. The repository's `skills.sh` remains available for a Codex-only destination; pass its `--dest` explicitly when that is the installation path:

```
./skills.sh --dest "$HOME/.codex/skills" --skill cmux-browser
```

Never commit home-directory skill copies, credentials, cookies, or saved browser state.

### Deep-dive references

| Reference | When to Use |
| --- | --- |
| [references/surface-discovery.md](https://skilld.dev/gh/manaflow-ai/cmux/cmux-browser/-/references/surface-discovery.md) | Find and target an existing browser surface without focus changes |
| [references/commands.md](https://skilld.dev/gh/manaflow-ai/cmux/cmux-browser/-/references/commands.md) | Full command mapping, aliases, `agent-browser` equivalents, viewport error codes |
| [references/snapshot-refs.md](https://skilld.dev/gh/manaflow-ai/cmux/cmux-browser/-/references/snapshot-refs.md) | Ref lifecycle and stale-ref troubleshooting |
| [references/authentication.md](https://skilld.dev/gh/manaflow-ai/cmux/cmux-browser/-/references/authentication.md) | Login/OAuth/2FA patterns and state save/load |
| [references/session-management.md](https://skilld.dev/gh/manaflow-ai/cmux/cmux-browser/-/references/session-management.md) | Multi-surface isolation and state persistence |
| [references/video-recording.md](https://skilld.dev/gh/manaflow-ai/cmux/cmux-browser/-/references/video-recording.md) | Recording status and practical alternatives |
| [references/proxy-support.md](https://skilld.dev/gh/manaflow-ai/cmux/cmux-browser/-/references/proxy-support.md) | Proxy behavior in WKWebView and workarounds |

### Ready-to-use templates

| Template | Description |
| --- | --- |
| [templates/form-automation.sh](https://github.com/manaflow-ai/cmux/blob/main/skills/cmux-browser/templates/form-automation.sh) | Snapshot/ref form fill loop (requires an explicit surface) |
| [templates/authenticated-session.sh](https://github.com/manaflow-ai/cmux/blob/main/skills/cmux-browser/templates/authenticated-session.sh) | Login once, save/load state (requires an explicit surface) |
| [templates/capture-workflow.sh](https://github.com/manaflow-ai/cmux/blob/main/skills/cmux-browser/templates/capture-workflow.sh) | Navigate and capture snapshots/screenshots (requires an explicit surface) |

Source: [SKILL.md on GitHub](https://github.com/manaflow-ai/cmux/blob/f97f1e948071d6adf779cf1c46fb7d508ca9a42a/skills/cmux-browser/SKILL.md)

## Third-party checks

<details>

<summary>1 alert16d5 checks · Risk SAFE</summary>



- Gen Agent Trust Hub16d

  The skill provides browser automation capabilities and correctly identifies the sensitivity of browser session data, providing best practices for securing it with restricted permissions and safe credential handling. It includes a mechanism for updates using an installer utility from the author's repository. The primary security consideration is the risk of indirect prompt injection inherent in processing content from arbitrary websites.
- Socket16d

  No alerts
- Snyk16d

  Risk: MEDIUM · 1 issue
- Runlayer6mo

  5/11 files flagged
- ZeroLeaks5mo

  Score: 93/100 · 2 sections analyzed

</details>

## Provenance

[Signed by skilld at f97f1e9.](https://github.com/manaflow-ai/cmux/commit/f97f1e948071d6adf779cf1c46fb7d508ca9a42a "f97f1e948071d6adf779cf1c46fb7d508ca9a42a") This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 days ago.

Activeupdated 3 weeks ago

## Topics

- [CLI](https://skilld.dev/skills/tag/cli "Command-line tools, scripting, shell")
- browser-automation
- cmux
- web-scraping
- form-automation
- wkwebview
- macos

## README badge

![README badge for manaflow-ai/cmux](https://skilld.dev/b/manaflow-ai/cmux?theme=light&label=0)

## What it does

Automates browser interactions within cmux webviews by opening pages, waiting for state changes, snapshotting the DOM, and performing actions like clicks and form fills using element references. Use this for tasks like form submission, data extraction, and navigation verification in cmux surfaces.

Generated from the current SKILL.md.

## Frequently asked

<details>

<summary>Does this skill work with WKWebView, or only Chrome?</summary>



It uses WKWebView. Some Chrome/CDP-only features like viewport emulation, network mocking, and trace recording are not supported, but core actions (click, fill, press, scroll, wait, snapshot) work.

</details>

<details>

<summary>How do I handle authentication and preserve login state across browser tasks?</summary>



Use the authenticated-session template or follow the authentication reference guide, which covers login flows, OAuth, 2FA patterns, and the save/load state workflow to persist credentials between surfaces.

</details>

<details>

<summary>What should I do if snapshot --interactive fails with a js_error?</summary>



Fall back to get url, get text body, or get html body to verify page state. If the issue persists, navigate to a simpler intermediate page and retry the task from there.

</details>

<details>

<summary>Can I run multiple browser tasks in parallel or do I need one surface per task?</summary>



Keep one surface per task unless you intentionally switch. Multi-surface isolation and state persistence patterns are covered in the session-management reference.

</details>

<details>

<summary>Does this work with the agent-browser skill or are they separate workflows?</summary>



This skill is specific to cmux webviews. It uses cmux CLI commands and surface references; the wait patterns are similar to agent-browser but the execution model is distinct to cmux.

</details>

Generated from the current SKILL.md. These answers refresh after source changes.

## Related skills

-
-
-
-
-
-