All skills
callstackincubator avatar

/agent-device

@23643af official

Automates interactions for Apple-platform apps (iOS, tvOS, macOS) and Android devices. Use when navigating apps, taking snapshots/screenshots, tapping, typing, scrolling, or extracting UI info across mobile, TV, and desktop targets.

Use this Skill: https://skilld.dev/gh/callstackincubator/agent-skills/agent-device

This session only. Nothing lands on disk.

referencesmacos-desktop.md

≈846 tokens on demand. Your agent reads this file only when SKILL.md points to it.

macOS Desktop

When to open this file

Open this file only when --platform macos is involved or the task needs frontmost-app, desktop, or menubar surfaces.

Main commands to reach for first

  • open <app> --platform macos
  • open --platform macos --surface frontmost-app|desktop|menubar
  • snapshot -i
  • get
  • is
  • click --button secondary

Most common mistake to avoid

Do not treat every macOS surface the same. Use the normal app surface when you want to act inside one app. Use frontmost-app, desktop, or menubar mainly to inspect what is visible before switching back to app for most interactions.

Canonical loop

agent-device open TextEdit --platform macos
agent-device snapshot
agent-device close

Surface rules

  • app: default surface and the normal choice for click, fill, press, scroll, screenshot, and record.
  • frontmost-app: inspect the currently focused app without naming it first.
  • desktop: inspect visible desktop windows across apps.
  • menubar: inspect the active app menu bar and system menu extras.

Use inspect-first surfaces to understand desktop-global UI, then switch back to app when you need to act in one app.

Snapshot expectations

  • snapshot -i should describe UI visible to a human.
  • desktop snapshots can include multiple windows from multiple apps.
  • menubar snapshots can include both app-menu items and system menu extras.
  • Finder-style rows, sidebar items, toolbar controls, search fields, and opened context menus should appear when visible.
  • Finder and other native apps may expose duplicate-looking row, cell, and child text nodes. Treat them as distinct AX nodes unless you have a stronger selector anchor.

Context menus

Context menus are not ambient UI. Open them explicitly, then re-snapshot.

agent-device click @e66 --button secondary --platform macos
agent-device snapshot -i

Expected loop:

  1. Snapshot visible content.
  2. Secondary-click the target item.
  3. Snapshot again.
  4. Interact with the new menu-item nodes.

Targeting rules

  • Prefer selectors or @ref values over raw coordinates.
  • On macOS, window position can vary across runs, so coordinate-only flows are fragile.
  • If the task only needs shared exploration rules, return to exploration.md.

Selector guidance:

  • Good selectors usually anchor on stable labels or app-owned identifiers such as label="Downloads" or role=menu-item label="Rename".
  • Avoid relying on framework-generated _NS:* identifiers as stable selectors.

Use snapshot --raw --platform macos only when debugging AX structure or collector filtering. Do not make raw snapshots the default agent loop.

Things not to rely on:

  • Mobile-only helpers such as install, reinstall, or push.
  • Desktop-global click or fill parity from desktop or menubar sessions.
  • Raw coordinate assumptions across runs.

Troubleshooting:

  • If visible content is missing from snapshot -i, re-snapshot after the UI settles.
  • If desktop is too broad, retry with frontmost-app.
  • If menubar is missing the expected menu, make the app frontmost first and retry.
  • If the wrong menu opened, retry secondary-clicking the row or cell wrapper rather than the nested text node.
  • If the app has multiple windows, make the correct window frontmost before relying on refs.

Source: SKILL.md on GitHub

No alerts17d3 checks · Risk SAFE
  • Gen Agent Trust Hub17d

    The analyzed skill provides documentation and guidance for an automated device interaction tool (agent-device) supporting iOS, Android, and macOS targets. It includes instruction references for bootstrapping, coordinate systems, debugging, UI exploration, macOS desktop specificity, remote tenancy (daemon over HTTP), and verification. No malicious behaviors or significant security concerns were detected; the skill primarily orchestrates local or remote mobile/desktop automation tools safely.

  • Socket17d

    No alerts

  • Snyk17d

    Risk: LOW · No issues

Signed by skilld at 23643af. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 weeks ago.

Activeupdated 6 months ago
  • Testing
  • ios
  • android
  • macos
  • tvos
  • mobile
  • ui-automation
  • device-interaction
  • snapshots
  • app-testing

README badge

README badge for callstackincubator/agent-skills/agent-device

Automates UI interaction and inspection for iOS, tvOS, macOS, and Android apps via snapshot, tap, type, scroll, and element targeting commands. Use this skill to navigate apps, extract UI state, and verify behavior across mobile and desktop platforms without writing native test code.

Generated from the current SKILL.md.

What platforms does this skill support?
iOS, tvOS, macOS, and Android. The skill routes interactions across all four platforms with platform-specific references for desktop and remote scenarios.
Do I need to set up the device or simulator first?
Yes. Always load bootstrap-install.md before acting to confirm the target, app install, and open app session in a deterministic way.
Should I use snapshot or snapshot -i?
Use plain snapshot to verify what is visible on screen. Use snapshot -i only when you need interactive refs like @e3 for a specific action or targeted query.
Can this skill handle web browsing or external lookups?
No. The skill does not browse the web or use external sources unless explicitly requested by the user.
How do I debug failures, capture recordings, or access logs?
Load references/debugging.md for logs and failure triage, or references/verification.md for screenshots, recordings, and perf data.

Generated from the current SKILL.md. These answers refresh after source changes.