All skills
callstackincubator avatar

/agent-device

@23643af official

Automates interactions for Apple-platform apps (iOS, tvOS, macOS) and Android devices. Use when navigating apps, taking snapshots/screenshots, tapping, typing, scrolling, or extracting UI info across mobile, TV, and desktop targets.

Use this Skill: https://skilld.dev/gh/callstackincubator/agent-skills/agent-device

This session only. Nothing lands on disk.

referencesverification.md

≈990 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Verification

When to open this file

Open this file when the task needs evidence, regression checks, replay maintenance, or startup performance measurements after the main interaction flow is already working.

Main commands to reach for first

  • screenshot
  • diff snapshot
  • record
  • replay -u
  • perf

Most common mistake to avoid

Do not use verification tools as the first exploration step. First get the app into the correct state with the normal interaction flow, then capture proof or maintain replay assets.

Canonical loop

agent-device open Settings --platform ios
# after using exploration to reach the state you want to verify
agent-device snapshot
agent-device screenshot /tmp/settings-proof.png
agent-device close

Structural verification with diff snapshot

Use diff snapshot when you need a compact view of how the UI changed between nearby states.

agent-device snapshot -i
agent-device press @e5
agent-device diff snapshot -i
  • Initialize the baseline at a stable point.
  • Perform the mutation.
  • Run diff snapshot to confirm the expected structural change.
  • Re-run full snapshot only when you need fresh refs.

Visual artifacts

Use screenshot when the proof needs a rendered image instead of a structural tree.

Session recording

Use record for debugging, documentation, or shareable verification artifacts.

agent-device record start ./recordings/ios.mov
agent-device open App
agent-device snapshot -i
agent-device press @e3
agent-device close
agent-device record stop
  • record supports iOS simulators, iOS devices, and Android.
  • On iOS, recording is a wrapper around simctl for simulators and the corresponding device capture path for physical devices.
  • On Android, recording is a wrapper around adb.
  • Recording writes a video artifact and a gesture-telemetry sidecar JSON.
  • On macOS hosts, touch overlay burn-in is available for supported recordings.
  • On non-macOS hosts, recording still succeeds but the video stays raw and record stop can return an overlayWarning.
  • If the agent already knows the interaction sequence and wants a more lifelike, uninterrupted recording, drive the flow with batch while recording instead of replanning between each step.

Example:

agent-device record start ./recordings/smoke.mov
agent-device batch --session sim --platform ios --steps-file /tmp/smoke-steps.json --json
agent-device record stop
  • Use this only after exploration has stabilized the flow.
  • Keep the batch short and add wait or is exists guards after mutating steps so the recorded flow still tracks realistic UI timing.

Replay maintenance

Use replay updates when selectors drift but the recorded scenario is still correct.

agent-device replay -u ./session.ad
agent-device test ./smoke --platform android
  • Prefer selector-based actions in recorded .ad replays.
  • Use test when you already have multiple .ad flows and need a quick regression pass after updating or recording them.
  • Keep the skill-level rule simple: use replay -u to maintain one script, use test to verify a folder or matcher of scripts.
  • Treat test as a human and CI-facing suite runner that an agent can invoke for verification, not as the main source of product documentation.
  • Failed runs keep suite artifacts under .agent-device/test-artifacts by default, which is usually enough for debugging without extra agent-side processing.
  • Use update mode for maintenance, not as a substitute for fixing a broken interaction strategy.

Performance checks

Use perf --json or metrics --json when you need startup timing for the active session.

agent-device open Settings --platform ios
agent-device perf --json
  • Current startup data is command round-trip timing around open.
  • It is not true first-frame or first-interactive telemetry.
  • fps, memory, and cpu are currently placeholders.

Source: SKILL.md on GitHub

No alerts17d3 checks · Risk SAFE
  • Gen Agent Trust Hub17d

    The analyzed skill provides documentation and guidance for an automated device interaction tool (agent-device) supporting iOS, Android, and macOS targets. It includes instruction references for bootstrapping, coordinate systems, debugging, UI exploration, macOS desktop specificity, remote tenancy (daemon over HTTP), and verification. No malicious behaviors or significant security concerns were detected; the skill primarily orchestrates local or remote mobile/desktop automation tools safely.

  • Socket17d

    No alerts

  • Snyk17d

    Risk: LOW · No issues

Signed by skilld at 23643af. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 weeks ago.

Activeupdated 6 months ago
  • Testing
  • ios
  • android
  • macos
  • tvos
  • mobile
  • ui-automation
  • device-interaction
  • snapshots
  • app-testing

README badge

README badge for callstackincubator/agent-skills/agent-device

Automates UI interaction and inspection for iOS, tvOS, macOS, and Android apps via snapshot, tap, type, scroll, and element targeting commands. Use this skill to navigate apps, extract UI state, and verify behavior across mobile and desktop platforms without writing native test code.

Generated from the current SKILL.md.

What platforms does this skill support?
iOS, tvOS, macOS, and Android. The skill routes interactions across all four platforms with platform-specific references for desktop and remote scenarios.
Do I need to set up the device or simulator first?
Yes. Always load bootstrap-install.md before acting to confirm the target, app install, and open app session in a deterministic way.
Should I use snapshot or snapshot -i?
Use plain snapshot to verify what is visible on screen. Use snapshot -i only when you need interactive refs like @e3 for a specific action or targeted query.
Can this skill handle web browsing or external lookups?
No. The skill does not browse the web or use external sources unless explicitly requested by the user.
How do I debug failures, capture recordings, or access logs?
Load references/debugging.md for logs and failure triage, or references/verification.md for screenshots, recordings, and perf data.

Generated from the current SKILL.md. These answers refresh after source changes.