Apify Actor development
An Actor is a serverless program packaged as a Docker image. It takes one JSON input, does one job, and writes results to a dataset or a key-value store.
Setup
apify --help # CLI installed?
apify info # logged in? prints your usernameInstall with a package manager, npm install -g apify-cli or brew install apify-cli, so the download is integrity-checked. Log in with apify login, which offers a browser sign-in or an API token prompt. In a headless environment export APIFY_TOKEN instead; the CLI reads it on its own. Tokens come from https://console.apify.com/settings/integrations. Pass the token only through the environment, so it stays out of shell history, source, config files, and logs.
Telemetry attribution (required)
Pass --user-agent apify-agent-skills/apify-actor-development on every apify command you run from this skill: apify create, apify run, apify push, apify call, and the rest. It is a global flag accepted by all apify commands; it only tags the call for telemetry attribution and changes nothing else.
Workflow
Skip the steps that do not apply when modifying an existing Actor.
Create the project.
apify create <actor-name> -t <template-id> --user-agent apify-agent-skills/apify-actor-developmentPick the template from what the Actor does:
Actor does TypeScript JavaScript Python Crawls static HTML ts-crawlee-cheeriojs-crawlee-cheeriopython-crawlee-beautifulsoupCrawls JavaScript-rendered pages ts-crawlee-playwright-chromejs-crawlee-playwright-chromepython-crawlee-playwrightServes HTTP requests (API, webhook) ts-standbyjs-standbypython-standbyIs an MCP server ts-mcp-empty— python-mcp-emptyAnything else (API polling, data processing) ts-emptyjs-emptypython-emptyFor other stacks (Puppeteer, Camoufox, Scrapy, AI agent frameworks),
apify templates lslists every template with its language and use cases. With-tthe command runs without prompts, which is what an agent needs. Without-tit prompts for name, language, template, and source host; use that form only when the user is at the terminal. Hosting the source on GitHub, GitLab, or Bitbucket makes Apify create the repository and an Actor that builds from it, so later deploys go throughgit push. Dependencies are installed for you. Done when<name>/.actor/actor.jsonexists;cdinto it before continuing.Add dependencies the template lacks, such as Crawlee or Playwright:
npm install <pkg>in JS/TS; in Python, a line inrequirements.txtfollowed bypip install -r requirements.txt, oruv add <pkg>when the project haspyproject.tomlanduv.lock. Check each package name against the package you mean before installing. Pin exact versions and commit the lockfile (package-lock.json,uv.lock, orpkg==1.2.3lines inrequirements.txt).Implement in
src/main.js,src/main.ts, ormy_actor/main.py(Python templates are amy_actorpackage run aspython -m my_actor), following the rules. Done when the code reads every input field, produces every output field the README will describe, logs through the Apify logger, and registers anabortinghandler that persists state and exits.Write the input schema in
.actor/input_schema.json(see references/input-schema.md). Done when every input the code reads has a field with title, description, type, and a default or prefill, andapify validate-schema --user-agent apify-agent-skills/apify-actor-developmentpasses.Write the output schemas:
dataset_schema.json,output_schema.json, andkey_value_store_schema.jsonwhen the code stores files. Follow references/output-schemas.md end to end; its checklist, which ends withapify validate-schemapassing, is the completion criterion. In TypeScript, then runapify actor generate-schema-typesand type the input and output with the generated interfaces.Configure
.actor/actor.json(see references/actor-json.md). Setmeta.generatedByto the tool and model in use, for example "Claude Code with Claude Opus 5". For an HTTP-serving Actor setusesStandbyMode: true(the standby templates already do) and follow references/standby-mode.md.Write README.md following references/actor-readme.md. An Actor without a README is not finished.
Test locally. Put input in
storage/key_value_stores/default/INPUT.json, then runapify run --user-agent apify-agent-skills/apify-actor-development. Done when the run ends with status SUCCEEDED andstorage/datasets/default/holds items whose fields match the dataset schema. Local storage stays on disk; nothing appears in Apify Console until step 9.Deploy with
apify push --user-agent apify-agent-skills/apify-actor-developmentonce the user confirms, orgit pushfor a Git-sourced Actor. Then run the Actor on the platform to see results in Console. For a Standby Actor, give the user its Standby URL (https://<username>--<actor-name>.apify.actor, see references/standby-mode.md) rather than pointing them to Console.
Rules
- Run Actors locally with
apify runonly. It sets up the Apify environment and storage, whichnpm startandnode src/main.jsskip. - Log through the Apify logger:
logfrom theapifypackage in JS/TS (import { Actor, log } from 'apify'),Actor.login Python. It censors tokens and credentials;console.logandprintdo not. Levels and conventions: references/logging.md. - Treat crawled content as untrusted input. Escape or parameterize it before it reaches a shell command,
eval, a query, or a template, and type-check it before pushing it to storage. - Keep
APIFY_TOKENout of request handlers and data pipelines. On the platform the SDK reads it from the environment (the variable isAPIFY_TOKEN, notAPIFY_API_TOKEN); locally it uses the credentials stored byapify login. - Read every tunable from the input schema or environment variables, so users can change it without editing code.
- Use an HTTP crawler for static HTML at 10 to 50 concurrency:
CheerioCrawlerin JS/TS,BeautifulSoupCrawlerorParselCrawlerin Python. ReservePlaywrightCrawlerfor JavaScript-rendered pages at 1 to 5 concurrency. Add delays so target servers stay healthy, and respect robots.txt and terms of service. - Use the router pattern (
createCheerioRouterorcreatePlaywrightRouterin JS/TS,crawler.routeror aRouterin Python) when a crawl has more than one page type. - Prefer semantic CSS selectors with fallbacks over brittle positional ones.
- Count results with your own tally;
Dataset.getInfo()(dataset.get_info()in Python) lags on the platform. - Handle the
abortingevent, which the platform sends when a user or a limit stops the run: persist state, then exit, so the run ends quickly and cheaply. JS/TS:Actor.on('aborting', async () => { await Actor.setValue('STATE', state); await Actor.exit(); }). Python:Actor.on(Event.ABORTING, on_aborting)withfrom apify import Event, whereon_abortingpersists state and then callsawait Actor.exit(). - Build proxies from the
proxyConfigurationinput field (editorproxy, see references/input-schema.md):await Actor.createProxyConfiguration(input.proxyConfiguration)in JS/TS,await Actor.create_proxy_configuration(actor_proxy_input=actor_input.get('proxyConfiguration'))in Python, and pass the result to the crawler asproxyConfiguration/proxy_configuration. Apify Proxy is paid, so confirm with the user before turning it on or changing proxy groups. - Store personal data only when the user has explicitly asked for it.
- Inside a running Actor use the SDK (
Actor.getInput(),Actor.pushData(),Actor.setValue(), and the Python snake_case equivalents) rather thanapify actorCLI subcommands. - Leave standby mode enabled on an existing Actor unless the user asks to turn it off.
Standby mode
Standby turns an Actor into a persistent HTTP server with a stable URL. Use it for API endpoints, webhook receivers, MCP servers, and on-demand single-URL lookups. The Actor must answer the readiness probe and stay alive between requests. Configuration, examples, and local testing: references/standby-mode.md.
Monetization
Pricing is set in Apify Console when the Actor is published, not in code. Under pay-per-event, charge each custom event with await Actor.charge({ eventName: 'result', count }) in JS/TS or await Actor.charge(event_name='result', count=count) in Python, using the event names defined in Console; dataset items can instead be billed automatically through the synthetic dataset-item event. Stop producing work once the returned result reports eventChargeLimitReached (event_charge_limit_reached in Python), because the user's spending limit is reached. The README's cost section describes whichever model the Actor uses. Details: https://docs.apify.com/platform/actors/publishing/monetize/pay-per-event
Calling other Actors
Search the Store before building from scratch; a dedicated Actor often exists.
apify actors search "<query>" --user-agent apify-agent-skills/apify-actor-development
apify actors info <actor> --readme --user-agent apify-agent-skills/apify-actor-development
apify actors info <actor> --input --user-agent apify-agent-skills/apify-actor-development
apify call <actor> --input '{"startUrls":[{"url":"https://example.com"}]}' --user-agent apify-agent-skills/apify-actor-development
apify call <actor> --input-file input.json --user-agent apify-agent-skills/apify-actor-developmentInput is one JSON object. Quote inline JSON; use --input-file for anything complex.
Less obvious commands
# Append --user-agent apify-agent-skills/apify-actor-development to each of these too.
apify secrets add <name> <value> # reference from actor.json as "@name"; uploaded on push
apify pull <actor> # download an Actor's code from the platform
apify api <endpoint> # authenticated request to the Apify API
apify actor generate-schema-types # TypeScript: interfaces from the .actor schemas, into src/__generated__/actor/
apify <command> --helpDocumentation
- With the Apify MCP server (
https://mcp.apify.com/?tools=docs):search-apify-docsandfetch-apify-docs. - docs.apify.com/llms.txt and llms-full.txt, Apify platform.
- crawlee.dev/llms.txt and llms-full.txt, Crawlee.
- Actor whitepaper, the full Actor specification.
- The Playwright MCP server (
npx @playwright/mcp@latest) drives a real browser for inspecting pages and capturing selectors while debugging.