All skills
apify avatar

/apify-actorization

@16bac29 official
by apifyapify/agent-skills2.4k stars
259

Convert existing projects into Apify Actors - serverless cloud programs. Actorize JavaScript/TypeScript (SDK with Actor.init/exit), Python (async context manager), or any language (CLI wrapper). Use when migrating code to Apify, wrapping CLI tools as Actors, or adding Actor SDK to existing projects.

Use this Skill: https://skilld.dev/gh/apify/agent-skills/apify-actorization

This session only. Nothing lands on disk.

referencesschemas-and-output.md

≈1.1k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Schemas and output configuration

Input schema

Map your application's inputs to .actor/input_schema.json. Validate against the JSON Schema from the @apify/json_schemas npm package (input.schema.json).

{
    "title": "My Actor Input",
    "type": "object",
    "schemaVersion": 1,
    "properties": {
        "startUrl": {
            "title": "Start URL",
            "type": "string",
            "description": "The URL to start processing from",
            "editor": "textfield",
            "prefill": "https://example.com"
        },
        "maxItems": {
            "title": "Max Items",
            "type": "integer",
            "description": "Maximum number of items to process",
            "default": 100,
            "minimum": 1
        }
    },
    "required": ["startUrl"]
}

Mapping guidelines

  • Command-line arguments → input schema properties
  • Environment variables → input schema or Actor env vars in actor.json
  • Config files → input schema with object/array types
  • Flatten deeply nested structures for better UX

Output schema

Choose where output goes, then write the schema files (output_schema.json, dataset_schema.json, key_value_store_schema.json) by invoking the apify-actor-development skill, whose references/output-schemas.md holds the rules.

For table-like data (multiple items)

  • Use Actor.pushData() (JS) or Actor.push_data() (Python)
  • Each item becomes a row in the dataset

For single files or blobs

  • Use key-value store: Actor.setValue() / Actor.set_value()
  • Get the public URL and include it in the dataset:
// Store file with public access
await Actor.setValue('report.pdf', pdfBuffer, { contentType: 'application/pdf' });

// Get the public URL
const storeInfo = await Actor.openKeyValueStore();
const publicUrl = `https://api.apify.com/v2/key-value-stores/${storeInfo.id}/records/report.pdf`;

// Include URL in dataset output
await Actor.pushData({ reportUrl: publicUrl });

For multiple files with a common prefix (collections)

// Store multiple files with a prefix
for (const [name, data] of files) {
    await Actor.setValue(`screenshots/${name}`, data, { contentType: 'image/png' });
}
// Files are accessible at: .../records/screenshots%2F{name}

Actor configuration (actor.json)

Configure .actor/actor.json. Validate against the JSON Schema from the @apify/json_schemas npm package (actor.schema.json).

{
    "actorSpecification": 1,
    "name": "my-actor",
    "title": "My Actor",
    "description": "Brief description of what the Actor does",
    "version": "1.0.0",
    "meta": {
        "templateId": "ts_empty",
        "generatedBy": "Claude Code with Claude Opus 4.5"
    },
    "inputSchema": "./input_schema.json",
    "dockerfile": "../Dockerfile"
}

Important: Fill in the generatedBy property with the tool/model used.

State management

Request queue - for pausable task processing

The request queue works for any task processing, not just web scraping. Use a dummy URL with custom uniqueKey and userData for non-URL tasks:

const requestQueue = await Actor.openRequestQueue();

// Add tasks to the queue (works for any processing, not just URLs)
await requestQueue.addRequest({
    url: 'https://placeholder.local',  // Dummy URL for non-scraping tasks
    uniqueKey: `task-${taskId}`,       // Unique identifier for deduplication
    userData: { itemId: 123, action: 'process' },  // Your custom task data
});

// Process tasks from the queue (with Crawlee)
const crawler = new BasicCrawler({
    requestQueue,
    requestHandler: async ({ request }) => {
        const { itemId, action } = request.userData;
        // Process your task using userData
        await processTask(itemId, action);
    },
});
await crawler.run();

// Or manually consume without Crawlee:
let request;
while ((request = await requestQueue.fetchNextRequest())) {
    await processTask(request.userData);
    await requestQueue.markRequestHandled(request);
}

Key-value store - for checkpoint state

// Save state
await Actor.setValue('STATE', { processedCount: 100 });

// Restore state on restart
const state = await Actor.getValue('STATE') || { processedCount: 0 };

Source: SKILL.md on GitHub

1 warning1d5 checks · Risk SAFE
  • Gen Agent Trust Hub1d

    The skill provides templates and instructions for building Apify Actors. It includes security best practices for handling secrets and untrusted data. One example uses a remote script execution pattern for tool installation in a Dockerfile, which is a potential security risk.

  • Socket1d

    No alerts

  • Snyk1d

    Risk: LOW · No issues

  • Runlayer7mo

    2/5 files flagged

  • ZeroLeaks5mo

    1 finding · Score: 82/100

Signed by skilld at 16bac29. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 2 weeks ago

README badge

README badge for apify/agent-skills/apify-actorization