All skills
openai avatar

/cloudflare-deploy

@bf9e226 official
by openaiopenai/skills28k stars
1,891

Deploy applications and infrastructure to Cloudflare using Workers, Pages, and related platform services. Use when the user asks to deploy, host, publish, or set up a project on Cloudflare.

Use this Skill: https://skilld.dev/gh/openai/skills/cloudflare-deploy

This session only. Nothing lands on disk.

referencesr2-data-catalogpatterns.md

≈1.5k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Common Patterns

Practical patterns for R2 Data Catalog with PyIceberg.

PyIceberg Connection

import os
from pyiceberg.catalog.rest import RestCatalog
from pyiceberg.exceptions import NamespaceAlreadyExistsError

catalog = RestCatalog(
    name="r2_catalog",
    warehouse=os.getenv("R2_WAREHOUSE"),      # bucket name
    uri=os.getenv("R2_CATALOG_URI"),          # catalog endpoint
    token=os.getenv("R2_TOKEN"),              # API token
)

# Create namespace (idempotent)
try:
    catalog.create_namespace("default")
except NamespaceAlreadyExistsError:
    pass

Pattern 1: Log Analytics Pipeline

Ingest logs incrementally, query by time/level.

import pyarrow as pa
from datetime import datetime
from pyiceberg.schema import Schema
from pyiceberg.types import NestedField, TimestampType, StringType, IntegerType
from pyiceberg.partitioning import PartitionSpec, PartitionField
from pyiceberg.transforms import DayTransform

# Create partitioned table (once)
schema = Schema(
    NestedField(1, "timestamp", TimestampType(), required=True),
    NestedField(2, "level", StringType(), required=True),
    NestedField(3, "service", StringType(), required=True),
    NestedField(4, "message", StringType(), required=False),
)

partition_spec = PartitionSpec(
    PartitionField(source_id=1, field_id=1000, transform=DayTransform(), name="day")
)

catalog.create_namespace("logs")
table = catalog.create_table(("logs", "app_logs"), schema=schema, partition_spec=partition_spec)

# Append logs (incremental)
data = pa.table({
    "timestamp": [datetime(2026, 1, 27, 10, 30, 0)],
    "level": ["ERROR"],
    "service": ["auth-service"],
    "message": ["Failed login"],
})
table.append(data)

# Query by time + level (leverages partitioning)
scan = table.scan(row_filter="level = 'ERROR' AND day = '2026-01-27'")
errors = scan.to_pandas()

Pattern 2: Time-Travel Queries

from datetime import datetime, timedelta

table = catalog.load_table(("logs", "app_logs"))

# Query specific snapshot
snapshot_id = table.current_snapshot().snapshot_id
data = table.scan(snapshot_id=snapshot_id).to_pandas()

# Query as of timestamp (yesterday)
yesterday_ms = int((datetime.now() - timedelta(days=1)).timestamp() * 1000)
data = table.scan(as_of_timestamp=yesterday_ms).to_pandas()

Pattern 3: Schema Evolution

from pyiceberg.types import StringType

table = catalog.load_table(("users", "profiles"))

with table.update_schema() as update:
    update.add_column("email", StringType(), required=False)
    update.rename_column("name", "full_name")
# Old readers ignore new columns, new readers see nulls for old data

Pattern 4: Partitioned Tables

from pyiceberg.partitioning import PartitionSpec, PartitionField
from pyiceberg.transforms import DayTransform, IdentityTransform

# Partition by day + country
partition_spec = PartitionSpec(
    PartitionField(source_id=1, field_id=1000, transform=DayTransform(), name="day"),
    PartitionField(source_id=2, field_id=1001, transform=IdentityTransform(), name="country"),
)
table = catalog.create_table(("events", "user_events"), schema=schema, partition_spec=partition_spec)

# Queries prune partitions automatically
scan = table.scan(row_filter="country = 'US' AND day = '2026-01-27'")

Pattern 5: Table Maintenance

from datetime import datetime, timedelta

table = catalog.load_table(("logs", "app_logs"))

# Compact → expire → cleanup (in order)
table.rewrite_data_files(target_file_size_bytes=128 * 1024 * 1024)
seven_days_ms = int((datetime.now() - timedelta(days=7)).timestamp() * 1000)
table.expire_snapshots(older_than=seven_days_ms, retain_last=10)
three_days_ms = int((datetime.now() - timedelta(days=3)).timestamp() * 1000)
table.delete_orphan_files(older_than=three_days_ms)

See api.md for detailed parameters.

Pattern 6: Concurrent Writes with Retry

from pyiceberg.exceptions import CommitFailedException
import time

def append_with_retry(table, data, max_retries=3):
    for attempt in range(max_retries):
        try:
            table.append(data)
            return
        except CommitFailedException:
            if attempt == max_retries - 1:
                raise
            time.sleep(2 ** attempt)

Pattern 7: Upsert Simulation

import pandas as pd
import pyarrow as pa

# Read → merge → overwrite (not atomic, use Spark MERGE INTO for production)
existing = table.scan().to_pandas()
new_data = pd.DataFrame({"id": [1, 3], "value": [100, 300]})
merged = pd.concat([existing, new_data]).drop_duplicates(subset=["id"], keep="last")
table.overwrite(pa.Table.from_pandas(merged))

Pattern 8: DuckDB Integration

import duckdb

arrow_table = table.scan().to_arrow()
con = duckdb.connect()
con.register("logs", arrow_table)
result = con.execute("SELECT level, COUNT(*) FROM logs GROUP BY level").fetchdf()

Pattern 9: Monitor Table Health

files = table.scan().plan_files()
avg_mb = sum(f.file_size_in_bytes for f in files) / len(files) / (1024**2)
print(f"Files: {len(files)}, Avg: {avg_mb:.1f}MB, Snapshots: {len(table.snapshots())}")

if avg_mb < 10 or len(files) > 1000:
    print("⚠️ Needs compaction")

Best Practices

Area Guideline
Partitioning Use day/hour for time-series; 100-1000 partitions; avoid high cardinality
File sizes Target 128-512MB; compact when avg <10MB or >10k files
Schema Add columns as nullable (required=False); batch changes
Maintenance Compact high-write daily/weekly; expire snapshots 7-30d; cleanup orphans after
Concurrency Reads automatic; writes to different partitions safe; retry same partition
Performance Filter on partitions; select only needed columns; batch appends 100MB+

Source: SKILL.md on GitHub

2 warnings17d5 checks · Risk SAFE
  • Gen Agent Trust Hub17d

    This skill provides comprehensive guidance for deploying and managing infrastructure on the Cloudflare platform. It includes extensive educational material on secure development practices, such as preventing SQL injection and managing secrets effectively. No malicious patterns or security risks were identified.

  • Socket17d

    2 alerts: gptAnomaly

  • Snyk17d

    Risk: LOW · No issues

  • Runlayer7mo

    310/310 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at bf9e226. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 months ago.

Activeupdated 8 months ago

README badge

README badge for openai/skills/cloudflare-deploy

Deploys applications and infrastructure to Cloudflare's platform, including Workers, Pages, D1, R2, Durable Objects, KV, and other services. Use decision trees to route to the right Cloudflare product based on compute, storage, AI, networking, security, or media needs.

Generated from the current SKILL.md.

Does this skill cover all Cloudflare products?
The skill is a consolidated index covering compute, storage, AI, networking, security, media, and developer tools on Cloudflare. It uses decision trees to route you to the right product reference, then loads detailed guidance for that product.
What authentication is required before deploying?
Run `npx wrangler whoami` to check if authenticated. For local deployment, use `wrangler login` (one-time OAuth). For CI/CD, set the `CLOUDFLARE_API_TOKEN` environment variable.
What should I do if deployment fails due to network issues?
Rerun the deploy with `sandbox_permissions=require_escalated` to grant elevated network access, which is required for outbound requests to Cloudflare during deployment.
How long does a Cloudflare deployment typically take?
Deployments may take several minutes. Use appropriate timeout values in your configuration or CI/CD environment.

Generated from the current SKILL.md. These answers refresh after source changes.