All skills
wshobson avatar

/python-background-jobs

@be57c0b
by Seth Hobsonwshobson/agents40k stars
4,281

Python background job patterns including task queues, workers, and event-driven architecture. Use when implementing async task processing, job queues, long-running operations, or decoupling work from request/response cycles.

Use this Skill: https://skilld.dev/gh/wshobson/agents/python-background-jobs

This session only. Nothing lands on disk.

referencesdetails.md

≈858 tokens on demand. Your agent reads this file only when SKILL.md points to it.

python-background-jobs — detailed worked examples

Advanced Patterns

Pattern 5: Dead Letter Queue

Handle permanently failed tasks for manual inspection.

@app.task(bind=True, max_retries=3)
def process_webhook(self, webhook_id: str, payload: dict) -> None:
    """Process webhook with DLQ for failures."""
    try:
        result = send_webhook(payload)
        if not result.success:
            raise WebhookFailedError(result.error)
    except Exception as e:
        if self.request.retries >= self.max_retries:
            # Move to dead letter queue for manual inspection
            dead_letter_queue.send({
                "task": "process_webhook",
                "webhook_id": webhook_id,
                "payload": payload,
                "error": str(e),
                "attempts": self.request.retries + 1,
                "failed_at": datetime.utcnow().isoformat(),
            })
            logger.error(
                "Webhook moved to DLQ after max retries",
                webhook_id=webhook_id,
                error=str(e),
            )
            return

        # Exponential backoff retry
        raise self.retry(exc=e, countdown=2 ** self.request.retries * 60)

Pattern 6: Status Polling Endpoint

Provide an endpoint for clients to check job status.

from fastapi import FastAPI, HTTPException

app = FastAPI()

@app.get("/jobs/{job_id}")
async def get_job_status(job_id: str) -> JobStatusResponse:
    """Get current status of a background job."""
    job = await jobs_repo.get(job_id)

    if job is None:
        raise HTTPException(404, f"Job {job_id} not found")

    return JobStatusResponse(
        job_id=job.id,
        status=job.status.value,
        created_at=job.created_at,
        started_at=job.started_at,
        completed_at=job.completed_at,
        result=job.result if job.status == JobStatus.SUCCEEDED else None,
        error=job.error if job.status == JobStatus.FAILED else None,
        # Helpful for clients
        is_terminal=job.status in (JobStatus.SUCCEEDED, JobStatus.FAILED),
    )

Pattern 7: Task Chaining and Workflows

Compose complex workflows from simple tasks.

from celery import chain, group, chord

# Simple chain: A → B → C
workflow = chain(
    extract_data.s(source_id),
    transform_data.s(),
    load_data.s(destination_id),
)

# Parallel execution: A, B, C all at once
parallel = group(
    send_email.s(user_email),
    send_sms.s(user_phone),
    update_analytics.s(event_data),
)

# Chord: Run tasks in parallel, then a callback
# Process all items, then send completion notification
workflow = chord(
    [process_item.s(item_id) for item_id in item_ids],
    send_completion_notification.s(batch_id),
)

workflow.apply_async()

Pattern 8: Alternative Task Queues

Choose the right tool for your needs.

RQ (Redis Queue): Simple, Redis-based

from rq import Queue
from redis import Redis

queue = Queue(connection=Redis())
job = queue.enqueue(send_email, "user@example.com", "Subject", "Body")

Dramatiq: Modern Celery alternative

import dramatiq
from dramatiq.brokers.redis import RedisBroker

dramatiq.set_broker(RedisBroker())

@dramatiq.actor
def send_email(to: str, subject: str, body: str) -> None:
    email_client.send(to, subject, body)

Cloud-native options:

  • AWS SQS + Lambda
  • Google Cloud Tasks
  • Azure Functions

Source: SKILL.md on GitHub

No alerts16d5 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    This skill provides educational patterns and best practices for implementing Python background jobs and task queues. It contains standard architectural guidance and code examples that do not present security risks.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: LOW · No issues

  • Runlayer7mo

    1/1 file flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at be57c0b. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 3 days ago.

Activeupdated 4 months ago
  • Python
  • celery
  • task-queue
  • background-jobs
  • async
  • redis
  • idempotency
  • job-state
  • event-driven

README badge

README badge for wshobson/agents/python-background-jobs

Provides patterns for decoupling long-running work from request/response cycles using task queues and background workers. Examples use Celery with Redis, covering job state management, idempotency, retries, and at-least-once delivery semantics.

Generated from the current SKILL.md.

Does this skill cover specific task queue libraries like Celery, RQ, or Dramatiq?
The skill uses Celery for examples as a widely adopted task queue, but acknowledges that RQ, Dramatiq, and cloud-native solutions like AWS SQS or GCP Tasks are equally valid choices. The patterns are library-agnostic.
How do I handle tasks that might be retried multiple times?
Design tasks to be idempotent using strategies like check-before-write, idempotency keys with external services, upsert patterns, or deduplication windows. The skill includes examples for each approach.
Should I return results immediately to the user or wait for the background job to complete?
Return immediately with a job ID for operations exceeding a few seconds, then let the user poll the job status endpoint. This decouples the request/response cycle from long-running work.
What happens if a worker crashes while processing a task?
Most queues guarantee at-least-once delivery, so tasks will be retried. Your task code must be idempotent to handle re-execution safely, and you should set appropriate soft and hard timeouts.
How do I distinguish between transient errors that should be retried and permanent failures?
The skill recommends using `autoretry_for` for transient errors like connection timeouts, and catching permanent failures separately (e.g. validation errors, declined payments) to avoid wasting retries.

Generated from the current SKILL.md. These answers refresh after source changes.