---
title: "skill by agenticluke · skilld"
canonical_url: "https://skilld.dev/gh/agenticluke/zero-downtime-deploys-plus"
meta:
  description: "Plan safe deploys for Kubernetes, Docker, Vercel, and cloud hosts. Covers rolling, blue-green, and canary deploys, health checks, rollback, scaling, and… From agenticluke/zero-downtime-deploys-plus."
  "og:description": "Plan safe deploys for Kubernetes, Docker, Vercel, and cloud hosts. Covers rolling, blue-green, and canary deploys, health checks, rollback, scaling, and… From agenticluke/zero-downtime-deploys-plus."
  "og:title": "skill by agenticluke"
  "twitter:description": "Plan safe deploys for Kubernetes, Docker, Vercel, and cloud hosts. Covers rolling, blue-green, and canary deploys, health checks, rollback, scaling, and… From agenticluke/zero-downtime-deploys-plus."
  "twitter:title": "skill by agenticluke"
---

`

[All skills](https://skilld.dev/skills)

[![agenticluke avatar](https://skilld.dev/_img/avatar?url=https%3A%2F%2Fgithub.com%2Fagenticluke.png%3Fsize%3D96)](https://skilld.dev/gh/agenticluke)

# **/skill**

[@e7329f8](https://github.com/agenticluke/zero-downtime-deploys-plus/commit/e7329f826399e8c8a28c8dd5a457c9345a419a9a "Your agent reads SKILL.md at commit e7329f8")

by [agenticluke](https://skilld.dev/gh/agenticluke)· [agenticluke](https://skilld.dev/gh/agenticluke)/ [zero-downtime-deploys-plus](https://skilld.dev/gh/agenticluke/zero-downtime-deploys-plus)

Plan safe deploys for Kubernetes, Docker, Vercel, and cloud hosts. Covers rolling, blue-green, and canary deploys, health checks, rollback, scaling, and updates with no downtime.

- 1 file
- 6.5 KB
- Updated 3 weeks ago
- [GitHub](https://github.com/agenticluke/zero-downtime-deploys-plus/blob/main/skill/SKILL.md "View SKILL.md on GitHub")

## SKILL.md

6.5 KB

**≈46** tokens always: the name and description. **≈1.6k** when used: this file.

## Deployment Patterns

This skill comes from ECC. Full credit goes to the original ECC author.

Use this skill to plan and check safe app deploys.

### When to Use

Use this skill when you need to:

- Deploy an app to Kubernetes, Docker, Vercel, or a cloud host.
- Update an app with no planned downtime.
- Pick a rolling, blue-green, or canary deploy.
- Add health checks.
- Set up auto scaling.
- Make a rollback plan.

### Before You Deploy

Check these facts first:

- How much downtime is allowed?
- Can the old and new app run at the same time?
- Will the new app change the database?
- How will you know the new app is healthy?
- How fast must you roll back?
- Can the host send only some users to the new app?

Do not claim zero downtime unless the full path supports it. This includes the app, database, network, and host.

### Pick a Deploy Plan

#### 1. Rolling Deploy

Replace old app copies a few at a time.

Use it when:

- Old and new copies can run at the same time.
- The app has more than one copy.
- A slow update is safe.

Kubernetes example:

```
spec:
  replicas: 3
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxSurge: 1
      maxUnavailable: 0
```

`maxSurge: 1` lets one extra pod start.

`maxUnavailable: 0` keeps all needed pods ready during the update.

Watch for:

- One app copy is not enough for no downtime.
- A bad health check may send work to a pod too soon.
- Old and new code may use the database in different ways.
- Long jobs may stop when an old pod shuts down.

#### 2. Blue-Green Deploy

Run two full sites:

- Blue is the live site.
- Green is the new site.

Test green first. Then send users to green.

Use it when:

- You need a fast switch.
- You need a fast rollback.
- You can pay for two full sites for a short time.

Watch for:

- Both sites may write to the same database.
- User sessions may break after the switch.
- Cache data may not match.
- Old jobs may still run after traffic moves.
- Database changes may make rollback unsafe.

Keep blue ready until green has been stable for the set check time.

#### 3. Canary Deploy

Send a small part of traffic to the new app first. Raise the share in steps.

A simple plan:

1. Send 5% of traffic to the new app.
2. Check errors, speed, and key user tasks.
3. Raise traffic to 25%.
4. Check again.
5. Raise traffic to 50%.
6. Move to 100% only if all checks pass.

Use it when:

- Your host can split traffic.
- You can track the new app by version.
- A small test group is safe.

Watch for:

- A small group may miss rare bugs.
- Signed-in users may need to stay on one version.
- Workers and timed jobs do not follow web traffic rules.
- Stop the rollout if you cannot tell old and new results apart.

#### 4. Replace All at Once

Stop the old app. Then start the new app.

Use it only when:

- Downtime is allowed.
- The app is small or used for tests.
- Other plans cost too much or add too much work.

State the planned downtime before using this plan.

### Health Checks

Use separate checks when the host supports them:

- **Start check:** Has the app finished starting?
- **Ready check:** Can the app take new work?
- **Live check:** Is the app stuck and in need of a restart?

A health check should:

- Return fast.
- Have a short wait limit.
- Test what users need.
- Fail when the app cannot serve safe work.

Do not make every check call many other services. One slow service could restart healthy app copies again and again.

Give slow apps enough start time. Do not send traffic before they are ready.

### Safe Database Changes

Use changes that work with both old and new code.

A safe order is:

1. Add the new table or field.
2. Deploy code that can use both old and new data.
3. Move or fill old data in small parts.
4. Deploy code that uses the new data.
5. Remove old data only after rollback is no longer needed.

Do not remove or rename a field before all old app copies are gone.

Back up key data before a risky change. Test that the backup can be restored.

### Shutdown and Long Jobs

Before stopping an old app copy:

- Stop sending it new work.
- Let open requests finish.
- Save or move long jobs.
- Set a clear shutdown time.
- Make repeated work safe when possible.

Make sure queue workers do not take the same job twice during a deploy.

### Auto Scaling

When auto scaling is used:

- Set a safe low and high copy count.
- Keep enough copies for a rolling deploy.
- Test sudden traffic growth.
- Check host limits.
- Do not scale from CPU alone if queue size or wait time matters more.
- Make sure new copies can start fast enough.

A deploy and a traffic spike can happen at the same time. Plan for both.

### Rollback Rules

Write rollback rules before the deploy.

Roll back when:

- Error rate rises past the set limit.
- Requests become too slow.
- Health checks keep failing.
- A key user task fails.
- Data is lost or changed in a bad way.

A code rollback may not undo a database change. Write a separate data repair plan when needed.

Keep the last good app image and settings until the new release is stable.

### Concrete Example

Task: Deploy version `2.4.0` to Kubernetes with no planned downtime.

Plan:

1. Run at least three app pods.
2. Add start, ready, and live checks.
3. Confirm versions `2.3.0` and `2.4.0` can use the same database.
4. Use `maxSurge: 1` and `maxUnavailable: 0`.
5. Start one new pod.
6. Wait until its ready check passes.
7. Send it traffic.
8. Replace the next old pod.
9. Watch errors, speed, restarts, and key user tasks.
10. Stop if a set limit fails.
11. Roll back to `2.3.0` if the fault does not clear.
12. Keep old database fields until rollback time has passed.

Success means all new pods are ready, key tasks work, and error and speed limits stay normal.

### Final Check

- The deploy plan fits the host.
- More than one app copy runs when no downtime is needed.
- Start, ready, and live checks work.
- Logs show the app version.
- Error and speed checks are ready.
- Key user tasks are tested.
- Old and new code can run together.
- Database changes are safe to roll out.
- Sessions, caches, jobs, and workers are covered.
- Auto scaling has safe limits.
- Stop and rollback rules are clear.
- The last good app image is kept.
- The plan was tested in a setup like the live site.
- The team knows who can stop or roll back the deploy.

Source: [SKILL.md on GitHub](https://github.com/agenticluke/zero-downtime-deploys-plus/blob/main/skill/SKILL.md)

## Third-party checks

No third-party reports yet.

## Provenance

[Signed by skilld at e7329f8.](https://github.com/agenticluke/zero-downtime-deploys-plus/commit/e7329f826399e8c8a28c8dd5a457c9345a419a9a "e7329f826399e8c8a28c8dd5a457c9345a419a9a") This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 3 weeks ago.

Activeupdated 3 weeks ago

## Capability

<dl>

<dt>origin</dt>
<dd>ECC</dd>

</dl>

## README badge

![README badge for agenticluke/zero-downtime-deploys-plus](https://skilld.dev/b/agenticluke/zero-downtime-deploys-plus?theme=light&label=0)

## Related skills

-
-
-
-
-
-