All skills
openshift avatar

/debug-cluster

@88dfa28 official
by openshiftopenshift/hypershift543 stars
576

Provides systematic debugging approaches for HyperShift hosted-cluster issues. Auto-applies when debugging cluster problems, investigating stuck deletions, or troubleshooting control plane issues.

Use this Skill: https://skilld.dev/gh/openshift/hypershift/debug-cluster

This session only. Nothing lands on disk.

aws-troubleshooting.md

≈600 tokens on demand. Your agent reads this file only when SKILL.md points to it.

AWS-Specific HyperShift Troubleshooting

This subskill provides AWS-specific debugging workflows for HyperShift hosted-cluster issues.

When to Use This Subskill

Use this when debugging AWS-specific issues such as:

  • CAPI AWS resource problems (AWSCluster, AWSMachine)
  • AWS infrastructure cleanup issues
  • AWS-specific machine deletion problems

Common AWS-Specific Issues

Issue: Machines stuck in "Deleting" phase with "WaitingForInfrastructureDeletion"

Cause: AWSCluster resource deleted before AWSMachine resources, blocking CAPI controller reconciliation

Root Cause: CAPI AWSMachine controller requires AWSCluster to be "ready" to proceed with deletion

Symptoms:

  • Machines show InfrastructureReady: 1 of 2 completed
  • CAPI logs show: "AWSCluster or AWSManagedControlPlane is not ready yet"
  • AWS instances still running but controller can't terminate them

Resolution:

  1. Gather cluster information:

    INFRA_ID=$(kubectl get hc -n <namespace> <cluster-name> -ojsonpath='{.spec.infraID}')
    BASE_DOMAIN=$(kubectl get hc -n <namespace> <cluster-name> -ojsonpath='{.spec.dns.baseDomain}')
    REGION=$(kubectl get hc -n <namespace> <cluster-name> -ojsonpath='{.spec.platform.aws.region}')
  2. Destroy AWS infrastructure:

    hypershift destroy infra aws \
      --name <cluster-name> \
      --aws-creds <path-to-credentials> \
      --base-domain $BASE_DOMAIN \
      --infra-id $INFRA_ID \
      --region $REGION
  3. Remove AWSMachine finalizers (infrastructure is already gone):

    kubectl patch awsmachine <name> -n <hcp-namespace> -p '{"metadata":{"finalizers":null}}' --type=merge
  4. Verify cleanup progresses: namespace should enter Terminating state, then be deleted

AWS-Specific HyperShift Reinstallation

When reinstalling HyperShift on AWS, you'll need these AWS-specific parameters:

Required Parameters

  • OIDC S3 bucket name
  • AWS credentials path
  • AWS region

AWS-Specific Installation Example

hypershift install \
  --oidc-storage-provider-s3-bucket-name $BUCKET_NAME \
  --oidc-storage-provider-s3-credentials $AWS_CREDS \
  --oidc-storage-provider-s3-region $REGION \
  --enable-defaulting-webhook true

Important: Ensure you use the same S3 bucket, credentials, and region as the original installation to maintain continuity with existing OIDC configurations.

Source: SKILL.md on GitHub

1 warning6mo4 checks · Risk SAFE
  • Gen Agent Trust Hub7mo

    The skill provides legitimate AWS HyperShift troubleshooting documentation and command-line examples. No security threats, malicious patterns, or data exfiltration attempts were detected.

  • Socket6mo

    No alerts

  • Snyk7mo

    Risk: LOW · No issues

  • Runlayer7mo

    2/2 files flagged

Signed by skilld at 88dfa28. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 11 months ago
  • Debugging
  • hypershift
  • openshift
  • kubernetes
  • cluster-management
  • capi
  • hosted-clusters
  • operators
  • troubleshooting

README badge

README badge for openshift/hypershift/debug-cluster

Provides systematic debugging approaches for HyperShift hosted-cluster issues, including stuck deletions, control plane problems, and NodePool lifecycle failures. Covers resource hierarchy, operator logs, finalizer inspection, and common failure modes with resolution steps for AWS and provider-agnostic scenarios.

Generated from the current SKILL.md.

Does this skill cover all HyperShift platforms or just specific cloud providers?
The skill provides provider-agnostic debugging workflows for common HyperShift issues. It includes a reference to an AWS-specific subskill for provider-specific troubleshooting, but the main workflows apply across platforms.
What scenarios does this skill help with?
It covers hosted-cluster deletion issues, stuck resources and finalizers, control plane problems, NodePool lifecycle issues, and missing or corrupted HyperShift CRDs requiring reinstallation.
Can this skill help me reinstall HyperShift if CRDs are corrupted?
Yes. The skill provides reinstallation guidance, but notes that AI assistants should guide users through the steps without executing the commands themselves. Reinstallation is positioned as a last resort that causes downtime.
Does this skill provide commands I can run directly?
Yes. The skill includes kubectl commands for inspecting resource status, checking finalizers, reviewing operator logs, and verifying installation state. It also includes bash one-liners for common inspection tasks.

Generated from the current SKILL.md. These answers refresh after source changes.