SkyPilot Troubleshooting Guide
Installation Issues
Cloud credentials not found
Error: sky check shows clouds as disabled
Solutions:
# AWS
aws configure
# Verify: aws sts get-caller-identity
# GCP
gcloud auth application-default login
# Verify: gcloud auth list
# Azure
az login
az account set -s <subscription-id>
# Kubernetes
export KUBECONFIG=~/.kube/config
kubectl get nodes
# Re-check after configuration
sky checkPermission errors
Error: PermissionError or AccessDenied
Solutions:
# AWS: Ensure IAM permissions include EC2, S3, IAM
# Required policies: AmazonEC2FullAccess, AmazonS3FullAccess, IAMFullAccess
# GCP: Ensure roles include Compute Admin, Storage Admin
gcloud projects add-iam-policy-binding PROJECT_ID \
--member="user:email@example.com" \
--role="roles/compute.admin"
# Azure: Ensure Contributor role on subscription
az role assignment create \
--assignee email@example.com \
--role Contributor \
--scope /subscriptions/SUBSCRIPTION_IDCluster Launch Issues
Quota exceeded
Error: Quota exceeded for resource
Solutions:
# Try different region
resources:
accelerators: A100:8
any_of:
- cloud: gcp
region: us-west1
- cloud: gcp
region: europe-west4
- cloud: aws
region: us-east-1
# Or request quota increase from cloud provider# Check quota before launching
sky show-gpus --cloud gcpGPU not available
Error: No resources available in region
Solutions:
# Use fallback accelerators
resources:
accelerators:
H100: 8
A100-80GB: 8
A100: 8
any_of:
- cloud: gcp
- cloud: aws
- cloud: azure# Check GPU availability
sky show-gpus A100
sky show-gpus --cloud awsInstance type not found
Error: Instance type 'xyz' not found
Solutions:
# Let SkyPilot choose instance automatically
resources:
accelerators: A100:8
cpus: 96+
memory: 512+
# Don't specify instance_type unless necessaryCluster stuck in INIT
Error: Cluster stays in INIT state
Solutions:
# Check cluster logs
sky logs mycluster --status
# SSH and check manually
ssh mycluster
journalctl -u sky-supervisor
# Terminate and retry
sky down mycluster
sky launch -c mycluster task.yamlSetup Command Issues
Setup script fails
Error: Setup commands fail during provisioning
Solutions:
# Add error handling and retries
setup: |
set -e # Exit on error
# Retry pip installs
for i in {1..3}; do
pip install torch transformers && break
echo "Retry $i..."
sleep 10
done
# Verify installation
python -c "import torch; print(torch.__version__)"Conda environment issues
Error: Conda not found or environment issues
Solutions:
setup: |
# Initialize conda for bash
source ~/.bashrc
# Or use full path
~/miniconda3/bin/conda create -n myenv python=3.10 -y
~/miniconda3/bin/conda activate myenvCUDA version mismatch
Error: CUDA driver version is insufficient
Solutions:
setup: |
# Install specific CUDA version
pip install torch==2.1.0+cu121 -f https://download.pytorch.org/whl/torch_stable.html
# Verify CUDA
python -c "import torch; print(torch.cuda.is_available())"Distributed Training Issues
Nodes can't communicate
Error: Connection refused between nodes
Solutions:
run: |
# Debug: Print all node IPs
echo "All nodes: $SKYPILOT_NODE_IPS"
echo "My rank: $SKYPILOT_NODE_RANK"
# Wait for all nodes to be ready
sleep 30
# Use correct master address
MASTER_ADDR=$(echo "$SKYPILOT_NODE_IPS" | head -n1)
echo "Master: $MASTER_ADDR"torchrun fails
Error: torch.distributed errors
Solutions:
run: |
# Ensure correct environment variables
export NCCL_DEBUG=INFO
export NCCL_IB_DISABLE=1 # Try if InfiniBand issues
torchrun \
--nnodes=$SKYPILOT_NUM_NODES \
--nproc_per_node=$SKYPILOT_NUM_GPUS_PER_NODE \
--node_rank=$SKYPILOT_NODE_RANK \
--master_addr=$(echo "$SKYPILOT_NODE_IPS" | head -n1) \
--master_port=12355 \
--rdzv_backend=c10d \
train.pyDeepSpeed hostfile errors
Error: Invalid hostfile or connection errors
Solutions:
run: |
# Create proper hostfile
echo "$SKYPILOT_NODE_IPS" | while read ip; do
echo "$ip slots=$SKYPILOT_NUM_GPUS_PER_NODE"
done > /tmp/hostfile
cat /tmp/hostfile # Debug
deepspeed --hostfile=/tmp/hostfile train.pyFile Mount Issues
Mount fails
Error: Failed to mount storage
Solutions:
# Verify bucket exists and credentials are valid
file_mounts:
/data:
source: s3://my-bucket/data
mode: MOUNT
# Check bucket access
# aws s3 ls s3://my-bucket/Slow file access
Problem: Reading from mount is very slow
Solutions:
# Use COPY mode for small datasets
file_mounts:
/data:
source: s3://bucket/data
mode: COPY # Pre-fetch to local disk
# Use MOUNT_CACHED for outputs
file_mounts:
/outputs:
name: outputs
store: s3
mode: MOUNT_CACHED # Cached writesStorage not persisting
Error: Data lost after cluster restart
Solutions:
# Use named storage (persists across clusters)
file_mounts:
/persistent:
name: my-persistent-storage
store: s3
mode: MOUNT
# Data in ~/sky_workdir is NOT persisted
# Always use file_mounts for persistent dataManaged Job Issues
Job keeps failing
Error: Job fails and doesn't recover
Solutions:
# Enable spot recovery
resources:
use_spot: true
spot_recovery: FAILOVER
# Add retry logic
max_restarts_on_errors: 5
# Implement checkpointing
run: |
python train.py \
--checkpoint-dir /checkpoints \
--resume-from-latestJob stuck in pending
Error: Job stays in PENDING state
Solutions:
# Check job controller status
sky jobs controller status
# View controller logs
sky jobs controller logs
# Restart controller if needed
sky jobs controller restartCheckpoint not resuming
Error: Training restarts from beginning
Solutions:
file_mounts:
/checkpoints:
name: training-checkpoints
store: s3
mode: MOUNT_CACHED
run: |
# Check for existing checkpoint
if [ -d "/checkpoints/latest" ]; then
RESUME_FLAG="--resume /checkpoints/latest"
else
RESUME_FLAG=""
fi
python train.py $RESUME_FLAG --checkpoint-dir /checkpointsSky Serve Issues
Service not accessible
Error: Cannot reach service endpoint
Solutions:
# Check service status
sky serve status my-service
# View replica logs
sky serve logs my-service
# Check readiness probe
sky serve status my-service --endpointReplicas keep crashing
Error: Replicas fail health checks
Solutions:
service:
readiness_probe:
path: /health
initial_delay_seconds: 120 # Increase for slow model loading
period_seconds: 30
timeout_seconds: 10
run: |
# Ensure health endpoint exists
python -c "
from fastapi import FastAPI
app = FastAPI()
@app.get('/health')
def health():
return {'status': 'ok'}
"Autoscaling not working
Problem: Service doesn't scale up/down
Solutions:
service:
replica_policy:
min_replicas: 1
max_replicas: 10
target_qps_per_replica: 2.0
upscale_delay_seconds: 30 # Faster scale up
downscale_delay_seconds: 60 # Faster scale down
# Monitor metrics
# sky serve status my-serviceSSH and Access Issues
Cannot SSH to cluster
Error: Connection refused or timeout
Solutions:
# Verify cluster is running
sky status
# Try with verbose output
ssh -v mycluster
# Check SSH key
ls -la ~/.ssh/sky-key*
# Regenerate SSH key if needed
sky launch -c test --dryrun # Regenerates keyPort forwarding fails
Error: Cannot forward ports
Solutions:
# Correct syntax
ssh -L 8080:localhost:8080 mycluster
# For Jupyter
ssh -L 8888:localhost:8888 mycluster
# Multiple ports
ssh -L 8080:localhost:8080 -L 6006:localhost:6006 myclusterCost and Billing Issues
Unexpected charges
Problem: Higher than expected costs
Solutions:
# Always terminate unused clusters
sky down --all
# Set autostop
sky autostop mycluster -i 30 --down
# Use spot instances
resources:
use_spot: trueSpot instance preempted
Error: Instance terminated unexpectedly
Solutions:
# Use managed jobs for automatic recovery
# sky jobs launch instead of sky launch
resources:
use_spot: true
spot_recovery: FAILOVER # Auto-failover to another region/cloud
# Always checkpoint frequently when using spotDebugging Commands
View cluster state
# Cluster status
sky status
sky status -a # Show all details
# Cluster resources
sky show-gpus
# Cloud credentials
sky checkView logs
# Task logs
sky logs mycluster
sky logs mycluster 1 # Specific job
# Managed job logs
sky jobs logs my-job
sky jobs logs my-job --follow
# Service logs
sky serve logs my-serviceInspect cluster
# SSH to cluster
ssh mycluster
# Check GPU status
nvidia-smi
# Check processes
ps aux | grep python
# Check disk space
df -hCommon Error Messages
| Error | Cause | Solution |
|---|---|---|
No launchable resources |
No available instances | Try different region/cloud |
Quota exceeded |
Cloud quota limit | Request increase or use different cloud |
Setup failed |
Script error | Check logs, add error handling |
Connection refused |
Network/firewall | Check security groups, wait for init |
CUDA OOM |
Out of GPU memory | Use larger GPU or reduce batch size |
Spot preempted |
Spot instance reclaimed | Use managed jobs for auto-recovery |
Mount failed |
Storage access issue | Check credentials and bucket exists |
Getting Help
- Documentation: https://docs.skypilot.co
- GitHub Issues: https://github.com/skypilot-org/skypilot/issues
- Slack: https://slack.skypilot.co
- Examples: https://github.com/skypilot-org/skypilot/tree/master/examples
Reporting Issues
Include:
- SkyPilot version:
sky --version - Python version:
python --version - Cloud provider and region
- Full error traceback
- Task YAML (sanitized)
- Output of
sky check