You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在K8s集群启用EC2 Cluster Autoscaler时能否复用旧EBS卷?

Can Longhorn reuse old AWS EBS volumes when Cluster Autoscaler replaces failed EC2 instances?

Absolutely, you can reuse the existing EBS volumes left behind by failed EC2 instances with Longhorn—this aligns perfectly with your goal of separating compute and storage, similar to how Portworx handles this scenario. Here's how to set this up and what to keep in mind:

Key Configuration Steps

1. Ensure EBS volumes aren’t automatically deleted

First, verify your Longhorn StorageClass uses the Retain reclaim policy. This ensures that when a PVC is deleted (e.g., due to node failure), the underlying EBS volume isn’t destroyed automatically.

Check your StorageClass with this command:

kubectl get sc longhorn -o yaml

Look for the reclaimPolicy field—if it’s set to Delete, update it to Retain (you can edit the SC directly with kubectl edit sc longhorn).

2. Align Cluster Autoscaler and ASG settings

When a node fails, Cluster Autoscaler will remove the faulty node from the cluster, and your ASG will spin up a replacement. To reuse old EBS volumes:

  • Keep replacement instances in the same AZ: AWS EBS volumes can’t be mounted across Availability Zones, so configure your ASG to launch new instances in the same AZ as the failed node.
  • Match node labels/taints: Ensure new instances have the same node labels and tolerations as the failed node. This helps Longhorn correctly associate the old volumes with the new node.
  • Grant necessary IAM permissions: The EC2 instance role needs permissions like ec2:AttachVolume, ec2:DescribeVolumes, and ec2:ModifyVolume to interact with EBS volumes.

3. Optimize Longhorn’s failure recovery settings

Tweak Longhorn’s settings to prioritize reusing existing volumes over rebuilding replicas:

  • In the Longhorn UI, go to Settings > Node and adjust the Replica Rebuild Wait Interval to a longer value (e.g., 30 minutes). This gives you time to mount old volumes before Longhorn starts rebuilding replicas from scratch.
  • Enable Allow Node Drain With Failed Disks to make Longhorn more flexible when handling volumes from failed nodes.

Unlike Portworx’s out-of-the-box integration with ASG, Longhorn doesn’t automatically reattach old volumes to new nodes. You can build a lightweight automation layer to handle this:

  • Create a Kubernetes controller or a script that listens for node deletion events.
  • For each deleted node, fetch the Longhorn volumes/replicas that were hosted on it, then retrieve their corresponding EBS volume IDs.
  • Once the replacement instance is up, attach those EBS volumes to the new node and notify Longhorn to re-recognize the replicas.

Important Limitations

  • AZ Lock: As mentioned, EBS volumes are AZ-specific—your ASG must respect this to reuse volumes.
  • Encrypted Volumes: If your EBS volumes are encrypted, the replacement instance needs access to the KMS key used for encryption.
  • Volume Cleanup: If you don’t reuse a volume, remember to manually delete it later (since we set Retain policy) to avoid unnecessary AWS costs.

内容的提问来源于stack exchange,提问作者breizh5729

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 17:12:43