Terraform部署EKS节点组时用户数据未传入启动模板求助
解决方案:通过Terraform为EKS托管节点组配置kubelet资源预留
问题根源
你当前的配置存在两个核心问题:
use_custom_launch_template = false导致EKS使用默认启动模板,不会注入你的自定义脚本pre_bootstrap_user_data/post_bootstrap_user_data是在默认bootstrap脚本前后执行代码,而非替换默认调用——默认bootstrap仍会以原生参数启动kubelet,覆盖你的自定义配置
推荐方案一:使用模块内置参数传递kubelet配置(简洁高效)
借助Terraform AWS EKS模块的bootstrap_extra_args参数,直接向默认bootstrap脚本传递自定义kubelet参数,无需编写完整脚本:
module "eks" { ................ eks_managed_node_group_defaults = { instance_types = ["m6i.large", "m5.large", "m5n.large", "m5zn.large"] disk_size = 100 } eks_managed_node_groups = { eks-ng1 = { min_size = 2 max_size = 12 desired_size = 2 instance_types = ["m6i.2xlarge", "m7i.2xlarge", "m6a.2xlarge", "c6i.2xlarge"] capacity_type = "SPOT" use_custom_launch_template = true # 必须启用自定义启动模板 create_launch_template = true # 让模块创建带自定义配置的启动模板 disk_size = 100 labels = { "node-managed-by" = "eks-ng1" } iam_role_additional_policies = { AmazonSSMManagedInstanceCore = "arn:aws:iam::aws:policy/AmazonSSMManagedInstanceCore" } ami_id = "ami-xxxxxxxxxx" enable_bootstrap_user_data = true # 直接传递kubelet额外参数给默认bootstrap脚本 bootstrap_extra_args = "--kubelet-extra-args '--max-pods=40 --system-reserved cpu=500m,memory=500Mi,ephemeral-storage=1Gi --kube-reserved cpu=1000m,memory=1024Mi,ephemeral-storage=3Gi'" } } tags = { "karpenter.sh/discovery" = local.cluster_name } }
方案二:完全自定义UserData(适合复杂场景)
如果需要在bootstrap前后添加额外操作(如安装依赖、修改系统配置),可以完全自定义UserData:
1. 创建自定义UserData模板(templates/user-data.tpl)
#!/bin/bash set -xe # 可选:添加pre-bootstrap自定义操作,例如调整系统内核参数 # sysctl -w vm.max_map_count=262144 # 执行EKS bootstrap脚本并传入自定义参数 /etc/eks/bootstrap.sh ${cluster_name} \ --kubelet-extra-args "--max-pods=40 --system-reserved cpu=500m,memory=500Mi,ephemeral-storage=1Gi --kube-reserved cpu=1000m,memory=1024Mi,ephemeral-storage=3Gi" # 可选:添加post-bootstrap自定义操作,例如重启监控服务 # systemctl restart prometheus-node-exporter
2. 修改Terraform配置
module "eks" { ................ eks_managed_node_group_defaults = { instance_types = ["m6i.large", "m5.large", "m5n.large", "m5zn.large"] disk_size = 100 } eks_managed_node_groups = { eks-ng1 = { min_size = 2 max_size = 12 desired_size = 2 instance_types = ["m6i.2xlarge", "m7i.2xlarge", "m6a.2xlarge", "c6i.2xlarge"] capacity_type = "SPOT" use_custom_launch_template = true create_launch_template = true disk_size = 100 labels = { "node-managed-by" = "eks-ng1" } iam_role_additional_policies = { AmazonSSMManagedInstanceCore = "arn:aws:iam::aws:policy/AmazonSSMManagedInstanceCore" } ami_id = "ami-xxxxxxxxxx" enable_bootstrap_user_data = false # 禁用默认bootstrap,使用自定义UserData # 编码并注入自定义UserData user_data = base64encode(templatefile("templates/user-data.tpl", { cluster_name = local.cluster_name })) } } tags = { "karpenter.sh/discovery" = local.cluster_name } }
验证步骤
- 执行
terraform apply更新配置 - 替换现有节点(旧节点不会自动应用新配置,可通过缩容再扩容或节点组滚动更新实现)
- 登录新节点验证:
- 查看kubelet参数:
ps aux | grep kubelet,确认包含自定义参数 - 检查节点可分配资源:
kubectl get nodes <node-name> -o jsonpath='{.status.allocatable}',确认CPU、内存、Pods数量符合预期 - 查看cloud-init日志:
cat /var/log/cloud-init-output.log,确认自定义脚本已执行
- 查看kubelet参数:
内容的提问来源于stack exchange,提问作者Ajay Renganathan
相关产品推荐
相关产品推荐

