关闭GKE节点池自动扩缩容后停止实例仍被重建的问题排查
问题描述
我编写了Python脚本,先关闭GKE集群节点池的自动扩缩容,再按可用区停止各MIG(托管实例组)的底层节点。但执行脚本时,停止的实例会立即被重建,状态从UNKNOWN变为READY;而通过控制台手动操作则无此问题。已确认节点池自动扩缩容已关闭,MIG级别的自动扩缩容和健康检查均禁用,仅节点池级存在自动扩缩容,请求排查原因及代码优化建议。
附脚本代码:
import subprocess import json import argparse import time def get_node_pool_name(cluster_name, project_name, region_name): try: cmd = [ 'gcloud', 'container', 'clusters', 'describe', cluster_name, '--project', project_name, '--region', region_name, '--format', 'json' ] ... ... return node_pool_name except Exception as e: print(f"Error occurred while getting node pool name: {str(e)}") raise def disable_autoscaler(cluster_name, project_name, region_name, node_pool_name): try: cmd = [ 'gcloud', 'container', 'node-pools', 'update', node_pool_name, '--cluster', cluster_name, '--project', project_name, '--region', region_name, '--no-enable-autoscaling' ] ... ... except Exception as e: print(f"Error occurred while disabling node pool autoscaler: {str(e)}") raise def get_instance_groups(cluster_name, project_name, region_name): try: cmd = [ 'gcloud', 'container', 'clusters', 'describe', cluster_name, '--project', project_name, '--region', region_name, '--format', 'json' ] ... ... return instance_groups except Exception as e: print(f"Error occurred while getting instance-groups: {str(e)}") raise def get_instances(instance_group_name, project_name, zone): try: cmd = [ 'gcloud', 'compute', 'instance-groups', 'list-instances', instance_group_name, '--project', project_name, '--zone', zone, '--format', 'json' ] ... ... return instance_names except Exception as e: print(f"Error occurred while getting instances: {str(e)}") raise def stop_instances(instance_names, project_name, zone): try: for instance in instance_names: cmd = [ 'gcloud', 'compute', 'instances', 'stop', instance, '--project', project_name, '--zone', zone ] subprocess.run(cmd, check=True) print(f"Instance stopped: {instance} in zone {zone}") except Exception as e: print(f"Error occurred while stopping instances: {str(e)}") raise def main(cluster_name, project_name, region_name): node_pool_name = get_node_pool_name(cluster_name, project_name, region_name) disable_autoscaler(cluster_name, project_name, region_name, node_pool_name) time.sleep(60) instance_groups = get_instance_groups(cluster_name, project_name, region_name) for instance_group, zone in instance_groups: instance_names = get_instances(instance_group, project_name, zone) stop_instances(instance_names, project_name, zone) time.sleep(30) if __name__ == "__main__": parser = argparse.ArgumentParser(description="Disable GKE autoscaler and stop instances.") parser.add_argument('cluster_name', type=str, help='Name of the GKE cluster') parser.add_argument('project_name', type=str, help='Google Cloud project name') parser.add_argument('region_name', type=str, help='Region name of the GKE cluster') args = parser.parse_args() main(args.cluster_name, args.project_name, args.region_name)
排查原因
- GKE配置同步延迟:脚本用固定
time.sleep(60)等待自动扩缩容关闭生效,但GKE的节点池配置变更需要同步到MIG控制平面,这个过程可能超过60秒。即使gcloud命令返回成功,底层MIG的实际规则可能未更新,导致停止实例后仍触发重建。 - 未验证自动扩缩容实际状态:仅执行
--no-enable-autoscaling命令,未确认节点池的自动扩缩容状态确实为禁用,可能存在命令执行成功但配置未落地的情况。 - GKE节点控制器的重建逻辑:GKE节点控制器会持续监控节点状态,若节点被标记为
NotReady,即使MIG自动扩缩容关闭,控制器可能因节点健康检查规则(如节点池默认的健康检测)触发重建。
代码优化建议
1. 轮询验证自动扩缩容状态
替换固定sleep,改为轮询检查节点池的自动扩缩容状态,确保配置真正生效:
def verify_autoscaler_disabled(cluster_name, project_name, region_name, node_pool_name): while True: cmd = [ 'gcloud', 'container', 'node-pools', 'describe', node_pool_name, '--cluster', cluster_name, '--project', project_name, '--region', region_name, '--format', 'value(autoscaling.enabled)' ] result = subprocess.run(cmd, capture_output=True, text=True, check=True) status = result.stdout.strip() if status == 'False': print("Autoscaler confirmed disabled.") break print("Waiting for autoscaler to disable...") time.sleep(10)
在disable_autoscaler函数执行后调用此函数,替代原有的time.sleep(60)。
2. 锁定MIG实例数量
停止实例前,将MIG的最小/最大实例数设置为当前实例数,彻底锁定实例数量,避免意外重建:
def lock_mig_size(instance_group_name, project_name, zone): # 获取当前MIG实例数 cmd_count = [ 'gcloud', 'compute', 'instance-groups', 'managed', 'describe', instance_group_name, '--project', project_name, '--zone', zone, '--format', 'value(targetSize)' ] result = subprocess.run(cmd_count, capture_output=True, text=True, check=True) target_size = int(result.stdout.strip()) # 设置最小/最大实例数为当前值 cmd_set = [ 'gcloud', 'compute', 'instance-groups', 'managed', 'set-autoscaling', instance_group_name, '--project', project_name, '--zone', zone, '--min-num-replicas', str(target_size), '--max-num-replicas', str(target_size), '--no-autoscaling' ] subprocess.run(cmd_set, check=True) print(f"MIG {instance_group_name} size locked to {target_size} instances.")
在stop_instances前调用此函数。
3. 批量停止实例
将单实例循环停止改为批量操作,减少gcloud调用次数,降低延迟:
def stop_instances(instance_names, project_name, zone): try: if not instance_names: print(f"No instances to stop in zone {zone}") return cmd = [ 'gcloud', 'compute', 'instances', 'stop', '--project', project_name, '--zone', zone ] + instance_names subprocess.run(cmd, check=True) print(f"Stopped instances: {', '.join(instance_names)} in zone {zone}") except Exception as e: print(f"Error occurred while stopping instances: {str(e)}") raise
4. 增强日志与错误排查
在所有subprocess.run调用中添加capture_output=True并打印输出,方便定位问题:
# 示例:在disable_autoscaler中添加日志 result = subprocess.run(cmd, capture_output=True, text=True, check=True) print(f"Autoscaler disable command output: {result.stdout}")
内容的提问来源于stack exchange,提问作者Schatt
相关产品推荐
相关产品推荐

