AWS ASG生命周期钩子+Lambda+Jenkins多实例重复触发问题求助
问题:ASG扩容时Lambda重复触发导致Jenkins任务重复执行
背景与期望流程
- ASG期望状态变更,启动新实例
- 自动扩缩容事件传入AWS EventBridge默认事件总线,通过规则筛选来源为
aws.autoscaling、类型为EC2 Instance-launch Lifecycle Action的事件 - 将事件负载转发至Lambda函数
- Python Lambda函数完成:发现实例私有IP、调用Jenkins远程API传入实例ID和IP、根据Jenkins调用结果通知ASG生命周期钩子执行
CONTINUE或ABANDON
当前配置信息
1. 生命周期钩子配置
{ "LifecycleHooks": [ { "LifecycleHookName": "LogAutoScalingEventHook", "AutoScalingGroupName": "prod-use1-test", "LifecycleTransition": "autoscaling:EC2_INSTANCE_LAUNCHING", "HeartbeatTimeout": 300, "GlobalTimeout": 30000, "DefaultResult": "ABANDON" } ] }
2. EventBridge规则模式
{ "Name": "LogAutoScalingEventRule", "Arn": "arn:aws:events:us-east-1:xxxxxxxxxxxx:rule/LogAutoScalingEventRule", "EventPattern": "{\"source\":[\"aws.autoscaling\"],\"detail-type\":[\"EC2 Instance-launch Lifecycle Action\"]}", "State": "ENABLED", "EventBusName": "default", "CreatedBy": "xxxxxxxxxxxx" }
3. EventBridge规则目标
{ "Targets": [ { "Id": "Id61f23d81-3a04-403b-bb73-9ea5f8b8e4d8", "Arn": "arn:aws:lambda:us-east-1:xxxxxxxxxxxx:function:jenkins-function" } ] }
4. Lambda基础设施配置
{ "Configuration": { "FunctionName": "jenkins-function", "FunctionArn": "arn:aws:lambda:us-east-1:xxxxxxxxxxxx:function:jenkins-function", "Runtime": "python3.9", "Role": "arn:aws:iam::xxxxxxxxxxxx:role/service-role/HOME_LogAutoScalingEventRole", "Handler": "index.jenkins_handler", "CodeSize": 960531, "Description": "", "Timeout": 3, "MemorySize": 128, "LastModified": "2023-04-05T15:56:50.713+0000", "CodeSha256": "8TNg+VrvSFiDxFCBohUNsBmuertrpqjyqQCqOdOf1ss=", "Version": "$LATEST", "VpcConfig": { "SubnetIds": [ "subnet-xxxxxxxxxxxxxx" ], "SecurityGroupIds": [ "sg-xxxxxxxxxxxxxx" ], "VpcId": "vpc-xxxxxxxxxxxxxx" }, "Environment": { "Variables": { "API_TOKEN": "jenkins_token", "JENKINS_PORT": "jenkins_port", "USERNAME": "jenkins_user", "JENKINS_URL": "jenkins_url" } }, "TracingConfig": { "Mode": "PassThrough" }, "RevisionId": "4118e398-9ef2-41f5-8d67-fd9bbf256cae", "State": "Active", "LastUpdateStatus": "Successful", "PackageType": "Zip", "Architectures": [ "x86_64" ], "EphemeralStorage": { "Size": 512 } }, "Code": { "RepositoryType": "S3", "Location": "code_location" } }
5. Lambda函数代码
import logging import json import os import boto3 import requests logger = logging.getLogger("asg-instance-launch") logger.setLevel(logging.DEBUG) RUNTIME_REGION = os.environ['AWS_REGION'] USERNAME = os.environ['USERNAME'] JENKINS_URL = os.environ['JENKINS_URL'] API_TOKEN = os.environ['API_TOKEN'] JENKINS_PORT = os.environ['JENKINS_PORT'] LIFECYCLE_KEY = "LifecycleHookName" ASG_KEY = "AutoScalingGroupName" EC2_KEY = "EC2InstanceId" LIFECYCLE_TOKEN = "LifecycleActionToken" def jenkins_handler(event, context): logger.debug(json.dumps(event, indent=2)) message = event['detail'] if LIFECYCLE_KEY in message and ASG_KEY in message: logger.debug("Jenkins API call") instance_id = message[EC2_KEY] instance_ip = discover_instance_ip(instance_id) life_cycle_hook = message[LIFECYCLE_KEY] auto_scaling_group = message[ASG_KEY] life_cycle_token = message[LIFECYCLE_TOKEN] url = f"http://{USERNAME}:{API_TOKEN}@{JENKINS_URL}:{JENKINS_PORT}/job/deployment/buildWithParameters?HOST_IP={instance_ip}&HOST_ID={instance_id}" print(url) response = requests.post(url) if response.status_code != 201: print(response.status_code) result = 'ABANDON' result = 'CONTINUE' notify_lifecycle(life_cycle_hook, auto_scaling_group, instance_id, life_cycle_token, result) return {} def notify_lifecycle(life_cycle_hook, auto_scaling_group, instance_id, life_cycle_token, result): asg_client = boto3.client('autoscaling', region_name=RUNTIME_REGION) try: response = asg_client.complete_lifecycle_action( LifecycleHookName=life_cycle_hook, AutoScalingGroupName=auto_scaling_group, LifecycleActionToken=life_cycle_token, LifecycleActionResult=result, InstanceId=instance_id ) logger.debug(response) except Exception as e: logger.error( "Lifecycle hook notified could not be executed: %s", str(e)) raise e def discover_instance_ip(instance_id): ec2_client = boto3.resource("ec2", region_name=RUNTIME_REGION) instance = ec2_client.Instance(instance_id) return instance.private_ip_address
实际现象
- 增加1个实例:Jenkins任务触发1次,Lambda执行1次,实例正常从
Pending:Wait转为InService - 增加2个实例:触发4次Jenkins任务,每个实例对应2次Lambda调用
- 增加3个实例:触发5次Jenkins任务,2个实例各对应2次Lambda调用,1个实例对应1次
解决方案
方案1:基于LifecycleActionToken实现幂等性
每个ASG生命周期钩子事件携带唯一的LifecycleActionToken,通过DynamoDB记录已处理的Token,避免重复执行:
- 创建DynamoDB表,主键设为
LifecycleActionToken(字符串类型) - 修改Lambda函数,在处理事件前检查Token是否已存在:
- 若存在,直接返回跳过后续逻辑
- 若不存在,写入表后执行Jenkins调用和生命周期通知
修改后的Lambda handler示例:
# 新增DynamoDB客户端初始化 dynamodb = boto3.resource('dynamodb', region_name=RUNTIME_REGION) processed_tokens_table = dynamodb.Table('ASGLifecycleProcessedTokens') def jenkins_handler(event, context): logger.debug(json.dumps(event, indent=2)) message = event['detail'] if LIFECYCLE_KEY in message and ASG_KEY in message: life_cycle_token = message[LIFECYCLE_TOKEN] # 检查Token是否已处理 try: response = processed_tokens_table.get_item(Key={'LifecycleActionToken': life_cycle_token}) if 'Item' in response: logger.debug(f"Token {life_cycle_token} already processed, skipping") return {} except Exception as e: logger.error(f"Failed to check processed tokens: {str(e)}") raise e # 标记Token为已处理 try: processed_tokens_table.put_item(Item={'LifecycleActionToken': life_cycle_token}) except Exception as e: logger.error(f"Failed to record processed token: {str(e)}") raise e # 原有逻辑继续执行 logger.debug("Jenkins API call") instance_id = message[EC2_KEY] instance_ip = discover_instance_ip(instance_id) life_cycle_hook = message[LIFECYCLE_KEY] auto_scaling_group = message[ASG_KEY] url = f"http://{USERNAME}:{API_TOKEN}@{JENKINS_URL}:{JENKINS_PORT}/job/deployment/buildWithParameters?HOST_IP={instance_ip}&HOST_ID={instance_id}" print(url) response = requests.post(url) # 修复逻辑错误:添加else分支设置CONTINUE if response.status_code != 201: print(response.status_code) result = 'ABANDON' else: result = 'CONTINUE' notify_lifecycle(life_cycle_hook, auto_scaling_group, instance_id, life_cycle_token, result) return {}
方案2:配置EventBridge批量事件处理
调整EventBridge规则的目标配置,开启批量事件发送,将同批次扩容事件批量传递给Lambda,配合幂等检查批量处理:
- 修改EventBridge规则的Targets配置,添加批量参数:
{ "Targets": [ { "Id": "Id61f23d81-3a04-403b-bb73-9ea5f8b8e4d8", "Arn": "arn:aws:lambda:us-east-1:xxxxxxxxxxxx:function:jenkins-function", "BatchParameters": { "BatchSize": 10, "MaximumBatchingWindowInSeconds": 10 } } ] }
- 修改Lambda函数以支持批量事件处理:
# 初始化DynamoDB客户端(同方案1) dynamodb = boto3.resource('dynamodb', region_name=RUNTIME_REGION) processed_tokens_table = dynamodb.Table('ASGLifecycleProcessedTokens') def jenkins_handler(event, context): logger.debug(json.dumps(event, indent=2)) # 处理批量Records for record in event.get('Records', []): try: message = json.loads(record['body'])['detail'] except (KeyError, json.JSONDecodeError): logger.error("Invalid event record format") continue if LIFECYCLE_KEY in message and ASG_KEY in message: life_cycle_token = message[LIFECYCLE_TOKEN] # 幂等检查 try: response = processed_tokens_table.get_item(Key={'LifecycleActionToken': life_cycle_token}) if 'Item' in response: logger.debug(f"Token {life_cycle_token} already processed, skipping") continue except Exception as e: logger.error(f"Failed to check processed tokens: {str(e)}") raise e try: processed_tokens_table.put_item(Item={'LifecycleActionToken': life_cycle_token}) except Exception as e: logger.error(f"Failed to record processed token: {str(e)}") raise e # 原有处理逻辑 logger.debug("Jenkins API call") instance_id = message[EC2_KEY] instance_ip = discover_instance_ip(instance_id) life_cycle_hook = message[LIFECYCLE_KEY] auto_scaling_group = message[ASG_KEY] url = f"http://{USERNAME}:{API_TOKEN}@{JENKINS_URL}:{JENKINS_PORT}/job/deployment/buildWithParameters?HOST_IP={instance_ip}&HOST_ID={instance_id}" print(url) response = requests.post(url) if response.status_code != 201: print(response.status_code) result = 'ABANDON' else: result = 'CONTINUE' notify_lifecycle(life_cycle_hook, auto_scaling_group, instance_id, life_cycle_token, result) return {}
方案3:修复Lambda逻辑错误
原Lambda代码中存在逻辑漏洞:无论Jenkins调用是否成功,最终都会设置result='CONTINUE',导致调用失败时无法触发ABANDON操作,需修正为:
if response.status_code != 201: print(response.status_code) result = 'ABANDON' else: result = 'CONTINUE'
内容的提问来源于stack exchange,提问作者bialy_rb
相关产品推荐
相关产品推荐

