You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AWS ASG生命周期钩子+Lambda+Jenkins多实例重复触发问题求助

问题:ASG扩容时Lambda重复触发导致Jenkins任务重复执行

背景与期望流程

  • ASG期望状态变更,启动新实例
  • 自动扩缩容事件传入AWS EventBridge默认事件总线,通过规则筛选来源为aws.autoscaling、类型为EC2 Instance-launch Lifecycle Action的事件
  • 将事件负载转发至Lambda函数
  • Python Lambda函数完成:发现实例私有IP、调用Jenkins远程API传入实例ID和IP、根据Jenkins调用结果通知ASG生命周期钩子执行CONTINUE或ABANDON

当前配置信息

1. 生命周期钩子配置

{
    "LifecycleHooks": [
        {
            "LifecycleHookName": "LogAutoScalingEventHook",
            "AutoScalingGroupName": "prod-use1-test",
            "LifecycleTransition": "autoscaling:EC2_INSTANCE_LAUNCHING",
            "HeartbeatTimeout": 300,
            "GlobalTimeout": 30000,
            "DefaultResult": "ABANDON"
        }
    ]
}

2. EventBridge规则模式

{
    "Name": "LogAutoScalingEventRule",
    "Arn": "arn:aws:events:us-east-1:xxxxxxxxxxxx:rule/LogAutoScalingEventRule",
    "EventPattern": "{\"source\":[\"aws.autoscaling\"],\"detail-type\":[\"EC2 Instance-launch Lifecycle Action\"]}",
    "State": "ENABLED",
    "EventBusName": "default",
    "CreatedBy": "xxxxxxxxxxxx"
}

3. EventBridge规则目标

{
    "Targets": [
        {
            "Id": "Id61f23d81-3a04-403b-bb73-9ea5f8b8e4d8",
            "Arn": "arn:aws:lambda:us-east-1:xxxxxxxxxxxx:function:jenkins-function"
        }
    ]
}

4. Lambda基础设施配置

{
    "Configuration": {
        "FunctionName": "jenkins-function",
        "FunctionArn": "arn:aws:lambda:us-east-1:xxxxxxxxxxxx:function:jenkins-function",
        "Runtime": "python3.9",
        "Role": "arn:aws:iam::xxxxxxxxxxxx:role/service-role/HOME_LogAutoScalingEventRole",
        "Handler": "index.jenkins_handler",
        "CodeSize": 960531,
        "Description": "",
        "Timeout": 3,
        "MemorySize": 128,
        "LastModified": "2023-04-05T15:56:50.713+0000",
        "CodeSha256": "8TNg+VrvSFiDxFCBohUNsBmuertrpqjyqQCqOdOf1ss=",
        "Version": "$LATEST",
        "VpcConfig": {
            "SubnetIds": [
                "subnet-xxxxxxxxxxxxxx"
            ],
            "SecurityGroupIds": [
                "sg-xxxxxxxxxxxxxx"
            ],
            "VpcId": "vpc-xxxxxxxxxxxxxx"
        },
        "Environment": {
            "Variables": {
                "API_TOKEN": "jenkins_token",
                "JENKINS_PORT": "jenkins_port",
                "USERNAME": "jenkins_user",
                "JENKINS_URL": "jenkins_url"
            }
        },
        "TracingConfig": {
            "Mode": "PassThrough"
        },
        "RevisionId": "4118e398-9ef2-41f5-8d67-fd9bbf256cae",
        "State": "Active",
        "LastUpdateStatus": "Successful",
        "PackageType": "Zip",
        "Architectures": [
            "x86_64"
        ],
        "EphemeralStorage": {
            "Size": 512
        }
    },
    "Code": {
        "RepositoryType": "S3",
        "Location": "code_location"
    }
}

5. Lambda函数代码

import logging
import json
import os
import boto3
import requests

logger = logging.getLogger("asg-instance-launch")
logger.setLevel(logging.DEBUG)

RUNTIME_REGION = os.environ['AWS_REGION']
USERNAME = os.environ['USERNAME']
JENKINS_URL = os.environ['JENKINS_URL']
API_TOKEN = os.environ['API_TOKEN']
JENKINS_PORT = os.environ['JENKINS_PORT']

LIFECYCLE_KEY = "LifecycleHookName"
ASG_KEY = "AutoScalingGroupName"
EC2_KEY = "EC2InstanceId"
LIFECYCLE_TOKEN = "LifecycleActionToken"


def jenkins_handler(event, context):
    logger.debug(json.dumps(event, indent=2))

    message = event['detail']
    if LIFECYCLE_KEY in message and ASG_KEY in message:
        logger.debug("Jenkins API call")
        instance_id = message[EC2_KEY]
        instance_ip = discover_instance_ip(instance_id)
        life_cycle_hook = message[LIFECYCLE_KEY]
        auto_scaling_group = message[ASG_KEY]
        life_cycle_token = message[LIFECYCLE_TOKEN]
        url = f"http://{USERNAME}:{API_TOKEN}@{JENKINS_URL}:{JENKINS_PORT}/job/deployment/buildWithParameters?HOST_IP={instance_ip}&HOST_ID={instance_id}"
        print(url)
        response = requests.post(url)
        if response.status_code != 201:
            print(response.status_code)
            result = 'ABANDON'
        result = 'CONTINUE'    
        notify_lifecycle(life_cycle_hook, auto_scaling_group, instance_id, life_cycle_token, result)
    return {}

def notify_lifecycle(life_cycle_hook, auto_scaling_group, instance_id, life_cycle_token, result):
    asg_client = boto3.client('autoscaling', region_name=RUNTIME_REGION)
    try:
        response = asg_client.complete_lifecycle_action(
            LifecycleHookName=life_cycle_hook,
            AutoScalingGroupName=auto_scaling_group,
            LifecycleActionToken=life_cycle_token,
            LifecycleActionResult=result,
            InstanceId=instance_id
        )
        logger.debug(response)
    except Exception as e:
        logger.error(
            "Lifecycle hook notified could not be executed: %s", str(e))
        raise e

def discover_instance_ip(instance_id):
    ec2_client = boto3.resource("ec2", region_name=RUNTIME_REGION)
    instance = ec2_client.Instance(instance_id)
    return instance.private_ip_address

实际现象

  • 增加1个实例:Jenkins任务触发1次,Lambda执行1次,实例正常从Pending:Wait转为InService
  • 增加2个实例:触发4次Jenkins任务,每个实例对应2次Lambda调用
  • 增加3个实例:触发5次Jenkins任务,2个实例各对应2次Lambda调用,1个实例对应1次

解决方案

方案1:基于LifecycleActionToken实现幂等性

每个ASG生命周期钩子事件携带唯一的LifecycleActionToken,通过DynamoDB记录已处理的Token,避免重复执行:

  1. 创建DynamoDB表,主键设为LifecycleActionToken(字符串类型)
  2. 修改Lambda函数,在处理事件前检查Token是否已存在:
    • 若存在,直接返回跳过后续逻辑
    • 若不存在,写入表后执行Jenkins调用和生命周期通知

修改后的Lambda handler示例:

# 新增DynamoDB客户端初始化
dynamodb = boto3.resource('dynamodb', region_name=RUNTIME_REGION)
processed_tokens_table = dynamodb.Table('ASGLifecycleProcessedTokens')

def jenkins_handler(event, context):
    logger.debug(json.dumps(event, indent=2))

    message = event['detail']
    if LIFECYCLE_KEY in message and ASG_KEY in message:
        life_cycle_token = message[LIFECYCLE_TOKEN]
        
        # 检查Token是否已处理
        try:
            response = processed_tokens_table.get_item(Key={'LifecycleActionToken': life_cycle_token})
            if 'Item' in response:
                logger.debug(f"Token {life_cycle_token} already processed, skipping")
                return {}
        except Exception as e:
            logger.error(f"Failed to check processed tokens: {str(e)}")
            raise e
        
        # 标记Token为已处理
        try:
            processed_tokens_table.put_item(Item={'LifecycleActionToken': life_cycle_token})
        except Exception as e:
            logger.error(f"Failed to record processed token: {str(e)}")
            raise e
        
        # 原有逻辑继续执行
        logger.debug("Jenkins API call")
        instance_id = message[EC2_KEY]
        instance_ip = discover_instance_ip(instance_id)
        life_cycle_hook = message[LIFECYCLE_KEY]
        auto_scaling_group = message[ASG_KEY]
        url = f"http://{USERNAME}:{API_TOKEN}@{JENKINS_URL}:{JENKINS_PORT}/job/deployment/buildWithParameters?HOST_IP={instance_ip}&HOST_ID={instance_id}"
        print(url)
        response = requests.post(url)
        # 修复逻辑错误:添加else分支设置CONTINUE
        if response.status_code != 201:
            print(response.status_code)
            result = 'ABANDON'
        else:
            result = 'CONTINUE'    
        notify_lifecycle(life_cycle_hook, auto_scaling_group, instance_id, life_cycle_token, result)
    return {}

方案2:配置EventBridge批量事件处理

调整EventBridge规则的目标配置,开启批量事件发送,将同批次扩容事件批量传递给Lambda,配合幂等检查批量处理:

  1. 修改EventBridge规则的Targets配置,添加批量参数:
{
    "Targets": [
        {
            "Id": "Id61f23d81-3a04-403b-bb73-9ea5f8b8e4d8",
            "Arn": "arn:aws:lambda:us-east-1:xxxxxxxxxxxx:function:jenkins-function",
            "BatchParameters": {
                "BatchSize": 10,
                "MaximumBatchingWindowInSeconds": 10
            }
        }
    ]
}
  1. 修改Lambda函数以支持批量事件处理:
# 初始化DynamoDB客户端(同方案1)
dynamodb = boto3.resource('dynamodb', region_name=RUNTIME_REGION)
processed_tokens_table = dynamodb.Table('ASGLifecycleProcessedTokens')

def jenkins_handler(event, context):
    logger.debug(json.dumps(event, indent=2))
    
    # 处理批量Records
    for record in event.get('Records', []):
        try:
            message = json.loads(record['body'])['detail']
        except (KeyError, json.JSONDecodeError):
            logger.error("Invalid event record format")
            continue
            
        if LIFECYCLE_KEY in message and ASG_KEY in message:
            life_cycle_token = message[LIFECYCLE_TOKEN]
            
            # 幂等检查
            try:
                response = processed_tokens_table.get_item(Key={'LifecycleActionToken': life_cycle_token})
                if 'Item' in response:
                    logger.debug(f"Token {life_cycle_token} already processed, skipping")
                    continue
            except Exception as e:
                logger.error(f"Failed to check processed tokens: {str(e)}")
                raise e
            
            try:
                processed_tokens_table.put_item(Item={'LifecycleActionToken': life_cycle_token})
            except Exception as e:
                logger.error(f"Failed to record processed token: {str(e)}")
                raise e
            
            # 原有处理逻辑
            logger.debug("Jenkins API call")
            instance_id = message[EC2_KEY]
            instance_ip = discover_instance_ip(instance_id)
            life_cycle_hook = message[LIFECYCLE_KEY]
            auto_scaling_group = message[ASG_KEY]
            url = f"http://{USERNAME}:{API_TOKEN}@{JENKINS_URL}:{JENKINS_PORT}/job/deployment/buildWithParameters?HOST_IP={instance_ip}&HOST_ID={instance_id}"
            print(url)
            response = requests.post(url)
            if response.status_code != 201:
                print(response.status_code)
                result = 'ABANDON'
            else:
                result = 'CONTINUE'    
            notify_lifecycle(life_cycle_hook, auto_scaling_group, instance_id, life_cycle_token, result)
    return {}

方案3:修复Lambda逻辑错误

原Lambda代码中存在逻辑漏洞:无论Jenkins调用是否成功,最终都会设置result='CONTINUE',导致调用失败时无法触发ABANDON操作,需修正为:

if response.status_code != 201:
    print(response.status_code)
    result = 'ABANDON'
else:
    result = 'CONTINUE'    

内容的提问来源于stack exchange,提问作者bialy_rb

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 11:42:38