多次触发AWS Lambda仅最后迭代告警及索引错误排查
问题根源与解决方案
错误原因
- 固定等待时间不可靠:硬编码
time.sleep(850)无法保证SSM命令一定执行完成。如果命令执行时间超过850秒,get_command_invocation返回的输出不完整,导致分割后的列表长度不足,触发IndexError。 - 输出处理逻辑脆弱:直接通过索引
[2]、[3]取值,未考虑输出为空、命令执行失败或输出格式变化的情况,容错性极差。 - 命令输出解析错误:用
split()分割所有空白字符,会将多行输出合并成一维列表,一旦某条命令无输出(比如没有匹配的过期文件),后续结果的索引会全部偏移。 - 判断逻辑有误:
int(last_output_outbound or last_output_VEIpda)的写法不符合需求,实际需要判断两个值中任意一个大于10就触发告警。
修复后的代码
import json, boto3, time def lambda_handler(event, context): client = boto3.client('ssm') instance_id = 'i-' # 替换为实际实例ID command_timeout = 900 # 命令超时时间(15分钟,与Lambda超时一致) poll_interval = 30 # 轮询间隔(30秒) # 发送SSM命令 response = client.send_command( InstanceIds=[instance_id], DocumentName='AWS-RunShellScript', Parameters={ 'commands': [ 'find /mnt1/lean-batch/bift-outbound/done -maxdepth 1 -mtime +365 -type f -exec ls -l {} \; | wc -l', 'find /mnt1/lean-batch/bift-outbound/done -maxdepth 1 -mtime +365 -type f -exec rm -f {} \;', 'find /mnt1/lean-batch/bift-outbound/done -maxdepth 1 -mtime +180 -type f -name "VeiPda.*" -exec ls -l {} \; | wc -l', 'find /mnt1/lean-batch/bift-outbound/done -maxdepth 1 -mtime +180 -type f -name "VeiPda.*" -exec rm -f {} \;', 'find /mnt1/lean-batch/bift-outbound/done -maxdepth 1 -mtime +365 -type f -exec ls -l {} \; | wc -l', 'find /mnt1/lean-batch/bift-outbound/done -maxdepth 1 -mtime +180 -type f -name "VeiPda.*" -exec ls -l {} \; | wc -l' ] } ) command_id = response['Command']['CommandId'] start_time = time.time() # 轮询等待命令完成 while True: elapsed_time = time.time() - start_time if elapsed_time >= command_timeout: raise TimeoutError(f"SSM命令执行超时,CommandId: {command_id}") try: invocation = client.get_command_invocation( CommandId=command_id, InstanceId=instance_id, ) except client.exceptions.InvocationDoesNotExist: # 命令尚未开始执行,继续等待 time.sleep(poll_interval) continue status = invocation['Status'] if status in ['Success', 'Failed', 'TimedOut']: break time.sleep(poll_interval) # 处理命令输出:按行分割,每个命令对应一行结果 output_lines = invocation['StandardOutputContent'].strip().split('\n') # 过滤空行(避免命令无输出导致的空字符串) output_lines = [line for line in output_lines if line.strip()] # 验证输出行数是否符合预期 if len(output_lines) < 6: Esito = "Failed" error_msg = f"命令输出不完整,仅获取到{len(output_lines)}行结果,预期6行" obj = { "Esito": Esito, "Failed_Log": error_msg + "\n" + invocation['StandardOutputContent'], "CommandStatus": status } print(json.dumps(obj)) return obj # 提取最后两次验证的结果(第5、6行,索引从0开始) try: last_output_outbound = int(output_lines[4].strip()) last_output_VEIpda = int(output_lines[5].strip()) except ValueError as e: Esito = "Failed" error_msg = f"结果转换为整数失败: {str(e)}" obj = { "Esito": Esito, "Failed_Log": error_msg + "\n" + invocation['StandardOutputContent'], "CommandStatus": status } print(json.dumps(obj)) return obj # 判断是否触发告警:任意一个结果大于10则标记为Failed if last_output_outbound > 10 or last_output_VEIpda > 10: Esito = "Failed" obj = { "Esito": Esito, "Failed_Log": invocation['StandardOutputContent'], "CommandStatus": status, "OutboundRemaining": last_output_outbound, "VEIpdaRemaining": last_output_VEIpda } print(json.dumps(obj)) return obj else: Esito = "OK" obj = { "Esito": Esito, "OK_Log": invocation['StandardOutputContent'], "CommandStatus": status, "OutboundRemaining": last_output_outbound, "VEIpdaRemaining": last_output_VEIpda } print(json.dumps(obj)) return obj
关键修改点说明
- 替换固定等待为轮询:通过循环检查SSM命令状态,直到命令完成或超时,确保获取完整输出。
- 按行解析输出:将输出按行分割,每个命令的结果对应一行,避免索引偏移问题。
- 添加容错处理:检查输出行数、处理整数转换失败的情况,避免因异常导致Lambda崩溃。
- 修正判断逻辑:明确判断两个剩余文件数是否有任意一个超过10,符合业务需求。
- 增加状态日志:返回命令执行状态和具体剩余数量,便于排查问题。
内容的提问来源于stack exchange,提问作者Ava
相关产品推荐
相关产品推荐

