如何延迟Lambda停止AWS EC2实例至用户不活跃或登出?
这个需求完全可以实现,核心思路是在Lambda执行关机操作前,先判断目标EC2实例上是否有用户活跃登录,再决定是立即关机、延迟重试还是直接跳过当天的关机任务。下面给你几个可行的实现方案,全部可以通过boto3代码完成:
前提准备
首先,你的现有Lambda代码里的print语句是Python2语法,建议改成Python3的格式(现在Lambda默认使用Python3.x):
print(f"Stopped your instance: {str(instances)}")
另外,要确保目标EC2实例已经安装并运行了SSM Agent(大多数AWS官方AMI默认自带),并且实例关联了带有AmazonSSMManagedInstanceCore权限的IAM角色,这样Lambda才能通过SSM远程执行命令检查用户会话。同时,Lambda的执行角色需要添加ssm:SendCommand、ssm:GetCommandInvocation、ec2:StopInstances这几个权限。
方案1:检查活跃会话后决定是否关机
这个方案会在Lambda触发后,先检查实例上的活跃用户数,只有当没有活跃用户时才执行关机,否则直接跳过本次操作。
修改后的Lambda代码示例:
import boto3 import time def check_active_sessions(region, instance_id): """检查指定EC2实例是否有活跃用户会话""" ssm = boto3.client('ssm', region_name=region) try: # 发送Shell命令,统计当前登录用户数 response = ssm.send_command( InstanceIds=[instance_id], DocumentName='AWS-RunShellScript', Parameters={'commands': ['who | wc -l']} ) command_id = response['Command']['CommandId'] # 等待命令执行完成(根据实例性能调整等待时间) time.sleep(5) # 获取命令执行结果 output = ssm.get_command_invocation( CommandId=command_id, InstanceId=instance_id ) active_user_count = int(output['StandardOutputContent'].strip()) # 返回是否无活跃用户 return active_user_count == 0 except Exception as e: print(f"检查实例{instance_id}活跃会话时出错: {str(e)}") # 若无法检查(比如SSM Agent未运行),可根据需求返回True(强制关机)或False(跳过) return False def lambda_handler(event, context): region = event.get('region') instances = event.get('instances') ec2 = boto3.client('ec2', region_name=region) for instance_id in instances: if check_active_sessions(region, instance_id): ec2.stop_instances(InstanceIds=[instance_id]) print(f"已停止实例: {instance_id}") else: print(f"实例{instance_id}存在活跃用户,跳过本次关机操作")
方案2:延迟重试直到用户登出
如果希望不是直接跳过,而是每隔一段时间重试检查,直到用户登出或达到最大重试次数再停止实例,可以在方案1的基础上增加轮询逻辑:
import boto3 import time def check_active_sessions(region, instance_id): # 复用方案1中的检查函数 ssm = boto3.client('ssm', region_name=region) try: response = ssm.send_command( InstanceIds=[instance_id], DocumentName='AWS-RunShellScript', Parameters={'commands': ['who | wc -l']} ) command_id = response['Command']['CommandId'] time.sleep(5) output = ssm.get_command_invocation(CommandId=command_id, InstanceId=instance_id) active_user_count = int(output['StandardOutputContent'].strip()) return active_user_count == 0 except Exception as e: print(f"检查实例{instance_id}活跃会话时出错: {str(e)}") return False def wait_for_no_active_users(region, instance_id, max_retries=4, delay_minutes=30): """轮询检查实例,直到无活跃用户或达到最大重试次数""" retries = 0 while retries < max_retries: if check_active_sessions(region, instance_id): return True retries += 1 print(f"实例{instance_id}仍有活跃用户,{delay_minutes}分钟后重试(第{retries}/{max_retries}次)") time.sleep(delay_minutes * 60) print(f"已达到最大重试次数,跳过实例{instance_id}的关机操作") return False def lambda_handler(event, context): region = event.get('region') instances = event.get('instances') ec2 = boto3.client('ec2', region_name=region) for instance_id in instances: if wait_for_no_active_users(region, instance_id): ec2.stop_instances(InstanceIds=[instance_id]) print(f"已停止实例: {instance_id}")
方案3:用Step Functions实现更可靠的重试流程
如果希望重试逻辑更可靠(比如Lambda超时后仍能继续重试),可以使用AWS Step Functions来编排整个流程:
- 由CloudWatch定时触发Step Functions状态机
- 状态机先调用Lambda检查实例活跃会话
- 若无活跃用户,执行关机Lambda;若有,等待指定时间后重试
- 达到最大重试次数后,终止流程(跳过关机)
这种方式比Lambda内部轮询更稳定,避免Lambda因超时中断重试逻辑。
关于你提到的Waiter
你说的AWS内置Waiter确实没有等待实例无活跃的类型,但我们可以通过自己编写轮询逻辑(比如方案2中的代码)来实现类似的效果,本质就是自定义的Waiter。
内容的提问来源于stack exchange,提问作者Andrew Drake

