Python AWS Lambda函数结构优化与变量作用域问题咨询
Hey there! Let's work through your Lambda issues step by step.
1. Fixing the NameError & Making Variables Accessible in evaluate_conditions()
The immediate error (name 'messages' is not defined) happens because you're calling the sqs() function but not capturing its return value. Plus, your count_instances function has structural bugs that prevent total_instances from being usable elsewhere. Here's how to fix this:
Step 1: Capture the return value from sqs()
Inside evaluate_conditions(), assign the result of sqs() to a variable:
messages = sqs()
Step 2: Fix the count_instances() function
Your EC2 filter syntax is incorrect (you merged two filters into one dictionary), and you're trying to use a client where you need a resource. Update the function like this:
# Count the number of active EC2 instances def count_instances(ec2_resource): total_instances = 0 # Fix filter structure: each filter is a separate dict instances = ec2_resource.instances.filter(Filters=[ { 'Name': 'instance-state-name', 'Values': ['running'] }, { 'Name': 'tag:Name', 'Values': ['NameOfInstance'] } ]) for _ in instances: total_instances += 1 return total_instances
Step 3: Pass the correct EC2 resource and capture its return value
In evaluate_conditions(), create an EC2 resource using your assumed role session, then call count_instances() and save the result:
ec2_resource = assumed_role_session.resource('ec2') total_instances = count_instances(ec2_resource)
Step 4: Remove the invalid global print statement
Delete this line outside of any function—it will throw an error because total_instances doesn't exist in the global scope:
print(f"Total number of active scan servers is: {total_instances}")
2. Code Structure Review & Refactoring Suggestions
Your current code has several structural issues that make it hard to maintain and prone to errors. A moderate refactor will make it cleaner, more reliable, and aligned with Lambda best practices. Here's what to adjust:
Key Issues in the Current Structure:
- Global
boto3clients are unused (you're using an assumed role session, so these clients don't have the right permissions) - The
sns()function is empty and doesn't do anything - Condition checks are overly verbose (you only care about one specific trigger condition)
- Assumed role logic is run globally (Lambda reuses containers, so credentials might expire on subsequent calls)
- Logging calls are incomplete (e.g.,
logger.info()with no message)
Refactored Full Code Example
import os import boto3 import logging # Configure logging once at the top logger = logging.getLogger(__name__) logger.setLevel(logging.INFO) def get_assumed_role_session(): """Helper function to get an assumed role session for cross-account access""" sts_client = boto3.client('sts') assumed_role_object = sts_client.assume_role( RoleArn=os.environ['ROLE_ARN'], RoleSessionName="AssumeRoleFromCloudOperations" ) credentials = assumed_role_object['Credentials'] return boto3.Session( aws_access_key_id=credentials['AccessKeyId'], aws_secret_access_key=credentials['SecretAccessKey'], aws_session_token=credentials['SessionToken'] ) def get_sqs_queue_message_count(session): """Get approximate number of messages in the target SQS queue""" sqs_client = session.client('sqs') queue_attrs = sqs_client.get_queue_attributes( QueueUrl=os.environ['SQS_QUEUE_URL'], AttributeNames=['ApproximateNumberOfMessages'] ) return int(queue_attrs["Attributes"]["ApproximateNumberOfMessages"]) def count_running_ec2_instances(session): """Count running EC2 instances with the specified Name tag""" ec2_resource = session.resource('ec2') instances = ec2_resource.instances.filter(Filters=[ {'Name': 'instance-state-name', 'Values': ['running']}, {'Name': 'tag:Name', 'Values': ['NameOfInstance']} ]) return sum(1 for _ in instances) # More concise way to count def send_sns_alert(session): """Send alert message to the target SNS topic""" sns_client = session.client('sns') try: sns_client.publish( TopicArn=os.environ['SNS_ARN'], Message='High number of SQS messages detected with insufficient EC2 processing instances.', Subject='SQS Queue Backlog Alert' ) logger.info("Successfully published alert to SNS topic") except Exception as e: logger.error(f"Failed to publish SNS alert: {str(e)}") def evaluate_conditions(event, context): """Main Lambda handler: check conditions and trigger alert if needed""" try: # Get assumed role session for all AWS service calls session = get_assumed_role_session() # Fetch metrics message_count = get_sqs_queue_message_count(session) instance_count = count_running_ec2_instances(session) logger.info(f"Current SQS message count: {message_count}, Running EC2 instances: {instance_count}") # Get threshold values from environment variables queue_threshold = int(os.environ['AVG_QUEUE_SIZE']) instance_threshold = int(os.environ['AVG_NR_OF_EC2_SCAN_SERVERS']) # Core condition: High queue messages AND low EC2 instances if message_count > queue_threshold and instance_count < instance_threshold: send_sns_alert(session) else: logger.info("Conditions not met - no alert needed") except Exception as e: logger.error(f"Error executing Lambda function: {str(e)}", exc_info=True) raise # Re-raise to let Lambda handle error reporting # Set the Lambda handler to this function handler = evaluate_conditions
Key Improvements in the Refactored Code:
- Modular functions: Each function handles one specific task (e.g., getting assumed role, counting instances) making code easier to test and debug
- Reusable session: All AWS service calls use the assumed role session, ensuring consistent permissions
- Simplified condition logic: Directly checks your core requirement (high queue + low instances) instead of unnecessary extra conditions
- Error handling: Added try/except blocks and meaningful logging to catch and report issues
- Cleaner counting: Uses a generator expression (
sum(1 for _ in instances)) to count EC2 instances more concisely - Proper logging: Logs key metrics and errors to help with debugging
内容的提问来源于stack exchange,提问作者Flo Flo

