如何使用AWS Lambda Functions将AWS S3 Bucket中的文件复制到EC2 Linux实例?请求提供创建步骤及示例代码(每日有文件传入S3 Bucket)
Got it, let's walk through how to set up an AWS Lambda function that copies new files from an S3 bucket to your EC2 Linux instance—perfect for handling those daily new files that show up. I’ll break this down into clear steps with example code to make this work smoothly, following AWS best practices for security and reliability.
First, let’s make sure your EC2 instance is ready to go:
- Your EC2 Linux instance should have the AWS CLI installed (Amazon Linux 2/CentOS/RHEL come with it pre-installed; for Ubuntu/Debian, you can install it via
sudo apt install awscli). - Attach an IAM role to your EC2 instance with these permissions:
AmazonS3ReadOnlyAccess(to pull files from your S3 bucket)AmazonSSMManagedInstanceCore(to let AWS Systems Manager manage the instance—this avoids needing SSH keys or opening port 22)
- Ensure the EC2 instance can reach AWS Systems Manager (SSM): Either give it internet access, or set up VPC endpoints for SSM and S3 if it’s in a private subnet. Also, confirm the SSM Agent is running (it’s pre-installed on Amazon Linux; for other distros, follow AWS docs to install it).
Lambda needs permissions to read S3 event data and send commands to your EC2 instance via SSM:
- Go to the IAM Console > Roles > Create role.
- Select Lambda as the trusted entity, then click Next.
- Attach these managed policies:
AmazonS3ReadOnlyAccess(to read S3 bucket details and file metadata)AmazonSSMFullAccess(or create a custom policy to restrict access to only your target EC2 instance—better for security)
- Name the role something like
Lambda_S3ToEC2_Copy_Roleand create it.
Now let’s build the Lambda function and link it to your S3 bucket:
- Open the Lambda Console > Create function > Select Author from scratch.
- Name your function (e.g.,
S3ToEC2FileCopy), choose Python 3.12 (or your preferred runtime), and select the IAM role you just created. - Add an S3 trigger:
- Click Add trigger, select S3 from the dropdown.
- Choose your target S3 bucket.
- For event types, select All object create events (or just "Put" if you only want to trigger on new file uploads).
- Optionally, set a prefix/suffix (e.g.,
daily_files/or.csv) to filter which files trigger the function. - Check "Recursive invocation" (but only if you don’t plan to copy files back to S3 from EC2—avoid loops!).
- Click Add.
Replace the default Lambda code with this example. It reads the S3 event, pulls the new file details, and sends a command to EC2 via SSM to copy the file locally.
First, set up environment variables for your Lambda function (in the Configuration tab > Environment variables):
EC2_INSTANCE_ID: Your EC2 instance’s ID (e.g.,i-0abc123def456ghi)EC2_TARGET_PATH: The directory on EC2 where you want to save files (e.g.,/home/ec2-user/daily_uploads/—make sure this directory exists on EC2!)
Then paste this code into the Lambda code editor:
import boto3 import os def lambda_handler(event, context): # Extract S3 file details from the trigger event s3_record = event['Records'][0]['s3'] bucket_name = s3_record['bucket']['name'] file_key = s3_record['object']['key'] # Initialize SSM client ssm_client = boto3.client('ssm') # Build the AWS CLI command to copy the file from S3 to EC2 copy_command = f"aws s3 cp s3://{bucket_name}/{file_key} {os.environ['EC2_TARGET_PATH']}" try: # Send the command to the EC2 instance via SSM response = ssm_client.send_command( InstanceIds=[os.environ['EC2_INSTANCE_ID']], DocumentName='AWS-RunShellScript', Parameters={'commands': [copy_command]} ) print(f"Successfully triggered file copy. Command ID: {response['Command']['CommandId']}") return { 'statusCode': 200, 'body': f"Queued copy of {file_key} to EC2 instance {os.environ['EC2_INSTANCE_ID']}" } except Exception as e: print(f"Error copying file: {str(e)}") # Optional: Add retry logic or send an alert here raise e
Let’s verify everything works:
- Upload a test file to your S3 bucket (matching any prefix/suffix you set).
- Check Lambda’s CloudWatch Logs (in the Lambda function’s Monitor tab > Logs) to confirm the function ran without errors.
- Log into your EC2 instance and navigate to the target path—your test file should be there!
Since we configured the S3 trigger to fire on all new object creations, any files uploaded to the bucket each day will automatically trigger the Lambda function. If you expect a high volume of files, consider:
- Adding batch processing logic to handle multiple files in one Lambda run.
- Setting up a dead-letter queue (DLQ) to catch failed copy attempts so you can retry them later.
- Using S3 event notifications with SQS to buffer events if Lambda can’t keep up with traffic.
- Least Privilege: Instead of using
AmazonSSMFullAccess, create a custom policy that only allows Lambda to send commands to your specific EC2 instance. - Error Handling: Add logic to retry failed SSM commands (e.g., using exponential backoff) or send an email/Slack alert when copies fail.
- EC2 Path Permissions: Ensure the user running the SSM command (default is
rootorec2-user) has write access to the target directory on EC2.
内容的提问来源于stack exchange,提问作者Ashish Joshi

