如何用Python等方法自动化生成AWS S3新摄入文件审计报告
Got it, let's walk through exactly how to build an automated audit report for new files added to your AWS S3 bucket—this will save you all that manual checking time. We'll use Python since it's straightforward with AWS's SDK, and I'll cover both the core report generation and how to automate the whole process.
Step 1: Setup Prerequisites
First, make sure you have these in place:
- Install the AWS SDK for Python:
pip install boto3 - Configure your AWS credentials (either via
~/.aws/credentialsfile, environment variables, or IAM roles if running on AWS services like Lambda) - Ensure your AWS identity has s3:ListBucket permission on the target bucket
Step 2: Core Python Script to Generate the Report
This script will list all files added in a specified time window (e.g., the last 24 hours), collect their filenames and timestamps, and save the data to a CSV report. We'll handle S3's pagination to make sure we get all files, even if the bucket has thousands of objects.
import boto3 from datetime import datetime, timedelta import csv def generate_s3_audit_report(bucket_name, hours_back=24, output_file="s3_audit_report.csv"): # Initialize S3 client s3 = boto3.client('s3') # Calculate the cutoff time (UTC, since S3 uses UTC timestamps) cutoff_time = datetime.utcnow() - timedelta(hours=hours_back) # List objects in the bucket, handling pagination paginator = s3.get_paginator('list_objects_v2') page_iterator = paginator.paginate(Bucket=bucket_name) audit_data = [] for page in page_iterator: if 'Contents' not in page: continue # Skip if no objects in this page for obj in page['Contents']: # Check if the file was added after our cutoff time if obj['LastModified'] >= cutoff_time: audit_data.append({ 'filename': obj['Key'], 'timestamp_utc': obj['LastModified'].strftime('%Y-%m-%d %H:%M:%S UTC') }) # Write data to CSV if audit_data: with open(output_file, 'w', newline='') as csvfile: fieldnames = ['filename', 'timestamp_utc'] writer = csv.DictWriter(csvfile, fieldnames=fieldnames) writer.writeheader() for row in audit_data: writer.writerow(row) print(f"Audit report generated: {output_file} (found {len(audit_data)} new files)") else: print("No new files found in the specified time window.") # Example usage if __name__ == "__main__": generate_s3_audit_report( bucket_name="your-target-bucket-name", hours_back=24, # Adjust this to your desired time range output_file="daily_s3_audit.csv" )
Key Notes on the Script:
- Time Zones: S3 stores
LastModifiedtimestamps in UTC. If you need local time, add a conversion step usingpytzor Python 3.9+'s zoneinfo module. - Filtering: Modify
hours_backto adjust the time window (e.g.,7*24for weekly reports). You could also filter by prefix if you only care about files in a specific folder. - Large Buckets: The paginator handles large buckets automatically, so you don't have to worry about hitting the 1000-object limit per API call.
Step 3: Automate the Process
Now that the script works manually, let's automate it so you don't have to run it every time. Here are two common options:
Option A: Run Locally with Cron (Linux/macOS)
If you want to run the report on your own machine, set up a cron job:
- Save the script as
s3_audit.py - Open crontab:
crontab -e - Add a line to run it daily at 9 AM (adjust as needed):
0 9 * * * /usr/bin/python3 /path/to/your/s3_audit.py
Option B: Serverless Automation with AWS Lambda
For a fully managed solution, use Lambda to trigger the script automatically:
- Create a new Lambda function with Python runtime
- Paste the core script (adjust to save the report to another S3 bucket instead of local file—since Lambda has temporary storage)
- Add a CloudWatch Events trigger (now called EventBridge) to run it on a schedule (e.g., daily)
- Make sure the Lambda execution role has:
s3:ListBucketon the source buckets3:PutObjecton the bucket where you'll save the report
Bonus: Enhance the Report
Want more details? You can add extra fields to the audit data:
- File size:
obj['Size'](in bytes) - Storage class:
obj['StorageClass'] - ETag (file hash):
obj['ETag']
Just update the audit_data dictionary and the CSV fieldnames to include these.
Hope this helps take the manual work out of your S3 audits! If you hit any snags with permissions, pagination, or automation, feel free to ask for more details.
内容的提问来源于stack exchange,提问作者abhi

