You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python等方法自动化生成AWS S3新摄入文件审计报告

Automating AWS S3 New File Audit Reports with Python

Got it, let's walk through exactly how to build an automated audit report for new files added to your AWS S3 bucket—this will save you all that manual checking time. We'll use Python since it's straightforward with AWS's SDK, and I'll cover both the core report generation and how to automate the whole process.

Step 1: Setup Prerequisites

First, make sure you have these in place:

  • Install the AWS SDK for Python: pip install boto3
  • Configure your AWS credentials (either via ~/.aws/credentials file, environment variables, or IAM roles if running on AWS services like Lambda)
  • Ensure your AWS identity has s3:ListBucket permission on the target bucket

Step 2: Core Python Script to Generate the Report

This script will list all files added in a specified time window (e.g., the last 24 hours), collect their filenames and timestamps, and save the data to a CSV report. We'll handle S3's pagination to make sure we get all files, even if the bucket has thousands of objects.

import boto3
from datetime import datetime, timedelta
import csv

def generate_s3_audit_report(bucket_name, hours_back=24, output_file="s3_audit_report.csv"):
    # Initialize S3 client
    s3 = boto3.client('s3')
    
    # Calculate the cutoff time (UTC, since S3 uses UTC timestamps)
    cutoff_time = datetime.utcnow() - timedelta(hours=hours_back)
    
    # List objects in the bucket, handling pagination
    paginator = s3.get_paginator('list_objects_v2')
    page_iterator = paginator.paginate(Bucket=bucket_name)
    
    audit_data = []
    for page in page_iterator:
        if 'Contents' not in page:
            continue  # Skip if no objects in this page
        
        for obj in page['Contents']:
            # Check if the file was added after our cutoff time
            if obj['LastModified'] >= cutoff_time:
                audit_data.append({
                    'filename': obj['Key'],
                    'timestamp_utc': obj['LastModified'].strftime('%Y-%m-%d %H:%M:%S UTC')
                })
    
    # Write data to CSV
    if audit_data:
        with open(output_file, 'w', newline='') as csvfile:
            fieldnames = ['filename', 'timestamp_utc']
            writer = csv.DictWriter(csvfile, fieldnames=fieldnames)
            
            writer.writeheader()
            for row in audit_data:
                writer.writerow(row)
        
        print(f"Audit report generated: {output_file} (found {len(audit_data)} new files)")
    else:
        print("No new files found in the specified time window.")

# Example usage
if __name__ == "__main__":
    generate_s3_audit_report(
        bucket_name="your-target-bucket-name",
        hours_back=24,  # Adjust this to your desired time range
        output_file="daily_s3_audit.csv"
    )

Key Notes on the Script:

  • Time Zones: S3 stores LastModified timestamps in UTC. If you need local time, add a conversion step using pytz or Python 3.9+'s zoneinfo module.
  • Filtering: Modify hours_back to adjust the time window (e.g., 7*24 for weekly reports). You could also filter by prefix if you only care about files in a specific folder.
  • Large Buckets: The paginator handles large buckets automatically, so you don't have to worry about hitting the 1000-object limit per API call.

Step 3: Automate the Process

Now that the script works manually, let's automate it so you don't have to run it every time. Here are two common options:

Option A: Run Locally with Cron (Linux/macOS)

If you want to run the report on your own machine, set up a cron job:

  1. Save the script as s3_audit.py
  2. Open crontab: crontab -e
  3. Add a line to run it daily at 9 AM (adjust as needed):
    0 9 * * * /usr/bin/python3 /path/to/your/s3_audit.py
    

Option B: Serverless Automation with AWS Lambda

For a fully managed solution, use Lambda to trigger the script automatically:

  1. Create a new Lambda function with Python runtime
  2. Paste the core script (adjust to save the report to another S3 bucket instead of local file—since Lambda has temporary storage)
  3. Add a CloudWatch Events trigger (now called EventBridge) to run it on a schedule (e.g., daily)
  4. Make sure the Lambda execution role has:
    • s3:ListBucket on the source bucket
    • s3:PutObject on the bucket where you'll save the report

Bonus: Enhance the Report

Want more details? You can add extra fields to the audit data:

  • File size: obj['Size'] (in bytes)
  • Storage class: obj['StorageClass']
  • ETag (file hash): obj['ETag']

Just update the audit_data dictionary and the CSV fieldnames to include these.

Hope this helps take the manual work out of your S3 audits! If you hit any snags with permissions, pagination, or automation, feel free to ask for more details.

内容的提问来源于stack exchange,提问作者abhi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 16:56:14