使用Python与Boto3统计AWS S3指定文件数量的问题排查
Fixes for S3 Object Counting Script
Here's the corrected script that addresses all issues causing zero counts:
import csv import boto3 # Base S3 URL base_s3_url = 's3://coursevideotesting/' # Replace this with your base S3 URL # Input and output CSV file names input_csv_file = 'ldt_ffw_course_videos_temp.csv' # Replace with your input CSV file name output_csv_file = 'file_count_result.csv' # Replace with your output CSV file name # Function to count 'file_*.ts' objects in a specific S3 folder def count_file_objects(s3_bucket, s3_prefix): s3 = boto3.client('s3') count = 0 continuation_token = None while True: # Build request parameters params = {'Bucket': s3_bucket, 'Prefix': s3_prefix} if continuation_token: params['ContinuationToken'] = continuation_token response = s3.list_objects_v2(**params) # Count matching objects if contents exist if 'Contents' in response: for obj in response['Contents']: # Extract filename from the object key filename = obj['Key'].split('/')[-1] # Check if filename starts with 'file_' and is a .ts file if filename.startswith('file_') and filename.endswith('.ts'): count += 1 # Handle pagination for more than 1000 objects if response.get('IsTruncated'): continuation_token = response['NextContinuationToken'] else: break return count # Read URLs from input CSV and check file counts with open(input_csv_file, mode='r') as infile, open(output_csv_file, mode='w', newline='') as outfile: reader = csv.DictReader(infile) fieldnames = ['URL', 'Actual Files', 'Expected Files'] writer = csv.DictWriter(outfile, fieldnames=fieldnames) writer.writeheader() for row in reader: # Get the relative path from CSV and ensure it ends with '/' folder_path = row['course_video_s3_url'] if not folder_path.endswith('/'): folder_path += '/' # Construct full S3 URL for output s3_url = base_s3_url + folder_path expected_files = int(row['course_video_ts_file_cnt']) # Pass the relative path as prefix (not full S3 URL) actual_files = count_file_objects('coursevideotesting', folder_path) writer.writerow({'URL': s3_url, 'Actual Files': actual_files, 'Expected Files': expected_files})
Key Issues in Original Code & Fixes
Incorrect Prefix Parameter
- Original code passed the full S3 URL (like
s3://coursevideotesting/.../144p/) as thePrefixtolist_objects_v2. ThePrefixshould only be the path inside the bucket, not including the bucket name ors3://scheme. - Fix: Extract the relative folder path from the CSV and use that as the prefix, ensuring it ends with
/to avoid partial matches.
- Original code passed the full S3 URL (like
Wrong Filename Matching Logic
- Original code checked if the entire object key started with
file_000.ts, which would only count the first shard (not all incremental shards likefile_001.ts,file_002.ts, etc.). - Fix: Split the object key to get the filename, then check if it starts with
file_and ends with.tsto match all video shards.
- Original code checked if the entire object key started with
No Pagination Handling
list_objects_v2returns a maximum of 1000 objects per request. If there are more than 1000 shards, the original code would miss them.- Fix: Add a loop to handle pagination using
ContinuationTokenandIsTruncatedflag to fetch all pages of results.
Missing Contents Check
- When no objects match the prefix,
Contentsis not present in the response. The original code usedresponse.get('Contents', [])which handles this, but the root cause was the incorrect prefix leading to no matches.
- When no objects match the prefix,
Why Your Original Response Returned None for Contents
The full S3 URL passed as Prefix didn't match any object paths in the bucket, so list_objects_v2 returned no contents. Using the correct relative path as the prefix will resolve this and return the expected objects.
内容的提问来源于stack exchange,提问作者Prakhar Verma
相关产品推荐
相关产品推荐

