You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python与Boto3统计AWS S3指定文件数量的问题排查

Fixes for S3 Object Counting Script

Here's the corrected script that addresses all issues causing zero counts:

import csv
import boto3

# Base S3 URL
base_s3_url = 's3://coursevideotesting/'  # Replace this with your base S3 URL

# Input and output CSV file names
input_csv_file = 'ldt_ffw_course_videos_temp.csv'  # Replace with your input CSV file name
output_csv_file = 'file_count_result.csv'  # Replace with your output CSV file name

# Function to count 'file_*.ts' objects in a specific S3 folder
def count_file_objects(s3_bucket, s3_prefix):
    s3 = boto3.client('s3')
    count = 0
    continuation_token = None
    
    while True:
        # Build request parameters
        params = {'Bucket': s3_bucket, 'Prefix': s3_prefix}
        if continuation_token:
            params['ContinuationToken'] = continuation_token
        
        response = s3.list_objects_v2(**params)
        
        # Count matching objects if contents exist
        if 'Contents' in response:
            for obj in response['Contents']:
                # Extract filename from the object key
                filename = obj['Key'].split('/')[-1]
                # Check if filename starts with 'file_' and is a .ts file
                if filename.startswith('file_') and filename.endswith('.ts'):
                    count += 1
        
        # Handle pagination for more than 1000 objects
        if response.get('IsTruncated'):
            continuation_token = response['NextContinuationToken']
        else:
            break
    
    return count

# Read URLs from input CSV and check file counts
with open(input_csv_file, mode='r') as infile, open(output_csv_file, mode='w', newline='') as outfile:
    reader = csv.DictReader(infile)
    fieldnames = ['URL', 'Actual Files', 'Expected Files']
    writer = csv.DictWriter(outfile, fieldnames=fieldnames)
    writer.writeheader()

    for row in reader:
        # Get the relative path from CSV and ensure it ends with '/'
        folder_path = row['course_video_s3_url']
        if not folder_path.endswith('/'):
            folder_path += '/'
        
        # Construct full S3 URL for output
        s3_url = base_s3_url + folder_path
        expected_files = int(row['course_video_ts_file_cnt'])
        # Pass the relative path as prefix (not full S3 URL)
        actual_files = count_file_objects('coursevideotesting', folder_path)
        
        writer.writerow({'URL': s3_url, 'Actual Files': actual_files, 'Expected Files': expected_files})

Key Issues in Original Code & Fixes

  1. Incorrect Prefix Parameter

    • Original code passed the full S3 URL (like s3://coursevideotesting/.../144p/) as the Prefix to list_objects_v2. The Prefix should only be the path inside the bucket, not including the bucket name or s3:// scheme.
    • Fix: Extract the relative folder path from the CSV and use that as the prefix, ensuring it ends with / to avoid partial matches.
  2. Wrong Filename Matching Logic

    • Original code checked if the entire object key started with file_000.ts, which would only count the first shard (not all incremental shards like file_001.ts, file_002.ts, etc.).
    • Fix: Split the object key to get the filename, then check if it starts with file_ and ends with .ts to match all video shards.
  3. No Pagination Handling

    • list_objects_v2 returns a maximum of 1000 objects per request. If there are more than 1000 shards, the original code would miss them.
    • Fix: Add a loop to handle pagination using ContinuationToken and IsTruncated flag to fetch all pages of results.
  4. Missing Contents Check

    • When no objects match the prefix, Contents is not present in the response. The original code used response.get('Contents', []) which handles this, but the root cause was the incorrect prefix leading to no matches.

Why Your Original Response Returned None for Contents

The full S3 URL passed as Prefix didn't match any object paths in the bucket, so list_objects_v2 returned no contents. Using the correct relative path as the prefix will resolve this and return the expected objects.

内容的提问来源于stack exchange,提问作者Prakhar Verma

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 04:10:56