求助:使用Python的Boto从S3下载snappy.parquet文件遇空文件夹问题
Hey there! Sorry to hear you're stuck with an empty folder after trying to download snappy.parquet files from Amazon S3—let's walk through some common fixes to get this sorted out.
Let's start with the most common culprits and work our way through:
1. Double-check your S3 path logic
Amazon S3 doesn't have real "folders"—it uses prefixes to organize objects. If you're targeting a prefix (like s3://my-bucket/data/) instead of the actual parquet file, you might end up creating an empty local folder if there are no objects under that prefix, or if your download tool isn't set to fetch recursively.
For example, if your file lives at s3://my-bucket/data/transactions.snappy.parquet, make sure you're pointing directly to that file, or using recursive download to grab all objects under the data/ prefix.
2. Fix your download command/code
Let's cover two common tools: AWS CLI and Python's boto3.
AWS CLI
If you're using the CLI, skip the folder-only path and either target the specific file:
aws s3 cp s3://your-bucket/path/to/transactions.snappy.parquet ./local-downloads/
Or use --recursive to pull all objects under a prefix (useful if you have multiple parquet files):
aws s3 cp s3://your-bucket/data/ ./local-downloads/ --recursive
Note: This will skip 0-byte "folder placeholder" objects that S3 sometimes uses to display empty folders.
Python Boto3
If you're coding with boto3, make sure you're actually iterating over the parquet objects and not just creating local folders. Here's a reliable snippet:
import boto3 import os # Set up S3 client s3 = boto3.client('s3') bucket = 'your-bucket-name' s3_prefix = 'path/to/your/parquet/files/' local_dir = './local-parquet-files/' # Create local directory if it doesn't exist os.makedirs(local_dir, exist_ok=True) # List all objects under the prefix response = s3.list_objects_v2(Bucket=bucket, Prefix=s3_prefix) if 'Contents' in response: for obj in response['Contents']: # Skip empty placeholder "folders" (size 0) if obj['Size'] > 0: # Get the filename from the S3 object key file_name = obj['Key'].split('/')[-1] local_file_path = os.path.join(local_dir, file_name) # Download the file s3.download_file(bucket, obj['Key'], local_file_path) print(f"Downloaded: {local_file_path}") else: print("No objects found under the specified S3 prefix.")
3. Verify the S3 object isn't empty
Head over to the S3 Console, navigate to your object, and check its size. If it's 0 bytes, that's why you're getting an empty file/folder. Double-check that the parquet file was uploaded correctly to S3 in the first place.
4. Convert to CSV once downloaded
Once you have the snappy.parquet file locally, use pandas to convert it to CSV. First install the required dependencies:
pip install pandas pyarrow
Then run this code:
import pandas as pd # Load the snappy-compressed parquet file df = pd.read_parquet('./local-parquet-files/transactions.snappy.parquet', engine='pyarrow') # Save as CSV (exclude index with index=False) df.to_csv('./output.csv', index=False)
Give these steps a try—chances are the issue is either a missing recursive flag, targeting the wrong S3 path, or accidentally downloading an empty placeholder object. Let me know if you hit any snags along the way!
内容的提问来源于stack exchange,提问作者EllaBV

