使用Boto3读取S3中CSV文件时遇UnicodeDecodeError问题求助
Hey there, let's work through that UnicodeDecodeError you're hitting when loading your CSV from S3. The error specifically calls out byte 0x96 failing to decode with UTF-8—here's what's going on and how to fix it:
Why This Happens
That 0x96 byte isn't a valid UTF-8 character, but it is the en dash (–) in the Windows-1252 encoding. This means your CSV file was almost certainly saved using Windows-1252 (or a similar single-byte encoding like ISO-8859-1) instead of UTF-8.
Solutions to Try
Try Windows-1252 encoding first
Since0x96is a common character in Windows-1252, update yourpd.read_csvline to use this encoding—it's the most likely quick fix:data = pd.read_csv(io.BytesIO(obj['Body'].read()), delimiter=',', engine='python', encoding='windows-1252')Auto-detect the file's actual encoding
If Windows-1252 doesn't work, use thechardetlibrary to automatically figure out the correct encoding. First install it (pip install chardet), then adjust your code:import chardet import boto3 import pandas as pd import io client = boto3.client('s3') obj = client.get_object(Bucket='bucket1', Key='file.csv') content_bytes = obj['Body'].read() # Detect the encoding detection_result = chardet.detect(content_bytes) detected_encoding = detection_result['encoding'] # Read the CSV with the detected encoding data = pd.read_csv(io.BytesIO(content_bytes), delimiter=',', engine='python', encoding=detected_encoding)Use a fallback encoding to load data first
If you just need to get the data loaded (even with some placeholder characters), useISO-8859-1—it's a single-byte encoding that accepts any byte without throwing errors:data = pd.read_csv(io.BytesIO(obj['Body'].read()), delimiter=',', engine='python', encoding='ISO-8859-1')Note: This might show odd characters for non-ASCII text, but you can clean them up once the data is loaded.
Explicitly handle decoding errors
If there are only a few problematic characters, you can tell pandas to replace or ignore them while sticking with UTF-8:# Replace un-decodable characters with � data = pd.read_csv(io.BytesIO(obj['Body'].read()), delimiter=',', engine='python', encoding='utf-8', errors='replace') # Or ignore un-decodable characters entirely data = pd.read_csv(io.BytesIO(obj['Body'].read()), delimiter=',', engine='python', encoding='utf-8', errors='ignore')
Start with the first solution—it's the most probable fix for your specific error. If that doesn't work, move on to auto-detecting the encoding.
内容的提问来源于stack exchange,提问作者Nasri

