You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Boto3读取S3中CSV文件时遇UnicodeDecodeError问题求助

Fixing UnicodeDecodeError When Reading CSV from S3

Hey there, let's work through that UnicodeDecodeError you're hitting when loading your CSV from S3. The error specifically calls out byte 0x96 failing to decode with UTF-8—here's what's going on and how to fix it:

Why This Happens

That 0x96 byte isn't a valid UTF-8 character, but it is the en dash (–) in the Windows-1252 encoding. This means your CSV file was almost certainly saved using Windows-1252 (or a similar single-byte encoding like ISO-8859-1) instead of UTF-8.

Solutions to Try

  • Try Windows-1252 encoding first
    Since 0x96 is a common character in Windows-1252, update your pd.read_csv line to use this encoding—it's the most likely quick fix:

    data = pd.read_csv(io.BytesIO(obj['Body'].read()), delimiter=',', engine='python', encoding='windows-1252')
    
  • Auto-detect the file's actual encoding
    If Windows-1252 doesn't work, use the chardet library to automatically figure out the correct encoding. First install it (pip install chardet), then adjust your code:

    import chardet
    import boto3
    import pandas as pd
    import io
    
    client = boto3.client('s3')
    obj = client.get_object(Bucket='bucket1', Key='file.csv')
    content_bytes = obj['Body'].read()
    
    # Detect the encoding
    detection_result = chardet.detect(content_bytes)
    detected_encoding = detection_result['encoding']
    
    # Read the CSV with the detected encoding
    data = pd.read_csv(io.BytesIO(content_bytes), delimiter=',', engine='python', encoding=detected_encoding)
    
  • Use a fallback encoding to load data first
    If you just need to get the data loaded (even with some placeholder characters), use ISO-8859-1—it's a single-byte encoding that accepts any byte without throwing errors:

    data = pd.read_csv(io.BytesIO(obj['Body'].read()), delimiter=',', engine='python', encoding='ISO-8859-1')
    

    Note: This might show odd characters for non-ASCII text, but you can clean them up once the data is loaded.

  • Explicitly handle decoding errors
    If there are only a few problematic characters, you can tell pandas to replace or ignore them while sticking with UTF-8:

    # Replace un-decodable characters with �
    data = pd.read_csv(io.BytesIO(obj['Body'].read()), delimiter=',', engine='python', encoding='utf-8', errors='replace')
    
    # Or ignore un-decodable characters entirely
    data = pd.read_csv(io.BytesIO(obj['Body'].read()), delimiter=',', engine='python', encoding='utf-8', errors='ignore')
    

Start with the first solution—it's the most probable fix for your specific error. If that doesn't work, move on to auto-detecting the encoding.

内容的提问来源于stack exchange,提问作者Nasri

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 06:39:38