You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python数据分析新手求助:Pandas读取CSV遇UTF-8解码错误

Fixing UTF-8 Decode Error with Pandas read_csv()

Hey there! I totally get forgetting how to fix this—encoding issues are the worst, especially when you're just starting out with pandas and data analysis. Let's get this sorted out for you.

What's causing the error?

The 'Utf-8' code can’t decode byte 0xcd in position 0: invalid continuation byte error means your CSV file isn't saved in UTF-8 encoding. Pandas uses UTF-8 by default, so when it hits a byte that doesn't fit that standard (like 0xcd here), it throws this error. This is super common with Chinese-language CSV files, which are often saved with other encodings like GBK or GB2312.

Step-by-step solutions

  • Try common Chinese encodings first
    Since your file path looks like it's on Windows, GBK is the most likely culprit. Update your code to specify the encoding:

    import pandas as pd
    df = pd.read_csv('C:/Users/36373748/files/zllr.csv', encoding='gbk')
    

    If GBK doesn't work, give GB2312 a shot—it's another common encoding for Chinese text:

    df = pd.read_csv('C:/Users/36373748/files/zllr.csv', encoding='gb2312')
    
  • Detect the exact encoding with chardet
    If you're not sure what encoding the file uses, you can use the chardet library to auto-detect it. First install it via pip:

    pip install chardet
    

    Then run this code to check the encoding:

    import chardet
    
    # Open the file in binary mode to read raw bytes
    with open('C:/Users/36373748/files/zllr.csv', 'rb') as f:
        encoding_result = chardet.detect(f.read())
    
    # Print the detected encoding (e.g., 'GB2312', 'ISO-8859-1')
    print(encoding_result['encoding'])
    

    Now use that detected encoding in your read_csv call:

    df = pd.read_csv('C:/Users/36373748/files/zllr.csv', encoding=encoding_result['encoding'])
    
  • Last resort: Ignore or replace invalid bytes
    If you can't find the right encoding (and don't mind losing some data), you can tell pandas to skip invalid bytes. This isn't ideal, but it's an option:

    # Ignore invalid bytes
    df = pd.read_csv('C:/Users/36373748/files/zllr.csv', encoding='utf-8', errors='ignore')
    
    # Replace invalid bytes with a placeholder
    df = pd.read_csv('C:/Users/36373748/files/zllr.csv', encoding='utf-8', errors='replace')
    

Quick tip for future reference

Windows often saves Chinese CSV files in GBK encoding, while macOS/Linux tend to use UTF-8. So next time you hit this with a Chinese file, start with GBK first—it'll save you time!

内容的提问来源于stack exchange,提问作者Nick Gao

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:32:21