Python数据分析新手求助:Pandas读取CSV遇UTF-8解码错误
Hey there! I totally get forgetting how to fix this—encoding issues are the worst, especially when you're just starting out with pandas and data analysis. Let's get this sorted out for you.
What's causing the error?
The 'Utf-8' code can’t decode byte 0xcd in position 0: invalid continuation byte error means your CSV file isn't saved in UTF-8 encoding. Pandas uses UTF-8 by default, so when it hits a byte that doesn't fit that standard (like 0xcd here), it throws this error. This is super common with Chinese-language CSV files, which are often saved with other encodings like GBK or GB2312.
Step-by-step solutions
Try common Chinese encodings first
Since your file path looks like it's on Windows, GBK is the most likely culprit. Update your code to specify the encoding:import pandas as pd df = pd.read_csv('C:/Users/36373748/files/zllr.csv', encoding='gbk')If GBK doesn't work, give GB2312 a shot—it's another common encoding for Chinese text:
df = pd.read_csv('C:/Users/36373748/files/zllr.csv', encoding='gb2312')Detect the exact encoding with chardet
If you're not sure what encoding the file uses, you can use thechardetlibrary to auto-detect it. First install it via pip:pip install chardetThen run this code to check the encoding:
import chardet # Open the file in binary mode to read raw bytes with open('C:/Users/36373748/files/zllr.csv', 'rb') as f: encoding_result = chardet.detect(f.read()) # Print the detected encoding (e.g., 'GB2312', 'ISO-8859-1') print(encoding_result['encoding'])Now use that detected encoding in your
read_csvcall:df = pd.read_csv('C:/Users/36373748/files/zllr.csv', encoding=encoding_result['encoding'])Last resort: Ignore or replace invalid bytes
If you can't find the right encoding (and don't mind losing some data), you can tell pandas to skip invalid bytes. This isn't ideal, but it's an option:# Ignore invalid bytes df = pd.read_csv('C:/Users/36373748/files/zllr.csv', encoding='utf-8', errors='ignore') # Replace invalid bytes with a placeholder df = pd.read_csv('C:/Users/36373748/files/zllr.csv', encoding='utf-8', errors='replace')
Quick tip for future reference
Windows often saves Chinese CSV files in GBK encoding, while macOS/Linux tend to use UTF-8. So next time you hit this with a Chinese file, start with GBK first—it'll save you time!
内容的提问来源于stack exchange,提问作者Nick Gao

