如何解决Python Pandas读取文件时的UnicodeEncodeError编码错误?
Hey there, let's break down what's going on here. First off—your error isn't from reading the Excel file at all. The problem pops up when you try to print the data to your Windows console, which uses cp1252 encoding by default. That encoding can't handle some of the Unicode characters in your dataset, hence the UnicodeEncodeError.
Also, quick note: the encoding parameter in pd.read_excel() doesn't actually do anything for .xlsx files. Those files are stored in a UTF-8 compliant format internally, so you can drop that parameter entirely from your code.
Here are a few straightforward fixes to try:
1. Switch Your Console to UTF-8
Windows consoles default to cp1252, but you can force them to use UTF-8 which supports all Unicode characters. Run these two commands in your terminal before launching your script:
chcp 65001 set PYTHONIOENCODING=utf-8
Then run your Python script again. This tells both the console and Python to use UTF-8 for output.
2. Print a Subset of the Data Instead of the Whole Thing
If you don't need to see every single row, just print the first few rows with head()—this minimizes the chance of hitting those unprintable characters:
print(data.head())
3. Encode Output to Handle Problematic Characters
If you really need to print the full dataset, you can encode the output to cp1252 and either ignore or replace characters that can't be displayed:
# Ignore unprintable characters print(data.to_string().encode('cp1252', errors='ignore').decode('cp1252')) # Replace unprintable characters with a ? print(data.to_string().encode('cp1252', errors='replace').decode('cp1252'))
4. Export to a UTF-8 CSV File
If viewing in the console isn't a must, export the data to a UTF-8 CSV file and open it in Excel or a text editor that supports UTF-8:
data.to_csv('output_data.csv', encoding='utf-8-sig', index=False)
The utf-8-sig encoding ensures Excel recognizes the file as UTF-8 without any weird character issues.
And here's your cleaned-up original code (since the encoding parameter is unnecessary for xlsx):
import pandas as pd excel_file = 'Task1/Data_task1.xlsx' data = pd.read_excel(excel_file) # Use one of the above methods to view the data print(data.head())
To sum up: The read operation worked fine—your console just can't display all the characters in your dataset. Any of these methods should get you past the error so you can work with your data.
内容的提问来源于stack exchange,提问作者srinivas muralidharan

