使用NumPy读取FIFA 18球员数据集时Python 3报UnicodeDecodeError
Hey there! That UnicodeDecodeError is super common when working with CSV datasets that aren’t saved in the default UTF-8 encoding—and the FIFA 18 player dataset is a classic example of this. Let’s break down how to fix it quickly.
Why this happens
NumPy’s np.genfromtxt() defaults to using UTF-8 encoding, but many older datasets (like this FIFA one) are saved with Latin-1 (or cp1252) encoding instead. The error pops up when the file has characters that don’t fit in UTF-8.
The quick fix
Just add the encoding parameter to your genfromtxt call, specifying either 'latin1' or 'cp1252'—both should work for this dataset:
import numpy as np # Use latin1 encoding to handle non-UTF-8 characters np_fifa = np.genfromtxt('Datasets/FIFA2018.csv', delimiter=',', encoding='latin1') print(np_fifa)
If that still doesn’t work...
If you’re still getting errors, you can check the actual encoding of your CSV file:
- Open it in a text editor like Notepad++ and look at the "Encoding" menu to see what it’s saved as.
- Or use the
chardetlibrary to auto-detect the encoding:import chardet with open('Datasets/FIFA2018.csv', 'rb') as f: result = chardet.detect(f.read()) print(result['encoding'])
Then use that detected encoding in the encoding parameter.
内容的提问来源于stack exchange,提问作者user8795229

