《Python for Data Analysis》第二章读取记录遇UnicodeDecodeError求助
Hey there! I’ve dealt with this exact problem while working through data analysis projects—nothing kills momentum faster than a decoding error mid-script. Let’s get this sorted out for you.
Why This Happens
That 0X92 byte is a dead giveaway: your JSON file is probably encoded in Windows-1252 (cp1252) instead of UTF-8. The default open() function uses UTF-8 by default, so when it hits a byte that doesn’t fit UTF-8’s rules (like that right single quote character ’ encoded as 0x92), it throws the error.
Solution 1: Specify the Correct Encoding Directly
Since 0x92 is common in cp1252, try opening the file with that encoding first:
import json records = [json.loads(line) for line in open(path, encoding='cp1252')]
This should handle that problematic character without issues.
Solution 2: Detect the File’s Encoding (If You’re Unsure)
If cp1252 doesn’t work, use the chardet library to auto-detect the file’s encoding:
- First install it (if you haven’t already):
pip install chardet - Then run this code to check the encoding:
import chardet with open(path, 'rb') as f: encoding_result = chardet.detect(f.read()) print(f"Detected encoding: {encoding_result['encoding']}") - Use the detected encoding in your original code:
records = [json.loads(line) for line in open(path, encoding=encoding_result['encoding'])]
Solution 3: Ignore Errors (Last Resort)
Only use this if you don’t care about losing the problematic characters (it can break JSON structure if the error is in a critical spot):
records = [json.loads(line) for line in open(path, encoding='utf-8', errors='ignore')]
I’d avoid this unless you’re sure the bad bytes don’t affect your data’s integrity.
内容的提问来源于stack exchange,提问作者setu bhavsar

