Python加载含多语言JSON文件遇UnicodeDecodeError,求解决方案
解决JSON文件多语言编码加载错误
核心原因
报错是因为Python默认使用系统编码(Windows下通常为cp1252)打开文件,而包含多语言的JSON文件一般采用UTF-8编码,两种编码不匹配导致无法解析非ASCII字符。
有效解决方案
1. 指定UTF-8编码打开文件(推荐)
修改代码,在open()函数中明确指定encoding='utf-8',同时使用with语句自动管理文件关闭:
import json with open('./DummyData/movie_details.json', encoding='utf-8') as file: data = json.load(file) print(data)
2. 临时忽略/替换无效字节(不推荐,可能丢失数据)
如果文件存在少量无效UTF-8字节,可通过errors参数跳过或替换错误内容:
- 忽略错误:
import json with open('./DummyData/movie_details.json', encoding='utf-8', errors='ignore') as file: data = json.load(file) print(data)
- 用
�替换无效字节:
import json with open('./DummyData/movie_details.json', encoding='utf-8', errors='replace') as file: data = json.load(file) print(data)
3. 自动检测文件编码
如果不确定文件实际编码,可使用chardet库检测后再加载:
- 先安装库:
pip install chardet
- 检测并加载:
import chardet import json # 检测文件编码 with open('./DummyData/movie_details.json', 'rb') as f: encoding_info = chardet.detect(f.read()) # 用检测到的编码打开文件 with open('./DummyData/movie_details.json', encoding=encoding_info['encoding']) as file: data = json.load(file) print(data)
内容的提问来源于stack exchange,提问作者Science Meetup
相关产品推荐
相关产品推荐

