You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

加载JSON文件时UnicodeDecodeError持续出现的解决咨询

解决JSON文件UTF-8解码失败的问题

针对你遇到的UnicodeDecodeError: 'utf-8' codec can't decode byte 0xb4错误,给出以下几个可行的解决方向:

1. 先确定文件的实际编码

VS Code指定编码只是用于编辑器显示,不会改变文件本身的编码格式。你需要先检测文件的真实编码:

  • 安装chardet库(pip install chardet),然后运行以下代码检测:
    import chardet
    import os
    
    downloads_path = os.path.expanduser("~/Downloads")
    filename = os.path.join(downloads_path, "DiscussData-Words_of_Putin_&_Medvedev__Term_Frequency_Analysis-transcripts_ru_2023-01-16.json")
    
    with open(filename, 'rb') as f:
        # 读取前100KB内容进行检测,足够判断编码
        result = chardet.detect(f.read(100000))
    print(f"检测到的编码: {result['encoding']}, 置信度: {result['confidence']}")
    
  • 从报错的0xb4字节来看,这个文件很可能是cp1251编码(俄语场景常用的Windows编码),该编码下0xb4对应´字符。

2. 使用检测到的编码打开文件

如果检测结果是cp1251,直接修改文件打开语句:

with open(filename, "r", encoding="cp1251") as file:
    speeches = json.load(file)

如果仍有少量无法解码的字符,可结合errors='replace'参数跳过错误(无法解码的字符会被替换为�):

with open(filename, "r", encoding="cp1251", errors='replace') as file:
    speeches = json.load(file)

3. 处理超大JSON文件(文件大小超300MB)

报错位置在314000015字节,说明文件极大,一次性加载可能导致内存问题,同时编码错误也会直接中断加载。可以用ijson库流式解析:

  • 安装ijson(pip install ijson),然后修改代码为:
    import json
    import os
    import re
    import ijson
    
    keywords = ["украин", "донбас", "донецк", "крым"]
    downloads_path = os.path.expanduser("~/Downloads")
    filename = os.path.join(downloads_path, "DiscussData-Words_of_Putin_&_Medvedev__Term_Frequency_Analysis-transcripts_ru_2023-01-16.json")
    filtered_speeches = []
    
    # 流式解析大JSON,避免一次性加载
    with open(filename, 'rb') as f:
        # 假设JSON是顶层数组,用'item'遍历每个元素
        parser = ijson.items(f, 'item')
        for speech in parser:
            text = speech["speech"]
            if any(re.search(r"\b{}\b".format(keyword), text, re.IGNORECASE) for keyword in keywords):
                filtered_speeches.append(speech)
    
    output_file = os.path.join(downloads_path, "filtered_speeches.json")
    with open(output_file, "w", encoding="utf-8") as file:
        json.dump(filtered_speeches, file, ensure_ascii=False, indent=2)
    
    print("Filtered speeches saved to", output_file)
    
  • 这种方式逐行解析,不会因为某一处编码错误导致整个程序崩溃,同时节省内存。

4. 修复UTF-8文件中的损坏字节

如果文件确实是UTF-8但存在少量损坏字节,可以用errors='surrogateescape'参数处理:

with open(filename, "r", encoding="utf-8", errors='surrogateescape') as file:
    speeches = json.load(file)

这个参数会将无法解码的字节转换为代理字符,后续写入文件时还能还原原始字节,适合UTF-8文件有个别损坏的场景。


内容的提问来源于stack exchange,提问作者researcher.ella

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 05:37:01