读取含多组data/error数组的JSON文件遇解码错误,求解决方案
问题
我有一个需要转换为CSV的JSON文件,用Python3读取时触发值错误,原因是文件里包含多组data和error数组——相当于多个JSON文件硬合并在一起,直接破坏了JSON的格式规范。
下面是文件里的片段示例:
{ "detail": "Could not find user with ids: [99].", "parameter": "ids", "resource_id": "99", "resource_type": "user", "title": "Not Found Error", "type": "https://api.twitter.com/2/problems/resource-not-found", "value": "99" } ] }{"data": [ { "created_at": "2006-04-13T01:08:55.000Z",
上面是Twitter返回的用户不存在的error数组片段,下面是成功获取的用户data数组片段。我试过用pandas读取,但和json.load()报了一样的错:
json.decoder.JSONDecodeError: Extra data: line 1601 column 2 (char 59799)
想请教两种解决思路:要么把文件按data和error拆成多个合法JSON文件,方便用json.load()读取;要么有没有工具能直接读取这种损坏的文件?
解决方案
方法一:拆分文件为单个合法JSON
写个简单的Python脚本就能完成拆分,核心是按JSON对象的边界}{分割内容,再补全可能缺失的括号,最后分别保存:
import json # 替换成你的目标文件路径 input_file = "your_damaged_json.json" with open(input_file, 'r', encoding='utf-8') as f: raw_content = f.read() # 把}{替换成}\n{,分割成多个JSON片段 json_snippets = raw_content.replace('}{', '}\n{').split('\n') # 过滤空行和空白内容 json_snippets = [s.strip() for s in json_snippets if s.strip()] for index, snippet in enumerate(json_snippets): # 尝试修复可能缺失的括号 fixed_snippet = snippet try: # 先尝试直接解析 json.loads(fixed_snippet) except json.JSONDecodeError: # 尝试补全数组的首尾括号 if not (fixed_snippet.startswith('[') or fixed_snippet.startswith('{')): fixed_snippet = f"[{fixed_snippet}]" elif fixed_snippet.endswith(']') and not fixed_snippet.startswith('['): fixed_snippet = f"[{fixed_snippet}" elif fixed_snippet.startswith('[') and not fixed_snippet.endswith(']'): fixed_snippet = f"{fixed_snippet}]" # 再次尝试解析,失败则跳过 try: json.loads(fixed_snippet) except: print(f"第{index+1}个片段无法修复,已跳过") continue # 保存为单个JSON文件 with open(f"json_part_{index+1}.json", 'w', encoding='utf-8') as out_f: out_f.write(fixed_snippet)
拆分后每个文件都是合法JSON,直接用json.load()或pandas读取都没问题。
方法二:直接读取处理(不生成中间文件)
如果不想拆分文件,可以用逐段解析的方式读取,不用一次性加载整个文件:
import json from json.decoder import WHITESPACE def read_multiple_json(file_path): with open(file_path, 'r', encoding='utf-8') as f: decoder = json.JSONDecoder() buffer = "" for line in f: buffer += line buffer = buffer.lstrip(WHITESPACE) while buffer: try: # 尝试解析当前缓冲区的JSON json_obj, parse_end = decoder.raw_decode(buffer) yield json_obj # 截断缓冲区,处理剩余内容 buffer = buffer[parse_end:].lstrip(WHITESPACE) except json.JSONDecodeError: # 当前缓冲区内容不足,继续读下一行 break # 使用示例:遍历所有JSON对象并处理 for obj in read_multiple_json("your_damaged_json.json"): if "data" in obj: # 处理data数组,比如直接写入CSV print("读取到data数组:", obj["data"]) elif "detail" in obj or "error" in obj: # 处理错误信息 print("读取到错误信息:", obj)
这个方法可以直接在遍历过程中将数据转换成CSV,省掉拆分文件的步骤。
工具替代方案
不想写代码的话,可以用jq命令行工具直接处理:
- 提取所有
data数组并保存到文件:
jq -s '.[] | select(.data != null)' your_damaged_json.json > all_data.json
- 提取所有错误相关内容并保存:
jq -s '.[] | select(.detail != null)' your_damaged_json.json > all_errors.json
jq会自动识别这种多JSON拼接的格式,把它们当成数组处理,直接提取你需要的内容。
内容的提问来源于stack exchange,提问作者BoostMatch
相关产品推荐
相关产品推荐

