合并JSON文件报错JSONDecodeError: Extra data的解决方法咨询
你的文件属于行式JSON(JSON Lines)格式,每行是独立的JSON对象,对应两种常用合并方案:
方案1:轻量文本拼接(适合已确认所有文件格式合法的场景)
这种方法不需要解析JSON内容,直接逐行拷贝内容,内存占用极低,处理大文件速度快:
import os # 配置项:按需修改文件路径、后缀规则 user = 'name' file_prefix = f'/local/data/{user}/testing_json_merge/testlog' # 示例按a-z顺序取前30个文件,若为数字编号可改为range(1,31) file_suffixes = [chr(ord('a') + i) for i in range(30)] output_path = f'/local/data/{user}/testing_json_merge/merged_log.json' with open(output_path, 'w', encoding='utf-8') as out_f: for suffix in file_suffixes: file_path = f'{file_prefix}{suffix}.json' if not os.path.exists(file_path): continue with open(file_path, 'r', encoding='utf-8') as in_f: for line in in_f: stripped_line = line.strip() # 跳过空行 if stripped_line: out_f.write(stripped_line + '\n')
方案2:带JSON校验的合并(学术场景优先推荐)
如果担心源文件存在格式错误的行,可逐行解析校验后再写入,确保合并后的文件所有内容都是合法JSON:
import json import os user = 'name' file_prefix = f'/local/data/{user}/testing_json_merge/testlog' file_suffixes = [chr(ord('a') + i) for i in range(30)] output_path = f'/local/data/{user}/testing_json_merge/merged_log.json' merged_lines = [] for suffix in file_suffixes: file_path = f'{file_prefix}{suffix}.json' if not os.path.exists(file_path): continue with open(file_path, 'r', encoding='utf-8') as f: for line in f: stripped_line = line.strip() if not stripped_line: continue # 校验JSON格式合法性 try: json.loads(stripped_line) merged_lines.append(stripped_line) except json.JSONDecodeError: # 遇到非法行打印提示跳过,避免中断合并 print(f"跳过文件{file_path}中的非法JSON行:{stripped_line[:100]}...") # 写入合并后的行式JSON with open(output_path, 'w', encoding='utf-8') as f: f.write('\n'.join(merged_lines))
补充说明
如果你后续导入模拟器需要的是外层包裹数组的标准JSON格式,把最后写入的部分改成如下代码即可:
# 输出为外层数组的标准JSON with open(output_path, 'w', encoding='utf-8') as f: f.write('[' + ',\n'.join(merged_lines) + ']')
内容的提问来源于stack exchange,提问作者Androo42
相关产品推荐
相关产品推荐

