使用Pandas读取由to_json生成的带缩进的JSON文件
问题描述
我用DataFrame.to_json()写入JSON文件时设置了indent参数,代码如下:
df.to_json(path_or_buf=file_json, orient="records", lines=True, indent=2)
关键问题出在indent=2,如果不设置这个参数,读写流程是正常的。现在我尝试用pd.read_json()读取该文件:
df = pd.read_json(file_json, lines=True)
但这种方式要求每行对应一个完整的JSON对象,而缩进导致每个对象被拆分成多行,读取失败。查看read_json的文档后没找到处理缩进的参数,请问怎么读取这个文件?最好不用自己编写复杂的读取逻辑。
解决方案
方法1:借助Python标准json模块预处理后转DataFrame
这是最简便的方案,无需手动编写解析逻辑:
import json import pandas as pd with open(file_json, "r", encoding="utf-8") as f: content = f.read() # 将多个多行JSON对象拼接为合法的JSON数组 json_array = "[" + content.replace("}\n{", "},{") + "]" data = json.loads(json_array) df = pd.DataFrame(data)
方法2:针对大文件的分块读取方案
如果文件体积过大,无法一次性加载到内存,可以分块读取单个JSON对象:
import pandas as pd import json df_list = [] current_obj = [] with open(file_json, "r", encoding="utf-8") as f: for line in f: stripped_line = line.strip() if not stripped_line: continue current_obj.append(stripped_line) # 检测到对象结尾时解析并存储 if stripped_line.endswith("}"): obj_str = "".join(current_obj) df_list.append(pd.DataFrame([json.loads(obj_str)])) current_obj = [] df = pd.concat(df_list, ignore_index=True)
后续优化建议
后续写入JSON时,避免同时使用lines=True和indent参数,二者设计目标冲突:
lines=True用于生成每行一个紧凑JSON对象的格式,适合按行快速读写indent用于生成带缩进的可读JSON,通常配合orient="records"生成单个JSON数组,此时直接用pd.read_json(file_json)就能正常读取
内容的提问来源于stack exchange,提问作者Soid
相关产品推荐
相关产品推荐

