You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas读取由to_json生成的带缩进的JSON文件

问题描述

我用DataFrame.to_json()写入JSON文件时设置了indent参数,代码如下:

df.to_json(path_or_buf=file_json, orient="records", lines=True, indent=2)

关键问题出在indent=2,如果不设置这个参数,读写流程是正常的。现在我尝试用pd.read_json()读取该文件:

df = pd.read_json(file_json, lines=True)

但这种方式要求每行对应一个完整的JSON对象,而缩进导致每个对象被拆分成多行,读取失败。查看read_json的文档后没找到处理缩进的参数,请问怎么读取这个文件?最好不用自己编写复杂的读取逻辑。

解决方案

方法1:借助Python标准json模块预处理后转DataFrame

这是最简便的方案,无需手动编写解析逻辑:

import json
import pandas as pd

with open(file_json, "r", encoding="utf-8") as f:
    content = f.read()
    # 将多个多行JSON对象拼接为合法的JSON数组
    json_array = "[" + content.replace("}\n{", "},{") + "]"
    data = json.loads(json_array)

df = pd.DataFrame(data)

方法2:针对大文件的分块读取方案

如果文件体积过大,无法一次性加载到内存,可以分块读取单个JSON对象:

import pandas as pd
import json

df_list = []
current_obj = []

with open(file_json, "r", encoding="utf-8") as f:
    for line in f:
        stripped_line = line.strip()
        if not stripped_line:
            continue
        current_obj.append(stripped_line)
        # 检测到对象结尾时解析并存储
        if stripped_line.endswith("}"):
            obj_str = "".join(current_obj)
            df_list.append(pd.DataFrame([json.loads(obj_str)]))
            current_obj = []

df = pd.concat(df_list, ignore_index=True)

后续优化建议

后续写入JSON时,避免同时使用lines=True和indent参数,二者设计目标冲突:

  • lines=True用于生成每行一个紧凑JSON对象的格式,适合按行快速读写
  • indent用于生成带缩进的可读JSON,通常配合orient="records"生成单个JSON数组,此时直接用pd.read_json(file_json)就能正常读取

内容的提问来源于stack exchange,提问作者Soid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.23 13:40:10