You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何删除嵌套JSON冗余字段,将其转换为指定的简化数据格式?

嵌套JSON简化处理方案

直接通过Python原生的json库配合多层循环提取目标字段即可,这种方式逻辑清晰,不会打乱原有数据对应关系,比字符串替换、强行转DataFrame的方式适配性更高:

import json

# 读取原始JSON文件
with open("你的原始文件路径.json", "r", encoding="utf-8") as f:
    raw_data = json.load(f)

result = []
# 遍历外层data数组
for item in raw_data["data"]:
    # 遍历每个段落,拿到context值
    for para in item["paragraphs"]:
        context = para["context"]
        # 遍历每个问答对,提取目标字段
        for qa in para["qas"]:
            # 提取所有回答的文本内容组成数组
            answer_texts = [ans["text"] for ans in qa["answers"]]
            result.append({
                "id": qa["id"],
                "context": context,
                "question": qa["question"],
                "answers": {"text": answer_texts}
            })

# 导出处理后的简化结果
with open("简化结果.json", "w", encoding="utf-8") as f:
    json.dump(result, f, ensure_ascii=False, indent=2)

上述代码会严格保留context和所属问答对的对应关系,自动过滤不需要的title、answer_start等字段,输出格式完全匹配你需要的结构。如果后续原JSON嵌套规则有小调整,只要修改对应层级的遍历逻辑即可,维护成本很低。

内容的提问来源于stack exchange,提问作者futuredataengineer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 04:06:07