You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从批量字典数据中提取指定Key并存储为新列表?

从多行JSON数据中提取指定字段

问题背景

我有如下格式的多行JSON数据:

{"id": 0, "key1": value, "key2": "value", "imp_key": {"sub_key1": 5, "sub_key2": 10, "sub_key3": 15}}
{"id": 1, "key1": value, "key2": "value", "imp_key": {"sub_key1": 5, "sub_key2": 10, "sub_key3": 15}}
{"id": 2, "key1": value, "key2": "value", "imp_key": {"sub_key1": 5, "sub_key2": 10, "sub_key3": 15}}

需要从中提取id和imp_key字段,将结果存储为一个新的JSON列表,期望输出如下:

[{"id": 0, "imp_key": {"sub_key1": 5, "sub_key2": 10, "sub_key3": 15}},
     {"id": 1, "imp_key": {"sub_key1": 5, "sub_key2": 10, "sub_key3": 15}},
     {"id": 2, "imp_key": {"sub_key1": 5, "sub_key2": 10, "sub_key3": 15}}
    ]

数据集规模较大,但格式完全一致,imp_key下的子键值均为数值类型。

解决方案(Python实现)

针对大数据集,提供两种处理方式:

1. 内存充足场景:一次性加载处理

如果数据集可以完整放入内存,直接逐行解析并提取字段:

import json

extracted_list = []

# 读取源数据文件
with open('source_data.json', 'r') as f:
    for line in f:
        stripped_line = line.strip()
        if not stripped_line:
            continue
        # 解析单行JSON
        item = json.loads(stripped_line)
        # 提取目标字段
        extracted_item = {
            "id": item["id"],
            "imp_key": item["imp_key"]
        }
        extracted_list.append(extracted_item)

# 可选:将结果写入文件或打印
with open('result.json', 'w') as out_f:
    json.dump(extracted_list, out_f, indent=4)

2. 内存受限场景:流式处理

如果数据集过大无法一次性加载,采用流式读写,避免占用过多内存:

import json

with open('source_data.json', 'r') as infile, open('result.json', 'w') as outfile:
    outfile.write('[')
    is_first_item = True
    for line in infile:
        stripped_line = line.strip()
        if not stripped_line:
            continue
        try:
            item = json.loads(stripped_line)
            extracted_item = {
                "id": item["id"],
                "imp_key": item["imp_key"]
            }
            if not is_first_item:
                outfile.write(',\n')
            else:
                is_first_item = False
            # 写入格式化后的条目
            outfile.write(json.dumps(extracted_item, indent=4))
        except json.JSONDecodeError as e:
            print(f"跳过无效行: {stripped_line}, 错误信息: {str(e)}")
            continue
    outfile.write('\n]')

异常处理补充

为了增强鲁棒性,建议添加异常捕获,处理可能的JSON解析错误或字段缺失:

try:
    item = json.loads(stripped_line)
    # 检查必要字段是否存在
    if "id" not in item or "imp_key" not in item:
        print(f"跳过缺失字段的行: {stripped_line}")
        continue
    extracted_item = {
        "id": item["id"],
        "imp_key": item["imp_key"]
    }
    # 后续写入逻辑
except (json.JSONDecodeError, KeyError) as e:
    print(f"处理行失败: {stripped_line}, 错误: {str(e)}")
    continue

内容的提问来源于stack exchange,提问作者Ubuntu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 18:39:28