如何从嵌套JSON中过滤tag含now的条目提取types数据生成DataFrame
多层嵌套JSON筛选提取生成DataFrame方案
你可以按以下逻辑处理多层嵌套的JSON结构,核心是逐层遍历+条件过滤,下面是可直接运行的完整代码:
import pandas as pd # 原始JSON数据,若存储为本地文件可直接用 pd.read_json("文件路径") 读取 raw_json = { "products": [ { "id": 12121, "product": "hair", "tag":"now, later", "types": [ { "product_id": 11111, "id": 22222 } ], "options": [ { "name": "Title" } ] }, { "id": 1313131, "product": "pillow", "tag":"later, never", "types": [ { "product_id": 33333, "id": 44444 } ], "options": [ { "name": "Title" } ] }, { "id": 14141414, "product": "face", "tag":"now, never", "types": [ { "product_id": 5555, "id": 7777 } ], "options": [ { "name": "Title" } ] } ] } result = [] # 遍历最外层products数组 for product in raw_json["products"]: # 拆分tag字段判断是否包含now,先做去空格处理避免格式问题导致匹配失败 tags = [t.strip() for t in product["tag"].split(",")] if "now" in tags: # 遍历types数组提取目标字段 for type_item in product["types"]: result.append({ "tag": "now", "product_id": type_item["product_id"], "id": type_item["id"] }) # 转换为DataFrame df = pd.DataFrame(result) print(df)
运行后输出和你预期完全一致:
tag product_id id 0 now 11111 22222 1 now 5555 7777
核心逻辑说明
- 先对最外层
products数组做遍历,每条数据对应一个商品的完整信息 - 对
tag字段做拆分去空格处理,避免因为逗号后带空格、存在多余空格等问题导致匹配错误 - 符合条件的商品再遍历其下的
types数组,兼容单个/多个type的场景,按要求拼装数据 - 最终将结构化的结果列表直接转换为DataFrame即可
内容的提问来源于stack exchange,提问作者ApacheOne
相关产品推荐
相关产品推荐

