如何使用json_normalize展平嵌套JSON,提取uuid与指定日期CardMarket价格?
你可以直接按路径手动提取所需字段,比调整json_normalize参数更高效,也不会出现大宽表的问题,完整代码如下:
import json import pandas as pd # 配置目标参数,可根据需求修改 TARGET_DATE = "2021-11-07" PRICE_KEYS = ["paper", "cardmarket", "retail", "normal", TARGET_DATE] RESULT_COL_NAME = f"paper.cardmarket.retail.normal.{TARGET_DATE}" # 加载JSON数据 with open("AllPrices.json", "r", encoding="utf-8") as f: raw_data = json.load(f) card_price_dict = raw_data["data"] # 遍历提取所需字段 output_data = [] for uuid, price_info in card_price_dict.items(): # 逐层取嵌套值,缺失则返回空 current_val = price_info target_price = None try: for k in PRICE_KEYS: current_val = current_val[k] target_price = current_val except KeyError: pass output_data.append({ "uuid": uuid, RESULT_COL_NAME: target_price }) # 转为DataFrame,可按需删除价格为空的行 df = pd.DataFrame(output_data) # 可选:过滤无价格的记录 # df = df.dropna(subset=[RESULT_COL_NAME])
如果需要用json_normalize实现,可先将uuid转为字段再处理:
# 将UUID从字典键转为记录字段 processed_list = [{"uuid": k, **v} for k, v in card_price_dict.items()] df = pd.json_normalize(processed_list)[["uuid", RESULT_COL_NAME]]
该方法会先展开所有嵌套字段,内存开销比手动提取高,仅适合数据量较小的场景。
内容的提问来源于stack exchange,提问作者fab
相关产品推荐
相关产品推荐

