You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pd.json_normalize处理openFDA NDC数据时出现KeyError求助

关于pd.json_normalize处理FDA药品NDC数据时的KeyError问题

问题背景

处理FDA公开的药品NDC JSON数据集时,使用pd.json_normalize展开嵌套数据遇到KeyError:

  • 单独执行pd.json_normalize(data['results'], record_path=["packaging"], meta=['product_ndc'])可正常生成预期列
  • 但尝试将active_ingredients嵌套进record_path,或把brand_name、generic_name加入meta列表时,均触发KeyError

复现代码

import pandas as pd
import json
import requests, zipfile, io, os

cwd = os.getcwd()
zip_url = 'https://download.open.fda.gov/drug/ndc/drug-ndc-0001-of-0001.json.zip'
r = requests.get(zip_url)
z = zipfile.ZipFile(io.BytesIO(r.content))
z.extractall(cwd)

with open('drug-ndc-0001-of-0001.json', 'r') as file:
    data = json.load(file)

# 报错代码1:错误嵌套record_path
# result = pd.json_normalize(data['results'], record_path=["packaging","active_ingredients"],meta=['product_ndc','brand_name','generic_name'])

# 报错代码2:未处理缺失的meta字段
packaging_data = pd.json_normalize(
    data['results'], 
    record_path=["packaging"], 
    meta=['product_ndc', 'brand_name', 'generic_name']
)

active_ingredients_data = pd.json_normalize(
    data['results'], 
    record_path=["active_ingredients"], 
    meta=['product_ndc', 'brand_name', 'generic_name']
)

错误原因

  1. 路径嵌套错误:active_ingredients与packaging是results下的平级字段,并非嵌套在packaging内部,用["packaging","active_ingredients"]作为record_path会导致找不到对应键。
  2. meta字段缺失:数据集中部分results条目未包含brand_name或generic_name字段,pd.json_normalize默认要求所有指定的meta字段必须存在,否则抛出KeyError。

解决方案

1. 分开处理平级嵌套列表

packaging和active_ingredients是同级独立列表,需分别调用pd.json_normalize展开。

2. 忽略缺失的meta字段

添加errors='ignore'参数,允许部分条目缺失指定的meta字段,缺失值自动填充为NaN。

修改后的可运行代码

import pandas as pd
import json
import requests, zipfile, io, os

cwd = os.getcwd()
zip_url = 'https://download.open.fda.gov/drug/ndc/drug-ndc-0001-of-0001.json.zip'
r = requests.get(zip_url)
z = zipfile.ZipFile(io.BytesIO(r.content))
z.extractall(cwd)

with open('drug-ndc-0001-of-0001.json', 'r') as file:
    data = json.load(file)

# 处理packaging数据,忽略缺失的meta字段
packaging_data = pd.json_normalize(
    data['results'], 
    record_path=["packaging"], 
    meta=['product_ndc', 'brand_name', 'generic_name'],
    errors='ignore'
)

# 处理active_ingredients数据
active_ingredients_data = pd.json_normalize(
    data['results'], 
    record_path=["active_ingredients"], 
    meta=['product_ndc', 'brand_name', 'generic_name'],
    errors='ignore'
)

内容的提问来源于stack exchange,提问作者Brad

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 07:57:21