Python中从JSON提取多组织指定字段(含多位置ID)的方法
问题:遍历JSON数据提取所有组织的指定字段
需求说明
我有如下JSON结构的数据(加载后d是列表类型,d[0]为字典),需要提取所有组织的以下字段:
code(作为org_id)name(作为org_name)- 所有关联的
location_id status
原始JSON结构(示例)
{ "organisations": { "total-items": "41477", "organisation": [ { "mappings": null, "active-request": "false", "identifiers": { "identifier": { "code": "ORG-100023310", "code-system": "100000167446", "code-system-name": " OMS Organization Identifier" } }, "name": "advanceCOR GmbH", "operational-attributes": { "created-on": "2016-10-18T15:38:34.322+02:00", "modified-on": "2022-11-02T08:23:13.989+01:00" }, "locations": { "location": [ { "location-id": { "link": { "href": "https://v1/locations/LOC-100052061" }, "id": "LOC-100052061" } }, { "location-id": { "link": { "href": "https://v1/locations/LOC-100032442" }, "id ": "LOC-100032442" } }, { "location-id": { "link": { "href": "https://v1/locations/LOC-100042003" }, "id": "LOC-100042003" } } ] }, "organisation-id": { "link": { "rel": "self", "href": "https://v1 /organisations/ORG-100023310" }, "id": "ORG-100023310" }, "status": "ACTIVE" }, { "mappings": null, "active-request": "false", "identifiers": { "identifier": { "code": "ORG-100004261", "code-system": "100000167446", "code-system-name": "OMS organization Identifier" } }, "name": "Beacon Pharmaceuticals Limited", "operational-attributes": { "created-on": "2016-10-18T14:48:16.293+02:00", "modified-on": "2022-10-12T08:26:24.645+02:00" }, "locations": { "location": [ { "location-id": { "link": { "href": "https://v1/locations/LOC-100005615" }, "id": "LOC-100005615" } }, { "location-id": { "link": { "href": "https://v1/locations/LOC-100000912" }, "id": "LOC-100000912" } }, { "location-id": { "link": { "href": "https://v1/locations/LOC-100043831" }, "id": "LOC-100043831" } } ] }, "organisation-id": { "link": { "rel": "self", "href": "https://v1/organisations/ORG-100004261" }, "id": "ORG-100004261" }, "status": "ACTIVE" } ] } }
现有代码问题
当前代码只能获取单个组织的单个位置ID,无法遍历所有组织及全部位置ID:
with open('organisations.json', encoding='utf-8') as f: d = json.load(f) print(d[0]['organisations']['organisation'][2]['identifiers']['identifier']['code']) #code print(d[0]['organisations']['organisation'][2]['name']) #name print(d[0]['organisations']['organisation'][2]['locations']) #location print(d[0]['organisations']['organisation'][2]['status']) #status
当前输出
ORG-100023310 advanceCOR GmbH {'location': {'location-id': {'link': {'href': 'https://v1/locations/LOC-100052061'}, 'id': ' LOC-100052061'}}} ACTIVE
预期输出
org_id org_name location_id status ORG-100023310 advanceCOR GmbH LOC-100052061, LOC-100032442, LOC-100042003 ACTIVE ORG-100004261 Beacon Pharmaceuticals Limited LOC-100005615, LOC-100000912, LOC-100043831 ACTIVE
解决方案
通过循环遍历所有组织,并对每个组织的位置列表再次循环提取ID,最后格式化输出即可实现需求:
修改后的代码
import json with open('organisations.json', encoding='utf-8') as f: d = json.load(f) # 获取完整的组织列表 organisations = d[0]['organisations']['organisation'] # 打印对齐的表头 print(f"{'org_id':<15} {'org_name':<30} {'location_id':<45} {'status'}") # 遍历每个组织 for org in organisations: # 提取组织ID org_id = org['identifiers']['identifier']['code'] # 提取组织名称 org_name = org['name'] # 提取所有关联的位置ID,兼容键名带空格的情况(如'id ') location_ids = [] for loc in org['locations']['location']: for key in loc['location-id']: if key.strip() == 'id': loc_id = loc['location-id'][key].strip() location_ids.append(loc_id) # 将位置ID用逗号拼接成字符串 location_str = ', '.join(location_ids) # 提取组织状态 status = org['status'] # 格式化输出对齐字段 print(f"{org_id:<15} {org_name:<30} {location_str:<45} {status}")
代码说明
- 遍历组织列表:用
for org in organisations替代固定索引访问,实现所有组织的遍历 - 提取位置ID:嵌套循环处理每个组织的位置列表,通过键名兼容处理JSON中存在的
'id '(带空格)问题 - 格式化输出:使用Python格式化字符串(
f-string)的对齐语法,让输出格式和预期一致 - 兼容性处理:通过
strip()去除ID字符串的首尾空格,避免输出杂乱
内容的提问来源于stack exchange,提问作者rshar
相关产品推荐
相关产品推荐

