如何筛选位于美国马里兰州贝塞斯达NINDS的临床试验?
修正NINDS临床试验数据筛选的方法
先排查拆分后地点字段的结构,再针对不同场景调整筛选逻辑:
1. 先确认数据结构
先打印字段样例,明确拆分后的数据类型:
# 查看前5行地点字段的格式 print(df['locations'].head())
2. 针对不同结构的筛选方案
场景1:地点字段是字符串列表(每个元素为地点字典/字符串)
如果每个地点是包含机构名、城市、州的字典(比如[{"name": "National Institute of Neurological Disorders and Stroke", "city": "Bethesda", "state": "MD"}, ...]),用apply遍历列表做精准匹配:
# 匹配机构全称+贝塞斯达+马里兰州,兼容NINDS缩写 ninds_trials = df[df['locations'].apply(lambda loc_list: any( (loc.get('name') in ['National Institute of Neurological Disorders and Stroke', 'NINDS']) and loc.get('city') == 'Bethesda' and loc.get('state') == 'MD' for loc in loc_list ))]
场景2:地点字段是分隔符拼接的字符串
统一转为小写消除大小写差异,再用关键词组合筛选:
# 统一转小写,避免匹配遗漏 df['locations_lower'] = df['locations'].str.lower() # 同时匹配贝塞斯达、马里兰州、NINDS相关关键词 ninds_trials = df[ df['locations_lower'].str.contains('bethesda') & df['locations_lower'].str.contains('md|maryland') & df['locations_lower'].str.contains('ninds|national institute of neurological disorders and stroke') ]
场景3:地点信息分散在多列
如果机构名、城市、州分别存在独立列(比如institution、city、state),直接组合条件:
ninds_trials = df[ (df['institution'].str.lower().str.contains('ninds|national institute of neurological disorders and stroke')) & (df['city'].str.lower() == 'bethesda') & (df['state'].str.lower() == 'md') ]
场景4:地点字段是多层嵌套JSON
先用json_normalize拉平嵌套结构,再筛选:
from pandas import json_normalize # 拉平地点字段,生成每个地点单独的行 flat_locations = json_normalize(df['locations']) # 合并原数据与拉平后的地点信息 df_flat = df.join(flat_locations) # 筛选目标地点 ninds_trials = df_flat[ (df_flat['name'].str.lower().str.contains('ninds')) & (df_flat['city'].str.lower() == 'bethesda') & (df_flat['state'].str.lower() == 'md') ]
常见避坑点
- 必须统一大小写:比如有的记录写
bethesda、有的写Bethesda,转小写/大写后再匹配 - 兼容机构名称的不同表述:比如全称、缩写
NINDS都要覆盖 - 多地点记录要确保至少有一个地点符合条件,不能用全量匹配
内容的提问来源于stack exchange,提问作者Rstudyer
相关产品推荐
相关产品推荐

