基于R数据框识别独立出现的NIH/NIA并生成逻辑结果列
识别Funding列中单独出现的NIH/NIA并新增逻辑判断列
你需要在数据集的Funding列中,找出满足以下任一条件的行,并新增一列存储TRUE/FALSE的逻辑判断结果:
- 行首单独出现NIH或NIA
- NIH或NIA单独位于两个分号之间(前后带有空格)
解决思路
用正则表达式精准匹配目标场景,结合R的grepl()函数实现逻辑判断。正则需覆盖三种核心场景:
- 行首直接是NIH或NIA:
^NIH|^NIA - 被分号+空格包裹的NIH或NIA:
; (NIH|NIA); - 出现在行尾、前面是分号+空格的NIH或NIA:
; (NIH|NIA)$
将这三种情况用|拼接,即可覆盖所有符合要求的场景。
代码实现
首先是你提供的测试数据框:
df <- data.frame(Funding = c( "NHI; National Institutes of Health, NIH; National Institute on Aging, NIA; Fogarty International Center, FIC; Institute for Translational Neuroscience, ITN; International Brain Research Organization, IBRO; Sveriges Tandläkarförbund, SDA; NIA; National Institute of Neurological Disorders and Stroke", "Canadian Partnership for Stroke Recovery; EuroImmun; National Institutes of Health, NIH; U.S. Department of Defense, DOD; National Institute on Aging, NIA; National Institute of Biomedical Imaging and Bioengineering, NIBIB; Michael J. Fox Foundation for Parkinson's Research, MJFF; Alzheimer's Association, AA; Alzheimer's Disease Neuroimaging Initiative, ADNI; Fondation Brain Canada; Northern California Institute for Research and Education, NCIRE; Sunnybrook Research Institute, SRI; Weston Brain Institute, WBI; Canadian Institutes of Health Research, IRSC; Alzheimer’s Research UK, ARUK", "National Institutes of Health, NIH; National Institute on Aging, NIA; National Institute of Diabetes and Digestive and Kidney Diseases, NIDDK; National Institute of Child Health and Human Development, NICHD; National Institute of Child Health and Human Development", "Institute for Neurological Research; Research and Development Grants for Dementia; NIA; Alzheimer's Association, AA; Japan Agency for Medical Research and Development, AMED; Consejo Nacional de Investigaciones Científicas y Técnicas, CONICET; Korea Health Industry Development Institute, KHIDI; Deutsches Zentrum für Neurodegenerative Erkrankungen, DZNE", "National Institutes of Health, NIH; National Institute on Aging, NIA; Alzheimer's Association, AA; Global Brain Health Institute; Rainwater Charitable Foundation, RCF" ))
执行以下代码新增逻辑判断列:
# 新增逻辑判断列,标记是否存在单独出现的NIH/NIA df$has_standalone_NIH_NIA <- grepl("^NIH|^NIA|; (NIH|NIA);|; (NIH|NIA)$", df$Funding)
结果说明
运行后查看df,会看到:
- 第1行和第4行的
has_standalone_NIH_NIA值为TRUE(第1行包含; NIA;,第4行包含; NIA;) - 其余行均为
FALSE(这些行的NIH/NIA是附在机构全称后的,并非单独出现)
内容的提问来源于stack exchange,提问作者always.learning
相关产品推荐
相关产品推荐

