如何从含多条目的DataFrame行提取字符串并筛选美国临床站点?
处理格式混乱的CSV临床试验站点数据:拆分条目并筛选美国站点
需求说明
你的CSV数据存在格式问题:部分行包含多个用|分隔的临床试验站点,部分站点地址被换行拆分,需要完成以下操作:
- 将多行多条目拆分为单个站点条目
- 筛选仅位于美国的站点
- 生成唯一站点列表
解决方案(基于Tidyverse)
以下是分步实现的代码,适配你的数据格式:
1. 读取并预处理数据
先清理数据中的行号、空行,合并被换行拆分的站点地址:
library(tidyverse) # 读取所有行(替换为你的文件路径) lines <- read_lines("your_clinical_sites.csv") # 过滤无效行:空行、表头、纯数字行号 clean_lines <- lines %>% str_trim() %>% keep(~ .x != "" && .x != "Locations" && !str_detect(.x, "^\\d+$")) # 合并被换行拆分的站点(比如跨两行的加拿大站点地址) merged_lines <- vector("character", length(clean_lines)) i <- 1 while(i <= length(clean_lines)) { current_line <- clean_lines[i] # 若当前行不是美国站点且下一行存在,合并两行 if(i < length(clean_lines) && !str_detect(current_line, "United States$")) { current_line <- str_c(current_line, clean_lines[i+1], sep = " ") i <- i + 2 } else { i <- i + 1 } merged_lines[i-1] <- current_line } # 去除NA值 merged_lines <- merged_lines[!is.na(merged_lines)] # 转换为数据框 site_df <- tibble(location = merged_lines)
2. 拆分多条目行并清洗
将每行的多个站点用|拆分,去除每个条目的前后空白:
split_sites <- site_df %>% separate_rows(location, sep = "\\|") %>% mutate(location = str_trim(location))
3. 筛选美国站点并去重
提取仅包含美国地址的站点,生成唯一列表:
us_unique_sites <- split_sites %>% filter(str_detect(location, "United States$")) %>% distinct(location) %>% arrange(location)
4. 输出结果
查看或导出处理后的站点列表:
# 打印所有结果 print(us_unique_sites, n = Inf) # 导出到文本文件(可选) write_lines(us_unique_sites$location, "us_clinical_sites_unique.txt")
效果验证
处理后会得到符合你需求的唯一美国临床试验站点列表,例如:
University of Pennsylvania, Philadelphia, Pennsylvania, United States
University of Texas Southwestern Medical Center - Dallas, Dallas, Texas, United States
Houston Methodist Cancer Center, Houston, Texas, United States
Hem-Onc Associates of the Treasure Coast, Port Saint Lucie, Florida, United States
Moffitt Cancer Center, Tampa, Florida, United States
...
内容的提问来源于stack exchange,提问作者rumple.weed37
相关产品推荐
相关产品推荐

