如何在R语言中按Type列规则清洗DataFrame的Name列?
问题描述
首先构造目标DataFrame的R代码:
Name <- c("cure center state hospital","state polyclinic cure center","state hospital main dispancer","first hospital number one") Type <- c("hospital", "center", "dispancer", "hospital") df <- data.frame(Name, Type, stringsAsFactors = FALSE)
同时有医疗机构类型列表:
types_list <- c("hospital", "center", "dispancer", "polyclinic")
需求规则:
- 检查
Name列中是否存在types_list内、且与当前行Type值不一致的类型词汇 - 若存在该类词汇,删除最后一个这类词汇之前的所有内容(保留该词汇之后的部分)
- 若不存在,保留原
Name内容
期望输出的结果如下:
| Name | Type | Name after |
|---|---|---|
| cure center state hospital | hospital | state hospital |
| state polyclinic cure center | center | cure center |
| state hospital main dispancer | dispancer | main dispancer |
| first hospital number one | hospital | first hospital number one |
解决方案
以下提供两种R实现方式:
方式一:使用tidyverse工具链(dplyr + stringr)
# 加载依赖包 library(dplyr) library(stringr) # 构造数据 Name <- c("cure center state hospital","state polyclinic cure center","state hospital main dispancer","first hospital number one") Type <- c("hospital", "center", "dispancer", "hospital") df <- data.frame(Name, Type, stringsAsFactors = FALSE) types_list <- c("hospital", "center", "dispancer", "polyclinic") # 定义处理逻辑函数 process_name <- function(name, type, types) { # 筛选出当前Type之外的目标类型词汇 target_types <- setdiff(types, type) # 构建匹配整个单词的正则表达式 pattern <- str_c("\\b(", str_c(target_types, collapse = "|"), ")\\b") # 获取所有匹配位置 matches <- str_locate_all(name, pattern)[[1]] if (nrow(matches) > 0) { # 截取最后一个匹配词汇之后的内容并去除首尾空格 str_trim(str_sub(name, matches[nrow(matches), "end"] + 1)) } else { # 无匹配则返回原内容 name } } # 应用函数生成新列 df <- df %>% mutate(Name_after = mapply(process_name, Name, Type, MoreArgs = list(types = types_list))) # 查看结果 print(df)
运行后输出:
Name Type Name_after 1 cure center state hospital hospital state hospital 2 state polyclinic cure center center cure center 3 state hospital main dispancer dispancer main dispancer 4 first hospital number one hospital first hospital number one
方式二:基础R实现(无需额外包)
# 构造数据 Name <- c("cure center state hospital","state polyclinic cure center","state hospital main dispancer","first hospital number one") Type <- c("hospital", "center", "dispancer", "hospital") df <- data.frame(Name, Type, stringsAsFactors = FALSE) types_list <- c("hospital", "center", "dispancer", "polyclinic") # 批量处理生成新列 df$Name_after <- mapply(function(name, type) { # 筛选目标类型词汇 target_types <- setdiff(types_list, type) # 拆分Name为单词列表 words <- strsplit(name, " ")[[1]] # 找到所有属于目标类型的单词索引 target_indices <- which(words %in% target_types) if (length(target_indices) > 0) { # 取最后一个目标单词之后的所有单词拼接成字符串 paste(words[(max(target_indices) + 1):length(words)], collapse = " ") } else { name } }, df$Name, df$Type) # 查看结果 print(df)
该方法通过拆分单词、定位目标词汇位置的逻辑实现,结果与方式一完全一致。
内容的提问来源于stack exchange,提问作者Polina Ermolaeva
相关产品推荐
相关产品推荐

