You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言中按Type列规则清洗DataFrame的Name列?

问题描述

首先构造目标DataFrame的R代码:

Name <- c("cure center state hospital","state polyclinic cure center","state hospital main dispancer","first hospital number one")
Type <- c("hospital", "center", "dispancer", "hospital")
df <- data.frame(Name, Type, stringsAsFactors = FALSE)

同时有医疗机构类型列表:

types_list <- c("hospital", "center", "dispancer", "polyclinic")

需求规则:

  • 检查Name列中是否存在types_list内、且与当前行Type值不一致的类型词汇
  • 若存在该类词汇,删除最后一个这类词汇之前的所有内容(保留该词汇之后的部分)
  • 若不存在,保留原Name内容

期望输出的结果如下:

NameTypeName after
cure center state hospitalhospitalstate hospital
state polyclinic cure centercentercure center
state hospital main dispancerdispancermain dispancer
first hospital number onehospitalfirst hospital number one
解决方案

以下提供两种R实现方式:

方式一:使用tidyverse工具链(dplyr + stringr)

# 加载依赖包
library(dplyr)
library(stringr)

# 构造数据
Name <- c("cure center state hospital","state polyclinic cure center","state hospital main dispancer","first hospital number one")
Type <- c("hospital", "center", "dispancer", "hospital")
df <- data.frame(Name, Type, stringsAsFactors = FALSE)

types_list <- c("hospital", "center", "dispancer", "polyclinic")

# 定义处理逻辑函数
process_name <- function(name, type, types) {
  # 筛选出当前Type之外的目标类型词汇
  target_types <- setdiff(types, type)
  # 构建匹配整个单词的正则表达式
  pattern <- str_c("\\b(", str_c(target_types, collapse = "|"), ")\\b")
  # 获取所有匹配位置
  matches <- str_locate_all(name, pattern)[[1]]
  
  if (nrow(matches) > 0) {
    # 截取最后一个匹配词汇之后的内容并去除首尾空格
    str_trim(str_sub(name, matches[nrow(matches), "end"] + 1))
  } else {
    # 无匹配则返回原内容
    name
  }
}

# 应用函数生成新列
df <- df %>%
  mutate(Name_after = mapply(process_name, Name, Type, MoreArgs = list(types = types_list)))

# 查看结果
print(df)

运行后输出:

Name      Type             Name_after
1    cure center state hospital  hospital         state hospital
2  state polyclinic cure center    center            cure center
3 state hospital main dispancer dispancer         main dispancer
4     first hospital number one  hospital first hospital number one

方式二:基础R实现(无需额外包)

# 构造数据
Name <- c("cure center state hospital","state polyclinic cure center","state hospital main dispancer","first hospital number one")
Type <- c("hospital", "center", "dispancer", "hospital")
df <- data.frame(Name, Type, stringsAsFactors = FALSE)

types_list <- c("hospital", "center", "dispancer", "polyclinic")

# 批量处理生成新列
df$Name_after <- mapply(function(name, type) {
  # 筛选目标类型词汇
  target_types <- setdiff(types_list, type)
  # 拆分Name为单词列表
  words <- strsplit(name, " ")[[1]]
  # 找到所有属于目标类型的单词索引
  target_indices <- which(words %in% target_types)
  
  if (length(target_indices) > 0) {
    # 取最后一个目标单词之后的所有单词拼接成字符串
    paste(words[(max(target_indices) + 1):length(words)], collapse = " ")
  } else {
    name
  }
}, df$Name, df$Type)

# 查看结果
print(df)

该方法通过拆分单词、定位目标词汇位置的逻辑实现,结果与方式一完全一致。

内容的提问来源于stack exchange,提问作者Polina Ermolaeva

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 18:42:52