You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于模式向量批量清洗数据框中的公司名称

批量映射不规范公司名称到规范名称

核心思路

通过构建映射字典+数据关联的方式实现批量匹配,无需手动维护case_when的匹配规则,新增关键词时仅需更新keywords向量即可。

实现步骤

  1. 基于keywords生成映射字典:包含清洗后的匹配键(统一小写、去除空格)和对应的规范名称
  2. 对原数据的不规范名称做相同规则的清洗
  3. 通过关联操作匹配规范名称,未匹配项可保留原名称或设为NA

代码示例

1. 加载依赖包

library(dplyr)
library(stringr)

2. 定义关键词与生成映射字典

如果规范名称有自定义格式(如样本中的Daveabc),可直接指定;若仅需首字母大写,用str_to_title自动处理:

# 目标关键词向量
keywords <- c("johnscompany","daveabc")

# 自定义规范名称(按需修改)
custom_clean_names <- c("Johnscompany", "Daveabc")

# 生成映射字典
mapping_df <- tibble(
  # 生成统一的匹配键:小写+去空格
  stripped = tolower(str_replace_all(keywords, fixed(" "), "")),
  # 绑定对应的规范名称
  clean_names = custom_clean_names
)

3. 批量清洗与匹配

# 样本数据
df <- structure(list(unclean_names = c("JohnscompaNy", "Johns company", 
"Dave ABC", "daveabc")), class = "data.frame", row.names = c(NA, -4L))

# 处理原数据
df_cleaned <- df %>%
  # 对不规范名称执行相同清洗规则,生成匹配键
  mutate(stripped = tolower(str_replace_all(unclean_names, fixed(" "), ""))) %>%
  # 关联映射字典,匹配规范名称
  left_join(mapping_df, by = "stripped") %>%
  # 未匹配项保留原名称(可选:若不需要可删除此步,未匹配项为NA)
  mutate(clean_names = if_else(is.na(clean_names), unclean_names, clean_names))

运行结果

> df_cleaned
  unclean_names      stripped clean_names
1   JohnscompaNy johnscompany Johnscompany
2  Johns company johnscompany Johnscompany
3       Dave ABC      daveabc      Daveabc
4        daveabc      daveabc      Daveabc

优势说明

  • 扩展性强:新增公司关键词时,只需更新keywords和对应的custom_clean_names,无需修改匹配逻辑代码
  • 逻辑清晰:通过数据关联替代硬编码的条件判断,可读性和维护性更高
  • 灵活适配:可根据需求调整清洗规则(如去除特殊字符)或规范名称格式

内容的提问来源于stack exchange,提问作者A03

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 11:33:14