You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何合并R语言数据框同一列的重复类别?以Education列为例

解决R中数据框重复类别合并问题

方法1:直接合并重复类别

如果已经明确需要合并的目标类别,可以用以下两种方式统一:

方式A:使用dplyr包

library(dplyr)

df_all <- df_all %>%
  mutate(Education = recode(Education,
                            # 将所有该类的变体映射为统一名称
                            "Less than Primary" = "Less than Primary",
                            "Tertiary or above" = "Tertiary or above"))

# 验证合并结果
table(df_all$Education)

方式B:使用基础R(无需额外包)

# 合并"Less than Primary"的所有变体
df_all$Education[grepl("Less than Primary", df_all$Education)] <- "Less than Primary"
# 合并"Tertiary or above"的所有变体
df_all$Education[grepl("Tertiary or above", df_all$Education)] <- "Tertiary or above"

# 验证合并结果
table(df_all$Education)

方法2:排查重复类别的真实差异

若想明确看似相同的类别为何被识别为不同,可以查看字符串的原始字节编码,找出潜在的不可见字符或编码差异:

# 获取Education列的唯一值
unique_edu <- unique(df_all$Education)

# 查看每个唯一值的原始字节
lapply(unique_edu, charToRaw)

根据输出的字节差异,针对性清理,比如移除不可见控制字符:

# 移除所有非打印字符
df_all$Education <- gsub("[^[:print:]]", "", df_all$Education)

# 验证结果
table(df_all$Education)

内容的提问来源于stack exchange,提问作者oohsehun

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 04:20:30