如何高效合并DataFrame中同核心词变体的频次并移除变体行
Great question! Your current approach gets the job done, but we can streamline this into a more readable, maintainable workflow—no more manual row indexing or separate update steps. Let’s walk through a couple of better options:
方法1:用dplyr(推荐,语法直观易扩展)
If you’re using the tidyverse, dplyr lets you handle grouping, aggregation, and cleanup in a single pipeline. This is especially handy if you later need to handle multiple core words instead of just "oh".
library(dplyr) # 第一步:标记哪些词属于"oh"的拼写变体 is_oh_variant <- grepl("^o[oh-]+$", test$Word) # 第二步:分组聚合,合并变体到核心词 test_clean <- test %>% # 创建分组键:变体归到"oh"组,其他词用自身作为组 mutate(group_key = ifelse(is_oh_variant, "oh", Word)) %>% group_by(group_key) %>% summarise( # 优先保留核心词"oh",非变体组保留原词 Word = first(Word[Word == "oh"]) %||% first(Word), # 累加频次列 Freq = sum(Freq), Freq_BNCc = sum(Freq_BNCc), # 保留核心词的c5标签,非变体组保留原标签 c5 = first(c5[Word == "oh"]) %||% first(c5) ) %>% ungroup() %>% select(-group_key) # 移除临时分组列 # 查看结果 test_clean
This will give you exactly the output you want, and it’s easy to adjust if you need to add more core words later (just expand the group_key logic).
方法2:基础R实现(无需额外包)
If you prefer sticking to base R, we can restructure your original logic into a cleaner, less error-prone flow:
# 标记"oh"变体和核心词 is_oh_variant <- grepl("^o[oh-]+$", test$Word) core_oh <- test$Word == "oh" # 拆分数据:核心行、变体行、非变体行 core_row <- test[core_oh, ] variant_rows <- test[is_oh_variant & !core_oh, ] non_variant_rows <- test[!is_oh_variant, ] # 累加变体的频次到核心行 core_row$Freq <- core_row$Freq + sum(variant_rows$Freq) core_row$Freq_BNCc <- core_row$Freq_BNCc + sum(variant_rows$Freq_BNCc) # 合并核心行和非变体行,还原原顺序 test_clean <- rbind(core_row, non_variant_rows) test_clean <- test_clean[match(c("oh", "right-oh", "o'clock", "o-b-i-t-r-y"), test_clean$Word), ] # 查看结果 test_clean
为什么这些方法比原来的更好?
- Less error-prone: No manual row indexing (like
test[-which(...)]) which can break if there are no matches (sincewhich()returns an empty vector). - More readable: The logic is explicit—anyone reading the code can immediately see we’re grouping variants and aggregating counts.
- Scalable: If you need to handle multiple core words (e.g., "yeah" and its variants), you can easily adjust the grouping logic instead of writing separate code for each core word.
内容的提问来源于stack exchange,提问作者Chris Ruehlemann

