在R中识别逗号分隔颜色组重复项并合并计数的高效方法
R中合并无序颜色组合的计数求和
原始数据集
sample_table = data.frame( colors = c("red", "blue", "red,blue", "blue, red"), counts = c(12, 10, 5,6) )
输出结果:
colors counts 1 red 12 2 blue 10 3 red,blue 5 4 blue, red 6
需求说明
需要将颜色顺序不同的组合(比如blue, red和red, blue,以及多颜色组合如red, blue, green与green, blue, red)视为同一组,对其counts字段求和合并,最终得到如下结果:
colors counts 1 red 12 2 blue 10 3 red,blue 11
现有实现代码
你已手动实现需求,代码如下:
standardize_colors <- function(color_string) { colors <- unlist(strsplit(color_string,",\\s*")) return(paste(sort(colors), collapse = ",")) } sample_table$standardized_colors <- sapply(sample_table$colors, standardize_colors) aggregated_table <- aggregate(counts ~ standardized_colors, data = sample_table, sum) print(aggregated_table)
更高效的实现方式
你的核心逻辑(拆分颜色字符串→排序→重组标准化字符串)是处理这类无序分组问题的通用正确方案,以下是几种更高效或更简洁的实现:
1. tidyverse工具链(dplyr + stringr)
适合管道式操作场景,代码更简洁,大数据集下性能优于基础R的sapply:
library(tidyverse) sample_table %>% mutate( standardized_colors = str_split(colors, ",\\s*") %>% map(sort) %>% map_chr(paste, collapse = ",") ) %>% group_by(standardized_colors) %>% summarise(counts = sum(counts)) %>% rename(colors = standardized_colors)
2. data.table(极致性能)
针对百万行以上的大规模数据集,data.table的运行速度会显著优于基础R和tidyverse:
library(data.table) setDT(sample_table)[, standardized_colors := paste(sort(unlist(strsplit(colors, ",\\s*"))), collapse = ","), by = seq_len(nrow(sample_table)) ][, .(counts = sum(counts)), by = standardized_colors][, colors := standardized_colors][, standardized_colors := NULL][]
3. 基础R优化版
用vapply替代sapply,指定返回类型后性能略优且更安全:
standardize_colors <- function(color_string) { paste(sort(strsplit(color_string, ",\\s*")[[1]]), collapse = ",") } sample_table$standardized_colors <- vapply(sample_table$colors, standardize_colors, character(1)) aggregated_table <- aggregate(counts ~ standardized_colors, sample_table, sum) colnames(aggregated_table)[1] <- "colors"
以上实现均保留了你核心的标准化逻辑,同时在代码简洁性或性能上做了优化,且都支持任意数量的颜色组合场景。
内容的提问来源于stack exchange,提问作者wulasa
相关产品推荐
相关产品推荐

