R语言如何对每个品牌标识的色相变量分箱并统计各分箱观测次数
基础R实现(符合你要求的cut()+split()组合逻辑)
首先要注意cut()默认生成右闭左开区间,你需要的是左闭右开区间,必须加right = FALSE参数保证区间匹配正确。
# 你已定义的分箱参数 breaks <- c(0,45,90,135,180,225,270,315,360) labels <- c("[0-45)","[45-90)", "[90-135)", "[135-180)", "[180-225)", "[225-270)","[270-315)", "[315-360)") # 1. 对Hue变量分箱 df$bin <- cut(df$Hue, breaks = breaks, labels = labels, right = FALSE) # 2. 按行业+品牌标识拆分数据集 grouped_df <- split(df, list(df$Industry, df$Logo), drop = TRUE) # 3. 对每个分组统计各分箱的观测次数,合并结果 result <- do.call(rbind, lapply(grouped_df, function(sub_df) { # 统计当前品牌各分箱频次 bin_counts <- table(sub_df$bin) # 拼接分组标识和频次结果 cbind( Ind = unique(sub_df$Industry), Logo = unique(sub_df$Logo), as.data.frame(t(bin_counts)) ) })) # 重置行号 rownames(result) <- NULL
运行后得到的result就是你需要的宽格式数据框,输出示例:
| Ind | Logo | [0-45) | [45-90) | [90-135) | [135-180) | [180-225) | [225-270) | [270-315) | [315-360) |
|---|---|---|---|---|---|---|---|---|---|
| Fossil | Petrox | 3 | 0 | 0 | 0 | 2 | 0 | 0 | 0 |
| Renewable | Windo | 1 | 0 | 0 | 0 | 0 | 0 | 1 | 1 |
注:你给出的预期输出中Logo列写为Petrol、Wind属于笔误,原始数据中对应名称为Petrox、Windo,以上输出以原始数据为准。
更简洁的基础R实现(无需split,直接用table生成交叉表)
如果不需要强制用split(),可以用更短的代码实现相同效果:
df$bin <- cut(df$Hue, breaks = breaks, labels = labels, right = FALSE) count_mat <- table(df$Industry, df$Logo, df$bin) result <- cbind(unique(df[,c("Industry", "Logo")]), as.data.frame.matrix(count_mat)) colnames(result)[1] <- "Ind" rownames(result) <- NULL
内容的提问来源于stack exchange,提问作者CMJohannessen
相关产品推荐
相关产品推荐

