基于最高百分比将行名映射为列名的R语言实现需求
解决列名替换需求的R实现方案
首先先明确你的输入数据和现有函数:
输入数据
voyages <- c("VIC0016", "VIC0016", "VIC0016", "VIC0016", "VIC0016", "VIC0016", "Truck", "VIC0016", "VIC0016", "VIC0016", "JUL0983", "BB11356", "VIC0022", "VIC0022", "ISK1981", "ISK1981", "ISK1981", "ISK1981", "ISK1981", "ISK1981", "ISK1981", "ISK1981", "ISK1981", "ISK1981", "ISK1981", "ISK1981") clusters <- c(5, 5, 5, 4, 4, 4, 1, 3, 4, 3, 5, 2, 4, 5, 6, 6, 6, 6, 6, 6, 6, 6, 6, 6, 6, 6)
现有生成百分比矩阵的函数
calculate.confusion <- function(voyages, clusters) { d <- data.frame(voyages, clusters) td <- as.data.frame(table(d)) # convert the raw counts into percentage of each voyage number pc <- matrix(ncol=max(clusters),nrow=0) for (i in 1:11) # 11 different voyage numbers { total <- sum(td[td$voyages==td$voyages[i],3]) #,3 is the third column, showing the frequencies pc <- rbind(pc, td[td$voyages==td$voyages[i],3]/total) } rownames(pc) <- td[1:11,1] colnames(pc)<-1:11 return(pc) }
先运行函数得到初始矩阵:
confusion_mat <- calculate.confusion(voyages, clusters)
接下来针对你的两个需求(每行最高百分比列用该行名命名、每个行名仅用一次),提供两种实现方案:
方案1:基础映射(先到先得处理冲突)
这个方案先给每行的最高百分比列分配行名,遇到多个行争夺同一列时,保留第一个匹配的行名,剩余行名分配给未被使用的列:
# 1. 找出每行百分比最高的列索引 max_col_indices <- apply(confusion_mat, 1, which.max) # 2. 建立初始映射:列索引 -> 对应的行名 col_row_mapping <- setNames(names(max_col_indices), max_col_indices) # 3. 处理未被分配的行和列 used_row_names <- unique(names(max_col_indices)) remaining_rows <- setdiff(rownames(confusion_mat), used_row_names) remaining_cols <- setdiff(colnames(confusion_mat), names(col_row_mapping)) # 给剩余列分配剩余行名(按顺序) col_row_mapping <- c(col_row_mapping, setNames(remaining_rows, remaining_cols)) # 4. 替换矩阵列名 colnames(confusion_mat) <- col_row_mapping[colnames(confusion_mat)]
方案2:优化映射(优先分配百分比最高的匹配)
如果出现多个行的最高百分比列相同,这个方案会为该列选择百分比最高的那个行名,剩余的行再去匹配次高百分比的列,逻辑更严谨:
需要用到tidyverse工具包,如果没安装先运行install.packages("tidyverse"):
library(tidyverse) # 1. 将矩阵转成长格式,方便按列(cluster)筛选最高百分比的行 long_format <- confusion_mat %>% as.data.frame() %>% mutate(voyage = rownames(.)) %>% pivot_longer(-voyage, names_to = "cluster", values_to = "percentage") %>% arrange(cluster, desc(percentage)) # 按cluster分组,百分比从高到低排序 # 2. 为每个cluster选择百分比最高的唯一voyage unique_mapping <- long_format %>% group_by(cluster) %>% slice(1) %>% # 取每组第一个(即百分比最高的) ungroup() %>% select(cluster, voyage) %>% deframe() # 转换为cluster -> voyage的映射 # 3. 处理剩余未分配的行和列 unused_voyages <- setdiff(rownames(confusion_mat), unique_mapping) unused_clusters <- setdiff(colnames(confusion_mat), names(unique_mapping)) # 分配剩余的对应关系(可根据需求调整顺序逻辑) unique_mapping <- c(unique_mapping, setNames(unused_voyages, unused_clusters)) # 4. 替换列名 colnames(confusion_mat) <- unique_mapping[colnames(confusion_mat)]
运行后,你的矩阵列名就会按照要求替换为voyage编号,且每个行名只出现一次。
内容的提问来源于stack exchange,提问作者Fleur Lolkema
相关产品推荐
相关产品推荐

