R语言转换CSV数据适配ggsankey的物种符号匹配问题
修正ggsankey所需DataFrame的转换代码
要生成ggsankey要求的包含x、node、next_x、next_node的DataFrame,核心是按物种顺序逐对匹配相邻物种的symbol,并单独处理最后一个物种的NA填充。以下是针对宽格式CSV数据的修正方案(假设你的CSV列是物种名,每行对应不同symbol的匹配关系):
修正后的代码
# 读取原始CSV数据(替换为你的文件路径) raw_data <- read.csv("your_data.csv", stringsAsFactors = FALSE) # 获取物种列表(CSV的列名即为物种顺序) species_order <- colnames(raw_data) # 初始化结果数据集 sankey_df <- data.frame() # 生成相邻物种的连接行 for (i in seq_along(species_order)[-length(species_order)]) { current_species <- species_order[i] next_species <- species_order[i + 1] # 提取当前物种和下一个物种的symbol列 current_nodes <- raw_data[[current_species]] next_nodes <- raw_data[[next_species]] # 构建当前连接段的数据集 segment_df <- data.frame( x = current_species, node = current_nodes, next_x = next_species, next_node = next_nodes, stringsAsFactors = FALSE ) sankey_df <- rbind(sankey_df, segment_df) } # 处理最后一个物种的行(next_x和next_node设为NA) last_species <- species_order[length(species_order)] last_nodes <- raw_data[[last_species]] last_segment <- data.frame( x = last_species, node = last_nodes, next_x = NA_character_, next_node = NA_character_, stringsAsFactors = FALSE ) sankey_df <- rbind(sankey_df, last_segment)
关键修正点
- 物种顺序匹配:直接用CSV列名作为
x和next_x的物种顺序,避免手动指定导致的顺序错误 - symbol一一对应:按行提取相邻物种的symbol列,保证
node和next_node是匹配的对应关系,解决了之前匹配错误的问题 - NA值正确设置:单独处理最后一个物种的所有行,统一将
next_x和next_node设为NA_character_(字符型NA,避免类型警告),解决了最后一行NA设置混乱的问题
示例验证
假设你的原始CSV是:
SpeciesA,SpeciesB,SpeciesC GeneX,GeneP,GeneM GeneY,GeneQ,GeneN GeneZ,GeneP,GeneO
运行代码后会生成符合要求的输出:
x node next_x next_node 1 SpeciesA GeneX SpeciesB GeneP 2 SpeciesA GeneY SpeciesB GeneQ 3 SpeciesA GeneZ SpeciesB GeneP 4 SpeciesB GeneP SpeciesC GeneM 5 SpeciesB GeneQ SpeciesC GeneN 6 SpeciesB GeneP SpeciesC GeneO 7 SpeciesC GeneM <NA> <NA> 8 SpeciesC GeneN <NA> <NA> 9 SpeciesC GeneO <NA> <NA>
内容的提问来源于stack exchange,提问作者Kshitij Behera
相关产品推荐
相关产品推荐

