You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Network D3 Sankey图链接数据框构建:跳过NA节点与自动数值修正

解决方案

完整优化代码

library(dplyr)
library(tidyr)
library(networkD3)

# 示例数据
First_Contact <- c("UTC", "UTC", "111", "111")
Second_Contact <- c(NA, "ED - ED RV", "UTC", "UTC")
Third_Contact <- c(NA, NA, "ED - ED RV", "ED - ED RV")
Final_Pathway_Outcome <- c("Discharged", "Discharged", "Discharged", "Discharged")

df <- data.frame(First_Contact, Second_Contact, Third_Contact, Final_Pathway_Outcome)

# 处理流程:将每行转为节点序列,过滤NA,生成连续链接并聚合
links_processed <- df %>%
  mutate(row = row_number()) %>%
  # 转换为长格式,保留行号与环节顺序
  pivot_longer(cols = -row, names_to = "step", values_to = "node") %>%
  # 按患者行分组,过滤无效NA节点
  group_by(row) %>%
  filter(!is.na(node)) %>%
  # 生成下一个有效节点作为目标,记录环节序号
  mutate(target_node = lead(node),
         source_step = match(step, names(df)),
         target_step = source_step + 1) %>%
  # 移除无后续节点的出院节点
  filter(!is.na(target_node)) %>%
  ungroup() %>%
  # 为节点添加环节后缀,区分同名称不同环节的节点
  mutate(source = paste0(node, "_", source_step),
         target = paste0(target_node, "_", target_step)) %>%
  # 聚合重复流程,统计患者数量作为value
  group_by(source, target) %>%
  summarise(value = n(), .groups = "drop")

# 生成节点数据
nodes_df <- data.frame(
  name = unique(c(links_processed$source, links_processed$target)),
  label = unique(c(links_processed$source, links_processed$target))
)

# 匹配networkD3所需的0起始节点ID
links_processed <- links_processed %>%
  mutate(source_id = match(source, nodes_df$name) - 1,
         target_id = match(target, nodes_df$name) - 1)

# 绘制Sankey图
sankeyNetwork(Links = links_processed, 
              Nodes = nodes_df, 
              Source = 'source_id', 
              Target = 'target_id', 
              Value = 'value', 
              NodeID = 'label', 
              fontSize = 16, 
              iterations = 0)

问题1:优雅处理NA节点

  • 核心逻辑:按患者行分组后,先过滤所有NA节点,再用lead()生成下一个有效节点的链接,直接跳过无效的中间NA环节。
  • 效果:像UTC -> NA -> NA -> Discharged这类流程会自动转换为UTC_1 -> Discharged_4,完全无需手动修改数据。
  • 关键处理代码:
    group_by(row) %>%
    filter(!is.na(node)) %>%
    mutate(target_node = lead(node),
           source_step = match(step, names(df)),
           target_step = source_step + 1) %>%
    filter(!is.na(target_node)) %>%
    ungroup()
    

问题2:自动聚合重复流程

  • 核心逻辑:通过group_by(source, target)分组后,用summarise(value = n())自动统计相同流转路径的患者数量,替代原代码逐行赋值value=1的低效方式。
  • 效果:针对大规模数据集,能高效完成重复流程的聚合,无需手动处理重复行。
  • 关键聚合代码:
    group_by(source, target) %>%
    summarise(value = n(), .groups = "drop")
    

内容的提问来源于stack exchange,提问作者James Cai

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 05:53:19