Python Plotly桑基图输出异常:节点与链接配置求助
问题
作为Python Plotly新手,我想创建桑基图展示各Topic的月度Size变化,已准备好如下pandas DataFrame:
year month Topic Size 1 2022 1 1.0 40 2 2022 1 2.0 36 3 2022 1 3.0 70 4 2022 1 4.0 42 5 2022 1 5.0 25 6 2022 1 6.0 48 7 2022 2 1.0 14 8 2022 2 2.0 34 9 2022 2 3.0 39 10 2022 2 4.0 210 11 2022 2 5.0 106 12 2022 2 6.0 70 13 2022 3 1.0 47 14 2022 3 2.0 45 15 2022 3 3.0 78 16 2022 3 4.0 114 17 2022 3 5.0 84 18 2022 3 6.0 78 19 2022 4 1.0 42 20 2022 4 2.0 45 21 2022 4 3.0 78 22 2022 4 4.0 28 23 2022 4 5.0 74 24 2022 4 6.0 44 25 2022 5 1.0 57 26 2022 5 2.0 45 27 2022 5 3.0 82 28 2022 5 4.0 94 29 2022 5 5.0 54 30 2022 5 6.0 24
参考示例写了代码,但不知道如何配置node_label、source_node、target_node变量,导致桑基图输出不正确,现有代码片段如下:
node_label = ? source_node = ? target_node = ? values = df['Size'] from webcolors import hex_to_rgb %matplotlib inline from plotly.offline import download_plotlyjs, init_notebook_mode, plot, iplot import plotly.graph_objects as go # Import the graphical object fig = go.Figure( data=[go.Sankey( # The plot we are interest # This part is for the node information node = dict( label = node_label ), # This part is for the link information link = dict( source = source_node, target = target_node, value = values ))]) # With this save the plots plot(fig, image_filename='sankey_plot_1', image='png', image_width=5000, image_height=3000) # And shows the plot fig.show()
解决方案
要实现需求,核心是明确桑基图的节点逻辑:每个节点代表某一月份的某个Topic,链接代表同一Topic从当月到下月的Size变化。以下是具体配置方法和完整代码:
1. 配置逻辑说明
- node_label:为每个(月份, Topic)组合生成唯一名称(如
2022-1 Topic1.0),确保不同时间的同一Topic能被区分。 - source_node/target_node:按Topic分组,将每个Topic的当月节点作为源,下月节点作为目标,用数字索引映射节点名称,满足Plotly的格式要求。
- values:取下月的Size作为链接值,让节点大小直观反映对应月份Topic的Size数值。
2. 完整可运行代码
import pandas as pd import plotly.graph_objects as go # 构造你的DataFrame(如果已有可跳过此部分) data = [ [2022,1,1.0,40],[2022,1,2.0,36],[2022,1,3.0,70],[2022,1,4.0,42],[2022,1,5.0,25],[2022,1,6.0,48], [2022,2,1.0,14],[2022,2,2.0,34],[2022,2,3.0,39],[2022,2,4.0,210],[2022,2,5.0,106],[2022,2,6.0,70], [2022,3,1.0,47],[2022,3,2.0,45],[2022,3,3.0,78],[2022,3,4.0,114],[2022,3,5.0,84],[2022,3,6.0,78], [2022,4,1.0,42],[2022,4,2.0,45],[2022,4,3.0,78],[2022,4,4.0,28],[2022,4,5.0,74],[2022,4,6.0,44], [2022,5,1.0,57],[2022,5,2.0,45],[2022,5,3.0,82],[2022,5,4.0,94],[2022,5,5.0,54],[2022,5,6.0,24] ] df = pd.DataFrame(data, columns=['year','month','Topic','Size']) # 生成节点唯一名称 df['node_name'] = df.apply(lambda x: f"2022-{x['month']} Topic{x['Topic']}", axis=1) # 构建节点标签和索引映射 unique_nodes = df['node_name'].unique() node_label = list(unique_nodes) node_index = {name: idx for idx, name in enumerate(node_label)} # 生成source、target和values列表 source_node = [] target_node = [] values = [] # 按Topic分组处理月度链接 for topic, group in df.groupby('Topic'): sorted_group = group.sort_values('month') # 按月份排序确保顺序正确 for i in range(len(sorted_group) - 1): # 获取源节点(当月)和目标节点(下月)的索引 source_idx = node_index[sorted_group.iloc[i]['node_name']] target_idx = node_index[sorted_group.iloc[i+1]['node_name']] # 链接值为下月的Size,对应Topic的月度变化 val = sorted_group.iloc[i+1]['Size'] source_node.append(source_idx) target_node.append(target_idx) values.append(val) # 绘制桑基图 fig = go.Figure(data=[go.Sankey( node=dict( label=node_label, pad=15, # 节点间距 thickness=20 # 节点厚度 ), link=dict( source=source_node, target=target_node, value=values ) )]) # 设置布局 fig.update_layout(title_text='各Topic月度Size变化桑基图', font_size=12) # 展示图表 fig.show() # 保存高清图片 fig.write_image('sankey_plot_1.png', width=5000, height=3000)
3. 关键变量解释
- node_label:包含所有唯一节点名称的列表,每个名称对应一个月度的Topic,清晰标识节点含义。
- node_index:字典,将节点名称映射为数字索引,Plotly的Sankey组件需要用数字索引指定链接的源和目标。
- source_node/target_node:分别存储每个链接的源节点索引和目标节点索引,实现同一Topic跨月份的连接。
- values:存储每个链接的数值,对应下月Topic的Size,让桑基图的流量大小匹配实际数据。
内容的提问来源于stack exchange,提问作者shan
相关产品推荐
相关产品推荐

