如何使用Python的Plotly库绘制弦图或桑基图做关联分析?
多字段共现可视化实现方案
依赖安装
首先安装需要的工具库:
pip install pandas plotly
1 数据预处理
先加载你的数据并计算字段两两共现的ID数量:
import pandas as pd import plotly.graph_objects as go # 替换为你的真实数据 df = pd.DataFrame({ 'ID': [1,2,3,4], 'ant_mi': ['Yes', 'No', 'No', 'Yes'], 'inf_mi': ['No', 'No', 'No', 'Yes'], 'lat_mi': ['Yes', 'No', 'Yes', 'No'], 'post_mi': ['No', 'No', 'Yes', 'No'] }) # 提取要统计的字段列 feature_cols = ['ant_mi', 'inf_mi', 'lat_mi', 'post_mi'] # 初始化共现矩阵 co_matrix = pd.DataFrame(0, index=feature_cols, columns=feature_cols) # 遍历每个ID统计共现 for _, row in df.iterrows(): # 筛选当前ID取值为Yes的字段 yes_features = [col for col in feature_cols if row[col] == 'Yes'] # 两两配对计数 for i in range(len(yes_features)): for j in range(i+1, len(yes_features)): a, b = yes_features[i], yes_features[j] co_matrix.loc[a, b] += 1 co_matrix.loc[b, a] += 1
2 绘制桑基图(推荐)
Plotly原生支持桑基图,无需额外配置,可视化效果直观,鼠标悬停在连接线上即可查看两个字段的共现ID数量:
# 构造节点和连接数据 node_labels = feature_cols source_list = [] target_list = [] value_list = [] for i, source in enumerate(node_labels): for j, target in enumerate(node_labels): if i < j and co_matrix.loc[source, target] > 0: source_list.append(i) target_list.append(j) value_list.append(co_matrix.loc[source, target]) # 生成桑基图 fig = go.Figure(data=[go.Sankey( node = dict( pad = 20, thickness = 25, line = dict(color = "#333", width = 0.5), label = node_labels ), link = dict( source = source_list, target = target_list, value = value_list ) )]) fig.update_layout(title = "多字段Yes共现桑基图", font_size = 13) fig.show()
3 绘制弦图
如果需要弦图形式的展示,可以使用plotly的figure_factory模块实现:
import plotly.figure_factory as ff matrix = co_matrix.values.tolist() fig = ff.create_chord(matrix, node_labels) fig.update_layout(title = "多字段Yes共现弦图") fig.show()
注意:如果调用
create_chord报错,可升级plotly到最新版本,或优先使用桑基图方案。
内容的提问来源于stack exchange,提问作者Nemra Khalil
相关产品推荐
相关产品推荐

