Python如何将学生多轮测试分类成绩转换为桑基图(Sankey plot)
多阶段测试成绩转化桑基图实现方案
需求与数据说明
现有学生四次测试的成绩等级数据,等级包含Poor、Fair、Good、Excellent四类,存在NaN缺失值,样例数据如下:
Roll no. 1 2 3 4 0 30 Good Fair Excellent Good 1 31 Poor Fair Good NaN 2 34 Excellent Good Poor Fair 3 35 Good Good Fair Good 4 36 NaN Fair Poor Fair 5 37 Excellent Good Excellent Excellent 6 39 Good Good Fair Excellent 7 42 Good Good Fair Fair 8 44 Fair Good Fair Poor 9 45 Good Good Good Good 10 46 Poor Good Fair Fair 11 50 Excellent Good Good Good
需要实现横轴依次为Test1到Test4四个阶段的桑基图,统计不同阶段成绩等级的转化人数。
桑基图核心参数逻辑
之前代码未达预期的核心原因是不同阶段的同等级成绩共用了同一节点索引,桑基图要求所有节点独立编号:
- nodes:所有节点的集合,4个测试每个对应4个等级,共16个独立节点,每个节点有对应标签、颜色、坐标
- source:每条转化流的起点节点索引,对应前一测试的成绩等级编号
- target:每条转化流的终点节点索引,对应后一测试的成绩等级编号
- value:对应两个等级之间的转化人数
完整实现代码
import pandas as pd import plotly.graph_objects as go # ---------------------- 1. 加载并预处理数据 ---------------------- # 样例数据构造,实际使用时替换为你的数据读取逻辑 data = [ [30, 'Good', 'Fair', 'Excellent', 'Good'], [31, 'Poor', 'Fair', 'Good', None], [34, 'Excellent', 'Good', 'Poor', 'Fair'], [35, 'Good', 'Good', 'Fair', 'Good'], [36, None, 'Fair', 'Poor', 'Fair'], [37, 'Excellent', 'Good', 'Excellent', 'Excellent'], [39, 'Good', 'Good', 'Fair', 'Excellent'], [42, 'Good', 'Good', 'Fair', 'Fair'], [44, 'Fair', 'Good', 'Fair', 'Poor'], [45, 'Good', 'Good', 'Good', 'Good'], [46, 'Poor', 'Good', 'Fair', 'Fair'], [50, 'Excellent', 'Good', 'Good', 'Good'] ] df = pd.DataFrame(data, columns=['Roll no.', 'Test1', 'Test2', 'Test3', 'Test4']) # 剔除存在缺失值的行,缺测学生无法计算连续转化 df = df.dropna(subset=['Test1', 'Test2', 'Test3', 'Test4']).reset_index(drop=True) # ---------------------- 2. 配置固定映射规则 ---------------------- # 等级到索引的映射,保证每个阶段的等级顺序统一 level_map = {'Poor':0, 'Fair':1, 'Good':2, 'Excellent':3} level_list = list(level_map.keys()) # 等级对应颜色 color_map = {'Poor':'#3498db', 'Fair':'#f1c40f', 'Good':'#2ecc71', 'Excellent':'#e67e22'} light_color_map = {'Poor':'#aed6f1', 'Fair':'#f9e79f', 'Good':'#abebc6', 'Excellent':'#fad7a0'} # 测试阶段列表 test_stages = ['Test1', 'Test2', 'Test3', 'Test4'] # ---------------------- 3. 构造nodes和link数据 ---------------------- node_labels = [] node_colors = [] node_x = [] node_y = [] # 每个阶段的x坐标,从左到右分布 x_pos = [0.1, 0.4, 0.7, 1.0] # 每个等级的y坐标,从上到下分布 y_pos = [0.8, 0.5, 0.2, 0.0] for i, stage in enumerate(test_stages): for level in level_list: node_labels.append(f"{stage}_{level}") node_colors.append(color_map[level]) node_x.append(x_pos[i]) node_y.append(y_pos[level_map[level]]) # 构造流数据 source = [] target = [] value = [] link_colors = [] # 遍历相邻两个测试阶段 for stage_idx in range(len(test_stages)-1): prev_stage = test_stages[stage_idx] next_stage = test_stages[stage_idx+1] # 计算两个阶段的转化交叉表 cross = pd.crosstab(df[prev_stage], df[next_stage]) # 遍历所有等级组合 for prev_level in level_list: for next_level in level_list: if prev_level not in cross.index or next_level not in cross.columns: cnt = 0 else: cnt = cross.loc[prev_level, next_level] if cnt > 0: # 计算起点和终点的节点索引 source_idx = stage_idx * 4 + level_map[prev_level] target_idx = (stage_idx + 1) * 4 + level_map[next_level] source.append(source_idx) target.append(target_idx) value.append(cnt) link_colors.append(light_color_map[prev_level]) # ---------------------- 4. 绘制桑基图 ---------------------- fig = go.Figure(data=[go.Sankey( node = dict( pad = 15, thickness = 20, line = dict(color = "black", width = 0.5), label = node_labels, color = node_colors, x = node_x, y = node_y ), link = dict( source = source, target = target, value = value, color = link_colors ) )]) fig.update_layout( title_text="四次测试成绩等级转化桑基图", font_size=12, width=900, height=600 ) fig.show()
结果说明
生成的桑基图从左到右对应Test1到Test4四个测试阶段,每个阶段的四个节点从上到下依次为Poor、Fair、Good、Excellent,流的宽度对应两个等级之间的转化人数,流的颜色和源节点等级颜色对应,便于追踪成绩变化趋势。
内容的提问来源于stack exchange,提问作者Pri D
相关产品推荐
相关产品推荐

