使用Plotly按每5列单独绘制箱线图未达预期,求解决方案
问题解决方法
原代码存在两个核心问题:
- 索引错误:
X[i]取的是数据集的第i行,而我们需要的是列数据,应改为X[:, i-1](Python是0索引,循环从1开始时需对应调整)。 - 未按需求分组:原循环是间隔5列取数据合并到一个图,没有实现“每5列单独一个图表”的要求。
以下提供两种符合需求的实现方案:
方案一:生成多个独立图表(每5列一组)
该方案会为每5列数据生成一个独立的HTML图表文件,便于单独查看每组数据分布:
import numpy as np from sklearn import datasets import plotly.graph_objs as go from plotly.offline import plot # 生成目标规模的数据集 X, y = datasets.make_classification(n_samples=300, n_features=70, n_classes=3, n_redundant=0, n_clusters_per_class=1, weights=[0.5, 0.3, 0.2], random_state=42) # 按每5列分组处理 for group_start in range(0, X.shape[1], 5): group_end = min(group_start + 5, X.shape[1]) box_plots = [] # 遍历当前组内的每一列 for col_idx in range(group_start, group_end): box_plots.append(go.Box( y=X[:, col_idx], name=f"列{col_idx + 1}", showlegend=False )) # 生成并保存独立图表 plot(box_plots, filename=f"列组_{group_start+1}_至_{group_end}.html")
方案二:生成包含多子图的汇总图表
该方案将所有5列分组的箱线图整合到一个大图表的子图中,便于整体对比所有组的数据:
import numpy as np from sklearn import datasets from plotly.subplots import make_subplots import plotly.graph_objs as go from plotly.offline import plot # 生成目标规模的数据集 X, y = datasets.make_classification(n_samples=300, n_features=70, n_classes=3, n_redundant=0, n_clusters_per_class=1, weights=[0.5, 0.3, 0.2], random_state=42) # 计算分组数量,处理非5的倍数列数 total_groups = X.shape[1] // 5 if X.shape[1] % 5 != 0: total_groups += 1 # 创建子图布局(示例为2行7列,可根据需求调整) fig = make_subplots( rows=2, cols=7, subplot_titles=[f"列组{i*5+1}至{min(i*5+5, X.shape[1])}" for i in range(total_groups)] ) # 填充每个子图的箱线图 for group_idx in range(total_groups): # 计算当前子图的位置 row = (group_idx // 7) + 1 col = (group_idx % 7) + 1 # 获取当前组的列范围 start_col = group_idx * 5 end_col = min(start_col + 5, X.shape[1]) # 添加当前组的所有箱线图 for col_idx in range(start_col, end_col): fig.add_trace( go.Box(y=X[:, col_idx], name=f"列{col_idx+1}", showlegend=False), row=row, col=col ) # 调整整体布局 fig.update_layout(height=800, width=1600, title_text="每5列分组箱线图汇总") # 保存汇总图表 plot(fig, filename="分组箱线图汇总.html")
内容的提问来源于stack exchange,提问作者ala mazahreh
相关产品推荐
相关产品推荐

