如何可视化值为不等长字符串列表的Python字典,展示KNN各试验对应准确率
可视化实现方案
推荐优先使用横向条形图展示不同准确率对应的试验数量,也可搭配散点图展示单个试验的准确率分布,两种实现代码如下:
第一步:数据预处理
首先将现有字典转换为更适合绘图的结构化数据:
import pandas as pd import matplotlib.pyplot as plt import seaborn as sns # 填入你生成的duplicate_dict duplicate_dict = {'0.9885185185185185': ['Nearest_Neighbors_6', 'Nearest_Neighbors_7', 'Nearest_Neighbors_13', 'Nearest_Neighbors_14', 'Nearest_Neighbors_27', 'Nearest_Neighbors_28'], '0.9890740740740741': ['Nearest_Neighbors_21', 'Nearest_Neighbors_35'], '0.9894444444444445': ['Nearest_Neighbors_20', 'Nearest_Neighbors_34'], '0.9903703703703703': ['Nearest_Neighbors_5', 'Nearest_Neighbors_12', 'Nearest_Neighbors_26'], '0.9922222222222222': ['Nearest_Neighbors_19', 'Nearest_Neighbors_33'], '0.9929629629629629': ['Nearest_Neighbors_4', 'Nearest_Neighbors_11', 'Nearest_Neighbors_25'], '0.9931481481481481': ['Nearest_Neighbors_18', 'Nearest_Neighbors_32'], '0.9938888888888889': ['Nearest_Neighbors_3', 'Nearest_Neighbors_10', 'Nearest_Neighbors_17', 'Nearest_Neighbors_24', 'Nearest_Neighbors_31'], '0.9955555555555555': ['Nearest_Neighbors_2', 'Nearest_Neighbors_9', 'Nearest_Neighbors_23'], '0.9957407407407407': ['Nearest_Neighbors_16', 'Nearest_Neighbors_30'], '0.997037037037037': ['Nearest_Neighbors', 'Nearest_Neighbors_8', 'Nearest_Neighbors_15', 'Nearest_Neighbors_22', 'Nearest_Neighbors_29']} # 转换为DataFrame格式 records = [] for score_str, models in duplicate_dict.items(): score = float(score_str) for model_name in models: records.append({'试验名称': model_name, '准确率': score}) df = pd.DataFrame(records) # 按准确率升序排序方便后续展示 df = df.sort_values('准确率', ascending=True)
第二步:可视化实现
方案1:横向条形图(展示不同准确率对应的试验数量)
适合快速查看不同准确率对应的试验分布密度:
# 统计每个准确率对应的模型数量 count_df = df.groupby('准确率')['试验名称'].count().reset_index(name='对应试验数量') count_df['准确率'] = count_df['准确率'].apply(lambda x: round(x,4)) # 保留4位小数更美观 plt.figure(figsize=(10,6)) sns.barplot(data=count_df, y='准确率', x='对应试验数量', palette='Blues_d', orient='h') plt.title('不同KNeighborsClassifier试验的准确率分布', fontsize=14) plt.xlabel('对应试验数量', fontsize=12) plt.ylabel('准确率', fontsize=12) plt.tight_layout() plt.show()
方案2:散点图(展示单个试验的准确率详情)
适合查看每个单独试验的具体得分情况:
plt.figure(figsize=(12,7)) sns.scatterplot(data=df, x='准确率', y='试验名称', s=100, color='darkblue') plt.title('各KNeighborsClassifier试验准确率详情', fontsize=14) plt.xlabel('准确率', fontsize=12) plt.ylabel('试验名称', fontsize=12) plt.grid(axis='x', linestyle='--', alpha=0.7) plt.tight_layout() plt.show()
其他说明
如果需要用热图展示,可额外提取每个试验的n_neighbors参数、距离度量参数作为两个维度,准确率作为热力值绘制二维热图,只需补充存储每个试验对应参数的列表即可实现。
内容的提问来源于stack exchange,提问作者Shihab Sharar
相关产品推荐
相关产品推荐

