Matplotlib绘制空心/实心散点图:相似点配色及文字颜色问题求助
散点图绘制问题修复方案
存在的问题
- 与种子词相似的点无法匹配对应种子词的颜色
- 指定Okabe and Ito调色板后,第一个种子词显示为白色
修复后的完整代码
import pandas as pd import numpy as np import matplotlib.pyplot as plt seed_terms = ['clean', 'recovery', 'spiral', 'tolerance', 'program'] embeddings_ex = np.random.rand(5, 10, 2) words_ex = [ ['quit', 'finally', 'pill', 'vomit', 'survive' ,'lil', 'chance' ,'chain', 'zero', 'quickly'], ['bullshit' ,'unrelated', 'everywhere', 'appear' ,'probably' ,'deal', 'mistake', 'window', 'comment', 'honest'], ['majority' ,'familiar', 'queer', 'edgy', 'skin', 'withdrawl' ,'sad', 'develop', 'perfectly', 'daughter'], ['snort', 'cheap', 'brain', 'teach' ,'shoot' ,'inject' ,'freak', 'type', 'black', 'absolute'], ['substitution', 'suboxone', 'country' ,'clinic', 'nerve', 'representation', '2', 'website' ,'youtuber', 'insane'] ] # Okabe and Ito调色板,取前5个对应5个种子词 colors = ['#FA4D4D', '#FBC93D', '#E37E3B', '#C13BE3', '#4B42FD'] fig, ax = plt.subplots(figsize=(16, 9)) # 按种子词分组绘制相似词点 for idx, (seed_word, embeds, sim_words) in enumerate(zip(seed_terms, embeddings_ex, words_ex)): # 绘制相似词的点,使用对应种子词的颜色,半透明 ax.scatter(embeds[:, 0], embeds[:, 1], s=45, marker='o', alpha=0.4, color=colors[idx], edgecolors='k') # 绘制种子词的实心点,颜色一致,更大更醒目 ax.scatter(embeds[0, 0], embeds[0, 1], marker='o', alpha=1.0, color=colors[idx], edgecolors='none', s=100) # 标注种子词,加粗字体 ax.annotate(seed_word, xy=(embeds[0, 0], embeds[0, 1]), xytext=(5, 2), textcoords='offset points', ha='right', va='bottom', size=11, fontweight='bold') # 标注相似词 for x, y, word in zip(embeds[:, 0], embeds[:, 1], sim_words): ax.annotate(word, xy=(x, y), xytext=(3, 1), textcoords='offset points', ha='right', va='bottom', size=6) # 添加图例 ax.legend(seed_terms, loc=4) # 移除边框和刻度 ax.grid(False) for spine in ax.spines.values(): spine.set_visible(False) ax.tick_params(axis='both', which='both', bottom=False, left=False, labelbottom=False, labelleft=False) plt.show()
修复说明
问题1:相似词点匹配种子词颜色
- 原代码先绘制了统一样式的空心黑边点,覆盖了后续分组颜色设置。现在改为按种子词分组遍历,每组相似词使用对应种子词的颜色绘制,确保颜色关联正确。
- 每组内先绘制相似词的半透明点,再绘制种子词的实心点,保证种子词不会被遮挡。
问题2:第一个种子词显示白色
- 原代码中种子词的坐标取值错误:
sc[i]取的是flatten后的前5个点(均来自第一组),导致种子词坐标重叠或错位,视觉上显示为白色(实际是被其他点覆盖)。现在直接从每组嵌入向量中取对应种子词的坐标(假设每组第一个点对应种子词,可根据实际数据调整),确保坐标正确且不重叠。 - 调色板只保留前5个颜色,与种子词数量对应,避免索引越界或颜色浪费。
内容的提问来源于stack exchange,提问作者laBouz
相关产品推荐
相关产品推荐

