You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Matplotlib绘制空心/实心散点图:相似点配色及文字颜色问题求助

散点图绘制问题修复方案

存在的问题

  • 与种子词相似的点无法匹配对应种子词的颜色
  • 指定Okabe and Ito调色板后,第一个种子词显示为白色

修复后的完整代码

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt

seed_terms = ['clean', 'recovery', 'spiral', 'tolerance', 'program']
embeddings_ex = np.random.rand(5, 10, 2)
words_ex = [
    ['quit', 'finally', 'pill', 'vomit', 'survive' ,'lil', 'chance' ,'chain', 'zero', 'quickly'],
    ['bullshit' ,'unrelated', 'everywhere', 'appear' ,'probably' ,'deal', 'mistake', 'window', 'comment', 'honest'],
    ['majority' ,'familiar', 'queer', 'edgy', 'skin', 'withdrawl' ,'sad', 'develop', 'perfectly', 'daughter'],
    ['snort', 'cheap', 'brain', 'teach' ,'shoot' ,'inject' ,'freak', 'type', 'black', 'absolute'],
    ['substitution', 'suboxone', 'country' ,'clinic', 'nerve', 'representation', '2', 'website' ,'youtuber', 'insane']
]

# Okabe and Ito调色板,取前5个对应5个种子词
colors = ['#FA4D4D', '#FBC93D', '#E37E3B', '#C13BE3', '#4B42FD']

fig, ax = plt.subplots(figsize=(16, 9))

# 按种子词分组绘制相似词点
for idx, (seed_word, embeds, sim_words) in enumerate(zip(seed_terms, embeddings_ex, words_ex)):
    # 绘制相似词的点,使用对应种子词的颜色,半透明
    ax.scatter(embeds[:, 0], embeds[:, 1], s=45, marker='o', alpha=0.4, color=colors[idx], edgecolors='k')
    # 绘制种子词的实心点,颜色一致,更大更醒目
    ax.scatter(embeds[0, 0], embeds[0, 1], marker='o', alpha=1.0, color=colors[idx], edgecolors='none', s=100)
    # 标注种子词,加粗字体
    ax.annotate(seed_word, xy=(embeds[0, 0], embeds[0, 1]), xytext=(5, 2), 
                textcoords='offset points', ha='right', va='bottom', size=11, fontweight='bold')
    # 标注相似词
    for x, y, word in zip(embeds[:, 0], embeds[:, 1], sim_words):
        ax.annotate(word, xy=(x, y), xytext=(3, 1), 
                    textcoords='offset points', ha='right', va='bottom', size=6)

# 添加图例
ax.legend(seed_terms, loc=4)

# 移除边框和刻度
ax.grid(False)
for spine in ax.spines.values():
    spine.set_visible(False)
ax.tick_params(axis='both', which='both', bottom=False, left=False, labelbottom=False, labelleft=False)

plt.show()

修复说明

问题1:相似词点匹配种子词颜色

  • 原代码先绘制了统一样式的空心黑边点,覆盖了后续分组颜色设置。现在改为按种子词分组遍历,每组相似词使用对应种子词的颜色绘制,确保颜色关联正确。
  • 每组内先绘制相似词的半透明点,再绘制种子词的实心点,保证种子词不会被遮挡。

问题2:第一个种子词显示白色

  • 原代码中种子词的坐标取值错误:sc[i]取的是flatten后的前5个点(均来自第一组),导致种子词坐标重叠或错位,视觉上显示为白色(实际是被其他点覆盖)。现在直接从每组嵌入向量中取对应种子词的坐标(假设每组第一个点对应种子词,可根据实际数据调整),确保坐标正确且不重叠。
  • 调色板只保留前5个颜色,与种子词数量对应,避免索引越界或颜色浪费。

内容的提问来源于stack exchange,提问作者laBouz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 02:35:48