如何在Seaborn中按两类变量实现带渐变的分类着色?
解决思路与实现方案
要实现同一图内叠加、按category区分主色系、按category_2数值生成渐变的需求,我们可以通过两种清晰可控的方式来完成:分批次绘图指定渐变调色板,或者构建组合分类变量并自定义颜色映射。
方法一:分批次绘图(直观可控)
这种方式可以明确给category=0和category=1各自分配一套渐变色系,同时保留size和style对category_2的区分:
%pylab inline import pandas as pd import numpy as np import seaborn as sns; sns.set();sns.set_style("whitegrid") # 生成示例数据(和你的代码一致) df = pd.DataFrame(np.random.randint(0,100,size=(100, 4)), columns=list('ABCD')) df['category'] = df['A'].mod(2) df['category_2'] = df['B'].mod(3) # 为每个category定义渐变调色板:category=0用蓝色渐变,category=1用红色渐变 # 使用_r后缀让数值大的category_2对应更深的颜色 palette_cat0 = sns.color_palette("Blues_r", n_colors=df['category_2'].nunique()) palette_cat1 = sns.color_palette("Reds_r", n_colors=df['category_2'].nunique()) # 创建绘图画布 fig, ax = plt.subplots(figsize=(8, 6)) # 绘制category=0的所有曲线 sns.lineplot( x='C', y='D', data=df[df['category'] == 0], hue='category_2', size='category_2', style='category_2', palette=palette_cat0, ax=ax ) # 绘制category=1的所有曲线,叠加在同一图上 sns.lineplot( x='C', y='D', data=df[df['category'] == 1], hue='category_2', size='category_2', style='category_2', palette=palette_cat1, ax=ax ) # 优化图例:合并重复的category_2标签,明确标注所属的category handles, labels = ax.get_legend_handles_labels() # 只保留前6个有效图例(每个category对应3个category_2值) new_legend_labels = [f'Cat 0 - {lab}' for lab in labels[:3]] + [f'Cat 1 - {lab}' for lab in labels[3:6]] ax.legend(handles[:6], new_legend_labels, title='Category & Category 2') plt.title('Gradient Line Plot with Overlaid Categories') plt.show()
方法二:构建组合分类变量(简洁高效)
我们可以把category和category_2组合成一个新变量,然后为每个组合自定义颜色,实现渐变效果:
%pylab inline import pandas as pd import numpy as np import seaborn as sns; sns.set();sns.set_style("whitegrid") # 生成示例数据 df = pd.DataFrame(np.random.randint(0,100,size=(100, 4)), columns=list('ABCD')) df['category'] = df['A'].mod(2) df['category_2'] = df['B'].mod(3) # 创建组合分类变量:格式为"category_category_2" df['cat_combined'] = df.apply(lambda row: f"{row['category']}_{row['category_2']}", axis=1) # 构建颜色映射:给每个组合分配对应渐变颜色 color_mapping = {} # 给category=0的三个category_2值分配蓝色渐变 for idx, c2_val in enumerate(df['category_2'].unique()): color_mapping[f"0_{c2_val}"] = sns.color_palette("Blues_r", 3)[idx] # 给category=1的三个category_2值分配红色渐变 for idx, c2_val in enumerate(df['category_2'].unique()): color_mapping[f"1_{c2_val}"] = sns.color_palette("Reds_r", 3)[idx] # 绘制单图 sns.lineplot( x='C', y='D', data=df, hue='cat_combined', size='category_2', style='category_2', palette=color_mapping ) # 优化图例标签 handles, labels = plt.gca().get_legend_handles_labels() clean_labels = [lab.replace('_', ' - Cat ') for lab in labels[:6]] plt.legend(handles[:6], clean_labels, title='Category - Category 2') plt.title('Combined Hue with Gradient Colors') plt.show()
关键说明
- 我们使用
Blues_r和Reds_r色系(_r表示反向渐变),让category_2数值越大,颜色越深,符合你“根据数值大小生成渐变色”的需求; - 两种方法都避免了
relplot的子图拆分问题,实现了曲线叠加,方便对比不同category的趋势; - 图例优化是为了避免重复的
category_2标签,让图表更易读。
内容的提问来源于stack exchange,提问作者Georg Heiler
相关产品推荐
相关产品推荐

