如何优化Seaborn lineplot的数值标注效果?
折线图标注优化问题与解决方案
数据集
| mentor_cnt | mentee_cnt | |
|---|---|---|
| 0 | 3 | 3 |
| 1 | 3 | 3 |
| 2 | 7 | 7 |
| 3 | 18 | 19 |
| 4 | 24 | 27 |
| 5 | 37 | 35 |
| 6 | 53 | 56 |
| 7 | 63 | 70 |
| 8 | 86 | 89 |
| 9 | 102 | 114 |
| 10 | 135 | 149 |
| 11 | 154 | 169 |
| 12 | 174 | 202 |
| 13 | 232 | 287 |
| 14 | 298 | 386 |
| 15 | 343 | 475 |
| 16 | 384 | 552 |
| 17 | 446 | 684 |
| 18 | 509 | 883 |
| 19 | 469 | 757 |
初始实现代码
import pandas as pd import matplotlib.pyplot as plt import seaborn as sns sns.set_style('darkgrid') fig, ax = plt.subplots(figsize=(10, 6)) sns.lineplot(x=df['month'], y=df['mentor_cnt'], color = 'b', label='Mentors', markers=True, marker='o') sns.lineplot(x=df['month'], y=df['mentee_cnt'], color = 'r', label='Mentee', markers=True, marker='o') plt.ylim(min(df['mentee_cnt'])-60, max(df['mentee_cnt'])+50) columns =['mentor_cnt', 'mentee_cnt'] colors = ['b', 'w'] facecolors = ['none', 'r'] x_position = [0, -10] y_position = [-18, 5] for col, fc, c, x_p, y_p in zip(columns, facecolors, colors, x_position, y_position): for x, y in zip(df['month'], df[col]): label = '{:.0f}'.format(y) plt.annotate( label, (x, y), textcoords='offset points', xytext=(x_p, y_p), ha='center', color=c ).set_bbox(dict(facecolor=fc, alpha=0.5, boxstyle='round', edgecolor='none')) ax.legend() plt.xlabel('Month') plt.ylabel('Number of persons') plt.title('Mentee, mentor') plt.show()
初始效果
导师(Mentors)的折线标注效果尚可,但学员(Mentee)的数值标注拥挤重叠,可读性差,尝试用set_bbox提升效果也不佳。
优化思路
通过仅标注关键分位数和末尾数值减少标注数量,避免拥挤,同时调整标注位置参数提升可读性。
优化后代码
fig, ax = plt.subplots(figsize=(10, 6)) sns.lineplot(x=df['month'], y=df['mentor_cnt'], color='b', label='Mentors', markers=True, marker='o') sns.lineplot(x=df['month'], y=df['mentee_cnt'], color='r', label='Mentee', markers=True, marker='o') plt.ylim(min(df['mentee_cnt'])-60, max(df['mentee_cnt'])+50) columns = ['mentor_cnt', 'mentee_cnt'] colors = ['b', 'r'] facecolors = ['none', 'r'] x_position = [0, -15] y_position = [-18, 3] for col, fc, c, x_p, y_p in zip(columns, facecolors, colors, x_position, y_position): for x, y in zip(df['month'], df[col]): # 生成包含分位数和末尾值的关键数值列表 lst = [0, 0.25, 0.5, 0.75, 1] describe_nearest = [] [describe_nearest.append(df[col].quantile(el, interpolation='nearest')) for el in lst] describe_nearest.append(df[col].values[-1::][0]) # 仅对关键数值添加标注 if y in describe_nearest: label = '{:.0f}'.format(y) plt.annotate( label, (x, y), textcoords='offset points', xytext=(x_p, y_p), ha='center', color=c ) ax.legend() plt.xlabel('Month') plt.ylabel('Number of persons') plt.title('Mentee, mentor') plt.show()
优化后效果
仅保留关键分位数(0、25%、50%、75%、100%分位)和最后一个数值的标注,标注布局清晰,学员折线的标注可读性显著提升。
内容的提问来源于stack exchange,提问作者John Doe
相关产品推荐
相关产品推荐

