catplot中stripplot与barplot价格差异原因及barplot轴调整方法
Airbnb纽约数据集可视化问题:stripplot与barplot的y轴差异及总价展示调整
数据处理代码
import numpy as np import pandas as pd import matplotlib.pyplot as plt import seaborn as sns %matplotlib inline sns.set_style('darkgrid') sns.set_theme() df = pd.read_csv("https://raw.githubusercontent.com/lokalhangatt/stackoverlow/refs/heads/main/airbnb_nyc_2019.csv") df.fillna({'reviews_per_month':0}, inplace=True)
手动聚合的总价结果
通过groupby计算各房型、各区域的价格总和:
df.groupby(['room_type', 'neighbourhood_group'])['price'].sum()
输出结果:
room_type neighbourhood_group Entire home/apt Bronx 48325 Brooklyn 1704633 Manhattan 3289707 Queens 308218 Staten Island 30597 Private room Bronx 43546 Brooklyn 775099 Manhattan 932111 Queens 241983 Staten Island 11711 Shared room Bronx 3588 Brooklyn 20868 Manhattan 42709 Queens 13666 Staten Island 517 Name: price, dtype: int64
问题
使用sns.stripplot()时,y轴显示的是原始价格的分布;但改用sns.barplot()后,y轴数值和上述手动聚合的总价完全不符,想知道两者差异的原因,以及如何调整barplot使其显示实际总价。
原因分析
Seaborn的barplot默认使用均值(estimator=np.mean)作为统计量,即计算每个分组内价格的平均值,而非总和;而stripplot是直接绘制原始数据点,所以y轴对应原始价格,两者统计逻辑完全不同,导致y轴数值差异。
调整barplot显示总价的两种方法
方法1:先聚合数据再绘图
先手动计算总价并转换为普通DataFrame,再传入barplot:
# 聚合总价并重置索引 price_total = df.groupby(['room_type', 'neighbourhood_group'])['price'].sum().reset_index() # 绘制barplot plt.figure(figsize=(12,6)) sns.barplot(data=price_total, x='room_type', y='price', hue='neighbourhood_group') plt.ylabel('总价') plt.title('各房型按区域的价格总和') plt.ticklabel_format(style='plain', axis='y') # 关闭科学计数法,清晰显示大数值 plt.show()
方法2:直接指定barplot的统计量
在barplot中通过estimator参数指定使用求和函数:
plt.figure(figsize=(12,6)) sns.barplot(data=df, x='room_type', y='price', hue='neighbourhood_group', estimator=np.sum) plt.ylabel('总价') plt.title('各房型按区域的价格总和') plt.ticklabel_format(style='plain', axis='y') plt.show()
两种方法都能得到和手动聚合一致的总价条形图,其中方法1更灵活,可提前对聚合数据做额外处理;方法2更简洁,适合直接基于原始数据绘图。
内容的提问来源于stack exchange,提问作者lokalhangatt
相关产品推荐
相关产品推荐

