如何处理pandas字符串型价格列并绘制分组折线图展示所有价格?
实现步骤
第一步:清洗价格列,转换为数值类型
首先处理带$符号的字符串格式价格,去掉符号后转为浮点型:
import pandas as pd import matplotlib.pyplot as plt # 示例数据 df = pd.DataFrame({'id':['A', 'B', 'C', 'D'], 'price':['$106.00', '$156.00', '$166.00', '$106.00']}) # 清洗价格列:去掉$符号后转数值 df['price'] = df['price'].str.replace('$', '', regex=False).astype(float)
第二步:价格分组处理
有两种常用分组方式适配不同的分析需求:
- 按实际价格值分组,统计每个价格的出现频次,适合价格取值较少的场景:
# 按价格分组计数后按价格排序 grouped_df = df.groupby('price')['id'].count().sort_index().reset_index(name='count')
- 按自定义区间分箱分组,适合价格取值较多、需要看整体集中区间的场景:
# 按20元为间隔划分区间,可根据实际数据调整区间范围 df['price_bin'] = pd.cut(df['price'], bins=[100,120,140,160,180], labels=['100-120','120-140','140-160','160-180']) # 按分箱分组,同时统计每个区间的样本量、平均价格 grouped_df = df.groupby('price_bin', observed=True)['price'].agg(['count','mean']).reset_index()
第三步:绘制符合需求的折线图
如果需要同时展示价格分布区间、标注整体平均价格,参考如下代码:
# x轴为价格分组区间,y轴为对应区间的样本数量 ax = grouped_df.plot.line(x='price_bin', y='count', marker='o', figsize=(8,5), title='价格分布折线图') # 添加平均价格参考线 avg_price = df['price'].mean() ax.axhline(y=avg_price, color='r', linestyle='--', label=f'平均价格: {avg_price:.2f}元') # 调整坐标轴标签和图例 ax.set_xlabel('价格区间') ax.set_ylabel('出现频次') ax.legend() plt.show()
如果需要y轴为价格、x轴为分组,直接展示各分组的平均价格折线,参考如下代码:
ax = grouped_df.plot.line(x='price_bin', y='mean', marker='s', color='green', figsize=(8,5)) ax.set_xlabel('价格区间') ax.set_ylabel('平均价格(元)') plt.show()
内容的提问来源于stack exchange,提问作者Pyyyyt
相关产品推荐
相关产品推荐

