You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在PySpark环境中使用Matplotlib与Pandas绘制直方图

没问题!既然你已经完成了PySpark DataFrame的采样并转换成了Pandas DataFrame,接下来用Matplotlib结合Pandas绘制直方图非常简单,我给你两种常用的实现方式,还可以根据你的数据特点做优化:

方法一:直接用Pandas内置的hist()方法

Pandas的Series/DataFrame自带直方图绘制方法,底层会调用Matplotlib,代码简洁直观:

# 先导入Matplotlib(如果还没导入的话)
import matplotlib.pyplot as plt

# 针对beer_freq列绘制直方图
sample_pd['beer_freq'].hist(bins=15, edgecolor='black')

# 添加可视化标签和标题,提升可读性
plt.title('Distribution of Client Beer Frequency')
plt.xlabel('Beer Frequency')
plt.ylabel('Number of Clients')

# 显示图形
plt.show()

参数说明:

  • bins=15:指定直方图的柱子数量,你可以根据数据分布调整这个数值;
  • edgecolor='black':给柱子加上黑色边框,避免相邻柱子混在一起,更清晰。

方法二:用Matplotlib原生hist()函数绘制

如果你需要更精细的控制,也可以直接用Matplotlib的原生函数:

import matplotlib.pyplot as plt

# 传入Pandas Series数据到Matplotlib的hist函数
plt.hist(sample_pd['beer_freq'], bins=15, edgecolor='black')

# 同样添加标签和标题
plt.title('Distribution of Client Beer Frequency')
plt.xlabel('Beer Frequency')
plt.ylabel('Number of Clients')

plt.show()

针对你的数据优化:处理大量0值的情况

从你给出的前10条数据来看,beer_freq有不少0值,默认的直方图可能会让非0值的分布被压缩。你可以自定义分箱范围,或者使用对数刻度来优化:

自定义分箱范围

sample_pd['beer_freq'].hist(bins=[0, 0.1, 0.2, 0.4, 0.6, 0.8, 1.0], edgecolor='black')
plt.title('Distribution of Client Beer Frequency (Custom Bins)')
plt.xlabel('Beer Frequency')
plt.ylabel('Number of Clients')
plt.show()

使用对数刻度(查看小数值区间的分布)

sample_pd['beer_freq'].hist(bins=15, edgecolor='black')
plt.title('Distribution of Client Beer Frequency (Log Scale)')
plt.xlabel('Beer Frequency')
plt.ylabel('Number of Clients (Log)')
plt.yscale('log')  # 将y轴设为对数刻度
plt.show()

这样就能更清晰地看到非0值的分布情况啦!

内容的提问来源于stack exchange,提问作者Elsa Li

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 04:21:14