You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark DataFrame绘制雪佛兰年度数量柱状图报错求助

问题解决方法

错误原因

直接将PySpark DataFrame传入pd.DataFrame()是不合法的操作——PySpark DataFrame需要通过toPandas()方法完成与Pandas DataFrame的转换。此外你的原始代码未做分组计数处理,直接用原数据列绘图无法实现"各年份车辆数量统计"的需求。

修正后的完整代码

import pandas as pd
import matplotlib.pyplot as plt

# 1. 筛选雪佛兰品牌数据,按年份分组统计车辆数量并排序
chevy_stats = df16.filter(df16.Brand == '雪佛兰') \
                  .groupBy('Model Year') \
                  .count() \
                  .orderBy('Model Year')

# 2. 将PySpark DataFrame转换为Pandas DataFrame
pd_chevy = chevy_stats.toPandas()

# 3. 绘制柱状图
pd_chevy.plot(x='Model Year', y='count', kind='bar', legend=False)
plt.ylabel('车辆数量')
plt.title('雪佛兰各年份车辆数量增长情况')
plt.show()

关键修改点

  • 新增filter()筛选雪佛兰品牌数据,避免混入其他品牌干扰结果
  • 通过groupBy('Model Year').count()完成核心的年份-数量统计,这是实现需求的必要步骤
  • 用toPandas()替代直接传入pd.DataFrame(),完成两种DataFrame的合法转换
  • 调整绘图的x/y轴参数,对应统计后的年份列和数量列,确保图表符合需求

内容的提问来源于stack exchange,提问作者AJS

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 22:32:33