PySpark DataFrame调用pivot报错:'DataFrame'对象无pivot属性
PySpark透视表报错:AttributeError: 'DataFrame' object has no attribute 'pivot'
问题场景
原始PySpark DataFrame结构:
| user_id | item_id | last_watch_dt | total_dur | watched_pct |
|---|---|---|---|---|
| 1 | 1 | 2021-05-11 | 4250 | 72 |
| 1 | 2 | 2021-05-11 | 80 | 99 |
| 2 | 3 | 2021-05-11 | 1000 | 80 |
| 2 | 4 | 2021-05-11 | 5000 | 40 |
尝试执行以下代码生成透视表:
df_new = df.pivot(index='user_id', columns='item_id', values='watched_pct')
报错信息:
AttributeError: 'DataFrame' object has no attribute 'pivot'
期望得到的透视表结果:
| 1 | 2 | 3 | 4 | |
|---|---|---|---|---|
| 1 | 72 | 99 | 0 | 0 |
| 2 | 0 | 0 | 80 | 40 |
问题原因
PySpark的DataFrame类没有直接提供pivot方法,这和Pandas的用法不同。在PySpark中,必须先通过groupBy指定透视的索引列,再调用pivot方法指定列名,最后配合聚合函数完成透视操作。
解决代码
# 先分组,再透视,聚合watched_pct,最后填充空值为0 df_new = df.groupBy('user_id') \ .pivot('item_id') \ .agg({'watched_pct': 'first'}) \ .fillna(0)
代码说明
groupBy('user_id'):按用户ID分组,作为透视表的行索引pivot('item_id'):将item_id的不同取值作为透视表的列agg({'watched_pct': 'first'}):因为每个user-item组合只有一条数据,用first聚合可以直接取到对应的watched_pct值,也可以用max/min,效果一致fillna(0):将透视后出现的空值(即用户未观看的item)替换为0,符合期望结果
执行上述代码后,就能得到你想要的透视表结构。
内容的提问来源于stack exchange,提问作者KlimShaman
相关产品推荐
相关产品推荐

