PySpark GroupBy与Pivot操作触发TypeError错误排查求助
问题排查与解决办法
嘿,我来帮你理清楚这个错误的原因,以及如何得到你想要的结果:
首先看你碰到的TypeError: _api() takes 1 positional argument but 2 were given,核心问题是你给count()传了参数,但PySpark里GroupedData对象的count()方法根本不需要参数——它默认统计每个分组的非空行数,不能像agg(count("col"))那样传列名。
另外还有个小疏漏:你代码里的groupby('device_id','country','brand')漏掉了date字段,但你期望的分组是包含date的,这会导致最终结果缺失date列,得补上。
两种正确实现方式
方式1:用pivot+count+列重命名
这种方式贴合你原本的思路,先按指定字段分组,pivot展开event_type,然后统计数量,最后把列名改成你要的xxx_count格式:
# 补全分组字段,pivot后调用无参count,再重命名列 df3 = df1.groupby('device_id', 'date', 'country', 'brand')\ .pivot("event_type")\ .count()\ .withColumnRenamed('load', 'load_count')\ .withColumnRenamed('close', 'close_count')\ .withColumnRenamed('display', 'display_count')
方式2:用agg直接统计(无需pivot)
如果你想更明确地指定统计ad_id的数量,或者担心pivot在event_type取值过多时性能下降,可以用agg配合条件统计,一步到位生成带目标列名的结果:
from pyspark.sql.functions import count, when df3 = df1.groupby('device_id', 'date', 'country', 'brand')\ .agg( count(when(col('event_type') == 'load', 'ad_id')).alias('load_count'), count(when(col('event_type') == 'close', 'ad_id')).alias('close_count'), count(when(col('event_type') == 'display', 'ad_id')).alias('display_count') )
再解释下错误根源
PySpark的GroupedData.count()是一个无参方法,作用是统计每个分组的总行数。你试图给它传"ad_id"参数,相当于给只接受1个参数(self)的方法传了2个参数,自然就触发了类型错误。如果要针对特定列统计非空值,应该用agg(count("ad_id"))的写法,而不是直接在count()里传参数。
内容的提问来源于stack exchange,提问作者Element
相关产品推荐
相关产品推荐

