PySpark中如何对多列使用describe函数获取统计信息?
PySpark中describe()仅显示单列统计的解决办法
问题原因
describe()方法默认仅对**数值类型(Numeric Type)**的列生成统计信息,你的Isball和Isboundary列大概率是布尔、字符串或其他非数值类型,因此被自动跳过,只输出了数值型的Runs列统计结果。
解决方案
方案1:显式指定统计列
直接将需要统计的列名作为参数传给describe(),这样会强制对指定列生成统计(非数值列的mean、stddev会显示null,但count、min、max会正常输出):
my_data.describe('Isball', 'Isboundary', 'Runs').show()
方案2:转换非数值列为数值型后统计
如果Isball和Isboundary是布尔型或可映射为数值的类型,先转换为数值列再统计,能得到完整的均值、标准差等统计值:
- 布尔型转整数(True=1,False=0):
from pyspark.sql.functions import col converted_df = my_data.withColumn('Isball_int', col('Isball').cast('int')) \ .withColumn('Isboundary_int', col('Isboundary').cast('int')) converted_df.describe('Isball_int', 'Isboundary_int', 'Runs').show()
- 字符串型(如"Y"/"N")映射为数值:
from pyspark.sql.functions import when mapped_df = my_data.withColumn('Isball_int', when(col('Isball') == 'Y', 1).otherwise(0)) \ .withColumn('Isboundary_int', when(col('Isboundary') == 'Y', 1).otherwise(0)) mapped_df.describe('Isball_int', 'Isboundary_int', 'Runs').show()
内容的提问来源于stack exchange,提问作者CodeByAM
相关产品推荐
相关产品推荐

