如何在PySpark中生成含各列值计数的目标DataFrame?
PySpark实现列值计数并排展示
你的需求是对原DataFrame的每一列分别统计不同值的出现次数,再将这些统计结果按行并排拼接成新的DataFrame。直接用普通join会因无共同关联键导致结果不符合预期,可通过添加行号作为关联键实现目标:
步骤1:创建输入DataFrame
先把输入数据转换成PySpark DataFrame:
from pyspark.sql import SparkSession from pyspark.sql.functions import monotonically_increasing_id spark = SparkSession.builder.appName("ColumnCount").getOrCreate() data = [ (10, 20, 11, 20), (20, 11, 10, 99), (10, 11, 20, 1), (30, 12, 20, 99), (10, 11, 20, 20), (40, 13, 15, 3), (30, 8, 11, 99) ] df = spark.createDataFrame(data, ["A", "B", "C", "D"])
步骤2:分别统计各列计数并添加行号
对每一列分组计数后,添加自增行号作为后续join的关联键:
# 统计A列计数并添加行号 df_a = df.groupBy("A").count().withColumnRenamed("count", "A_Count").withColumn("rn", monotonically_increasing_id()) # 统计B列计数并添加行号 df_b = df.groupBy("B").count().withColumnRenamed("count", "B_Count").withColumn("rn", monotonically_increasing_id()) # 统计C列计数并添加行号 df_c = df.groupBy("C").count().withColumnRenamed("count", "C_Count").withColumn("rn", monotonically_increasing_id()) # 统计D列计数并添加行号 df_d = df.groupBy("D").count().withColumnRenamed("count", "D_Count").withColumn("rn", monotonically_increasing_id())
步骤3:通过行号关联所有统计结果
将四个统计后的DataFrame按行号rn关联,最后移除行号列:
result_df = df_a.join(df_b, on="rn").join(df_c, on="rn").join(df_d, on="rn").drop("rn") result_df.show()
执行后即可得到目标DataFrame:
+---+-------+---+-------+---+-------+---+-------+ | A|A_Count| B|B_Count| C|C_Count| D|D_Count| +---+-------+---+-------+---+-------+---+-------+ | 10| 3| 8| 1| 10| 1| 1| 1| | 20| 1| 11| 3| 11| 2| 3| 1| | 30| 2| 12| 1| 15| 1| 20| 2| | 40| 1| 13| 1| 20| 3| 99| 3| +---+-------+---+-------+---+-------+---+-------+
说明
之前join未得到正确结果,是因为统计后的各DataFrame没有共同关联字段。添加行号后,每个统计结果的行可通过行号一一对应,从而实现并排展示的效果。
内容的提问来源于stack exchange,提问作者Satyam Singh
相关产品推荐
相关产品推荐

