You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark单列透视多聚合:如何高效转换目标DataFrame?

Solution in PySpark

Here's an efficient, generalizable way to achieve your desired pivot transformation using PySpark's built-in functions:

Step 1: Compute Feature Averages by Day

First, calculate the average of each feature grouped by the Day column. This gives a condensed DataFrame with daily averages for each feature:

from pyspark.sql import functions as F

# Calculate daily averages
avg_by_day = df.groupBy("Day").agg(
    F.avg("Feature A").alias("AvgA"),
    F.avg("Feature B").alias("AvgB")
)

For your sample data, this produces:

+----+----+----+
| Day|AvgA|AvgB|
+----+----+----+
|  01|  32|   2|
|  02|  44|   4|
|  03|  52|   3|
+----+----+----+

Step 2: Unpivot Feature Columns into Rows

Use the stack function to convert the wide-format average columns (AvgA, AvgB) into a long-format with rows for each feature-day pair:

# Unpivot features into rows
unpivoted = avg_by_day.select(
    F.expr("stack(2, 'AvgA', AvgA, 'AvgB', AvgB) as (`Feat\\Day`, Value)"),
    "Day"
)

This results in:

+------------+-----+----+
|  Feat\Day  |Value| Day|
+------------+-----+----+
|        AvgA|   32|  01|
|        AvgB|    2|  01|
|        AvgA|   44|  02|
|        AvgB|    4|  02|
|        AvgA|   52|  03|
|        AvgB|    3|  03|
+------------+-----+----+

Step 3: Pivot on Day to Get Final Structure

Finally, group by the feature name and pivot on Day to reshape the data into your target format:

# Pivot to get days as columns
result = unpivoted.groupBy("Feat\\Day").pivot("Day").agg(F.first("Value"))

The output matches your desired DataFrame:

+------------+---+---+---+
|  Feat\Day  | 01| 02| 03|
+------------+---+---+---+
|        AvgA| 32| 44| 52|
|        AvgB|  2|  4|  3|
+------------+---+---+---+

Why This Works (and Is Efficient)

  • Optimized Spark Operations: All steps use PySpark's native, optimized functions (groupBy, agg, stack, pivot) instead of manual concatenation or slow loops.
  • Generalizable: To add more features (e.g., AvgC, AvgD), simply extend the stack function with additional entries (e.g., stack(3, 'AvgA', AvgA, 'AvgB', AvgB, 'AvgC', AvgC)).
  • Scalable: This approach works efficiently even with large datasets, as Spark handles distributed processing under the hood.

内容的提问来源于stack exchange,提问作者dangus poochie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 02:12:12