PySpark单列透视多聚合:如何高效转换目标DataFrame?
Solution in PySpark
Here's an efficient, generalizable way to achieve your desired pivot transformation using PySpark's built-in functions:
Step 1: Compute Feature Averages by Day
First, calculate the average of each feature grouped by the Day column. This gives a condensed DataFrame with daily averages for each feature:
from pyspark.sql import functions as F # Calculate daily averages avg_by_day = df.groupBy("Day").agg( F.avg("Feature A").alias("AvgA"), F.avg("Feature B").alias("AvgB") )
For your sample data, this produces:
+----+----+----+ | Day|AvgA|AvgB| +----+----+----+ | 01| 32| 2| | 02| 44| 4| | 03| 52| 3| +----+----+----+
Step 2: Unpivot Feature Columns into Rows
Use the stack function to convert the wide-format average columns (AvgA, AvgB) into a long-format with rows for each feature-day pair:
# Unpivot features into rows unpivoted = avg_by_day.select( F.expr("stack(2, 'AvgA', AvgA, 'AvgB', AvgB) as (`Feat\\Day`, Value)"), "Day" )
This results in:
+------------+-----+----+ | Feat\Day |Value| Day| +------------+-----+----+ | AvgA| 32| 01| | AvgB| 2| 01| | AvgA| 44| 02| | AvgB| 4| 02| | AvgA| 52| 03| | AvgB| 3| 03| +------------+-----+----+
Step 3: Pivot on Day to Get Final Structure
Finally, group by the feature name and pivot on Day to reshape the data into your target format:
# Pivot to get days as columns result = unpivoted.groupBy("Feat\\Day").pivot("Day").agg(F.first("Value"))
The output matches your desired DataFrame:
+------------+---+---+---+ | Feat\Day | 01| 02| 03| +------------+---+---+---+ | AvgA| 32| 44| 52| | AvgB| 2| 4| 3| +------------+---+---+---+
Why This Works (and Is Efficient)
- Optimized Spark Operations: All steps use PySpark's native, optimized functions (
groupBy,agg,stack,pivot) instead of manual concatenation or slow loops. - Generalizable: To add more features (e.g.,
AvgC,AvgD), simply extend thestackfunction with additional entries (e.g.,stack(3, 'AvgA', AvgA, 'AvgB', AvgB, 'AvgC', AvgC)). - Scalable: This approach works efficiently even with large datasets, as Spark handles distributed processing under the hood.
内容的提问来源于stack exchange,提问作者dangus poochie
相关产品推荐
相关产品推荐

