Spark MinMaxScaler按天归一化DataFrame并保留时间字段的实现
result by day with MinMaxScaler While Keeping time Field Got it, let's work through this problem together! You need to normalize the result column per day using MinMaxScaler, and keep the corresponding time values intact. Here's a straightforward way to do this with pandas and scikit-learn:
Step 1: Import Required Libraries
First, make sure you have pandas and scikit-learn installed, then import them:
import pandas as pd from sklearn.preprocessing import MinMaxScaler
Step 2: Create Your Sample DataFrame
Let's start with the data you provided:
df = pd.DataFrame({ 'day': [1, 1, 1, 2, 2], 'time': [6, 7, 8, 6, 10], 'result': [0.5, 10.2, 5.7, 11.0, 22.3] })
Step 3: Normalize result by day (With time Preserved)
We'll use pandas' groupby to split the data by day, then apply MinMaxScaler to each group's result column. The key is to keep all original columns (day, time, result) while adding the normalized values.
Option 1: Using groupby.apply() (More Readable)
Define a helper function to handle each group, then apply it:
scaler = MinMaxScaler() def normalize_single_day(group): # MinMaxScaler expects 2D input, so reshape the result column normalized_vals = scaler.fit_transform(group['result'].values.reshape(-1, 1)) # Add the normalized values as a new column in the group group['normalized_result'] = normalized_vals return group # Apply the function to each day group and combine back into a DataFrame normalized_df = df.groupby('day').apply(normalize_single_day)
Option 2: Using groupby.transform() (More Concise)
If you prefer a one-liner, use transform to apply the scaler directly:
scaler = MinMaxScaler() df['normalized_result'] = df.groupby('day')['result'].transform( lambda x: scaler.fit_transform(x.values.reshape(-1, 1)).flatten() )
Step 4: Check the Result
Running either option will give you a DataFrame with all original columns plus the normalized result:
day time result normalized_result 0 1 6 0.5 0.000000 1 1 7 10.2 1.000000 2 1 8 5.7 0.543046 3 2 6 11.0 0.000000 4 2 10 22.3 1.000000
Key Notes:
- Each day's
resultvalues are scaled independently to the [0, 1] range (since we fit the scaler per group). - The
timecolumn stays perfectly aligned with its originalresultvalue—no data gets lost or misaligned. - Using
reshape(-1, 1)is necessary becauseMinMaxScalerrequires input to be a 2D array (even if it's a single column).
内容的提问来源于stack exchange,提问作者YAKOVM

