Spark MLlib数据集标准化与Min-Max缩放的数学原理解析请求
Hey there! Let's break down exactly how those scaled values are calculated for both StandardScaler and MinMaxScaler using your dataset. I'll walk through the math step by step with your actual data points so you can see where each number comes from.
First, note that Spark's StandardScaler uses default parameters: withMean=false and withStd=true. This means each feature value is scaled by dividing it by the population standard deviation of that feature (we divide by the total number of data points, not n-1, which is used for sample standard deviation).
The formula for each scaled value is:
scaled_value = original_value / population_standard_deviation_of_feature
Let's calculate this for each feature dimension using your dataset:
Your features column has 5 vectors:[1.0,0.1,-1.0], [2.0,1.1,1.0], [1.0,0.1,-1.0], [2.0,1.1,1.0], [3.0,10.1,3.0]
Feature 1 (first element in each vector)
Original values: 1.0, 2.0, 1.0, 2.0, 3.0
- Mean = (1+2+1+2+3)/5 = 1.8
- Sum of squared differences from mean: (1-1.8)² + (2-1.8)² + (1-1.8)² + (2-1.8)² + (3-1.8)² = 0.64 + 0.04 + 0.64 + 0.04 + 1.44 = 2.8
- Population standard deviation = √(2.8/5) ≈ 0.7483
- Scaled values:
- 1.0 / 0.7483 ≈ 1.1952 (matches your first row's scaled value)
- 2.0 / 0.7483 ≈ 2.3905 (matches your second row's scaled value)
- 3.0 / 0.7483 ≈ 3.5857 (matches your fifth row's scaled value)
Feature 2 (second element in each vector)
Original values: 0.1, 1.1, 0.1, 1.1, 10.1
- Mean = (0.1+1.1+0.1+1.1+10.1)/5 = 2.5
- Sum of squared differences from mean: (0.1-2.5)² + (1.1-2.5)² + (0.1-2.5)² + (1.1-2.5)² + (10.1-2.5)² = 5.76 + 1.96 + 5.76 + 1.96 + 57.76 = 73.2
- Population standard deviation = √(73.2/5) ≈ 3.8262
- Scaled values:
- 0.1 / 3.8262 ≈ 0.0234 (matches your first row's scaled value)
- 1.1 / 3.8262 ≈ 0.2571 (matches your second row's scaled value)
- 10.1 / 3.8262 ≈ 2.3610 (matches your fifth row's scaled value)
Feature 3 (third element in each vector)
Original values: -1.0, 1.0, -1.0, 1.0, 3.0
- Mean = (-1+1-1+1+3)/5 = 0.6
- Sum of squared differences from mean: (-1-0.6)² + (1-0.6)² + (-1-0.6)² + (1-0.6)² + (3-0.6)² = 2.56 + 0.16 + 2.56 + 0.16 + 5.76 = 11.2
- Population standard deviation = √(11.2/5) ≈ 1.4967
- Scaled values:
- -1.0 / 1.4967 ≈ -0.5976 (matches your first row's scaled value)
- 1.0 / 1.4967 ≈ 0.5976 (matches your second row's scaled value)
- 3.0 / 1.4967 ≈ 1.7928 (matches your fifth row's scaled value)
You configured MinMaxScaler with setMin(5) and setMax(10). The formula for each scaled value is:
scaled_value = min_target + (max_target - min_target) * (original_value - min_feature) / (max_feature - min_feature)
Where:
min_target = 5,max_target =10min_feature= minimum value of the feature across all data pointsmax_feature= maximum value of the feature across all data points
Again, let's calculate for each feature dimension:
Feature 1
Original values: 1.0,2.0,1.0,2.0,3.0 → min_feature=1.0, max_feature=3.0
- Scaled values:
- 5 + (10-5)*(1.0-1.0)/(3.0-1.0) = 5 + 0 = 5.0 (matches first row)
- 5 + 5*(2.0-1.0)/(3.0-1.0) =5 + 2.5=7.5 (matches second row)
-5 +5*(3.0-1.0)/(3.0-1.0)=5+5=10.0 (matches fifth row)
Feature 2
Original values: 0.1,1.1,0.1,1.1,10.1 → min_feature=0.1, max_feature=10.1
- Scaled values:
-5 +5*(0.1-0.1)/(10.1-0.1)=5+0=5.0 (matches first row)
-5 +5*(1.1-0.1)/(10.1-0.1)=5+0.5=5.5 (matches second row)
-5 +5*(10.1-0.1)/(10.1-0.1)=5+5=10.0 (matches fifth row)
Feature 3
Original values: -1.0,1.0,-1.0,1.0,3.0 → min_feature=-1.0, max_feature=3.0
- Scaled values:
-5 +5*(-1.0 - (-1.0))/(3.0 - (-1.0))=5+0=5.0 (matches first row)
-5 +5*(1.0 - (-1.0))/(3.0 - (-1.0))=5+2.5=7.5 (matches second row)
-5 +5*(3.0 - (-1.0))/(3.0 - (-1.0))=5+5=10.0 (matches fifth row)
内容的提问来源于stack exchange,提问作者Kumar-Sandeep

