R语言KNN算法前含数值与因子型数据集的标准化问询
Hey there! 刚接触数据分析就敢上手KNN和3万条数据,这起步相当扎实啊!针对你卡在标准化这一步的问题,我给你梳理一套完整的流程,包括自定义函数的实现,完全适配你的场景:
一、先搞懂为什么KNN必须做标准化
KNN是基于距离计算的算法,数值型变量的量纲差异会直接干扰距离权重——比如age(范围0-100)和收入类变量(范围0-10000),如果不标准化,收入的数值波动会完全盖过年龄的影响,导致模型结果偏差。而因子型变量不需要标准化,但要做编码处理才能被KNN识别。
二、标准化完整流程方案
1. 第一步:拆分变量类型
先把你提取的7个变量分成两类:
- 数值型:比如你提到的
age(int类型),还有其他连续/整数型变量 - 因子型:分类变量(比如性别、婚姻状态这类离散值)
2. 选择标准化方法
根据变量分布选适合的方式:
- Z-score标准化:把变量转化为均值=0、标准差=1的分布,适合变量近似正态分布的情况,公式:
(x - 均值) / 标准差 - Min-Max标准化:把变量缩放到[0,1]区间,适合变量无明显分布规律的情况,公式:
(x - 最小值) / (最大值 - 最小值)
KNN对两种方法都兼容,你可以根据变量情况二选一。
3. 自定义标准化函数(附Python/R双版本)
Python版本(适配pandas数据框)
这个函数会自动区分变量类型,处理数值型标准化+因子型独热编码:
import pandas as pd import numpy as np def custom_standardize(df, num_method='z-score'): # 拆分数值型与因子型列 num_cols = df.select_dtypes(include=['int64', 'float64']).columns cat_cols = df.select_dtypes(include=['object', 'category']).columns # 处理数值型变量 num_df = df[num_cols].copy() if num_method == 'z-score': for col in num_cols: mean_val = num_df[col].mean() std_val = num_df[col].std() num_df[col] = (num_df[col] - mean_val) / std_val elif num_method == 'min-max': for col in num_cols: min_val = num_df[col].min() max_val = num_df[col].max() num_df[col] = (num_df[col] - min_val) / (max_val - min_val) # 处理因子型变量(独热编码,避免标签编码的顺序误导) cat_df = pd.get_dummies(df[cat_cols], drop_first=True) # 合并处理后的结果 standardized_df = pd.concat([num_df, cat_df], axis=1) return standardized_df
R版本(适配data.frame)
custom_standardize <- function(df, num_method = "z-score") { # 识别数值型与因子型变量 num_cols <- sapply(df, is.numeric) cat_cols <- sapply(df, is.factor) # 处理数值型变量 num_df <- df[, num_cols, drop = FALSE] if (num_method == "z-score") { num_df <- scale(num_df) # R原生scale函数就是Z-score标准化 } else if (num_method == "min-max") { num_df <- apply(num_df, 2, function(x) (x - min(x)) / (max(x) - min(x))) } # 处理因子型变量(独热编码) cat_df <- model.matrix(~ . -1, data = df[, cat_cols, drop = FALSE]) # 合并结果 standardized_df <- cbind(as.data.frame(num_df), as.data.frame(cat_df)) return(standardized_df) }
4. 应用到你的数据集
假设你提取的7个变量存在selected_data数据框里,直接调用函数即可:
Python示例
# 用Z-score标准化 standardized_data = custom_standardize(selected_data, num_method='z-score') # 或者用Min-Max标准化 # standardized_data = custom_standardize(selected_data, num_method='min-max')
R示例
standardized_data <- custom_standardize(selected_data, num_method = "z-score")
5. 验证标准化结果
做完后可以简单检查:
- Z-score标准化的数值型变量,均值接近0、标准差接近1
- Min-Max标准化的数值型变量,范围在[0,1]之间
- 因子型变量已转化为0/1的独热编码列
三、关键注意事项
- 避免数据泄露:如果后续要拆分训练/测试集,一定要用**训练集的统计量(均值、最值等)**来标准化测试集,不能用整个数据集的统计量!比如自定义函数时,可以先保存训练集的统计参数,再用它处理测试集。
- 因子型编码选择:如果因子是有序的(比如学历:小学<中学<大学),可以用标签编码;如果是无序的(比如性别、职业),优先用独热编码,避免模型误以为变量有顺序关系。
内容的提问来源于stack exchange,提问作者Fred11
相关产品推荐
相关产品推荐

