R中Recipe的step函数归一化数值预测变量无变化问题
R语言recipe包step函数执行后数据无变化的解决方法
问题描述
使用R语言的recipe包调用step_range函数对iris数据集的数值预测变量做0-1归一化,执行后无报错,但用head()查看结果时,数据完全没有变化,两次尝试的代码及输出如下:
初始尝试代码
iris_recipe<-iris %>% recipe(Species ~ .) %>% step_range(recipe, all_numeric_predictors(), min = 0, max=1) head(iris_recipe)
初始尝试输出
$var_info # A tibble: 5 × 4 variable type role source <chr> <chr> <chr> <chr> 1 Sepal.Length numeric predictor original 2 Sepal.Width numeric predictor original 3 Petal.Length numeric predictor original 4 Petal.Width numeric predictor original 5 Species nominal outcome original $term_info # A tibble: 5 × 4 variable type role source <chr> <chr> <chr> <chr> 1 Sepal.Length numeric predictor original 2 Sepal.Width numeric predictor original 3 Petal.Length numeric predictor original 4 Petal.Width numeric predictor original 5 Species nominal outcome original $steps $steps[[1]] $terms <list_of<quosure>> [[1]] <quosure> expr: ^recipe env: 0x000002b201dc23e0 [[2]] <quosure> expr: ^all_numeric_predictors() env: 0x000002b201dc23e0 $role [1] NA $trained [1] FALSE $min [1] 0 $max [1] 1 $ranges NULL $skip [1] FALSE $id [1] "range_5JGit" attr(,"class") [1] "step_range" "step" $template # A tibble: 150 × 5 Sepal.Length Sepal.Width Petal.Length Petal.Width Species <dbl> <dbl> <dbl> <dbl> <fct> 1 5.1 3.5 1.4 0.2 setosa 2 4.9 3 1.4 0.2 setosa 3 4.7 3.2 1.3 0.2 setosa 4 4.6 3.1 1.5 0.2 setosa 5 5 3.6 1.4 0.2 setosa 6 5.4 3.9 1.7 0.4 setosa 7 4.6 3.4 1.4 0.3 setosa 8 5 3.4 1.5 0.2 setosa 9 4.4 2.9 1.4 0.2 setosa 10 4.9 3.1 1.5 0.1 setosa # … with 140 more rows # ℹ Use `print(n = ...)` to see more rows $levels NULL $retained [1] NA
特定尝试代码
iris_recipe2<-iris %>% recipe(Species ~ .) %>% step_range(recipe, Sepal.Length, min=0, max=1) head(iris_recipe2)
特定尝试输出
$var_info # A tibble: 5 × 4 variable type role source <chr> <chr> <chr> <chr> 1 Sepal.Length numeric predictor original 2 Sepal.Width numeric predictor original 3 Petal.Length numeric predictor original 4 Petal.Width numeric predictor original 5 Species nominal outcome original $term_info # A tibble: 5 × 4 variable type role source <chr> <chr> <chr> <chr> 1 Sepal.Length numeric predictor original 2 Sepal.Width numeric predictor original 3 Petal.Length numeric predictor original 4 Petal.Width numeric predictor original 5 Species nominal outcome original $steps $steps[[1]] $terms <list_of<quosure>> [[1]] <quosure> expr: ^recipe env: 0x000002b2073fe328 [[2]] <quosure> expr: ^Sepal.Length env: 0x000002b2073fe328 $role [1] NA $trained [1] FALSE $min [1] 0 $max [1] 1 $ranges NULL $skip [1] FALSE $id [1] "range_p7JFo" attr(,"class") [1] "step_range" "step" $template # A tibble: 150 × 5 Sepal.Length Sepal.Width Petal.Length Petal.Width Species <dbl> <dbl> <dbl> <dbl> <fct> 1 5.1 3.5 1.4 0.2 setosa 2 4.9 3 1.4 0.2 setosa 3 4.7 3.2 1.3 0.2 setosa 4 4.6 3.1 1.5 0.2 setosa 5 5 3.6 1.4 0.2 setosa 6 5.4 3.9 1.7 0.4 setosa 7 4.6 3.4 1.4 0.3 setosa 8 5 3.4 1.5 0.2 setosa 9 4.4 2.9 1.4 0.2 setosa 10 4.9 3.1 1.5 0.1 setosa # … with 140 more rows # ℹ Use `print(n = ...)` to see more rows $levels NULL $retained [1] NA
问题原因
- step_range参数错误:使用管道
%>%时,前一步的recipe对象会自动传入step_range,不需要手动在step_range的第一个参数写recipe,这个错误导致函数识别的处理对象错误。 - 未执行训练与应用流程:
recipe包的step系列函数只是定义数据处理的流程,并没有实际计算处理规则(比如归一化的min/max值)和应用到数据上,必须通过prep()训练流程,再用bake()或juice()提取处理后的数据。
正确解决方案
代码示例
# 加载包 library(recipes) library(dplyr) # 1. 定义处理流程(注意step_range的参数不再写recipe) iris_recipe <- iris %>% recipe(Species ~ .) %>% step_range(all_numeric_predictors(), min = 0, max = 1) # 2. 训练处理流程(计算归一化所需的统计量) iris_recipe_trained <- prep(iris_recipe, training = iris) # 3. 提取处理后的数据(juice用于提取训练集处理结果,bake可用于新数据) iris_processed <- juice(iris_recipe_trained) # 查看结果 head(iris_processed)
输出示例
Sepal.Length Sepal.Width Petal.Length Petal.Width Species 1 0.2222222 0.6250000 0.0677966 0.0416667 setosa 2 0.1666667 0.4166667 0.0677966 0.0416667 setosa 3 0.1111111 0.5000000 0.0508475 0.0416667 setosa 4 0.0833333 0.4583333 0.0847458 0.0416667 setosa 5 0.1944444 0.6666667 0.0677966 0.0416667 setosa 6 0.3055556 0.7916667 0.1186441 0.1250000 setosa
关键说明
prep():根据训练数据计算处理步骤所需的统计量(比如归一化的最小值、最大值),将流程标记为已训练(trained = TRUE)。juice():从已训练的流程中提取处理后的训练数据。bake():如果有新数据,用已训练的流程处理新数据,比如bake(iris_recipe_trained, new_data = new_iris_data)。
内容的提问来源于stack exchange,提问作者Miner_problems
相关产品推荐
相关产品推荐

