You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何阻止tidymodels将数值型字段转为字符型?

问题描述

我在使用tidymodels的recipe构建模型时遇到了一个问题,流程如下:

  1. 导入数据:myData = read.csv('file'),此时各字段的数值/非数值类型均正确;
  2. 数据处理后拆分数据集:
train_test_split <- initial_split(data=myManipulatedData, prop=MySplitPct)
train_data <- train_test_split %>% training()
test_data <- train_test_split %>% testing()

验证过test_data中的数值字段类型正确,例如y_coordinate的summary()输出显示为数值型统计结果;
3. 创建recipe:

my_recipe <- recipe(outcome ~ numeric_1 + numeric_2 + ... + string_1 + ..., data=test_data) %>%
  update_role(numeric_n, numeric_m, new_role="ID")

但执行summary(my_recipe)时,所有数值字段的type列都显示为<chr [2]>,明明test_data里的类型是对的,请问这是什么原因?

解决方案

1. 纠正recipe的创建数据源(核心问题)

你当前用测试集test_data创建recipe是错误的操作——tidymodels的最佳实践要求必须基于训练集train_data构建recipe。预处理逻辑(包括类型识别、统计量计算)都应该只从训练集获取,避免数据泄露。用测试集初始化recipe可能导致类型识别异常,甚至后续模型训练出现偏差。

修改后的代码:

my_recipe <- recipe(outcome ~ numeric_1 + numeric_2 + ... + string_1 + ..., data=train_data) %>%
  update_role(numeric_n, numeric_m, new_role="ID")

2. 确认数据类型的准确性

不要只依赖summary()检查类型,用str(test_data)或glimpse(test_data)查看变量的原始类型——summary()对字符型的数字不会输出数值统计,所以如果它显示了Min/Max等信息,说明变量确实是数值型,但str()能更直观确认类型是否正确。

3. 更新recipes包版本

老版本的recipes包可能存在类型识别的bug,执行以下命令更新到最新版:

update.packages("recipes")

4. 逐步排查问题

如果上述方法无效,尝试简化recipe逐步排查:

  • 先创建不含update_role的基础recipe:
    simple_recipe <- recipe(outcome ~ numeric_1, data=train_data)
    summary(simple_recipe)
    
    查看numeric_1的type是否显示为numeric;
  • 若正常,再逐步添加其他变量和update_role步骤,定位导致类型异常的环节。

内容的提问来源于stack exchange,提问作者Elliott Barinberg

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 07:37:39