R语言:在mutate中使用ifelse引用外部数据框进行性别预测时的索引错误解决
解决R中dplyr mutate里基于匹配数据框行的预测问题
嘿,我完全懂你碰到的这个麻烦——这个错误的核心原因是向量长度不匹配:你在mutate里写model$height == height时,右边的height是来自test_height的整列(长度526),但model只有28行,用长逻辑向量去索引短数据框,自然会触发下标不匹配的报错。
咱们来用两种更靠谱的方式解决这个问题:
方法1:用left_join合并匹配数据(推荐,更符合tidyverse风格)
先把model里对应身高的女性概率合并到测试数据集里,再做预测,这样每一行都能拿到自己对应的概率,完全不会有长度问题:
library(tidyverse) library(dslabs) library(caret) # 划分训练/测试集 ind <- createDataPartition(heights$height, 1, p=0.5, list = FALSE) test_height <- heights %>% slice(ind) train_height <- heights %>% slice(-ind) # 生成身高-女性概率模型 model <- train_height %>% mutate(height=round(height)) %>% group_by(height) %>% summarise(pf = mean(sex=="Female"), .groups = "drop") # .groups参数避免分组警告 # 合并概率并生成预测结果 y_hat <- test_height %>% mutate(height=round(height)) %>% left_join(model, by = "height") %>% mutate(pred = ifelse(pf > 0.5, "Female", "Male"))
如果测试集中出现了训练集里没有的身高值,left_join会把对应的pf设为NA,你可以用case_when补充处理逻辑:
y_hat <- test_height %>% mutate(height=round(height)) %>% left_join(model, by = "height") %>% mutate(pred = case_when( pf > 0.5 ~ "Female", pf <= 0.5 ~ "Male", is.na(pf) ~ "Male" # 对未见过的身高,默认预测为Male ))
方法2:用match函数匹配对应索引
如果你不想用合并的方式,也可以用match函数直接找到每个测试身高在model中的位置,再提取对应的概率:
y_hat <- test_height %>% mutate(height=round(height)) %>% mutate(pred = ifelse(model$pf[match(height, model$height)] > 0.5, "Female", "Male"))
match(height, model$height)会返回测试集中每个身高在model$height里的位置索引,用这个索引取pf值,就能保证向量长度完全匹配。
内容的提问来源于stack exchange,提问作者Hans
相关产品推荐
相关产品推荐

