R语言:基于DataFrame的tota_score列处理expected_score列数据
解决R语言DataFrame的两个数据处理需求
需求回顾
- 过滤掉
tota_score列值为NA的行 - 仅保留
expected_score中**大于对应行tota_score**的值,其余情况(包括小于等于、本身为NA)都设为NA
现有代码的问题
你写的ifelse(!is.na(expected_score) > !is.na(tota_score), NA, expected_score)逻辑完全错误——!is.na()返回布尔值(TRUE/FALSE),比较布尔值时会自动转成1和0,这根本不是在判断expected_score和tota_score的数值大小,自然得不到预期结果。
正确实现代码
使用dplyr包按顺序完成两个需求:
library(dplyr) # 初始数据框 df <- data.frame(tota_score=c(4.5,12.2,4.6,9.2,12.2,NA,4.5,12.2,4.6,9.2,12.2,36.4), expected_score=c(4.5,12.1,NA,10,12.2,36.4,5,12.5,NA,9.2,16,NA), Region1=c("All region",NA,NA,"All region","All region",NA,"All region",NA,NA,"All region","All region",NA), Region2=c("EAST","EAST","EAST","EAST","EAST",NA,"EAST","EAST","EAST","EAST","EAST",NA), Region3=c("West",NA,"West","West","West","West","West",NA,"West","West","West","West")) # 数据处理 df_processed <- df %>% # 过滤tota_score为NA的行 filter(!is.na(tota_score)) %>% # 仅保留expected_score大于tota_score的值,其余设为NA mutate(expected_score = ifelse(expected_score > tota_score, expected_score, NA)) # 查看处理结果 print(df_processed)
代码解释
filter(!is.na(tota_score)):直接筛除tota_score为NA的行,完成第一个需求;mutate(...)中的ifelse判断:- 当
expected_score数值大于对应行的tota_score时,保留原数值; - 所有不满足条件的情况(包括
expected_score小于等于tota_score、expected_score本身为NA),统一设为NA,完全匹配第二个需求。
- 当
处理后结果
tota_score expected_score Region1 Region2 Region3 1 4.5 NA All region EAST West 2 12.2 NA <NA> EAST <NA> 3 4.6 NA <NA> EAST West 4 9.2 10.0 All region EAST West 5 12.2 NA All region EAST West 6 4.5 5.0 All region EAST West 7 12.2 12.5 <NA> EAST <NA> 8 4.6 NA <NA> EAST West 9 9.2 NA All region EAST West 10 12.2 16.0 All region EAST West 11 36.4 NA <NA> <NA> West
内容的提问来源于stack exchange,提问作者potro
相关产品推荐
相关产品推荐

