如何动态移除数据框中同时为数值型且含Overall的列?
动态移除同时满足数值型且列名含"Overall"的列
我有数十份调查数据文件,每份包含若干数值型列与字符型列,需动态移除同时满足以下两个条件的列:
- 列是数值型;
- 列名包含"Overall"。
不能直接删除所有含"Overall"的列,因需保留的某字符型列标题也含该词;也无法按列名或位置删除,因不同文件的列位置与名称不统一,且并非所有文件都存在目标列。
示例数据框
#### reproducible example #### columns <- c("rating A", "rating B", "Student Overall Rating", "feedback 1", "feedback 2", "Student Overall Feedback") c1 <- c(4, 4, 3) c2 <- c(5, 4, 4) c3 <- c(4.5, 4, 3.5) c4 <- c("blah", "blah", "blah") c5 <- c("blah", "blah", "blah") c6 <- c("blahblah", "blahblah", "blahblah") df <- as.data.frame(cbind(c1, c2, c3, c4, c5, c6)) names(df) <- columns df$`rating A` <- as.numeric(df$`rating A`) df$`rating B` <- as.numeric(df$`rating B`) df$`Student Overall Rating` <- as.numeric(df$`Student Overall Rating`) str(df) # shows relative structure I am dealing with
运行str(df)输出:
'data.frame': 3 obs. of 6 variables: $ rating A : num 4 4 3 $ rating B : num 5 4 4 $ Student Overall Rating : num 4.5 4 3.5 $ feedback 1 : chr "blah" "blah" "blah" $ feedback 2 : chr "blah" "blah" "blah" $ Student Overall Feedback: chr "blahblah" "blahblah" "blahblah"
错误尝试及报错
尝试1
df <- df %>% select(!intersect(is.numeric(df), df %like% "Overall"))
报错信息:
Error in `select()`: ! Can't subset columns with `intersect(is.numeric(df), df %like% "Overall")`. ✖ `intersect(is.numeric(df), df %like% "Overall")` must be numeric or character, not `FALSE`.
尝试2
df <- df %>% select(!where(is.numeric | contains("Overall")))
报错信息:
Error in `select()`: ! Problem while evaluating `where(is.numeric | contains("Overall"))`. Caused by error in `is.numeric | contains("Overall")`: ! operations are possible only for numeric, logical or complex types
期望结果
移除数值型"Student Overall Rating"列后的数据框:
'data.frame': 3 obs. of 5 variables: $ rating A : num 4 4 3 $ rating B : num 5 4 4 $ feedback 1 : chr "blah" "blah" "blah" $ feedback 2 : chr "blah" "blah" "blah" $ Student Overall Feedback: chr "blahblah" "blahblah" "blahblah"
解决方案
要实现同时满足两个条件的列筛选并移除,可通过以下两种方式实现:
方法1:用dplyr的select()+where()自定义判断逻辑
在where()中传入自定义函数,同时检查列的类型和列名:
library(dplyr) df_cleaned <- df %>% select(!where(function(col) is.numeric(col) & grepl("Overall", cur_column()))) str(df_cleaned)
方法2:先提取目标列名再移除
先筛选出所有符合条件的列名,再通过列名批量移除:
cols_to_remove <- names(df)[sapply(df, is.numeric) & grepl("Overall", names(df))] df_cleaned <- df %>% select(-all_of(cols_to_remove)) str(df_cleaned)
两种方法都能动态识别所有满足「数值型+列名含Overall」的列并移除,同时保留字符型的含Overall列,适配不同文件的列结构差异,无需手动处理每份文件。
内容的提问来源于stack exchange,提问作者Brandon Signorino
相关产品推荐
相关产品推荐

