无法从R语言数据集移除指定非国家条目行的问题排查
问题:无法移除数据集中的非国家条目
我正在处理全球心理健康障碍数据集,为提升图表准确性,希望移除其中的非国家条目。以下是我使用的代码:
Mental_health_Depression_disorder_Data <- mentalhealth #finding the average rate for Schizophrenia globally onlyschizophrenia <- mentalhealth %>% select(Entity, Year, `Schizophrenia (%)`) onlyschizophrenianona <- na.omit(onlyschizophrenia) test1 <- onlyschizophrenianona %>% filter(`Schizophrenia (%)` < 1) %>% mutate(`Schizophrenia (%)` = `Schizophrenia (%)` * 100) %>% select(Entity, `Schizophrenia (%)`, Year) #non-countries noncountries <- c( "Andean Latin America", "Central African Republic", "Central Asia", "Central Europe, Eastern Europe, and Central Asia", "Central Latin America", "Central Sub-Saharan Africa", "East Asia", "Eastern Europe", "Eastern Sub-Saharan Africa", "High SDI", "High-income", "High-income Asia Pacific", "High-middle SDI", "Latin America and Caribbean", "Low SDI", "Low-middle SDI", "North Africa and Middle East", "North America", "Northern Ireland", "Oceania", "South Asia", "Southeast Asia", "Southeast Asia, East Asia, and Oceania", "Southern Latin America", "Southern Sub-Saharan Africa", "United States Virgin Islands", "Western Europe", "Western Sub-Saharan Africa" ) #removing non-countries #method 1 cleandata1 <- test1[!(row.names(test1) %in% noncountries), ] view(cleandata1) #method 2 trial4 <- subset(test1 , test1$Entity != noncountries) view(trial4)
尝试的两种方法都未能成功移除test1中的指定非国家条目,请问问题出在哪里?
问题分析与解决方法
两种方法的错误原因
- 方法1错误:
row.names(test1)获取的是数据的数字行索引,但你要排除的非国家名称存在Entity列里,用行名去匹配noncountries完全不对应,自然筛选无效。 - 方法2错误:R中
!=是逐元素比较,当你用test1$Entity != noncountries时,noncountries是向量,会触发循环补齐机制——只会保留Entity不等于noncountries第一个元素的行,而非排除所有在列表中的元素,导致大部分目标条目没被过滤。
正确解决方法
方法一:用dplyr语法(和现有代码风格统一)
# 精准排除noncountries中的所有条目 cleandata <- test1 %>% filter(!(Entity %in% noncountries)) View(cleandata)
方法二:用base R的subset函数
cleandata_base <- subset(test1, !(Entity %in% noncountries)) View(cleandata_base)
核心逻辑是用%in%判断Entity列的元素是否在noncountries列表中,加上!取反,就能准确移除所有指定的非国家条目。
内容的提问来源于stack exchange,提问作者Parvitha Ramesh Rao
相关产品推荐
相关产品推荐

