R语言中将字符型year转为因子时出现NA问题的排查与解决
问题:将year列转换为因子时全部得到NA的原因
我有如下数据集:
years <- c("2010", "2011", "2012", "2013", "2014", "2015", "2016", "2017", "2018", "2019", "2020") n_cohorts <- length(years) df <- structure(list(label2plot = structure(c(1L, 2L, 1L, 2L, 1L, 2L, 1L, 2L, 1L, 2L, 1L, 2L, 1L, 2L, 1L, 2L, 1L, 2L, 1L, 2L, 1L, 2L ), levels = c("aaa", "bbb"), class = c("ordered", "factor")), var_abs = c(717L, 569L, 860L, 752L, 713L, 575L, 918L, 724L, 946L, 764L, 951L, 764L, 784L, 691L, 672L, 610L, 833L, 671L, 773L, 620L, 532L, 293L), var_rel = c(0.557542768273717, 0.442457231726283, 0.533498759305211, 0.466501240694789, 0.553571428571429, 0.446428571428571, 0.559074299634592, 0.440925700365408, 0.553216374269006, 0.446783625730994, 0.554518950437318, 0.445481049562682, 0.531525423728814, 0.468474576271186, 0.52418096723869, 0.47581903276131, 0.553856382978723, 0.446143617021277, 0.554917444364681, 0.445082555635319, 0.644848484848485, 0.355151515151515 ), year = c("2020", "2020", "2019", "2019", "2018", "2018", "2017", "2017", "2016", "2016", "2015", "2015", "2014", "2014", "2013", "2013", "2012", "2012", "2011", "2011", "2010", "2010" )), row.names = c(NA, -22L), class = "data.frame") cohorts <- length(years)
我希望将其中的year列转换为因子,执行了如下代码:
df$year2 <- factor(df$year, levels = c(1:n_cohorts), labels = years)
但得到的结果中year2列全部为NA,无法解释原因,请问我哪里操作错误?结果如下:
structure(list(label2plot = structure(c(1L, 2L, 1L, 2L, 1L, 2L, 1L, 2L, 1L, 2L, 1L, 2L, 1L, 2L, 1L, 2L, 1L, 2L, 1L, 2L, 1L, 2L ), levels = c("aaa", "bbb"), class = c("ordered", "factor")), var_abs = c(717L, 569L, 860L, 752L, 713L, 575L, 918L, 724L, 946L, 764L, 951L, 764L, 784L, 691L, 672L, 610L, 833L, 671L, 773L, 620L, 532L, 293L), var_rel = c(0.557542768273717, 0.442457231726283, 0.533498759305211, 0.466501240694789, 0.553571428571429, 0.446428571428571, 0.559074299634592, 0.440925700365408, 0.553216374269006, 0.446783625730994, 0.554518950437318, 0.445481049562682, 0.531525423728814, 0.468474576271186, 0.52418096723869, 0.47581903276131, 0.553856382978723, 0.446143617021277, 0.554917444364681, 0.445082555635319, 0.644848484848485, 0.355151515151515), year = c("2020", "2020", "2019", "2019", "2018", "2018", "2017", "2017", "2016", "2016", "2015", "2015", "2014", "2014", "2013", "2013", "2012", "2012", "2011", "2011", "2010", "2010"), year2 = structure(c(NA_integer_, NA_integer_, NA_integer_, NA_integer_, NA_integer_, NA_integer_, NA_integer_, NA_integer_, NA_integer_, NA_integer_, NA_integer_, NA_integer_, NA_integer_, NA_integer_, NA_integer_, NA_integer_, NA_integer_, NA_integer_, NA_integer_, NA_integer_, NA_integer_, NA_integer_ ), levels = c("2010", "2011", "2012", "2013", "2014", "2015", "2016", "2017", "2018", "2019", "2020"), class = "factor")), row.names = c(NA, -22L), class = "data.frame")
错误原因与解决方案
错误原因
你在factor()函数中设置的levels参数是c(1:n_cohorts),也就是数字序列1到11,但原始的year列是字符串类型的年份值(比如"2020"、"2019"等),两者的取值完全不匹配,R无法找到对应关系,因此所有值都被转换为NA。
正确解决方案
只需要将levels参数设置为years即可,因为years里的元素就是year列的所有取值,同时还能指定你想要的因子顺序(2010到2020):
df$year2 <- factor(df$year, levels = years)
执行这段代码后,year2列会正确对应原始year的取值,并且因子的顺序按照years的顺序排列,不会出现NA。
内容的提问来源于stack exchange,提问作者Dierforth
相关产品推荐
相关产品推荐

