R语言处理FIFA数据集时如何绘制同列变量图表并排除NA值
解决方案
你只需要提前过滤掉原数据中League列为NA的观测即可,修改后代码如下:
# 先过滤掉League列是NA的行,仅保留你需要的12个联赛数据 df_filtered <- df %>% filter(!is.na(League)) # 基于过滤后的数据统计各联赛平均年龄 plotageleague <- df_filtered %>% group_by(League) %>% summarise(age = mean(Age, na.rm = TRUE)) options(repr.plot.width = 12, repr.plot.height = 8) ggplot()+ geom_histogram(df_filtered, mapping = aes(Age, fill = League), bins = 15)+ geom_vline(plotageleague, mapping = aes(xintercept = age), color = "red", size = 1.5)+ geom_text(plotageleague, mapping = aes(x = age+3, y = 65, label = round(age,digits = 2)))+ facet_wrap(League~.)+ theme_minimal()+ theme(legend.position = "bottom")+ labs(y = "Frequency", title = "", caption = "")
- 代码中给
mean()函数加了na.rm=TRUE参数,避免Age列存在缺失值时平均年龄计算结果为NA - 可以自行调整
geom_histogram的bins参数优化年龄分布的展示效果 - 如果不需要单独存储过滤后的数据集,也可以直接在绘图/统计的数据源位置嵌套
filter(!is.na(League))逻辑
内容的提问来源于stack exchange,提问作者Enes Kilic
相关产品推荐
相关产品推荐

