基于R语言ggplot2绘制双年份并列柱状图的技术需求
实现双年份犯罪统计分组柱状图
先明确你的原始数据情况:
原始数据集
df2010(2010年犯罪统计)
| type | Count |
|---|---|
| hate | 20'000 |
| drug | 15'600 |
| rape | 30'000 |
| murder | 32'000 |
| robbery | 45'700 |
| theft | 65'000 |
df2020(2020年犯罪统计)
| type | Count |
|---|---|
| hate | 42'000 |
| rape | 55'000 |
| murder | 16'000 |
| burglary | 11'800 |
| theft | 87'000 |
要把两张单年份柱状图合并成分组对比图,核心是先把数据转换成ggplot2适配的长格式,补全缺失年份的犯罪数据,再绘制分组柱形。
步骤1:数据预处理
首先修正Count列的格式(去掉单引号转成数值),再给数据集添加年份标识,最后合并并补全缺失值:
# 加载所需包 library(ggplot2) library(dplyr) library(tidyr) # 处理Count列:去掉单引号,转换为数值类型 df2010$Count <- as.numeric(gsub("'", "", df2010$Count)) df2020$Count <- as.numeric(gsub("'", "", df2020$Count)) # 给数据集添加年份标签 df2010 <- df2010 %>% mutate(year = "2010") df2020 <- df2020 %>% mutate(year = "2020") # 若你的2020年数据集实际名为crime_type_2020,替换此处即可 # 全连接合并,保留所有出现过的犯罪类型 combined_df <- full_join(df2010, df2020, by = "type") # 规整成长格式,补全缺失年份的案件数为0 combined_df <- combined_df %>% pivot_longer(cols = c(Count.x, Count.y), names_to = "temp_col", values_to = "Count") %>% mutate(year = ifelse(temp_col == "Count.x", "2010", "2020")) %>% select(type, year, Count) %>% replace_na(list(Count = 0))
步骤2:绘制分组柱状图
用position_dodge()实现同类型双柱分组,同时优化X轴标签可读性:
ggplot(combined_df, aes(x = type, y = Count, fill = year)) + geom_bar(stat = "identity", position = position_dodge(width = 0.8), width = 0.7) + xlab("犯罪类型") + ylab("案件数量") + ggtitle("2010 vs 2020 犯罪类型分布对比") + scale_fill_manual(values = c("2010" = "steelblue", "2020" = "coral")) + # 用不同颜色区分年份 theme(axis.text.x = element_text(angle = 90, vjust = 0.5, hjust = 1), legend.title = element_blank()) # 去掉冗余的图例标题
关键说明
full_join保证所有出现过的犯罪类型都被保留,哪怕某年份无对应数据- 缺失值补0后,无数据的年份会显示高度为0的柱子(若不想显示0柱,可保留NA,但X轴仍会显示该类型)
position_dodge()控制分组柱子的间距,避免重叠- 自定义颜色让两年的数据对比更直观
内容的提问来源于stack exchange,提问作者Maurice
相关产品推荐
相关产品推荐

