R语言ggplot绘制X轴为时间、Y轴为二分类变量1计数的散点/折线图
R语言ggplot2实现时间维度二分类变量计数可视化
前置依赖加载
首先加载绘图和数据处理需要的R包,未安装的话先运行安装命令:
# 安装依赖(首次运行执行即可) # install.packages(c("tidyverse", "lubridate")) library(dplyr) library(ggplot2) library(lubridate)
数据预处理
你的数据中时间列为created_at,二分类结果列为取值0/1的morality字段,预处理需要完成日期格式转换、时间粒度聚合、正向取值(值为1)计数三个步骤:
# 载入内置可复现示例数据,使用自己的真实数据时替换为对应读取逻辑即可 df <- structure(list(X.1 = 0:5, X = c(502026L, 198322L, 711188L, 563672L, 993641L, 474508L), tweet_id = c(867481042428579840, 469268704732393536, 915248573083553792, 689948979740725248, 1003463365811953664, 958533305716101120), user_username = c("GerryConnolly", "SenatorMenendez", "RepJayapal", "RoyBlunt", "SenJeffMerkley", "RepChrisStewart" ), text = c(".@governorva demonstrates compassion that potus lacks. trump's immigration eo still threatens to tear this family apart. #freeliliana", "hoy,repet<ed> mi llamado a mis colegas rep. de la c<e1>mara para q hagan lo correcto y aprueben una #reformamigraotira #cir", "@repadamsmith @reproybalallard the incarceration system for immigrants operates in the shadows, at a huge profit for corporations. our bill phases them out in 3 years.", "now isn't the time to accept syrian & iraqi refugees into our country w/o proper system for vetting. rt if you agree", "mr. president, the only <93>horrible law<94> is your policy. you have the power to change it. if you saw what i saw today, you would. never before has america deliberately inflicted cruelty on children to deter asylum seekers from finding refuge here. never. and we never should.", "republicans and democrats need to work together and reform our immigration policies. #sotu" ), created_at = c("2017-05-24", "2014-05-22", "2017-10-03", "2016-01-20", "2018-06-04", "2018-01-31"), user_name = c("Rep. Gerry Connolly", "Senator Bob Menendez", "Rep. Pramila Jayapal", "Senator Roy Blunt", "Senator Jeff Merkley", "Rep. Chris Stewart"), user_verified = c("True", "True", "True", "True", "True", "True"), morality_binary = c(0.78794396, 0.06992793, 0.75065666, 0.7655833, 0.85510856, 0.52538866), morality = c(1, 0, 1, 1, 1, 1)), row.names = c(NA, 6L), class = "data.frame") # 预处理生成绘图用数据集 plot_df <- df %>% # 将字符串格式的时间转换为标准日期格式 mutate(created_at = ymd(created_at)) %>% # 定义时间统计粒度:按年统计用year(),按月/季度替换为floor_date(created_at, unit = "month"/"quarter") mutate(time_unit = year(created_at)) %>% # 按时间单位分组,统计二分类变量取值为1的观测数 group_by(time_unit) %>% summarise(morality_1_count = sum(morality == 1, na.rm = TRUE))
绘图实现
以下代码生成带散点标记、可选趋势拟合线的折线图,匹配目标视觉效果:
ggplot(plot_df, aes(x = time_unit, y = morality_1_count)) + # 绘制主折线 geom_line(color = "#2c3e50", linewidth = 1) + # 绘制折线上的数值散点标记 geom_point(color = "#e74c3c", size = 3) + # 可选:添加loess平滑趋势线和置信区间,不需要可删除这一行 geom_smooth(method = "loess", se = TRUE, color = "#3498db", fill = "#bdc3c7", alpha = 0.3) + # 坐标轴刻度适配 scale_x_continuous(breaks = seq(min(plot_df$time_unit), max(plot_df$time_unit), by = 1)) + # 标签设置 labs( x = "时间", y = "二分类变量取值为1的观测计数", title = "道德类推文数量时间变化趋势" ) + # 主题风格调整 theme_bw() + theme( plot.title = element_text(hjust = 0.5, size = 14), axis.text = element_text(size = 11), axis.title = element_text(size = 12) )
自定义调整说明
- 若仅需要散点图,删除
geom_line()和geom_smooth()图层即可;若仅需要折线图,删除geom_point()图层即可。 - 若你使用的是0-1区间的概率值(如示例中的
morality_binary列)而非提前生成的0/1二分类值,可在预处理步骤中添加mutate(morality = ifelse(morality_binary >= 0.5, 1, 0)),按自定义阈值生成二分类结果后再统计。 - 调整时间粒度时,ggplot会自动识别日期类型的
time_unit字段,生成适配的时间轴刻度标签。
内容的提问来源于stack exchange,提问作者Quantizer
相关产品推荐
相关产品推荐

