You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中创建分立场的推文小时级平均情感得分时间线图

罗诉韦德案推文情感时间线可视化解决方案

1. 数据预处理与聚合

首先需要将时间戳转换为标准格式,然后按小时和立场分组计算统计量,同时过滤掉样本量过小的小时组,避免结果受极端波动影响:

# 加载依赖包
library(dplyr)
library(lubridate)
library(ggplot2)
library(plotly)

# 转换时间戳为POSIXct格式(确保时间列可被正确识别)
tweets$POSIX <- ymd_hms(tweets$POSIX)

# 按小时分组计算平均得分与推文数量,并筛选目标时间范围
hourly_data <- tweets %>%
  # 筛选2022年6月22日00:00至7月3日23:59的推文(罗诉韦德案2022年6月24日宣布)
  filter(between(POSIX, ymd_hms("2022-06-22 00:00:00"), ymd_hms("2022-07-03 23:59:59"))) %>%
  # 将时间截断到小时级别
  mutate(hourly_time = floor_date(POSIX, unit = "hour")) %>%
  # 按小时+立场分组
  group_by(hourly_time, stance) %>%
  # 计算平均情感得分、该小时推文数量
  summarise(
    avg_score = mean(score, na.rm = TRUE),
    tweet_count = n(),
    .groups = "drop"
  ) %>%
  # 过滤每小时推文数少于10的组(阈值可根据需求调整)
  filter(tweet_count >= 10)

2. 静态折线图(ggplot2)

生成包含两条立场线的时间轴图,同时用点的大小直观展示每小时推文数量:

ggplot(hourly_data, aes(x = hourly_time, y = avg_score, color = stance)) +
  geom_line(linewidth = 1) +
  geom_point(aes(size = tweet_count), alpha = 0.7) +
  # 自定义立场颜色
  scale_color_manual(values = c("choice" = "#1f77b4", "life" = "#ff7f0e")) +
  # 图表标签与标题
  labs(
    x = "时间(小时)",
    y = "平均情感得分",
    color = "立场",
    size = "每小时推文数",
    title = "罗诉韦德案后不同立场推文每小时平均情感得分"
  ) +
  # 优化主题样式
  theme_minimal() +
  theme(
    plot.title = element_text(hjust = 0.5, size = 14, face = "bold"),
    axis.text.x = element_text(angle = 45, hjust = 1)
  )

3. 交互式折线图(plotly)

如果需要hover查看详细信息(比如具体时间、推文数量),可将ggplot对象转换为plotly交互式图表:

# 转换为交互式图表
ggplotly() %>%
  layout(
    title = list(text = "罗诉韦德案后不同立场推文每小时平均情感得分", x = 0.5),
    xaxis = list(title = "时间(小时)", tickangle = 45),
    yaxis = list(title = "平均情感得分")
  )

进阶:控制推文数量影响的补充方案

如果需要更严谨地控制样本量影响,可以添加95%置信区间,直观展示得分的可靠性:

# 计算带置信区间的统计量
hourly_data_ci <- tweets %>%
  filter(between(POSIX, ymd_hms("2022-06-22 00:00:00"), ymd_hms("2022-07-03 23:59:59"))) %>%
  mutate(hourly_time = floor_date(POSIX, unit = "hour")) %>%
  group_by(hourly_time, stance) %>%
  summarise(
    avg_score = mean(score, na.rm = TRUE),
    se_score = sd(score, na.rm = TRUE)/sqrt(n()),
    tweet_count = n(),
    .groups = "drop"
  ) %>%
  filter(tweet_count >= 10) %>%
  mutate(
    lower_ci = avg_score - 1.96*se_score,
    upper_ci = avg_score + 1.96*se_score
  )

# 绘制带置信区间的折线图
ggplot(hourly_data_ci, aes(x = hourly_time, y = avg_score, color = stance)) +
  geom_line(linewidth = 1) +
  # 添加置信区间填充
  geom_ribbon(aes(ymin = lower_ci, ymax = upper_ci, fill = stance), alpha = 0.2) +
  scale_color_manual(values = c("choice" = "#1f77b4", "life" = "#ff7f0e")) +
  scale_fill_manual(values = c("choice" = "#1f77b4", "life" = "#ff7f0e")) +
  labs(
    x = "时间(小时)",
    y = "平均情感得分",
    color = "立场",
    fill = "立场",
    title = "不同立场推文每小时平均情感得分(带95%置信区间)"
  ) +
  theme_minimal() +
  theme(plot.title = element_text(hjust = 0.5))

内容的提问来源于stack exchange,提问作者Brenna

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 18:10:51