You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用dplyr按子组计算特定情感-政客指向的帖子数量与占比

问题

我有一个由Reddit帖子组成的数据集,每行包含帖子内容、日期、基于ML预测的情感(mood字段)以及帖子指向的特定政客(directed_to_whom字段)。

数据示例

post        date            mood         directed_to_whom 
Cartman   2012-09-03.       negative           Romney
Cartman  2012-09-06.        negative           Romney
Cartman  2012-09-13.        negative           Romney 
Cartman    2012-09-15.      neutral           Bush
Mackey   2012-09-03.       negative           Bush
Mackey  2012-09-08.        neutral            Bush
Mackey  2012-09-13.        neutral            post
Garrison   2012-09-03.      negative          Romney
Garrison  2012-09-04.       negative          pre
Garrison  2012-09-04.       negative          pre
Garrison  2012-09-05.     negative           Obama

已有代码(月度情感占比图表)

我已经用ggplot生成了展示不同时间段内负面、中性、正面帖子月度占比的图表,代码如下:

ggplot(both_group, aes(x = as.Date(month_year), fill = sentiment ,y = sentiment_percentage)) +
    geom_bar(stat = "identity", position=position_dodge()) + 
    scale_x_date(date_breaks = "1 month", date_labels = "%b %Y") + 
    xlab("Sentiment") + 
    theme(plot.title = element_text(size = 18, face = "bold")) +
    scale_y_continuous (name = "Sentiment share") +
    theme_classic()+
    theme(plot.title = element_text(size = 5, face = "bold"),
          axis.text.x = element_text(angle = 90, vjust = 0.5))

需求

现在我想创建一个变量,用来统计指向奥巴马的负面帖子或者指向罗姆尼的正面帖子的数量/占比,不确定是否可行,想请教实现方法。


解决方案

当然可以实现这个需求,你只需要通过条件筛选标记出符合要求的帖子,再基于标记结果统计数量或占比即可。下面是具体的实现步骤:

1. 创建目标帖子标记变量

先用dplyr的mutate函数新增一个布尔型变量,标记符合条件的帖子:

library(dplyr)

# 假设你的数据集名为df
df <- df %>%
  mutate(target_post = case_when(
    mood == "negative" & directed_to_whom == "Obama" ~ TRUE,
    mood == "positive" & directed_to_whom == "Romney" ~ TRUE,
    TRUE ~ FALSE
  ))

2. 统计目标帖子总数

直接对标记变量求和,就能得到符合条件的帖子总数:

target_total_count <- sum(df$target_post, na.rm = TRUE)

如果需要按月度分组统计数量,可以结合lubridate处理日期后聚合:

library(lubridate)

monthly_target_count <- df %>%
  # 处理日期格式(原数据日期末尾有个点,需指定格式)
  mutate(month_year = floor_date(as.Date(date, format = "%Y-%m-%d."), "month")) %>%
  group_by(month_year) %>%
  summarise(target_count = sum(target_post, na.rm = TRUE))

3. 统计目标帖子占比

计算目标帖子在总帖子中的占比:

target_overall_ratio <- target_total_count / nrow(df)

如果需要按月度统计占比(目标帖子占当月总帖子的比例):

monthly_target_ratio <- df %>%
  mutate(month_year = floor_date(as.Date(date, format = "%Y-%m-%d."), "month")) %>%
  group_by(month_year) %>%
  summarise(
    total_monthly_posts = n(),
    target_count = sum(target_post, na.rm = TRUE),
    target_ratio = target_count / total_monthly_posts
  )

4. 可视化扩展(可选)

如果要把月度占比结果可视化,比如生成折线图,可以用ggplot:

ggplot(monthly_target_ratio, aes(x = month_year, y = target_ratio)) +
  geom_line(color = "darkblue", linewidth = 1) +
  scale_x_date(date_breaks = "1 month", date_labels = "%b %Y") +
  labs(x = "Month", y = "Target Post Ratio", title = "Monthly Ratio of Target Posts") +
  theme_classic() +
  theme(axis.text.x = element_text(angle = 90, vjust = 0.5))

内容的提问来源于stack exchange,提问作者nesta1990

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 10:45:36