You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用dplyr将Reddit帖子观测指标按id转换为面板格式

问题描述

我有一份基于Reddit帖子的数据集,每行对应一条帖子,包含以下字段:

  • username:发帖用户名
  • month_year:帖子发布年月
  • collective_action:二元变量,帖子含集体行动呼吁则标记为1,否则为0
  • directed_to_whom:帖子指向的对象(如不同政府部门)

需要将数据转换为以username为统计单位的面板格式,每个username对应唯一id,并计算两个指标:

  1. 每个用户发布的帖子中含集体行动的比例(collective_action为1的帖子数/总帖子数)
  2. 每个用户的帖子中指向Congress的比例(指向Congress的帖子数/总帖子数)

数据示例

dput(df[1:10,c(1,4,7,8,14)]) # 打印指定列的数据示例

输出:

structure(list(id = c(213L, 365L, 411L, 192L, 154L, 443L, 453L, 
462L, 213L, 213L), username = c("Cartman", "Cartman", 
"Cartman", "Kyle profleski", "Kyle profleski", 
"Cartman", "Kyle profleski Kyle profleski", "Kyle profleski", 
"Cartman", "Cartman"), collective_action = c(0, 1, 0, 0, 
1, 1, 0, 0, 0, 1), directed_to_whom = c("Congress", 
"Congress", "Senate", "Congress", "Congress", "president", 
"president", "Senate", "Congress", "Senate"), month_year = structure(c(2011.41666666667, 
2011.41666666667, 2011.41666666667, 2011.41666666667, 2011.41666666667, 
2011.41666666667, 2011.41666666667, 2011.41666666667, 2011.41666666667, 
2011.41666666667), class = "yearmon")), class = c("grouped_df", 
"tbl_df", "tbl", "data.frame"), row.names = c(NA, -10L), groups = structure(list(
    uusername = c("Cartman", "Cartman", 
"Cartman", "Kyle profleski", "Kyle profleski", 
"Cartman", "Kyle profleski Kyle profleski", "Kyle profleski", 
"Cartman", "Cartman"), .rows = structure(list(
        5L, 4L, c(1L, 9L, 10L), 2L, 3L, 6L, 7L, 8L), ptype = integer(0), class = c("vctrs_list_of", 
    "vctrs_vctr", "list"))), class = c("tbl_df", "tbl", "data.frame"
), row.names = c(NA, -8L), .drop = TRUE))

期望输出格式

id.    username        collec_action_prop   direc_to_Congress__prop
521.   Cartman            0.71                        0.57
501.   Mr.Mackey            0.4                       .32
499    Kyle profleski       0.3                       0.09
解决方案

使用dplyr包按用户名分组计算指标,并生成唯一用户ID:

library(dplyr)

# 取消原数据的分组状态(避免分组逻辑干扰)
df_ungrouped <- df %>% ungroup()

# 生成用户级面板数据
user_panel <- df_ungrouped %>%
  group_by(username) %>%
  summarise(
    # 计算集体行动帖子比例
    collec_action_prop = mean(collective_action, na.rm = TRUE),
    # 计算指向Congress的帖子比例
    direc_to_Congress__prop = mean(directed_to_whom == "Congress", na.rm = TRUE),
    .groups = "drop"
  ) %>%
  # 生成唯一用户ID
  mutate(id = row_number()) %>%
  # 调整列顺序匹配期望输出
  select(id, username, collec_action_prop, direc_to_Congress__prop)

# 查看结果
print(user_panel)

代码说明

  1. 取消分组:原数据为grouped_df格式,先取消分组确保计算逻辑正确
  2. 比例计算:
    • 二元变量collective_action的均值即为1的占比,直接用mean()计算
    • 通过directed_to_whom == "Congress"生成布尔值列,其均值就是指向Congress的帖子比例
  3. 生成唯一ID:用row_number()为每个用户分配唯一标识
  4. 列顺序调整:按期望输出的列排列结果

额外处理建议

如果数据中存在用户名格式不一致的情况(比如示例中的Kyle profleski Kyle profleski),可以先清理用户名:

library(stringr)
df_ungrouped <- df_ungrouped %>%
  mutate(username = str_squish(username)) # 去除多余空格和换行

内容的提问来源于stack exchange,提问作者nesta1990

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 05:27:25