You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CEO任职前后企业得分统计量计算:求高效实现方法

高效计算CEO任职前后3年的企业得分统计量

可以用dplyr结合tidyr的管道式操作,一次性完成分组、时间段标记、统计量计算,大幅减少手动步骤。以下是具体实现:

步骤1:构造示例数据

library(tidyverse)

# 构造示例数据框
df <- tibble(
  ceoID = c("48195", "48195", "48195", "48195", "48195", "48195", 
            "48195", "10203", "10203", "10203", "10203", "10203", "22879", 
            "22879", "22879"), 
  firmID = c("001004", "001004", "001004", "001004", 
             "001004", "001004", "001004", "001013", "001013", "001013", "001013", 
             "001013", "001013", "001013", "001013"), 
  ceoYEAR = c(2018, 2018, 2018, 2018, 2018, 2018, 2018, 2003, 2003, 2003, 2003, 2003, 2001, 
              2001, 2001), 
  scoreYEAR = c(2003, 2004, 2013, 2015, 2016, 2017, 
                2018, 1999, 2001, 2002, 2003, 2004, 1999, 2001, 2002), 
  scores = c(0, -2, 0, 0, 0, 0, 0, 0, -1, -1, 0, 1, 0, -1, -1)
)

步骤2:高效计算统计量

核心思路是:

  • 按ceoID和firmID分组(区分同一CEO的不同任职企业场景)
  • 标记得分年份所属的任职前3年(ceoYEAR-3 ≤ scoreYEAR ≤ ceoYEAR-1)、任职后3年(ceoYEAR+1 ≤ scoreYEAR ≤ ceoYEAR+3)
  • 对每个分组和时间段计算mean、median、sum,缺失数据自动返回NA
  • 转成宽格式方便对比前后数据
result <- df %>%
  # 按CEO、企业、任职年份分组
  group_by(ceoID, firmID, ceoYEAR) %>%
  # 标记目标时间段,其余年份排除
  mutate(period = case_when(
    scoreYEAR %in% (ceoYEAR - 3):(ceoYEAR - 1) ~ "pre_3y",
    scoreYEAR %in% (ceoYEAR + 1):(ceoYEAR + 3) ~ "post_3y"
  )) %>%
  filter(!is.na(period)) %>%
  # 计算各时间段统计量
  summarise(
    mean_score = mean(scores, na.rm = TRUE),
    median_score = median(scores, na.rm = TRUE),
    sum_score = sum(scores, na.rm = TRUE),
    .groups = "drop"
  ) %>%
  # 转宽格式整合结果
  pivot_wider(
    names_from = period,
    values_from = c(mean_score, median_score, sum_score),
    values_fill = list(mean_score = NA, median_score = NA, sum_score = NA)
  )

print(result)

输出结果说明

运行后会得到每个CEO-企业组合的任职前后3年得分统计量:

  • 无对应时间段数据时,统计量自动填充NA
  • 比如示例中CEO 48195在企业001004的任职年份为2018,仅2015-2017属于前3年,后3年无数据,因此后3年统计量均为NA
  • CEO 10203在企业001013的任职年份为2003,前3年仅2001、2002有数据,后3年仅2004有数据,对应统计量会基于现有数据计算

这种管道式操作无需手动拆分分组,代码简洁可复用,适配大规模数据处理场景。

内容的提问来源于stack exchange,提问作者LearningR

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 00:11:05