如何在R中将起止年份事件数据集转换为面板数据集?
解决方案:将事件数据集转换为district-year面板数据
先修正原数据笔误
原数据中第3行的end_year值20010应为2010,先修正数据:
# 修正后的原始数据 project_id <- c(1,2,3,4) district_id <- c(5,6,7,8) start_year <- c(2000, 2011, 2006, 2004) end_year <- c(2002, 2020, 2010, 2015) # 修正20010为2010 sector <- c("education", "infrastructure", "education", "infrastructure") df <- data.frame(project_id, district_id, start_year, end_year, sector)
方法1:使用tidyverse工具(适合常规规模数据)
利用dplyr分组+tidyr的unnest()函数,将每个项目的起止年份序列展开为单独行:
library(tidyverse) panel_df <- df %>% # 为每个项目生成从start_year到end_year的年份序列 mutate(year = map2(start_year, end_year, seq)) %>% # 将年份序列拆分为单独行 unnest(year) %>% # 移除不再需要的起止年份列 select(-start_year, -end_year) %>% # 按district_id和year排序(可选) arrange(district_id, year) # 查看结果 print(panel_df)
方法2:使用data.table(适合超大规模数据集)
data.table的展开操作效率远高于tidyverse,处理百万级数据更高效:
library(data.table) # 转换为data.table格式 dt <- as.data.table(df) # 生成面板数据 panel_dt <- dt[, .(year = seq(start_year, end_year)), by = .(project_id, district_id, sector)] %>% setorder(district_id, year) # 查看结果 print(panel_dt)
目标输出示例(以项目3为例)
project_id district_id year sector 1 3 7 2006 education 2 3 7 2007 education 3 3 7 2008 education 4 3 7 2009 education 5 3 7 2010 education
内容的提问来源于stack exchange,提问作者KC15
相关产品推荐
相关产品推荐

