在R中对大规模数据集实现动态分组基础计算的技术问询
数据处理需求
数据集背景
我有一个名为df的数据集,包含数万条观测记录,分类变量有100余种。数据集记录了不同个体(id)在指定年份、以指定价格将不同类型患者送往不同地点的信息,示例数据如下:
year <- c(2010, 2010, 2010, 2010, 2011, 2011, 2011, 2010, 2011) id <- c("A", "A" , "A" , "A" , "A" , "A" , "A", "B", "B") type <- c("kid", "kid", "adult", "kid", "kid", "dog", "cat", "kid", "kid") place <- c("hosp", "hosp", "house", "hosp", "hosp", "hosp", "house", "hosp", "hosp") price <- c(2, 3, 6, 5, 1, 2, 3, 4, 5) df <- data.frame(year, id, type, place, price)
具体处理需求
需按id-year分组对df进行基础统计计算,具体要求:
- 按患者类型生成经验变量:取值为该
id拥有该类型的年份数量; - 按地点生成经验变量:取值为该
id拥有该地点的年份数量; - 计算每个
id在当年的平均就诊价格; - 生成标记变量:该
id是否在次年(t+1)出现,取值为0(否)或1(是)。
期望输出
最终期望得到类似df_new的数据集:
year <- c("2010", "2011", "2010", "2011") id <- c("A", "A", "B", "B") exp_type_kid <- c(1, 2, 1, 2) exp_type_adult <- c(1, 1, 0, 0) exp_type_dog <- c(0, 1, 0, 0) exp_type_cat <- c(0, 1, 0, 0) exp_place_hosp <- c(1, 2, 1, 2) exp_place_house <- c(1, 2, 0, 0) avg_price <- c(4, 2, 4, 5) id_repeat_next_year <- c(1, 0, 1, 0) df_new <- data.frame(year, id, exp_type_kid, exp_type_adult, exp_type_dog, exp_type_cat, exp_place_hosp, exp_place_house, avg_price, id_repeat_next_year)
补充说明
数据集可能包含更多不连续年份,示例如下:
year <- c(2010, 2010, 2010, 2010, 2011, 2011, 2011, 2009, 2010, 2015, 2017) id <- c("A", "A" , "A" , "A" , "A" , "A" , "A", "B", "B", "B", "B") type <- c("kid", "kid", "adult", "kid", "kid", "dog", "cat", "kid", "kid", "kid", "kid") place <- c("hosp", "hosp", "house", "hosp", "hosp", "hosp", "house", "hosp", "hosp", "hosp", "hosp") price <- c(2, 3, 6, 5, 1, 2, 3, 4, 4, 4, 4) df <- data.frame(year, id, type, place, price)
内容的提问来源于stack exchange,提问作者vog
相关产品推荐
相关产品推荐

