R语言使用dplyr按年份州汇总灾害数据并生成唯一灾害类型列
你要的唯一灾害类型拼接字段可以直接在dplyr::summarize()中通过toString(unique(disastertype))实现,完整处理代码如下:
library(dplyr) # 构造原始数据集 df <- structure(list(year = c(1998, 1998, 1998, 1998, 1998), country = c("US", "US", "US", "US", "US"), state = c("Texas", "Texas", "California", "New York", "New York"), deaths = c(12, 5, 9, 10, 18), injured = c(3, 1, 3, 5, 9), disastertype = c("Hurricane", "Hurricane", "Wild fire", "Flood", "Epidemic")), class = "data.frame", row.names = c(NA, -5L)) # 按年份、州分组汇总 result <- df %>% group_by(year, state) %>% summarize( # 所有灾害类型拼接(对应预期结果里的disastertype字段) disastertype = toString(disastertype), # 唯一灾害类型拼接 u_disastertype = toString(unique(disastertype)), # 死亡人数总和 deaths = sum(deaths, na.rm = TRUE), # 受伤人数总和 injured = sum(injured, na.rm = TRUE), # 组内灾害事件总数 n_disasters = n(), # 组内不同灾害类型数量 n_distinct = n_distinct(disastertype) ) %>% ungroup()
toString()会自动把多个字符串用, 拼接,配合unique()就能得到每组去重后的灾害类型列表,运行上述代码得到的结果和你给出的预期结构完全一致。
内容的提问来源于stack exchange,提问作者flxflks
相关产品推荐
相关产品推荐

