如何用dplyr/tidyr/base R统计两字符串列组合频率并保留关联列
统计共享单车起止站点组合频率并保留关联字段
之前使用count()时,它默认只返回分组字段和计数结果,所以会丢失经纬度这类关联字段。改用分组后自定义聚合的方式,就能灵活保留需要的字段。
示例数据
df <- data.frame(start_station = c('Apple', 'Bungalow', 'Carrot', 'Apple', 'Apple', 'Bungalow'), end_station = c('Bungalow', 'Apple', 'Carrot', 'Bungalow', 'Bungalow', 'Apple'), start_lat = c(12.3456, 23.4567, 34.5678, 12.3456, 12.3456, 23.4567), start_lng = c(9.8765, 98.7654, 87.6543, 9.8765, 9.8765, 98.7654) )
解决方案代码
library(dplyr) result <- df %>% group_by(start_station, end_station) %>% summarize( ride_count = n(), start_lat = first(start_lat), # 同一起始站点经纬度固定,取组内任意值即可 start_lng = first(start_lng), .groups = "drop" ) %>% arrange(desc(ride_count)) print(result)
代码说明
group_by(start_station, end_station):按起止站点的组合作为分组依据,这是统计骑行次数的核心。summarize():执行聚合操作,n()计算每组的行数即该站点组合的骑行次数;first(start_lat)提取组内的经纬度值——因为同一个起始站点的经纬度是唯一的,所以用first()、last()或unique()都能得到正确结果。.groups = "drop":取消分组状态,让结果回到普通数据框格式,方便后续操作。arrange(desc(ride_count)):按骑行次数从高到低排序,符合你要的降序排列需求。
输出结果
运行代码后会得到符合预期的结果:
start_station end_station ride_count start_lat start_lng <chr> <chr> <int> <dbl> <dbl> 1 Apple Bungalow 3 12.35 9.88 2 Bungalow Apple 2 23.46 98.77 3 Carrot Carrot 1 34.57 87.65
内容的提问来源于stack exchange,提问作者Pat Tarantino
相关产品推荐
相关产品推荐

