You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用dplyr/tidyr/base R统计两字符串列组合频率并保留关联列

统计共享单车起止站点组合频率并保留关联字段

之前使用count()时,它默认只返回分组字段和计数结果,所以会丢失经纬度这类关联字段。改用分组后自定义聚合的方式,就能灵活保留需要的字段。

示例数据

df <- data.frame(start_station = c('Apple', 'Bungalow', 'Carrot', 'Apple', 'Apple', 'Bungalow'),
                 end_station = c('Bungalow', 'Apple', 'Carrot', 'Bungalow', 'Bungalow', 'Apple'),
                 start_lat = c(12.3456, 23.4567, 34.5678, 12.3456, 12.3456, 23.4567),
                 start_lng = c(9.8765, 98.7654, 87.6543, 9.8765, 9.8765, 98.7654)
)

解决方案代码

library(dplyr)

result <- df %>%
  group_by(start_station, end_station) %>%
  summarize(
    ride_count = n(),
    start_lat = first(start_lat),  # 同一起始站点经纬度固定,取组内任意值即可
    start_lng = first(start_lng),
    .groups = "drop"
  ) %>%
  arrange(desc(ride_count))

print(result)

代码说明

  • group_by(start_station, end_station):按起止站点的组合作为分组依据,这是统计骑行次数的核心。
  • summarize():执行聚合操作,n()计算每组的行数即该站点组合的骑行次数;first(start_lat)提取组内的经纬度值——因为同一个起始站点的经纬度是唯一的,所以用first()、last()或unique()都能得到正确结果。
  • .groups = "drop":取消分组状态,让结果回到普通数据框格式,方便后续操作。
  • arrange(desc(ride_count)):按骑行次数从高到低排序,符合你要的降序排列需求。

输出结果

运行代码后会得到符合预期的结果:

start_station end_station ride_count start_lat start_lng
  <chr>         <chr>            <int>     <dbl>     <dbl>
1 Apple         Bungalow             3     12.35      9.88
2 Bungalow      Apple                2     23.46     98.77
3 Carrot        Carrot               1     34.57     87.65

内容的提问来源于stack exchange,提问作者Pat Tarantino

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 04:01:13