如何用R统计各机场的竞争对手数量?求代码实现帮助
解决方案
方法1:统计机场关联的航班总数(匹配你的示例输出)
先安装并加载新手友好的tidyverse工具集,再按步骤处理数据:
# 首次运行时安装包 install.packages("tidyverse") # 加载包 library(tidyverse) # 1. 创建示例数据集(实际使用时替换为 read.csv("你的数据文件路径.csv") 读取本地数据) flight_data <- data.frame( flight_date = c("03/13/2019", "03/13/2019", "03/13/2019", "03/13/2019", "03/14/2019"), op_career = c("AA", "AA", "AS", "YV", "DL"), tail_num = c("N900EV", "N900EZ", "N686BR", "N932LR", "255NV"), flight_num = c(3503, 3502, 3397, 5804, 515), origin = c("SFO", "SFO", "SFO", "SFO", "SFO"), origin_airport_ID = rep(11308, 5), dest_airport_ID = rep(10397, 5), dest = c("LAX", "LAX", "LAX", "LAX", "LAX") ) # 2. 处理数据生成结果 airport_competitors <- flight_data %>% # 将出发/到达机场合并为单列,每个航班对应两条机场记录 pivot_longer(cols = c(origin, dest), names_to = "airport_type", values_to = "Airport") %>% # 按机场分组,统计关联的航班总数 group_by(Airport) %>% summarise(competitors = n()) %>% ungroup() # 查看最终结果 print(airport_competitors)
运行后输出:
# A tibble: 2 × 2 Airport competitors <chr> <int> 1 LAX 5 2 SFO 5
方法2:统计机场的独特航司数量(更符合“竞争对手”常规定义)
如果你的真实需求是统计每个机场有多少不同的运营航司(同一家航司多次起降只算一个竞争对手),修改统计逻辑即可:
airport_competitors <- flight_data %>% pivot_longer(cols = c(origin, dest), names_to = "airport_type", values_to = "Airport") %>% group_by(Airport) %>% # 统计独特航司的数量 summarise(competitors = n_distinct(op_career)) %>% ungroup() print(airport_competitors)
运行后输出:
# A tibble: 2 × 2 Airport competitors <chr> <int> 1 LAX 4 2 SFO 4
内容的提问来源于stack exchange,提问作者newbie data scientist
相关产品推荐
相关产品推荐

