基于行程起止目的列创建tourID的R语言实现求助
需求:为行程数据生成tourID字段
作为R语言新手,我需要创建tourID字段,规则是同一personID下,从home出发、最终回到home的连续行程归为同一个tourID。
输入数据示例
| personID | TripID | satartpurpose | endpurpose |
|---|---|---|---|
| 1 | 1 | home | shopping |
| 1 | 2 | shopping | home |
| 2 | 1 | home | work |
| 2 | 2 | work | home |
| 2 | 3 | home | visit friend |
| 2 | 4 | visit friend | home |
期望输出结果
| personID | TripID | startpurpose | endpurpose | tourID |
|---|---|---|---|---|
| 1 | 1 | home | shopping | 1 |
| 1 | 2 | shopping | home | 1 |
| 2 | 1 | home | work | 2 |
| 2 | 2 | work | home | 2 |
| 2 | 3 | home | visit friend | 3 |
| 2 | 4 | visit friend | home | 3 |
R代码实现
这里用dplyr包处理,语法直观适合新手:
# 先安装dplyr包(首次运行需执行) # install.packages("dplyr") library(dplyr) # 构造示例输入数据 trip_data <- tibble( personID = c(1,1,2,2,2,2), TripID = c(1,2,1,2,3,4), startpurpose = c("home", "shopping", "home", "work", "home", "visit friend"), endpurpose = c("shopping", "home", "work", "home", "visit friend", "home") ) # 生成tourID result <- trip_data %>% # 按用户和行程ID排序,确保行程顺序正确 arrange(personID, TripID) %>% # 按用户分组处理 group_by(personID) %>% # 标记行程起点是否为home,再累加生成tourID mutate( is_tour_start = startpurpose == "home", tourID = cumsum(is_tour_start) ) %>% # 移除临时标记列 select(-is_tour_start) %>% # 取消分组 ungroup() # 查看最终结果 print(result)
代码说明:
arrange(personID, TripID):确保每个用户的行程按顺序排列,这是生成正确tourID的基础group_by(personID):针对每个用户单独处理行程数据is_tour_start = startpurpose == "home":标记每个行程是否为一个tour的起点(从home出发)tourID = cumsum(is_tour_start):对起点标记进行累加,每遇到一个起点,tourID就递增1,这样同一个tour内的行程会共享同一个ID
内容的提问来源于stack exchange,提问作者Tina
相关产品推荐
相关产品推荐

