You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言中基于转移日期数据构建面板数据集?

在R中将Origin-Destination时间数据转换为ID-Time位置面板数据集

当然可以!用R的tidyverse工具链(dplyr + tidyr)就能轻松实现这个需求,下面是针对你的数据集的具体解决方案:

步骤1:加载所需包并导入原始数据

首先加载处理数据的核心包,然后输入你的原始数据集:

library(tidyverse)

# 你的原始数据集
raw_df <- tibble(
  ID = c(1, 2, 2, 3, 3),
  origin = c("a", "b", "a", "c", "b"),
  destination = c("b", "a", "c", "b", "c"),
  time = c(2, 1, 4, 1, 3)
)

步骤2:处理位置时间区间

我们需要先为每个ID确定不同位置的生效时间区间:

  1. 按ID分组并按时间排序,确保行程顺序正确
  2. 计算每个行程的下一个时间节点,用来确定当前位置的结束时间
  3. 生成初始位置的时间区间(如果第一个行程不是从时间1开始)
  4. 生成每个目的地的时间区间
# 全局最大时间(你也可以按每个ID的最大时间调整)
max_global_time <- 4

# 处理行程的时间节点
processed_df <- raw_df %>%
  group_by(ID) %>%
  arrange(time) %>% # 按时间排序每个ID的行程
  mutate(
    next_time = lead(time, default = max_global_time + 1), # 获取下一个行程的时间,默认值设为全局最大时间+1
    max_time = max_global_time
  ) %>%
  ungroup()

# 生成初始位置的时间区间(仅当第一个行程的时间大于1时)
initial_locations <- processed_df %>%
  group_by(ID) %>%
  slice(1) %>% # 取每个ID的第一条行程
  filter(time > 1) %>% # 只保留那些不是从时间1开始移动的ID
  transmute(
    ID,
    location = origin,
    start_time = 1,
    end_time = time - 1
  ) %>%
  ungroup()

# 生成每个目的地的时间区间
dest_locations <- processed_df %>%
  transmute(
    ID,
    location = destination,
    start_time = time,
    end_time = ifelse(next_time - 1 >= time, next_time - 1, max_time)
  )

步骤3:展开为面板数据集

将时间区间展开为每个时间点的单独行,最后整理成你需要的格式:

panel_df <- bind_rows(initial_locations, dest_locations) %>%
  group_by(ID, location, start_time, end_time) %>%
  mutate(time = list(start_time:end_time)) %>% # 将时间区间转为序列列表
  unnest(time) %>% # 展开列表为单独行
  select(ID, location, time) %>% # 保留需要的列
  arrange(ID, time) # 按ID和时间排序

# 查看结果
print(panel_df)

运行后你会得到完全符合需求的面板数据集:

# A tibble: 12 × 3
# Groups:   ID [3]
      ID location  time
   <dbl> <chr>    <int>
 1     1 a            1
 2     1 b            2
 3     1 b            3
 4     1 b            4
 5     2 a            1
 6     2 a            2
 7     2 a            3
 8     2 c            4
 9     3 b            1
10     3 b            2
11     3 c            3
12     3 c            4

逻辑说明

这个方案的核心是先明确每个位置的生效时间范围:

  • 如果一个ID的第一个行程是在时间t从origin移动到destination,那么origin的生效时间是1到t-1(如果t>1)
  • 每个destination的生效时间是从当前行程的时间t,到下一个行程的时间-1;如果是最后一个行程,则生效到全局最大时间

内容的提问来源于stack exchange,提问作者Kenji

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:51:13