You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言:按Run将流转位置数据转换为宽格式的问题

R中将流转数据从长格式转换为宽格式的正确方法

问题背景

现有记录物品流转起始与终点位置的长格式数据:

run = c(1, 2, 3, 3, 4, 5, 5, 5, 6, 7, 7, 7, 8, 9, 10, 10, 11)
start_site = c("A", "C", "A", "B", "A", "B", "C", "A", "B", "C", "B", "A", "A", "A", "A", "B", "C")
end_site = c("B", "B", "B", "C", "C", "C", "A", "C", "A", "B", "A", "C", "B", "C", "B", "C", "B")
df = data.frame(run, start_site, end_site)

期望转换为宽格式,按run分组,每个流转阶段对应一列,结果如下:

# 期望输出
      run  first_location second_location third_location fourth_location
 [1,] "1"  "A"            "B"             NA             NA              
 [2,] "2"  "C"            "B"             NA             NA              
 [3,] "3"  "A"            "B"             "C"            NA              
 [4,] "4"  "A"            "C"             NA             NA              
 [5,] "5"  "B"            "C"             "A"            "C"             
 [6,] "6"  "B"            "A"             NA             NA              
 [7,] "7"  "C"            "B"             "A"            "C"             
 [8,] "8"  "A"            "B"             NA             NA              
 [9,] "9"  "A"            "C"             NA             NA              
[10,] "10" "A"            "B"             "C"            NA              
[11,] "11" "C"            "B"             NA             NA     

错误原因分析

你之前尝试的代码错误在于使用*melt*将start_site和end_site都转换为长格式,导致每个流转步骤被拆分为两行(起始位置一行、结束位置一行)。在后续分组添加行号时,这些拆分后的行都会被计数,最终生成了重复的位置列(比如run=3中出现了两次B),不符合流转链的逻辑。

正确解决方案

我们需要先按run分组,提取每个组的完整流转位置链,再将其拆分为宽格式列。以下是具体代码:

library(dplyr)
library(tidyr)

# 生成目标宽格式数据
df_wide <- df %>%
  group_by(run) %>%
  # 提取每个run的完整流转位置链:第一个起始位置 + 所有结束位置
  summarise(locations = c(first(start_site), end_site)) %>%
  # 将位置链拆分为多列
  unnest_wider(locations, names_sep = "_") %>%
  # 重命名列名为指定格式
  rename_with(
    ~ case_match(
      .x,
      "locations_1" ~ "first_location",
      "locations_2" ~ "second_location",
      "locations_3" ~ "third_location",
      "locations_4" ~ "fourth_location"
    )
  )

# 查看结果
df_wide

运行后将得到与期望完全一致的输出:每个run的流转位置按顺序排列为first_location到fourth_location,无多余重复列。

内容的提问来源于stack exchange,提问作者NM_

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 21:35:59