You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用pivot_longer将情绪调研宽格式数据转换为指定长格式

背景

我们要求每位参与者识别多种情绪,随后采集每种情绪的对应数据,数据中包含参与者识别的第1种、第2种情绪等列,以及每种情绪对应的各随访问题的独立列,宽格式数据示例如下:

rows <- 1:4
cols <- c("PID", "Stage", "Emo1_", "Emo2_", 
          "Emo1_Intense", "Emo2_Intense", 
          "Emo1_Desc", "Emo2_Desc", "Keyword")
df  <- data.frame(matrix(NA, 
                         nrow = length(rows), 
                         ncol = length(cols), 
                         dimnames = list(rows, cols)))
df$PID <- c("A-001", "A-002", "A-003", "A-004")
df$Stage <- c("Beginning", "End", "Middle", "Middle")
df$Emo1_ <- c("Fear", "Sadness", "Happy", "Anger")
df$Emo2_ <- c("Content", "Depressed", "Lost", "Sad")
df$Emo1_Intense <- 5:8
df$Emo2_Intense <- 1:4
df$Emo1_Desc <- c("E", "F", "G", "H")
df$Emo2_Desc <- c("A", "B", "C", "D")
df$Keyword <- c("Bus", "Ceiling", "Chainsaw", "Floor")

原数据预览:

#    PID     Stage   Emo1_     Emo2_ Emo1_Intense Emo2_Intense Emo1_Desc Emo2_Desc  Keyword
#1 A-001 Beginning    Fear   Content            5            1         E         A      Bus
#2 A-002       End Sadness Depressed            6            2         F         B  Ceiling
#3 A-003    Middle   Happy      Lost            7            3         G         C Chainsaw
#4 A-004    Middle   Anger       Sad            8            4         H         D    Floor

问题

需要将宽格式数据转换为长格式,转换后单列对应:

  • 情绪命名的顺位
  • 具体情绪名称
  • 对应情绪的各随访问题结果
    目标格式示例如下:
rows <- 1:8
cols <- c("PID", "Stage", "Number", "Emo", "Intense", "Desc", "Keyword")
df  <- data.frame(matrix(NA, 
                         nrow = length(rows), 
                         ncol = length(cols), 
                         dimnames = list(rows, cols)))
df$PID <- sort(rep(c("A-001", "A-002", "A-003", "A-004"), 2))
df$Stage <- sort(rep(c("Beginning", "Middle", "Middle", "End"), 2))
df$Number <- rep(1:2, 4)
df$Emo <- c("Fear", "Content", "Sadness", "Depressed", "Happy", "Lost", "Anger", "Sad")
df$Intense <- c(5,1,6,2,7,3,4,8)
df$Desc <- c("E", "A", "F", "B", "G", "C", "H", "D")
df$Keyword <- rep(c("Bus", "Ceiling", "Chainsaw", "Floor"),2)

目标格式预览:

#    PID     Stage Number       Emo Intense Desc  Keyword
#1 A-001 Beginning      1      Fear       5    E      Bus
#2 A-001 Beginning      2   Content       1    A  Ceiling
#3 A-002       End      1   Sadness       6    F Chainsaw
#4 A-002       End      2 Depressed       2    B    Floor
#5 A-003    Middle      1     Happy       7    G      Bus
#6 A-003    Middle      2      Lost       3    C  Ceiling
#7 A-004    Middle      1     Anger       4    H Chainsaw
#8 A-004    Middle      2       Sad       8    D    Floor

注:上述目标示例中的Keyword列为笔误,实际转换后会保留原宽表中每个PID对应的唯一Keyword值。

解决方案

使用tidyr的pivot_longer函数,通过正则匹配提取列名中的情绪顺位和属性字段即可完成转换:

library(tidyverse)

df_long <- df %>%
  # 选择所有以Emo开头的待转换列
  pivot_longer(
    cols = starts_with("Emo"),
    # 正则分组:第一组匹配情绪顺位数字,第二组匹配情绪属性
    names_pattern = "Emo(\\d+)_*(.*)",
    # 第一组值存入Number列,第二组作为新列名,对应值填充到对应列
    names_to = c("Number", ".value")
  ) %>%
  # 原EmoX_列对应属性为空,重命名为情绪名称列Emo
  rename(Emo = "")

如果提前调整原数据列名(将Emo1_改为Emo1_Name、Emo2_改为Emo2_Name),转换逻辑会更清晰,无需额外重命名:

# 调整列名后转换代码
colnames(df) <- str_replace(colnames(df), "Emo(\\d+)_$", "Emo\\1_Name")

df_long <- df %>%
  pivot_longer(
    cols = starts_with("Emo"),
    names_pattern = "Emo(\\d+)_(.*)",
    names_to = c("Number", ".value")
  ) %>%
  rename(Emo = Name)

转换后输出符合要求,支持任意数量的情绪顺位扩展,无需手动调整逻辑。

内容的提问来源于stack exchange,提问作者William Mitchell

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 00:06:04