You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R代码处理Qualtrics随机选项顺序的调研数据?

处理Qualtrics随机选项顺序的调研数据

方法一:使用tidyverse包(推荐,代码更易读维护)

先加载tidyverse工具包,它整合了数据处理常用的dplyr和tidyr,能高效处理这类格式转换问题:

library(tidyverse)

# 你的示例数据
respondents <- c("Respondent 1", "Respondent 2", "Respondent 3")
tv_show_1 <- c("excellent", "good", "bad")
tv_show_2 <- c("good", "bad", "neutral")
tv_show_3 <- c("neutral", "good", "neutral")
tv_show_DO <- c("Friends|Seinfeld|Full House", "Seinfeld|Friends|Full House", "Seinfeld|Full House|Friends")

df <- data.frame(respondents, tv_show_1, tv_show_2, tv_show_3, tv_show_DO)

# 1. 拆分剧集顺序列,得到每个位置对应的剧集名称
df_shows <- df %>%
  separate(tv_show_DO, into = paste0("show_pos_", 1:3), sep = "\\|") %>%
  select(respondents, starts_with("show_pos_"))

# 2. 将评分列转为长格式,保留位置编号
df_ratings_long <- df %>%
  select(respondents, starts_with("tv_show_")) %>%
  pivot_longer(cols = starts_with("tv_show_"),
               names_to = "position",
               values_to = "rating",
               names_prefix = "tv_show_") %>%
  mutate(position = as.integer(position))

# 3. 将剧集位置列转为长格式,匹配位置与剧集名
df_shows_long <- df_shows %>%
  pivot_longer(cols = starts_with("show_pos_"),
               names_to = "position",
               values_to = "tv_show",
               names_prefix = "show_pos_") %>%
  mutate(position = as.integer(position))

# 4. 合并数据并转成宽格式,得到最终结果
final_df <- df_ratings_long %>%
  left_join(df_shows_long, by = c("respondents", "position")) %>%
  select(respondents, tv_show, rating) %>%
  pivot_wider(names_from = tv_show, values_from = rating)

print(final_df)

代码说明

  • separate:把tv_show_DO按竖线|拆分,生成对应位置的剧集列;
  • pivot_longer:把分散的评分列(tv_show_1/2/3)转成每行对应一个位置的评分,方便后续匹配;
  • left_join:通过受访者ID和位置编号,把评分和对应的剧集名关联;
  • pivot_wider:把长格式数据转回宽格式,最终每列对应一部剧的评分。

方法二:基础R方法(无需加载外部包)

如果不想加载工具包,用基础R也能实现:

# 你的示例数据
respondents <- c("Respondent 1", "Respondent 2", "Respondent 3")
tv_show_1 <- c("excellent", "good", "bad")
tv_show_2 <- c("good", "bad", "neutral")
tv_show_3 <- c("neutral", "good", "neutral")
tv_show_DO <- c("Friends|Seinfeld|Full House", "Seinfeld|Friends|Full House", "Seinfeld|Full House|Friends")

df <- data.frame(respondents, tv_show_1, tv_show_2, tv_show_3, tv_show_DO)

# 拆分每个受访者的剧集顺序为列表
show_lists <- strsplit(df$tv_show_DO, "\\|")

# 创建空结果数据框
final_df_base <- data.frame(respondents = df$respondents)

# 遍历每行,匹配评分与剧集
for(i in 1:nrow(df)){
  current_shows <- show_lists[[i]]
  current_ratings <- c(df$tv_show_1[i], df$tv_show_2[i], df$tv_show_3[i])
  final_df_base[i, current_shows] <- current_ratings
}

print(final_df_base)

代码说明

  • strsplit:把每个受访者的剧集字符串拆分成向量;
  • 循环遍历每行,把对应位置的评分赋值给以剧集名为列名的单元格,直接完成匹配。

注意事项

  • 如果剧集数量不是3部(比如5部),只需修改对应位置的数字(比如paste0("show_pos_", 1:5)、c(df$tv_show_1[i], ..., df$tv_show_5[i]));
  • 实际处理Qualtrics数据时,注意列名可能有前缀(比如Q1_tv_show_1),需要调整代码中列名匹配的规则(比如starts_with("Q1_tv_show_"));
  • 不管评分是字符型(excellent/good)还是数值型,处理逻辑完全一致。

内容的提问来源于stack exchange,提问作者user25563153

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 05:33:11