基于R(tidyverse)计算有序因子问卷数据的多维度得分
用tidyverse计算Likert问卷的两种计分规则下的总分与子量表得分
问题背景
现有一份11题的问卷数据,每题是4级有序因子(对应Less than usual/No more than usual/More than usual/Much more than usual),需要按两种规则计算得分:
- Likert计分:因子水平依次赋值0、1、2、3
- Binary计分:因子水平依次赋值0、0、1、1
需计算6个指标:total_likert(全题Likert总分)、total_binary(全题Binary总分)、total_ss1_likert(子量表1:q1-q7的Likert总分)、total_ss1_binary(子量表1的Binary总分)、total_ss2_likert(子量表2:q8-q11的Likert总分)、total_ss2_binary(子量表2的Binary总分)
解决方案
核心思路是先将有序因子转换为对应计分的数值,再按子量表范围计算行总和。利用tidyverse的across和rowSums可高效完成:
library(tidyverse) # 复现数据 df <- tibble(id = c(1, 2, 3, 4, 5), q1 = c(3, 4, 2, 3, 3), q2 = c(4, 4, 2, 3, 2), q3 = c(3, 3, 2, 2, 3), q4 = c(2, 2, 3, 2, 1), q5 = c(3, 3, 3, 3, 3), q6 = c(4, 3, 2, 2, 2), q7 = c(1, 2, 2, 2, 2), q8 = c(3, 3, 3, 2, 1), q9 = c(3, 4, 4, 2, 1), q10 = c(2, 4, 3, 2, 1), q11 = c(2, 3, 2, 2, 1)) %>% mutate(across(q1:q11, ~factor(.x, levels = c(1, 2, 3, 4), labels = c("Less than usual", "No more than usual", "More than usual", "Much more than usual"), ordered = TRUE))) # 计算各得分指标 df_scores <- df %>% mutate( # 子量表1(q1-q7)的两种计分总分 total_ss1_likert = rowSums(across(q1:q7, ~as.numeric(.x) - 1)), total_ss1_binary = rowSums(across(q1:q7, ~ifelse(as.numeric(.x) >= 3, 1, 0))), # 子量表2(q8-q11)的两种计分总分 total_ss2_likert = rowSums(across(q8:q11, ~as.numeric(.x) - 1)), total_ss2_binary = rowSums(across(q8:q11, ~ifelse(as.numeric(.x) >= 3, 1, 0))), # 全量表总分 total_likert = total_ss1_likert + total_ss2_likert, total_binary = total_ss1_binary + total_ss2_binary ) # 查看结果 select(df_scores, id, starts_with("total_"))
结果说明
运行上述代码后,df_scores会包含原始数据和6个计算好的得分指标:
as.numeric(.x) - 1将有序因子的水平(1-4)转换为Likert计分的0-3ifelse(as.numeric(.x) >= 3, 1, 0)将前两个水平(1、2)赋值为0,后两个(3、4)赋值为1,实现Binary计分rowSums(across(...))按行计算指定列的数值之和,得到子量表总分
示例输出:
# A tibble: 5 × 7 id total_ss1_likert total_ss1_binary total_ss2_likert total_ss2_binary total_likert total_binary <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> 1 1 11 4 6 2 17 6 2 2 13 4 9 3 22 7 3 3 8 2 5 2 13 4 4 4 7 2 3 1 10 3 5 5 6 2 1 0 7 2
内容的提问来源于stack exchange,提问作者Braden
相关产品推荐
相关产品推荐

