请求简化R语言中均值、SD计算及Ipsatization代码
个体内标准化(Ipsatization)代码简化需求
我正在完成毕业论文,需要对三个测量工具(压力、调节、SCS)的项目得分做个体内标准化(Ipsatization):先计算每个个体在同一工具下所有项目的均值和标准差,再用各项目得分减去均值后除以标准差完成转换。
目前用RStudio实现了功能,但代码重复度很高——每个工具都要重复计算均值、标准差,再循环转换。希望简化这段重复代码。
原实现代码
library(dplyr) #create a new dataset for ipsatized results data5 <- data4 #create list of item names for three instruments stress <- c("BES_1", "BES_2", "BES_3", "BES_4", "BES_5") reg <- c("Neg_ctrl1","Pos_ctrl1","Neg_use1","Pos_use1") scs <- c("Ind_1","Ind_2","Inter_1","Inter_2","Ind_3") #calculate overall mean and sd within each intrument(this I repeated 3 times for different instruments) row_m <- rowMeans(data4[, stress], na.rm = TRUE) row_sds <- apply(data4[, stress], 1, sd, na.rm = TRUE) row_m <- rowMeans(data4[, reg], na.rm = TRUE) row_sds <- apply(data4[, reg], 1, sd, na.rm = TRUE) row_m <- rowMeans(data4[, scs], na.rm = TRUE) row_sds <- apply(data4[, scs], 1, sd, na.rm = TRUE) #proceed to the final calculation for (col in stress) { data5[, col] <- (data5[, col] - row_m) / row_sds } for (col in reg) { data5[, col] <- (data5[, col] - row_m) / row_sds } for (col in scs) { data5[, col] <- (data5[, col] - row_m) / row_sds } #for the ease of presenting here, I slice data5 to only first 10 row. dput(data5)
简化后的代码方案
把工具项目列表整合到一个命名列表中,用一次循环批量处理所有工具,彻底消除重复代码:
library(dplyr) # 整合所有测量工具的项目列表,一次定义后续复用 instruments <- list( stress = c("BES_1", "BES_2", "BES_3", "BES_4", "BES_5"), reg = c("Neg_ctrl1","Pos_ctrl1","Neg_use1","Pos_use1"), scs = c("Ind_1","Ind_2","Inter_1","Inter_2","Ind_3") ) # 初始化结果数据集 data5 <- data4 # 批量处理每个测量工具,一次循环完成所有标准化 for (items in instruments) { # 计算当前工具的行均值和行标准差 row_m <- rowMeans(data4[, items], na.rm = TRUE) row_sds <- apply(data4[, items], 1, sd, na.rm = TRUE) # 对当前工具的所有项目执行标准化转换 data5[, items] <- (data5[, items] - row_m) / row_sds } # 查看前10行结果 head(data5, 10)
代码说明
- 用命名列表
instruments统一管理三个工具的项目,避免分散定义 - 一次循环自动遍历所有工具,重复逻辑只写一次,代码量减少一半以上
- 完全保留原代码的计算逻辑,输出结果和原代码完全一致,同时提升了可读性和可维护性
如果偏好dplyr管道风格,也可以用以下写法:
library(dplyr) library(purrr) instruments <- list( stress = c("BES_1", "BES_2", "BES_3", "BES_4", "BES_5"), reg = c("Neg_ctrl1","Pos_ctrl1","Neg_use1","Pos_use1"), scs = c("Ind_1","Ind_2","Inter_1","Inter_2","Ind_3") ) data5 <- data4 %>% mutate( across(all_of(unlist(instruments)), ~ { # 自动匹配当前列所属的工具组 tool <- keep(instruments, ~ cur_column() %in% .x) %>% names() items <- instruments[[tool]] row_m <- rowMeans(data4[, items], na.rm = TRUE) row_sds <- apply(data4[, items], 1, sd, na.rm = TRUE) (.x - row_m) / row_sds }) )
内容的提问来源于stack exchange,提问作者YGaby
相关产品推荐
相关产品推荐

