R语言t.test()执行报错:观测值不足问题的代码修改咨询
问题与解决方案
问题背景
开展调研数据分析,通过随机生成器EM05分配了EM01、EM02、EM03三种实验处理,希望分别对EM01/EM02/EM03组的einst_index进行t检验。已完成数据预处理(含项目得分反转、构建einst_index),但执行t.test()时触发错误:
Error in t.test.default(mig_neutral_stim$einst_index, mig_neutral_stim$EM05, : not enough 'x' observations
错误原因
- 检验逻辑错误:原代码错误地将分组变量
EM05作为t检验的第二个输入变量,且设置paired = TRUE。实际需求是独立样本t检验,比较不同实验处理组的einst_index均值差异,而非配对检验。 - 数据处理错误:原代码仅过滤出
EM03组数据,且尝试将字符串类型的EM05转为数值型,导致该列全为NA,引发观测值不足的错误。 - 语法错误:过滤缺失值的代码缺少闭合括号,导致数据处理不完整。
修正后的代码
library(tidyverse) library(dplyr) library(report) library(psych) library(readxl) # 读取数据集 mig <- read_excel("data/data_ExperimentMigration2024_2024-07-04_20-34.xlsx") # 定义得分反转函数,简化重复代码 reverse_score <- function(x) { case_when( x == 1 ~ 5, x == 2 ~ 4, x == 3 ~ 3, x == 4 ~ 2, x == 5 ~ 1 ) } # 执行项目得分反转 mig <- mig %>% mutate( new_ME03_09 = reverse_score(ME03_09), new_ME03_04 = reverse_score(ME03_04), new_ME03_02 = reverse_score(ME03_02), new_ME03_01 = reverse_score(ME03_01) ) # 构建einst_index指数 mig_selected <- mig %>% select(new_ME03_01, new_ME03_02, ME03_03, new_ME03_04, ME03_05, ME03_07, new_ME03_09, ME03_10) %>% mutate(across(everything(), as.numeric)) mig <- mig %>% mutate(einst_index = rowMeans(mig_selected, na.rm = TRUE)) # 查看指数的均值与标准差 mig %>% summarise( mean_einst_index = mean(einst_index, na.rm = TRUE), sd_einst_index = sd(einst_index, na.rm = TRUE) ) # 检查数据结构 str(mig) # 确保einst_index为数值型 mig$einst_index <- as.numeric(mig$einst_index) # 1. EM01 vs EM02 独立样本t检验 t_test_em01_em02 <- t.test(einst_index ~ EM05, data = mig, subset = EM05 %in% c("EM01", "EM02")) print(t_test_em01_em02) report(t_test_em01_em02) # 2. EM01 vs EM03 独立样本t检验 t_test_em01_em03 <- t.test(einst_index ~ EM05, data = mig, subset = EM05 %in% c("EM01", "EM03")) print(t_test_em01_em03) report(t_test_em01_em03) # 3. EM02 vs EM03 独立样本t检验 t_test_em02_em03 <- t.test(einst_index ~ EM05, data = mig, subset = EM05 %in% c("EM02", "EM03")) print(t_test_em02_em03) report(t_test_em02_em03) # 可选:检查每组样本量,确保每组有足够观测值 mig %>% count(EM05)
关键修正点
- 简化得分反转逻辑:定义
reverse_score函数,避免重复编写反转规则。 - 修正检验逻辑:使用公式形式
einst_index ~ EM05执行独立样本t检验,通过subset参数指定要比较的两组。 - 修复语法问题:补全原代码中缺失的闭合括号,同时通过
rowMeans的na.rm=TRUE处理缺失值,无需额外过滤。 - 添加样本量检查:通过
count(EM05)确认每组的观测数量,提前排查样本量不足的问题。
内容的提问来源于stack exchange,提问作者Alina Dzaferovic
相关产品推荐
相关产品推荐

