R语言报错:data.frame参数行数不一致问题求助
问题分析与修复方案
错误根源
- 变量名拼写错误:处理ColumnA检验结果时,误写为
shapiro_result_pos$statistic,但实际定义的变量是shapiro_result_A。不存在的变量会返回空值,导致新创建的dataframe列长度不匹配(0行 vs 1行),触发报错。 - 未应用筛选条件:原需求是依据列X(取值yes/no)和列Z(取值Characteristic A)筛选数据后再做检验,但当前代码直接对整列进行检验,不符合要求。
- 数据集列表未命名:创建
Exp_list时未指定元素名称,导致循环中names(Exp_list)[i]无法获取正确的数据集名称。
修复步骤
- 修正变量名拼写错误,将
shapiro_result_pos$statistic改为shapiro_result_A$statistic。 - 添加数据筛选逻辑,在执行Shapiro检验前按X和Z的条件过滤数据集。
- 为
Exp_list的元素命名,确保能正确记录每个数据集的名称。 - 优化结果收集方式,避免全局赋值
<<-,改为函数返回检验结果后统一合并,减少副作用。
修正后的完整代码
# 创建空结果数据集 results_df <- data.frame(DataframeName = character(), TestType = character(), WStatistic = numeric(), PValue = numeric(), stringsAsFactors = FALSE) # 为数据集列表命名,确保能正确获取名称 Exp_list <- list(Exp1 = Exp1, Exp2 = Exp2, Exp3 = Exp3) # 定义检验函数:返回单次检验的结果行,避免全局赋值 perform_normality_test <- function(df, dataframe_name) { # 按条件筛选数据:X为yes,Z为Characteristic A filtered_df <- df[df$X == "yes" & df$Z == "Characteristic A", ] if (nrow(filtered_df) < 3) { cat("数据集", dataframe_name, "筛选后数据量不足,无法执行正态性检验。\n") return(NULL) # 返回空,不添加结果 } else { results <- list() # 处理ColumnA shapiro_result_A <- shapiro.test(filtered_df$ColumnA) results[[1]] <- data.frame(DataframeName = dataframe_name, TestType = "ColumnA", WStatistic = shapiro_result_A$statistic, PValue = shapiro_result_A$p.value, stringsAsFactors = FALSE) cat("数据集", dataframe_name, "的ColumnA Shapiro-Wilk检验结果:\n") cat("W统计量:", shapiro_result_A$statistic, "\n") cat("p值:", shapiro_result_A$p.value, "\n\n") # 处理ColumnB shapiro_result_B <- shapiro.test(filtered_df$ColumnB) results[[2]] <- data.frame(DataframeName = dataframe_name, TestType = "ColumnB", WStatistic = shapiro_result_B$statistic, PValue = shapiro_result_B$p.value, stringsAsFactors = FALSE) cat("数据集", dataframe_name, "的ColumnB Shapiro-Wilk检验结果:\n") cat("W统计量:", shapiro_result_B$statistic, "\n") cat("p值:", shapiro_result_B$p.value, "\n\n") # 返回两个检验结果的组合 do.call(rbind, results) } } # 遍历数据集列表,收集所有结果 for (i in 1:length(Exp_list)) { df <- Exp_list[[i]] dataframe_name <- names(Exp_list)[i] cat("正在处理数据集:", dataframe_name, "\n") test_results <- perform_normality_test(df, dataframe_name) # 如果有结果,合并到results_df if (!is.null(test_results)) { results_df <- rbind(results_df, test_results) } } print(results_df)
额外说明
- 筛选条件可根据实际情况调整(比如X的取值大小写、Z的特征值匹配规则)。
- 使用
do.call(rbind, results)合并函数内的检验结果,比多次调用rbind更高效。 - 当筛选后数据量小于3时,Shapiro-Wilk检验无法执行,函数会返回空并提示信息,避免后续错误。
内容的提问来源于stack exchange,提问作者Hieuwd
相关产品推荐
相关产品推荐

