You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言报错:data.frame参数行数不一致问题求助

问题分析与修复方案

错误根源

  • 变量名拼写错误:处理ColumnA检验结果时,误写为shapiro_result_pos$statistic,但实际定义的变量是shapiro_result_A。不存在的变量会返回空值,导致新创建的dataframe列长度不匹配(0行 vs 1行),触发报错。
  • 未应用筛选条件:原需求是依据列X(取值yes/no)和列Z(取值Characteristic A)筛选数据后再做检验,但当前代码直接对整列进行检验,不符合要求。
  • 数据集列表未命名:创建Exp_list时未指定元素名称,导致循环中names(Exp_list)[i]无法获取正确的数据集名称。

修复步骤

  1. 修正变量名拼写错误,将shapiro_result_pos$statistic改为shapiro_result_A$statistic。
  2. 添加数据筛选逻辑,在执行Shapiro检验前按X和Z的条件过滤数据集。
  3. 为Exp_list的元素命名,确保能正确记录每个数据集的名称。
  4. 优化结果收集方式,避免全局赋值<<-,改为函数返回检验结果后统一合并,减少副作用。

修正后的完整代码

# 创建空结果数据集
results_df <- data.frame(DataframeName = character(),
                         TestType = character(),
                         WStatistic = numeric(),
                         PValue = numeric(),
                         stringsAsFactors = FALSE)

# 为数据集列表命名,确保能正确获取名称
Exp_list <- list(Exp1 = Exp1, Exp2 = Exp2, Exp3 = Exp3)

# 定义检验函数:返回单次检验的结果行,避免全局赋值
perform_normality_test <- function(df, dataframe_name) {
  # 按条件筛选数据:X为yes,Z为Characteristic A
  filtered_df <- df[df$X == "yes" & df$Z == "Characteristic A", ]
  
  if (nrow(filtered_df) < 3) { 
    cat("数据集", dataframe_name, "筛选后数据量不足,无法执行正态性检验。\n")
    return(NULL)  # 返回空,不添加结果
  } else { 
    results <- list()
    
    # 处理ColumnA
    shapiro_result_A <- shapiro.test(filtered_df$ColumnA)
    results[[1]] <- data.frame(DataframeName = dataframe_name,
                               TestType = "ColumnA",
                               WStatistic = shapiro_result_A$statistic,
                               PValue = shapiro_result_A$p.value,
                               stringsAsFactors = FALSE)
    
    cat("数据集", dataframe_name, "的ColumnA Shapiro-Wilk检验结果:\n")
    cat("W统计量:", shapiro_result_A$statistic, "\n")
    cat("p值:", shapiro_result_A$p.value, "\n\n")
    
    # 处理ColumnB
    shapiro_result_B <- shapiro.test(filtered_df$ColumnB)
    results[[2]] <- data.frame(DataframeName = dataframe_name,
                               TestType = "ColumnB",
                               WStatistic = shapiro_result_B$statistic,
                               PValue = shapiro_result_B$p.value,
                               stringsAsFactors = FALSE)
    
    cat("数据集", dataframe_name, "的ColumnB Shapiro-Wilk检验结果:\n")
    cat("W统计量:", shapiro_result_B$statistic, "\n")
    cat("p值:", shapiro_result_B$p.value, "\n\n")
    
    # 返回两个检验结果的组合
    do.call(rbind, results)
  }
}

# 遍历数据集列表,收集所有结果
for (i in 1:length(Exp_list)) { 
  df <- Exp_list[[i]]
  dataframe_name <- names(Exp_list)[i]
  cat("正在处理数据集:", dataframe_name, "\n")
  test_results <- perform_normality_test(df, dataframe_name)
  # 如果有结果,合并到results_df
  if (!is.null(test_results)) {
    results_df <- rbind(results_df, test_results)
  }
}

print(results_df)

额外说明

  • 筛选条件可根据实际情况调整(比如X的取值大小写、Z的特征值匹配规则)。
  • 使用do.call(rbind, results)合并函数内的检验结果,比多次调用rbind更高效。
  • 当筛选后数据量小于3时,Shapiro-Wilk检验无法执行,函数会返回空并提示信息,避免后续错误。

内容的提问来源于stack exchange,提问作者Hieuwd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 01:12:08