You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言DataFrame无法进行正态性检验的问题求助

Fixing Normality Test Issues with Your R DataFrame

Hey there! Let's break down why you can't run normality tests on your DataFrame and get it sorted out quickly.

The Root Problem

Looking at your str(df) output, both trtGroup and cd are stored as Factor types. Normality tests (like Shapiro-Wilk or Kolmogorov-Smirnov) require variables to be numeric—factors (categorical data) can't be tested for normality, which is why you're hitting roadblocks.

How to Fix It

Let's adjust your data processing steps to ensure numeric columns stay (or get converted to) numeric types:

  1. Fix the Initial Data Import & Conversion
    Your read.csv uses stringsAsFactors=FALSE (great call!), but the as.data.frame(lapply(df, unlist)) step might be converting character columns back to factors (older R versions default to stringsAsFactors=TRUE in as.data.frame). Let's rewrite that code to preserve numeric types:

    # Read CSV with explicit settings to avoid unwanted factors
    df <- read.csv('cdSH2015Fall.csv', dec = ".", na.strings = c("na"), header=TRUE, row.names=NULL, stringsAsFactors=FALSE)
    
    # Convert factor/character columns to numeric (skip categorical columns like trtGroup)
    # First, identify which columns should be numeric (exclude your grouping variable)
    numeric_columns <- setdiff(names(df), "trtGroup")
    
    # Batch convert to numeric, handling factors properly
    df[numeric_columns] <- lapply(df[numeric_columns], function(col) {
      # If the column is a factor, convert to character first to avoid level codes
      if (is.factor(col)) {
        col <- as.character(col)
      }
      as.numeric(col)
    })
    
  2. Verify the Data Structure
    Run str(df) again—you should see numeric columns (labeled num) instead of factors for all variables you want to test for normality. For example, cd should now show as num instead of Factor w/ 2 levels "0","1".

  3. Run Your Normality Test
    Now you can run tests like Shapiro-Wilk on numeric columns. For example, testing the cd column (with NA handling):

    # Shapiro-Wilk test for normality (omit NA values)
    shapiro.test(na.omit(df$cd))
    

Quick Notes

  • You don't need to run normality tests on categorical variables like trtGroup—those tests only apply to continuous numeric data.
  • If you still see non-numeric values in your columns, double-check your CSV file for stray characters (like spaces or symbols) that might be forcing columns to be read as characters/factors.

内容的提问来源于stack exchange,提问作者Delia Shelton

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:14:27