R语言分位数归一化技术咨询:qqnorm函数及数据转换问题
Hey there! Let's work through your quantile normalization questions step by step—this is a common gotcha with R's qqnorm() function and messy real-world data, so you're not alone.
$x in qqnorm() First up: when you run qqnorm(x, plot.it = F), the function returns a list with two key components:
$y: The original sample quantiles (your input data's values sorted and mapped to quantile positions)$x: The theoretical quantiles from a standard normal distribution (mean 0, standard deviation 1) that align with your sample's quantile structure.
That $x is exactly the normalized values you want—it’s what transforms your raw data to fit the quantiles of a normal distribution. To see this clearly, run str(qqnorm(x, plot.it = F)) on your test subset: you’ll see the list structure, and extracting $x pulls out the normalized numeric output (a matrix if your input is a data frame, a vector if it’s a single column).
The issue boils down to two problems with your actual input data:
- Your "NA" values are strings, not real R missing values: R recognizes unquoted
NAas a missing value, but"NA"is just a character string. This forces columnsC,D, andEto be character-type instead of numeric. qqnorm()requires numeric input: It can’t process character vectors, so trying to run it on your originalydata frame throws an error.
Here’s how to fix your data and run the normalization correctly:
# Load your real dataset x <- data.frame(A =c("A","B","C","D"), B = c("X", "Y", "Z","L"), C=c(193973, 185750, 185511, "NA"), D = c(56433,52298, 53040, "NA"), E = c(4668, 6074,6246, "NA")) # Select columns to transform y <- x[,3:5] # Step 1: Replace "NA" strings with actual R NA values y[y == "NA"] <- NA # Step 2: Convert columns from character to numeric (the warning about NAs is expected and harmless) y <- apply(y, 2, as.numeric) # Now run quantile normalization on each column ytransformed <- apply(y, 2, function(col) qqnorm(col, plot.it = FALSE)$x)
The apply() function runs qqnorm() on each column individually, since qqnorm() works best with vectors. This will give you a matrix of normalized values matching your original y structure.
To confirm your normalization worked, you can compare QQ plots (to check normality) and density plots (to compare distribution shapes before/after):
QQ Plots (Check Normality Fit)
QQ plots show how well your data aligns with a normal distribution. After normalization, points should lie close to the diagonal reference line:
# Set up a 2-row, 3-column plot grid par(mfrow = c(2, 3)) for(i in 1:ncol(y)){ # Original data QQ plot qqnorm(y[,i], main = paste("Original: ", colnames(y)[i])) qqline(y[,i], col = "red", lwd = 2) # Transformed data QQ plot qqnorm(ytransformed[,i], main = paste("Normalized: ", colnames(y)[i])) qqline(ytransformed[,i], col = "blue", lwd = 2) } # Reset plot layout to default par(mfrow = c(1, 1))
Density Plots (Compare Distribution Shapes)
These visualize how the shape of your data changes after normalization:
# Set up a 1-row, 3-column grid par(mfrow = c(1, 3)) for(i in 1:ncol(y)){ # Plot original density plot(density(y[,i], na.rm = TRUE), main = colnames(y)[i], col = "red", lwd = 2, xlab = "Value") # Overlay transformed density lines(density(ytransformed[,i], na.rm = TRUE), col = "blue", lwd = 2) # Add legend for clarity legend("topright", legend = c("Original", "Normalized"), col = c("red", "blue"), lwd = 2) } par(mfrow = c(1, 1))
After normalization, the transformed data’s QQ plot should follow the diagonal much more closely, and the density should resemble a standard normal bell curve (centered around 0, symmetric shape).
内容的提问来源于stack exchange,提问作者star

