如何在R语言中绘制Q-Q图?含手动实现多数据集Q-Q图任务
Got it, let's walk through how to complete this task properly, step by step. You've got a good start with fitting the normal distribution to your data, but we need to expand it to generate the required samples and build custom Q-Q plots without relying on qqnorm() or qqplot().
First, let's clean up the code for calculating the mean and standard deviation of your original data—we can use R's built-in functions for clarity and accuracy:
# Original observed data y <- c(194, 209, 205, 180, 196, 178, 214, 199, 224, 230) # Fit normal distribution parameters (sample mean and standard deviation) mu <- mean(y) sigma <- sd(y) # Print the fitted parameters for reference cat("Fitted Normal Distribution:\nMean =", round(mu, 2), "\nStandard Deviation =", round(sigma, 2), "\n")
We'll generate 3 independent samples, each with 100 values from the fitted normal distribution. Adding a seed ensures your results are reproducible:
# Set random seed for consistent results set.seed(123) # Generate 3 datasets (100 observations each) num_datasets <- 3 sample_size <- 100 samples <- replicate(num_datasets, rnorm(n = sample_size, mean = mu, sd = sigma)) # Check the structure: 100 rows (observations) × 3 columns (datasets) str(samples)
Q-Q plots work by comparing sorted sample values to theoretical quantiles from the reference distribution (our fitted normal). Here's a custom function to create these plots, plus a loop to plot all 3 datasets:
# Define a function to create a custom Q-Q plot custom_qqplot <- function(sample_data, dist_mean, dist_sd, plot_title) { # 1. Sort the sample observations sorted_sample <- sort(sample_data) # 2. Calculate empirical quantile positions (using (i - 0.5)/n for continuity correction) n <- length(sorted_sample) quantile_positions <- (1:n - 0.5) / n # 3. Get theoretical quantiles from the fitted normal distribution theoretical_quantiles <- qnorm(quantile_positions, mean = dist_mean, sd = dist_sd) # 4. Create the scatter plot plot(theoretical_quantiles, sorted_sample, xlab = "Theoretical Normal Quantiles", ylab = "Sorted Sample Values", main = plot_title, pch = 16, col = "#2c3e50") # 5. Add a reference line (ideal fit line: y = mu + sigma*x) abline(a = dist_mean, b = dist_sd, col = "#e74c3c", lwd = 2) } # Plot all 3 datasets in a 3-row grid par(mfrow = c(3, 1), mar = c(4, 4, 2, 1)) # Adjust plot margins for readability for (i in 1:num_datasets) { custom_qqplot(samples[, i], mu, sigma, title = paste("Q-Q Plot: Dataset", i)) } par(mfrow = c(1, 1)) # Reset plotting layout to default
Quick explanation of the Q-Q logic:
- We sort the sample to get empirical order statistics
- The
quantile_positionsformula adjusts for the fact that we can't have a quantile at exactly 0 or 1 for finite samples qnorm()gives us the values we'd expect from a perfect normal distribution at those positions- The red reference line shows where points would lie if the sample perfectly matched the fitted normal
If you want to double-check the fits, you can plot histograms with the fitted normal density curve for each dataset:
# Plot histograms with fitted density curves par(mfrow = c(3, 1), mar = c(4, 4, 2, 1)) for (i in 1:num_datasets) { hist(samples[, i], breaks = 15, freq = FALSE, main = paste("Histogram: Dataset", i), xlab = "Observed Value", col = "#bdc3c7") curve(dnorm(x, mu, sigma), add = TRUE, col = "#2980b9", lwd = 2) } par(mfrow = c(1, 1))
This completes all your requirements: we've fitted the normal distribution to your original data, generated 3 samples of 100 observations each, and created custom Q-Q plots without using R's built-in qqnorm() or qqplot() functions.
内容的提问来源于stack exchange,提问作者Kate

