如何优化加州住房数据集PCA双标图(Biplot)的可读性?
Got it, let's fix that messy biplot you're getting from biplot(housingpr, scale = 0). Here are actionable, R-specific tweaks to make your plot clear and informative:
Add meaningful axis labels with variance explained
Default axis labels just say "PC1" and "PC2"—that doesn't tell you how much variance each component captures. First calculate the variance percentages:variance <- summary(housingpr)$importance[2, 1:2] pc1_label <- paste0("Principal Component 1 (", round(variance[1]*100, 1), "% Variance)") pc2_label <- paste0("Principal Component 2 (", round(variance[2]*100, 1), "% Variance)")Then pass these to your biplot:
biplot(housingpr, scale = 0, xlab = pc1_label, ylab = pc2_label, main = "PCA Biplot - California Housing")Fix overlapping variable labels with custom plotting
The base Rbiplot()function often crams variable labels together. Instead, build the plot manually to control arrow and label positions:# Extract PCA results pca_results <- prcomp(your_housing_data, scale. = TRUE) # Replace with your raw data frame pc_scores <- pca_results$x[, 1:2] loadings <- pca_results$rotation[, 1:2] # Plot observations with low opacity to reduce clutter plot(pc_scores, pch = 16, col = rgb(0,0,0, alpha = 0.3), xlab = pc1_label, ylab = pc2_label) # Add scaled arrows for variables arrow_scale <- 2 # Adjust this to make arrows longer/shorter arrows(0, 0, loadings[,1]*arrow_scale, loadings[,2]*arrow_scale, length = 0.1, col = "darkred") # Place labels slightly outside arrows to avoid overlap text(loadings[,1]*arrow_scale*1.1, loadings[,2]*arrow_scale*1.1, rownames(loadings), col = "darkred", cex = 0.8)Use the
ggbiplotpackage for flexible, clean visuals
If you're open to ggplot2-based tools,ggbiplotis a game-changer for PCA plots. It handles label placement better and supports grouping:# Install if needed: install.packages("ggbiplot") library(ggbiplot) # Plot with optimized labels and optional grouping (e.g., by ocean proximity) ggbiplot(housingpr, obs.scale = 1, var.scale = 1, groups = your_housing_data$ocean_proximity, # Replace with your categorical variable ellipse = TRUE, circle = TRUE) + theme_minimal() + theme(legend.position = "bottom", panel.grid = element_blank()) + labs(title = "PCA Biplot - California Housing Dataset", x = pc1_label, y = pc2_label)Tweak element sizes to reduce clutter
In base R, use thecexparameter to shrink labels and points—critical for your 20k+ observations:biplot(housingpr, scale = 0, cex = 0.7, xlab = pc1_label, ylab = pc2_label)Smaller text and points mean less overlap without losing key information.
Remove unnecessary grid lines
Grid lines can make a dense plot look busier. Turn them off in base R withgrid = FALSE:biplot(housingpr, scale = 0, grid = FALSE, cex = 0.7, xlab = pc1_label, ylab = pc2_label)
内容的提问来源于stack exchange,提问作者user13601395

