在R中绘制因子分析聚类可视化图求助:已得到4个潜变量
Hey there! I get that you’ve run a factor analysis in R with 4 latent variables and want a visualization similar to the one you shared to make your results more intuitive. Let’s walk through two solid approaches to build that—one focused on clear factor loading plots, and another that mimics the clustered heatmap style you’re targeting.
First, Prep Your Data & Install Required Packages
First, make sure you have these packages installed (they’ll handle both visualization and data manipulation):
install.packages(c("factoextra", "ggplot2", "reshape2", "ggdendro", "gridExtra"))
Then, extract the factor loadings from your factor2 object into a usable data frame (we’ll filter out any loadings below your 0.3 cutoff too):
# Convert factor loadings to a data frame loadings_df <- as.data.frame(unclass(factor2$loadings)) # Add variable names as a column loadings_df$Variable <- rownames(loadings_df) # Remove row names rownames(loadings_df) <- NULL # Keep only variables with at least one loading ≥ 0.3 (matches your cutoff) loadings_df <- loadings_df[rowSums(abs(loadings_df[,1:4])) >= 0.3, ]
Approach 1: Simple Factor Loading Scatter Plot (Great for Factor Relationships)
If you want to visualize how variables map to your factors (focusing on the first two factors, which explain the most variance), use the factoextra package’s built-in function for factor analysis:
library(factoextra) # Plot loadings for MR1 and MR2, colored by variable contribution fviz_fa_var(factor2, col.var = "contrib", # Color variables by their contribution to factors gradient.cols = c("#00AFBB", "#E7B800", "#FC4E07"), repel = TRUE, # Prevent label overlap title = "Factor Loadings: MR1 vs MR2")
This plot shows each variable’s loading on the first two factors, with darker colors representing variables that contribute more to the factors. It’s perfect for spotting which variables align with which latent factors.
Approach 2: Clustered Heatmap (Matches Your Target Visual)
To replicate the clustered style of the image you shared (with a dendrogram showing variable clusters alongside factor loadings), use this combination of clustering and heatmap code:
library(ggplot2) library(reshape2) library(ggdendro) library(gridExtra) # Step 1: Cluster variables based on their factor loadings d <- dist(scale(loadings_df[,1:4]), method = "euclidean") hc <- hclust(d, method = "ward.D2") # Ward's method for tight clusters # Step 2: Get the ordered variable names from the cluster tree order_vars <- hc$labels[hc$order] # Step 3: Reshape data for the heatmap loadings_df$Variable <- factor(loadings_df$Variable, levels = order_vars) loadings_melt <- melt(loadings_df, id.vars = "Variable", variable.name = "Factor", value.name = "Loading") # Step 4: Build the dendrogram and heatmap, then combine them # Dendrogram plot dendro_plot <- ggdendrogram(hc, rotate = TRUE, theme_dendro = FALSE) + labs(y = "Cluster Distance") + theme(axis.text.y = element_blank(), axis.ticks.y = element_blank()) # Heatmap of factor loadings heatmap_plot <- ggplot(loadings_melt, aes(x = Factor, y = Variable, fill = Loading)) + geom_tile(color = "white") + scale_fill_gradient2(low = "#2c7bb6", mid = "white", high = "#d7191c", midpoint = 0, limit = c(-1,1), name = "Factor Loading") + theme_minimal() + theme(axis.text.y = element_text(size = 8)) # Combine both plots grid.arrange(dendro_plot, heatmap_plot, widths = c(1, 3))
This output will have a dendrogram on the left showing how variables cluster based on their factor loadings, and a heatmap on the right where each cell’s color represents the strength and direction of the variable’s loading on each factor. It’s exactly the kind of intuitive, clustered visualization you’re looking for.
Quick Notes
- If your
factor2object was created withfactanal()(common for exploratory factor analysis), both approaches will work seamlessly. - Adjust the clustering method (e.g.,
method = "complete"instead of"ward.D2") or color palettes if you want to tweak the look to match your preferences.
内容的提问来源于stack exchange,提问作者Gabriela Simona

