如何用R中factanal的旧因子载荷计算新数据的因子得分?
Solution for Calculating Factor Scores on New Data with
factanal Results Got it, let's tackle this problem. When you use factanal with scores="regression", it only computes scores for the training data you passed in. To get scores for new data, you'll need to replicate the scaling and factor score calculation logic using the parameters from your fitted factanal object. Here's how to do it step by step:
First, let's recap what we need from the fitted model:
- The mean and standard deviation of the original 30 variables (from your training data
mydata[,-1]) — because factor analysis relies on standardized variables, and we need to apply the same scaling tonew_data. - The rotated factor loadings from
fitted_data$loadings(since you usedrotation="varimax"). - The regression-style score coefficients, which we can compute using the loadings.
Here's the complete code:
# Step 1: Save scaling parameters from training data (the 30 variables used in factanal) train_vars <- mydata[,-1] train_means <- colMeans(train_vars) train_sds <- apply(train_vars, 2, sd) # Step 2: Extract rotated factor loadings from the fitted factanal object loadings_mat <- as.matrix(fitted_data$loadings) # Step 3: Calculate regression-style score coefficients # Formula matches the "regression" score method used in factanal score_coef <- loadings_mat %*% solve(t(loadings_mat) %*% loadings_mat) # Step 4: Standardize the new data using training data's mean and sd # Critical: Ensure new_data has the SAME 30 variables in the SAME order as train_vars! new_data_std <- scale(new_data, center = train_means, scale = train_sds) # Step 5: Compute factor scores for new_data new_factor_scores <- new_data_std %*% score_coef # Optional: Rename columns to match your original factor labels colnames(new_factor_scores) <- paste0("Factor", 1:7) # Now combine scores with new_data and use your pre-trained lm model for prediction new_data_with_scores <- cbind(new_data, new_factor_scores) predictions <- predict(mod1, newdata = new_data_with_scores)
Key Notes:
- Variable Alignment: Double-check that
new_datahas exactly the same 30 variables in the same order asmydata[,-1]— misalignment will break the scaling and loading matching. - Consistent Scaling: We use the training data's mean and standard deviation to standardize
new_data, not the new data's own stats. This ensures consistency with how the original factor analysis was computed. - Score Method Validation: If you want to confirm this works, apply the code to your original training data — the resulting scores should match
fitted_data$scores(up to tiny floating-point differences).
内容的提问来源于stack exchange,提问作者badders
相关产品推荐
相关产品推荐

