使用R Studio中ggplot绘制多元回归参数的技术问题(已找到解决方案,见文末)
Got it, let's adjust your code to use the paper's exact regression coefficients instead of letting geom_smooth() fit a new model to your data. Here's a step-by-step solution:
Step 1: Build a prediction dataset
We'll create a grid of values for your x-axis variable (EDUC), include both levels of the binary LANG variable, and fix EXP and ysm to the mean values reported in the paper. This ensures we're calculating predictions exactly as the paper specifies.
Step 2: Compute predicted values with the paper's coefficients
We'll plug the fixed parameters directly into the regression formula from the paper to calculate predicted LNEARN for every combination of EDUC and LANG.
Step 3: Plot data points + fixed-coefficient lines
Keep your original scatter plot, then replace the data-driven geom_smooth() with our custom-calculated regression lines.
Full Working Code
library(ggplot2) library(ggthemes) # For theme_pander() # 1. Create prediction data grid # Use the full range of EDUC from your dataset educ_range <- seq(min(df_clean$EDUC), max(df_clean$EDUC), by = 0.5) pred_df <- expand.grid( EDUC = educ_range, LANG = c(0, 1), EXP = 22.76, # Paper's reported mean for EXP ysm = 15.43 # Paper's reported mean for ysm ) # 2. Calculate predicted LNEARN using the paper's exact formula pred_df$pred_LNEARN <- 5.056 + 0.049 * pred_df$EDUC + 0.023 * pred_df$EXP - 0.037 * (pred_df$EXP^2 / 100) + 0.028 * pred_df$ysm - 0.041 * (pred_df$ysm^2 / 100) + pred_df$LANG # LANG's coefficient is 1 as per the paper's model notation # 3. Generate the final plot ggplot(df_clean, aes(x = EDUC, y = LNEARN)) + # Keep original data points with color by LANG and size by ysm geom_point(aes(color = factor(LANG), size = ysm)) + # Add the fixed-coefficient regression lines geom_line(data = pred_df, aes(y = pred_LNEARN, color = factor(LANG)), linewidth = 1) + # Beautification (fixed a small legend label typo from your original code) xlab("Years of Formal Education") + ylab("Log of Earnings") + ggtitle("Education's Potential Impact On Immigrant Earnings") + labs(subtitle = "1990 US Census Data", color = "Language", size = "Years Since Migration") + theme_pander()
Key Notes:
- The
pred_LNEARNcalculation strictly follows the paper's regression equation, using every coefficient exactly as provided. - We lock
EXPandysmto their reported means to isolate the relationship betweenEDUC,LANG, andLNEARNas the paper presents it. - Replaced
geom_smooth()withgeom_line()using precomputed predictions to ensure the lines match the paper's model, not a new fit to your data. - Fixed a small typo in your original code: the size legend was incorrectly labeled "Education" (it maps to
ysm, years since migration).
更新:已找到解决方案,答案见上方代码
内容的提问来源于stack exchange,提问作者JellyfishSunset

