You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R Studio中ggplot绘制多元回归参数的技术问题(已找到解决方案,见文末)

Got it, let's adjust your code to use the paper's exact regression coefficients instead of letting geom_smooth() fit a new model to your data. Here's a step-by-step solution:

Step 1: Build a prediction dataset

We'll create a grid of values for your x-axis variable (EDUC), include both levels of the binary LANG variable, and fix EXP and ysm to the mean values reported in the paper. This ensures we're calculating predictions exactly as the paper specifies.

Step 2: Compute predicted values with the paper's coefficients

We'll plug the fixed parameters directly into the regression formula from the paper to calculate predicted LNEARN for every combination of EDUC and LANG.

Step 3: Plot data points + fixed-coefficient lines

Keep your original scatter plot, then replace the data-driven geom_smooth() with our custom-calculated regression lines.


Full Working Code

library(ggplot2)
library(ggthemes) # For theme_pander()

# 1. Create prediction data grid
# Use the full range of EDUC from your dataset
educ_range <- seq(min(df_clean$EDUC), max(df_clean$EDUC), by = 0.5)
pred_df <- expand.grid(
  EDUC = educ_range,
  LANG = c(0, 1),
  EXP = 22.76,  # Paper's reported mean for EXP
  ysm = 15.43   # Paper's reported mean for ysm
)

# 2. Calculate predicted LNEARN using the paper's exact formula
pred_df$pred_LNEARN <- 5.056 +
  0.049 * pred_df$EDUC +
  0.023 * pred_df$EXP -
  0.037 * (pred_df$EXP^2 / 100) +
  0.028 * pred_df$ysm -
  0.041 * (pred_df$ysm^2 / 100) +
  pred_df$LANG  # LANG's coefficient is 1 as per the paper's model notation

# 3. Generate the final plot
ggplot(df_clean, aes(x = EDUC, y = LNEARN)) +
  # Keep original data points with color by LANG and size by ysm
  geom_point(aes(color = factor(LANG), size = ysm)) +
  # Add the fixed-coefficient regression lines
  geom_line(data = pred_df, aes(y = pred_LNEARN, color = factor(LANG)), linewidth = 1) +
  # Beautification (fixed a small legend label typo from your original code)
  xlab("Years of Formal Education") + 
  ylab("Log of Earnings") + 
  ggtitle("Education's Potential Impact On Immigrant Earnings") + 
  labs(subtitle = "1990 US Census Data", color = "Language", size = "Years Since Migration") +
  theme_pander()

Key Notes:

  • The pred_LNEARN calculation strictly follows the paper's regression equation, using every coefficient exactly as provided.
  • We lock EXP and ysm to their reported means to isolate the relationship between EDUC, LANG, and LNEARN as the paper presents it.
  • Replaced geom_smooth() with geom_line() using precomputed predictions to ensure the lines match the paper's model, not a new fit to your data.
  • Fixed a small typo in your original code: the size legend was incorrectly labeled "Education" (it maps to ysm, years since migration).

更新:已找到解决方案,答案见上方代码

内容的提问来源于stack exchange,提问作者JellyfishSunset

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 19:27:37