基于ggplot的多组最小二乘二次拟合绘图技术咨询
问题
我需要为black、white、other三个种族组分别绘制log_wage与exp的二次拟合曲线,已知两者存在二次曲线关系。已经有一个计算二次拟合系数的函数:
quad_fit <- function(data_sub) { return(lm(log_wage~exp+I(exp^2),data=data_sub)$coefficients) } quad_fit(salary_data)
拟合公式为$\hat{Y} = a1 + a2x + a3x²$(其中$\hat{Y}$是log_wage,$x$是exp)。要求用ggplot完成绘图(base R绘图只能得部分分数),需要包含图例和合适标签。
我尝试了针对black组的代码,但不清楚stat_smooth的公式该怎么写,代码如下:
blackfit <- quad_fit(salary_data[salary_data$race == "black",]) whitefit <- quad_fit(salary_data[salary_data$race == "white",]) otherfit <- quad_fit(salary_data[salary_data$race == "other",]) yblack <- blackfit[1] + blackfit[2]*salary_data$exp + blackfit[3]*(salary_data$exp)^2 ywhite <- whitefit[1] + whitefit[2]*salary_data$exp + whitefit[3]*(salary_data$exp)^2 yother <- otherfit[1] + otherfit[2]*salary_data$exp + otherfit[3]*(salary_data$exp)^2 soloblack <- salary_data[salary_data$race == "black",] solowhite <- salary_data[salary_data$race == "white",] soloother <- salary_data[salary_data$race == "other",] ggplot(data = soloblack) + geom_point(aes(x = exp, y = log_wage)) + stat_smooth(aes(y = log_wage, x = exp), formula = y ~ yblack)
解决方案
你的思路方向是对的,但其实不用手动拆分数据集再单独计算拟合值——ggplot可以帮你自动按种族分组完成二次拟合,或者你也可以先整理好所有种族的拟合数据再绘图,下面给你两种可行的方法:
方法1:直接用ggplot的stat_smooth自动分组拟合
这是最简洁的方式,不需要手动调用quad_fit函数,直接在ggplot中指定二次公式,并通过color = race让ggplot按种族分组处理:
library(ggplot2) ggplot(salary_data, aes(x = exp, y = log_wage, color = race)) + # 添加散点,按种族区分颜色(alpha参数降低点的透明度避免重叠) geom_point(alpha = 0.5) + # 添加二次拟合曲线,指定线性回归方法与二次公式 stat_smooth(method = "lm", formula = y ~ x + I(x^2), se = FALSE, size = 1) + # 添加清晰的标签与标题 labs( x = "工作经验 (exp)", y = "对数工资 (log_wage)", color = "种族", title = "不同种族的对数工资与工作经验二次拟合曲线" ) + # 可选:用简约主题优化图表观感 theme_minimal()
关键说明:
method = "lm"指定用线性回归拟合,配合formula = y ~ x + I(x^2)就会拟合二次曲线se = FALSE关闭置信区间(如果需要保留置信区间,直接删除这个参数即可)color = race会自动为每个种族分配不同颜色的散点和拟合线,同时生成对应的图例
方法2:用你的quad_fit函数先计算拟合值再绘图
如果你一定要用自己写的quad_fit函数,那可以先为每个种族生成匹配的拟合数据序列,再合并绘图:
library(ggplot2) library(dplyr) # 1. 按种族拆分数据集,批量计算拟合值 split_data <- split(salary_data, salary_data$race) fit_data_list <- lapply(names(split_data), function(race_group) { data_sub <- split_data[[race_group]] fit_coef <- quad_fit(data_sub) # 注意:用当前种族子集的exp数据计算拟合值,避免数据不匹配 data_sub$fit_log_wage <- fit_coef[1] + fit_coef[2] * data_sub$exp + fit_coef[3] * (data_sub$exp)^2 data_sub$race <- race_group return(data_sub) }) # 合并所有种族的拟合数据 fit_data <- do.call(rbind, fit_data_list) # 2. 用ggplot绘制散点与拟合曲线 ggplot(fit_data, aes(x = exp, color = race)) + geom_point(aes(y = log_wage), alpha = 0.5) + geom_line(aes(y = fit_log_wage), size = 1) + labs( x = "工作经验 (exp)", y = "对数工资 (log_wage)", color = "种族", title = "不同种族的对数工资与工作经验二次拟合曲线" ) + theme_minimal()
修正你之前的错误:你之前用整个数据集的salary_data$exp计算拟合值,这会导致拟合值与对应种族的散点数据不匹配,这里改用每个种族子集的exp数据就解决了这个问题。
两种方法都能生成符合要求的图表,第一种更高效,第二种适合需要手动控制拟合过程的场景。
内容的提问来源于stack exchange,提问作者Dario Gentiletti
相关产品推荐
相关产品推荐

