如何基于梯度下降法实现含参数a、b的泛函优化?
梯度下降法实现线性模型参数优化(C#)
核心逻辑梳理
你要实现的本质是线性回归任务(示例中的a*x + b是线性预测模型),泛函最小化的目标是让模型预测值与真实样本的误差尽可能小。L1/L2范数就是两种常用的误差衡量标准:
- L2范数(均方误差MSE):计算误差平方的平均值,是光滑可导的误差函数,梯度计算简单,是最常用的选择
- L1范数(平均绝对误差MAE):计算误差绝对值的平均值,对异常值的鲁棒性更强,但绝对值函数在0点不可导,实现时需要用符号函数近似梯度
分步实现与代码示例
1. 基础框架与参数初始化
先定义模型类,初始化学习率、最大迭代次数、收敛阈值等超参数,以及待优化的a、b参数。
2. L2范数(MSE)的梯度下降实现
MSE的计算公式为:
MSE = (1/N) * Σ(预测值 - 真实值)²
对参数a、b的偏导数(梯度)为:
- ∂MSE/∂a = (2/N) * Σ(x*(a*x + b - y_true))
- ∂MSE/∂b = (2/N) * Σ(a*x + b - y_true)
对应的C#代码:
using System; using System.Collections.Generic; public class LinearRegressionGD { public double A { get; private set; } public double B { get; private set; } private readonly double _learningRate; private readonly int _maxIterations; private readonly double _tolerance; public LinearRegressionGD(double learningRate = 0.01, int maxIterations = 10000, double tolerance = 1e-6) { _learningRate = learningRate; _maxIterations = maxIterations; _tolerance = tolerance; A = 0; B = 0; } // 基于L2范数(MSE)训练模型 public void TrainMSE(List<double> x, List<double> y) { if (x.Count != y.Count) throw new ArgumentException("X与Y的样本数量必须一致"); int sampleCount = x.Count; double lastError = double.MaxValue; for (int iter = 0; iter < _maxIterations; iter++) { double gradA = 0; double gradB = 0; double currentError = 0; // 遍历所有样本计算梯度与当前误差 for (int i = 0; i < sampleCount; i++) { double yPred = A * x[i] + B; double error = yPred - y[i]; currentError += error * error; gradA += x[i] * error; gradB += error; } // 计算平均梯度与平均误差 gradA *= 2.0 / sampleCount; gradB *= 2.0 / sampleCount; currentError /= sampleCount; // 更新参数:向梯度反方向移动 A -= _learningRate * gradA; B -= _learningRate * gradB; // 检查收敛:误差变化小于阈值则停止迭代 if (Math.Abs(currentError - lastError) < _tolerance) break; lastError = currentError; } } // 预测方法 public double Predict(double x) { return A * x + B; } } // 使用示例 class Program { static void Main() { var xSamples = new List<double> { 1, 2, 3, 4, 5 }; var ySamples = new List<double> { 2.1, 3.9, 6.2, 7.8, 10.1 }; var model = new LinearRegressionGD(learningRate: 0.01, maxIterations: 10000); model.TrainMSE(xSamples, ySamples); Console.WriteLine($"优化后参数: a = {model.A:F4}, b = {model.B:F4}"); Console.WriteLine($"x=6的预测值: {model.Predict(6):F4}"); } }
3. L1范数(MAE)的梯度下降实现
MAE的计算公式为:
MAE = (1/N) * Σ|预测值 - 真实值|
由于绝对值函数在0点不可导,用符号函数近似梯度:
- ∂MAE/∂a = (1/N) * Σ(x * sign(a*x + b - y_true))
- ∂MAE/∂b = (1/N) * Σ(sign(a*x + b - y_true))
其中sign(z):z>0返回1,z<0返回-1,z=0返回0
只需修改训练方法:
// 基于L1范数(MAE)训练模型 public void TrainMAE(List<double> x, List<double> y) { if (x.Count != y.Count) throw new ArgumentException("X与Y的样本数量必须一致"); int sampleCount = x.Count; double lastError = double.MaxValue; for (int iter = 0; iter < _maxIterations; iter++) { double gradA = 0; double gradB = 0; double currentError = 0; for (int i = 0; i < sampleCount; i++) { double yPred = A * x[i] + B; double error = yPred - y[i]; currentError += Math.Abs(error); // 符号函数近似梯度 double sign = error > 0 ? 1 : error < 0 ? -1 : 0; gradA += x[i] * sign; gradB += sign; } gradA /= sampleCount; gradB /= sampleCount; currentError /= sampleCount; A -= _learningRate * gradA; B -= _learningRate * gradB; if (Math.Abs(currentError - lastError) < _tolerance) break; lastError = currentError; } }
关键注意事项
- 学习率调整:过大会导致参数震荡不收敛,过小会让训练速度极慢,建议从0.001、0.01、0.1等数值开始测试
- 数据归一化:如果x的取值范围过大(比如0-1000),建议先将x归一化到[0,1]或[-1,1]区间,避免梯度爆炸或训练效率低下
- 收敛判断:可以同时使用误差变化阈值和最大迭代次数,防止模型陷入无限迭代
内容的提问来源于stack exchange,提问作者JAAHARA-SEI XS TUMALUNGWA
相关产品推荐
相关产品推荐

