You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于梯度下降法实现含参数a、b的泛函优化?

梯度下降法实现线性模型参数优化(C#)

核心逻辑梳理

你要实现的本质是线性回归任务(示例中的a*x + b是线性预测模型),泛函最小化的目标是让模型预测值与真实样本的误差尽可能小。L1/L2范数就是两种常用的误差衡量标准:

  • L2范数(均方误差MSE):计算误差平方的平均值,是光滑可导的误差函数,梯度计算简单,是最常用的选择
  • L1范数(平均绝对误差MAE):计算误差绝对值的平均值,对异常值的鲁棒性更强,但绝对值函数在0点不可导,实现时需要用符号函数近似梯度

分步实现与代码示例

1. 基础框架与参数初始化

先定义模型类,初始化学习率、最大迭代次数、收敛阈值等超参数,以及待优化的a、b参数。

2. L2范数(MSE)的梯度下降实现

MSE的计算公式为:

MSE = (1/N) * Σ(预测值 - 真实值)²

对参数a、b的偏导数(梯度)为:

  • ∂MSE/∂a = (2/N) * Σ(x*(a*x + b - y_true))
  • ∂MSE/∂b = (2/N) * Σ(a*x + b - y_true)

对应的C#代码:

using System;
using System.Collections.Generic;

public class LinearRegressionGD
{
    public double A { get; private set; }
    public double B { get; private set; }
    private readonly double _learningRate;
    private readonly int _maxIterations;
    private readonly double _tolerance;

    public LinearRegressionGD(double learningRate = 0.01, int maxIterations = 10000, double tolerance = 1e-6)
    {
        _learningRate = learningRate;
        _maxIterations = maxIterations;
        _tolerance = tolerance;
        A = 0;
        B = 0;
    }

    // 基于L2范数(MSE)训练模型
    public void TrainMSE(List<double> x, List<double> y)
    {
        if (x.Count != y.Count)
            throw new ArgumentException("X与Y的样本数量必须一致");

        int sampleCount = x.Count;
        double lastError = double.MaxValue;

        for (int iter = 0; iter < _maxIterations; iter++)
        {
            double gradA = 0;
            double gradB = 0;
            double currentError = 0;

            // 遍历所有样本计算梯度与当前误差
            for (int i = 0; i < sampleCount; i++)
            {
                double yPred = A * x[i] + B;
                double error = yPred - y[i];
                currentError += error * error;
                gradA += x[i] * error;
                gradB += error;
            }

            // 计算平均梯度与平均误差
            gradA *= 2.0 / sampleCount;
            gradB *= 2.0 / sampleCount;
            currentError /= sampleCount;

            // 更新参数:向梯度反方向移动
            A -= _learningRate * gradA;
            B -= _learningRate * gradB;

            // 检查收敛:误差变化小于阈值则停止迭代
            if (Math.Abs(currentError - lastError) < _tolerance)
                break;

            lastError = currentError;
        }
    }

    // 预测方法
    public double Predict(double x)
    {
        return A * x + B;
    }
}

// 使用示例
class Program
{
    static void Main()
    {
        var xSamples = new List<double> { 1, 2, 3, 4, 5 };
        var ySamples = new List<double> { 2.1, 3.9, 6.2, 7.8, 10.1 };

        var model = new LinearRegressionGD(learningRate: 0.01, maxIterations: 10000);
        model.TrainMSE(xSamples, ySamples);

        Console.WriteLine($"优化后参数: a = {model.A:F4}, b = {model.B:F4}");
        Console.WriteLine($"x=6的预测值: {model.Predict(6):F4}");
    }
}

3. L1范数(MAE)的梯度下降实现

MAE的计算公式为:

MAE = (1/N) * Σ|预测值 - 真实值|

由于绝对值函数在0点不可导,用符号函数近似梯度:

  • ∂MAE/∂a = (1/N) * Σ(x * sign(a*x + b - y_true))
  • ∂MAE/∂b = (1/N) * Σ(sign(a*x + b - y_true))
    其中sign(z):z>0返回1,z<0返回-1,z=0返回0

只需修改训练方法:

// 基于L1范数(MAE)训练模型
public void TrainMAE(List<double> x, List<double> y)
{
    if (x.Count != y.Count)
        throw new ArgumentException("X与Y的样本数量必须一致");

    int sampleCount = x.Count;
    double lastError = double.MaxValue;

    for (int iter = 0; iter < _maxIterations; iter++)
    {
        double gradA = 0;
        double gradB = 0;
        double currentError = 0;

        for (int i = 0; i < sampleCount; i++)
        {
            double yPred = A * x[i] + B;
            double error = yPred - y[i];
            currentError += Math.Abs(error);
            // 符号函数近似梯度
            double sign = error > 0 ? 1 : error < 0 ? -1 : 0;
            gradA += x[i] * sign;
            gradB += sign;
        }

        gradA /= sampleCount;
        gradB /= sampleCount;
        currentError /= sampleCount;

        A -= _learningRate * gradA;
        B -= _learningRate * gradB;

        if (Math.Abs(currentError - lastError) < _tolerance)
            break;

        lastError = currentError;
    }
}

关键注意事项

  • 学习率调整:过大会导致参数震荡不收敛,过小会让训练速度极慢,建议从0.001、0.01、0.1等数值开始测试
  • 数据归一化:如果x的取值范围过大(比如0-1000),建议先将x归一化到[0,1]或[-1,1]区间,避免梯度爆炸或训练效率低下
  • 收敛判断:可以同时使用误差变化阈值和最大迭代次数,防止模型陷入无限迭代

内容的提问来源于stack exchange,提问作者JAAHARA-SEI XS TUMALUNGWA

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 20:51:27