You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于sklearn的二维数组多项式回归系数求解问题

二维特征下多项式拟合问题的解决方法

你的核心问题是:默认的PolynomialFeatures(2)会生成所有特征的交叉项(如x₁x₂、x₁x₃等),但你的目标模型仅包含每个特征自身的一次项和二次项,没有交叉项,这导致拟合后的系数矩阵包含大量冗余项,无法直接对应你需要的a、b、c。同时你合成的y是多输出数组(n_samples, n_features),需要匹配对应的建模逻辑。

问题分析

在一维场景中,PolynomialFeatures(2)生成的特征是[1, x, x²],刚好匹配你的模型Y=a+bX+cX²;但二维场景下,它会生成所有可能的组合项,包括交叉项,特征数量从n_features暴涨到1 + n_features + n_features*(n_features+1)/2,完全偏离了你需要的模型结构。

解决方案

方法一:对每个特征单独拟合(最直观,和一维逻辑一致)

循环处理每个输入-输出特征对,单独拟合二次模型,直接得到每个特征对应的a、b、c:

import numpy as np
from sklearn.preprocessing import PolynomialFeatures
from sklearn.linear_model import LinearRegression
from sklearn.pipeline import make_pipeline

n_subfeatures = 4
n_components_per_subf = 3
n_features = n_subfeatures * n_components_per_subf
n_samples = 100

np.random.seed(42)
x = np.random.randn(n_samples, n_features)
y = 0.3 + 1.7 * x - 0.5 * x**2

# 存储每个特征的系数(a, b, c)
coef_results = []
for idx in range(n_features):
    # 针对单个特征构建管道
    pipe = make_pipeline(PolynomialFeatures(2), LinearRegression())
    pipe.fit(x[:, idx:idx+1], y[:, idx])
    
    # 提取系数:PolynomialFeatures生成的特征顺序是[1, x, x²]
    a = pipe[1].intercept_
    b = pipe[1].coef_[0][1]
    c = pipe[1].coef_[0][2]
    coef_results.append((round(a,4), round(b,4), round(c,4)))
    print(f"特征{idx}: a={a:.4f}, b={b:.4f}, c={c:.4f}")

方法二:手动构造目标特征矩阵(一次拟合所有输出)

直接构造仅包含每个特征一次项和二次项的特征矩阵,避免交叉项,然后用多输出线性回归拟合:

import numpy as np
from sklearn.linear_model import LinearRegression

n_subfeatures = 4
n_components_per_subf = 3
n_features = n_subfeatures * n_components_per_subf
n_samples = 100

np.random.seed(42)
x = np.random.randn(n_samples, n_features)
y = 0.3 + 1.7 * x - 0.5 * x**2

# 构造特征矩阵:每个特征的一次项 + 二次项
X_poly = np.hstack([x, x**2])

# 拟合多输出线性回归(自动添加偏置项)
model = LinearRegression(fit_intercept=True)
model.fit(X_poly, y)

# 提取系数
a = model.intercept_.round(4)  # 每个输出的a值
b = model.coef_[:, :n_features].round(4)  # 每个输出对应的b值
c = model.coef_[:, n_features:].round(4)  # 每个输出对应的c值

print("所有特征的a系数:", a)
print("所有特征的b系数:\n", b)
print("所有特征的c系数:\n", c)

验证效果

两种方法都能精确还原你合成数据时设置的系数:a≈0.3,b≈1.7,c≈-0.5,完全匹配你的目标模型。

内容的提问来源于stack exchange,提问作者booo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.15 22:47:09