You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Linear SVC的系数而非支持向量导出至PMML?

如何导出含线性SVC系数的PMML并输出原始得分

核心问题原因

sklearn的SVC(kernel="linear")在sklearn2pmml中默认会以支持向量机(Support Vector Machine)的PMML类型导出,仅存储支持向量而非线性系数;而线性SVM的决策函数本质是w·x + b(系数与特征的点积加截距),要获取系数需让PMML以线性分类模型的形式存储。

解决方案1:替换为LinearSVC(推荐)

LinearSVC是sklearn专为线性分类优化的模型,训练后直接暴露coef_(系数)和intercept_(截距)属性,sklearn2pmml会将其导出为PMML的RegressionModel类型,自然包含系数数据。同时可配置输出原始得分而非仅分类结果。

代码示例

from sklearn.pipeline import Pipeline
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler
from sklearn.svm import LinearSVC
from sklearn2pmml import make_pmml_pipeline, sklearn2pmml

# 构建包含预处理和LinearSVC的Pipeline
pipeline = Pipeline([
    ("imputer", SimpleImputer()),
    ("scaler", StandardScaler()),
    # 调整参数对齐SVC(kernel='linear')的行为,按需设置loss、dual等
    ("classifier", LinearSVC(loss="hinge", dual=True, random_state=42))
])

# 训练模型
pipeline.fit(X_train, y_train)

# 导出PMML,指定输出原始得分
pmml_pipeline = make_pmml_pipeline(
    pipeline,
    active_fields=X_train.columns.tolist(),
    target_fields=["y"],
    output_fields=[
        {"name": "predicted_class", "optype": "categorical", "dataType": "integer"},
        {"name": "raw_score", "optype": "continuous", "dataType": "double", "feature": "rawScore"}
    ]
)

sklearn2pmml(pmml_pipeline, "LinearSVC_Pipeline.pmml")

解决方案2:基于SVC(kernel='linear')的系数构建等效线性模型

若必须使用SVC(kernel='linear'),可在训练后提取其coef_和intercept_,用自定义线性分类器替换Pipeline中的SVC,再导出PMML。

代码示例

from sklearn.pipeline import Pipeline
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler
from sklearn.svm import SVC
from sklearn.base import BaseEstimator, ClassifierMixin
import numpy as np
from sklearn2pmml import make_pmml_pipeline, sklearn2pmml

# 自定义线性分类器,复用SVC的系数
class LinearSVCWrapper(BaseEstimator, ClassifierMixin):
    def __init__(self, svc_model):
        self.coef_ = svc_model.coef_
        self.intercept_ = svc_model.intercept_
    
    def decision_function(self, X):
        # 计算原始得分:w·x + b
        return X @ self.coef_.T + self.intercept_
    
    def predict(self, X):
        scores = self.decision_function(X)
        return np.where(scores >= 0, 1, 0)  # 二分类场景,多分类需调整逻辑

# 先训练原始SVC模型
svc = SVC(kernel="linear", random_state=42)
pipeline = Pipeline([
    ("imputer", SimpleImputer()),
    ("scaler", StandardScaler()),
    ("classifier", svc)
])
pipeline.fit(X_train, y_train)

# 替换为自定义线性模型
pipeline[-1] = LinearSVCWrapper(svc)

# 导出含系数和原始得分的PMML
pmml_pipeline = make_pmml_pipeline(
    pipeline,
    active_fields=X_train.columns.tolist(),
    target_fields=["y"],
    output_fields=[
        {"name": "predicted_class", "optype": "categorical", "dataType": "integer"},
        {"name": "raw_score", "optype": "continuous", "dataType": "double", "feature": "rawScore"}
    ]
)

sklearn2pmml(pmml_pipeline, "SVC_Linear_With_Coefs.pmml")

关键注意事项

  • LinearSVC与SVC(kernel='linear')的优化目标略有差异,需调整参数(如loss、dual)保证模型性能一致。
  • 多分类场景下,coef_为(n_classes, n_features)数组,自定义模型的decision_function需返回每个类别的得分,PMML会自动适配多分类输出。
  • 导出的PMML文件中,系数存储在RegressionModel的RegressionTable下,可直接提取用于本地计算特征与系数的点积得分。

内容的提问来源于stack exchange,提问作者Hannah Oosterhuis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 11:10:31