You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何绘制Logistic Regression模型的Sigmoid函数及参数设置

问题描述

处理成人收入数据集,训练逻辑回归模型后准确率达0.8051589,尝试绘制Sigmoid函数以理解模型,但无法确定绘制代码中start、end、num_points参数的取值,且现有绘制代码存在逻辑问题。

训练代码:

import pandas as pd
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
import numpy as np
import matplotlib.pyplot as plt

df = pd.read_csv('adultdata_encoded.csv')

# 特征矩阵与目标变量
X = df[['education-num', 'workclass_encoded', 'hours-per-week', 'sex_encoded', 'relationship_encoded', 'occupation_encoded', 'maritalstatus_encoded', 'race_encoded', 'nativecountry_encoded']].values
y = df['income_encoded'].values

# 划分训练集与测试集
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

# 训练逻辑回归模型
logreg = LogisticRegression()
logreg.fit(X_train, y_train)

# 预测与计算准确率
y_pred = logreg.predict(X_test)
accuracy = accuracy_score(y_test, y_pred)
for i in range(0, len(y_pred)):
    print(y_pred[i], y_test[i])
print("Accuracy:", accuracy)

尝试的绘制代码(存在逻辑问题):

coef = logreg.coef_[0]
intercept = logreg.intercept_

# 定义sigmoid函数
def sigmoid(x):
    z = np.dot(coef, x) + intercept
    return 1 / (1 + np.exp(-z))


# 生成输入值范围
start = 0  # 起始点待定义
end = 1  # 终止点待定义
num_points = 10  # 点数待定义
x = np.linspace(start, end, num_points)

# 计算预测概率
y = sigmoid(x)

# 绘图
plt.plot(x, y)
plt.xlabel('Input')
plt.ylabel('Probability')
plt.title('Sigmoid Function')
plt.grid(True)
plt.show()
解决方案

问题核心

你的模型是多特征逻辑回归,Sigmoid函数的输入是所有特征的线性组合 z = coef·X + intercept,而非单个特征值。直接用单变量x传入sigmoid会导致维度不匹配,同时参数取值也无法对应实际特征分布。

以下提供两种实用的绘制方案:


方案1:针对单个特征绘制Sigmoid曲线(固定其他特征)

选择一个你关注的特征(比如education-num),将其他特征固定为训练集的均值/中位数,生成该特征的合理取值范围,观察其对预测概率的影响。

参数取值规则:

  • start:取该特征在训练集中的最小值
  • end:取该特征在训练集中的最大值
  • num_points:设为100~200,保证曲线平滑

修正代码:

coef = logreg.coef_[0]
intercept = logreg.intercept_[0]

# 选择关注的特征索引(比如0对应education-num)
target_feature_idx = 0
feature_name = df.columns[target_feature_idx]

# 计算其他特征的均值(固定为均值)
other_features_mean = X_train.mean(axis=0)
other_features_mean = other_features_mean.reshape(1, -1)

# 定义适配多特征的sigmoid计算函数
def sigmoid_single_feature(feature_values, other_features_mean, coef, intercept):
    # 复制固定特征值,替换目标特征的取值
    X = np.tile(other_features_mean, (len(feature_values), 1))
    X[:, target_feature_idx] = feature_values
    z = np.dot(X, coef) + intercept
    return 1 / (1 + np.exp(-z))

# 生成目标特征的取值范围
start = X_train[:, target_feature_idx].min()
end = X_train[:, target_feature_idx].max()
num_points = 150
x_values = np.linspace(start, end, num_points)

# 计算对应的概率
y_probs = sigmoid_single_feature(x_values, other_features_mean, coef, intercept)

# 绘图
plt.plot(x_values, y_probs)
plt.xlabel(f'Feature: {feature_name}')
plt.ylabel('Probability of income >50K')
plt.title(f'Sigmoid Curve for {feature_name} (Other Features Fixed to Mean)')
plt.grid(True)
plt.show()

方案2:绘制标准Sigmoid函数(理解函数形状)

如果只是想理解Sigmoid函数本身的形状,直接针对线性组合值z绘制即可,无需关联具体特征:

参数取值规则:

  • start:设为-10(Sigmoid在z<-10时趋近于0)
  • end:设为10(Sigmoid在z>10时趋近于1)
  • num_points:设为100,保证曲线平滑

代码:

# 标准Sigmoid函数
def sigmoid(z):
    return 1 / (1 + np.exp(-z))

# 生成z的取值范围
start = -10
end = 10
num_points = 100
z_values = np.linspace(start, end, num_points)

# 计算对应的概率
y_probs = sigmoid(z_values)

# 绘图
plt.plot(z_values, y_probs)
plt.xlabel('Linear Combination z = coef·X + intercept')
plt.ylabel('Probability')
plt.title('Standard Sigmoid Function')
plt.grid(True)
plt.axvline(x=0, color='r', linestyle='--', label='z=0 (Probability=0.5)')
plt.legend()
plt.show()

内容的提问来源于stack exchange,提问作者cvbalbas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 14:45:13