You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

自定义Logistic Regression预测时遇TypeError问题求助

解决Logistic Regression实现中的TypeError: float与NoneType相乘问题

错误根源

你的train函数存在严重的缩进逻辑错误,导致函数返回None,最终在predict时触发类型错误:

  • for循环仅执行了最后一次迭代的y_pred和error计算,梯度更新、权重调整代码都在循环外部,完全没有实现迭代训练。
  • return w被放在if epoch % 100 == 0的分支中,当训练到最后一个epoch(比如1000轮的最后一轮是999),不满足模100等于0的条件,函数没有返回值,默认返回None。后续调用predict时,w是None,执行np.dot(X, w)就会触发float和NoneType相乘的错误。

修正后的完整代码

以下是修正了缩进问题并优化部分细节的代码:

import numpy as np
import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score, classification_report, confusion_matrix

# 修正函数名拼写:load_cvs -> load_csv
def load_csv(filename):
    data = []
    labels = []
    with open(filename,  'r') as f:
        for line in f:
            # 处理可能的换行符
            line = line.strip()
            if not line:
                continue
            items = line.split(",")
            data.append([float(items[0]),float(items[1]),float(items[2]),float(items[3]),float(items[4]),float(items[5])])
            labels.append(int(items[6]))
    return np.array(data), np.array(labels)

# 数据加载与预处理
col_names = ['Glucose','BloodPressure','SkinThickness','Insulin','BMI','Age','Outcome']
df = pd.read_csv('diabetes.csv', header=None, names=col_names)
# 注意:用0填充缺失值可能影响模型效果,建议根据数据特性处理,比如用均值/中位数
nandf = df.fillna(value=0)

feature_cols = ['Glucose','BloodPressure','SkinThickness','Insulin','BMI','Age']
X = nandf[feature_cols]
y = nandf.Outcome

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=16)

def sigmoid(z):
     return 1 / (1 + np.exp(-z))

def predict(X, w):
    z = np.dot(X, w)
    return sigmoid(z)

def cost(y, y_pred):
    # 添加微小值避免log(0)的警告/错误
    epsilon = 1e-10
    y_pred = np.clip(y_pred, epsilon, 1 - epsilon)
    return -np.mean(y * np.log(y_pred) + (1 - y) * np.log(1 - y_pred))

def train(X, y, epochs=1000, lr=0.02):
    # 添加偏置项
    X = np.c_[np.ones((X.shape[0], 1)), X]
    # 初始化权重
    w = np.random.randn(X.shape[1])
    
    for epoch in range(epochs):
        y_pred = predict(X, w)
        error = y_pred - y
        # 计算梯度并更新权重(缩进进循环内部)
        gradient = np.dot(X.T, error) / y.size
        w -= lr * gradient
        
        # 每100轮打印一次损失(缩进进循环内部)
        if epoch % 100 == 0:
            c = cost(y, y_pred)
            print(f'Epoch {epoch}: cost = {c:.4f}')
    
    # 训练完成后返回最终权重(移到循环外部)
    return w

# 训练模型
w = train(X_train, y_train, epochs=1000, lr=0.02)

# 生成测试集预测
X_test_with_bias = np.c_[np.ones((X_test.shape[0], 1)), X_test]
y_pred_proba = predict(X_test_with_bias, w)
# 转换为分类标签(概率>0.5为1,否则为0)
y_pred = (y_pred_proba > 0.5).astype(int)

# 评估模型
print("\n模型评估结果:")
print(f"准确率: {accuracy_score(y_test, y_pred):.4f}")
print("\n分类报告:")
print(classification_report(y_test, y_pred))
print("\n混淆矩阵:")
print(confusion_matrix(y_test, y_pred))

关键修正点说明

  1. 修复缩进问题:将梯度计算、权重更新、损失打印的代码全部缩进至for循环内部,确保每轮epoch都执行训练逻辑。
  2. 调整return位置:将return w移到for循环外部,保证训练完成后返回最终训练好的权重,不会返回None。
  3. 避免log(0)错误:在cost函数中添加了np.clip和微小值epsilon,防止计算np.log(0)时出现数值错误。
  4. 优化数据加载:修正了函数名拼写错误,添加了空行过滤逻辑,避免读取空行导致的错误。
  5. 补充模型评估:添加了分类标签转换和模型评估代码,方便验证模型效果。

内容的提问来源于stack exchange,提问作者Abood 4433

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 05:30:50