You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Jupyter Notebook中Logistic Regression训练出现NaN/inf值报错求助

问题根因
  • np.nan_to_num() 属于返回新对象的非inplace操作,你没有将处理后的结果赋值回X,原始特征矩阵中的空值、无穷值没有被实际修改,是触发报错的直接原因。
  • 特征、标签处理逻辑完全混乱:
    1. 独热编码时将标签列Relationship也放入特征编码逻辑,存在严重特征泄露,训练出的模型没有预测价值
    2. 你同时生成了独热编码后的X,又在训练测试集拆分时调用了原始未编码的特征列,最后拟合又混用了独热编码后的X,逻辑完全不自洽
    3. 二分类场景下不需要对标签y做独热编码,直接使用原始的0/1标签列即可
  • 代码未导入numpy库就直接调用np开头的方法,存在语法错误
修正后可运行代码
# 导入依赖
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
%matplotlib inline
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split

# 读取数据
df = pd.read_csv('boo.csv')

# 拆分特征和标签,标签直接用原始的0/1列
y = df['Relationship']
X = df[['Month','Year','Victim Age', 'Perpetrator Age', 'Victim Sex', 'Victim Race', 'Perpetrator Sex', 'Crime Type', 'Perpetrator Race']]

# 对特征做独热编码,不要处理标签列
X = pd.get_dummies(X, drop_first=True)

# 处理空值、无穷值,注意要赋值回X
X = np.nan_to_num(X, nan=0, posinf=0, neginf=0)

# 拆分训练测试集
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.1, random_state=42)

# 训练模型,指定solver避免收敛警告
model = LogisticRegression(solver='liblinear', max_iter=1000)
model.fit(X_train, y_train)

# 可选:查看测试集准确率
print(f"测试集准确率:{model.score(X_test, y_test):.2f}")
额外优化建议
  • 空值处理建议不要直接填充0,可以根据特征含义选择中位数、众数填充,避免引入数据偏差
  • 数值类特征Victim Age、Perpetrator Age建议做标准化处理,提升逻辑回归的收敛速度和效果
  • 如果标签存在类别不平衡,建议加入class_weight参数调整权重

内容的提问来源于stack exchange,提问作者user16543065

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 07:42:03