You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Isolation Forest训练葡萄酒数据集准确率为0.0的问题求助

问题描述

编辑说明:我正在学习如何发布优质问题,欢迎大家提出意见

我尝试使用IsolationForest()训练葡萄酒数据集,目标是将训练好的模型用于另一个质量调整后的数据集,预测其质量值并筛选出所有质量为8和9的葡萄酒。

但目前遇到了问题,分类报告显示准确率为0.0:

print(classification_report(y_test, prediction))

              precision    recall  f1-score   support

          -1       0.00      0.00      0.00       0.0
           1       0.00      0.00      0.00       0.0
           3       0.00      0.00      0.00     866.0
           4       0.00      0.00      0.00     829.0
           5       0.00      0.00      0.00     841.0
           6       0.00      0.00      0.00     861.0
           7       0.00      0.00      0.00     822.0
           8       0.00      0.00      0.00     886.0
           9       0.00      0.00      0.00     851.0

    accuracy                           0.00    5956.0
   macro avg       0.00      0.00      0.00    5956.0
weighted avg       0.00      0.00      0.00    5956.0

我不确定这是超参数问题、数据处理错误还是参数设置错误,已尝试使用SMOTE和不使用SMOTE的情况,期望至少达到90%的准确率。

我的代码如下:

from sklearn.preprocessing import OrdinalEncoder
from sklearn.ensemble import IsolationForest
import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from imblearn.over_sampling import SMOTE
from sklearn.metrics import classification_report,confusion_matrix

df = pd.read_csv('wines.csv')

df.head(5)

ordinalEncoder = OrdinalEncoder(handle_unknown='use_encoded_value', unknown_value=-99).fit(df[['color']])
df[['color']] = ordinalEncoder.transform(df[['color']])

df.info()

df['color'] = df['color'].astype(int)

df.head(3)

stm = SMOTE(k_neighbors=4)
x_smote = df.drop('quality',axis=1)
y_smote = df['quality']
x_smote,y_smote = stm.fit_resample(x_smote,y_smote)

print(x_smote.shape,y_smote.shape)

x_smote.columns

scaler = StandardScaler()
X = scaler.fit_transform(x_smote)
y = y_smote

X.shape, y.shape

x_train, x_test, y_train, y_test = train_test_split(X,y,test_size=0.3)

from sklearn.ensemble import IsolationForest
from sklearn.metrics import hamming_loss

iforest = IsolationForest(n_estimators=200, max_samples=0.1, contamination=0.10, max_features=1.0, bootstrap=False, n_jobs=-1, 
                            random_state=None, verbose=0, warm_start=False)

iforest_fit = iforest.fit(x_train,y_train)

prediction = iforest_fit.predict(x_test)

print (prediction.shape, y_test.shape)

y.value_counts()

prediction

print(confusion_matrix(y_test, prediction))
hamming_loss(y_test, prediction)

from sklearn.metrics import classification_report
print(classification_report(y_test, prediction))

问题根源与解决方案

1. 核心错误:用异常检测模型做分类任务

IsolationForest是异常检测算法,不是分类算法。它的predict()方法只会输出1(正常样本)或-1(异常样本),根本无法预测你的多分类标签(3-9)。这就是分类报告里3-9类全部准确率为0的原因——模型输出的只有1和-1,和真实标签完全不匹配。

2. 解决方案步骤

(1)换用分类模型

如果你要预测葡萄酒质量(多分类任务),应该用随机森林、XGBoost、LightGBM这类分类模型,替换掉IsolationForest:

from sklearn.ensemble import RandomForestClassifier

# 替换原IsolationForest训练代码
clf = RandomForestClassifier(n_estimators=200, n_jobs=-1, random_state=42)
clf.fit(x_train_scaled, y_train_smote)
prediction = clf.predict(x_test_scaled)

print(classification_report(y_test, prediction))

(2)修正SMOTE与数据缩放的逻辑

你当前先对全量数据做SMOTE再拆分训练测试集,会导致数据泄露。正确流程是:

  • 先拆分训练集和测试集
  • 只对训练集做SMOTE处理
  • 用训练集拟合的Scaler去转换测试集

修正后的数据处理代码:

# 先拆分数据集,避免数据泄露
x = df.drop('quality',axis=1)
y = df['quality']
x_train, x_test, y_train, y_test = train_test_split(x,y,test_size=0.3, random_state=42)

# 仅对训练集做SMOTE
stm = SMOTE(k_neighbors=4)
x_train_smote, y_train_smote = stm.fit_resample(x_train, y_train)

# 用训练集拟合Scaler,分别转换训练/测试集
scaler = StandardScaler()
x_train_scaled = scaler.fit_transform(x_train_smote)
x_test_scaled = scaler.transform(x_test)

(3)针对目标优化:筛选8和9类

如果最终目标是筛选质量8和9的葡萄酒,可以把任务转为二分类,让模型更聚焦目标类别:

# 转换标签为二分类:8/9标记为1,其他为0
y_train_bin = np.where(y_train_smote.isin([8,9]), 1, 0)
y_test_bin = np.where(y_test.isin([8,9]), 1, 0)

# 训练二分类模型
clf = RandomForestClassifier(n_estimators=200, n_jobs=-1, random_state=42)
clf.fit(x_train_scaled, y_train_bin)

# 用概率输出调整阈值,减少误判
y_proba = clf.predict_proba(x_test_scaled)[:,1]
prediction_bin = np.where(y_proba > 0.7, 1, 0)

print(classification_report(y_test_bin, prediction_bin))

3. 额外提示

葡萄酒质量数据集的特征与标签相关性不算极强,要达到90%准确率可能需要配合特征工程(比如组合特征、筛选重要特征);另外不要混淆异常检测和分类任务的适用场景,两者设计目标完全不同。

内容的提问来源于stack exchange,提问作者Gabriel Rodrigues

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 17:24:15