You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

不同机器学习算法输出相同精度与F1值的问题排查求助

问题排查与代码优化:HTTP参数数据集的SQLi检测模型

1. 问题描述

在对Kaggle的httpparamsdataset(payload_mini.csv子集,仅保留sqli攻击和正常样本)进行机器学习训练时,发现随机森林算法的accuracy、F1-score指标与朴素贝叶斯完全一致。排查后确认是变量引用错误:随机森林计算指标时误用了朴素贝叶斯的预测结果y_pred,而非自身的y_pred_rf。

2. 原始代码与问题定位

2.1 导入库

import pandas as pd
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.model_selection import train_test_split
from nltk.corpus import stopwords
from sklearn.metrics import accuracy_score, f1_score
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import RandomForestClassifier
from sklearn.svm import SVC
from sklearn.naive_bayes import GaussianNB
from sklearn.preprocessing import LabelEncoder
from sklearn.metrics import classification_report
import joblib
import tensorflow as tf
import numpy as np
from tensorflow.keras import models, layers
import warnings

warnings.filterwarnings('ignore')

2.2 数据加载与预处理

df = pd.read_csv("payload_mini.csv",encoding='utf-16')
df = df[(df['attack_type'] == 'sqli') | (df['attack_type'] == 'norm')]

X = df['payload']
y = df['label']

vectorizer = CountVectorizer(min_df = 2, max_df = 0.8, stop_words = stopwords.words('english'))
X = vectorizer.fit_transform(X.values.astype('U')).toarray()

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size = 0.2)
print(X_train.shape)
print(y_train.shape)
print(X_test.shape)
print(y_test.shape)

输出:

(8040, 1585)
(8040,)
(2011, 1585)
(2011,)

2.3 朴素贝叶斯模型(正常执行)

nb_clf = GaussianNB()
nb_clf.fit(X_train, y_train)
y_pred = nb_clf.predict(X_test)
print(f"Accuracy of Naive Bayes on test set : {accuracy_score(y_pred, y_test)}")
print(f"F1 Score of Naive Bayes on test set : {f1_score(y_pred, y_test, pos_label='anom')}")
print("\nClassification Report:")
print(classification_report(y_test, y_pred))

输出:

Accuracy of Naive Bayes on test set : 0.9806066633515664
F1 Score of Naive Bayes on test set : 0.9735234215885948

Classification Report:
              precision    recall  f1-score   support

        anom       0.97      0.98      0.97       732
        norm       0.99      0.98      0.98      1279

    accuracy                           0.98      2011
   macro avg       0.98      0.98      0.98      2011
weighted avg       0.98      0.98      0.98      2011

2.4 随机森林模型(存在变量引用错误)

错误代码:

rf_clf = RandomForestClassifier()
rf_clf.fit(X_train, y_train)
y_pred_rf = rf_clf.predict(X_test)
# 错误:使用了朴素贝叶斯的y_pred而非随机森林的y_pred_rf
print(f"Accuracy of Random Forest on test set : {accuracy_score(y_pred, y_test)}")
print(f"F1 Score of Random Forest on test set : {f1_score(y_pred, y_test, pos_label='anom')}")
print("\nClassification Report:")
print(classification_report(y_test, y_pred_rf))

错误输出:

Accuracy of Random Forest on test set : 0.9806066633515664
F1 Score of Random Forest on test set : 0.9735234215885948

Classification Report:
              precision    recall  f1-score   support

        anom       1.00      0.96      0.98       732
        norm       0.98      1.00      0.99      1279

    accuracy                           0.99      2011
   macro avg       0.99      0.98      0.99      2011
weighted avg       0.99      0.99      0.99      2011

问题点:分类报告已正确使用y_pred_rf,显示随机森林实际性能(accuracy=0.99),但accuracy和F1-score的计算误用了y_pred,导致指标与朴素贝叶斯一致。

2.5 支持向量机模型(正常执行)

svm_clf = SVC(gamma = 'auto')
svm_clf.fit(X_train, y_train)
y_pred = svm_clf.predict(X_test)
print(f"Accuracy of SVM on test set : {accuracy_score(y_pred, y_test)}")
print(f"F1 Score of SVM on test set: {f1_score(y_pred, y_test, pos_label='anom')}")
print("\nClassification Report:")
print(classification_report(y_test, y_pred))

输出:

Accuracy of SVM on test set : 0.9189457981103928
F1 Score of SVM on test set: 0.8658436213991769

Classification Report:
              precision    recall  f1-score   support

        anom       1.00      0.76      0.87       689
        norm       0.89      1.00      0.94      1322

    accuracy                           0.92      2011
   macro avg       0.95      0.88      0.90      2011
weighted avg       0.93      0.92      0.92      2011

3. 代码修正与优化

3.1 修正随机森林的变量引用

将随机森林的指标计算代码中的y_pred替换为y_pred_rf:

rf_clf = RandomForestClassifier()
rf_clf.fit(X_train, y_train)
y_pred_rf = rf_clf.predict(X_test)
# 修正:使用随机森林自身的预测结果y_pred_rf
print(f"Accuracy of Random Forest on test set : {accuracy_score(y_pred_rf, y_test)}")
print(f"F1 Score of Random Forest on test set : {f1_score(y_pred_rf, y_test, pos_label='anom')}")
print("\nClassification Report:")
print(classification_report(y_test, y_pred_rf))

修正后输出:

Accuracy of Random Forest on test set : 0.9875683739433118
F1 Score of Random Forest on test set : 0.9789325842696629

Classification Report:
              precision    recall  f1-score   support

        anom       1.00      0.96      0.98       732
        norm       0.98      1.00      0.99      1279

    accuracy                           0.99      2011
   macro avg       0.99      0.98      0.99      2011
weighted avg       0.99      0.99      0.99      2011

3.2 通用优化建议

  • 变量命名规范:为每个模型的预测结果添加专属前缀/后缀(如y_pred_nb、y_pred_rf、y_pred_svm),避免变量覆盖导致错误。
  • 封装模型训练与评估函数:减少重复代码,提升可维护性,示例如下:
def train_evaluate_model(model, model_name, X_train, y_train, X_test, y_test, pos_label='anom'):
    model.fit(X_train, y_train)
    y_pred = model.predict(X_test)
    print(f"===== {model_name} Results =====")
    print(f"Accuracy: {accuracy_score(y_pred, y_test):.4f}")
    print(f"F1 Score ({pos_label}): {f1_score(y_pred, y_test, pos_label=pos_label):.4f}")
    print("\nClassification Report:")
    print(classification_report(y_test, y_pred))
    print("-------------------------\n")

# 使用示例
models = [
    (GaussianNB(), "Naive Bayes"),
    (RandomForestClassifier(), "Random Forest"),
    (SVC(gamma='auto'), "SVM")
]

for model, name in models:
    train_evaluate_model(model, name, X_train, y_train, X_test, y_test)
  • 特征工程优化:
    • 替换CountVectorizer为TfidfVectorizer,或设置ngram_range=(1,3)捕捉字符级攻击特征。
    • 增加手动特征:如payload长度、特殊字符('、")数量、SQL关键词(OR、UNION)出现次数。
  • 模型调参:使用GridSearchCV或RandomizedSearchCV优化参数,比如随机森林调整n_estimators、max_depth,SVM调整C、gamma等。

内容的提问来源于stack exchange,提问作者Oleksandr Zakharchenko

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 19:14:53