不同机器学习算法输出相同精度与F1值的问题排查求助
问题排查与代码优化:HTTP参数数据集的SQLi检测模型
1. 问题描述
在对Kaggle的httpparamsdataset(payload_mini.csv子集,仅保留sqli攻击和正常样本)进行机器学习训练时,发现随机森林算法的accuracy、F1-score指标与朴素贝叶斯完全一致。排查后确认是变量引用错误:随机森林计算指标时误用了朴素贝叶斯的预测结果y_pred,而非自身的y_pred_rf。
2. 原始代码与问题定位
2.1 导入库
import pandas as pd from sklearn.feature_extraction.text import CountVectorizer from sklearn.model_selection import train_test_split from nltk.corpus import stopwords from sklearn.metrics import accuracy_score, f1_score from sklearn.linear_model import LogisticRegression from sklearn.ensemble import RandomForestClassifier from sklearn.svm import SVC from sklearn.naive_bayes import GaussianNB from sklearn.preprocessing import LabelEncoder from sklearn.metrics import classification_report import joblib import tensorflow as tf import numpy as np from tensorflow.keras import models, layers import warnings warnings.filterwarnings('ignore')
2.2 数据加载与预处理
df = pd.read_csv("payload_mini.csv",encoding='utf-16') df = df[(df['attack_type'] == 'sqli') | (df['attack_type'] == 'norm')] X = df['payload'] y = df['label'] vectorizer = CountVectorizer(min_df = 2, max_df = 0.8, stop_words = stopwords.words('english')) X = vectorizer.fit_transform(X.values.astype('U')).toarray() X_train, X_test, y_train, y_test = train_test_split(X, y, test_size = 0.2) print(X_train.shape) print(y_train.shape) print(X_test.shape) print(y_test.shape)
输出:
(8040, 1585) (8040,) (2011, 1585) (2011,)
2.3 朴素贝叶斯模型(正常执行)
nb_clf = GaussianNB() nb_clf.fit(X_train, y_train) y_pred = nb_clf.predict(X_test) print(f"Accuracy of Naive Bayes on test set : {accuracy_score(y_pred, y_test)}") print(f"F1 Score of Naive Bayes on test set : {f1_score(y_pred, y_test, pos_label='anom')}") print("\nClassification Report:") print(classification_report(y_test, y_pred))
输出:
Accuracy of Naive Bayes on test set : 0.9806066633515664 F1 Score of Naive Bayes on test set : 0.9735234215885948 Classification Report: precision recall f1-score support anom 0.97 0.98 0.97 732 norm 0.99 0.98 0.98 1279 accuracy 0.98 2011 macro avg 0.98 0.98 0.98 2011 weighted avg 0.98 0.98 0.98 2011
2.4 随机森林模型(存在变量引用错误)
错误代码:
rf_clf = RandomForestClassifier() rf_clf.fit(X_train, y_train) y_pred_rf = rf_clf.predict(X_test) # 错误:使用了朴素贝叶斯的y_pred而非随机森林的y_pred_rf print(f"Accuracy of Random Forest on test set : {accuracy_score(y_pred, y_test)}") print(f"F1 Score of Random Forest on test set : {f1_score(y_pred, y_test, pos_label='anom')}") print("\nClassification Report:") print(classification_report(y_test, y_pred_rf))
错误输出:
Accuracy of Random Forest on test set : 0.9806066633515664 F1 Score of Random Forest on test set : 0.9735234215885948 Classification Report: precision recall f1-score support anom 1.00 0.96 0.98 732 norm 0.98 1.00 0.99 1279 accuracy 0.99 2011 macro avg 0.99 0.98 0.99 2011 weighted avg 0.99 0.99 0.99 2011
问题点:分类报告已正确使用
y_pred_rf,显示随机森林实际性能(accuracy=0.99),但accuracy和F1-score的计算误用了y_pred,导致指标与朴素贝叶斯一致。
2.5 支持向量机模型(正常执行)
svm_clf = SVC(gamma = 'auto') svm_clf.fit(X_train, y_train) y_pred = svm_clf.predict(X_test) print(f"Accuracy of SVM on test set : {accuracy_score(y_pred, y_test)}") print(f"F1 Score of SVM on test set: {f1_score(y_pred, y_test, pos_label='anom')}") print("\nClassification Report:") print(classification_report(y_test, y_pred))
输出:
Accuracy of SVM on test set : 0.9189457981103928 F1 Score of SVM on test set: 0.8658436213991769 Classification Report: precision recall f1-score support anom 1.00 0.76 0.87 689 norm 0.89 1.00 0.94 1322 accuracy 0.92 2011 macro avg 0.95 0.88 0.90 2011 weighted avg 0.93 0.92 0.92 2011
3. 代码修正与优化
3.1 修正随机森林的变量引用
将随机森林的指标计算代码中的y_pred替换为y_pred_rf:
rf_clf = RandomForestClassifier() rf_clf.fit(X_train, y_train) y_pred_rf = rf_clf.predict(X_test) # 修正:使用随机森林自身的预测结果y_pred_rf print(f"Accuracy of Random Forest on test set : {accuracy_score(y_pred_rf, y_test)}") print(f"F1 Score of Random Forest on test set : {f1_score(y_pred_rf, y_test, pos_label='anom')}") print("\nClassification Report:") print(classification_report(y_test, y_pred_rf))
修正后输出:
Accuracy of Random Forest on test set : 0.9875683739433118 F1 Score of Random Forest on test set : 0.9789325842696629 Classification Report: precision recall f1-score support anom 1.00 0.96 0.98 732 norm 0.98 1.00 0.99 1279 accuracy 0.99 2011 macro avg 0.99 0.98 0.99 2011 weighted avg 0.99 0.99 0.99 2011
3.2 通用优化建议
- 变量命名规范:为每个模型的预测结果添加专属前缀/后缀(如
y_pred_nb、y_pred_rf、y_pred_svm),避免变量覆盖导致错误。 - 封装模型训练与评估函数:减少重复代码,提升可维护性,示例如下:
def train_evaluate_model(model, model_name, X_train, y_train, X_test, y_test, pos_label='anom'): model.fit(X_train, y_train) y_pred = model.predict(X_test) print(f"===== {model_name} Results =====") print(f"Accuracy: {accuracy_score(y_pred, y_test):.4f}") print(f"F1 Score ({pos_label}): {f1_score(y_pred, y_test, pos_label=pos_label):.4f}") print("\nClassification Report:") print(classification_report(y_test, y_pred)) print("-------------------------\n") # 使用示例 models = [ (GaussianNB(), "Naive Bayes"), (RandomForestClassifier(), "Random Forest"), (SVC(gamma='auto'), "SVM") ] for model, name in models: train_evaluate_model(model, name, X_train, y_train, X_test, y_test)
- 特征工程优化:
- 替换
CountVectorizer为TfidfVectorizer,或设置ngram_range=(1,3)捕捉字符级攻击特征。 - 增加手动特征:如payload长度、特殊字符(
'、")数量、SQL关键词(OR、UNION)出现次数。
- 替换
- 模型调参:使用
GridSearchCV或RandomizedSearchCV优化参数,比如随机森林调整n_estimators、max_depth,SVM调整C、gamma等。
内容的提问来源于stack exchange,提问作者Oleksandr Zakharchenko
相关产品推荐
相关产品推荐

