You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Gaussian Naive Bayes交叉验证结果异常,求问题排查

问题

我用sklearn实现了基础版高斯贝叶斯模型,但K折交叉验证结果异常:

  • 当测试集占比为0.2时,10折CV的首折准确率仅约0.367,后续折准确率大幅提升至接近1.0;
  • 当测试集占比调整为0.6时,10折CV准确率全部为1.0。

以下是代码实现:

import numpy as np
import pandas as pd
from sklearn.naive_bayes import GaussianNB
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score, confusion_matrix, recall_score, precision_score, classification_report
import csv
from sklearn.model_selection import cross_val_score

column_names = ['AS', 'fh', 'class2']
df = pd.read_csv("C:/Users/Jans/Music/docx/222/test.csv",  sep=';', header = 0, names = column_names)

x = df.drop(['AS', 'class2'], axis=1)
df['class2'] = df['class2'].astype(int)
y = df['class2'].values

x_train, x_test, y_train, y_test = train_test_split(x, y, test_size=0.2, shuffle = False, random_state = None)

model = GaussianNB()
model.fit(x_train, y_train.astype('int'))

k_fold_acc = cross_val_score(model, x_train, y_train, cv=10)
k_fold_mean = k_fold_acc.mean()
for i in k_fold_acc:
    print(i)
print("accuracy K Fold CV:" + str(k_fold_mean))

grid_predictions = model.predict(x_test)

测试集占比0.2时的10折CV结果:

0.36714285714285716
0.8271428571428572
0.9785714285714285
0.9357142857142857
0.9628571428571429
0.9957081545064378
1.0
1.0
0.994277539341917
0.9842632331902719
accuracy K Fold CV:0.90456774984672

测试集占比调整为0.6时的10折CV结果:

1.0
1.0
1.0
1.0
1.0
1.0
1.0
1.0
1.0
1.0
accuracy K Fold CV:1.0

问题分析与解决

核心问题

  1. 数据未打乱导致分布偏差:你在train_test_split中设置了shuffle=False,数据按原始顺序直接拆分。如果你的数据集是按标签排序的(比如前半段全是类别0,后半段全是类别1),拆分后的训练集会出现严重的类别分布失衡:

    • 当test_size=0.2时,训练集的前1/10折可能只包含少数类样本,模型无法学习到有效模式,导致准确率极低;后续折覆盖了足够多的目标类别,准确率自然飙升。
    • 当test_size=0.6时,训练集刚好包含了所有类别的样本且分布均匀,因此每折CV都能得到满分。
  2. 交叉验证逻辑冗余:先拆分训练集再做CV的意义不大,且会放大数据未打乱的负面影响。直接对全数据集做CV,或先打乱再拆分,结果会更可靠。

修正方案

  • 强制开启数据打乱:在train_test_split或cross_val_score中设置shuffle=True,确保每个子集的类别分布与整体一致。
  • 优化CV逻辑:要么直接用全数据集做交叉验证,要么先打乱拆分后再对训练集做CV。

修正后的代码

import numpy as np
import pandas as pd
from sklearn.naive_bayes import GaussianNB
from sklearn.model_selection import train_test_split, cross_val_score
from sklearn.metrics import accuracy_score

column_names = ['AS', 'fh', 'class2']
df = pd.read_csv("C:/Users/Jans/Music/docx/222/test.csv", sep=';', header=0, names=column_names)

# 数据预处理
df['class2'] = df['class2'].astype(int)
x = df.drop(['AS', 'class2'], axis=1)
y = df['class2'].values

# 方案1:打乱后拆分,再对训练集做CV
x_train, x_test, y_train, y_test = train_test_split(x, y, test_size=0.2, shuffle=True, random_state=42)
model = GaussianNB()
# 交叉验证时也可设置shuffle=True进一步保证分布一致
k_fold_acc = cross_val_score(model, x_train, y_train, cv=10, shuffle=True, random_state=42)

print("训练集10折CV准确率:")
for acc in k_fold_acc:
    print(f"{acc:.4f}")
print(f"平均准确率:{k_fold_acc.mean():.4f}")

# 测试集评估
model.fit(x_train, y_train)
test_acc = accuracy_score(y_test, model.predict(x_test))
print(f"测试集准确率:{test_acc:.4f}")

# 方案2:直接对全数据集做CV(更简洁)
# model = GaussianNB()
# k_fold_acc = cross_val_score(model, x, y, cv=10, shuffle=True, random_state=42)
# print("全数据集10折CV准确率:")
# for acc in k_fold_acc:
#     print(f"{acc:.4f}")
# print(f"平均准确率:{k_fold_acc.mean():.4f}")

额外检查建议

  • 确认原始数据集是否按标签排序:执行print(df['class2'].head(30))查看前30行标签分布,若呈现明显的有序性,说明数据未打乱是核心诱因。

内容的提问来源于stack exchange,提问作者Questions123

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 04:44:56