You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用StratifiedKFold交叉验证时遇布尔数组类型错误,如何解决?

解决StratifiedKFold交叉验证中"Boolean array expected for the condition, not float64"错误

尝试在数据集上使用StratifiedKFold进行交叉验证时,触发如下错误:

ValueError: Boolean array expected for the condition, not float64

原代码

import pandas as pd
import numpy as np
from sklearn.model_selection import StratifiedKFold
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score

from imblearn.over_sampling import SMOTE
cleanedDataset = `pd.read_csv("train_numeric_shuffled_50000_cleaned_90.csv")`

#providing input and output features
x=cleanedDataset.drop(['Id','Response'], axis=1)
y=cleanedDataset['Response']

#applico Stratified K-fold con K=4
skf = StratifiedKFold(n_splits=4)

#stampo risultati dei 4 fold
for i, (train_index, test_index) in enumerate(skf.split(x, y)):
        print(f"Fold {i}:")
        print(f"  Train: index={train_index}")
        print(f"  Test:  index={test_index}")

#uso la colonna response come Target
target = cleanedDataset.loc[:,'Response']

#definizione train model
model = LogisticRegression()
def train_model(train, test, fold_no):

    x_train = train[x]
    y_train = train[y]
    x_test = test[x]
    x_test = test[y]
    model.fit(X_train,y_train)
    predictions = model.predict(X_test)
    print('Fold',str(fold_no),'Accuracy:',accuracy_score(y_test,predictions))

#stampo valori accuratezza algoritmo
fold_no =1
for train_index, test_index in skf.split(cleanedDataset, target):
    train = cleanedDataset.loc[train_index,:]
    test = cleanedDataset.loc[test_index,:]
    train_model(train,test,fold_no)
    fold_no += 1

错误回溯信息

ValueError  Traceback (most recent call last)
    ~\AppData\Local\Temp\ipykernel_8004\1316530102.py in <module>
          4     train = cleanedDataset.loc[train_index,:]
          5     test = cleanedDataset.loc[test_index,:]
    ----> 6     train_model(train,test,fold_no)
          7     fold_no += 1
    
    ~\AppData\Local\Temp\ipykernel_8004\3643313375.py in train_model(train, test, fold_no)
          3 def train_model(train, test, fold_no):
          4 
    ----> 5     X_train = train[x]
          6     y_train = train[y]
          7     X_test = test[x]

~\anaconda3\lib\site-packages\pandas\core\frame.py in __getitem__(self, key)
   3490         # Do we have a (boolean) DataFrame?
   3491         if isinstance(key, DataFrame):
-> 3492             return self.where(key)
   3493 
   3494         # Do we have a (boolean) 1d indexer?

~\anaconda3\lib\site-packages\pandas\util\_decorators.py in wrapper(*args, **kwargs)
    309                     stacklevel=stacklevel,
    310                 )
-> 311             return func(*args, **kwargs)
    312 
    313         return wrapper

~\anaconda3\lib\site-packages\pandas\core\frame.py in where(self, cond, other, inplace, axis, level, errors, try_cast)
  10962         try_cast=lib.no_default,
  10963     ):
> 10964         return super().where(cond, other, inplace, axis, level, errors, try_cast)
  10965 
  10966     @deprecate_nonkeyword_arguments(

~\anaconda3\lib\site-packages\pandas\core\generic.py in where(self, cond, other, inplace, axis, level, errors, try_cast)
   9313             )
   9314 
-> 9315         return self._where(cond, other, inplace, axis, level, errors=errors)
   9316 
   9317     @doc(

~\anaconda3\lib\site-packages\pandas\core\generic.py in _where(self, cond, other, inplace, axis, level, errors)
   9074                 for dt in cond.dtypes:
   9075                     if not is_bool_dtype(dt):
-> 9076                         raise ValueError(msg.format(dtype=dt))
   9077         else:
   9078             # GH#21947 we have an empty DataFrame/Series, could be object-dtype

ValueError: Boolean array expected for the condition, not float64

错误根源与修正方案

1. 核心错误:用DataFrame/Series作为索引触发布尔筛选

原代码中x = cleanedDataset.drop(['Id','Response'], axis=1)得到的是DataFrame对象,train[x]会被pandas误判为布尔条件筛选逻辑——pandas会尝试将DataFrame作为布尔数组使用,但你的数据是float类型,因此抛出类型不匹配错误。

2. 其他代码问题

  • 变量名大小写不一致(x_train和X_train)
  • 错误赋值x_test = test[y],覆盖了测试集特征
  • 未定义y_test变量
  • 两次调用skf.split的输入不一致(第一次用特征集+目标,第二次用全量数据集)

修正后完整代码

import pandas as pd
import numpy as np
from sklearn.model_selection import StratifiedKFold
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score

from imblearn.over_sampling import SMOTE

# 修正:去掉多余的反引号
cleanedDataset = pd.read_csv("train_numeric_shuffled_50000_cleaned_90.csv")

# 改用列名列表来定义特征和目标,避免用DataFrame/Series作为索引
feature_cols = cleanedDataset.columns.drop(['Id', 'Response'])
target_col = 'Response'

x = cleanedDataset[feature_cols]
y = cleanedDataset[target_col]

# 初始化StratifiedKFold
skf = StratifiedKFold(n_splits=4)

# 打印折分索引(可选,只展示前5个避免输出过长)
for i, (train_index, test_index) in enumerate(skf.split(x, y)):
    print(f"Fold {i}:")
    print(f"  Train: index={train_index[:5]}...")
    print(f"  Test:  index={test_index[:5]}...")

# 定义模型,增加max_iter避免收敛警告
model = LogisticRegression(max_iter=1000)

def train_model(train, test, fold_no):
    # 用列名提取特征和标签
    x_train = train[feature_cols]
    y_train = train[target_col]
    x_test = test[feature_cols]
    y_test = test[target_col]
    
    model.fit(x_train, y_train)
    predictions = model.predict(x_test)
    print(f'Fold {fold_no} Accuracy: {accuracy_score(y_test, predictions):.4f}')

# 运行交叉验证,统一用特征集x和目标变量y作为split输入
fold_no = 1
for train_index, test_index in skf.split(x, y):
    train = cleanedDataset.loc[train_index, :]
    test = cleanedDataset.loc[test_index, :]
    train_model(train, test, fold_no)
    fold_no += 1

关键修改点

  • 将特征和目标的定义改为列名列表,避免直接用DataFrame/Series作为索引
  • 修正train_model函数中的变量赋值错误,确保正确提取训练/测试的特征与标签
  • 统一skf.split的输入参数为特征集x和目标变量y,保持逻辑一致性
  • 给LogisticRegression增加max_iter=1000,避免默认迭代次数不足导致的收敛警告

内容的提问来源于stack exchange,提问作者Martina Pascucci

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 03:25:18