You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

调用FN1函数执行MinMaxScaler时触发ValueError:需至少一个数组或dtype

问题排查:MinMaxScaler触发ValueError: at least one array or dtype is required

问题背景

FN1函数代码

def FN1(I,trainInput,trainOutput,dim):
         reducedfeatures=[]
         max_fold=3
         cv = StratifiedKFold(n_splits=max_fold,shuffle=True,random_state=20181224)
         fold_N=0
         headers = list(trainInput.columns)
         Accuracy= np.zeros((max_fold))
         min_max_scaler = MinMaxScaler()
    
         for index in range(0,dim):
           if (I[index]==1):
               reducedfeatures.append(index)

         X=trainInput.iloc[:,reducedfeatures]
         y=trainOutput

         bag = KNeighborsClassifier(n_neighbors=5)
         for train_index, test_index in cv.split(X,y):
             X_train, X_test =  X.iloc[train_index], X.iloc[test_index]
             y_train, y_test = y[train_index], y[test_index]
             X_train = min_max_scaler.fit_transform(X_train)
             X_test = min_max_scaler.transform(X_test)
        
          
         bag.fit(X_train, y_train)
         bag_test_pred = bag.predict(X_test)
         acc = accuracy_score(y_test, bag_test_pred)
         Accuracy[fold_N]= acc
         fold_N = fold_N +1             
        

         acc_train = float(Accuracy.mean())
         fitness=0.99*(1-acc_train)+0.01*sum(I)/(dim)

         return fitness

调用参数

  • I:长度为51的数组,元素为0或1(示例值:[1. 0. 0. 0. 1. 0. 0. 0. 0. 0. 0. 1. 1. 1. 1. 0. 1. 1. 1. 1. 1. 1. 1. 1., 0. 1. 1. 1. 0. 1. 1. 1. 0. 0. 1. 1. 1. 0. 0. 1. 1. 0. 1. 0. 1. 1. 1. 1., 1. 1. 0.])
  • trainInput:形状为(3662,51)的DataFrame,包含数值特征
  • trainOutput:形状为(3662,)的全0标签ndarray
  • dim:51

错误栈

V.VVVV.V is optimizing  "FN1"
Traceback (most recent call last):
  File "C:\Users\sa\Desktop\ISSAA\main.py", line 69, in <module>
    x=slctr.selector(i,func_details,PopulationSize,Iterations,completeData)
  File "C:\Users\sa\Desktop\ISSAA\selector.py", line 69, in selector
    x=ssaelo.SSAELO(getattr(fitnessFUNs, function_name),lb,ub,dim,popSize,Iter,features,labels)
  File "C:\Users\sa\Desktop\ISSAA\SSAELO.py", line 255, in SSAELO
    SalpPositions[:, i] = main(objf, N, Max_iteration, 0.5, 0.5, lb, ub, trainInput, trainOutput, dim)
  File "C:\Users\sa\Desktop\ISSAA\SSAELO.py", line 118, in main
    obj = [cal_obj(pop[i], trainInput, trainOutput, dim) for i in range(npop)]  # objectives
  File "C:\Users\sa\Desktop\ISSAA\SSAELO.py", line 118, in <listcomp>
    obj = [cal_obj(pop[i], trainInput, trainOutput, dim) for i in range(npop)]  # objectives
  File "C:\Users\sa\Desktop\ISSAA\fitnessFUNs.py", line 49, in FN1
    X_train = min_max_scaler.fit_transform(X_train)
  File "C:\Users\sa\anaconda3\lib\site-packages\sklearn\base.py", line 852, in fit_transform
    return self.fit(X, **fit_params).transform(X)
  File "C:\Users\sa\anaconda3\lib\site-packages\sklearn\preprocessing\_data.py", line 416, in fit
    return self.partial_fit(X, y)
  File "C:\Users\sa\anaconda3\lib\site-packages\sklearn\preprocessing\_data.py", line 453, in partial_fit
    X = self._validate_data(
  File "C:\Users\sa\anaconda3\lib\site-packages\sklearn\base.py", line 566, in _validate_data
    X = check_array(X, **check_params)
  File "C:\Users\sa\anaconda3\lib\site-packages\sklearn\utils\validation.py", line 665, in check_array
    dtype_orig = np.result_type(*dtypes_orig)
  File "<__array_function__ internals>", line 5, in result_type
ValueError: at least one array or dtype is required

错误原因分析

  1. 空特征集问题:当传入的I数组中没有值为1的元素时,reducedfeatures会是空列表,导致X = trainInput.iloc[:,reducedfeatures]生成空DataFrame。后续交叉验证拆分出的X_train也为空,MinMaxScaler.fit_transform无法处理空数据集,触发报错。
  2. 循环逻辑错误:原代码中模型训练、预测、准确率计算的代码位于cv.split循环外部,仅会处理最后一次循环生成的X_train/X_test,且fold_N自增逻辑错误,导致Accuracy数组仅第一个元素被赋值,完全不符合交叉验证的预期流程。
  3. 分层交叉验证不适用全类别标签:trainOutput是全0的ndarray,而StratifiedKFold要求目标变量至少包含两个类别才能实现分层拆分,这种情况下使用该方法会导致拆分逻辑异常,甚至间接引发数据处理错误。

解决方法

1. 增加空特征集校验

在生成reducedfeatures后,检查是否为空,若为空则直接返回惩罚性的fitness值(避免后续流程报错):

if not reducedfeatures:
    # 没有选中任何特征,返回最大可能的fitness(fitness越小性能越好)
    return 1.0

2. 修复交叉验证循环逻辑

将模型训练、预测、准确率计算的代码移入cv.split循环内部,确保每个fold都完成完整的训练和评估,同时每个fold独立初始化scaler避免数据泄露:

for train_index, test_index in cv.split(X,y):
    X_train, X_test = X.iloc[train_index], X.iloc[test_index]
    y_train, y_test = y[train_index], y[test_index]
    # 每个fold重新初始化scaler
    fold_scaler = MinMaxScaler()
    X_train = fold_scaler.fit_transform(X_train)
    X_test = fold_scaler.transform(X_test)
    
    bag.fit(X_train, y_train)
    bag_test_pred = bag.predict(X_test)
    acc = accuracy_score(y_test, bag_test_pred)
    Accuracy[fold_N] = acc
    fold_N += 1

3. 替换交叉验证方法

由于trainOutput是全0标签,改用普通KFold替代StratifiedKFold:

from sklearn.model_selection import KFold
cv = KFold(n_splits=max_fold, shuffle=True, random_state=20181224)

完整修改后的FN1函数

def FN1(I, trainInput, trainOutput, dim):
    reducedfeatures = []
    max_fold = 3
    # 替换为KFold适配全类别标签
    cv = KFold(n_splits=max_fold, shuffle=True, random_state=20181224)
    fold_N = 0
    Accuracy = np.zeros((max_fold))
    
    for index in range(dim):
        if I[index] == 1:
            reducedfeatures.append(index)
    
    # 校验空特征集
    if not reducedfeatures:
        return 1.0
    
    X = trainInput.iloc[:, reducedfeatures]
    y = trainOutput

    bag = KNeighborsClassifier(n_neighbors=5)
    
    for train_index, test_index in cv.split(X, y):
        X_train, X_test = X.iloc[train_index], X.iloc[test_index]
        y_train, y_test = y[train_index], y[test_index]
        # 每个fold独立初始化scaler,防止数据泄露
        fold_scaler = MinMaxScaler()
        X_train = fold_scaler.fit_transform(X_train)
        X_test = fold_scaler.transform(X_test)
        
        bag.fit(X_train, y_train)
        bag_test_pred = bag.predict(X_test)
        acc = accuracy_score(y_test, bag_test_pred)
        Accuracy[fold_N] = acc
        fold_N += 1             
    
    acc_train = float(Accuracy.mean())
    fitness = 0.99 * (1 - acc_train) + 0.01 * sum(I) / dim

    return fitness

内容的提问来源于stack exchange,提问作者Mhd33

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 13:05:04