You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

集成学习权重优化:Tensor转译错误排查与技术实现求助

问题描述

我用三个预训练模型做集成学习,每个模型对训练集输出类别置信度的NumPy数组(示例如下):

# 针对A、B、C三类
# 某模型对4张训练图像的预测数组:
[[0.2, 0.2, 0.6],
 [0.1, 0.8, 0.1], 
 [0.4, 0.4, 0.2], 
 [0.3, 0.3, 0.4]] 

计划通过加权求和生成最终预测:M = w1*M1 + w2*M2 + w3*M3(w3=1-w1-w2),用神经网络结合独热编码标签优化损失求解权重。独热编码标签示例:

[[0.0, 0.0, 1.0],
 [0.0, 1.0, 0.0],
 [1.0, 0.0, 0.0],
 [0.0, 1.0, 0.0]]

将三个模型的置信度存入pandas DataFrame作为输入,和独热标签一起训练ANN时,出现错误:

Failed to convert a NumPy array to a Tensor (Unsupported object type tensorflow.python.framework.ops.EagerTensor)

相关代码如下:

集成模型定义

model_ensemble  = tf.keras.models.Sequential()
model_ensemble.add(tf.keras.layers.Dense(units=10, activation='relu'))
model_ensemble.add(tf.keras.layers.Dense(units=10, activation='relu'))
model_ensemble.add(tf.keras.layers.Dense(units=3, activation='softmax'))
model_ensemble.compile(optimizer = 'adam', loss = 'categorical_crossentropy', metrics = ['accuracy'])
model_ensemble.fit(X, a, batch_size = 32, epochs = 25)

独热编码标签生成

a= pd.DataFrame(x_train['label'] ).astype('int').astype('category')
print(a)
a= tf.keras.utils.to_categorical(a)

预测标签生成逻辑

label1= [np.asarray([])]
label2= [np.asarray([])]
label3= [np.asarray([])]
for i in x_train['name']:
    Path= path.join('/content/data/train', I)
    img = tf.io.read_file(Path)
    tensor = tf.io.decode_image(img, channels=3, dtype=tf.dtypes.float32)
    tensor = tf.image.resize(tensor, [224, 224])
    input_tensor = tf.expand_dims(tensor, axis=0)
    result1=model1.predict(input_tensor)
    result2=model2.predict(input_tensor)
    result3=model3.predict(input_tensor)
    label1.append(tf.convert_to_tensor(result1))
    label2.append(tf.convert_to_tensor(result2))
    label3.append(tf.convert_to_tensor(result3))

错误原因

  1. 输入特征格式不兼容:生成预测时,将model.predict()返回的NumPy数组转成了EagerTensor并存入列表,最终组成的DataFrame包含Tensor对象,Keras拟合时无法将混合格式的DataFrame转换为统一的Tensor输入。
  2. 独热编码输入错误:tf.keras.utils.to_categorical要求输入是整数类型的一维数组,但代码中先转成了category类型,会导致编码过程出现格式异常。
  3. 初始化列表存在空数组:label1= [np.asarray([])]的初始化方式会在列表开头加入空数组,后续拼接后破坏输入数据的形状,加剧格式转换问题。

解决方案

1. 修正预测特征生成逻辑

直接保留model.predict()返回的NumPy数组,不转成Tensor,同时正确初始化列表:

# 初始化空列表,避免加入空数组
label1 = []
label2 = []
label3 = []
for name in x_train['name']:
    Path = path.join('/content/data/train', name)  # 修正变量名拼写错误:I改为name
    img = tf.io.read_file(Path)
    tensor = tf.io.decode_image(img, channels=3, dtype=tf.dtypes.float32)
    tensor = tf.image.resize(tensor, [224, 224])
    input_tensor = tf.expand_dims(tensor, axis=0)
    # 保留NumPy数组,去掉多余的batch维度
    result1 = model1.predict(input_tensor).squeeze()
    result2 = model2.predict(input_tensor).squeeze()
    result3 = model3.predict(input_tensor).squeeze()
    label1.append(result1)
    label2.append(result2)
    label3.append(result3)

# 转换为NumPy数组并拼接成输入X
M1 = np.array(label1)
M2 = np.array(label2)
M3 = np.array(label3)
X = np.concatenate([M1, M2, M3], axis=1)  # 每个样本对应9个特征(3个模型×3类置信度)

2. 修正独热编码生成

直接使用整数数组输入to_categorical:

# 提取整数标签,无需转换为category类型
labels = x_train['label'].astype('int').values
a = tf.keras.utils.to_categorical(labels, num_classes=3)  # 指定类别数,避免自动推断出错

3. 重新训练集成模型

此时X和a都是标准NumPy数组,可直接用于拟合,同时明确输入形状:

model_ensemble  = tf.keras.models.Sequential()
model_ensemble.add(tf.keras.layers.Dense(units=10, activation='relu', input_shape=(9,)))  # 对应9个输入特征
model_ensemble.add(tf.keras.layers.Dense(units=10, activation='relu'))
model_ensemble.add(tf.keras.layers.Dense(units=3, activation='softmax'))
model_ensemble.compile(optimizer = 'adam', loss = 'categorical_crossentropy', metrics = ['accuracy'])
model_ensemble.fit(X, a, batch_size = 32, epochs = 25)

内容的提问来源于stack exchange,提问作者napstablook

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 11:25:26