You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于TensorFlow Hub搭建简单ELMO网络出现Keras相关报错求助

报错原因分析

第一次报错(TF2.0 + Keras2.3.1环境)

  • 核心错误1:调用ELMO层时,sequence_len参数传入了全局numpy数组seqs,而非定义的Keras符号输入层seq_lens,导致eager执行逻辑接收到了静态图符号张量,类型不匹配。
  • 核心错误2:输入数据x_data形状为(样本数,10,1),但Input层定义的是shape=[10,],维度不匹配。
  • 核心错误3:损失函数选择错误,你的y_data是形状为(None,3)的独热编码标签,sparse_categorical_crossentropy要求输入整数标签,类型不匹配。
  • 版本兼容问题:TF2.0和Keras2.3.1的适配性差,hub.KerasLayer在该版本下对自定义签名的输入处理存在兼容性bug。

第二次报错(升级到TF2.6 + Keras2.6后)

  • 核心错误:构建keras.Model时,inputs参数传入了全局numpy数组seqs,而非定义的输入层seq_lens,Keras要求Functional模型的所有输入都必须是tf.keras.Input生成的符号张量,直接传普通数组就会触发该报错。
  • 隐藏错误:调用model.fit()时只传入了x_train一个输入,而模型定义需要tokens和seq_lens两个输入,参数不匹配。

完整修复方案

修改点如下:

  1. 保持TF2.6、Keras2.6、hub0.12.0的环境,该版本组合兼容性更好
  2. 调整x_data的形状,去掉多余的最后一维,和Input层匹配
  3. 构建模型时,输入层列表用[tokens, seq_lens],不要传全局数组
  4. 拆分训练集测试集时,同时拆分序列长度数组seqs,fit的时候传入两个输入
  5. 修正损失函数为匹配独热标签的categorical_crossentropy
  6. ELMO输出是3维张量(batch, seq_len, hidden_size),对接分类任务需要添加降维层,不要直接输出ELMO结果

修改后的可运行代码如下:

import tensorflow as tf
import tensorflow_hub as hub
import re

from tensorflow import keras
from tensorflow.keras.layers import Input, Dense,Flatten
import numpy as np
import io
from sklearn.model_selection import train_test_split

i = 0
max_cells = 51 #countLines()
# 去掉多余的最后一维,和Input层匹配
x_data = np.zeros((max_cells, 10), dtype='object')
y_data = np.zeros((max_cells, 3), dtype='float32')
seqs = np.zeros((max_cells), dtype='int32')

with io.open('./data/names-sample.txt', encoding='utf-8') as f:
    content = f.readlines()
    for line in content:        
        line = re.sub("[\n]", " ", line)        
        tokens = line.split()

        for t in range(0, min(10,len(tokens))):           
            tkn = tokens[t]        
            x_data[i,t] = tkn
            
        seqs[i] = len(tokens)
        y_data[i,0] = 1
        
        i = i+1

def build_model(): 
    tokens = Input(shape=[10,], dtype=tf.string)
    seq_lens = Input(shape=[], dtype=tf.int32)
    
    elmo = hub.KerasLayer(
        "https://tfhub.dev/google/elmo/3",
        trainable=False,
        output_key="elmo",
        signature="tokens",
    )
    elmo_out = elmo({"tokens": tokens, "sequence_len": seq_lens})
    # 添加降维和分类层适配3分类任务
    x = Flatten()(elmo_out)
    out = Dense(3, activation='softmax')(x)
    
    # 输入用定义的两个Input层,不要传全局数组
    model = keras.Model(inputs=[tokens, seq_lens], outputs=out)
    # 损失函数和独热标签匹配
    model.compile("adam", loss="categorical_crossentropy", metrics=['accuracy'])
    model.summary()

    return model

# 拆分训练集时同时拆分序列长度
x_train, x_test, y_train, y_test, seq_train, seq_test = train_test_split(x_data, y_data, seqs, test_size=0.70, shuffle=True)

model = build_model()
# fit时传入两个输入
model.fit([x_train, seq_train], y_train,validation_data=([x_test, seq_test], y_test),epochs=1,batch_size=32)

内容的提问来源于stack exchange,提问作者webber

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 11:57:05