使用HuggingFace TFBertModel与AutoTokenizer构建模型的输入问题
问题根因
你遇到的所有报错本质上都是同一个问题:HuggingFace原生分词器只能接收实际字符串输入,无法直接处理Keras构建计算图时的符号张量,同时你原有代码还存在两个隐藏逻辑错误:
pooler_output是BERT的<[BOS_never_used_51bce0c785ca2f68081bfa7d91973934]>输出,为二维张量,无序列维度,无法直接传入要求三维输入的LSTM层- 原代码中
Concatenate(axis=-1)([X, input_layer])的input_layer未定义
最优解决方案(推荐)
把分词预处理放到数据加载阶段实现,不要嵌入模型计算图,既规避张量兼容问题,也能利用CPU并行预处理提升训练效率。
步骤1:改造数据预处理管道
import tensorflow as tf def preprocess_single_sample(text, label): # 内部分词逻辑,运行在eager模式下可直接调用numpy() def run_tokenize(text): text_str = text.numpy().decode("utf-8") encoded = tokenizer( text_str, add_special_tokens=True, max_length=110, padding="max_length", truncation=True ) return encoded["input_ids"], encoded["attention_mask"], encoded["token_type_ids"] # 用tf.py_function封装分词逻辑,兼容数据管道 input_ids, attention_mask, token_type_ids = tf.py_function( func=run_tokenize, inp=[text], Tout=[tf.int32, tf.int32, tf.int32] ) # 固定张量形状,避免后续模型推断形状失败 input_ids.set_shape([110]) attention_mask.set_shape([110]) token_type_ids.set_shape([110]) return (input_ids, attention_mask, token_type_ids), label # 示例:构建训练数据管道 train_dataset = tf.data.Dataset.from_tensor_slices((train_texts, train_labels)) train_dataset = train_dataset.map( preprocess_single_sample, num_parallel_calls=tf.data.AUTOTUNE ).batch(32).shuffle(1000)
步骤2:修改模型输入结构
def build_classifier_model(): # 直接接收分词后的三个张量输入 input_ids = tf.keras.layers.Input(shape=(110,), dtype=tf.int32, name="input_ids") attention_mask = tf.keras.layers.Input(shape=(110,), dtype=tf.int32, name="attention_mask") token_type_ids = tf.keras.layers.Input(shape=(110,), dtype=tf.int32, name="token_type_ids") bert_output = bert([input_ids, attention_mask, token_type_ids]) # 用last_hidden_state替代pooler_output,获取序列维度输出适配LSTM seq_output = bert_output["last_hidden_state"] X = tf.keras.layers.Bidirectional( tf.keras.layers.LSTM(64, return_sequences=True, dropout=0.1, recurrent_dropout=0.1) )(seq_output) X = tf.keras.layers.MaxPooling1D(20)(X) X = tf.keras.layers.SpatialDropout1D(0.4)(X) X = tf.keras.layers.Flatten()(X) X = tf.keras.layers.Dense(128, activation="relu")(X) X = tf.keras.layers.Dropout(0.25)(X) X = tf.keras.layers.Dense(2, activation="softmax")(X) return tf.keras.Model(inputs=[input_ids, attention_mask, token_type_ids], outputs=X)
端到端部署方案(可选)
如果需要模型直接接收字符串输入实现端到端部署,可以使用HuggingFace Fast分词器自带的TF层转换能力,无需自行封装预处理层:
# 加载Fast版本分词器 tokenizer = AutoTokenizer.from_pretrained(model_name, use_fast=True) # 直接转换为TensorFlow层 tf_tokenizer_layer = tokenizer.to_tf_layer( max_length=110, padding="max_length", truncation=True ) def build_end2end_classifier(): text_input = tf.keras.layers.Input(shape=(), dtype=tf.string, name="text") encoder_inputs = tf_tokenizer_layer(text_input) bert_output = bert(encoder_inputs) seq_output = bert_output["last_hidden_state"] # 后续层逻辑和上述一致 X = tf.keras.layers.Bidirectional( tf.keras.layers.LSTM(64, return_sequences=True, dropout=0.1, recurrent_dropout=0.1) )(seq_output) X = tf.keras.layers.MaxPooling1D(20)(X) X = tf.keras.layers.SpatialDropout1D(0.4)(X) X = tf.keras.layers.Flatten()(X) X = tf.keras.layers.Dense(128, activation="relu")(X) X = tf.keras.layers.Dropout(0.25)(X) X = tf.keras.layers.Dense(2, activation="softmax")(X) return tf.keras.Model(inputs=text_input, outputs=X)
内容的提问来源于stack exchange,提问作者Gerardo Zinno
相关产品推荐
相关产品推荐

