TensorFlow微调Longformer模型后保存加载报错排查
Longformer二分类模型微调后保存加载报错问题
环境与任务说明
- 运行环境:Google Colab、TensorFlow 2.8.2
- 使用预训练模型:allenai/longformer-base-4096
- 任务类型:文本二分类,标签分为两类:不相关(0)、相关(1)
- 训练限制:模型占满GPU显存,batch_size仅能设置为1,否则触发OOM错误,需先解决模型保存加载问题才能开展后续评估工作。
训练代码
tokenizer_path = "allenai/longformer-base-4096" tokenizer = AutoTokenizer.from_pretrained(tokenizer_path) model_path = "allenai/longformer-base-4096" config = AutoConfig.from_pretrained(model_path) config.attention_window = 128 config.num_labels = 2 config.id2label = {0: "irrelevant", 1: "relevant"} longformer = TFAutoModel.from_pretrained(model_path, config=config) def tokenize(texts): tokenized_texts = tokenizer(texts, truncation=True, padding=True, return_tensors="np") return tokenized_texts def format_input(texts, labels): inputs = tokenize(texts).data labels = np.asarray(labels).astype('float16').reshape((-1,1)) return inputs, labels train_inputs, train_labels = format_input(train_texts, train_labels) # train_texts为字符串数组,train_labels为取值0/1的标签数组,验证集、测试集格式与训练集一致 val_inputs, val_labels = format_input(val_texts, val_labels) test_inputs, test_labels = format_input(test_texts, test_labels) # 模型层定义 input_ids = tf.keras.layers.Input((None,), dtype=np.int32, name="input_ids") attention_mask = tf.keras.layers.Input((None,), dtype=np.int32, name="attention_mask") longformer_output = longformer.longformer(input_ids=input_ids, attention_mask=attention_mask) cls_output = longformer_output["last_hidden_state"][:,0,:] hidden = tf.keras.layers.Dense(32, activation="tanh")(cls_output) output = tf.keras.layers.Dense(1, activation="sigmoid")(hidden) model = tf.keras.Model(inputs=[input_ids, attention_mask], outputs=[output]) loss = tf.keras.losses.BinaryCrossentropy() metrics = [ tf.metrics.BinaryAccuracy(), ] # 确认第2层为Longformer层 print(model.layers[2]) # 冻结Longformer层 model.layers[2].trainable = False epochs = 100 steps_per_epoch = len(train_labels) num_train_steps = steps_per_epoch * epochs num_warmup_steps = int(0.1*num_train_steps) init_lr = 3e-5 optimizer = optimization.create_optimizer(init_lr=init_lr, num_train_steps=num_train_steps, num_warmup_steps=num_warmup_steps, optimizer_type='adamw') model.compile(optimizer=optimizer, loss=loss, metrics=metrics) early_stopping_callback = tf.keras.callbacks.EarlyStopping(monitor='loss', patience=0) history = model.fit( x=train_inputs, y=train_labels, batch_size=1, validation_batch_size=1, validation_data=(val_inputs, val_labels), epochs=epochs, class_weight=class_weight, steps_per_epoch=steps_per_epoch, callbacks=[ early_stopping_callback, ] ) trained_model_save_path = "tf_longformer_cls" model.save(trained_model_save_path, include_optimizer=False)
模型结构摘要
Model: "model_1" __________________________________________________________________________________________________ Layer (type) Output Shape Param # Connected to ================================================================================================== input_ids (InputLayer) [(None, 4096)] 0 [] attention_mask (InputLayer) [(None, 4096)] 0 [] longformer (TFLongformerMainLa multiple 148659456 ['input_ids[0][0]', yer) 'attention_mask[0][0]'] tf.__operators__.getitem_1 (Sl (None, 768) 0 ['longformer[1][0]'] icingOpLambda) dense_2 (Dense) (None, 32) 24608 ['tf.__operators__.getitem_1[0][0 ]'] dense_3 (Dense) (None, 1) 33 ['dense_2[0][0]'] ================================================================================================== Total params: 148,684,097 Trainable params: 148,684,097 Non-trainable params: 0
报错详情
模型加载阶段报错
训练完成后保存模型到Google Drive,使用如下代码加载时触发ValueError:
new_model = tf.keras.models.load_model(trained_model_save_path)
报错内容如下:
ValueError: The two structures don't have the same nested structure. First structure: type=tuple str=(({'input_ids': TensorSpec(shape=(None, 5), dtype=tf.int32, name='input_ids/input_ids'), 'global_attention_mask': TensorSpec(shape=(None, 5), dtype=tf.int32, name=None), 'attention_mask': TensorSpec(shape=(None, 5), dtype=tf.int32, name=None)}, None, None, None, None, None, None, None, None, None, False), {}) Second structure: type=tuple str=((TensorSpec(shape=(None, None), dtype=tf.int32, name='input_ids'), TensorSpec(shape=(None, None), dtype=tf.int32, name='attention_mask'), None, None, None, None, None, None, None, None, False), {}) More specifically: Substructure "type=dict str={'input_ids': TensorSpec(shape=(None, 5), dtype=tf.int32, name='input_ids/input_ids'), 'global_attention_mask': TensorSpec(shape=(None, 5), dtype=tf.int32, name=None), 'attention_mask': TensorSpec(shape=(None, 5), dtype=tf.int32, name=None)}" is a sequence, while substructure "type=TensorSpec str=TensorSpec(shape=(None, None), dtype=tf.int32, name='input_ids')" is not Entire first structure: (({'input_ids': ., 'global_attention_mask': ., 'attention_mask': .}, ., ., ., ., ., ., ., ., ., .), {}) Entire second structure: ((., ., ., ., ., ., ., ., ., ., .), {})
自行修改代码后的报错
初步判断报错与输入形状不匹配有关,检索到BERT同类问题可通过修改传参方式解决:用位置传参替代关键字传参给嵌入层传入输入。因此将代码中longformer_output = longformer.longformer(input_ids=input_ids, attention_mask=attention_mask)
修改为longformer_output = longformer.longformer([input_ids, attention_mask])
但修改后保存模型阶段就触发OperatorNotAllowedInGraphError,报错内容如下:
OperatorNotAllowedInGraphError: Exception encountered when calling layer "longformer" (type TFLongformerMainLayer). using a `tf.Tensor` as a Python `bool` is not allowed: AutoGraph did convert this function. This might indicate you are trying to use an unsupported feature. Call arguments received: • args=(['tf.Tensor(shape=(None, None), dtype=int32)', 'tf.Tensor(shape=(None, None), dtype=int32)'],) • kwargs={'training': 'False'}
内容的提问来源于stack exchange,提问作者vlsb
相关产品推荐
相关产品推荐

