You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow微调Longformer模型后保存加载报错排查

Longformer二分类模型微调后保存加载报错问题

环境与任务说明

  • 运行环境:Google Colab、TensorFlow 2.8.2
  • 使用预训练模型:allenai/longformer-base-4096
  • 任务类型:文本二分类,标签分为两类:不相关(0)、相关(1)
  • 训练限制:模型占满GPU显存,batch_size仅能设置为1,否则触发OOM错误,需先解决模型保存加载问题才能开展后续评估工作。

训练代码

tokenizer_path = "allenai/longformer-base-4096"
tokenizer = AutoTokenizer.from_pretrained(tokenizer_path)

model_path = "allenai/longformer-base-4096"
config = AutoConfig.from_pretrained(model_path)

config.attention_window = 128
config.num_labels = 2
config.id2label = {0: "irrelevant", 1: "relevant"}

longformer = TFAutoModel.from_pretrained(model_path, config=config)

def tokenize(texts):
    tokenized_texts = tokenizer(texts, truncation=True, padding=True, return_tensors="np")
    return tokenized_texts

def format_input(texts, labels):
    inputs = tokenize(texts).data
    labels = np.asarray(labels).astype('float16').reshape((-1,1))

    return inputs, labels

train_inputs, train_labels = format_input(train_texts, train_labels) # train_texts为字符串数组,train_labels为取值0/1的标签数组,验证集、测试集格式与训练集一致
val_inputs, val_labels = format_input(val_texts, val_labels)
test_inputs, test_labels = format_input(test_texts, test_labels)

# 模型层定义
input_ids = tf.keras.layers.Input((None,), dtype=np.int32, name="input_ids")
attention_mask = tf.keras.layers.Input((None,), dtype=np.int32, name="attention_mask")

longformer_output = longformer.longformer(input_ids=input_ids, attention_mask=attention_mask)
cls_output = longformer_output["last_hidden_state"][:,0,:]

hidden = tf.keras.layers.Dense(32, activation="tanh")(cls_output)
output = tf.keras.layers.Dense(1, activation="sigmoid")(hidden)
model = tf.keras.Model(inputs=[input_ids, attention_mask], outputs=[output])

loss = tf.keras.losses.BinaryCrossentropy()

metrics = [
    tf.metrics.BinaryAccuracy(),
]

# 确认第2层为Longformer层
print(model.layers[2])

# 冻结Longformer层
model.layers[2].trainable = False

epochs = 100

steps_per_epoch = len(train_labels)
num_train_steps = steps_per_epoch * epochs
num_warmup_steps = int(0.1*num_train_steps)

init_lr = 3e-5
optimizer = optimization.create_optimizer(init_lr=init_lr,
                                          num_train_steps=num_train_steps,
                                          num_warmup_steps=num_warmup_steps,
                                          optimizer_type='adamw')

model.compile(optimizer=optimizer, loss=loss, metrics=metrics)

early_stopping_callback = tf.keras.callbacks.EarlyStopping(monitor='loss', patience=0)

history = model.fit(
    x=train_inputs,
    y=train_labels,
    batch_size=1,
    validation_batch_size=1,
    validation_data=(val_inputs, val_labels),
    epochs=epochs,
    class_weight=class_weight,
    steps_per_epoch=steps_per_epoch,
    callbacks=[
        early_stopping_callback,
    ]
)

trained_model_save_path = "tf_longformer_cls"
model.save(trained_model_save_path, include_optimizer=False)

模型结构摘要

Model: "model_1"
__________________________________________________________________________________________________
 Layer (type)                   Output Shape         Param #     Connected to                     
==================================================================================================
 input_ids (InputLayer)         [(None, 4096)]       0           []                               
                                                                                                  
 attention_mask (InputLayer)    [(None, 4096)]       0           []                               
                                                                                                  
 longformer (TFLongformerMainLa  multiple            148659456   ['input_ids[0][0]',              
 yer)                                                             'attention_mask[0][0]']         
                                                                                                  
 tf.__operators__.getitem_1 (Sl  (None, 768)         0           ['longformer[1][0]']             
 icingOpLambda)                                                                                   
                                                                                                  
 dense_2 (Dense)                (None, 32)           24608       ['tf.__operators__.getitem_1[0][0
                                                                 ]']                              
                                                                                                  
 dense_3 (Dense)                (None, 1)            33          ['dense_2[0][0]']                
                                                                                                  
==================================================================================================
Total params: 148,684,097
Trainable params: 148,684,097
Non-trainable params: 0

报错详情

模型加载阶段报错

训练完成后保存模型到Google Drive,使用如下代码加载时触发ValueError:

new_model = tf.keras.models.load_model(trained_model_save_path)

报错内容如下:

ValueError: The two structures don't have the same nested structure.

First structure: type=tuple str=(({'input_ids': TensorSpec(shape=(None, 5), dtype=tf.int32, name='input_ids/input_ids'), 'global_attention_mask': TensorSpec(shape=(None, 5), dtype=tf.int32, name=None), 'attention_mask': TensorSpec(shape=(None, 5), dtype=tf.int32, name=None)}, None, None, None, None, None, None, None, None, None, False), {})

Second structure: type=tuple str=((TensorSpec(shape=(None, None), dtype=tf.int32, name='input_ids'), TensorSpec(shape=(None, None), dtype=tf.int32, name='attention_mask'), None, None, None, None, None, None, None, None, False), {})

More specifically: Substructure "type=dict str={'input_ids': TensorSpec(shape=(None, 5), dtype=tf.int32, name='input_ids/input_ids'), 'global_attention_mask': TensorSpec(shape=(None, 5), dtype=tf.int32, name=None), 'attention_mask': TensorSpec(shape=(None, 5), dtype=tf.int32, name=None)}" is a sequence, while substructure "type=TensorSpec str=TensorSpec(shape=(None, None), dtype=tf.int32, name='input_ids')" is not
Entire first structure:
(({'input_ids': ., 'global_attention_mask': ., 'attention_mask': .}, ., ., ., ., ., ., ., ., ., .), {})
Entire second structure:
((., ., ., ., ., ., ., ., ., ., .), {})

自行修改代码后的报错

初步判断报错与输入形状不匹配有关,检索到BERT同类问题可通过修改传参方式解决:用位置传参替代关键字传参给嵌入层传入输入。因此将代码中
longformer_output = longformer.longformer(input_ids=input_ids, attention_mask=attention_mask)
修改为
longformer_output = longformer.longformer([input_ids, attention_mask])
但修改后保存模型阶段就触发OperatorNotAllowedInGraphError,报错内容如下:

OperatorNotAllowedInGraphError: Exception encountered when calling layer "longformer" (type TFLongformerMainLayer).

using a `tf.Tensor` as a Python `bool` is not allowed: AutoGraph did convert this function. This might indicate you are trying to use an unsupported feature.

Call arguments received:
  • args=(['tf.Tensor(shape=(None, None), dtype=int32)', 'tf.Tensor(shape=(None, None), dtype=int32)'],)
  • kwargs={'training': 'False'}

内容的提问来源于stack exchange,提问作者vlsb

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 01:06:23