TFBertModel调用方法末尾方括号的作用是什么?
问题:TFBertModel调用语句末尾[0]的含义是什么?
我在代码中使用bert-base-cased预训练模型时,无法理解以下代码行末尾方括号[0]的作用:
input_ids = Input(shape=(max_len,), dtype=tf.int32, name="input_ids") input_mask = Input(shape=(max_len,), dtype=tf.int32, name="attention_mask") embeddings = bert(input_ids, attention_mask = input_mask)[0]
对应的模型声明代码如下:
from transformers import AutoTokenizer,TFBertModel tokenizer = AutoTokenizer.from_pretrained('bert-base-cased') bert = TFBertModel.from_pretrained('bert-base-cased')
我查阅Hugging Face文档时,发现TFBertForNextSentencePrediction的调用示例中也使用了类似写法,但未说明其作用:
>>> import tensorflow as tf >>> from transformers import BertTokenizer, TFBertForNextSentencePrediction >>> tokenizer = BertTokenizer.from_pretrained('bert-base-uncased') >>> model = TFBertForNextSentencePrediction.from_pretrained('bert-base-uncased') >>> prompt = "In Italy, pizza served in formal settings, such as at a restaurant, is presented unsliced." >>> next_sentence = "The sky is blue due to the shorter wavelength of blue light." >>> encoding = tokenizer(prompt, next_sentence, return_tensors='tf') >>> logits = model(encoding['input_ids'], token_type_ids=encoding['token_type_ids'])[0] >>> assert logits[0][0] < logits[0][1] # 下一句为随机无关内容
解答
这里的[0]是用来提取模型返回元组中的第一个元素,也就是我们需要的核心输出部分。
Hugging Face的TensorFlow系列BERT模型(如TFBertModel、TFBertForNextSentencePrediction)在调用时,返回的是一个元组对象,里面封装了模型不同维度的输出结果:
针对
TFBertModel:- 元组第一个元素:
last_hidden_state→ 模型最后一层输出的所有token级嵌入,形状为(batch_size, sequence_length, hidden_size),这正是你代码里要赋值给embeddings的序列级特征。 - 元组第二个元素:
pooler_output→ <[BOS_never_used_51bce0c785ca2f68081bfa7d91973934]>token经过全连接层和tanh激活后的句子级嵌入,形状为(batch_size, hidden_size),常用于句子分类等任务。
- 元组第一个元素:
针对
TFBertForNextSentencePrediction:
元组第一个元素是下一句预测任务的logits(原始分类输出),后续元素则包含序列嵌入、池化嵌入等额外输出(取决于配置)。
简单来说,[0]就是直接从模型返回的多结果元组中,取出我们当前任务需要的那一部分输出。如果不需要其他额外输出,只取第一个元素即可满足需求。
内容的提问来源于stack exchange,提问作者Majd
相关产品推荐
相关产品推荐

