如何从Hugging Face Transformer问答示例代码中获取答案置信度得分
问题解答
现有代码完全可以实现置信度得分返回,你只需要在原有逻辑基础上增加少量计算即可:
实现原理
你代码中已经获取到的answer_start_scores和answer_end_scores是问答模型输出的原始logits值,分别对应输入序列中每个token作为答案起始、结束位置的打分,只需要将这两个打分通过softmax转换为概率分布,取答案对应起止位置的概率相乘,即可得到该答案的置信度,这个计算逻辑和Hugging Face官方pipeline的置信度计算规则一致。
修改后的核心代码
from transformers import AutoTokenizer, TFAutoModelForQuestionAnswering import tensorflow as tf tokenizer = AutoTokenizer.from_pretrained("bert-large-uncased-whole-word-masking-finetuned-squad") model = TFAutoModelForQuestionAnswering.from_pretrained("bert-large-uncased-whole-word-masking-finetuned-squad") text = r""" 🤗 Transformers (前称pytorch-transformers和pytorch-pretrained-bert)提供了面向自然语言理解(NLU)和自然语言生成(NLG)的通用架构(包括BERT、GPT-2、RoBERTa、XLM、DistilBert、XLNet等),支持100多种语言的超过32个预训练模型,同时可在TensorFlow 2.0和PyTorch之间无缝切换使用。 """ questions = [ "Transformers库中提供了多少个预训练模型?", "Transformers库提供了什么能力?", "Transformers库支持哪些框架之间的互操作?", ] for question in questions: inputs = tokenizer.encode_plus(question, text, add_special_tokens=True, return_tensors="tf") input_ids = inputs["input_ids"].numpy()[0] text_tokens = tokenizer.convert_ids_to_tokens(input_ids) answer_start_scores, answer_end_scores = model(inputs) answer_start = tf.argmax( answer_start_scores, axis=1 ).numpy()[0] # 取打分最高的位置作为答案起始下标 answer_end = ( tf.argmax(answer_end_scores, axis=1) + 1 ).numpy()[0] # 取打分最高的位置作为答案结束下标 answer = tokenizer.convert_tokens_to_string(tokenizer.convert_ids_to_tokens(input_ids[answer_start:answer_end])) # 新增置信度计算逻辑 start_probs = tf.nn.softmax(answer_start_scores, axis=1) end_probs = tf.nn.softmax(answer_end_scores, axis=1) start_conf = tf.reduce_max(start_probs).numpy() end_conf = tf.reduce_max(end_probs).numpy() total_confidence = start_conf * end_conf print(f"问题: {question}") print(f"答案: {answer}") print(f"置信度: {total_confidence:.4f}\n")
内容的提问来源于stack exchange,提问作者Abhishek K
相关产品推荐
相关产品推荐

