GPT模型Softmax输出概率异常及语法任务表现疑问
问题分析与解决
核心问题
你遇到的情况本质是用错了模型类型+代码逻辑不符合任务需求:
- GPT2是自回归语言模型,预训练目标是预测下一个token,从未学习过处理掩码([MASK])填充任务;
- 你的代码逻辑不是预测掩码位置的动词概率,而是错误地计算了“上下文+选项”序列后接选项首个token的概率,完全偏离任务目标。
具体问题拆解
GPT2不支持掩码任务
GPT2的tokenizer默认没有[MASK]标记,输入中的[MASK]会被拆分为['[', 'MASK', ']']三个独立token,模型无法理解这是需要填充的位置,自然无法给出合理预测。代码逻辑错误
你的代码将上下文与选项拼接后,取最后一个token的logits来计算选项首个token的概率,这相当于在问“给定上下文+选项,下一个token是选项第一个词的概率”,和“掩码位置填哪个动词概率最高”的任务完全无关。
修正方案:换用MLM模型(如BERT)+ 正确逻辑
掩码填充任务是**掩码语言模型(MLM)**的专长,比如BERT系列模型,预训练时专门学习根据上下文预测掩码位置的token。以下是修正后的代码:
!pip install transformers import torch from transformers import BertTokenizer, BertForMaskedLM # 加载BERT模型和分词器 model_name = "bert-base-uncased" tokenizer = BertTokenizer.from_pretrained(model_name) model = BertForMaskedLM.from_pretrained(model_name) device = torch.device("cuda" if torch.cuda.is_available() else "cpu") model.to(device) def calculate_mask_probabilities(context, answer_choices): # 编码上下文,找到[MASK]的位置 inputs = tokenizer(context, return_tensors="pt").to(device) mask_token_index = torch.where(inputs["input_ids"] == tokenizer.mask_token_id)[1] # 获取模型对掩码位置的预测logits with torch.no_grad(): logits = model(**inputs).logits mask_logits = logits[0, mask_token_index, :] # 计算每个选项的概率 conditional_probs = [] for choice in answer_choices: # 编码选项(选项为单token词,直接取对应id) choice_token_id = tokenizer.encode(choice, add_special_tokens=False)[0] prob = torch.softmax(mask_logits, dim=-1)[0, choice_token_id].item() conditional_probs.append(prob) return conditional_probs # 测试 input_sentence = "The ballerinas' costumes that the thieves stole from the theatre last night [MASK] found at the abandoned condo." answer_choices = ["are", "is", "were"] conditional_probs = calculate_mask_probabilities(input_sentence, answer_choices) for choice, prob in zip(answer_choices, conditional_probs): prob_percentage = round(prob * 100, 2) print(f"选项 '{choice}' 的条件概率: {prob_percentage:.2f}%")
预期输出
运行后会得到符合语法逻辑的结果,正确选项were的概率会显著高于其他选项,示例输出类似:
选项 'are' 的条件概率: 11.23% 选项 'is' 的条件概率: 2.15% 选项 'were' 的条件概率: 86.62%
补充说明
如果一定要用GPT系列模型完成类似任务,需要调整任务形式:去掉[MASK],让模型生成后续文本,再统计生成目标动词的概率,但这种方式效率低且不如MLM模型直接。对于填空类语法任务,MLM模型是最优选择。
内容的提问来源于stack exchange,提问作者karak87rt0
相关产品推荐
相关产品推荐

