Mistral-7B后续文本是否影响前置Token的Logits?实验疑问
为什么Mistral-7B输入序列长度不同时,前置Token的Logits会出现差异?
我原本认为Transformer内部的掩码机制会阻止后续Token影响前置Token的Logits,但使用Mistral-7B + torch.bfloat16进行实验时发现了异常:
### TEST # 假设已完成输入的tokenize并保存到"inp_rej",形状为[1, 192] model1.eval() new_inp = inp_rej[0, :173] with torch.no_grad(): new_out1 = model1.generate(new_inp.unsqueeze(0), temperature=0, max_length=256, return_dict_in_generate=True, output_scores=True) temp_out1 = model1(new_inp.unsqueeze(0)) comp_out1 = model1(inp_rej) a1 = torch.softmax(new_out1['scores'][0], dim=-1).max() a2 = torch.softmax(temp_out1.logits[0][-1], dim=-1).max() a3 = torch.softmax(comp_out1.logits[0, len(new_inp) - 1], dim=-1).max() print(a1 - a2) # tensor(0., device='cuda:0'), 符合预期 print(a1 - a3) # tensor(0.0300, device='cuda:0'), 为何出现差异?
实验中a1与a3的差值为0.03,甚至执行abs(temp_out1.logits[0][0] - comp_out1.logits[0][0]).mean()时,输出为tensor(0.0122, device='cuda:0')。长序列场景下该差异更显著:
for N in [30, 60, 90, 120, 150]: new_inp_rej = inp_rej.clone()[0:N] model1.eval() new_inp = new_inp_rej[0, :N - 20] with torch.no_grad(): new_out1 = model1.generate(new_inp.unsqueeze(0), temperature=0, max_length=256, return_dict_in_generate=True, output_scores=True) temp_out1 = model1(new_inp.unsqueeze(0)) comp_out1 = model1(inp_rej) a1 = torch.softmax(new_out1['scores'][0], dim=-1).max() a2 = torch.softmax(temp_out1.logits[0][-1], dim=-1).max() a3 = torch.softmax(comp_out1.logits[0, len(new_inp) - 1], dim=-1).max() diff = (a1 - a3).item() if diff != 0: print(N, ":", diff) # 结果: # 60 : 3.6954879760742188e-06 # 90 : 0.0025225281715393066 # 120 : 0.00031453371047973633
请问为何会出现上述差异?
内容的提问来源于stack exchange,提问作者Hellowhatsup
相关产品推荐
相关产品推荐

