带Adapter的RoBERTa蒸馏:SequenceClassifierOutput损失为生成器问题
问题:带Adapter的RoBERTa模型蒸馏时loss为生成器而非张量的原因
背景
对带有Adapter的RoBERTa模型执行知识蒸馏,修改了权重蒸馏函数以让学生模型继承教师模型的Adapter参数,但使用自定义DistillationTrainer训练时出现报错,核心问题是学生模型输出的outputs_student.loss是生成器对象而非张量,仅用logits计算的交叉熵部分运行正常。
相关代码
def distill_weights(teacher, student): """ Recursively copies the weights of the (teacher) to the (student). This function is meant to be first called on a RobertaFor... model, but is then called on every children of that model recursively. The only part that's not fully copied is the encoder, of which only half is copied. """ # If the part is an entire RoBERTa model or a RobertaFor..., unpack and iterate if isinstance(teacher, RobertaModel) or type(teacher).__name__.startswith('RobertaFor'): for teacher_part, student_part in zip(teacher.children(), student.children()): distill_weights(teacher_part, student_part) # Else if the part is an encoder, copy one out of every layer elif isinstance(teacher, RobertaEncoder): teacher_encoding_layers = [layer for layer in next(teacher.children())] student_encoding_layers = [layer for layer in next(student.children())] for i in range(len(student_encoding_layers)): student_encoding_layers[i].load_state_dict(teacher_encoding_layers[2*i].state_dict()) # Else the part is a head or something else, copy the state_dict else: student.load_state_dict(teacher.state_dict(), strict=False) def distill_roberta_based(teacher_model): """ Distilates a RoBERTa (teacher_model) like would DistilBERT for a BERT model. The student model has the same configuration, except for the number of hidden layers, which is // by 2. The student layers are initilized by copying one out of two layers of the teacher, starting with layer 0. The head of the teacher is also copied. """ # Set student configuration configuration = teacher_model.config.to_dict() configuration['num_hidden_layers'] //= 2 configuration = RobertaConfig.from_dict(configuration) # create student model student_model = type(teacher_model)(configuration) distill_weights(teacher=teacher_model, student=student_model) return student_model # 蒸馏模型训练类 class DistillationTrainer(Trainer): def __init__(self, *args, teacher_model=None, **kwargs): super().__init__(*args, **kwargs) self.teacher = teacher_model # 将教师模型移动到学生模型同一设备 self._move_model_to_device(self.teacher,self.model.device) self.teacher.eval() def compute_loss(self, model, inputs, return_outputs = False): """ 用于蒸馏BERT类模型的损失函数,结合教师logits、学生logits和标签计算多部分损失,可设置温度系数 """ outputs_student = model(**inputs) print(outputs_student) student_loss = outputs_student.loss # 计算教师模型输出(不计算梯度) with torch.no_grad(): outputs_teacher = self.teacher(**inputs) # 确保logits尺寸一致 assert outputs_student.logits.size() == outputs_teacher.logits.size() # 分类损失(任务特定损失) loss_function = CrossEntropyLoss() # 温度缩放与softmax student_logits = F.softmax(outputs_student.logits / self.args.temperature, dim=-1) teacher_logits = F.softmax(outputs_teacher.logits / self.args.temperature, dim=-1) loss_logits = loss_function(student_logits, teacher_logits) # 返回加权总损失 loss = self.args.alpha * student_loss + (1. - self.args.alpha) * loss_logits return (loss, outputs_student) if return_outputs else loss # 创建学生模型 student_model_adapter = distill_roberta_based(teacher_model) # 激活并训练Adapter student_model_adapter.set_active_adapters('parallel') student_model_adapter.train_adapter('parallel') # 初始化训练器 trainer = DistillationTrainer( student_model_adapter, training_args, teacher_model=teacher_model, train_dataset=tokenized_datasets["train"], eval_dataset=tokenized_datasets["validation"], data_collator=data_collator, tokenizer=tokenizer, compute_metrics=compute_metrics, ) trainer.args._n_gpu = 4
期望与实际输出
期望输出
SequenceClassifierOutput(loss=tensor([0.6899, 0.6902, 0.6926, 0.6913, 0.6906, 0.6904, 0.6922, 0.6917], device='cuda:0', grad_fn=<GatherBackward>), logits=tensor([[-1.2512e-03, -9.7885e-03], [ 6.2714e-03, -5.7755e-03],.....])
实际输出
SequenceClassifierOutput(loss=<generator object gather.<locals>.gather_map.<locals>.<genexpr> at 0x7f5bb4fbe9d0>, logits=tensor([[-0.0150, 0.0075], [-0.0122, 0.0181],...
问题原因
- Adapter训练模式的loss逻辑修改:当调用
train_adapter('parallel')时,Adapter框架会将模型的loss输出从单张量改为生成器对象。这是因为在Adapter训练模式下,框架需要单独计算Adapter参数的梯度,避免更新主模型的权重,生成器用于分步返回各个Adapter的损失分量。 - 未处理生成器类型的loss:你的
DistillationTrainer直接读取outputs_student.loss,但此时这个变量是未解析的生成器,而非可直接运算的张量,导致后续加权计算时出现类型不匹配的报错。 - 学生模型Adapter配置初始化不完整:你通过
distill_weights函数复制教师模型的Adapter权重,但学生模型是基于修改层数后的配置创建的,并未显式初始化Adapter的相关配置,这会导致模型的loss生成逻辑异常,进一步触发生成器类型的输出。
解决建议
- 解析生成器loss:将
student_loss = outputs_student.loss修改为student_loss = sum(outputs_student.loss),把生成器中的所有损失张量求和得到可用于计算的总损失。 - 显式初始化学生模型的Adapter:创建学生模型后,手动添加与教师模型一致的Adapter配置,而非仅通过权重复制,示例代码:
# 创建学生模型后添加Adapter配置 student_model_adapter.add_adapter("parallel", config=teacher_model.get_adapter_config("parallel")) - 验证训练模式配置:确认
train_adapter的调用正确,确保只有Adapter参数处于训练状态,主模型参数已冻结(若有需求)。
内容的提问来源于stack exchange,提问作者Mara de Jess Garcia Santiago
相关产品推荐
相关产品推荐

