You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

带Adapter的RoBERTa蒸馏:SequenceClassifierOutput损失为生成器问题

问题:带Adapter的RoBERTa模型蒸馏时loss为生成器而非张量的原因

背景

对带有Adapter的RoBERTa模型执行知识蒸馏,修改了权重蒸馏函数以让学生模型继承教师模型的Adapter参数,但使用自定义DistillationTrainer训练时出现报错,核心问题是学生模型输出的outputs_student.loss是生成器对象而非张量,仅用logits计算的交叉熵部分运行正常。

相关代码

def distill_weights(teacher, student):
    """
    Recursively copies the weights of the (teacher) to the (student).
    This function is meant to be first called on a RobertaFor... model, but is then called on every children of that model recursively.
    The only part that's not fully copied is the encoder, of which only half is copied.
    """
    # If the part is an entire RoBERTa model or a RobertaFor..., unpack and iterate
    if isinstance(teacher, RobertaModel) or type(teacher).__name__.startswith('RobertaFor'):
        for teacher_part, student_part in zip(teacher.children(), student.children()):
            distill_weights(teacher_part, student_part)
    # Else if the part is an encoder, copy one out of every layer
    elif isinstance(teacher, RobertaEncoder):
            teacher_encoding_layers = [layer for layer in next(teacher.children())]
            student_encoding_layers = [layer for layer in next(student.children())]
            for i in range(len(student_encoding_layers)):
                student_encoding_layers[i].load_state_dict(teacher_encoding_layers[2*i].state_dict())
    # Else the part is a head or something else, copy the state_dict
    else:
        student.load_state_dict(teacher.state_dict(), strict=False)


def distill_roberta_based(teacher_model):
    """
    Distilates a RoBERTa (teacher_model) like would DistilBERT for a BERT model.
    The student model has the same configuration, except for the number of hidden layers, which is // by 2.
    The student layers are initilized by copying one out of two layers of the teacher, starting with layer 0.
    The head of the teacher is also copied.
    """
    # Set student configuration
    configuration = teacher_model.config.to_dict()
    configuration['num_hidden_layers'] //= 2
    configuration = RobertaConfig.from_dict(configuration)
    
    # create student model
    student_model = type(teacher_model)(configuration)
    distill_weights(teacher=teacher_model, student=student_model)

    return student_model

# 蒸馏模型训练类
class DistillationTrainer(Trainer):
    def __init__(self, *args, teacher_model=None, **kwargs):
        super().__init__(*args, **kwargs)
        
        self.teacher = teacher_model
        # 将教师模型移动到学生模型同一设备
        self._move_model_to_device(self.teacher,self.model.device)
        self.teacher.eval()

    
    def compute_loss(self, model, inputs, return_outputs = False):
        """
        用于蒸馏BERT类模型的损失函数,结合教师logits、学生logits和标签计算多部分损失,可设置温度系数
        """
        outputs_student =  model(**inputs)
        print(outputs_student)
        student_loss    = outputs_student.loss
        
        # 计算教师模型输出(不计算梯度)
        with torch.no_grad():
            outputs_teacher = self.teacher(**inputs)
        
        # 确保logits尺寸一致
        assert outputs_student.logits.size() == outputs_teacher.logits.size()
                                

        # 分类损失(任务特定损失)
        loss_function = CrossEntropyLoss()
        
        # 温度缩放与softmax
        student_logits = F.softmax(outputs_student.logits / self.args.temperature, dim=-1)
        teacher_logits = F.softmax(outputs_teacher.logits / self.args.temperature, dim=-1)
        loss_logits = loss_function(student_logits, teacher_logits)

        # 返回加权总损失
        loss = self.args.alpha * student_loss + (1. - self.args.alpha) * loss_logits
        return (loss, outputs_student) if return_outputs else loss

# 创建学生模型
student_model_adapter = distill_roberta_based(teacher_model)
# 激活并训练Adapter
student_model_adapter.set_active_adapters('parallel')
student_model_adapter.train_adapter('parallel')  

# 初始化训练器
trainer = DistillationTrainer(
    student_model_adapter,
    training_args,
    teacher_model=teacher_model,
    train_dataset=tokenized_datasets["train"],
    eval_dataset=tokenized_datasets["validation"],
    data_collator=data_collator,
    tokenizer=tokenizer,
    compute_metrics=compute_metrics,
)
trainer.args._n_gpu = 4

期望与实际输出

期望输出

SequenceClassifierOutput(loss=tensor([0.6899, 0.6902, 0.6926, 0.6913, 0.6906, 0.6904, 0.6922, 0.6917],
       device='cuda:0', grad_fn=<GatherBackward>), logits=tensor([[-1.2512e-03, -9.7885e-03],
        [ 6.2714e-03, -5.7755e-03],.....])

实际输出

SequenceClassifierOutput(loss=<generator object gather.<locals>.gather_map.<locals>.<genexpr> at 0x7f5bb4fbe9d0>, logits=tensor([[-0.0150,  0.0075],
        [-0.0122,  0.0181],...

问题原因

  1. Adapter训练模式的loss逻辑修改:当调用train_adapter('parallel')时,Adapter框架会将模型的loss输出从单张量改为生成器对象。这是因为在Adapter训练模式下,框架需要单独计算Adapter参数的梯度,避免更新主模型的权重,生成器用于分步返回各个Adapter的损失分量。
  2. 未处理生成器类型的loss:你的DistillationTrainer直接读取outputs_student.loss,但此时这个变量是未解析的生成器,而非可直接运算的张量,导致后续加权计算时出现类型不匹配的报错。
  3. 学生模型Adapter配置初始化不完整:你通过distill_weights函数复制教师模型的Adapter权重,但学生模型是基于修改层数后的配置创建的,并未显式初始化Adapter的相关配置,这会导致模型的loss生成逻辑异常,进一步触发生成器类型的输出。

解决建议

  • 解析生成器loss:将student_loss = outputs_student.loss修改为student_loss = sum(outputs_student.loss),把生成器中的所有损失张量求和得到可用于计算的总损失。
  • 显式初始化学生模型的Adapter:创建学生模型后,手动添加与教师模型一致的Adapter配置,而非仅通过权重复制,示例代码:
    # 创建学生模型后添加Adapter配置
    student_model_adapter.add_adapter("parallel", config=teacher_model.get_adapter_config("parallel"))
    
  • 验证训练模式配置:确认train_adapter的调用正确,确保只有Adapter参数处于训练状态,主模型参数已冻结(若有需求)。

内容的提问来源于stack exchange,提问作者Mara de Jess Garcia Santiago

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 07:55:35