训练LayoutLMv2(DataParallel)遇RuntimeError:输入输出索引设备不匹配
问题场景
在AWS SageMaker单节点4GPU实例上,使用torch.nn.DataParallel训练LayoutLMv2模型时,执行代码行outputs = model(**{key: torch.squeeze(value.to(self.training_device), 1) for key, value in inputs.items()})触发错误:
RuntimeError: Input, output and indices must be on the current device
已知inputs是包含dict_keys(['input_ids', 'token_type_ids', 'attention_mask', 'bbox', 'labels', 'image'])的字典,无法直接调用inputs = inputs.to(self.training_device),且已尝试将模型和输入移至GPU但问题依旧。
自定义Trainer类代码如下:
class LanguageModelTrainer(Trainer): def __init__(self, *args, class_weights=None, l1_coef=0, **kwargs): self.class_weights = class_weights self.l1_coef = l1_coef self.training_device = torch.device("cuda:0" if torch.cuda.is_available() else "cpu") super().__init__(*args, **kwargs) def compute_loss(self, model, inputs, return_outputs=False): model = torch.nn.DataParallel(model) model.to(self.training_device) loss_fn = torch.nn.CrossEntropyLoss(weight=self.class_weights) outputs = model(**{key: torch.squeeze(value.to(self.training_device), 1) for key, value in inputs.items()}) loss = loss_fn(outputs.logits.view(-1, model.config.num_labels), inputs["labels"].view(-1)) if self.l1_coef: l1_norm = sum(torch.linalg.norm(p.flatten(), 1) for n, p in model.named_parameters() if 'bias' not in n) loss += self.l1_coef * l1_norm return (loss, outputs) if return_outputs else loss
错误原因分析
- DataParallel重复包装:在
compute_loss内每次重新用DataParallel包装模型,会导致模型参数分布混乱,设备映射逻辑冲突。 - 张量设备不一致:仅将输入张量移至GPU,但计算损失时使用的原始
inputs["labels"]仍在CPU,引发设备不匹配错误。 - 模型设备配置逻辑错误:
model.to(self.training_device)仅指定单GPU,但DataParallel默认会使用所有可用GPU,且重复迁移设备会打乱参数分布。
修复方案
- 把
DataParallel模型包装逻辑移至__init__方法,避免每次计算损失时重复创建,确保模型参数稳定分布在多GPU上。 - 确保所有输入张量(包括labels)都移至训练设备,计算损失时使用已迁移的标签张量。
- 移除
compute_loss内重复的模型设备迁移操作,初始化阶段完成设备配置即可。 - 访问DataParallel包装后的模型参数和配置时,需通过
model.module层级获取。
修改后的完整代码
class LanguageModelTrainer(Trainer): def __init__(self, *args, class_weights=None, l1_coef=0, **kwargs): self.class_weights = class_weights self.l1_coef = l1_coef self.training_device = torch.device("cuda:0" if torch.cuda.is_available() else "cpu") super().__init__(*args, **kwargs) # 初始化时完成DataParallel包装,适配多GPU环境 if torch.cuda.device_count() > 1: self.model = torch.nn.DataParallel(self.model) self.model.to(self.training_device) # 将类别权重同步到训练设备 if self.class_weights is not None: self.class_weights = self.class_weights.to(self.training_device) def compute_loss(self, model, inputs, return_outputs=False): loss_fn = torch.nn.CrossEntropyLoss(weight=self.class_weights) # 统一处理所有输入张量的设备迁移与维度压缩 device_inputs = {} for key, value in inputs.items(): device_inputs[key] = torch.squeeze(value.to(self.training_device), 1) outputs = model(**device_inputs) # 使用已迁移到GPU的labels计算损失 loss = loss_fn(outputs.logits.view(-1, model.module.config.num_labels), device_inputs["labels"].view(-1)) if self.l1_coef: # 访问DataParallel包装后的模型参数需通过model.module l1_norm = sum(torch.linalg.norm(p.flatten(), 1) for n, p in model.module.named_parameters() if 'bias' not in n) loss += self.l1_coef * l1_norm return (loss, outputs) if return_outputs else loss
内容的提问来源于stack exchange,提问作者MichiganMagician
相关产品推荐
相关产品推荐

