You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Conformer音频模型训练报错:无法创建张量问题求助

解决Conformer训练中的ValueError:input_ids过度嵌套问题

问题根源

你遇到的错误核心是:prepare_dataset函数返回的input_ids和labels是带批量维度的二维张量,而Hugging Face数据集要求每个样本的特征是一维序列(或一维张量),这导致DataCollator合并批量时出现嵌套结构(list中包含二维张量而非一维序列),无法创建统一形状的批量张量。

具体来说,你调用tokenizer时用了return_tensors="pt",直接返回了[batch_size, seq_len]的二维张量,但数据集每个样本只应对应单个序列的一维数据;另外手动对音频做pad_sequence的操作,也和processor的内置逻辑冲突,容易引发额外问题。

修复方案

1. 修改prepare_dataset函数

移除return_tensors="pt",让tokenizer返回列表形式的序列;同时去掉手动音频padding,交给processor统一处理:

MAX_TRANSCRIPTION_LENGTH = 128

def prepare_dataset(batch):
    speech, transcription = batch["audio"], batch["text"]
    
    # 直接返回原始音频数组,交给processor处理padding/truncation
    input_values = [s["array"] for s in speech]
    
    with tokenizer.as_target_tokenizer():
        # 去掉return_tensors="pt",返回一维列表形式的序列
        tokenized_output = tokenizer(
            transcription,
            padding=False,  # 禁用当前padding,交给DataCollator统一处理
            truncation=True,
            add_special_tokens=True,
            max_length=MAX_TRANSCRIPTION_LENGTH
        )
    
    return {
        "input_values": input_values,
        "labels": tokenized_output["input_ids"]
    }

2. 微调DataCollator(可选但推荐)

在标签处理部分加上truncation=True,确保标签长度严格统一:

def __call__(self, features: List[Dict[str, Union[List[int], torch.Tensor]]]) -> Dict[str, torch.Tensor]:
    if not features:
        return {}
    
    input_features = [{"input_values": feature["input_values"]} for feature in features]
    label_features = [{"input_ids": feature["labels"]} for feature in features]

    # 处理音频输入的padding/truncation
    batch = self.processor.pad(
        input_features,
        padding=self.padding,
        max_length=self.max_length,
        truncation=True,
        pad_to_multiple_of=self.pad_to_multiple_of,
        return_tensors="pt",
    )
    
    # 处理标签的padding/truncation,新增truncation=True
    with self.processor.as_target_processor():
        labels_batch = self.processor.pad(
            label_features,
            padding=self.padding,
            max_length=self.max_length_labels,
            truncation=True,
            pad_to_multiple_of=self.pad_to_multiple_of_labels,
            return_tensors="pt",
        )

    # 替换padding为-100,避免计算损失
    labels = labels_batch["input_ids"].masked_fill(labels_batch.attention_mask.ne(1), -100)
    batch["labels"] = labels

    return batch

关键说明

  • 去掉return_tensors="pt"后,tokenizer为每个样本返回一维整数列表,符合Hugging Face数据集要求,DataCollator可正确合并成批量张量。
  • 音频的padding/truncation完全交给processor处理,避免手动操作和processor逻辑冲突,确保输入格式统一。
  • 移除冗余的两次tokenizer调用,一次调用即可生成所需的标签序列。

内容的提问来源于stack exchange,提问作者moonface16

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 20:55:44