RuntimeError报错求助:torch.cat()期望非空张量列表
解决RuntimeError: torch.cat(): expected a non-empty list of Tensors错误
这个错误的核心原因是你要拼接的token_id或attention_masks列表是空的,导致torch.cat没有可拼接的张量。以下是排查和解决步骤:
1. 排查根本原因
- 先检查
text列表是否为空:打印len(text),如果结果是0,说明没有输入数据,循环一次都没执行,两个列表自然是空的。 - 再检查
text里的样本是否全是空字符串:如果所有样本都是空的,循环里的处理逻辑无法生成有效张量,也会导致列表为空。
2. 修改代码解决问题
在你的代码中添加校验和过滤逻辑,避免空列表进入torch.cat:
# 先校验输入文本列表是否为空 if not text: raise ValueError("输入的text列表为空,请确认数据是否正确加载") token_id = [] attention_masks = [] def preprocessing(input_text, tokenizer): ''' Returns <class transformers.tokenization_utils_base.BatchEncoding> with the following fields: - input_ids: list of token ids - token_type_ids: list of token type ids - attention_mask: list of indices (0,1) specifying which tokens should considered by the model (return_attention_mask = True). ''' return tokenizer.encode_plus( input_text, add_special_tokens = True, max_length = 32, pad_to_max_length = True, return_attention_mask = True, return_tensors = 'pt' ) for sample in text: # 过滤空字符串样本,避免无效处理 if not sample.strip(): continue encoding_dict = preprocessing(sample, tokenizer) token_id.append(encoding_dict['input_ids']) attention_masks.append(encoding_dict['attention_mask']) # 拼接前再次校验,确保有有效张量 if not token_id or not attention_masks: raise ValueError("没有有效样本被处理,请检查text中的内容是否为非空字符串") token_id = torch.cat(token_id, dim = 0) attention_masks = torch.cat(attention_masks, dim = 0) labels = torch.tensor(labels)
3. 额外注意事项
- 确保
labels的长度和处理后的token_id行数一致,否则后续训练会出现维度不匹配的错误。 - 如果使用较新版本的Hugging Face Tokenizer,
pad_to_max_length已被弃用,建议替换为padding='max_length'。
内容的提问来源于stack exchange,提问作者Abdulhadi Alamoudi
相关产品推荐
相关产品推荐

