解决'AttributeError: torch.Size对象无as_list属性'的NLP推理问题
问题原因
AttributeError: 'torch.Size' object has no attribute 'as_list' 的核心原因是张量框架不兼容:
- 你用HuggingFace的
AutoTokenizer生成了PyTorch张量(通过return_tensors='pt'参数) - 但加载的是Keras(TensorFlow)模型,Keras无法识别PyTorch的
torch.Size对象,它需要的是TensorFlow张量或numpy数组。
另外还有两个次要问题需要同步处理:
- 提示的序列长度超过模型限制(3002>256),需要确保
max_len和模型训练时的输入长度一致,且不超过PhoBERT的默认256 - 代码里手动覆盖了模型输出
output = torch.randn(1, 3),这会导致模型推理结果被丢弃,完全没用
修复后的完整代码
推理函数修改
def infer(text, tokenizer, max_len=256): # 改成256适配PhoBERT默认长度 class_names = ['thế giới', 'thể thao', 'văn hóa', 'vi tính'] model = keras.models.load_model('./models/cnn_nlp_text_classification_4_classer.h5') encoded_review = tokenizer.encode_plus( text, max_length=max_len, truncation=True, add_special_tokens=True, padding='max_length', return_attention_mask=True, return_token_type_ids=False, return_tensors='tf', # 改成返回TensorFlow张量,或者用'np'返回numpy数组 ) input_ids = encoded_review['input_ids'] attention_mask = encoded_review['attention_mask'] # 移除手动生成随机张量的代码,使用真实模型输出 output = model([input_ids, attention_mask]) # Keras模型输出是TensorFlow张量,转成numpy数组再处理 y_pred = output.numpy().argmax(axis=1)[0] print(f'Text: {text}') print(f'Category: {class_names[y_pred]}')
调用代码
from tensorflow import keras from transformers import AutoTokenizer # 补充导入AutoTokenizer tokenizer = AutoTokenizer.from_pretrained("vinai/phobert-base") infer('Chào mừng mọi người đến với bản tin thể thao 360 tối nay !', tokenizer)
关键修改点
- 张量类型统一:将
return_tensors='pt'改为'tf'(返回TensorFlow张量)或'np'(返回numpy数组),确保输入和模型框架匹配 - 修正序列长度:将
max_len设为256,匹配PhoBERT预训练模型的默认最大序列长度,避免索引错误提示 - 恢复真实模型输出:删除
output = torch.randn(1, 3)这行无效代码,使用模型实际推理结果 - 适配Keras输出处理:Keras模型输出是TensorFlow张量,用
.numpy()转成numpy数组后,用argmax获取预测类别索引
内容的提问来源于stack exchange,提问作者michaelkenshijpv97
相关产品推荐
相关产品推荐

