本地加载经save_pretrained保存的Huggingface模型调用报错求助
问题描述
能成功运行以下Hugging Face代码完成文本分类:
from transformers import pipeline, TFAutoModel classifier = pipeline(task="text-classification", model="SamLowe/roberta-base-go_emotions", top_k=None) sentences = ["This has been a good day"] classifier(sentences)
但将模型保存到本地后,用TFAutoModel加载并直接预测时触发错误:
classifier.save_pretrained("local_go_emotions") model = TFAutoModel.from_pretrained("local_go_emotions") model(sentences)
错误信息:
Data of type <class 'str'> is not allowed only (<class 'tensorflow.python.framework.tensor.Tensor'>, <class 'bool'>, <class 'int'>, <class 'transformers.utils.generic.ModelOutput'>, <class 'tuple'>, <class 'list'>, <class 'dict'>, <class 'numpy.ndarray'>) is accepted for input_ids.
错误原因
- 封装层级差异:
pipeline是一站式封装工具,内部集成了文本转张量的tokenizer和适配分类任务的模型;而TFAutoModel仅加载基础预训练模型,无法直接处理原始字符串,必须传入经过tokenizer处理后的张量数据。 - 模型类选错:文本分类任务需要加载
TFAutoModelForSequenceClassification(适配分类任务的模型类),而非通用的TFAutoModel——后者仅输出模型隐藏层特征,不输出分类结果。 - 缺少文本预处理:直接给模型传字符串列表,模型无法识别,必须先通过tokenizer将文本转换为
input_ids、attention_mask等模型可接受的张量格式。
解决方法
方法一:用pipeline直接加载本地模型(最简便)
保存后直接复用pipeline加载本地模型,继续使用封装好的预测接口:
from transformers import pipeline # 加载本地保存的模型 classifier = pipeline(task="text-classification", model="local_go_emotions", top_k=None) sentences = ["This has been a good day"] # 直接执行预测 print(classifier(sentences))
方法二:手动加载tokenizer与适配模型类(更灵活)
若需要自定义处理流程,可手动加载tokenizer和分类模型,完成文本预处理后再预测:
from transformers import TFAutoModelForSequenceClassification, AutoTokenizer import tensorflow as tf # 加载tokenizer和分类专用模型 tokenizer = AutoTokenizer.from_pretrained("local_go_emotions") model = TFAutoModelForSequenceClassification.from_pretrained("local_go_emotions") sentences = ["This has been a good day"] # 将文本转换为模型所需的张量格式 inputs = tokenizer(sentences, return_tensors="tf", padding=True, truncation=True) # 执行预测并转换为概率值 outputs = model(**inputs) predictions = tf.nn.softmax(outputs.logits, axis=-1) print(predictions)
内容的提问来源于stack exchange,提问作者Saurabh Verma
相关产品推荐
相关产品推荐

