FastAPI部署情感分析API遇ValueError:特征数不匹配求助
问题:FastAPI部署情感分析模型时特征维度不匹配报错
问题详情
部署情感分析模型后,用Postman测试时持续收到internal server error,终端报错显示特征维度不匹配。相关代码、测试数据及报错信息如下:
部署代码
from fastapi import FastAPI from pydantic import BaseModel import pickle import json import numpy as np app = FastAPI() class ModelInput(BaseModel): review: str # Loading the saved model sentiment_model = pickle.load(open('sentiment_analyzer_new.sav', 'rb')) @app.post('/sentiment_prediction') async def sentiment_pred(input_parameters: ModelInput): rev = input_parameters.review input_list = [rev] # Reshape the input data to have a shape of (1, -1) input_array = np.array(input_list).reshape(1, -1) # Print for debugging print("Review Text:", rev) print("Input Array Shape:", input_array.shape) predictions = sentiment_model.predict(input_array) # Assuming 'predictions' is an array of predictions if predictions[0] == 1: sentiment = 'Positive' else: sentiment = 'Negative' response = { "Review": rev, "Sentiment": sentiment } return response
测试用JSON
{ "review": "This is a positive review." }
注:原测试JSON末尾多了逗号,属于无效格式,需删除
报错信息
Traceback (most recent call last): File "c:\users\user\anaconda3\lib\site-packages\uvicorn\protocols\http\h11_impl.py", line 408, in run_asgi result = await app( # type: ignore[func-returns-value] File "c:\users\user\anaconda3\lib\site-packages\uvicorn\middleware\proxy_headers.py", line 84, in __call__ return await self.app(scope, receive, send) File "c:\users\user\anaconda3\lib\site-packages\fastapi\applications.py", line 289, in __call__ await super().__call__(scope, receive, send) File "c:\users\user\anaconda3\lib\site-packages\starlette\applications.py", line 122, in __call__ await self.middleware_stack(scope, receive, send) File "c:\users\user\anaconda3\lib\site-packages\starlette\middleware\errors.py", line 184, in __call__ raise exc File "c:\users\user\anaconda3\lib\site-packages\starlette\middleware\errors.py", line 162, in __call__ await self.app(scope, receive, _send) File "c:\users\user\anaconda3\lib\site-packages\starlette\middleware\exceptions.py", line 79, in __call__ raise exc File "c:\users\user\anaconda3\lib\site-packages\starlette\middleware\exceptions.py", line 68, in __call__ await self.app(scope, receive, sender) File "c:\users\user\anaconda3\lib\site-packages\fastapi\middleware\asyncexitstack.py", line 20, in __call__ raise e File "c:\users\user\anaconda3\lib\site-packages\fastapi\middleware\asyncexitstack.py", line 17, in __call__ await self.app(scope, receive, send) File "c:\users\user\anaconda3\lib\site-packages\starlette\routing.py", line 718, in __call__ await route.handle(scope, receive, send) File "c:\users\user\anaconda3\lib\site-packages\starlette\routing.py", line 276, in handle await self.app(scope, receive, send) File "c:\users\user\anaconda3\lib\site-packages\starlette\routing.py", line 66, in app response = await func(request) File "c:\users\user\anaconda3\lib\site-packages\fastapi\routing.py", line 273, in app raw_response = await run_endpoint_function( File "c:\users\user\anaconda3\lib\site-packages\fastapi\routing.py", line 190, in run_endpoint_function return await dependant.call(**values) File "C:\Users\USER\Desktop\Sentiment analysis\api.py", line 28, in sentiment_pred predictions = sentiment_model.predict(input_array) File "c:\users\user\anaconda3\lib\site-packages\sklearn\linear_model\_base.py", line 307, in predict scores = self.decision_function(X) File "c:\users\user\anaconda3\lib\site-packages\sklearn\linear_model\_base.py", line 286, in decision_function raise ValueError("X has %d features per sample; expecting %d" ValueError: X has 1 features per sample; expecting 6209089
错误原因
你只序列化保存了模型本身,但训练模型时用到的**文本特征提取器(如TF-IDF Vectorizer、CountVectorizer)**没有一同保存。模型训练时是将文本转换成了6209089维的特征向量,现在直接传入原始文本,即使reshape也只是将单个字符串转为1维数组,完全不符合模型期望的特征维度。
解决步骤
1. 重新保存模型与特征提取器
训练模型时,将特征提取器和模型打包成字典一起序列化保存:
from sklearn.feature_extraction.text import TfidfVectorizer from sklearn.linear_model import LogisticRegression import pickle # 示例训练流程(根据你的实际训练代码调整) vectorizer = TfidfVectorizer() X_train_vec = vectorizer.fit_transform(X_train) # X_train为训练文本数据 model = LogisticRegression() model.fit(X_train_vec, y_train) # y_train为训练标签 # 保存特征提取器和模型 with open('sentiment_model_with_vectorizer.sav', 'wb') as f: pickle.dump({'vectorizer': vectorizer, 'model': model}, f)
2. 修改FastAPI部署代码
加载包含特征提取器和模型的文件,先用提取器处理输入文本,再传入模型预测:
from fastapi import FastAPI from pydantic import BaseModel import pickle app = FastAPI() class ModelInput(BaseModel): review: str # 加载包含特征提取器和模型的文件 with open('sentiment_model_with_vectorizer.sav', 'rb') as f: saved_objects = pickle.load(f) vectorizer = saved_objects['vectorizer'] sentiment_model = saved_objects['model'] @app.post('/sentiment_prediction') async def sentiment_pred(input_parameters: ModelInput): rev = input_parameters.review # 用特征提取器将文本转换为模型所需的特征向量 input_vec = vectorizer.transform([rev]) # 直接预测(transform返回的是符合维度要求的稀疏矩阵) predictions = sentiment_model.predict(input_vec) sentiment = 'Positive' if predictions[0] == 1 else 'Negative' return { "Review": rev, "Sentiment": sentiment }
3. 额外注意
- 确保训练和部署使用的特征提取器完全一致,避免特征维度偏差。
- 测试时使用格式正确的JSON,删除原测试数据末尾的多余逗号。
内容的提问来源于stack exchange,提问作者Apata Oyinlade
相关产品推荐
相关产品推荐

