You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

FastAPI部署情感分析API遇ValueError:特征数不匹配求助

问题:FastAPI部署情感分析模型时特征维度不匹配报错

问题详情

部署情感分析模型后,用Postman测试时持续收到internal server error,终端报错显示特征维度不匹配。相关代码、测试数据及报错信息如下:

部署代码

from fastapi import FastAPI
from pydantic import BaseModel
import pickle
import json
import numpy as np
app = FastAPI()

class ModelInput(BaseModel):
    review: str  
    
# Loading the saved model
sentiment_model = pickle.load(open('sentiment_analyzer_new.sav', 'rb'))

@app.post('/sentiment_prediction')
async def sentiment_pred(input_parameters: ModelInput):
    
    rev = input_parameters.review  
    
    input_list = [rev]
    
    # Reshape the input data to have a shape of (1, -1)
    input_array = np.array(input_list).reshape(1, -1)
    
    # Print for debugging
    print("Review Text:", rev)
    print("Input Array Shape:", input_array.shape)
    
    predictions = sentiment_model.predict(input_array)
    
    # Assuming 'predictions' is an array of predictions
    if predictions[0] == 1:
        sentiment = 'Positive'
    else:
        sentiment = 'Negative'
    
    response = {
        "Review": rev,
        "Sentiment": sentiment
    }
    
    return response

测试用JSON

{
    "review": "This is a positive review."
}

注:原测试JSON末尾多了逗号,属于无效格式,需删除

报错信息

Traceback (most recent call last):
  File "c:\users\user\anaconda3\lib\site-packages\uvicorn\protocols\http\h11_impl.py", line 408, in run_asgi
    result = await app(  # type: ignore[func-returns-value]
  File "c:\users\user\anaconda3\lib\site-packages\uvicorn\middleware\proxy_headers.py", line 84, in __call__
    return await self.app(scope, receive, send)
  File "c:\users\user\anaconda3\lib\site-packages\fastapi\applications.py", line 289, in __call__
    await super().__call__(scope, receive, send)
  File "c:\users\user\anaconda3\lib\site-packages\starlette\applications.py", line 122, in __call__
    await self.middleware_stack(scope, receive, send)
  File "c:\users\user\anaconda3\lib\site-packages\starlette\middleware\errors.py", line 184, in __call__
    raise exc
  File "c:\users\user\anaconda3\lib\site-packages\starlette\middleware\errors.py", line 162, in __call__
    await self.app(scope, receive, _send)
  File "c:\users\user\anaconda3\lib\site-packages\starlette\middleware\exceptions.py", line 79, in __call__
    raise exc
  File "c:\users\user\anaconda3\lib\site-packages\starlette\middleware\exceptions.py", line 68, in __call__
    await self.app(scope, receive, sender)
  File "c:\users\user\anaconda3\lib\site-packages\fastapi\middleware\asyncexitstack.py", line 20, in __call__
    raise e
  File "c:\users\user\anaconda3\lib\site-packages\fastapi\middleware\asyncexitstack.py", line 17, in __call__
    await self.app(scope, receive, send)
  File "c:\users\user\anaconda3\lib\site-packages\starlette\routing.py", line 718, in __call__
    await route.handle(scope, receive, send)
  File "c:\users\user\anaconda3\lib\site-packages\starlette\routing.py", line 276, in handle
    await self.app(scope, receive, send)
  File "c:\users\user\anaconda3\lib\site-packages\starlette\routing.py", line 66, in app
    response = await func(request)
  File "c:\users\user\anaconda3\lib\site-packages\fastapi\routing.py", line 273, in app
    raw_response = await run_endpoint_function(
  File "c:\users\user\anaconda3\lib\site-packages\fastapi\routing.py", line 190, in run_endpoint_function
    return await dependant.call(**values)
  File "C:\Users\USER\Desktop\Sentiment analysis\api.py", line 28, in sentiment_pred
    predictions = sentiment_model.predict(input_array)
  File "c:\users\user\anaconda3\lib\site-packages\sklearn\linear_model\_base.py", line 307, in predict
    scores = self.decision_function(X)
  File "c:\users\user\anaconda3\lib\site-packages\sklearn\linear_model\_base.py", line 286, in decision_function
    raise ValueError("X has %d features per sample; expecting %d"
ValueError: X has 1 features per sample; expecting 6209089

错误原因

你只序列化保存了模型本身,但训练模型时用到的**文本特征提取器(如TF-IDF Vectorizer、CountVectorizer)**没有一同保存。模型训练时是将文本转换成了6209089维的特征向量,现在直接传入原始文本,即使reshape也只是将单个字符串转为1维数组,完全不符合模型期望的特征维度。

解决步骤

1. 重新保存模型与特征提取器

训练模型时,将特征提取器和模型打包成字典一起序列化保存:

from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
import pickle

# 示例训练流程(根据你的实际训练代码调整)
vectorizer = TfidfVectorizer()
X_train_vec = vectorizer.fit_transform(X_train)  # X_train为训练文本数据
model = LogisticRegression()
model.fit(X_train_vec, y_train)  # y_train为训练标签

# 保存特征提取器和模型
with open('sentiment_model_with_vectorizer.sav', 'wb') as f:
    pickle.dump({'vectorizer': vectorizer, 'model': model}, f)

2. 修改FastAPI部署代码

加载包含特征提取器和模型的文件,先用提取器处理输入文本,再传入模型预测:

from fastapi import FastAPI
from pydantic import BaseModel
import pickle

app = FastAPI()

class ModelInput(BaseModel):
    review: str  

# 加载包含特征提取器和模型的文件
with open('sentiment_model_with_vectorizer.sav', 'rb') as f:
    saved_objects = pickle.load(f)
vectorizer = saved_objects['vectorizer']
sentiment_model = saved_objects['model']

@app.post('/sentiment_prediction')
async def sentiment_pred(input_parameters: ModelInput):
    rev = input_parameters.review  
    # 用特征提取器将文本转换为模型所需的特征向量
    input_vec = vectorizer.transform([rev])
    # 直接预测(transform返回的是符合维度要求的稀疏矩阵)
    predictions = sentiment_model.predict(input_vec)
    
    sentiment = 'Positive' if predictions[0] == 1 else 'Negative'
    
    return {
        "Review": rev,
        "Sentiment": sentiment
    }

3. 额外注意

  • 确保训练和部署使用的特征提取器完全一致,避免特征维度偏差。
  • 测试时使用格式正确的JSON,删除原测试数据末尾的多余逗号。

内容的提问来源于stack exchange,提问作者Apata Oyinlade

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 00:01:05