在Vertex AI导入Scikit-learn模型时遇AttributeError求解决
解决Vertex AI导入Scikit-learn Pipeline时的
AttributeError: Can't get attribute 'preprocess_text'错误 问题根源
你的Scikit-learn Pipeline依赖自定义的preprocess_text函数,序列化(pickle)模型时,Pipeline仅保存该函数的模块引用(如utils.preprocess_text),而非函数代码本身。Vertex AI的默认服务脚本(model_server.py)加载模型时,无法在服务环境中找到该函数——要么你的utils模块未被包含在部署包中,要么模块路径与本地环境不一致。
修复步骤
1. 标准化自定义函数的模块结构
- 将
preprocess_text函数单独存入utils.py文件,确保逻辑完整:# utils.py def preprocess_text(text): # 此处编写你的预处理逻辑,比如小写转换、去停用词等 processed_text = text.lower().strip() # 其他处理步骤... return processed_text - 训练脚本中从该模块导入函数:
from utils import preprocess_text complaints_clf_pipeline = Pipeline( [ ("preprocess", text.TfidfVectorizer(preprocessor=preprocess_text, ngram_range=(1, 2))), ("clf", naive_bayes.MultinomialNB(alpha=0.3)), ] )
2. 打包模型与依赖模块
训练完成后保存模型,确保utils.py与模型文件处于同一目录:
import joblib # 保存训练好的Pipeline joblib.dump(complaints_clf_pipeline, "./model.joblib")
此时目录结构应为:
./ ├── model.joblib └── utils.py
3. 部署时包含依赖模块
将上述目录打包为ZIP文件(确保ZIP根目录直接包含model.joblib和utils.py,不要嵌套子目录),再上传至Vertex AI模型资源:
- 用gcloud命令行部署时,指定
--artifact-uri指向包含这两个文件的GCS路径,或直接上传ZIP包。 - 用Vertex AI控制台部署时,上传ZIP文件需保证内部结构符合要求。
4. (可选)使用自定义预测脚本增强可控性
若默认服务仍有问题,可编写predict.py脚本明确模型加载逻辑:
# predict.py import joblib import os from utils import preprocess_text MODEL_FILENAME = "model.joblib" def load_model(): model_path = os.path.join(os.environ["AIP_STORAGE_URI"], MODEL_FILENAME) return joblib.load(model_path) def predict(instances, model): # 解析输入实例中的文本 texts = [instance["text"] for instance in instances] predictions = model.predict(texts) # 格式化返回结果 return [{"prediction": pred} for pred in predictions]
部署时指定该脚本作为预测入口,确保utils.py、predict.py和model.joblib在同一部署包中。
常见坑点
- 不要直接在Jupyter Notebook代码单元中定义
preprocess_text并序列化模型,Notebook的模块环境与Vertex AI服务环境不一致,会导致引用失效。 - 确保
utils.py无本地特有库或路径依赖,额外依赖需通过requirements.txt指定,保证Vertex AI服务环境能安装。
内容的提问来源于stack exchange,提问作者VertexEnjoyer
相关产品推荐
相关产品推荐

