Docker环境加载Keras自定义LSTM模型.pkl文件失败求助
Docker中加载含Keras自定义LSTM模型的Scikit-learn Pipeline时权重丢失
我用Keras开发了一个双输出的自定义LSTM模型:一个输出是连续预测值,另一个是多分类标签。为简化流程,我把模型封装进Scikit-learn Pipeline,并用joblib保存整个流水线。由于Scikit-learn的Keras wrapper只支持序列模型,所以我自己实现了自定义模型,本地运行完全正常,还搭了验证API和小型网站测试。
但构建Docker镜像后,加载流水线时出现权重丢失的错误:
Layer 'lstm_cell' expected 3 variables, but received 0 variables during loading. Expected: ['lstm/lstm_cell/kernel:0', 'lstm/lstm_cell/recurrent_kernel:0', 'lstm/lstm_cell/bias:0']
我的项目文件夹结构:
. └── END2END/ ├── pipelines/ │ ├── scripts │ ├── __init__.py │ └── myscripts.py ├── train_models/ │ ├── __init__.py │ ├── script2.py │ └── my_model_.pkl ├── api.py └── Dockerfile
模型定义代码
def ModelLSTM(input_shape, hidden_units: int, output_1: int, output_2: int, lr: float, model_name: str) -> keras.Model: """ LSTM model - single multioutput model ### Arguments: - input_shape : input shapes - hidden_units : hidden units (LSTM units) - output_1 : number of outputs (forecasting outputs) - output_2 : number of outputs (failures outputs) - lr : learning rate - name: name of the model ### Returns: model (keras.Model)""" # Data Input inputs_ = keras.Input( shape=(input_shape[1], input_shape[2]), name='inputs' ) # Lstm Layer hiddens_layers = layers.LSTM(hidden_units, return_sequences=True)(inputs_) # output forecasting output_forecast = layers.Dense(output_1, activation='linear', name='forecast')(hiddens_layers) # output failures output_failures = layers.Dense(output_2, activation='sigmoid', name='failures')(hiddens_layers) # Create Model model = keras.Model(inputs=inputs_, outputs=[output_forecast, output_failures], name=model_name) model.compile(loss={ 'forecast': 'mean_squared_error', 'failures': custom_multilabel_loss}, optimizer= tf.keras.optimizers.Adam(learning_rate=lr), metrics ={ 'forecast': 'mae', 'failures': 'accuracy'}) return model
Dockerfile代码
# Variables ARG WorkingDirectory=/project # Pull Docker Image FROM python:3.11-bullseye # Set up Working Environment WORKDIR $WorkingDirectory # Create an "application" directory RUN mkdir -p $WorkingDirectory/application RUN export PYTHONPATH=$WorkingDirectory # Copy Files COPY . $WorkingDirectory/application # Update Requirements RUN pip install --upgrade pip RUN pip install --no-cache-dir -r $WorkingDirectory/application/requirements/requirements.txt EXPOSE 8000 # Start API Server CMD ["sh", "-c", "python ${ADD_DIR}/application/main.py"]
流水线保存与加载代码
def save_pipeline(*, pipeline_to_saved: Pipeline, name: str) -> None: """ Saved the versioned model and overwrite any previos saved models, this ensure there is only one trained model that can be called ### Arguments: - pipeline_to_save (`sklearn.Pipeline`) : pipeline already trained """ # Versioned file name save_file_name = f"{name}_{_version}.pkl" save_path = TRAINED_MODEL_DIR / save_file_name # Re-write or delete model file remove_old_pipelines(files_to_keep=save_file_name) # saved joblib.dump(pipeline_to_saved, save_path) def load_pipeline(*, file_name: str) -> Pipeline: """ Load Pipeline from trained model ### Arguments: - file_name (str): pipeline file name ### Returns: train_pipeline : Pipeline object """ file_path = TRAINED_MODEL_DIR / f"{file_name}_{_version}.pkl" print(file_path) trained_pipeline = joblib.load(filename=file_path) return trained_pipeline
我已经尝试过:
- 重新训练两次模型,本地加载、运行、预测都完全正常
- 将pkl文件移到api.py同目录,甚至硬编码文件路径,问题依旧
我知道可以拆分流水线,用keras.load_weights单独加载权重再插入流水线,但既然本地完全正常,想弄明白Docker环境下为什么会出现这个问题,是Docker处理pkl文件的方式有问题,还是需要先压缩文件?
内容的提问来源于stack exchange,提问作者Marvin Garcia
相关产品推荐
相关产品推荐

