如何解决TensorFlow Recommenders模型保存失败的问题?
解决TensorFlow Recommenders模型保存失败问题
问题描述
保存TensorFlow Recommenders框架的自定义子类模型时出现报错,无法使用HDF5格式保存模型。
报错信息
File "C:\Users\guera\AppData\Local\Programs\Python\Python310\lib\site-packages\keras\utils\traceback_utils.py", line 67, in error_handler raise e.with_traceback(filtered_tb) from None File "C:\Users\guera\AppData\Local\Programs\Python\Python310\lib\site-packages\keras\saving\save.py", line 142, in save_model raise NotImplementedError(NotImplementedError: Saving the model to HDF5 format requires the model to be a Functional model or a Sequential model. It does not work for subclassed models, because such models are defined via the body of a Python method, which isn't safely serializable. Consider saving to the Tensorflow SavedModel format (by setting save_format="tf") or using save_weights.
相关代码
主函数代码
import os import pprint import tempfile from typing import Dict, Text import numpy as np import tensorflow as tf import tensorflow_datasets as tfds import joblib import tensorflow_recommenders as tfrs from src.models.movie_lens import MovieLensModel def content_based_filtering(user_id, table1, table2, user_id_label, target_label): # Ratings data. ratings = tfds.load(table1, split="train") # Features of all the available movies. movies = tfds.load(table2, split="train") # Select the basic features. ratings = ratings.map(lambda x: { "{}".format(target_label): x[target_label], "{}".format(user_id_label): x[user_id_label] }) movies = movies.map(lambda x: x[target_label]) user_ids_vocabulary = tf.keras.layers.StringLookup(mask_token=None) user_ids_vocabulary.adapt(ratings.map(lambda x: x[user_id_label])) movie_titles_vocabulary = tf.keras.layers.StringLookup(mask_token=None) movie_titles_vocabulary.adapt(movies) # Define user and movie models. user_model = tf.keras.Sequential([ user_ids_vocabulary, tf.keras.layers.Embedding(user_ids_vocabulary.vocab_size(), 64) ]) movie_model = tf.keras.Sequential([ movie_titles_vocabulary, tf.keras.layers.Embedding(movie_titles_vocabulary.vocab_size(), 64) ]) # Define your objectives. task = tfrs.tasks.Retrieval(metrics=tfrs.metrics.FactorizedTopK( movies.batch(128).map(movie_model) ) ) # Create a retrieval model. model = MovieLensModel(user_model, movie_model, task) model.compile(optimizer=tf.keras.optimizers.Adagrad(0.5)) # Train for 3 epochs. model.fit(ratings.batch(4096), epochs=3) model.save('content_model.h5') # Use brute-force search to set up retrieval using the trained representations. index = tfrs.layers.factorized_top_k.BruteForce(model.user_model) index.index_from_dataset( movies.batch(100).map(lambda title: (title, model.movie_model(title)))) # Get some recommendations. _, titles = index(np.array([str(user_id)])) # print('Recommendation content based filtering\n') return titles[0, :3].numpy()
MovieLensModel子类代码
import os import pprint import tempfile from typing import Dict, Text import numpy as np import tensorflow as tf import tensorflow_datasets as tfds import tensorflow_recommenders as tfrs class MovieLensModel(tfrs.Model): def __init__(self, user_model, movie_model, task): super().__init__() self.movie_model: tf.keras.Model = movie_model self.user_model: tf.keras.Model = user_model self.task: tf.keras.layers.Layer = task def compute_loss(self, features: Dict[Text, tf.Tensor], training=False) -> tf.Tensor: # We pick out the user features and pass them into the user model. user_embeddings = self.user_model(features["user_id"]) # And pick out the movie features and pass them into the movie model, # getting embeddings back. positive_movie_embeddings = self.movie_model(features["movie_title"]) # The task computes the loss and the metrics. return self.task(user_embeddings, positive_movie_embeddings)
解决方案
报错已明确说明原因:HDF5(.h5)格式仅支持Functional或Sequential模型,无法保存通过Python方法定义的子类模型。提供两种解决方式:
方式一:使用TensorFlow SavedModel格式保存完整模型
将原代码中的model.save('content_model.h5')替换为:
# SavedModel格式会生成一个文件夹,无需后缀 model.save('content_model', save_format='tf')
或直接省略save_format参数(默认就是tf格式):
model.save('content_model')
加载模型时使用:
loaded_model = tf.keras.models.load_model('content_model')
方式二:仅保存模型权重
如果只需要保存模型的训练参数,不需要完整结构,可以使用权重保存:
保存权重
model.save_weights('content_model_weights.h5')
加载权重
需要先重新构建与训练时完全一致的模型结构,再加载权重:
# 重复训练时的模型构建步骤:创建词汇表、user_model、movie_model、task ratings = tfds.load(table1, split="train") movies = tfds.load(table2, split="train") # ... 省略其他构建步骤,确保和训练时完全相同 ... user_model = tf.keras.Sequential([ user_ids_vocabulary, tf.keras.layers.Embedding(user_ids_vocabulary.vocab_size(), 64) ]) movie_model = tf.keras.Sequential([ movie_titles_vocabulary, tf.keras.layers.Embedding(movie_titles_vocabulary.vocab_size(), 64) ]) task = tfrs.tasks.Retrieval(metrics=tfrs.metrics.FactorizedTopK( movies.batch(128).map(movie_model) )) # 实例化模型并编译 loaded_model = MovieLensModel(user_model, movie_model, task) loaded_model.compile(optimizer=tf.keras.optimizers.Adagrad(0.5)) # 加载权重 loaded_model.load_weights('content_model_weights.h5')
额外注意:保存检索索引
如果需要复用BruteForce检索索引,也需要单独保存:
# 保存索引 index.save('retrieval_index') # 加载索引 loaded_index = tf.keras.models.load_model('retrieval_index')
内容的提问来源于stack exchange,提问作者Amoungui Serge
相关产品推荐
相关产品推荐

