TensorFlow Recommenders训练报错:期望字符串却获int64张量
问题解决步骤
1. 核心问题分析
报错根源:
- 训练数据集
ratings_tf仅包含movieId(int64类型)和userId,缺失电影标题特征 movie_model依赖字符串类型的电影标题输入,但compute_loss方法中错误传入了整个features字典(含int类型的movieId),导致StringLookup层触发类型不匹配错误
2. 数据预处理修正:关联ratings与电影标题
首先需要将ratings中的movieId与movies_df的电影标题关联,确保训练数据包含标题字段:
# 建立movieId到带年份标题的映射 movies_df['full_title'] = movies_df['title'] + ' (' + movies_df['year'].astype(str) + ')' movie_id_to_title = movies_df.set_index('id')['full_title'].to_dict() # 给ratings添加标题字段并过滤无效记录 ratings['full_title'] = ratings['movieId'].map(movie_id_to_title) ratings = ratings.dropna(subset=['full_title']) ratings['userId'] = ratings['userId'].astype(str)
3. 修正训练数据集构造
调整ratings_tf的构造,保留userId和full_title的字典结构:
ratings_tf = tf.data.Dataset.from_tensor_slices(dict(ratings[['userId', 'full_title']]))
4. 修正模型的compute_loss方法
修改MovieLensModel中的损失计算逻辑,让movie_model接收正确的标题特征:
def compute_loss(self, features: Dict[Text, tf.Tensor], training=False) -> tf.Tensor: user_embeddings = self.user_model(features["userId"]) # 传入标题特征而非整个features字典 movie_embeddings = self.movie_model(features["full_title"]) return self.task(user_embeddings, movie_embeddings)
5. 修正电影词汇表适配数据
确保movie_titles_vocabulary适配带年份的标题数据集:
movies_tf = tf.data.Dataset.from_tensor_slices(dict(movies_df[['full_title']])).map(lambda x: x["full_title"]) movie_titles_vocabulary = tf.keras.layers.StringLookup(mask_token=None) movie_titles_vocabulary.adapt(movies_tf)
6. 完整修正后的核心代码
# 数据预处理 movies_df['release_date'] = pd.to_datetime(movies_df['release_date']) movies_df['year'] = movies_df['release_date'].dt.year movies_df['full_title'] = movies_df['title'] + ' (' + movies_df['year'].astype(str) + ')' movie_id_to_title = movies_df.set_index('id')['full_title'].to_dict() ratings['full_title'] = ratings['movieId'].map(movie_id_to_title) ratings = ratings.dropna(subset=['full_title']) ratings['userId'] = ratings['userId'].astype(str) # 构造数据集 ratings_tf = tf.data.Dataset.from_tensor_slices(dict(ratings[['userId', 'full_title']])) movies_tf = tf.data.Dataset.from_tensor_slices(dict(movies_df[['full_title']])).map(lambda x: x["full_title"]) # 词汇表适配 user_ids_vocabulary = tf.keras.layers.StringLookup(mask_token=None) user_ids_vocabulary.adapt(ratings_tf.map(lambda x: x["userId"])) movie_titles_vocabulary = tf.keras.layers.StringLookup(mask_token=None) movie_titles_vocabulary.adapt(movies_tf) # 模型定义 class MovieLensModel(tfrs.Model): def __init__(self, user_model: tf.keras.Model, movie_model: tf.keras.Model, task: tfrs.tasks.Retrieval): super().__init__() self.user_model = user_model self.movie_model = movie_model self.task = task def compute_loss(self, features: Dict[Text, tf.Tensor], training=False) -> tf.Tensor: user_embeddings = self.user_model(features["userId"]) movie_embeddings = self.movie_model(features["full_title"]) return self.task(user_embeddings, movie_embeddings) user_model = tf.keras.Sequential([ user_ids_vocabulary, tf.keras.layers.Embedding(user_ids_vocabulary.vocabulary_size(), 64) ]) movie_model = tf.keras.Sequential([ movie_titles_vocabulary, tf.keras.layers.Embedding(movie_titles_vocabulary.vocabulary_size(), 64) ]) task = tfrs.tasks.Retrieval(metrics=tfrs.metrics.FactorizedTopK(movies_tf.map(movie_model))) model = MovieLensModel(user_model, movie_model, task) model.compile(optimizer=tf.keras.optimizers.Adagrad(0.5)) # 训练模型 model.fit(ratings_tf.batch(4096), epochs=3)
内容的提问来源于stack exchange,提问作者Stevi G
相关产品推荐
相关产品推荐

