You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow Recommenders训练报错:期望字符串却获int64张量

问题解决步骤

1. 核心问题分析

报错根源:

  • 训练数据集ratings_tf仅包含movieId(int64类型)和userId,缺失电影标题特征
  • movie_model依赖字符串类型的电影标题输入,但compute_loss方法中错误传入了整个features字典(含int类型的movieId),导致StringLookup层触发类型不匹配错误

2. 数据预处理修正:关联ratings与电影标题

首先需要将ratings中的movieId与movies_df的电影标题关联,确保训练数据包含标题字段:

# 建立movieId到带年份标题的映射
movies_df['full_title'] = movies_df['title'] + ' (' + movies_df['year'].astype(str) + ')'
movie_id_to_title = movies_df.set_index('id')['full_title'].to_dict()

# 给ratings添加标题字段并过滤无效记录
ratings['full_title'] = ratings['movieId'].map(movie_id_to_title)
ratings = ratings.dropna(subset=['full_title'])
ratings['userId'] = ratings['userId'].astype(str)

3. 修正训练数据集构造

调整ratings_tf的构造,保留userId和full_title的字典结构:

ratings_tf = tf.data.Dataset.from_tensor_slices(dict(ratings[['userId', 'full_title']]))

4. 修正模型的compute_loss方法

修改MovieLensModel中的损失计算逻辑,让movie_model接收正确的标题特征:

def compute_loss(self, features: Dict[Text, tf.Tensor], training=False) -> tf.Tensor:
    user_embeddings = self.user_model(features["userId"])
    # 传入标题特征而非整个features字典
    movie_embeddings = self.movie_model(features["full_title"])
    return self.task(user_embeddings, movie_embeddings)

5. 修正电影词汇表适配数据

确保movie_titles_vocabulary适配带年份的标题数据集:

movies_tf = tf.data.Dataset.from_tensor_slices(dict(movies_df[['full_title']])).map(lambda x: x["full_title"])
movie_titles_vocabulary = tf.keras.layers.StringLookup(mask_token=None)
movie_titles_vocabulary.adapt(movies_tf)

6. 完整修正后的核心代码

# 数据预处理
movies_df['release_date'] = pd.to_datetime(movies_df['release_date'])
movies_df['year'] = movies_df['release_date'].dt.year
movies_df['full_title'] = movies_df['title'] + ' (' + movies_df['year'].astype(str) + ')'

movie_id_to_title = movies_df.set_index('id')['full_title'].to_dict()
ratings['full_title'] = ratings['movieId'].map(movie_id_to_title)
ratings = ratings.dropna(subset=['full_title'])
ratings['userId'] = ratings['userId'].astype(str)

# 构造数据集
ratings_tf = tf.data.Dataset.from_tensor_slices(dict(ratings[['userId', 'full_title']]))
movies_tf = tf.data.Dataset.from_tensor_slices(dict(movies_df[['full_title']])).map(lambda x: x["full_title"])

# 词汇表适配
user_ids_vocabulary = tf.keras.layers.StringLookup(mask_token=None)
user_ids_vocabulary.adapt(ratings_tf.map(lambda x: x["userId"]))

movie_titles_vocabulary = tf.keras.layers.StringLookup(mask_token=None)
movie_titles_vocabulary.adapt(movies_tf)

# 模型定义
class MovieLensModel(tfrs.Model):
    def __init__(self, user_model: tf.keras.Model, movie_model: tf.keras.Model, task: tfrs.tasks.Retrieval):
        super().__init__()
        self.user_model = user_model
        self.movie_model = movie_model
        self.task = task

    def compute_loss(self, features: Dict[Text, tf.Tensor], training=False) -> tf.Tensor:
        user_embeddings = self.user_model(features["userId"])
        movie_embeddings = self.movie_model(features["full_title"])
        return self.task(user_embeddings, movie_embeddings)

user_model = tf.keras.Sequential([
    user_ids_vocabulary,
    tf.keras.layers.Embedding(user_ids_vocabulary.vocabulary_size(), 64)
])

movie_model = tf.keras.Sequential([
    movie_titles_vocabulary,
    tf.keras.layers.Embedding(movie_titles_vocabulary.vocabulary_size(), 64)
])

task = tfrs.tasks.Retrieval(metrics=tfrs.metrics.FactorizedTopK(movies_tf.map(movie_model)))
model = MovieLensModel(user_model, movie_model, task)
model.compile(optimizer=tf.keras.optimizers.Adagrad(0.5))

# 训练模型
model.fit(ratings_tf.batch(4096), epochs=3)

内容的提问来源于stack exchange,提问作者Stevi G

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 22:35:20