如何解决TensorFlow中int64转string不支持的报错?
问题:TensorFlow Recommenders数据处理时类型转换报错
我想基于TensorFlow Recommenders搭建一个小型推荐系统,按照官方指引开发,但在数据处理环节遇到报错。我从远程数据库获取数据,相关代码如下:
数据库读取引擎代码
import sqlalchemy as engines import tensorflow as tf import pandas as pd class BaseEngine(): """BaseEngine class""" def __init__(self, url: str): self.engines = engines.create_engine(url) # 获取数据表 def get_table(self, name): data = pd.read_sql_query("SELECT * from {}".format(name), self.engines) # data.to_csv('data/{}'.format(name)) return data # 提取多特征 def multi_features(self, tablename: str): data = tf.data.Dataset.from_tensor_slices(dict(self.get_table(tablename))) return data # 提取特征 def feature(self, tablename: str): data = tf.data.Dataset.from_tensor_slices(dict(self.get_table(tablename))) return data
主程序代码
from random import shuffle from myengine.engines.ormengines.mysql_engine import MysqlEngine import tensorflow as tf import numpy as np engine = MysqlEngine('sqlite:///base.db') ratings = engine.feature('ratings') movies = engine.feature('movies') for x in ratings.take(1).as_numpy_iterator(): print(x) for x in movies.take(1).as_numpy_iterator(): print(x) # 数据拆分 tf.random.set_seed(42) shuffled = ratings.shuffle(100_000, seed=42, reshuffle_each_iteration=False) train = shuffled.take(80_000) test = shuffled.skip(80_000).take(20_000) def iter_x(e): return { 'userId': e['userId'], 'movieId': e['movieId'] } def iter_y(e): return { 'title': e['title'], 'movieId': e['movieId'] } def mapping(x): return x['userId'] # 数据分批与特征映射 movies = movies.batch(1_000).map(iter_y) ratings = ratings.batch(1_000_000).map(iter_x) print(ratings) print(movies) user_ids_vocabulary = tf.keras.layers.StringLookup(mask_token=None) user_ids_vocabulary.adapt(ratings.map(mapping))
报错信息
2022-08-19 22:58:28.366801: W tensorflow/core/framework/op_kernel.cc:1722] OP_REQUIRES failed at cast_op.cc:121 : UNIMPLEMENTED: Cast int64 to string is not supported Traceback (most recent call last): File "C:\Users\guera\OneDrive\Documents\recommender\project\app.py", line 48, in <module> user_ids_vocabulary.adapt(ratings.map(mapping)) File "C:\Users\guera\AppData\Local\Programs\Python\Python310\lib\site-packages\keras\layers\preprocessing\string_lookup.py", line 396, in adapt super().adapt(data, batch_size=batch_size, steps=steps) File "C:\Users\guera\AppData\Local\Programs\Python\Python310\lib\site-packages\keras\engine\base_preprocessing_layer.py", line 249, in adapt self._adapt_function(iterator) File "C:\Users\guera\AppData\Local\Programs\Python\Python310\lib\site-packages\tensorflow\python\util\traceback_utils.py", line 153, in error_handler raise e.with_traceback(filtered_tb) from None File "C:\Users\guera\AppData\Local\Programs\Python\Python310\lib\site-packages\tensorflow\python\eager\execute.py", line 54, in quick_execute tensors = pywrap_tfe.TFE_Py_Execute(ctx._handle, device_name, op_name, tensorflow.python.framework.errors_impl.UnimplementedError: Graph execution error: Detected at node 'Cast' defined at (most recent call last): File "C:\Users\guera\OneDrive\Documents\recommender\project\app.py", line 48, in <module> user_ids_vocabulary.adapt(ratings.map(mapping)) File "C:\Users\guera\AppData\Local\Programs\Python\Python310\lib\site-packages\keras\layers\preprocessing\string_lookup.py", line 396, in adapt super().adapt(data, batch_size=batch_size, steps=steps) File "C:\Users\guera\AppData\Local\Programs\Python\Python310\lib\site-packages\keras\engine\base_preprocessing_layer.py", line 249, in adapt self._adapt_function(iterator) File "C:\Users\guera\AppData\Local\Programs\Python\Python310\lib\site-packages\keras\engine\base_preprocessing_layer.py", line 118, in adapt_step self.update_state(data) File "C:\Users\guera\AppData\Local\Programs\Python\Python310\lib\site-packages\keras\layers\preprocessing\index_lookup.py", line 531, in update_state data = utils.ensure_tensor(data, dtype=self.vocabulary_dtype) File "C:\Users\guera\AppData\Local\Programs\Python\Python310\lib\site-packages\keras\layers\preprocessing\preprocessing_utils.py", line 33, in ensure_tensor inputs = tf.cast(inputs, dtype) Node: 'Cast' Cast int64 to string is not supported [[{{node Cast}}]] [Op:__inference_adapt_step_117]
解决方案
报错核心是**Cast int64 to string is not supported**:你使用StringLookup层处理userId,但userId是int64数值类型,TensorFlow不支持直接将整数强制转为字符串类型。
方法一:将userId转换为字符串类型
可以在数据读取阶段或Dataset处理阶段把userId转成字符串:
- 在数据库读取时转换:
def get_table(self, name): data = pd.read_sql_query("SELECT * from {}".format(name), self.engines) # 针对ratings表转换userId类型 if name == 'ratings': data['userId'] = data['userId'].astype(str) return data
- 或者在Dataset的map操作中转换:
def mapping(x): return tf.strings.as_string(x['userId'])
方法二:改用IntegerLookup层处理数值ID
如果userId本身是整数类型,没必要转成字符串,直接用专门处理整数的IntegerLookup层更高效:
# 替换原有的StringLookup为IntegerLookup user_ids_vocabulary = tf.keras.layers.IntegerLookup(mask_token=None) user_ids_vocabulary.adapt(ratings.map(mapping))
推荐使用方法二,因为IntegerLookup是为整数ID设计的预处理层,更贴合数据类型,性能也更好。
内容的提问来源于stack exchange,提问作者Amoungui Serge
相关产品推荐
相关产品推荐

