BERT模型训练调用model.fit()抛出“Invalid dtype: object”错误求助
BERT模型训练调用model.fit()抛出“Invalid dtype: object”错误求助
问题描述
我尝试用Keras NLP的BERT模型做二分类任务,代码能正常运行到model.compile(),但调用model.fit()时抛出了ValueError: Invalid dtype: object错误。我的代码和错误信息如下:
模型构建代码:
from tf_keras.src.layers.serialization import activation import keras_nlp import tensorflow as tf # bert layers text_input = tf.keras.layers.Input(shape=(), dtype=tf.string ,name="text") preprocessor = keras_nlp.models.BertPreprocessor.from_preset("bert_base_en_uncased",trainable=True) preprocessed_text = preprocessor(text_input) encoder = keras_nlp.models.BertBackbone.from_preset("bert_base_en_uncased") outputs = encoder(preprocessed_text) # neural network layers l = tf.keras.layers.Dropout(0.1, name='dropout')(outputs['pooled_output']) l = tf.keras.layers.Dense(1, activation='sigmoid', name='output')(l) # construct final model model = tf.keras.Model(inputs=[text_input], outputs=[l]) model.summary() METRICS = [ tf.keras.metrics.BinaryAccuracy(name='accuracy'), tf.keras.metrics.Precision(name='prediction'), tf.keras.metrics.Recall(name='recall') ] model.compile(optimizer='adam', loss='binary_crossentropy', metrics=METRICS)
训练代码及错误:
model_bert=model.fit(X_train, y_train, epochs=10)
错误栈:
--------------------------------------------------------------------------- ValueError Traceback (most recent call last) /var/folders/my/psclr1fj1jq4cwk4tl2w9wzr0000gn/T/ipykernel_68633/736105977.py in <module> 1 # X_train = tf.data.Dataset.from_tensor_slices(X_train) 2 # y_train = tf.data.Dataset.from_tensor_slices(y_train) ----> 3 model_bert=model.fit(X_train, 4 y_train, 5 epochs=10) ~/opt/anaconda3/lib/python3.9/site-packages/keras/src/utils/traceback_utils.py in error_handler(*args, **kwargs) 120 # To get the full stack trace, call: 121 # `keras.config.disable_traceback_filtering()` --> 122 raise e.with_traceback(filtered_tb) from None 123 finally: 124 del filtered_tb ~/opt/anaconda3/lib/python3.9/site-packages/optree/ops.py in tree_map(func, tree, is_leaf, none_is_leaf, namespace, *rests) 750 leaves, treespec = _C.flatten(tree, is_leaf, none_is_leaf, namespace) 751 flat_args = [leaves] + [treespec.flatten_up_to(r) for r in rests] --> 752 return treespec.unflatten(map(func, *flat_args)) 753 754 ValueError: Invalid dtype: object
我一直在查资料但没找到解决办法,这是我学习LLM的测试项目,求帮忙解决最后这个问题!
解决方案
这个错误的核心原因是你的训练数据X_train或y_train的 dtype 是object,而TensorFlow无法直接处理这种类型的数据。下面是具体的修复步骤:
1. 先排查数据类型
先打印出数据的类型,确认问题:
print("X_train dtype:", X_train.dtype if hasattr(X_train, 'dtype') else type(X_train)) print("y_train dtype:", y_train.dtype if hasattr(y_train, 'dtype') else type(y_train))
如果输出里出现object,就需要做类型转换。
2. 转换输入数据为TensorFlow兼容的类型
处理文本数据
X_train:
如果X_train是pandas的object类型Series/列表,需要转换成字符串类型的Tensor:import tensorflow as tf import pandas as pd # 若用pandas,先转成string类型列 if isinstance(X_train, pd.Series): X_train = X_train.astype('string').values # 转换成TensorFlow字符串张量 X_train = tf.convert_to_tensor(X_train, dtype=tf.string)处理标签数据
y_train:
二分类的标签需要是数值类型(float32或int32),不能是object:y_train = tf.convert_to_tensor(y_train, dtype=tf.float32)
3. 推荐用tf.data.Dataset加载数据(更稳定)
对于文本任务,用TensorFlow的Dataset API包装数据会更兼容,也方便后续做batch、shuffle等操作:
# 包装成Dataset train_dataset = tf.data.Dataset.from_tensor_slices((X_train, y_train)) # 设置batch大小(根据你的显存调整) train_dataset = train_dataset.batch(32) # 用Dataset训练模型 model_bert = model.fit(train_dataset, epochs=10)
额外注意点
- 你的代码里把
BertPreprocessor设为trainable=True,但预处理器主要负责分词、生成输入ID等静态操作,一般不需要训练,建议改成trainable=False,避免不必要的问题。 - 如果你的数据量很大,还可以给Dataset加
shuffle和prefetch操作,提升训练效率:train_dataset = train_dataset.shuffle(buffer_size=1000).batch(32).prefetch(tf.data.AUTOTUNE)
备注:内容来源于stack exchange,提问作者MrBDude
相关产品推荐
相关产品推荐

