TensorFlow使用自定义数据集报错:NumPy数组转Tensor不支持int对象类型
报错根因与修复方案
核心错误点
- 错误1:你手动将
Pclass、SibSp、Parch三个特征转为object类型,TensorFlow无法直接将存储为object类型的整数值转为张量,触发Unsupported object type int报错。 - 错误2:
categorical_column_with_vocabulary_list生成的分类特征列不能直接传入LinearClassifier,必须用tf.feature_column.indicator_column包裹为one-hot编码格式才能被识别。 - 错误3:
make_input_fn返回逻辑不符合estimator接口要求,train/evaluate方法需要接收无参可调用的输入函数,你当前直接返回Dataset对象,会触发调用错误。 - 错误4:分类特征的词汇表你用了全量数据集的唯一值,存在训练数据泄露问题,正确做法是仅用训练集的唯一值构建词汇表。
修复后可运行代码
import tensorflow as tf import pandas as pd import numpy as np # 构造数据集 df = {'Survived': [0, 1, 1, 1, 0], 'Pclass': [3, 1, 3, 1, 3], 'Sex': ['male', 'female', 'female', 'female', 'male'], 'Age': [22.0, 38.0, 26.0, 35.0, 35.0], 'SibSp': [1, 1, 0, 1, 0], 'Parch': [0, 0, 0, 0, 0], 'Fare': [7.2500, 71.2833, 7.9250, 53.1000, 8.0500], 'Embarked': ['S', 'C', 'S', 'S', 'S']} df = pd.DataFrame(df) df.dropna(inplace=True) # 不需要将数值类离散特征转为object,保留原int类型即可 # df['Pclass'] = df['Pclass'].astype('object') # df['SibSp'] = df['SibSp'].astype('object') # df['Parch'] = df['Parch'].astype('object') # 切分数据集 train, test = np.split(df.sample(frac=1, random_state=42), [int(0.8*len(df))]) y_train_labels = train.pop('Survived') y_test_labels = test.pop('Survived') numerical_columns = ['Age','Fare'] categorical_columns = ['Sex','Embarked','Pclass','Parch','SibSp'] feature_column = [] for feature in categorical_columns: # 仅用训练集数据构建词汇表,避免数据泄露 vocabulary = train[feature].unique() cat_col = tf.feature_column.categorical_column_with_vocabulary_list(feature, vocabulary) # 用indicator_column包裹分类特征列 feature_column.append(tf.feature_column.indicator_column(cat_col)) for feature in numerical_columns: feature_column.append(tf.feature_column.numeric_column(feature, dtype=tf.float32)) def make_input_fn(data_df, label_df, num_epochs=20, shuffle=True, batch_size=32): def input_function(): ds = tf.data.Dataset.from_tensor_slices((dict(data_df), label_df)) if shuffle: ds = ds.shuffle(1000) ds = ds.batch(batch_size).repeat(num_epochs) return ds # 返回函数本身,不要提前调用 return input_function train_input_fn = make_input_fn(train, y_train_labels) eval_input_fn = make_input_fn(test, y_test_labels, num_epochs=1, shuffle=False) linear_est = tf.estimator.LinearClassifier(feature_columns=feature_column) linear_est.train(train_input_fn) result = linear_est.evaluate(eval_input_fn) print(result)
其他潜在问题
- 当前提供的最小复现样本量过少,仅5条数据,训练后的模型无实际预测价值,替换为全量泰坦尼克号数据集后效果会正常。
- 切分数据集时建议固定random_state,保证实验可复现。
- 如果后续分类特征的词汇量过大,可以替换
indicator_column为embedding_column降低特征维度。
内容的提问来源于stack exchange,提问作者Jack
相关产品推荐
相关产品推荐

