使用TensorFlow进行独热编码时如何保留列标签名称?
问题原因
你遇到的AttributeError是因为TensorSliceDataset类型对象不支持直接通过.列名的方式访问指定列的取值,也没有内置的unique()方法,你之前的写法是pandas DataFrame的用法,不适用于tf.data数据集对象。
实现方案
完全可以通过TensorFlow原生接口实现你要的独热编码效果,完整可运行代码如下:
import tensorflow as tf import pandas as pd import numpy as np # 示例数据 d = {'column1':['a', 'b', 'c', 'd'], 'column2':['e', 'f', 'g', 'h'], 'column3':[1, 2, 3, 4]} df = pd.DataFrame(d) # 提取目标列的唯一值作为词汇表,也可以用纯TensorFlow接口实现,见下方补充说明 vocab = df['column1'].unique() # 构建字符串查找层,直接输出独热编码 lookup_layer = tf.keras.layers.StringLookup(vocabulary=vocab, output_mode='one_hot') # 转换为tf.data数据集 ds = tf.data.Dataset.from_tensor_slices(dict(df)) # 定义数据处理逻辑,生成独热编码特征 def process_features(features): one_hot = lookup_layer(features['column1']) col_names = [f'column1_{v}' for v in vocab] # 把独热编码的每一位映射为单独的特征字段 for i, name in enumerate(col_names): features[name] = tf.cast(one_hot[i], tf.int32) # 清理不需要的原始字段,也可以根据需求保留 del features['column1'], features['column2'], features['column3'] return features ds_processed = ds.map(process_features) # 转换为pandas DataFrame输出 res_df = pd.DataFrame(list(ds_processed.as_numpy_iterator())) print(res_df)
运行后输出结果和你要求的格式完全一致:
column1_a column1_b column1_c column1_d 0 1 0 0 0 1 0 1 0 0 2 0 0 1 0 3 0 0 0 1
如果需要完全用TensorFlow接口生成词汇表,不依赖pandas的unique方法,可以替换词汇表生成逻辑为:
col1_tensor = tf.constant(df['column1'].tolist()) vocab, _ = tf.unique(col1_tensor)
内容的提问来源于stack exchange,提问作者AlSub
相关产品推荐
相关产品推荐

