调用Keras to_categorical将字符串转独热编码时报错的技术问询
问题描述
需要将以下字符串转换为分类形式或独热编码:
string1 = "Interstitial markings are diffusely prominent throughout both lungs. Heart size is normal. Pulmonary XXXX normal." st1 = string1.split()
使用以下代码尝试转换时出现错误:
from numpy import array from numpy import argmax from keras.utils import to_categorical # define example data = array(st1) print(data) encoded = to_categorical(data) print(encoded) # invert encoding inverted = argmax(encoded[0]) print(inverted)
错误信息:
['Interstitial' 'markings' 'are' 'diffusely' 'prominent' 'throughout' 'both' 'lungs.' 'Heart' 'size' 'is' 'normal.' 'Pulmonary' 'XXXX''normal.'] --------------------------------------------------------------------------- ValueError Traceback (most recent call last) <ipython-input-15-b034d9393342> in <module> 5 data = array(st1) 6 print(data) ----> 7 encoded = to_categorical(data) 8 print(encoded) 9 # invert encoding /usr/local/lib/python3.7/dist-packages/keras/utils/np_utils.py in to_categorical(y, num_classes, dtype) 60 [0. 0. 0. 0.] 61 """ ---> 62 y = np.array(y, dtype='int') 63 input_shape = y.shape 64 if input_shape and input_shape[-1] == 1 and len(input_shape) > 1: ValueError: invalid literal for int() with base 10: 'Interstitial'
错误原因
keras.utils.to_categorical的输入必须是整数类型的标签,它无法直接处理字符串数据。报错是因为代码试图将字符串自动转换为整数,而字符串无法被解析为十进制整数。
解决方案
要实现字符串到独热编码的转换,需要先将字符串映射为整数标签,再进行独热编码。以下是两种可行的实现方式:
方法1:LabelEncoder + to_categorical
先用sklearn.preprocessing.LabelEncoder把字符串转成整数标签,再用to_categorical生成独热编码:
from numpy import array, argmax from sklearn.preprocessing import LabelEncoder from keras.utils import to_categorical # 原始数据处理 string1 = "Interstitial markings are diffusely prominent throughout both lungs. Heart size is normal. Pulmonary XXXX normal." st1 = string1.split() data = array(st1) # 1. 将字符串转换为整数标签 label_encoder = LabelEncoder() integer_labels = label_encoder.fit_transform(data) # 2. 对整数标签做独热编码 onehot_encoded = to_categorical(integer_labels) print("独热编码结果:") print(onehot_encoded) # 反转编码:从独热编码还原回原始字符串 first_encoded_idx = argmax(onehot_encoded[0]) original_string = label_encoder.inverse_transform([first_encoded_idx])[0] print("\n还原第一个编码的原始字符串:", original_string)
方法2:使用sklearn的OneHotEncoder
直接用sklearn.preprocessing.OneHotEncoder完成转换(注意需要先转整数标签,因为OneHotEncoder也不直接处理字符串):
from numpy import array from sklearn.preprocessing import OneHotEncoder, LabelEncoder # 原始数据处理 string1 = "Interstitial markings are diffusely prominent throughout both lungs. Heart size is normal. Pulmonary XXXX normal." st1 = string1.split() data = array(st1).reshape(-1, 1) # OneHotEncoder要求输入为二维数组 # 1. 字符串转整数标签 label_encoder = LabelEncoder() integer_labels = label_encoder.fit_transform(data.flatten()).reshape(-1, 1) # 2. 生成独热编码 onehot_encoder = OneHotEncoder(sparse_output=False) onehot_encoded = onehot_encoder.fit_transform(integer_labels) print("独热编码结果:") print(onehot_encoded) # 反转编码 first_encoded_idx = onehot_encoded[0].argmax() original_string = label_encoder.inverse_transform([first_encoded_idx])[0] print("\n还原第一个编码的原始字符串:", original_string)
注意事项
- 重复出现的字符串(比如示例中的
normal.)会被映射为同一个整数标签,对应的独热编码也完全相同。 - 独热编码的维度等于数据中唯一字符串的数量,每个位置对应一个唯一的字符串。
内容的提问来源于stack exchange,提问作者user19686684
相关产品推荐
相关产品推荐

