You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

调用Keras to_categorical将字符串转独热编码时报错的技术问询

问题描述

需要将以下字符串转换为分类形式或独热编码:

string1 = "Interstitial markings are diffusely prominent throughout both lungs. Heart size is normal. Pulmonary XXXX normal."
st1 = string1.split()

使用以下代码尝试转换时出现错误:

from numpy import array
from numpy import argmax
from keras.utils import to_categorical
# define example
data = array(st1)
print(data)
encoded = to_categorical(data)
print(encoded)
# invert encoding
inverted = argmax(encoded[0])
print(inverted) 

错误信息:

['Interstitial' 'markings' 'are' 'diffusely' 'prominent' 'throughout' 'both' 'lungs.' 'Heart' 'size' 'is' 'normal.' 'Pulmonary' 'XXXX''normal.']
---------------------------------------------------------------------------
ValueError                                Traceback (most recent call last)
<ipython-input-15-b034d9393342> in <module>
  5 data = array(st1)
  6 print(data)
----> 7 encoded = to_categorical(data)
  8 print(encoded)
  9 # invert encoding

/usr/local/lib/python3.7/dist-packages/keras/utils/np_utils.py in to_categorical(y, num_classes, dtype)
 60   [0. 0. 0. 0.]
 61   """
---&gt; 62   y = np.array(y, dtype='int')
 63   input_shape = y.shape
 64   if input_shape and input_shape[-1] == 1 and len(input_shape) > 1:

ValueError: invalid literal for int() with base 10: 'Interstitial'
错误原因

keras.utils.to_categorical的输入必须是整数类型的标签,它无法直接处理字符串数据。报错是因为代码试图将字符串自动转换为整数,而字符串无法被解析为十进制整数。

解决方案

要实现字符串到独热编码的转换,需要先将字符串映射为整数标签,再进行独热编码。以下是两种可行的实现方式:

方法1:LabelEncoder + to_categorical

先用sklearn.preprocessing.LabelEncoder把字符串转成整数标签,再用to_categorical生成独热编码:

from numpy import array, argmax
from sklearn.preprocessing import LabelEncoder
from keras.utils import to_categorical

# 原始数据处理
string1 = "Interstitial markings are diffusely prominent throughout both lungs. Heart size is normal. Pulmonary XXXX normal."
st1 = string1.split()
data = array(st1)

# 1. 将字符串转换为整数标签
label_encoder = LabelEncoder()
integer_labels = label_encoder.fit_transform(data)

# 2. 对整数标签做独热编码
onehot_encoded = to_categorical(integer_labels)
print("独热编码结果:")
print(onehot_encoded)

# 反转编码:从独热编码还原回原始字符串
first_encoded_idx = argmax(onehot_encoded[0])
original_string = label_encoder.inverse_transform([first_encoded_idx])[0]
print("\n还原第一个编码的原始字符串:", original_string)

方法2:使用sklearn的OneHotEncoder

直接用sklearn.preprocessing.OneHotEncoder完成转换(注意需要先转整数标签,因为OneHotEncoder也不直接处理字符串):

from numpy import array
from sklearn.preprocessing import OneHotEncoder, LabelEncoder

# 原始数据处理
string1 = "Interstitial markings are diffusely prominent throughout both lungs. Heart size is normal. Pulmonary XXXX normal."
st1 = string1.split()
data = array(st1).reshape(-1, 1)  # OneHotEncoder要求输入为二维数组

# 1. 字符串转整数标签
label_encoder = LabelEncoder()
integer_labels = label_encoder.fit_transform(data.flatten()).reshape(-1, 1)

# 2. 生成独热编码
onehot_encoder = OneHotEncoder(sparse_output=False)
onehot_encoded = onehot_encoder.fit_transform(integer_labels)
print("独热编码结果:")
print(onehot_encoded)

# 反转编码
first_encoded_idx = onehot_encoded[0].argmax()
original_string = label_encoder.inverse_transform([first_encoded_idx])[0]
print("\n还原第一个编码的原始字符串:", original_string)
注意事项
  • 重复出现的字符串(比如示例中的normal.)会被映射为同一个整数标签,对应的独热编码也完全相同。
  • 独热编码的维度等于数据中唯一字符串的数量,每个位置对应一个唯一的字符串。

内容的提问来源于stack exchange,提问作者user19686684

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 09:05:19