TensorFlow中‘Cast string to float is not supported’报错排查求助
嘿,我一眼就揪出问题所在了——你把注意力全放在特征列上,但其实错误根源是你的标签列sp_call是字符串类型!
看错误栈里的关键线索:
UnimplementedError: Cast string to float is not supported
[[node linear/head/ToFloat (...) = CastDstT=DT_FLOAT, SrcT=DT_STRING,...]]
这里明明白白说的是labels(也就是你的sp_call列)是字符串类型,TensorFlow的LinearClassifier在计算损失时需要把标签转成数值型,但字符串转float是不支持的操作。
另外,你的代码里还有个小笔误:特征列列表里写的是im_red,但前面定义的数值列变量名是imred,这个也会导致找不到对应的特征列,得先修正这个细节。
具体解决步骤:
先确认标签列的类型和取值
运行下面的代码看看sp_call的真实情况:print(tree_data['sp_call'].dtype) print(tree_data['sp_call'].unique())你大概率会看到它是
object类型(字符串),比如类似['classA', 'classB']这样的类别取值。将标签转换为数值型
针对二分类任务,有两种简单的转换方式:- 方式一:用pandas的
map手动映射(适合你知道所有类别取值的情况)# 假设你的标签是类似'positive'/'negative'或者两个明确的字符串类别 labels = tree_data['sp_call'].map({'类别1': 0, '类别2': 1}) - 方式二:用sklearn的
LabelEncoder自动编码(适合类别较多或不确定的情况)from sklearn.preprocessing import LabelEncoder le = LabelEncoder() labels = le.fit_transform(tree_data['sp_call'])
转换后,标签会变成0和1的整数类型,TensorFlow就能正常处理了。
- 方式一:用pandas的
修正特征列的笔误
把feat_cols = [b3_sum,im3b3_s,im_red,sp1,sp2]改成feat_cols = [b3_sum,im3b3_s,imred,sp1,sp2],确保变量名和前面定义的一致。
修改后的完整代码示例:
import tensorflow as tf import pandas as pd from sklearn.model_selection import train_test_split from sklearn.preprocessing import LabelEncoder tree_data_file = r'\\David\f\first_test_feature_cols_v2.csv' tree_data = pd.read_csv(tree_data_file) ### Create feature columns for continuous data b3_sum = tf.feature_column.numeric_column('b3_sum') im3b3_s = tf.feature_column.numeric_column('im3b3_s') imred = tf.feature_column.numeric_column('imred') # Create feature columns for categorical data sp1 = tf.feature_column.categorical_column_with_hash_bucket('sp1',hash_bucket_size=10) sp2 = tf.feature_column.categorical_column_with_hash_bucket('sp2',hash_bucket_size=10) # 修正笔误:im_red -> imred feat_cols = [b3_sum,im3b3_s,imred,sp1,sp2] #TRAIN TEST SPLIT x_data = tree_data.drop('sp_call', axis=1) # 处理标签列:将字符串转成数值 le = LabelEncoder() labels = le.fit_transform(tree_data['sp_call']) X_train, X_test, Y_train, Y_test = train_test_split(x_data, labels, test_size = 0.3) input_func = tf.estimator.inputs.pandas_input_fn(x = X_train, y=Y_train, batch_size=10, num_epochs=1000, shuffle=True) model = tf.estimator.LinearClassifier(feature_columns=feat_cols, n_classes=2) model.train(input_fn=input_func, steps=1000)
这样应该就能解决你的问题了,下次遇到类似错误,先看错误栈里明确指出的是labels还是features,定位起来会更快~
内容的提问来源于stack exchange,提问作者D_C

