You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow嵌入训练报错:NumPy数组转Tensor失败

问题排查与解决:TensorFlow无法转换NumPy数组为Tensor

问题描述

训练TensorFlow模型时触发报错:

ValueError: Failed to convert a NumPy array to a Tensor (Unsupported object type float).

样本数据

fld0,fld1,fld2,fld3,fld4,fld5,fld6,fld7,fld8,fld9,fld10,fld11,fld12,fld13,fld14,fld15,fld16,fld17,fld18,fld19,fld20,fld21,fld22,fld23,fld24,fld25,fld26,fld27,fld28,fld29,fld30,fld31,fld32,fld33,fld34,fld35,fld36,fld37,fld38,fld39,fld40,fld41,fld42,fld43,fld44,fld45,fld46,fld47,fld48,fld49,fld50,fld51,fld52,fld53,fld54,fld55,fld56,fld57,fld58,fld59,fld60,fld61,fld62,fld63,fld64,fld65,fld66,fld67,fld68,fld69,fld70,fld71,fld72,fld73,fld74,fld75,fld76,fld77,fld78,fld79,fld80,fld81,fld82,fld83,fld84,fld85,fld86,fld87,fld88,fld89,fld90,fld91,fld92,fld93,fld94,fld95,fld96,fld97,fld98,fld99,fld100,fld101,fld102,fld103,fld104,fld105,fld106,fld107,fld108,fld109,fld110,fld111,fld112,fld113,fld114,fld115,fld116,fld117,fld118,fld119,fld120,fld121,fld122,fld123,fld124,fld125,fld126,fld127,fld128,fld129,fld130,fld131,fld132,fld133
0.5713509314139188,1,1,0,0,1,0.49030538979462923,[ 0.0756 0.0756 0.1176 0.0672 0.0588 0.0756 0.0672 0.0504 0.0336 0.1008 0.0252 0.0252 0.0252 0.0672 0.0252 0.0252 0.0168 0.0084 0.0000 0.0000 0.0000 0.0084 0.0084 0.0000 0.0000 0.0000 0.0000 0.0000 0.0000 0.0000 0.0000 0.0000 0.0000 0.0000 0.0000 0.0000 0.0000 0.0000 0.0000 0.0420 ],0,0,0,0,0,0,1,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,1,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,1,0,0,0,0,0,0,0,0,1,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0

运行环境

  • pandas 2.2.3
  • Python 3.10.13
  • TensorFlow(版本见代码输出)

代码实现

import tensorflow as tf
import pandas as pd
import numpy as np


def convert(item):
    item = item[1:-1]    # remove `[ ]`
    item = item.strip()  # remove spaces at the end
    item = np.fromstring(item, sep=' ')  # convert string to `numpy.array`
    return item

print("TensorFlow Version: "+tf.__version__)

X_train = pd.read_csv('training.csv',converters={'fld7':convert})
y_train = pd.read_csv('training_labels.csv')
print(f"{X_train.shape=}")
print(f"{y_train.shape=}")

x_row=X_train.iloc[0]
x_val=x_row['fld7']
print(f"{type(x_row)=} {x_row=}")
print(f"{type(x_val)=} {x_val=}")

model = tf.keras.Sequential([
    tf.keras.layers.Dense(1024, activation='relu'),
    tf.keras.layers.Dense(2048, activation='relu'),
    tf.keras.layers.Dense(2048, activation='relu'),
    tf.keras.layers.Dense(1, activation='sigmoid')
])

model.compile(
    loss=tf.keras.losses.binary_crossentropy,
    optimizer=tf.keras.optimizers.Adam(learning_rate=0.003),
    metrics=[
        tf.keras.metrics.BinaryAccuracy(name='accuracy'),
        tf.keras.metrics.Precision(name='precision'),
        tf.keras.metrics.Recall(name='recall')
    ]
    )

history = model.fit(X_train, y_train, epochs=25)

程序运行输出

X_train.shape=(10, 134)
y_train.shape=(10, 1)
type(x_row)=<class 'pandas.core.series.Series'> x_row=fld0      0.571351
fld1             1
fld2             1
fld3             0
fld4             0
            ...   
fld129           0
fld130           0
fld131           0
fld132           0
fld133           0
Name: 0, Length: 134, dtype: object
type(x_val)=<class 'numpy.ndarray'> x_val=array([0.0756, 0.0756, 0.1176, 0.0672, 0.0588, 0.0756, 0.0672, 0.0504,
       0.0336, 0.1008, 0.0252, 0.0252, 0.0252, 0.0672, 0.0252, 0.0252,
       0.0168, 0.0084, 0.    , 0.    , 0.    , 0.0084, 0.0084, 0.    ,
       0.    , 0.    , 0.    , 0.    , 0.    , 0.    , 0.    , 0.    ,
       0.    , 0.    , 0.    , 0.    , 0.    , 0.    , 0.    , 0.042 ])

2024-10-08 20:14:52.337621: I external/local_xla/xla/stream_executor/cuda/cuda_executor.cc:901] successful NUMA node read from SysFS had negative value (-1), but there must be at least one NUMA node, so returning NUMA node zero. See more at https://github.com/torvalds/linux/blob/v6.0/Documentation/ABI/testing/sysfs-bus-pci#L344-L355
2024-10-08 20:14:52.339066: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1929] Created device /job:localhost/replica:0/task:0/device:GPU:0 with 20723 MB memory:  -> device: 0, name: NVIDIA A10G, pci bus id: 0000:00:1e.0, compute capability: 8.6
2024-10-08 20:14:52.339066: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1929] Created device /job:localhost/replica:0/task:0/device:GPU:0 with 20723 MB memory:  -> device: 0, name: NVIDIA A10G, pci bus id: 0000:00:1e.0, compute capability: 8.6
Traceback (most recent call last):
  File "/home/ubuntu/new_model/testcase.py", line 44, in <module>
    history = model.fit(X_train, y_train, epochs=25)
  File "/opt/tensorflow/lib/python3.10/site-packages/keras/src/utils/traceback_utils.py", line 70, in error_handler
    raise e.with_traceback(filtered_tb) from None
  File "/opt/tensorflow/lib/python3.10/site-packages/tensorflow/python/framework/constant_op.py", line 103, in convert_to_eager_tensor
    return ops.EagerTensor(value, ctx.device_name, dtype)
ValueError: Failed to convert a NumPy array to a Tensor (Unsupported object type float).

报错原因分析

核心问题在于:虽然fld7列被转换为numpy.ndarray,但整个X_train的DataFrame中,该列的每个元素是一个数组对象,导致DataFrame的整体数据类型为object(从输出的dtype: object可以确认)。TensorFlow无法直接处理这种混合类型的输入——因为每一行同时包含普通数值和数组,无法转换为形状统一的张量。

解决方法

方法1:将fld7列展开为独立特征列

把fld7中的数组拆分成多个单独的数值列,让整个DataFrame成为纯数值类型,TensorFlow可以直接处理:

# 提取fld7列,将每个数组转为一行,生成新的DataFrame
fld7_expanded = pd.DataFrame(X_train['fld7'].tolist(), index=X_train.index)
# 为展开的列命名,避免重复
fld7_expanded.columns = [f'fld7_{i}' for i in range(fld7_expanded.shape[1])]
# 删除原有的fld7列,合并展开后的列
X_train_processed = pd.concat([X_train.drop('fld7', axis=1), fld7_expanded], axis=1)
# 确保所有列都是float类型
X_train_processed = X_train_processed.astype(float)

# 用处理后的数据集训练模型
history = model.fit(X_train_processed, y_train, epochs=25)

方法2:将整个数据集转换为统一的NumPy数组

把每一行的普通数值和fld7数组拼接成一个一维数组,最终生成二维NumPy数组作为模型输入:

# 定义函数处理单行数据:拼接普通特征和fld7数组
def process_single_row(row):
    # 提取除fld7外的所有特征,转为float数组
    other_features = row.drop('fld7').values.astype(float)
    # 拼接其他特征和fld7数组
    return np.concatenate([other_features, row['fld7']])

# 遍历所有行,生成二维数组
X_train_np = np.array([process_single_row(row) for _, row in X_train.iterrows()])

# 用NumPy数组训练模型
history = model.fit(X_train_np, y_train, epochs=25)

验证修改效果

两种方法都能将输入数据转换为TensorFlow可接受的统一数值类型张量,调用model.fit时不会再触发类型转换错误。

内容的提问来源于stack exchange,提问作者George Sexton

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 11:27:02