You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas concat合并CSV后列转为object,致Keras训练报错求助

解决合并CSV后Keras训练报错:ValueError: Failed to convert a NumPy array to a Tensor (Unsupported object type int)

问题场景

将拆分至10个独立.csv文件的数据集分别导入后,使用pd.concat合并为单个DataFrame,所有列的数据类型被转换为object。调用Keras的model.fit训练模型时触发报错:

ValueError: Failed to convert a NumPy array to a Tensor (Unsupported object type int)

仅导入单个.csv文件、不执行合并操作时,代码可正常运行。相关代码如下:

df1 = pd.read_csv("CSE-CIC-IDS2018/02-14-2018.csv", low_memory=False, nrows=10000)
df2 = pd.read_csv("CSE-CIC-IDS2018/02-15-2018.csv", low_memory=False, nrows=10000)
df3 = pd.read_csv("CSE-CIC-IDS2018/02-16-2018.csv", low_memory=False, nrows=10000)

df = pd.concat([df1,df2,df3])
df.drop(columns=['Timestamp','Flow ID', 'Src IP', 'Src Port', 'Dst IP'], inplace=True)

x = df.drop(["Label"], axis=1)  # 修正原代码语法错误:缺少闭合括号
y = df["Label"].apply(lambda x:0 if x=="Benign" else 1)

x_train, x_remaining, y_train, y_remaining = train_test_split(x,y, train_size=0.60, random_state=4)
x_val, x_test, y_val, y_test = train_test_split(x_remaining, y_remaining, test_size=0.5, random_state=4)

# Model
model = Sequential()
model.add(Dense(16, activation="relu"))#, input_dim=len(x_train.columns)))
model.add(Dense(32, activation="relu"))
model.add(Dense(units=1, activation="sigmoid"))

# Compiler
model.compile(loss="binary_crossentropy", optimizer="Adam", metrics="accuracy")

model.fit(x_train, y_train, epochs=10, batch_size=512)

问题根源

合并后列类型变为object,是因为不同CSV文件中同一列可能存在混合数据类型(如部分行是int、部分是str),pd.concat会将这类列统一转为object类型,而TensorFlow无法直接处理object类型的输入数组。

解决方案

1. 合并后强制转换数值列类型

显式将所有数值列转换为浮点型(避免整数与浮点混合的问题),并处理转换产生的空值:

# 获取除Label外的所有列作为数值列
numeric_cols = df.columns.drop('Label')
# 转换为float64,无法转换的设为NaN
df[numeric_cols] = df[numeric_cols].apply(pd.to_numeric, errors='coerce')
# 用0填充NaN(可根据业务调整填充策略,比如均值、中位数)
df.fillna(0, inplace=True)

2. 清理列中的非数值异常值

部分CSV可能包含字符串形式的异常值(如"?"、"NA"),需先清理再转换:

for col in numeric_cols:
    # 替换所有非数字、小数点、负号的字符为空
    df[col] = df[col].replace(r'[^0-9.-]', '', regex=True)
    # 转换为数值类型
    df[col] = pd.to_numeric(df[col], errors='coerce')
# 填充空值
df.fillna(0, inplace=True)

3. 确保输入为纯数值数组

在传入模型前,将训练数据转换为指定类型的NumPy数组:

x_train = x_train.astype('float32').values
y_train = y_train.astype('int32').values

4. 可选:导入时指定列类型

如果已知各列的正确数据类型,可在pd.read_csv时直接指定,从源头避免类型混乱:

# 构建数值列的类型字典,假设所有数值列应为float64
dtype_dict = {col: 'float64' for col in ['Dst Port', 'Protocol', 'Flow Duration', ...]}  # 替换为实际列名
df1 = pd.read_csv("CSE-CIC-IDS2018/02-14-2018.csv", low_memory=False, nrows=10000, dtype=dtype_dict)
df2 = pd.read_csv("CSE-CIC-IDS2018/02-15-2018.csv", low_memory=False, nrows=10000, dtype=dtype_dict)
df3 = pd.read_csv("CSE-CIC-IDS2018/02-16-2018.csv", low_memory=False, nrows=10000, dtype=dtype_dict)

内容的提问来源于stack exchange,提问作者deadpixels

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 18:15:27