训练含文本向量化的自编码器后预测报错:输入形状不兼容
构建了一个用于日志异常检测的自编码器神经网络,处理包含文本和数值数据的CSV日志,训练完成并保存模型后,调用predict时出现输入形状不兼容错误,错误信息如下:
WARNING:tensorflow:Model was constructed with shape (None, 972) for input KerasTensor(type_spec=TensorSpec(shape=(None, 972), dtype=tf.float32, name='input_1'), name='input_1', description="created by layer 'input_1'"), but it was called on an input with incompatible shape (None, 1).
ValueError: Exception encountered when calling layer "model_1" (type Functional).
Input 0 of layer "dense" is incompatible with the layer: expected axis -1 of input shape to have value 972, but received input with shape (None, 1) Call arguments received: • inputs=tf.Tensor(shape=(None, 1), dtype=string) • training=False • mask=None
训练代码:
import numpy as np import pandas as pd import tensorflow as tf from tensorflow.keras.models import Sequential from tensorflow.keras.layers import Dense, Input import matplotlib.pyplot as plt columnsToUse = [ "_source.log.OriginalTimestamp", "_source.log.SenderID", "_source.host.ip", "_source.log.Severity", "_source.log.message_text", "_source.host.mac", "_source.log.offset", ] df = pd.read_csv("log.csv", usecols=columnsToUse) df = df.replace(r'^\s*$', np.nan, regex=True).dropna() log_features = df.copy() inputs = {} for name, column in log_features.items(): dtype = column.dtype if dtype == object: dtype = tf.string else: dtype = tf.float32 inputs[name] = tf.keras.Input(shape=(1,), name=name, dtype=dtype) numeric_inputs = { name: input for name, input in inputs.items() if input.dtype == tf.float32 } x = tf.keras.layers.Concatenate()(list(numeric_inputs.values())) norm = tf.keras.layers.Normalization() norm.adapt(np.array(df[numeric_inputs.keys()])) all_numeric_inputs = norm(x) preprocessed_inputs = [all_numeric_inputs] for name, input in inputs.items(): if input.dtype == tf.float32: continue lookup = tf.keras.layers.StringLookup(vocabulary=np.unique(log_features[name])) one_hot = tf.keras.layers.CategoryEncoding(num_tokens=lookup.vocabulary_size()) x = lookup(input) x = one_hot(x) preprocessed_inputs.append(x) preprocessed_inputs_cat = tf.keras.layers.Concatenate()(preprocessed_inputs) log_preprocessing = tf.keras.Model(inputs, preprocessed_inputs_cat) tf.keras.utils.plot_model( model=log_preprocessing, to_file="preprocessing.png", rankdir="LR", dpi=72, show_shapes=True, ) log_features_dict = {name: np.array(value) for name, value in log_features.items()} features_dict = {name: values[:1] for name, values in log_features_dict.items()} print(log_preprocessing(features_dict)) processed_features = log_preprocessing(log_features_dict) features_len = log_preprocessing.output_shape[-1] def log_model(): input_layer = Input((features_len)) layer = Dense(7)(input_layer) layer = Dense(5)(layer) layer = Dense(5)(layer) layer = Dense(7)(layer) output_layer = Dense(features_len)(layer) model = tf.keras.Model(input_layer, output_layer) model.compile( loss=tf.keras.losses.MSE, optimizer=tf.keras.optimizers.Adam(), metrics=tf.keras.metrics.MSE ) tf.keras.utils.plot_model( model=model, to_file="model.png", rankdir="LR", dpi=72, show_shapes=True, ) return model final_model = log_model() final_model.fit(x=processed_features, y=processed_features, epochs=30) final_model.save('autoencoder')
测试代码:
import numpy as np import pandas as pd import tensorflow as tf model = tf.keras.models.load_model('autoencoder7557') df = pd.read_csv("testlog.csv") predictions = model.predict(df)
使用训练数据集测试时也出现相同错误。
问题核心是你只保存了自编码器模型final_model,但没有把预处理层log_preprocessing整合到最终模型中,导致加载后的模型只能接收预处理后的972维特征,而你传入的是原始CSV数据,形状不匹配。
修复步骤
训练时整合预处理层与自编码器
将预处理模型的输出直接作为自编码器的输入,构建端到端的完整模型,这样保存后就能直接接收原始数据。修改训练代码末尾部分:# 替换原有的final_model定义与训练逻辑 def build_full_model(preprocessing_model, encoder_model): inputs = preprocessing_model.input preprocessed = preprocessing_model(inputs) outputs = encoder_model(preprocessed) return tf.keras.Model(inputs, outputs) # 定义自编码器结构 encoder = log_model() # 构建完整的端到端模型 full_model = build_full_model(log_preprocessing, encoder) full_model.compile( loss=tf.keras.losses.MSE, optimizer=tf.keras.optimizers.Adam(), metrics=tf.keras.metrics.MSE ) # 用原始数据字典训练(而非预处理后的特征) full_model.fit(x=log_features_dict, y=processed_features, epochs=30) # 保存完整模型 full_model.save('full_autoencoder')预测时使用完整模型
加载整合后的模型,传入原始DataFrame转换的字典格式数据:import numpy as np import pandas as pd import tensorflow as tf model = tf.keras.models.load_model('full_autoencoder') df = pd.read_csv("testlog.csv") # 转换为模型要求的字典格式 test_data_dict = {name: np.array(value) for name, value in df.items()} # 确保列名与训练时一致,缺失列补默认值,多余列删除 required_cols = ["_source.log.OriginalTimestamp", "_source.log.SenderID", "_source.host.ip", "_source.log.Severity", "_source.log.message_text", "_source.host.mac", "_source.log.offset"] test_data_dict = {col: test_data_dict[col] for col in required_cols if col in test_data_dict} predictions = model.predict(test_data_dict)额外注意事项
- 训练时
StringLookup层用了np.unique(log_features[name]),若测试数据包含训练时未出现的类别,会触发错误。建议添加oov_token参数处理未知类别:lookup = tf.keras.layers.StringLookup(vocabulary=np.unique(log_features[name]), oov_token="[UNK]") - 测试数据需和训练数据做相同的缺失值处理:替换空值为NaN并删除,避免预处理出错。
- 训练时
内容的提问来源于stack exchange,提问作者Awetgomk

