You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow回归模型训练优测试差,归一化全数据集遇属性错误求助

TensorFlow回归模型测试表现差及归一化相关属性错误排查

问题背景

  • 调优后的TensorFlow回归模型训练阶段表现良好,但测试阶段结果极差,怀疑是测试集未做归一化导致。
  • 尝试先对全数据集归一化再划分训练/测试集,触发属性错误。

原代码示例

数据处理代码

#concatenate the surface data and single_downhole_col into a single dataframe
training_Data =[]
training_Data = pd.concat([surface_Data, single_downhole_col], axis=1)
#print('training data shape:',training_Data.shape)
#print(training_Data.head())

#normalize the data using keras
model_normalizer_layer = tf.keras.layers.Normalization(axis=-1)
model_normalizer_layer.adapt(training_Data)
normalized_training_Data = model_normalizer_layer(training_Data)

#convert the data frame to array
dataset = normalized_training_Data.copy()
dataset.tail()

#create a training and test set
train_dataset = dataset.sample(frac=0.8, random_state=0)
test_dataset = dataset.drop(train_dataset.index)

#check the data
train_dataset.describe().transpose()

#split features from labels
train_features = train_dataset.copy()
test_features = test_dataset.copy()

模型构建代码

def build_and_compile_model(data):
    model = keras.Sequential([
        model_normalizer_layer,
        layers.Dense(260, input_dim=401,activation='relu'),
        layers.Dense(80, activation='relu'),
        #layers.Dense(40, activation='relu'),
        layers.Dense(1)
    ])

问题根源分析

  1. 数据类型不匹配:Normalization层处理后返回的是Tensor对象,而非Pandas DataFrame,因此调用sample()、drop()、describe()等DataFrame专属方法会触发属性错误。
  2. 数据泄露风险:先归一化全量数据再划分训练/测试集,会让测试集的统计信息混入训练过程,直接导致模型泛化能力下降,这也是测试表现差的核心原因之一。

两种正确解决方案

方案1:将归一化层嵌入模型(推荐)

这种方式下,归一化层会自动学习训练集的均值和方差,推理时用相同统计量处理测试数据,彻底避免数据泄露:

# 1. 先拼接数据,再划分训练/测试集(关键步骤)
training_Data = pd.concat([surface_Data, single_downhole_col], axis=1)
train_dataset = training_Data.sample(frac=0.8, random_state=0)
test_dataset = training_Data.drop(train_dataset.index)

# 2. 分离特征与标签(假设最后一列为标签)
train_features = train_dataset.iloc[:, :-1]
train_labels = train_dataset.iloc[:, -1]
test_features = test_dataset.iloc[:, :-1]
test_labels = test_dataset.iloc[:, -1]

# 3. 用训练特征适配归一化层
model_normalizer_layer = tf.keras.layers.Normalization(axis=-1)
model_normalizer_layer.adapt(train_features)

# 4. 构建并编译模型
def build_and_compile_model():
    model = tf.keras.Sequential([
        model_normalizer_layer,
        tf.keras.layers.Dense(260, activation='relu'),  # 无需指定input_dim,归一化层会自动推断输入形状
        tf.keras.layers.Dense(80, activation='relu'),
        tf.keras.layers.Dense(1)
    ])
    model.compile(optimizer='adam', loss='mean_squared_error')
    return model

# 训练模型
model = build_and_compile_model()
model.fit(train_features, train_labels, epochs=15, validation_split=0.1)

# 测试时直接传入原始测试特征,归一化层自动处理
test_loss = model.evaluate(test_features, test_labels)

方案2:手动用训练集统计量归一化

若不需要将归一化层嵌入模型,可手动计算训练集的均值和方差,用其归一化训练/测试集:

training_Data = pd.concat([surface_Data, single_downhole_col], axis=1)

# 先划分训练/测试集
train_dataset = training_Data.sample(frac=0.8, random_state=0)
test_dataset = training_Data.drop(train_dataset.index)

# 分离特征与标签
train_features = train_dataset.iloc[:, :-1]
train_labels = train_dataset.iloc[:, -1]
test_features = test_dataset.iloc[:, :-1]
test_labels = test_dataset.iloc[:, -1]

# 计算训练特征的均值和方差
mean = train_features.mean(axis=0)
std = train_features.std(axis=0)

# 用训练集统计量归一化(绝对不能用测试集的均值方差)
train_features_normalized = (train_features - mean) / std
test_features_normalized = (test_features - mean) / std

# 后续用归一化后的特征训练、测试模型

核心注意事项

  • 禁止先归一化全量数据再划分数据集:这会严重破坏模型的泛化能力,是测试表现差的关键诱因。
  • Tensor对象不支持DataFrame方法,若需转换可使用pd.DataFrame(normalized_tensor.numpy()),但不推荐这种方式。

内容的提问来源于stack exchange,提问作者TheNewGuy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 18:18:30