使用TensorFlow Probability实现多项式线性回归结果异常求助
解决TensorFlow Probability中sts.Sum+LinearRegression多项式回归拟合偏差问题
核心问题分析
你的问题核心原因是特征未标准化:x²特征的数值范围(比如x=4时x²=16)远大于x⁰(恒为1)和x¹(0-4),优化过程中模型会优先调整大特征的权重,同时通过拉低x⁰的权重(等效偏置项)强行拟合数据,最终导致整体拟合偏差。另外仅5个样本的小数据集也会加剧优化的不稳定性。
解决方案步骤
1. 标准化特征
对所有输入特征(x⁰、x¹、x²)做标准化处理,让每个特征均值为0、方差为1,消除数值范围差异对优化的影响。
2. 调整优化参数
- 降低Adam优化器的学习率(小数据集建议用0.001,避免默认0.01导致的震荡)
- 增加训练步数,让变分推断充分收敛
- 保留x⁰特征的标准化合理性(手动避免其标准差为0的情况)
完整可运行代码示例
import tensorflow as tf import tensorflow_probability as tfp import numpy as np tfd = tfp.distributions sts = tfp.sts # 假设x序列为[0,1,2,3,4],对应给定的y值 x = np.arange(5, dtype=np.float32) y = np.array([0.1, 0.9, 4.1, 8.9, 16.1], dtype=np.float32) # 构建多项式特征矩阵:x⁰, x¹, x² features = np.stack([np.ones_like(x), x, x**2], axis=1) # 标准化特征(关键步骤) mean = np.mean(features, axis=0) std = np.std(features, axis=0) # 避免x⁰特征标准差为0,手动设置为1 std[0] = 1.0 normalized_features = (features - mean) / std # 转换为TensorFlow张量 norm_features_tensor = tf.convert_to_tensor(normalized_features) y_tensor = tf.convert_to_tensor(y) # 构建要求的STS模型:sts.Sum包含LinearRegression组件 linear_reg = sts.LinearRegression( design_matrix=norm_features_tensor, name="polynomial_regression" ) model = sts.Sum([linear_reg], observed_time_series=y_tensor) # 配置变分推断与优化器 num_train_steps = 2000 optimizer = tf.optimizers.Adam(learning_rate=0.001) # 初始化变分后验分布 variational_dist = sts.build_factored_surrogate_posterior(model=model) # 训练逻辑 @tf.function(autograph=False, jit_compile=False) def train_step(): with tf.GradientTape() as tape: # 计算变分推断的损失 loss = -variational_dist.log_prob(model.prior.sample()) + \ model.log_prob(y_tensor, variational_dist.sample()) grads = tape.gradient(loss, variational_dist.trainable_variables) optimizer.apply_gradients(zip(grads, variational_dist.trainable_variables)) return loss # 执行训练 for step in range(num_train_steps): loss = train_step() if step % 200 == 0: print(f"Step {step}, Loss: {loss.numpy():.4f}") # 提取后验权重并转换回原始特征空间 posterior_mean = variational_dist.mean() norm_weights = posterior_mean["polynomial_regression/weights"] # 反标准化得到原始特征对应的权重 original_weights = norm_weights / std # 计算等效偏置项(由标准化偏移带来的常数项) bias = -tf.reduce_sum(norm_weights * mean / std) print(f"原始特征权重(x⁰, x¹, x²): {original_weights.numpy()}") print(f"等效偏置项: {bias.numpy():.4f}") # 验证拟合效果 x_test = np.linspace(0, 4, 5) test_features = np.stack([np.ones_like(x_test), x_test, x_test**2], axis=1) y_pred = test_features @ original_weights.numpy() + bias.numpy() print("\n拟合对比:") for xi, yi, yp in zip(x, y, y_pred): print(f"x={xi}, 真实y={yi}, 预测y={yp:.4f}")
效果说明
执行上述代码后,x⁰的权重会接近0,拟合曲线与真实数据的偏差会大幅降低。标准化让每个特征的梯度更新处于同一量级,解决了大特征主导优化的问题;调整后的优化参数也适配了小数据集的收敛需求。
内容的提问来源于stack exchange,提问作者SU3
相关产品推荐
相关产品推荐

