You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python三层PMC神经网络始终输出训练集均值问题排查求助

三层感知器输出恒为训练集均值的问题分析

问题描述

开发三层多层感知器(PMC)拟合函数时,遇到神经网络对所有输入均返回训练集均值的问题,已排查代码多日未找到原因,需分析故障点。

代码

import pandas as pd
import numpy as np

class pmc_3_layers:
    
    def __init__(self, n1, n2, n3):
        # 各层神经元数量
        self.n1 = n1
        self.n2 = n2
        self.n3 = n3
        
        # LeCun随机权重初始化
        self.w = []
        self.w.append(np.random.default_rng().uniform(-(2.4/n1), (2.4/n1),(n2, n1+1)))
        self.w.append(np.random.default_rng().uniform(-(2.4/n1), (2.4/n1),(n3, n2+1)))
    
    def forward(self, variables_updated):
        
        # Sigmoid激活函数
        gfunc = np.vectorize(lambda a : 1/(1+np.exp(-a)))
        
        # 第一层计算
        i1 = self.w[0]@variables_updated
        y1 = gfunc(i1)
        y1 = np.insert(y1, 0, -1, axis = 0)
        
        # 第二层计算
        i2 = self.w[1]@y1
        y2 = gfunc(i2)
        
        return i1, y1, i2, y2
        
    def backward(self, variable, classe, i1, y1, i2, y2):
        
        # Sigmoid导数
        glinhafunc = np.vectorize(lambda a : np.exp(-a)/((1+np.exp(-a))**2))
        
        # 第二层梯度计算
        glinha2 = glinhafunc(i2)
        grad2 = (classe - y2)*glinha2
        
        if y1.ndim <= 1:
            self.w[1] = self.w[1] + self.taxa_aprendizado*grad2@y1.reshape(1, -1)
        else:
            self.w[1] = self.w[1] + self.taxa_aprendizado*grad2@y1.T
        
        # 第一层梯度计算
        glinha1 = glinhafunc(i1)
        
        if glinha1.ndim<=1:
            grad1 = -glinha1.reshape(-1,1)@grad2*self.w[1][:, 1:]
        else:
            grad1 = -glinha1.T@grad2*self.w[1][:, 1:]
        
        if grad1.ndim<=1:
            self.w[0] = self.w[0] + self.taxa_aprendizado*grad1.reshape(-1,1)@variable.reshape(1, -1)
        else:
            self.w[0] = self.w[0] + self.taxa_aprendizado*grad1.T@variable.reshape(1, -1)
        
    
    def eqm(self):
        # 均方误差计算
        eqm = 0
            
        for i in range(len(self.variables_updated)):
            i1, y1, i2, y2 = self.forward(self.variables_updated[i])
            for j in range(self.n3):
                eqm = eqm - (((self.output[j] - y2[j])**2)/2)
            
        eqm = eqm/len(self.variables_updated)
        return eqm
        
    def train(self, variables, output, taxa_aprendizado, precision):
        
        self.output = output
        self.variables = variables
        self.variables_updated = np.insert(self.variables, 0, np.full((1, len(self.variables)), -1), axis = 1)
        self.taxa_aprendizado = taxa_aprendizado
        self.precision = precision
        
        self.epoch = 1
        
        print('初始权重矩阵:')
        print(self.w)
        print('\n')
        
        while True:
            print('\n')
            print('######################################')
            print('开始第%i轮训练' %(self.epoch))
            print('\n')
            
            previous_eqm = self.eqm()
            
            for i in range(len(self.variables_updated)):
                i1, y1, i2, y2 = self.forward(self.variables_updated[i])
                self.backward(self.variables_updated[i], output[i], i1, y1, i2, y2)
            
            current_eqm = self.eqm()
            self.current_eqm = current_eqm
            
            if (abs(current_eqm - previous_eqm) <= precision):
                print("训练结束,共执行%s轮" % (self.epoch))
                print('\n')
                print("最终均方误差:%f" % (self.current_eqm))
                print('\n')
                print('最终权重矩阵:')
                print(self.w)
                break
            
            self.epoch = self.epoch + 1

    def predict(self, variables):
        # 预测函数
        variables_updated = np.insert(variables, 0, np.full((1, len(variables)), -1), axis = 1)
        resultados = np.zeros((len(variables_updated), self.n3))
        
        for i in range(len(variables_updated)):
            i1, y1, i2, y2 = self.forward(variables_updated[i])
            for j in range(len(y2)):
                resultados[i][j] = y2[j]
        
        return resultados

核心问题及修复方案

1. 均方误差(EQM)计算完全错误

原代码用减法累加损失,导致EQM为负数,训练终止条件判断完全失效,模型无法正常收敛:

# 错误代码
eqm = eqm - (((self.output[j] - y2[j])**2)/2)

修复:改为加法累加平方差,同时修正样本索引错误:

# 正确代码
eqm = eqm + ((self.output[i][j] - y2[j])**2)
# 最后计算均值
eqm = eqm / (len(self.variables_updated) * self.n3)

2. 反向传播梯度推导错误

第一层梯度的符号、矩阵乘法逻辑全部错误,导致权重更新方向完全偏离,模型无法学习:

# 错误代码
grad1 = -glinha1.reshape(-1,1)@grad2*self.w[1][:, 1:]

修复:正确计算隐藏层误差项,基于上层梯度和权重传递:

# 正确代码
if glinha1.ndim <= 1:
    grad1 = (self.w[1][:, 1:].T @ grad2.reshape(-1,1)).flatten() * glinha1
else:
    grad1 = (self.w[1][:, 1:].T @ grad2) * glinha1.T

3. 权重初始化参数错误

第二层权重初始化误用输入层维度n1,LeCun初始化应基于当前层输入维度n2:

# 错误代码
self.w.append(np.random.default_rng().uniform(-(2.4/n1), (2.4/n1),(n3, n2+1)))

修复:

self.w.append(np.random.default_rng().uniform(-(2.4/n2), (2.4/n2),(n3, n2+1)))

4. 训练数据偏置项插入错误

原代码插入偏置项时维度错误,导致每个样本的偏置值异常:

# 错误代码
self.variables_updated = np.insert(self.variables, 0, np.full((1, len(self.variables)), -1), axis = 1)

修复:

self.variables_updated = np.insert(self.variables, 0, -1, axis=1)

同理,predict方法中的偏置插入也需做相同修正。

5. 激活函数冗余处理

无需用np.vectorize包装sigmoid及其导数,直接用numpy广播机制即可,提升效率:

# 替换原gfunc
def gfunc(a):
    return 1/(1+np.exp(-a))

# 替换原glinhafunc
def glinhafunc(a):
    sig = gfunc(a)
    return sig * (1 - sig)

总结

EQM计算错误和反向传播梯度错误是导致模型输出恒为训练集均值的核心原因:前者让训练提前终止或无法判断收敛,后者导致权重更新完全无效,模型只能输出训练集均值作为“最优”预测。

内容的提问来源于stack exchange,提问作者Lucas Alponti

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 11:35:23