You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

手写Deep Neural Network训练准确率停滞问题求助

问题描述

从零实现深度神经网络时遇到训练困境:训练过程中准确率始终固定,比如一直维持在50.00835414342936%,无论迭代多少epoch都没有变化,而用TensorFlow实现的相同模型能达到70%-79%的准确率。以下是我的实现代码(本科毕业论文相关工作),恳请帮忙修复:

y_temp= []
y_temp2 = []

for i in range(len(y_train)):
    temp = ([y_train.iloc[i]])
    y_temp.append(temp)
y = np.array((y_temp), dtype=np.float128)

for i in range(len(y_test)):
    temp2= ([y_test.iloc[i]])
    y_temp2.append(temp2)
y2 = np.array((y_temp2), dtype=np.float128)


def sigmoid(self, x):
    t = t.astype(np.float128)
    return 1/(1+np.exp(-t))

def sigmoid_derivative(self, x):
    t = t.astype(np.float128)
    return sigmoid(t) * (1-sigmoid(t))

def relu(t):
    t = t.astype(np.float128)
    return np.maximum(0, t)

def relu_derivative(t):
    t = t.astype(np.float128)
    t[t <= 0] = 0
    t[t > 0] = 1
    return t


class DeepNeuralNetwork:
    def __init__(self, perceptron, x, y, hidden_layer, lr, bias):
        np.random.seed(42)
        self.hl = hidden_layer
        self.weights = {}
        self.input = x
        self.perceptron_input = self.input.shape[1]
        self.perceptron_output = 1
        self.lr = lr
        self.y = y

        # inisialisasi weight for each neuron 
        for i in range(hidden_layer): 
            self.weights[i] = np.random.rand(self.perceptron_input,perceptron) * sqrt(2.0 / self.perceptron_input)  
        self.weights[hidden_layer] = np.random.randn(perceptron, self.perceptron_output) * np.sqrt(1. / perceptron) 

        self.y = y
        self.output = np.zeros(y.shape)



    #forward propagation
    def forward_prop(self, hidden_layer):
        self.layer = {}


        self.layer[0] = relu(np.dot(self.input, self.weights[0])) 

        for i in range(hidden_layer+1): 
            if i != 0: 
                #perhitungan activation function relu
                self.layer[i] = relu(np.dot(self.layer[i-1], self.weights[i]))
            if i == hidden_layer:
                self.layer[i] = sigmoid(np.dot(self.layer[i-1], self.weights[i]))

        return self.layer[i]



    def backward_prop(self, hidden_layer):
        dW = {}
        E = {}

        #backpropagation output layer
        E[hidden_layer] = (self.y - self.output) * sigmoid_derivative(self.output)
        self.layer[hidden_layer-1] = np.multiply(E[hidden_layer], np.int64(self.layer[hidden_layer]>0))
        dW[hidden_layer] = np.dot(self.layer[hidden_layer-1].T, E[hidden_layer]) #turunan dari weight

        for i in reversed(range(hidden_layer)): 
            if i != 0:
                E[i] = np.dot(E[i+1], self.weights[i+1].T) * relu_derivative(self.layer[i])
                self.layer[i] = np.multiply(E[i], np.int64(self.layer[i]>0))
                dW[i] = np.dot(self.layer[i-1].T,E[i])
            if i == 0:
                E[0]= np.dot(E[1], self.weights[1].T) * relu_derivative(self.layer[0])
                self.layer[i] = np.multiply(E[i], np.int64(self.layer[i]>0))
                dW[0] = np.dot(self.input.T, E[0])
        for i in range(hidden_layer+1):
            self.weights[i] += self.lr * dW[i]  #lr


    def train(self, epoch):
        for i in range(epoch):
            self.output = self.forward_prop(self.hl) 
            self.backward_prop(self.hl)
            self.hitung_akurasi(self.input, self.y)


    def hitung_akurasi(self, X, y):
        predictions = []
        counter = 0

        for i in range(len(y)): #loop
            if self.output[i] >= 0.5: #if output greaterthan 0.5
                prediksi = 1 
            else:
                prediksi = 0
            predictions.append(prediksi)


            if predictions[i] == y[i]:
                counter = counter+1

        akurasi = (counter/len(y)) * 100    

        cm = confusion_matrix(predictions, y)
        true_positive = cm[1,1]
        true_negative = cm[0,0]
        false_positive = cm[0,1]
        false_negative = cm[1,0]

        print('Accuracy: '+ str(akurasi))



hl = 4
epoch = 50
lr = 0.01
nn = DeepNeuralNetwork(6, x_train_scaled, y, hl, lr, 0)
nn.train(epoch)

输出示例:

Output: Accuracy: 50.00835414342936 Accuracy: 50.00835414342936 Accuracy: 50.00835414342936 ........................... 

代码错误修复与解释

1. 激活函数定义错误

sigmoid和sigmoid_derivative被定义成了类方法格式(带self参数),但实际是类外普通函数;且函数内部错误使用未定义的t变量,应该替换为传入的参数:

def sigmoid(t):
    t = t.astype(np.float128)
    return 1/(1+np.exp(-t))

def sigmoid_derivative(t):
    t = t.astype(np.float128)
    return sigmoid(t) * (1-sigmoid(t))

2. 前向传播逻辑混乱

原循环中同时处理隐藏层和输出层,会导致最后一层的激活函数被重复覆盖。正确逻辑是:前hidden_layer层用ReLU,最后一层(输出层)用Sigmoid,修改后的forward_prop:

def forward_prop(self):
    self.layer = {}
    # 第一层隐藏层
    self.layer[0] = relu(np.dot(self.input, self.weights[0])) 
    # 中间隐藏层
    for i in range(1, self.hl): 
        self.layer[i] = relu(np.dot(self.layer[i-1], self.weights[i]))
    # 输出层
    self.layer[self.hl] = sigmoid(np.dot(self.layer[self.hl-1], self.weights[self.hl]))
    return self.layer[self.hl]

调用时直接用self.forward_prop(),无需传入hidden_layer参数。

3. 反向传播破坏前向传播结果

反向传播中错误修改了self.layer的存储值,这会破坏前向传播保存的中间特征,导致梯度计算完全错误。删除所有修改self.layer的语句,仅保留误差和梯度计算:

def backward_prop(self):
    dW = {}
    E = {}

    # 输出层反向传播
    E[self.hl] = (self.y - self.output) * sigmoid_derivative(self.output)
    dW[self.hl] = np.dot(self.layer[self.hl-1].T, E[self.hl])

    # 隐藏层反向传播
    for i in reversed(range(self.hl)): 
        if i != 0:
            E[i] = np.dot(E[i+1], self.weights[i+1].T) * relu_derivative(self.layer[i])
            dW[i] = np.dot(self.layer[i-1].T, E[i])
        else:
            E[i] = np.dot(E[i+1], self.weights[i+1].T) * relu_derivative(self.layer[i])
            dW[i] = np.dot(self.input.T, E[i])
    
    # 更新权重
    for i in range(self.hl+1):
        self.weights[i] += self.lr * dW[i]

4. 缺失必要导入

代码中使用了sqrt但未导入,需在代码开头添加:

from math import sqrt

或替换为np.sqrt以使用numpy的平方根函数。

5. 训练循环与准确率计算优化

原hitung_akurasi用循环逐个判断效率低,改用向量运算简化;同时train方法中调用forward_prop和backward_prop时无需传参:

def train(self, epoch):
    for i in range(epoch):
        self.output = self.forward_prop() 
        self.backward_prop()
        self.hitung_akurasi(self.input, self.y)

def hitung_akurasi(self, X, y):
    predictions = (self.output >= 0.5).astype(int)
    counter = np.sum(predictions == y)
    akurasi = (counter / len(y)) * 100    
    print(f'Accuracy: {akurasi:.2f}')

6. 初始化冗余代码清理

类初始化中重复赋值self.y = y,删除其中一行即可。


内容的提问来源于stack exchange,提问作者Coco Along

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 00:32:09