You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从零搭建的神经网络损失曲线异常且出现NaN,求技术排查

问题背景

从零搭建了一个神经网络,结构为:Input > Layer1(sigmoid激活)> Layer2 > Output(softmax激活)。基础代码开发完成后,运行得到异常损失曲线,迭代次数较多时输出还会出现NaN值。已自行检查反向传播推导过程与代码实现,但未定位到问题所在。

完整代码

import numpy as np
import pandas as pd
import tensorflow as tf
import sklearn as sk
from sklearn import preprocessing
from tensorflow.keras.datasets import mnist
import matplotlib.pyplot as plt

(x_train,y_train),(x_test,y_test)=mnist.load_data()

n_train = 100
x_train = x_train[0:n_train].reshape(-1,784)/255
y_train0 = y_train[0:n_train].reshape(-1,)
y_train1 = pd.get_dummies(y_train0)
y_train = np.array( y_train1.astype(int) )


nx = 784 #input size (nx,1)
n1 = 20   #neurons in 1st layer
n_class = 10  #neurons in 2nd layer


lambda0 = 0.00  #center loss parameter
alpha =  0.0001  #gradient decent parameter

#layer weights
#The network is simple 
#Input, n1 neurons with sigmoid activation, 10 neurons and softmax output
W1 = 0.001*np.random.rand( n1, nx )
W2 = 0.001*np.random.rand( n_class, n1 )
b2 = 0.001*np.random.rand( n_class,1  )
b1 = 0.001*np.random.rand( n1,1  )
centers = np.random.rand(n_class,1)


def softmax(x):
 
    exp_sum = np.sum( np.exp(x) )
    
    return np.exp(x)/( exp_sum )

def sigmoid(x):
    y = 1/(1+ np.exp(-x) )
    
    return y



def test():
    count = 0
    for t3 in range(0,100):
        
        #forward propagation #Layer 1
        Z1 = np.dot( W1 , x_train[t3]).reshape(-1,1)  
        A1 = sigmoid( Z1 ).reshape(-1,1)         
        #forward propagation #Layer 2
        Z2 = np.dot( W2, A1)  + b2
        Output = softmax(Z2)
        
        
        arg_max =  np.argmax(Output)
        
        if arg_max == np.argmax(y_train[t3:t3+1] ):
            count = count +1
            
    
    print(count/100)        


def new_center(x,y0,c0):
    c = np.zeros( (n_class, 1) )
    y = np.argmax(y0, axis=-1) #convert to one column array
    for i in range(n_class):
        indx = np.where(y == i)[0] #choose all points that equal class i
        a1 = np.mean( x[ [indx] ].reshape(-1,n_class)  )
        c[i,:] = a1
        
    return c-c0

iter = 50000
sets = 100
loss_temp = np.zeros(( sets ,1))
loss = []
Zi_vector =  np.zeros(( sets ,n_class))
y_pred_vector = np.zeros(( sets ,n_class))
loss = np.zeros(( iter,1))

print_counter = 0
for t1 in range (0,iter):
    
    for t2 in range(0,sets):
    
        #forward propagation #Layer 1 & 2
        Z1 = np.dot( W1 , x_train[t2:t2+1].T ).reshape(-1,1)  + b1
        A1 = sigmoid( Z1 ).reshape(-1,1)         
        #forward propagation #Layer 2
        Z2 = np.dot( W2, A1)  + b2
        Output = softmax(Z2)


        #back propagation #Layer 2
        dy = Output - y_train[t2:t2+1].T #+ lambda0*(Z2- centers)
        
        dE_dZ2 = dy
        dE_dW2 = dy*A1.T
        dE_db2 = dy
        W2T = W2.T
        
        
        #backpropagation #Layer 1
        dE_dZ2T = np.zeros((n1,1))
        dA1_dZ1 = Z1*(1-Z1) 
        
        for temp1 in range(0,n1):
            dE_dZ2T[temp1] = np.dot( dy.T, W2T[temp1].T )
            
        dE_dZ2T__dA1_dZ1 = dE_dZ2T*dA1_dZ1  
        dE_dW1 = dE_dZ2T__dA1_dZ1 * x_train[t2].T
        
        
        # #For regularization
        # L2_W2 = W2*0.001
        # L2_W1 = W1*0.001
        # L2_b2 = b2*0.001
        # L2_b1 = b1*0.001
        
        #weight update
        W1 = W1 - alpha*(dE_dW1  )               
        W2 = W2 - alpha*( dE_dW2   )
        b2 = b2 - alpha*(dy   )
        b1 = b1 - alpha*( dE_dZ2T__dA1_dZ1   )
        
        
        loss_temp[t2] = (  -np.sum( y_train[t2:t2+1]*np.log10(Output.T) ) )
        

        Zi_vector[t2] = Z2.reshape(-1,)
        y_pred_vector[t2,:] = Output.reshape(-1,n_class)


    #centers = new_center(Zi_vector,y_train, centers)        
    
    loss[t1] = ( np.mean(loss_temp)) 
    loss_temp = np.zeros(( 100 ,1))


    print_counter = print_counter + 1
    if print_counter > 100:
        print(t1) 
        print_counter = 1 

        
test()

plt.plot(loss )
plt.xlabel('iterations')
plt.ylabel('loss')
plt.show()

排查思路与解决方案

1. 数值稳定性修复(解决NaN核心问题)

  • softmax溢出处理:原softmax实现未做数值稳定,当输入值较大时np.exp(x)会溢出为无穷大,导致后续计算出现NaN。修改为:
    def softmax(x):
        x_max = np.max(x)  # 减去输入最大值避免指数溢出
        exp_x = np.exp(x - x_max)
        exp_sum = np.sum(exp_x)
        return exp_x / exp_sum
    
  • 损失函数修正:原代码用np.log10计算损失,当输出趋近于0时会产生负无穷,改用自然对数(交叉熵标准实现)并加极小值避免log(0):
    loss_temp[t2] = -np.sum(y_train[t2:t2+1] * np.log(Output.T + 1e-8))
    

2. 反向传播逻辑错误修正

  • sigmoid导数计算错误:原代码用Z1*(1-Z1)计算sigmoid导数,正确的导数应为A1*(1-A1)(因为A1 = sigmoid(Z1)),修正:
    dA1_dZ1 = A1 * (1 - A1)
    
  • 梯度计算优化:原循环计算dE_dZ2T冗余且易出错,改用矩阵乘法简化:
    dE_dZ2T = np.dot(W2.T, dy)
    

3. 训练参数与流程优化

  • 权重初始化优化:原初始化方式导致权重全为正,神经元激活偏向一侧,改用适合sigmoid的Xavier初始化:
    W1 = np.random.randn(n1, nx) * np.sqrt(1/nx)
    W2 = np.random.randn(n_class, n1) * np.sqrt(1/n1)
    b1 = np.zeros((n1, 1))  # 偏置初始化为0更稳妥
    b2 = np.zeros((n_class, 1))
    
  • 学习率调整:原alpha=0.0001过小,收敛极慢,建议先尝试alpha=0.01,再根据损失曲线微调。
  • 迭代次数调整:5万次全量迭代次数过多,易导致数值不稳定,建议改为小批量梯度下降(每次取10-20个样本),并将迭代次数降至1000-5000次。

4. 测试函数修正

测试函数中Z1计算遗漏偏置b1,导致测试与训练前向传播不一致,修正:

Z1 = np.dot(W1, x_train[t3].reshape(-1,1)) + b1

内容的提问来源于stack exchange,提问作者gingerorange

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 19:34:50