You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于DQfD的强化学习项目:向量化后Loss不收敛求助

DQfD项目向量化QmaxExp后训练不收敛问题排查

我正在开发一个基于Deep Q-learning from Demonstrations(DQfD)的强化学习项目,原训练过程因损失函数中的QmaxExp函数包含for循环,速度极慢。

原QmaxExp函数

def QmaxExp(state,model,Expert_action,OutsideConditions,setpoint,Inputdf):
   maxValue = -1000000
   for i in range(len(Inputdf)):
      Actions = np.array([Inputdf['col1'][i],Inputdf['col2'][i],Inputdf['col3'][i],Inputdf['col4'][i]])
       conditions = np.array(list(OutsideConditions.values()))
       Inputs = np.concatenate((state,conditions,[setpoint]))
       Qvalue = model(Inputs.reshape((1,18)))[0,i]
       Value = Qvalue + Lfunction(Actions,Expert_action) * 0.01
       if Value > maxValue:
           maxValue = Value
   return maxValue

向量化后的QmaxExp函数

def QmaxExp(state, model, Expert_action, OutsideConditions, setpoint, Inputdf):
   conditions = np.array(list(OutsideConditions.values()))
   Inputs = np.concatenate((state, conditions, [setpoint]))
   Qvalues = model(Inputs.reshape((1, 18)))[0]
   Actions = np.array(Inputdf[['col1', 'col2', 'col3', 'col4']])
   Lfunctie_value = (np.subtract(Actions,Expert_action)**2).sum(axis=1) #use np.sum(...) when using old function
   Values = Qvalues + Lfunctie_value * 0.01
   return np.max(Values)

问题现象

向量化版本速度提升40倍,单步测试时返回的maxValue及损失值与原版本完全一致,但训练阶段出现明显差异:原版本能收敛至loss≈20,向量化版本始终无法收敛,loss维持在≈330左右。

相关代码片段

训练循环代码

with tf.GradientTape() as tape:
    rest of the code...
    loss += custom_loss(model,modelTarget,Current_state,Next_state,actions,Expert_action,Reward,gamma,OutsideConditions,Setpoint,Inputdf,HeatingInput,OtherInput)
    lossNew += custom_lossNew(model,modelTarget,Current_state,Next_state,actions,Expert_action,Reward,gamma,OutsideConditions,Setpoint,Inputdf,HeatingInput,OtherInput)
        # loss = tf.convert_to_tensor(loss)
    loss = loss/Batch_size
    lossNew = lossNew/Batch_size
    gradients = tape.gradient(loss, model.trainable_variables)
    optimizer.apply_gradients(zip(gradients, model.trainable_variables))

自定义损失函数

def custom_loss(model,modelTarget,state,new_state,action,Expert_action,reward,gamma,OutsideConditions,setpoint,Inputdf,HeatingInput,OtherInput):
   JDQ = (reward + gamma*  QmaxT1(modelTarget,new_state,OutsideConditions,setpoint) - Qvalue(state,action,model,OutsideConditions,setpoint,HeatingInput,OtherInput,Inputdf))**2
   JE = QmaxExp(state,model,Expert_action,OutsideConditions,setpoint,Inputdf) - QExp(state,model,Expert_action,OutsideConditions,setpoint,HeatingInput,OtherInput,Inputdf)
   JL2 = (reward - Qvalue(state,action,model,OutsideConditions,setpoint,HeatingInput,OtherInput,Inputdf))**2

   lambda1 = 1
   lambda2 = 1
   lambda3 = 1

   Loss = lambda1 * JDQ + lambda2 * JE + lambda3 * JL2
   return Loss

L函数定义

def Lfunction(action,Expert_action):
   return np.sum(np.subtract(action,Expert_action)**2)#.sum(axis=1) #use np.sum(...) when using old function

模型结构为:输入层 + 25节点隐藏层 + 625节点输出层。

请问向量化后的QmaxExp函数可能存在什么问题导致训练不收敛?


内容的提问来源于stack exchange,提问作者olivier

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 08:57:45