基于DQfD的强化学习项目:向量化后Loss不收敛求助
DQfD项目向量化QmaxExp后训练不收敛问题排查
我正在开发一个基于Deep Q-learning from Demonstrations(DQfD)的强化学习项目,原训练过程因损失函数中的QmaxExp函数包含for循环,速度极慢。
原QmaxExp函数
def QmaxExp(state,model,Expert_action,OutsideConditions,setpoint,Inputdf): maxValue = -1000000 for i in range(len(Inputdf)): Actions = np.array([Inputdf['col1'][i],Inputdf['col2'][i],Inputdf['col3'][i],Inputdf['col4'][i]]) conditions = np.array(list(OutsideConditions.values())) Inputs = np.concatenate((state,conditions,[setpoint])) Qvalue = model(Inputs.reshape((1,18)))[0,i] Value = Qvalue + Lfunction(Actions,Expert_action) * 0.01 if Value > maxValue: maxValue = Value return maxValue
向量化后的QmaxExp函数
def QmaxExp(state, model, Expert_action, OutsideConditions, setpoint, Inputdf): conditions = np.array(list(OutsideConditions.values())) Inputs = np.concatenate((state, conditions, [setpoint])) Qvalues = model(Inputs.reshape((1, 18)))[0] Actions = np.array(Inputdf[['col1', 'col2', 'col3', 'col4']]) Lfunctie_value = (np.subtract(Actions,Expert_action)**2).sum(axis=1) #use np.sum(...) when using old function Values = Qvalues + Lfunctie_value * 0.01 return np.max(Values)
问题现象
向量化版本速度提升40倍,单步测试时返回的maxValue及损失值与原版本完全一致,但训练阶段出现明显差异:原版本能收敛至loss≈20,向量化版本始终无法收敛,loss维持在≈330左右。
相关代码片段
训练循环代码
with tf.GradientTape() as tape: rest of the code... loss += custom_loss(model,modelTarget,Current_state,Next_state,actions,Expert_action,Reward,gamma,OutsideConditions,Setpoint,Inputdf,HeatingInput,OtherInput) lossNew += custom_lossNew(model,modelTarget,Current_state,Next_state,actions,Expert_action,Reward,gamma,OutsideConditions,Setpoint,Inputdf,HeatingInput,OtherInput) # loss = tf.convert_to_tensor(loss) loss = loss/Batch_size lossNew = lossNew/Batch_size gradients = tape.gradient(loss, model.trainable_variables) optimizer.apply_gradients(zip(gradients, model.trainable_variables))
自定义损失函数
def custom_loss(model,modelTarget,state,new_state,action,Expert_action,reward,gamma,OutsideConditions,setpoint,Inputdf,HeatingInput,OtherInput): JDQ = (reward + gamma* QmaxT1(modelTarget,new_state,OutsideConditions,setpoint) - Qvalue(state,action,model,OutsideConditions,setpoint,HeatingInput,OtherInput,Inputdf))**2 JE = QmaxExp(state,model,Expert_action,OutsideConditions,setpoint,Inputdf) - QExp(state,model,Expert_action,OutsideConditions,setpoint,HeatingInput,OtherInput,Inputdf) JL2 = (reward - Qvalue(state,action,model,OutsideConditions,setpoint,HeatingInput,OtherInput,Inputdf))**2 lambda1 = 1 lambda2 = 1 lambda3 = 1 Loss = lambda1 * JDQ + lambda2 * JE + lambda3 * JL2 return Loss
L函数定义
def Lfunction(action,Expert_action): return np.sum(np.subtract(action,Expert_action)**2)#.sum(axis=1) #use np.sum(...) when using old function
模型结构为:输入层 + 25节点隐藏层 + 625节点输出层。
请问向量化后的QmaxExp函数可能存在什么问题导致训练不收敛?
内容的提问来源于stack exchange,提问作者olivier
相关产品推荐
相关产品推荐

