NFSP双神经网络配置疑问:激活函数选型与实验复现求助
Great question—let's break this down against the original NFSP paper by Heinrich & Silver to validate your network setup:
Key Observations from the Paper
The paper specifies rectified linear activations (ReLU) for hidden layers, but explicitly differentiates output layer requirements based on each network's purpose:
1. Optimal Response Network (br-model)
Your current use of relu for the output layer is incorrect here. The BR network is an off-policy RL model that estimates action values Q(s,a)—these values can be positive or negative (since they represent the expected cumulative reward of taking an action in a state). Using ReLU would clamp all negative Q-values to 0, which distorts the value estimation and will prevent proper learning of optimal responses.
The correct output activation for the BR network is a linear (identity) activation, as it allows the model to output unrestricted Q-values.
2. Average Response Network (ar-model)
Your choice of softmax here is perfect. The AR network performs supervised classification to mimic your agent's average behavior, so it needs to output a valid probability distribution over actions. Softmax converts raw logits into 0-1 probabilities that sum to 1, which aligns exactly with this task.
Corrected Code Snippets
Here's how to adjust your BR model's output layer:
# Best Response Network (corrected) def _build_best_response_model(self): input_ = Input(shape=self.s_dim, name='input') hidden = Dense(self.n_hidden, activation='relu')(input_) # Use linear activation for Q-value output out = Dense(3, activation='linear')(hidden) model = Model(inputs=input_, outputs=out, name="br-model") model.compile(loss='mean_squared_error', optimizer=Adam(lr=self.lr_br), metrics=['accuracy']) return model
Your Average Response Network code is already correct:
# Average Response Network (unchanged, correct) def _build_avg_response_model(self): input_ = Input(shape=self.s_dim, name='input') hidden = Dense(self.n_hidden, activation='relu')(input_) out = Dense(3, activation='softmax')(hidden) model = Model(inputs=input_, outputs=out, name="ar-model") model.compile(loss='categorical_crossentropy', optimizer=Adam(lr=self.lr_ar), metrics=['accuracy']) return model
This adjustment should fix a critical issue in your BR network's value estimation, which was likely contributing to your inability to reproduce the paper's results.
内容的提问来源于stack exchange,提问作者David Joos

