连续动作空间下REINFORCE算法实现咨询(humanoid-v2环境)
Absolutely! REINFORCE (and policy gradient methods broadly) work perfectly well with continuous action spaces, and you can absolutely implement REINFORCE for OpenAI Gym's humanoid-v2 environment. Let’s walk through how this works and what you need to adjust from the discrete action space version:
The core idea of REINFORCE (maximizing the expected discounted reward via policy gradients) stays the same, but we need to tweak how we represent and sample actions:
- Policy Representation: Instead of outputting a categorical probability distribution (for discrete actions), we use a continuous probability distribution (most commonly a Gaussian distribution) to model the action space. The policy network will output the mean (
mu) and standard deviation (sigma) of this distribution for each action dimension. - Action Sampling: We sample actions from the Gaussian distribution defined by
muandsigma. To make this sampling process differentiable (so we can compute gradients through it), we use the reparameterization trick:action = mu + sigma * epsilon, whereepsilonis sampled from a standard normal distribution. - Gradient Calculation: Instead of using the log-probability of a discrete action, we use the log-probability density function (log-PDF) of the continuous action under our Gaussian distribution. The policy gradient update remains
log_pdf(action) * discounted_return, with a negative sign added because we minimize loss in PyTorch/TensorFlow.
humanoid-v2 Here’s a simplified PyTorch implementation tailored to humanoid-v2:
First, import dependencies:
import gym import torch import torch.nn as nn import torch.optim as optim import numpy as np from torch.distributions import Normal
Define the policy network (outputs Gaussian distribution parameters):
class PolicyNetwork(nn.Module): def __init__(self, state_dim, action_dim): super().__init__() # Humanoid has a high-dimensional observation space, so we use two hidden layers self.fc1 = nn.Linear(state_dim, 128) self.fc2 = nn.Linear(128, 64) self.mu_head = nn.Linear(64, action_dim) self.log_sigma_head = nn.Linear(64, action_dim) def forward(self, x): x = torch.tanh(self.fc1(x)) x = torch.tanh(self.fc2(x)) mu = self.mu_head(x) # Output log(sigma) instead of sigma directly to ensure sigma stays positive log_sigma = self.log_sigma_head(x) sigma = torch.exp(log_sigma) return mu, sigma
Implement the REINFORCE training loop:
def train_reinforce(env, policy_net, optimizer, num_episodes, gamma=0.99): for episode_idx in range(num_episodes): state = env.reset() episode_log_probs = [] episode_rewards = [] # Run one episode while True: # Convert state to tensor state_tensor = torch.tensor(state, dtype=torch.float32).unsqueeze(0) # Get distribution parameters from policy mu, sigma = policy_net(state_tensor) dist = Normal(mu, sigma) # Sample action and compute its log PDF action = dist.sample() log_prob = dist.log_prob(action).sum() # Sum over action dimensions # Step environment next_state, reward, done, _ = env.step(action.detach().numpy()[0]) # Store data for update episode_log_probs.append(log_prob) episode_rewards.append(reward) state = next_state if done: break # Compute discounted rewards (REINFORCE core) discounted_rewards = [] running_reward = 0 # Iterate rewards in reverse to compute cumulative discounted return for r in reversed(episode_rewards): running_reward = r + gamma * running_reward discounted_rewards.insert(0, running_reward) # Normalize rewards to stabilize training (critical for complex envs like Humanoid) discounted_rewards = torch.tensor(discounted_rewards, dtype=torch.float32) discounted_rewards = (discounted_rewards - discounted_rewards.mean()) / (discounted_rewards.std() + 1e-9) # Calculate policy loss loss = 0 for log_prob, disc_reward in zip(episode_log_probs, discounted_rewards): # Negative sign because we want to maximize the objective (PyTorch minimizes loss) loss += -log_prob * disc_reward # Update policy network optimizer.zero_grad() loss.backward() optimizer.step() # Log progress every 10 episodes if episode_idx % 10 == 0: print(f"Episode {episode_idx:4d} | Total Reward: {sum(episode_rewards):.2f}") # Initialize environment and components env = gym.make('Humanoid-v2') state_dim = env.observation_space.shape[0] action_dim = env.action_space.shape[0] policy_net = PolicyNetwork(state_dim, action_dim) optimizer = optim.Adam(policy_net.parameters(), lr=3e-4) # Start training (expect 1000+ episodes to see meaningful progress) train_reinforce(env, policy_net, optimizer, num_episodes=2000)
humanoid-v2 - Training Stability: REINFORCE has high variance, so reward normalization is almost mandatory here. You might also want to try adding a baseline (like a value network) to reduce variance—this evolves into the Actor-Critic algorithm, which converges faster than vanilla REINFORCE.
- Network Capacity: Humanoid has a 376-dimensional observation space and 17-dimensional action space, so the network needs enough capacity to model the complex policy. Feel free to increase hidden layer sizes or add more layers if training is slow.
- Exploration: The learned
sigmacontrols exploration. If the agent isn’t exploring enough, you can initializelog_sigma_headto output higher values (e.g.,nn.init.constant_(self.log_sigma_head.weight, 0.0)gives initial sigma=1). - Training Time: Expect to run thousands of episodes before the agent starts walking coherently—continuous control tasks are computationally intensive!
内容的提问来源于stack exchange,提问作者Pythoncoder

