You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

连续动作空间下REINFORCE算法实现咨询(humanoid-v2环境)

Absolutely! REINFORCE (and policy gradient methods broadly) work perfectly well with continuous action spaces, and you can absolutely implement REINFORCE for OpenAI Gym's humanoid-v2 environment. Let’s walk through how this works and what you need to adjust from the discrete action space version:

Key Adaptations for Continuous Action Spaces

The core idea of REINFORCE (maximizing the expected discounted reward via policy gradients) stays the same, but we need to tweak how we represent and sample actions:

  • Policy Representation: Instead of outputting a categorical probability distribution (for discrete actions), we use a continuous probability distribution (most commonly a Gaussian distribution) to model the action space. The policy network will output the mean (mu) and standard deviation (sigma) of this distribution for each action dimension.
  • Action Sampling: We sample actions from the Gaussian distribution defined by mu and sigma. To make this sampling process differentiable (so we can compute gradients through it), we use the reparameterization trick: action = mu + sigma * epsilon, where epsilon is sampled from a standard normal distribution.
  • Gradient Calculation: Instead of using the log-probability of a discrete action, we use the log-probability density function (log-PDF) of the continuous action under our Gaussian distribution. The policy gradient update remains log_pdf(action) * discounted_return, with a negative sign added because we minimize loss in PyTorch/TensorFlow.
Implementing REINFORCE on humanoid-v2

Here’s a simplified PyTorch implementation tailored to humanoid-v2:

First, import dependencies:

import gym
import torch
import torch.nn as nn
import torch.optim as optim
import numpy as np
from torch.distributions import Normal

Define the policy network (outputs Gaussian distribution parameters):

class PolicyNetwork(nn.Module):
    def __init__(self, state_dim, action_dim):
        super().__init__()
        # Humanoid has a high-dimensional observation space, so we use two hidden layers
        self.fc1 = nn.Linear(state_dim, 128)
        self.fc2 = nn.Linear(128, 64)
        self.mu_head = nn.Linear(64, action_dim)
        self.log_sigma_head = nn.Linear(64, action_dim)
        
    def forward(self, x):
        x = torch.tanh(self.fc1(x))
        x = torch.tanh(self.fc2(x))
        mu = self.mu_head(x)
        # Output log(sigma) instead of sigma directly to ensure sigma stays positive
        log_sigma = self.log_sigma_head(x)
        sigma = torch.exp(log_sigma)
        return mu, sigma

Implement the REINFORCE training loop:

def train_reinforce(env, policy_net, optimizer, num_episodes, gamma=0.99):
    for episode_idx in range(num_episodes):
        state = env.reset()
        episode_log_probs = []
        episode_rewards = []
        
        # Run one episode
        while True:
            # Convert state to tensor
            state_tensor = torch.tensor(state, dtype=torch.float32).unsqueeze(0)
            # Get distribution parameters from policy
            mu, sigma = policy_net(state_tensor)
            dist = Normal(mu, sigma)
            # Sample action and compute its log PDF
            action = dist.sample()
            log_prob = dist.log_prob(action).sum()  # Sum over action dimensions
            
            # Step environment
            next_state, reward, done, _ = env.step(action.detach().numpy()[0])
            
            # Store data for update
            episode_log_probs.append(log_prob)
            episode_rewards.append(reward)
            state = next_state
            
            if done:
                break
        
        # Compute discounted rewards (REINFORCE core)
        discounted_rewards = []
        running_reward = 0
        # Iterate rewards in reverse to compute cumulative discounted return
        for r in reversed(episode_rewards):
            running_reward = r + gamma * running_reward
            discounted_rewards.insert(0, running_reward)
        
        # Normalize rewards to stabilize training (critical for complex envs like Humanoid)
        discounted_rewards = torch.tensor(discounted_rewards, dtype=torch.float32)
        discounted_rewards = (discounted_rewards - discounted_rewards.mean()) / (discounted_rewards.std() + 1e-9)
        
        # Calculate policy loss
        loss = 0
        for log_prob, disc_reward in zip(episode_log_probs, discounted_rewards):
            # Negative sign because we want to maximize the objective (PyTorch minimizes loss)
            loss += -log_prob * disc_reward
        
        # Update policy network
        optimizer.zero_grad()
        loss.backward()
        optimizer.step()
        
        # Log progress every 10 episodes
        if episode_idx % 10 == 0:
            print(f"Episode {episode_idx:4d} | Total Reward: {sum(episode_rewards):.2f}")

# Initialize environment and components
env = gym.make('Humanoid-v2')
state_dim = env.observation_space.shape[0]
action_dim = env.action_space.shape[0]

policy_net = PolicyNetwork(state_dim, action_dim)
optimizer = optim.Adam(policy_net.parameters(), lr=3e-4)

# Start training (expect 1000+ episodes to see meaningful progress)
train_reinforce(env, policy_net, optimizer, num_episodes=2000)
Critical Notes for humanoid-v2
  • Training Stability: REINFORCE has high variance, so reward normalization is almost mandatory here. You might also want to try adding a baseline (like a value network) to reduce variance—this evolves into the Actor-Critic algorithm, which converges faster than vanilla REINFORCE.
  • Network Capacity: Humanoid has a 376-dimensional observation space and 17-dimensional action space, so the network needs enough capacity to model the complex policy. Feel free to increase hidden layer sizes or add more layers if training is slow.
  • Exploration: The learned sigma controls exploration. If the agent isn’t exploring enough, you can initialize log_sigma_head to output higher values (e.g., nn.init.constant_(self.log_sigma_head.weight, 0.0) gives initial sigma=1).
  • Training Time: Expect to run thousands of episodes before the agent starts walking coherently—continuous control tasks are computationally intensive!

内容的提问来源于stack exchange,提问作者Pythoncoder

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 10:10:34