You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于机器学习实现双人追逐游戏AI训练的技术咨询

Hey there! Let's break down your problem step by step since you're looking to dive into machine learning with a practical game-based example—great choice for hands-on learning! I'll cover your core questions, actionable solutions, tool recommendations, and adjusted code to get you started.


Your Background & Core Goals

First, let's recap to make sure we're aligned:

  • You're not a professional programmer, but have basic Python skills
  • You care about general machine learning methods, the game is just a test case
  • Game Rules:
    • A rectangular arena with two players (A and B)
    • A must move to the nearest border where B starts between A and the border
    • B must move to A's position
    • A wins if reaching the target border without being touched; B wins if catching A
  • Desired Outcomes:
    a. Start with both players moving randomly
    b. AI learns autonomously to make A move toward its target and B move toward A
    c. Export AI rules for a human-vs-AI mode (human controls B, AI controls A)
  • Core Question: Does manually adjusting movement weights with time constraints count as "autonomous AI learning"?

First: Manual Weight Tuning ≠ Autonomous Learning

Short answer: No. Autonomous learning means the model improves its strategy from environment feedback (rewards/punishments) without hardcoded rules from you. Manually adjusting weights is heuristic programming—you're defining the rules, not letting the AI learn from experience. To hit goal (b), you need a reinforcement learning (RL) approach, where the AI learns through trial and error.


Solution Approach: Reinforcement Learning (RL)

RL is perfect for this game—it's designed for agents that learn to make decisions by interacting with an environment. Here's a structured plan:

1. Define the Game Environment (Critical!)

You need to formalize key components for the RL agent:

  • State Space: Represent the game state with discrete values (e.g., A's coordinates (xA, yA), B's coordinates (xB, yB), A's target border)
  • Action Space: 4 possible moves for each player (up, left, down, right)
  • Reward Function:
    • For A: +100 reward for reaching the target border; -100 for being caught; -1 per step (encourage fast movement)
    • For B: +100 for catching A; -100 for letting A escape; -1 per step (encourage fast catches)
  • Termination Conditions: A reaches the border, B catches A, or max steps are exceeded

2. Start with Single-Agent RL (Q-Learning)

Begin by training Player A first (simpler than multi-agent training):

  • Use Q-Learning: A value-based RL algorithm that maintains a "Q-table" to store the expected reward of every action in every state
  • Use ε-greedy exploration: Sometimes the agent picks a random action to explore, sometimes it picks the action with the highest Q-value to exploit learned knowledge

3. Scale to Multi-Agent Learning

Once A is trained, you can:

  • Train B to counter A's strategy using the same Q-Learning framework
  • Or use adversarial multi-agent RL (e.g., Actor-Critic) to let A and B learn against each other simultaneously

4. Export AI Rules for Human vs AI Mode

After training, convert the Q-table into a simple decision function: given A's current position and B's position, look up the action with the highest Q-value and execute it.


Python Tools & Libraries

  • numpy: Handle numerical computations and state representation
  • matplotlib: Visualize movement paths (you're already using this!)
  • stable-baselines3: Pre-built RL algorithms for fast prototyping (note: Python 2.7 is obsolete—upgrade to Python 3.8+ to use modern libraries)
  • gym: Customize your game environment for RL training

Adjusted Example Code (Python 3.x, Q-Learning for Player A)

This replaces your Python 2.7 random movement code with a trainable RL agent:

import numpy as np
import matplotlib.pyplot as plt
import random

# Game parameters
RECT_WIDTH = 8
RECT_HEIGHT = 12
MAX_STEPS = 100
NUM_EPISODES = 500
EPSILON = 0.1  # Exploration rate (10% random actions)
ALPHA = 0.1    # Learning rate
GAMMA = 0.9    # Discount factor for future rewards

# Action definitions: 0=Up, 1=Left, 2=Down, 3=Right
ACTIONS = [0, 1, 2, 3]

def get_target_border(xA, yA, xB, yB):
    """Calculate A's target border: nearest border where B is between A and the border"""
    # Calculate distance to each border
    dist_top = yA
    dist_left = xA
    dist_bottom = RECT_HEIGHT - yA
    dist_right = RECT_WIDTH - xA
    
    # Check if B is between A and each border
    valid_borders = []
    if yB < yA: valid_borders.append((dist_top, 0))  # B is between A and top border
    if xB < xA: valid_borders.append((dist_left, 1)) # B is between A and left border
    if yB > yA: valid_borders.append((dist_bottom, 2)) # B is between A and bottom border
    if xB > xA: valid_borders.append((dist_right, 3)) # B is between A and right border
    
    # Return the nearest valid border
    return min(valid_borders, key=lambda x: x[0])[1]

def get_reward(state):
    xA, yA, xB, yB, target_border = state
    done = False
    reward = -1  # Default step penalty
    
    # Check if B catches A
    if abs(xA - xB) < 1 and abs(yA - yB) < 1:
        reward = -100
        done = True
    # Check if A reaches target border
    elif (target_border == 0 and yA <= 0) or \
         (target_border == 1 and xA <= 0) or \
         (target_border == 2 and yA >= RECT_HEIGHT) or \
         (target_border == 3 and xA >= RECT_WIDTH):
        reward = 100
        done = True
    
    return reward, done

def step(state, action):
    xA, yA, xB, yB, target_border = state
    
    # Execute A's action
    if action == 0:
        yA = max(0, yA - 1)
    elif action == 1:
        xA = max(0, xA - 1)
    elif action == 2:
        yA = min(RECT_HEIGHT, yA + 1)
    elif action == 3:
        xA = min(RECT_WIDTH, xA + 1)
    
    # B moves randomly (we'll train B later)
    b_action = random.choice(ACTIONS)
    if b_action == 0:
        yB = max(0, yB - 1)
    elif b_action == 1:
        xB = max(0, xB - 1)
    elif b_action == 2:
        yB = min(RECT_HEIGHT, yB + 1)
    elif b_action == 3:
        xB = min(RECT_WIDTH, xB + 1)
    
    new_state = (xA, yA, xB, yB, target_border)
    reward, done = get_reward(new_state)
    return new_state, reward, done

def init_state():
    xA = random.randint(0, RECT_WIDTH)
    yA = random.randint(0, RECT_HEIGHT)
    xB = random.randint(0, RECT_WIDTH)
    yB = random.randint(0, RECT_HEIGHT)
    target_border = get_target_border(xA, yA, xB, yB)
    return (xA, yA, xB, yB, target_border)

# Initialize Q-table: [xA, yA, xB, yB, target_border, action]
Q = np.zeros((RECT_WIDTH+1, RECT_HEIGHT+1, RECT_WIDTH+1, RECT_HEIGHT+1, 4, len(ACTIONS)))

# Train the Q-Learning agent
for episode in range(NUM_EPISODES):
    state = init_state()
    total_reward = 0
    
    for step_idx in range(MAX_STEPS):
        xA, yA, xB, yB, target_border = state
        
        # ε-greedy action selection
        if random.uniform(0, 1) < EPSILON:
            action = random.choice(ACTIONS)
        else:
            action = np.argmax(Q[xA, yA, xB, yB, target_border])
        
        # Execute action and update state
        next_state, reward, done = step(state, action)
        xA_next, yA_next, xB_next, yB_next, target_next = next_state
        
        # Update Q-table using Bellman equation
        Q[xA, yA, xB, yB, target_border, action] += ALPHA * (
            reward + GAMMA * np.max(Q[xA_next, yA_next, xB_next, yB_next, target_next]) - 
            Q[xA, yA, xB, yB, target_border, action]
        )
        
        total_reward += reward
        state = next_state
        
        if done:
            break
    
    # Print progress every 50 episodes
    if episode % 50 == 0:
        print(f"Episode {episode:4d} | Total Reward: {total_reward:6d}")

# Test the trained AI
def test_trained_ai():
    state = init_state()
    x_path = [state[0]]
    y_path = [state[1]]
    
    for step_idx in range(MAX_STEPS):
        xA, yA, xB, yB, target_border = state
        # Choose optimal action from Q-table
        action = np.argmax(Q[xA, yA, xB, yB, target_border])
        next_state, reward, done = step(state, action)
        
        x_path.append(next_state[0])
        y_path.append(next_state[1])
        state = next_state
        
        if done:
            break
    
    # Plot results
    plt.figure(figsize=(8, 12))
    ax = plt.gca()
    ax.set_aspect('equal')
    ax.set_xlim(0, RECT_WIDTH)
    ax.set_ylim(0, RECT_HEIGHT)
    
    plt.plot(x_path, y_path, marker='o', label="AI Player A's Path")
    plt.scatter(state[2], state[3], color='red', s=50, label="Player B")
    plt.title("Trained AI Player A Movement")
    plt.legend()
    plt.show()

test_trained_ai()

Next Steps for Improvement

  • Upgrade Python: Your original code uses Python 2.7, which is no longer supported. Switch to Python 3.8+ to access modern ML libraries.
  • Train Player B: Adapt the Q-Learning code to train B to catch A, using rewards for successful catches.
  • Continuous State Space: If you want smoother movement, use DQN (Deep Q-Network) to handle continuous coordinate values instead of discrete integers.
  • Visualize Learning: Plot total reward per episode to see how the AI improves over time.

内容的提问来源于stack exchange,提问作者Matteo VR

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 04:04:02