基于机器学习实现双人追逐游戏AI训练的技术咨询
Hey there! Let's break down your problem step by step since you're looking to dive into machine learning with a practical game-based example—great choice for hands-on learning! I'll cover your core questions, actionable solutions, tool recommendations, and adjusted code to get you started.
Your Background & Core Goals
First, let's recap to make sure we're aligned:
- You're not a professional programmer, but have basic Python skills
- You care about general machine learning methods, the game is just a test case
- Game Rules:
- A rectangular arena with two players (A and B)
- A must move to the nearest border where B starts between A and the border
- B must move to A's position
- A wins if reaching the target border without being touched; B wins if catching A
- Desired Outcomes:
a. Start with both players moving randomly
b. AI learns autonomously to make A move toward its target and B move toward A
c. Export AI rules for a human-vs-AI mode (human controls B, AI controls A) - Core Question: Does manually adjusting movement weights with time constraints count as "autonomous AI learning"?
First: Manual Weight Tuning ≠ Autonomous Learning
Short answer: No. Autonomous learning means the model improves its strategy from environment feedback (rewards/punishments) without hardcoded rules from you. Manually adjusting weights is heuristic programming—you're defining the rules, not letting the AI learn from experience. To hit goal (b), you need a reinforcement learning (RL) approach, where the AI learns through trial and error.
Solution Approach: Reinforcement Learning (RL)
RL is perfect for this game—it's designed for agents that learn to make decisions by interacting with an environment. Here's a structured plan:
1. Define the Game Environment (Critical!)
You need to formalize key components for the RL agent:
- State Space: Represent the game state with discrete values (e.g., A's coordinates
(xA, yA), B's coordinates(xB, yB), A's target border) - Action Space: 4 possible moves for each player (up, left, down, right)
- Reward Function:
- For A: +100 reward for reaching the target border; -100 for being caught; -1 per step (encourage fast movement)
- For B: +100 for catching A; -100 for letting A escape; -1 per step (encourage fast catches)
- Termination Conditions: A reaches the border, B catches A, or max steps are exceeded
2. Start with Single-Agent RL (Q-Learning)
Begin by training Player A first (simpler than multi-agent training):
- Use Q-Learning: A value-based RL algorithm that maintains a "Q-table" to store the expected reward of every action in every state
- Use ε-greedy exploration: Sometimes the agent picks a random action to explore, sometimes it picks the action with the highest Q-value to exploit learned knowledge
3. Scale to Multi-Agent Learning
Once A is trained, you can:
- Train B to counter A's strategy using the same Q-Learning framework
- Or use adversarial multi-agent RL (e.g., Actor-Critic) to let A and B learn against each other simultaneously
4. Export AI Rules for Human vs AI Mode
After training, convert the Q-table into a simple decision function: given A's current position and B's position, look up the action with the highest Q-value and execute it.
Python Tools & Libraries
numpy: Handle numerical computations and state representationmatplotlib: Visualize movement paths (you're already using this!)stable-baselines3: Pre-built RL algorithms for fast prototyping (note: Python 2.7 is obsolete—upgrade to Python 3.8+ to use modern libraries)gym: Customize your game environment for RL training
Adjusted Example Code (Python 3.x, Q-Learning for Player A)
This replaces your Python 2.7 random movement code with a trainable RL agent:
import numpy as np import matplotlib.pyplot as plt import random # Game parameters RECT_WIDTH = 8 RECT_HEIGHT = 12 MAX_STEPS = 100 NUM_EPISODES = 500 EPSILON = 0.1 # Exploration rate (10% random actions) ALPHA = 0.1 # Learning rate GAMMA = 0.9 # Discount factor for future rewards # Action definitions: 0=Up, 1=Left, 2=Down, 3=Right ACTIONS = [0, 1, 2, 3] def get_target_border(xA, yA, xB, yB): """Calculate A's target border: nearest border where B is between A and the border""" # Calculate distance to each border dist_top = yA dist_left = xA dist_bottom = RECT_HEIGHT - yA dist_right = RECT_WIDTH - xA # Check if B is between A and each border valid_borders = [] if yB < yA: valid_borders.append((dist_top, 0)) # B is between A and top border if xB < xA: valid_borders.append((dist_left, 1)) # B is between A and left border if yB > yA: valid_borders.append((dist_bottom, 2)) # B is between A and bottom border if xB > xA: valid_borders.append((dist_right, 3)) # B is between A and right border # Return the nearest valid border return min(valid_borders, key=lambda x: x[0])[1] def get_reward(state): xA, yA, xB, yB, target_border = state done = False reward = -1 # Default step penalty # Check if B catches A if abs(xA - xB) < 1 and abs(yA - yB) < 1: reward = -100 done = True # Check if A reaches target border elif (target_border == 0 and yA <= 0) or \ (target_border == 1 and xA <= 0) or \ (target_border == 2 and yA >= RECT_HEIGHT) or \ (target_border == 3 and xA >= RECT_WIDTH): reward = 100 done = True return reward, done def step(state, action): xA, yA, xB, yB, target_border = state # Execute A's action if action == 0: yA = max(0, yA - 1) elif action == 1: xA = max(0, xA - 1) elif action == 2: yA = min(RECT_HEIGHT, yA + 1) elif action == 3: xA = min(RECT_WIDTH, xA + 1) # B moves randomly (we'll train B later) b_action = random.choice(ACTIONS) if b_action == 0: yB = max(0, yB - 1) elif b_action == 1: xB = max(0, xB - 1) elif b_action == 2: yB = min(RECT_HEIGHT, yB + 1) elif b_action == 3: xB = min(RECT_WIDTH, xB + 1) new_state = (xA, yA, xB, yB, target_border) reward, done = get_reward(new_state) return new_state, reward, done def init_state(): xA = random.randint(0, RECT_WIDTH) yA = random.randint(0, RECT_HEIGHT) xB = random.randint(0, RECT_WIDTH) yB = random.randint(0, RECT_HEIGHT) target_border = get_target_border(xA, yA, xB, yB) return (xA, yA, xB, yB, target_border) # Initialize Q-table: [xA, yA, xB, yB, target_border, action] Q = np.zeros((RECT_WIDTH+1, RECT_HEIGHT+1, RECT_WIDTH+1, RECT_HEIGHT+1, 4, len(ACTIONS))) # Train the Q-Learning agent for episode in range(NUM_EPISODES): state = init_state() total_reward = 0 for step_idx in range(MAX_STEPS): xA, yA, xB, yB, target_border = state # ε-greedy action selection if random.uniform(0, 1) < EPSILON: action = random.choice(ACTIONS) else: action = np.argmax(Q[xA, yA, xB, yB, target_border]) # Execute action and update state next_state, reward, done = step(state, action) xA_next, yA_next, xB_next, yB_next, target_next = next_state # Update Q-table using Bellman equation Q[xA, yA, xB, yB, target_border, action] += ALPHA * ( reward + GAMMA * np.max(Q[xA_next, yA_next, xB_next, yB_next, target_next]) - Q[xA, yA, xB, yB, target_border, action] ) total_reward += reward state = next_state if done: break # Print progress every 50 episodes if episode % 50 == 0: print(f"Episode {episode:4d} | Total Reward: {total_reward:6d}") # Test the trained AI def test_trained_ai(): state = init_state() x_path = [state[0]] y_path = [state[1]] for step_idx in range(MAX_STEPS): xA, yA, xB, yB, target_border = state # Choose optimal action from Q-table action = np.argmax(Q[xA, yA, xB, yB, target_border]) next_state, reward, done = step(state, action) x_path.append(next_state[0]) y_path.append(next_state[1]) state = next_state if done: break # Plot results plt.figure(figsize=(8, 12)) ax = plt.gca() ax.set_aspect('equal') ax.set_xlim(0, RECT_WIDTH) ax.set_ylim(0, RECT_HEIGHT) plt.plot(x_path, y_path, marker='o', label="AI Player A's Path") plt.scatter(state[2], state[3], color='red', s=50, label="Player B") plt.title("Trained AI Player A Movement") plt.legend() plt.show() test_trained_ai()
Next Steps for Improvement
- Upgrade Python: Your original code uses Python 2.7, which is no longer supported. Switch to Python 3.8+ to access modern ML libraries.
- Train Player B: Adapt the Q-Learning code to train B to catch A, using rewards for successful catches.
- Continuous State Space: If you want smoother movement, use DQN (Deep Q-Network) to handle continuous coordinate values instead of discrete integers.
- Visualize Learning: Plot total reward per episode to see how the AI improves over time.
内容的提问来源于stack exchange,提问作者Matteo VR

