无精确输出标签时如何训练神经网络?Python/Keras步行机器人方案咨询
Great question! When you don't have precise labeled outputs (the kind backpropagation relies on) but instead a holistic fitness score (like how far your walker bot travels or how stable it stays upright), you want Neuroevolution—combining neural networks with evolutionary algorithms. This approach is tailor-made for your walker bot use case, and I’ll walk you through how to implement it with Python and Keras.
Why Backpropagation Isn’t the Right Fit Here
Backpropagation needs a clear target output for every input to calculate error gradients. For your walker bot, you don’t have "correct" joint angles or motor commands for every sensor reading—you only have a high-level score of how well the bot performed overall. Neuroevolution solves this by treating neural network weights as "genes" and evolving them over generations, just like natural selection.
Core Neuroevolution Workflow for Your Walker Bot
This is the loop you’ll implement:
- Initialize a population: Create a group of neural networks with random weights (each network controls one walker bot).
- Evaluate fitness: Let each bot walk in its environment, then assign a fitness score (e.g., distance traveled, time upright, energy efficiency).
- Select elite performers: Keep the top-scoring bots (their weights are the "best genes" of the generation).
- Generate new population: Breed new bots by crossing weights of elite parents and adding small random mutations to avoid stagnation.
- Repeat: Iterate until your bot’s fitness meets your goals.
Python/Keras Implementation
Below is a working example you can adapt to your specific walker bot’s sensors and actuators.
Step 1: Setup Dependencies
import numpy as np from keras.models import Sequential from keras.layers import Dense from keras.models import load_model
Step 2: Define Your Neural Network
This network takes sensor inputs (e.g., joint angles, accelerometer data) and outputs motor commands (e.g., joint torque/angle targets).
def create_walker_model(input_dim, output_dim): """Create a simple feedforward network for the walker bot.""" model = Sequential() model.add(Dense(32, activation='relu', input_dim=input_dim)) model.add(Dense(16, activation='relu')) model.add(Dense(output_dim, activation='tanh')) # Output range: -1 to 1 for motor controls return model
Step 3: Helper Functions for Weight Manipulation
We need to convert Keras model weights to flat vectors (for easy genetic operations) and back.
def model_to_weights(model): """Convert Keras model weights to a single flat numpy array.""" return np.concatenate([weight.flatten() for weight in model.get_weights()]) def weights_to_model(weights, base_model): """Restore a Keras model from a flat weight vector.""" weight_shapes = [w.shape for w in base_model.get_weights()] current_idx = 0 restored_weights = [] for shape in weight_shapes: weight_size = np.prod(shape) restored_weights.append(weights[current_idx:current_idx+weight_size].reshape(shape)) current_idx += weight_size base_model.set_weights(restored_weights) return base_model
Step 4: Fitness Evaluation
Replace this with your actual walker bot simulation/real-world code. The goal is to assign a higher score to better-performing bots.
def evaluate_walker_fitness(weights, input_dim, output_dim): """Calculate fitness for a single walker bot.""" model = create_walker_model(input_dim, output_dim) model = weights_to_model(weights, model) # Replace this with your actual environment logic total_distance = 0 max_steps = 150 # Simulate 150 steps of walking for _ in range(max_steps): # Get sensor input (e.g., joint angles, accelerometer data) sensor_input = np.random.rand(1, input_dim) # Replace with real sensor data # Get motor command from the network motor_cmd = model.predict(sensor_input, verbose=0) # Update bot state and calculate distance traveled total_distance += np.mean(np.abs(motor_cmd)) * 0.1 # Example: reward smooth, consistent movement return total_distance
Step 5: Evolutionary Operations
def select_elite(population, fitness_scores, elite_size): """Select the top-performing bots from the population.""" sorted_indices = np.argsort(fitness_scores)[::-1] # Sort descending by fitness return [population[i] for i in sorted_indices[:elite_size]] def crossover(parent1_weights, parent2_weights): """Combine weights from two parents to create a child.""" crossover_point = np.random.randint(1, len(parent1_weights)) return np.concatenate([parent1_weights[:crossover_point], parent2_weights[crossover_point:]]) def mutate(weights, mutation_rate=0.01, mutation_scale=0.1): """Add small random noise to weights to encourage exploration.""" mutation_mask = np.random.rand(len(weights)) < mutation_rate mutation_noise = np.random.normal(0, mutation_scale, len(weights)) weights[mutation_mask] += mutation_noise[mutation_mask] return weights
Step 6: Main Training Loop
def train_walker(input_dim=8, output_dim=4, pop_size=50, elite_size=10, generations=100): """Main neuroevolution training loop.""" # Initialize population with random weights population = [] for _ in range(pop_size): model = create_walker_model(input_dim, output_dim) population.append(model_to_weights(model)) best_fitness_history = [] for gen in range(generations): # Evaluate all bots in the population fitness_scores = [evaluate_walker_fitness(w, input_dim, output_dim) for w in population] best_fitness = max(fitness_scores) best_fitness_history.append(best_fitness) print(f"Generation {gen+1} | Best Fitness: {best_fitness:.2f}") # Select elite performers elite = select_elite(population, fitness_scores, elite_size) # Generate new population: keep elite + breed new bots new_population = elite.copy() while len(new_population) < pop_size: # Pick random elite parents parent1 = elite[np.random.randint(len(elite))] parent2 = elite[np.random.randint(len(elite))] # Breed and mutate child = crossover(parent1, parent2) child = mutate(child) new_population.append(child) population = new_population # Save the best model final_fitness_scores = [evaluate_walker_fitness(w, input_dim, output_dim) for w in population] best_idx = np.argmax(final_fitness_scores) best_weights = population[best_idx] best_model = create_walker_model(input_dim, output_dim) best_model = weights_to_model(best_weights, best_model) best_model.save("walker_best_model.h5") return best_model, best_fitness_history # Run the training (adjust input/output dims to match your bot) best_walker_model, fitness_history = train_walker(input_dim=8, output_dim=4)
Key Tips for Success
- Tune your fitness function: This is the most critical part. Make sure it rewards exactly the behavior you want (e.g., prioritize distance over speed, penalize falls).
- Adjust population/elite sizes: Too small a population risks getting stuck in local optima; too large slows down training. Aim for 30-100 bots per generation.
- Tweak mutation rates: Start with a 1-2% mutation rate and small scale (0.1). Increase if evolution stagnates, decrease if good solutions are broken.
- Align simulation with reality: If training in a simulator, make sure sensor data and physics match your real walker bot to avoid "simulator drift."
内容的提问来源于stack exchange,提问作者Rasmus Anker Fossen Nordal

