请求将多进程代码改为单进程实现(适配Windows运行环境)
Fixing Multiprocessing Compatibility on Windows: Single-Process Refactor
Got it, let's get this code working smoothly on Windows by stripping out the multiprocessing bits and switching to a single-process implementation. The core issue with your original code on Windows is how Python spawns processes (it uses spawn instead of fork), which can cause pickling errors and unexpected behavior with shared objects like your classifier or dataset.
Here's a cleaned-up, single-process version that retains your original logic but avoids all multiprocessing headaches:
def fitness(self, generation): def calculate_individual_fitness(individual): # Subset columns based on the individual's feature selection selected_cols = [self.dataset.columns[j] for j in range(len(individual)) if individual[j] == 1] X_individual = self.dataset[selected_cols] # Split into train/test sets X_train, X_test, y_train, y_test = train_test_split( X_individual, self.response, test_size=self.test_size ) # Train and score the classifier self.algorithm.fit(X_train, y_train) return self.algorithm.score(X_test, y_test) # Initialize array to hold all fitness scores fitness_scores = np.zeros(generation.shape[0]) # Iterate over every individual in the generation for idx, individual in enumerate(generation): fitness_scores[idx] = calculate_individual_fitness(individual) return fitness_scores
Key Changes Made:
- Removed all multiprocessing code: No more
mp.Queue,mp.Process, or process startup/join calls. This eliminates Windows-specific spawn/pickling issues entirely. - Simplified helper function:
calculate_individual_fitnessdirectly returns the score for a single individual instead of pushing results to a queue. - Straightforward iteration: We loop through each individual in the generation directly, calculating fitness one by one and storing results in a pre-allocated array (just like your original code did internally).
If you prefer to keep the subarray splitting logic (to mirror your original structure), here's an alternative version that maintains that pattern without processes:
def fitness(self, generation): def process_subgeneration(subgeneration): fitness_array = np.zeros(subgeneration.shape[0]) for i in range(subgeneration.shape[0]): individual = subgeneration[i, :] selected_cols = [self.dataset.columns[j] for j in range(len(individual)) if individual[j] == 1] X_individual = self.dataset[selected_cols] X_train, X_test, y_train, y_test = train_test_split( X_individual, self.response, test_size=self.test_size ) self.algorithm.fit(X_train, y_train) fitness_array[i] = self.algorithm.score(X_test, y_test) return fitness_array # Split generation into subarrays (same as before, no processes) subarrays = np.array_split(generation, self.nprocs) all_fitness = [] # Process each subarray sequentially for subarray in subarrays: all_fitness.extend(process_subgeneration(subarray)) return np.array(all_fitness)
Notes for Windows Users:
- This code will run without any spawn/pickling errors since it's entirely single-threaded/single-process.
- While it won't have the parallel speedup of multiprocessing, it's fully compatible with Windows. If you later want to add back parallelism without
multiprocessingissues, considerconcurrent.futures.ThreadPoolExecutor(though keep in mind the GIL may limit gains for CPU-heavy model training).
内容的提问来源于stack exchange,提问作者Moy Zark
相关产品推荐
相关产品推荐

