使用Scoop并行运行线性Regression时遇NameError问题求助
The NameError: name 'y_vals' is not defined happens because Scoop's worker processes don't inherit the main process's global variables. When you run your code with python -m scoop, each worker starts as a separate Python process with its own isolated memory space—they can't access y_vals or X_vals defined in the main process's if __name__ == '__main__': block.
Fix 1: Pass Data Directly as Arguments to the Worker Function
Instead of relying on global variables, modify your linreg function to accept the target and feature data as parameters. Since futures.map passes each element of your iterable to the function, you can package your data into tuples alongside your dummy values to ensure every worker gets the data it needs.
Here's the revised code:
import pandas as pd import numpy as np import random from scoop import futures import statsmodels.api as sm from time import time def linreg(args): # Unpack the tuple: dummy value, target data, feature data _, y_vals, X_vals = args model = sm.OLS(y_vals, X_vals).fit() return model if __name__ == '__main__': random.seed(42) # Generate 10M-row dataset vals = pd.DataFrame(np.random.normal(loc=3, scale=100, size=(10000000, 5))) vals.columns = ['dep', 'ind1', 'ind2', 'ind3', 'ind4'] y_vals = vals['dep'] X_vals = vals[['ind1', 'ind2', 'ind3', 'ind4']] # Serial baseline run bt = time() model_vals = list(map(linreg, [(1, y_vals, X_vals), (2, y_vals, X_vals), (3, y_vals, X_vals)])) mval = model_vals[0] print("Serial Regression Summary:") print(mval.summary()) serial_time = time() - bt print(f"Serial Time: {serial_time:.2f}s\n") # Parallel run with Scoop bt1 = time() model_vals_1 = list(futures.map(linreg, [(1, y_vals, X_vals), (2, y_vals, X_vals), (3, y_vals, X_vals)])) mval_1 = model_vals_1[0] print("Parallel Regression Summary:") print(mval_1.summary()) parallel_time = time() - bt1 print(f"Parallel Time: {parallel_time:.2f}s\n") print(f"Serial vs Parallel Times: {serial_time:.2f}s | {parallel_time:.2f}s")
Key Changes Explained:
- Redesigned
linregfunction: It now takes a single tuple argument containing the dummy value, target data, and feature data. This ensures every worker process receives the dataset directly, no global variable dependencies. - Updated iterables: Instead of passing just
[1,2,3], we pass tuples that include the data. This makes both serial and parallel calls consistent in how they access the dataset. - Removed global variables: Eliminates cross-process variable access issues entirely.
Fix 2: Use Scoop's Shared Memory (For Large Datasets)
If your 10M-row dataset is too large to copy to each worker, use scoop.shared to store the data once and make it accessible to all workers. This avoids redundant data duplication across processes:
import pandas as pd import numpy as np import random from scoop import futures, shared import statsmodels.api as sm from time import time def linreg(_): # Retrieve shared data in the worker process y_vals = shared.get('y_vals') X_vals = shared.get('X_vals') model = sm.OLS(y_vals, X_vals).fit() return model if __name__ == '__main__': random.seed(42) vals = pd.DataFrame(np.random.normal(loc=3, scale=100, size=(10000000, 5))) vals.columns = ['dep', 'ind1', 'ind2', 'ind3', 'ind4'] y_vals = vals['dep'] X_vals = vals[['ind1', 'ind2', 'ind3', 'ind4']] # Store data in Scoop's shared memory pool shared.set('y_vals', y_vals) shared.set('X_vals', X_vals) # Serial baseline run bt = time() model_vals = list(map(linreg, [1,2,3])) mval = model_vals[0] print("Serial Regression Summary:") print(mval.summary()) serial_time = time() - bt print(f"Serial Time: {serial_time:.2f}s\n") # Parallel run with Scoop bt1 = time() model_vals_1 = list(futures.map(linreg, [1,2,3])) mval_1 = model_vals_1[0] print("Parallel Regression Summary:") print(mval_1.summary()) parallel_time = time() - bt1 print(f"Parallel Time: {parallel_time:.2f}s\n") print(f"Serial vs Parallel Times: {serial_time:.2f}s | {parallel_time:.2f}s")
This approach is more memory-efficient for large datasets, as the data is serialized once and shared across all workers instead of being copied multiple times.
Critical Reminder:
Always run your script with python -m scoop your_script.py to enable Scoop's parallelism. Without the -m scoop flag, futures.map falls back to Python's built-in serial map, which won't give you the parallel execution you need.
内容的提问来源于stack exchange,提问作者Umbp

