You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scoop并行运行线性Regression时遇NameError问题求助

Solving NameError with Scoop Parallel Linear Regression

The NameError: name 'y_vals' is not defined happens because Scoop's worker processes don't inherit the main process's global variables. When you run your code with python -m scoop, each worker starts as a separate Python process with its own isolated memory space—they can't access y_vals or X_vals defined in the main process's if __name__ == '__main__': block.

Fix 1: Pass Data Directly as Arguments to the Worker Function

Instead of relying on global variables, modify your linreg function to accept the target and feature data as parameters. Since futures.map passes each element of your iterable to the function, you can package your data into tuples alongside your dummy values to ensure every worker gets the data it needs.

Here's the revised code:

import pandas as pd
import numpy as np
import random
from scoop import futures
import statsmodels.api as sm
from time import time

def linreg(args):
    # Unpack the tuple: dummy value, target data, feature data
    _, y_vals, X_vals = args
    model = sm.OLS(y_vals, X_vals).fit()
    return model

if __name__ == '__main__':
    random.seed(42)
    # Generate 10M-row dataset
    vals = pd.DataFrame(np.random.normal(loc=3, scale=100, size=(10000000, 5)))
    vals.columns = ['dep', 'ind1', 'ind2', 'ind3', 'ind4']
    y_vals = vals['dep']
    X_vals = vals[['ind1', 'ind2', 'ind3', 'ind4']]
    
    # Serial baseline run
    bt = time()
    model_vals = list(map(linreg, [(1, y_vals, X_vals), (2, y_vals, X_vals), (3, y_vals, X_vals)]))
    mval = model_vals[0]
    print("Serial Regression Summary:")
    print(mval.summary())
    serial_time = time() - bt
    print(f"Serial Time: {serial_time:.2f}s\n")
    
    # Parallel run with Scoop
    bt1 = time()
    model_vals_1 = list(futures.map(linreg, [(1, y_vals, X_vals), (2, y_vals, X_vals), (3, y_vals, X_vals)]))
    mval_1 = model_vals_1[0]
    print("Parallel Regression Summary:")
    print(mval_1.summary())
    parallel_time = time() - bt1
    print(f"Parallel Time: {parallel_time:.2f}s\n")
    
    print(f"Serial vs Parallel Times: {serial_time:.2f}s | {parallel_time:.2f}s")

Key Changes Explained:

  1. Redesigned linreg function: It now takes a single tuple argument containing the dummy value, target data, and feature data. This ensures every worker process receives the dataset directly, no global variable dependencies.
  2. Updated iterables: Instead of passing just [1,2,3], we pass tuples that include the data. This makes both serial and parallel calls consistent in how they access the dataset.
  3. Removed global variables: Eliminates cross-process variable access issues entirely.

Fix 2: Use Scoop's Shared Memory (For Large Datasets)

If your 10M-row dataset is too large to copy to each worker, use scoop.shared to store the data once and make it accessible to all workers. This avoids redundant data duplication across processes:

import pandas as pd
import numpy as np
import random
from scoop import futures, shared
import statsmodels.api as sm
from time import time

def linreg(_):
    # Retrieve shared data in the worker process
    y_vals = shared.get('y_vals')
    X_vals = shared.get('X_vals')
    model = sm.OLS(y_vals, X_vals).fit()
    return model

if __name__ == '__main__':
    random.seed(42)
    vals = pd.DataFrame(np.random.normal(loc=3, scale=100, size=(10000000, 5)))
    vals.columns = ['dep', 'ind1', 'ind2', 'ind3', 'ind4']
    y_vals = vals['dep']
    X_vals = vals[['ind1', 'ind2', 'ind3', 'ind4']]
    
    # Store data in Scoop's shared memory pool
    shared.set('y_vals', y_vals)
    shared.set('X_vals', X_vals)
    
    # Serial baseline run
    bt = time()
    model_vals = list(map(linreg, [1,2,3]))
    mval = model_vals[0]
    print("Serial Regression Summary:")
    print(mval.summary())
    serial_time = time() - bt
    print(f"Serial Time: {serial_time:.2f}s\n")
    
    # Parallel run with Scoop
    bt1 = time()
    model_vals_1 = list(futures.map(linreg, [1,2,3]))
    mval_1 = model_vals_1[0]
    print("Parallel Regression Summary:")
    print(mval_1.summary())
    parallel_time = time() - bt1
    print(f"Parallel Time: {parallel_time:.2f}s\n")
    
    print(f"Serial vs Parallel Times: {serial_time:.2f}s | {parallel_time:.2f}s")

This approach is more memory-efficient for large datasets, as the data is serialized once and shared across all workers instead of being copied multiple times.

Critical Reminder:

Always run your script with python -m scoop your_script.py to enable Scoop's parallelism. Without the -m scoop flag, futures.map falls back to Python's built-in serial map, which won't give you the parallel execution you need.

内容的提问来源于stack exchange,提问作者Umbp

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 09:26:07