Python线程实现滚动R平方回归时报错:需传入可迭代对象而非int
Hey there, let's break down this error and get your parallel rolling regression working smoothly!
Why You're Seeing This Error
The TypeError: ParallelRegression() argument after * must be an iterable, not int pops up because of how you're passing arguments to your thread function. When using threading.Thread, the args parameter expects an iterable (like a tuple) even if you're only passing one value. If you write something like args=(totalThreads) without a trailing comma, Python treats that single integer as the iterable itself—and tries to unpack it into multiple arguments for ParallelRegression, which your function doesn't accept.
Step 1: Fix Thread Argument Passing
First, let's correct how you start your threads. Instead of:
thread = threading.Thread(target=ParallelRegression, args=(totalThreads))
You need to add a trailing comma to make args a proper tuple:
thread = threading.Thread(target=ParallelRegression, args=(totalThreads,))
Step 2: Adjust Your Parallel Logic
Looking at your code snippet, your ParallelRegression function's loop might not be splitting work across threads correctly. Right now, each thread would repeat the same work for range(threadnum) iterations. Instead, we should assign each thread a specific column to process so we get actual parallelization.
Here's a complete, fixed example with proper thread work splitting and rolling regression:
import threading import statsmodels.api as sm import pandas as pd import numpy as np # Sample data (replace with your actual DataFrame) df = pd.DataFrame(np.random.randn(100, 4)) # 100 rows, 4 columns (col0 is your reference column) window_size = 20 # Define your rolling window size total_threads = 3 # Initialize result array: rows = number of rolling windows, columns = number of target columns results = np.zeros((len(df) - window_size + 1, total_threads)) def parallel_regression(col_offset): """Each thread handles one target column (col_offset = 1,2,3 for cols 1,2,3)""" target_col = col_offset for window_idx in range(len(df) - window_size + 1): # Extract the current rolling window window = df.iloc[window_idx:window_idx+window_size, :] # Prepare X (add constant for intercept) and y X = sm.add_constant(window.iloc[:, 0]) y = window.iloc[:, target_col] # Fit OLS and store R-squared model = sm.OLS(y, X).fit() results[window_idx, col_offset - 1] = model.rsquared # Start threads, each assigned to a different target column threads = [] for col in range(1, total_threads + 1): # Pass column index as a tuple (note the trailing comma!) thread = threading.Thread(target=parallel_regression, args=(col,)) threads.append(thread) thread.start() # Wait for all threads to finish before accessing results for thread in threads: thread.join() # Check your computed rolling R-squared values print("Rolling R-squared results:\n", results)
Key Notes to Avoid Issues
- Thread Safety: In this setup, each thread writes to a unique column in the
resultsarray, so there's no race condition (multiple threads writing to the same spot). If you ever have shared resources beyond this, usethreading.Lock()to prevent conflicts. - Rolling Window Logic: Make sure your window loop runs from
0tolen(df) - window_size + 1to cover all valid rolling windows without index errors. - Function Clarity: Renamed the function and variables for readability—clear naming helps catch logic bugs faster.
内容的提问来源于stack exchange,提问作者Catherine Zhang

