如何基于不同DataFrame的y、x1、x2实现多元线性回归?
Hey there! Let's work through your multi-linear regression problem where your dependent variable (y from df_A) and independent variables (x1 from df_B, x2 from df_C) live in separate DataFrames. I totally get why this feels tricky—most guides assume all variables are in one dataframe, but we can sort this out easily.
First: Make Sure Your Indexes Are Aligned
Before doing anything else, you need to confirm that all three DataFrames share the exact same index (since you want to pair rows by index). Run this quick check:
import pandas as pd # Verify indexes match across all three DataFrames print("df_A and df_B indexes match:", df_A.index.equals(df_B.index)) print("df_A and df_C indexes match:", df_A.index.equals(df_C.index))
If either check returns False, align them first using reindex or join to avoid missing values later:
# Align df_B and df_C to df_A's index (drop rows that don't match) df_B = df_B.reindex(df_A.index).dropna() df_C = df_C.reindex(df_A.index).dropna() df_A = df_A.reindex(df_B.index) # Sync df_A to the now-aligned B/C indexes
Option 1: Combine DataFrames First (Recommended)
The simplest fix is to merge all three into a single DataFrame, which lets you use standard regression workflows. Here's how to do it cleanly:
# Rename columns to avoid confusion, then concatenate along columns combined_df = pd.concat( [ df_A.rename(columns={df_A.columns[0]: "y"}), # Rename df_A's column to 'y' df_B.rename(columns={df_B.columns[0]: "x1"}), # Rename df_B's column to 'x1' df_C.rename(columns={df_C.columns[0]: "x2"}) # Rename df_C's column to 'x2' ], axis=1 ) # Drop any rows with missing values (from misaligned indexes) combined_df = combined_df.dropna()
Now you can run your multi-linear regression like normal. Using statsmodels for detailed output:
import statsmodels.api as sm # Add a constant term for the intercept X = sm.add_constant(combined_df[["x1", "x2"]]) y = combined_df["y"] # Fit the model and print results model = sm.OLS(y, X).fit() print(model.summary())
Or using scikit-learn if you prefer that API:
from sklearn.linear_model import LinearRegression X = combined_df[["x1", "x2"]] y = combined_df["y"] model = LinearRegression() model.fit(X, y) print(f"Intercept: {model.intercept_}") print(f"x1 Coefficient: {model.coef_[0]}") print(f"x2 Coefficient: {model.coef_[1]}")
Option 2: Loop Through Groups (If You Need Per-Index Regression)
If your index represents groups (e.g., repeated index values for different subsets of data) and you want to run a regression for each group, here's how to do that after combining the DataFrames:
# Group the combined DataFrame by its index grouped_data = combined_df.groupby(combined_df.index) # Loop through each group and run regression for group_id, group in grouped_data: print(f"\n--- Regression Results for Group: {group_id} ---") # Skip groups with too few samples (need at least 3 for 2 variables + intercept) if len(group) < 3: print(f"Group {group_id} has insufficient samples ({len(group)} rows) — skipping.") continue X = sm.add_constant(group[["x1", "x2"]]) y = group["y"] model = sm.OLS(y, X).fit() print(model.summary())
Why Your Original Loop Logic Might Have Failed
Chances are you were trying to pull single rows (one y, one x1, one x2) per index and fit a model—this doesn't work because linear regression needs multiple samples to estimate coefficients. The fixes above either use all your data at once or group samples into meaningful subsets to fit valid models.
内容的提问来源于stack exchange,提问作者ian_chan

