R中deparse(substitute())的Python等效实现及模型预测列添加方法
Here's a clean, Pythonic solution that mirrors your R functionality—allowing you to add model predictions as a new column named after the model (or a custom name) to a pandas DataFrame. We'll use scikit-learn for modeling (the standard for Python ML) and the inspect module to automatically retrieve the model's variable name, just like deparse(substitute(model)) in R.
Step 1: Setup and Sample Data
First, let's create a sample DataFrame similar to your R example:
import pandas as pd import numpy as np from sklearn.linear_model import LinearRegression, Ridge import inspect # Reproducible sample data np.random.seed(42) df = pd.DataFrame({ 'a': np.arange(1, 6) + np.random.normal(0, 1, 5), 'b': np.arange(6, 11) + np.random.normal(0, 1, 5), 'y': np.arange(11, 16) + np.random.normal(0, 1, 5) }) # Define features and target X = df[['a', 'b']] y = df['y']
Step 2: The Reusable Function
This function handles generating predictions, naming the column after the model (or using a custom name), and adding it to your DataFrame:
def add_predictions(df, model, col_name=None, X=None): # Auto-infer model name if no column name is provided if col_name is None: # Get local variables from the caller's scope caller_locals = inspect.currentframe().f_back.f_locals # Find which variable refers to the model object for name, obj in caller_locals.items(): if obj is model: col_name = name break # Fallback if model name can't be inferred if col_name is None: col_name = "predictions" # Determine which features to use for prediction if X is not None: # Use explicitly provided features pred_features = X elif hasattr(model, 'feature_names_in_'): # Use feature names stored in scikit-learn models (when fit on pandas DataFrames) pred_features = df[model.feature_names_in_] else: raise ValueError( "Could not infer feature columns. Please pass the 'X' parameter explicitly, " "or use a model that stores feature names (e.g., scikit-learn models fit with pandas DataFrames)." ) # Generate and add predictions to the DataFrame df[col_name] = model.predict(pred_features) return df
Step 3: Usage Examples
Let's fit two models and add their predictions to the DataFrame:
Example 1: Linear Regression
# Fit a linear regression model ols = LinearRegression() ols.fit(X, y) # Add predictions (column will be named 'ols') df = add_predictions(df, ols)
Example 2: Ridge Regression
# Fit a ridge regression model ridge_model = Ridge(alpha=1.0) ridge_model.fit(X, y) # Add predictions with a custom column name df = add_predictions(df, ridge_model, col_name="ridge_predictions")
Example 3: Explicitly Pass Features (for custom models)
If you're using a model that doesn't store feature names, pass the features directly:
# Hypothetical custom model class CustomModel: def predict(self, X): return X.sum(axis=1) custom_model = CustomModel() df = add_predictions(df, custom_model, col_name="custom_preds", X=X)
Key Notes
- Automatic Naming: The function uses
inspectto find the model's variable name in the caller's scope, just like R'sdeparse(substitute(model)). This works for most common use cases (models stored as local variables). - Scikit-Learn Integration: Scikit-learn models fit on pandas DataFrames store feature names in
feature_names_in_, so the function can automatically select the right columns for prediction. - Flexibility: You can always override the column name with
col_nameor pass explicit features withXfor custom models.
After running these examples, your DataFrame will have columns like ols, ridge_predictions, and custom_preds with the respective model outputs.
内容的提问来源于stack exchange,提问作者Gaurav Bansal

