You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

多项式特征线性回归模型咨询:测试误差与多阶实现

Answers to Your Concrete Compressive Strength Regression Questions


1. Is it reasonable for the polynomial regression model to have higher test error than a simple linear model?

Absolutely, this is a common scenario and totally reasonable—here’s a breakdown of why this might happen:

  • Overfitting: Polynomial models add complexity (especially as n increases). While they can capture non-linear patterns in training data, they often learn noise or one-off quirks specific to the training set instead of the true underlying relationship. This leads to strong training performance but poor generalization to unseen test data.
  • Critical Code Order Mistake: Looking at your code, you’re splitting new_X (polynomial features) before creating it. That means you’re splitting the original features first, then generating polynomial features on the entire dataset (including test data)—this causes data leakage, where test set information contaminates your feature creation. Fix this by first generating new_X from the original X, then splitting into train/test sets.
  • Incorrect Shuffle Parameter: You used shuffle='False' (a string) instead of shuffle=False (a boolean). In Python, non-empty strings evaluate to True, so this would actually shuffle your data against your intent. If your dataset is ordered (e.g., by sample collection date), a misconfigured shuffle could create a train/test split that doesn’t reflect the overall data distribution, skewing error scores.

2. How to implement the test_poly_regression(X_train, y_train, X_test, y_test, n=2) function?

Here’s a clean, reusable function that fixes potential issues (like data leakage) and encapsulates all steps for polynomial regression testing:

import pandas as pd
from sklearn.linear_model import LinearRegression
from sklearn.preprocessing import StandardScaler
from sklearn import metrics

def test_poly_regression(X_train, y_train, X_test, y_test, n=2):
    # Generate polynomial features for train/test sets separately (no leakage!)
    X_train_poly = pd.concat([X_train**(i+1) for i in range(n)], axis=1)
    X_test_poly = pd.concat([X_test**(i+1) for i in range(n)], axis=1)
    
    # Optional: Rename columns for readability (helps with debugging)
    col_names = []
    for degree in range(1, n+1):
        for col in X_train.columns:
            col_names.append(f"{col}^{degree}")
    X_train_poly.columns = col_names
    X_test_poly.columns = col_names
    
    # Scale features (fit only on training data to avoid leakage)
    scaler = StandardScaler()
    X_train_scaled = scaler.fit_transform(X_train_poly)
    X_test_scaled = scaler.transform(X_test_poly)
    
    # Train the model
    ols = LinearRegression()
    ols.fit(X_train_scaled, y_train)
    
    # Make predictions and calculate errors
    train_pred = ols.predict(X_train_scaled)
    test_pred = ols.predict(X_test_scaled)
    
    train_mae = metrics.mean_absolute_error(y_train, train_pred)
    test_mae = metrics.mean_absolute_error(y_test, test_pred)
    
    # Print results
    print(f"Polynomial Degree = {n}")
    print(f"Training MAE: {train_mae:.4f}")
    print(f"Test MAE: {test_mae:.4f}\n")
    
    return train_mae, test_mae
Key Improvements:
  • No Data Leakage: We generate polynomial features and fit the scaler only on the training set. The test set is transformed using training-set logic, which is the correct practice for validating model generalization.
  • Readable Column Names: Instead of generic numbers, columns are named like cement^2 to make it easier to track which features contribute to the model.
  • Flexibility: Call the function with different n values (e.g., test_poly_regression(X_train, y_train, X_test, y_test, n=3)) to compare how polynomial complexity impacts error scores.
  • Return Values: The function returns MAE values so you can store them for further analysis (like plotting error vs. degree to find the optimal n).

内容的提问来源于stack exchange,提问作者Gvasiles

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 07:43:24