You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中对日期与ID分组为列?多索引数据可用于Scikit Learn线性回归吗?

Answer

First off, let's tackle your two main questions: using your MultiIndex DataFrame with Scikit-learn, and restructuring it into the desired row-wise format.

Can your MultiIndex structure be used for Scikit-learn Linear Regression?

Yes, but it’s not the most straightforward or compatible option. Scikit-learn expects input data (X) to be a 2D array-like structure (e.g., a flat-column DataFrame) where each column is a single feature. While some Scikit-learn functions can handle MultiIndex columns, many utilities (like feature selection, preprocessing pipelines, or model interpretation) work better with flat, human-readable column names.

To make your MultiIndex DataFrame ML-ready, flatten the column names by joining the two levels:

# Flatten MultiIndex columns
pivot02_flat = pivot02.copy()
pivot02_flat.columns = ['_'.join(col) for col in pivot02_flat.columns]

# Example: Separate features and target (if predicting Number Of Awards)
target_cols = [col for col in pivot02_flat.columns if col.startswith('Number Of Awards')]
feature_cols = [col for col in pivot02_flat.columns if not col.startswith('Number Of Awards')]

X = pivot02_flat[feature_cols]
y = pivot02_flat[target_cols]

# Fit a linear regression model
from sklearn.linear_model import LinearRegression
model = LinearRegression()
model.fit(X, y)

Restructuring to your desired row-wise format

You want each row to contain GeographyCode, Number of Awards, Secondary School Stage, SCQF Level, and DateCode. This is a "tidy" long data format, which is easy to achieve by stacking the DateCode level from your columns into the index, then resetting the index to turn those levels into columns:

# Stack the DateCode column level into the row index
long_format_df = pivot02.stack(level='DateCode').reset_index()

# Rename columns to match your desired structure
long_format_df.columns = ['GeographyCode', 'DateCode', 'Number Of Awards', 'SCQF Level', 'Secondary School Stage']

Let me break this down:

  • stack(level='DateCode') moves the DateCode values from the column MultiIndex into the row index, creating a new MultiIndex of (GeographyCode, DateCode). The columns now become your three core metrics.
  • reset_index() converts both index levels (GeographyCode and DateCode) into regular columns, giving you exactly the row structure you described.

ML with the long format

If you prefer to use the long format for your model (e.g., predicting Number Of Awards based on time and other metrics), you’ll need to encode the DateCode (since it’s a string). For time-based data, converting it to an ordinal value (e.g., 0 for 2002/2003, 1 for 2003/2004) is often useful:

# Encode DateCode as ordinal values to capture time sequence
date_order = sorted(long_format_df['DateCode'].unique())
long_format_df['DateOrdinal'] = long_format_df['DateCode'].map({date: i for i, date in enumerate(date_order)})

# Prepare features and target
X = long_format_df[['DateOrdinal', 'SCQF Level', 'Secondary School Stage']]
y = long_format_df['Number Of Awards']

# Fit model
model = LinearRegression()
model.fit(X, y)

This approach treats each GeographyCode + DateCode combination as a separate sample, which is ideal if you want to model how metrics change over time per geography.

内容的提问来源于stack exchange,提问作者user3617574

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:16:49