You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Pandas groupby生成的Series转为回归用宽表DataFrame?

Reshaping MultiIndex Series to Wide Format for Sales Data

Hey there! I see you're trying to reshape your grouped sales data into a wide table where shop_id is the index and date_block_num are the columns — let's fix that error and get you the right format.

Why Your Original Code Failed

The error AttributeError: 'Series' object has no attribute 'date_block_num' happens because ts is a MultiIndex Series, not a DataFrame. The date_block_num and shop_id are part of the series's index, not top-level attributes. So you can't access them directly like ts.date_block_num.

Solution 1: Use unstack() (Simplest Method)

Since you already have the grouped Series ts, the easiest way to reshape it into your desired wide format is using pandas' unstack() method. This will pivot one level of the MultiIndex into columns:

# Unstack the date_block_num level to turn it into columns
wide_df = ts.unstack(level='date_block_num')

# Optional: Fill missing values (shops with no sales in a date block) with 0
wide_df = wide_df.fillna(0)

This will give you exactly the structure you want:

date_block_num     0       1       2       ...      33
shop_id
0                5578.0    ...     ...     ...    1972.0
1                2947.0    ...     ...     ...    1263.0
2                1146.0    ...     ...     ...    2316.0
...               ...      ...     ...     ...      ...

Solution 2: Pivot Directly from Raw Data

If you want to skip the intermediate groupby step, you can use pivot_table directly on the original sales DataFrame to get the wide table in one go:

wide_df = sales.pivot_table(
    index='shop_id',
    columns='date_block_num',
    values='item_cnt_day',
    aggfunc='sum',
    fill_value=0  # Optional: fills missing sales entries with 0
)

This achieves the same result without first creating the ts Series — perfect if you want a more streamlined workflow.

Quick Notes

  • The unstack() method defaults to pivoting the last level of the MultiIndex. Specifying level='date_block_num' makes your code more readable, especially if you ever adjust the index order later.
  • Filling missing values with 0 is helpful for regression analysis, as most modeling libraries can't handle NaN values.

内容的提问来源于stack exchange,提问作者Revolucion for Monica

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 17:32:35