You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为按年份分组后的每个影片条目填充对应年份值?

How to Fill Corresponding Year Values for Each Entry After Grouping by Year and Name?

Problem Context

I’ve grouped my film cast data by year and name using cast.groupby(['year','name']).size(), and the output looks like this:

year name
1894 Blanche Bayliss 1
Chauncey Depew 1
William Courtenay 1
1900 Orrie Perry 1
Reg Perry 1
1905 Armand Dranem 1
1906 Battling Nelson 1
E.J. Tait 1
... (remaining entries)

The issue is that the year value only shows up at the start of each group, and I need every row to explicitly display its matching year.

Solutions

Method 1: Convert to a Regular DataFrame with reset_index()

This is the most straightforward fix. The output you’re seeing uses a MultiIndex (hierarchical index), which omits repeated values for readability. Using reset_index() will turn those index levels into regular columns, so every row gets its own year value:

# Generate grouped counts and convert to a full DataFrame
result_df = cast.groupby(['year', 'name']).size().reset_index(name='count')
  • reset_index(): Takes the year and name index levels and converts them into proper columns in the DataFrame.
  • name='count': Renames the default size column to a descriptive label (you can use any name here, like appearances).

After running this, you’ll get a standard DataFrame where every row has a complete year, name, and count entry—no more omitted year values.

Method 2: Fill Missing Year Values in the Grouped Structure

If you want to retain the grouped logic but fill in the hidden year values, you can convert the grouped Series to a DataFrame and use forward filling to propagate the year down each group:

# Convert the grouped Series to a DataFrame
grouped_df = cast.groupby(['year', 'name']).size().to_frame(name='count')
# Reset index to make year/name columns (year will show NaN for subsequent group rows)
grouped_df = grouped_df.reset_index()
# Forward fill the year column to replace NaNs with the group's year
grouped_df['year'] = grouped_df['year'].ffill()

This works because resetting the index turns those "hidden" year values into NaN entries. The ffill() method then replaces each NaN with the last valid year value, which is exactly the year for that group.

内容的提问来源于stack exchange,提问作者Tom Cheng

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:42:34