如何按season列自定义顺序重排pandas DataFrame的行数据
问题原因
df.set_index('season')操作默认不会修改原DataFrame,你没有将操作结果赋值给新变量或添加inplace=True参数,所以后续reindex操作还是基于DataFrame原有的默认整数索引执行,找不到2008-2019这类索引值,就会返回全NaN的结果。- 你原来的
reindex逻辑也有小问题:set_index之后season会变成索引,不会再作为普通列存在,和你预期的输出结构不符。
正确实现方案
方案1:修复reindex写法(适配你原有的操作逻辑)
先把set_index的结果赋值,reindex之后再把season从索引恢复成普通列即可:
# 先将season设为索引并赋值给新对象 df_season_idx = df.set_index('season') # 按自定义顺序重排索引,再把season恢复为普通列 custom_order = [2008, 2009, 2010, 2011, 2012, 2013, 2014, 2015, 2016, 2017, 2018, 2019] df_sorted = df_season_idx.reindex(custom_order).reset_index()
方案2:使用分类变量排序(更简洁,无需修改索引)
将season列设置为按你指定顺序排序的分类类型,再直接排序即可,后续做赛季相关的分组、筛选操作也会默认遵循该自定义顺序:
import pandas as pd custom_order = [2008, 2009, 2010, 2011, 2012, 2013, 2014, 2015, 2016, 2017, 2018, 2019] # 将season转换为按自定义顺序排序的分类类型 df['season'] = pd.Categorical(df['season'], categories=custom_order, ordered=True) # 按season列排序,相同赛季的行会保留原有的相对顺序 df_sorted = df.sort_values('season').reset_index(drop=True)
内容的提问来源于stack exchange,提问作者Faheeem Sajjad
相关产品推荐
相关产品推荐

