You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python Pandas中保留每日数据DataFrame的每月首尾行?

问题:保留Pandas DataFrame中每月的首尾行

我有一个包含每日数据的Python Pandas DataFrame,结构如下:

print(dataframe1.head(30))
                    Date   Close* month_initial  month day  year
date_final                                                      
2022-09-23  Sep 23, 2022  3693.23           Sep    9.0  23  2022
2022-09-22  Sep 22, 2022  3757.99           Sep    9.0  22  2022
2022-09-21  Sep 21, 2022  3789.93           Sep    9.0  21  2022
2022-09-20  Sep 20, 2022  3855.93           Sep    9.0  20  2022
2022-09-19  Sep 19, 2022  3899.89           Sep    9.0  19  2022
2022-09-16  Sep 16, 2022  3873.33           Sep    9.0  16  2022
2022-09-15  Sep 15, 2022  3901.35           Sep    9.0  15  2022
2022-09-14  Sep 14, 2022  3946.01           Sep    9.0  14  2022
2022-09-13  Sep 13, 2022  3932.69           Sep    9.0  13  2022
2022-09-12  Sep 12, 2022  4110.41           Sep    9.0  12  2022
2022-09-09  Sep 09, 2022  4067.36           Sep    9.0  09  2022
2022-09-08  Sep 08, 2022  4006.18           Sep    9.0  08  2022
2022-09-07  Sep 07, 2022  3979.87           Sep    9.0  07  2022
2022-09-06  Sep 06, 2022  3908.19           Sep    9.0  06  2022
2022-09-02  Sep 02, 2022  3924.26           Sep    9.0  02  2022
2022-09-01  Sep 01, 2022  3966.85           Sep    9.0  01  2022
2022-08-31  Aug 31, 2022  3955.00           Aug    8.0  31  2022
2022-08-30  Aug 30, 2022  3986.16           Aug    8.0  30  2022
2022-08-29  Aug 29, 2022  4030.61           Aug    8.0  29  2022
2022-08-26  Aug 26, 2022  4057.66           Aug    8.0  26  2022
2022-08-25  Aug 25, 2022  4199.12           Aug    8.0  25  2022
2022-08-24  Aug 24, 2022  4140.77           Aug    8.0  24  2022
2022-08-23  Aug 23, 2022  4128.73           Aug    8.0  23  2022
2022-08-22  Aug 22, 2022  4137.99           Aug    8.0  22  2022
2022-08-19  Aug 19, 2022  4228.48           Aug    8.0  19  2022
2022-08-18  Aug 18, 2022  4283.74           Aug    8.0  18  2022
2022-08-17  Aug 17, 2022  4274.04           Aug    8.0  17  2022
2022-08-16  Aug 16, 2022  4305.20           Aug    8.0  16  2022
2022-08-15  Aug 15, 2022  4297.14           Aug    8.0  15  2022
2022-08-12  Aug 12, 2022  4280.15           Aug    8.0  12  2022

我希望保留每个月的第一行和最后一行,尝试了以下代码但未得到预期结果:

import pandas as pd
dataframe1.set_index("date_final", inplace=True)
resultDf = dataframe1.groupby([dataframe1.index.year, dataframe1.index.month]).agg(["first", "last"])
resultDf.index.rename(["year", "month"], inplace=True)
resultDf.reset_index(inplace=True)
resultDf

解决方案

你之前的代码使用agg(["first", "last"])会对每个列单独提取首尾值,生成多层列索引,最终结果不是完整的行记录,而是各列首尾值的组合,这不符合需求。以下两种方法可以正确保留每月的完整首尾行:

方法一:分组取首尾行后合并

通过groupby的head(1)和tail(1)分别获取每个月的第一行和最后一行,再合并结果并排序:

import pandas as pd

# 确保date_final是datetime类型索引(如果尚未设置)
# dataframe1['date_final'] = pd.to_datetime(dataframe1['date_final'])
# dataframe1.set_index('date_final', inplace=True)

# 提取每月首行和末行
month_first = dataframe1.groupby([dataframe1.index.year, dataframe1.index.month]).head(1)
month_last = dataframe1.groupby([dataframe1.index.year, dataframe1.index.month]).tail(1)

# 合并并按日期排序
resultDf = pd.concat([month_first, month_last]).sort_index()

方法二:使用apply返回首尾行

通过groupby.apply直接返回每组的第一行和最后一行,代码更简洁:

resultDf = dataframe1.groupby([dataframe1.index.year, dataframe1.index.month])\
                     .apply(lambda x: x.iloc[[0, -1]])\
                     .reset_index(drop=True)

这两种方法都会返回包含每月完整首尾行的DataFrame,符合你的需求。


内容的提问来源于stack exchange,提问作者adrCoder

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 03:15:50