如何在Python中筛选DataFrame指定日期及下方的行子集
解决方法
先明确执行给定代码后,DataFrame的索引是倒序排列的(最新日期在最上方),且存在重复索引值(如2014-04-07出现三次)。针对你的需求,分两种场景处理:
场景1:获取指定日期及下方所有行
比如指定2014-04-07时,返回包含该日期的所有行及其下方内容。
实现步骤
- 找到目标日期在索引中第一次出现的位置
- 从该位置开始切片到DataFrame末尾
代码示例
import pandas as pd # 原始数据构建 df = pd.DataFrame({'Date' : ['2014-03-27', '2014-03-28', '2014-03-31', '2014-04-01', '2014-04-02', '2014-04-03', '2014-04-04', '2014-04-07', '2014-04-07', '2014-04-07'], 'income': [1849.04, 1857.62, 1872.34, 1885.52, 1890.9, 1888.77, 1865.09, 1845.04, 1235.04, 2323] }) df.set_index('Date',inplace=True) df = df.iloc[::-1] # 指定目标日期 target_date = "2014-04-07" # 定位第一个匹配日期的索引位置 pos = df.index.get_loc(target_date).argmax() # 切片获取结果 result = df.iloc[pos:] print(result)
执行后会返回所有包含2014-04-07的行及下方所有内容,符合需求。
场景2:获取指定日期下方的所有行(不包含指定日期本身)
比如指定2014-04-03时,返回2014-04-02及下方的所有行。
实现步骤
- 找到目标日期第一次出现的位置
- 从该位置的下一个索引开始切片到末尾
代码示例
# 指定目标日期 target_date = "2014-04-03" # 定位第一个匹配日期的索引位置 pos = df.index.get_loc(target_date).argmax() # 切片跳过目标日期行,获取下方内容 result = df.iloc[pos+1:] print(result)
执行后会返回从2014-04-02开始的所有行,满足需求。
补充说明
- 使用
df.index.get_loc(target_date).argmax()是因为存在重复索引,get_loc会返回布尔数组,argmax()可快速定位第一个匹配的位置。 - 如果索引是唯一值,直接用
pos = df.index.get_loc(target_date)获取位置即可,逻辑同样适用。
内容的提问来源于stack exchange,提问作者TRex
相关产品推荐
相关产品推荐

