Pandas数据集如何删除首次低于0.2后回升至0.2以上的后续行数据
实现方案
核心逻辑
通过两步定位完成数据截断:
- 找到序列中首次出现数值<0.2的行位置
- 从该位置向后查找首次出现数值>0.2的行位置,仅保留该位置之前的所有数据
如果全程未出现低于0.2的数值,或低于0.2后从未回升到0.2以上,则保留全量数据。
代码实现
import pandas as pd # 示例数据构造 data = { "Date and Time": ["2020-06-07 00:00", "2020-06-07 00:01", "2020-06-07 00:02", "2020-06-07 00:03", "2020-06-07 00:04", "2020-06-07 00:05", "2020-06-07 00:06", "2020-06-07 00:07", "2020-06-07 00:08", "2020-06-07 00:09", "2020-06-07 00:10", "2020-06-07 00:11", "2020-06-07 00:12", "2020-06-07 00:13", "2020-06-07 00:14", "2020-06-07 00:15", "2020-06-07 00:16", "2020-06-07 00:17", "2020-06-07 00:18", "2020-06-07 00:19", "2020-06-07 00:20", "2020-06-07 00:21", "2020-06-07 00:22", "2020-06-07 00:23", "2020-06-07 00:24", "2020-06-07 00:25", "2020-06-07 00:26", "2020-06-07 00:27", "2020-06-07 00:28", "2020-06-07 00:29"], "Value": [16.2, 15.1, 13.8, 12.0, 11.9, 12.1, 10.8, 9.8, 8.3, 6.2, 4.3, 4.2, 4.2, 3.3, 1.8, 0.1, 0.05, 0.15, 0.1, 0.18, 0.25, 1, 4, 8, 12.0, 12.0, 12.0, 12.0, 12.0, 12.0], } df = pd.DataFrame(data) # 筛选逻辑 threshold = 0.2 # 判断是否存在低于阈值的记录 has_below = (df['Value'] < threshold).any() if not has_below: df_filtered = df.copy() else: # 首次低于阈值的索引 first_below_idx = (df['Value'] < threshold).idxmax() # 查找低于阈值后首次回升到阈值以上的索引 after_below_above = df.loc[first_below_idx:, 'Value'] > threshold if not after_below_above.any(): df_filtered = df.copy() else: first_above_idx = after_below_above.idxmax() # 截断到回升前的最后一条记录 df_filtered = df.loc[:first_above_idx - 1] # 输出结果验证 print(df_filtered)
验证说明
运行上述代码后输出的df_filtered和你给出的期望结果完全一致,会保留从起始行到索引19(对应时间00:19)的所有记录,自动删除后续回升阶段的数据。
内容的提问来源于stack exchange,提问作者Mel
相关产品推荐
相关产品推荐

