You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas数据集如何删除首次低于0.2后回升至0.2以上的后续行数据

实现方案

核心逻辑

通过两步定位完成数据截断:

  • 找到序列中首次出现数值<0.2的行位置
  • 从该位置向后查找首次出现数值>0.2的行位置,仅保留该位置之前的所有数据
    如果全程未出现低于0.2的数值,或低于0.2后从未回升到0.2以上,则保留全量数据。

代码实现

import pandas as pd

# 示例数据构造
data = {
  "Date and Time": ["2020-06-07 00:00", "2020-06-07 00:01", "2020-06-07 00:02", "2020-06-07 00:03", "2020-06-07 00:04", "2020-06-07 00:05", "2020-06-07 00:06", "2020-06-07 00:07", "2020-06-07 00:08", "2020-06-07 00:09", "2020-06-07 00:10", "2020-06-07 00:11", "2020-06-07 00:12", "2020-06-07 00:13", "2020-06-07 00:14", "2020-06-07 00:15", "2020-06-07 00:16", "2020-06-07 00:17", "2020-06-07 00:18", "2020-06-07 00:19", "2020-06-07 00:20", "2020-06-07 00:21", "2020-06-07 00:22", "2020-06-07 00:23", "2020-06-07 00:24", "2020-06-07 00:25", "2020-06-07 00:26", "2020-06-07 00:27", "2020-06-07 00:28", "2020-06-07 00:29"],
  "Value": [16.2, 15.1, 13.8, 12.0, 11.9, 12.1, 10.8, 9.8, 8.3, 6.2, 4.3, 4.2, 4.2, 3.3, 1.8, 0.1, 0.05, 0.15, 0.1, 0.18, 0.25, 1, 4, 8, 12.0, 12.0, 12.0, 12.0, 12.0, 12.0],
}
df = pd.DataFrame(data)

# 筛选逻辑
threshold = 0.2
# 判断是否存在低于阈值的记录
has_below = (df['Value'] < threshold).any()

if not has_below:
    df_filtered = df.copy()
else:
    # 首次低于阈值的索引
    first_below_idx = (df['Value'] < threshold).idxmax()
    # 查找低于阈值后首次回升到阈值以上的索引
    after_below_above = df.loc[first_below_idx:, 'Value'] > threshold
    if not after_below_above.any():
        df_filtered = df.copy()
    else:
        first_above_idx = after_below_above.idxmax()
        # 截断到回升前的最后一条记录
        df_filtered = df.loc[:first_above_idx - 1]

# 输出结果验证
print(df_filtered)

验证说明

运行上述代码后输出的df_filtered和你给出的期望结果完全一致,会保留从起始行到索引19(对应时间00:19)的所有记录,自动删除后续回升阶段的数据。

内容的提问来源于stack exchange,提问作者Mel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 21:39:04