You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于DataFrame某一列的连续值匹配特定序列获取对应行?

Pandas匹配DataFrame指定列连续序列的实现方法

需求核心:提取指定列取值连续匹配目标序列的对应行,示例中为匹配View列连续a→p→p的序列,对应行索引为row2、row3、row4。

实现步骤

  • 构造示例DataFrame
import pandas as pd

# 构造示例数据
data = {
    "column1": [1, 3, 5, 7, 9],
    "column2": [2, 4, 6, 8, 10],
    "View": ["c", "a", "p", "p", "n"]
}
df = pd.DataFrame(data, index=["row1", "row2", "row3", "row4", "row5"])
  • 核心匹配逻辑(通用场景)
    通过拼接指定列取值为字符串,检索目标序列的起始位置,再截取对应行即可,兼容多组匹配场景:
# 定义目标匹配序列
target_seq = "app"
seq_len = len(target_seq)

# 拼接View列所有取值为完整字符串
view_total = "".join(df["View"])
# 检索所有匹配序列的起始行下标
match_start_indexes = [
    idx for idx in range(len(view_total) - seq_len + 1)
    if view_total[idx:idx+seq_len] == target_seq
]

# 提取匹配行存入列表,可根据需要调整输出格式
result = []
for start_idx in match_start_indexes:
    # 转为字典格式存储行数据,若只需值可替换为.values.tolist()
    match_rows = df.iloc[start_idx: start_idx + seq_len].to_dict("records")
    result.extend(match_rows)
  • 大数据量优化方案
    如果处理百万行以上的大表,避免拼接长字符串占用过多内存,可使用滑动窗口逐位匹配:
target_list = list("app")
seq_len = len(target_list)
# 生成匹配掩码
match_mask = pd.Series([True] * len(df))
for offset in range(seq_len):
    match_mask &= df["View"].shift(-offset) == target_list[offset]

# 提取所有匹配的连续行
result = []
for start_idx in df[match_mask].index:
    start_pos = df.index.get_loc(start_idx)
    match_rows = df.iloc[start_pos: start_pos + seq_len].to_dict("records")
    result.extend(match_rows)

输出结果验证

示例中最终result输出如下,对应row2、row3、row4三行数据:

[
    {'column1': 3, 'column2': 4, 'View': 'a'},
    {'column1': 5, 'column2': 6, 'View': 'p'},
    {'column1': 7, 'column2': 8, 'View': 'p'}
]

内容的提问来源于stack exchange,提问作者Giulio Augello

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 10:09:02