You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于DataFrame两行数据统计序列匹配次数的技术请求

统计DataFrame中的序列匹配次数

嘿,我来帮你搞定这个序列统计的问题!首先先把你提供的DataFrame数据整理成更清晰的表格,方便我们分析:

TijdnummerschaapcodeModifiercommentstatus
2.97111stilstaanNASTART
5.45711ruiken aan objectNANAPOINT
10.70311stilstaanNASTOP
10.70411lopenNASTART
12.95911lopenNASTOP
12.96011stilstaanNASTART
22.73211ruiken aan objectNANAPOINT
29.38311stilstaanNASTOP
29.38411lopenNASTART
42.56811lopenNASTOP
42.56911ruiken aan objectNANAPOINT
49.20611lopenNASTOP

下面我分两种常见的需求给你实现代码,你可以根据自己的实际目标调整:

需求1:统计通用状态序列START → POINT → STOP的次数

如果你只关心状态的连续匹配,不管具体行为,用这个方法:

import pandas as pd

# 构造你的DataFrame
data = [
    [2.971, 1, 1, "stilstaan", pd.NA, pd.NA, "START"],
    [5.457, 1, 1, "ruiken aan object", pd.NA, pd.NA, "POINT"],
    [10.703, 1, 1, "stilstaan", pd.NA, pd.NA, "STOP"],
    [10.704, 1, 1, "lopen", pd.NA, pd.NA, "START"],
    [12.959, 1, 1, "lopen", pd.NA, pd.NA, "STOP"],
    [12.960, 1, 1, "stilstaan", pd.NA, pd.NA, "START"],
    [22.732, 1, 1, "ruiken aan object", pd.NA, pd.NA, "POINT"],
    [29.383, 1, 1, "stilstaan", pd.NA, pd.NA, "STOP"],
    [29.384, 1, 1, "lopen", pd.NA, pd.NA, "START"],
    [42.568, 1, 1, "lopen", pd.NA, pd.NA, "STOP"],
    [42.569, 1, 1, "ruiken aan object", pd.NA, pd.NA, "POINT"],
    [49.206, 1, 1, "lopen", pd.NA, pd.NA, "STOP"]
]

df = pd.DataFrame(data, columns=["Tijd", "nummer", "schaap", "code", "Modifier", "comment", "status"])

# 定义要匹配的状态序列
target_status_seq = ["START", "POINT", "STOP"]
window_len = len(target_status_seq)
match_count = 0

# 遍历滑动窗口检查匹配
for i in range(len(df) - window_len + 1):
    current_window = df["status"].iloc[i:i+window_len].tolist()
    if current_window == target_status_seq:
        match_count += 1

print(f"匹配到的状态序列次数:{match_count}")

运行后会输出匹配到的状态序列次数:2——对应表格里第1-3行、第6-8行这两组序列。

需求2:统计特定行为+状态的序列(比如stilstaan(START) → ruiken aan object(POINT) → stilstaan(STOP))

如果需要更精确的匹配,同时结合行为(code字段)和状态,用这个方法:

# 定义目标序列,每个元素是(code, status)的元组
target_full_seq = [("stilstaan", "START"), ("ruiken aan object", "POINT"), ("stilstaan", "STOP")]
window_len = len(target_full_seq)
match_count = 0

for i in range(len(df) - window_len + 1):
    current_window = list(zip(df["code"].iloc[i:i+window_len], df["status"].iloc[i:i+window_len]))
    if current_window == target_full_seq:
        match_count += 1

print(f"匹配到的特定行为+状态序列次数:{match_count}")

结果同样是2次,和上面的结果一致,因为你的数据里只有那两组符合这个特定序列。

额外补充:灵活调整序列长度

如果你的目标是其他长度的序列(比如2行的lopen(START) → lopen(STOP)),只需要修改target_full_seq和window_len即可。举个例子:

target_seq = [("lopen", "START"), ("lopen", "STOP")]
window_len = 2
match_count = 0

for i in range(len(df) - window_len + 1):
    current_window = list(zip(df["code"].iloc[i:i+window_len], df["status"].iloc[i:i+window_len]))
    if current_window == target_seq:
        match_count += 1

print(f"lopen START→STOP的序列次数:{match_count}")

运行后会输出lopen START→STOP的序列次数:2,对应表格里第4-5行、第9-10行。


内容的提问来源于stack exchange,提问作者Marien

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:04:33