基于DataFrame两行数据统计序列匹配次数的技术请求
统计DataFrame中的序列匹配次数
嘿,我来帮你搞定这个序列统计的问题!首先先把你提供的DataFrame数据整理成更清晰的表格,方便我们分析:
| Tijd | nummer | schaap | code | Modifier | comment | status |
|---|---|---|---|---|---|---|
| 2.971 | 1 | 1 | stilstaan | NA | START | |
| 5.457 | 1 | 1 | ruiken aan object | NA | NA | POINT |
| 10.703 | 1 | 1 | stilstaan | NA | STOP | |
| 10.704 | 1 | 1 | lopen | NA | START | |
| 12.959 | 1 | 1 | lopen | NA | STOP | |
| 12.960 | 1 | 1 | stilstaan | NA | START | |
| 22.732 | 1 | 1 | ruiken aan object | NA | NA | POINT |
| 29.383 | 1 | 1 | stilstaan | NA | STOP | |
| 29.384 | 1 | 1 | lopen | NA | START | |
| 42.568 | 1 | 1 | lopen | NA | STOP | |
| 42.569 | 1 | 1 | ruiken aan object | NA | NA | POINT |
| 49.206 | 1 | 1 | lopen | NA | STOP |
下面我分两种常见的需求给你实现代码,你可以根据自己的实际目标调整:
需求1:统计通用状态序列START → POINT → STOP的次数
如果你只关心状态的连续匹配,不管具体行为,用这个方法:
import pandas as pd # 构造你的DataFrame data = [ [2.971, 1, 1, "stilstaan", pd.NA, pd.NA, "START"], [5.457, 1, 1, "ruiken aan object", pd.NA, pd.NA, "POINT"], [10.703, 1, 1, "stilstaan", pd.NA, pd.NA, "STOP"], [10.704, 1, 1, "lopen", pd.NA, pd.NA, "START"], [12.959, 1, 1, "lopen", pd.NA, pd.NA, "STOP"], [12.960, 1, 1, "stilstaan", pd.NA, pd.NA, "START"], [22.732, 1, 1, "ruiken aan object", pd.NA, pd.NA, "POINT"], [29.383, 1, 1, "stilstaan", pd.NA, pd.NA, "STOP"], [29.384, 1, 1, "lopen", pd.NA, pd.NA, "START"], [42.568, 1, 1, "lopen", pd.NA, pd.NA, "STOP"], [42.569, 1, 1, "ruiken aan object", pd.NA, pd.NA, "POINT"], [49.206, 1, 1, "lopen", pd.NA, pd.NA, "STOP"] ] df = pd.DataFrame(data, columns=["Tijd", "nummer", "schaap", "code", "Modifier", "comment", "status"]) # 定义要匹配的状态序列 target_status_seq = ["START", "POINT", "STOP"] window_len = len(target_status_seq) match_count = 0 # 遍历滑动窗口检查匹配 for i in range(len(df) - window_len + 1): current_window = df["status"].iloc[i:i+window_len].tolist() if current_window == target_status_seq: match_count += 1 print(f"匹配到的状态序列次数:{match_count}")
运行后会输出匹配到的状态序列次数:2——对应表格里第1-3行、第6-8行这两组序列。
需求2:统计特定行为+状态的序列(比如stilstaan(START) → ruiken aan object(POINT) → stilstaan(STOP))
如果需要更精确的匹配,同时结合行为(code字段)和状态,用这个方法:
# 定义目标序列,每个元素是(code, status)的元组 target_full_seq = [("stilstaan", "START"), ("ruiken aan object", "POINT"), ("stilstaan", "STOP")] window_len = len(target_full_seq) match_count = 0 for i in range(len(df) - window_len + 1): current_window = list(zip(df["code"].iloc[i:i+window_len], df["status"].iloc[i:i+window_len])) if current_window == target_full_seq: match_count += 1 print(f"匹配到的特定行为+状态序列次数:{match_count}")
结果同样是2次,和上面的结果一致,因为你的数据里只有那两组符合这个特定序列。
额外补充:灵活调整序列长度
如果你的目标是其他长度的序列(比如2行的lopen(START) → lopen(STOP)),只需要修改target_full_seq和window_len即可。举个例子:
target_seq = [("lopen", "START"), ("lopen", "STOP")] window_len = 2 match_count = 0 for i in range(len(df) - window_len + 1): current_window = list(zip(df["code"].iloc[i:i+window_len], df["status"].iloc[i:i+window_len])) if current_window == target_seq: match_count += 1 print(f"lopen START→STOP的序列次数:{match_count}")
运行后会输出lopen START→STOP的序列次数:2,对应表格里第4-5行、第9-10行。
内容的提问来源于stack exchange,提问作者Marien
相关产品推荐
相关产品推荐

