You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python统计DataFrame中子列表出现频率并获取关联ID

查找DataFrame中连续子序列的出现频率及对应ID

给定如下结构的DataFrame:

ID   action
1      A
2      B
3      C
4      D
5      E
6      A
7      B
8      C
...

需求:查找子列表[A,B,C]的出现频率,同时获取每次匹配中各元素对应的ID,输出可采用字典或列表形式,参考输出如下:

出现频率为2:
[A:1,B:2,C:3]
[A:6,B:7,C:8]

解决方案(Python pandas实现)

import pandas as pd

# 替换为你的实际DataFrame
df = pd.DataFrame({
    'ID': [1,2,3,4,5,6,7,8],
    'action': ['A','B','C','D','E','A','B','C']
})

target = ['A', 'B', 'C']
target_length = len(target)
match_records = []

# 遍历所有可能的滑动窗口
for idx in range(len(df) - target_length + 1):
    # 取出当前窗口的action序列
    current_window = df['action'].iloc[idx:idx+target_length].tolist()
    if current_window == target:
        # 生成当前匹配的action与ID映射
        record = {df['action'].iloc[idx+i]: df['ID'].iloc[idx+i] for i in range(target_length)}
        match_records.append(record)

# 输出结果
print(f"出现频率为{len(match_records)}:")
for item in match_records:
    print([f"{k}:{v}" for k, v in item.items()])

说明

  • 代码通过滑动窗口遍历DataFrame的action列,逐一匹配目标子序列
  • 匹配成功时,记录对应action和ID的映射关系
  • 最终输出匹配频率及每次匹配的详细信息,实际使用时只需替换示例中的DataFrame为你的数据即可

内容的提问来源于stack exchange,提问作者whoiswho

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 20:54:22