Python统计DataFrame中子列表出现频率并获取关联ID
查找DataFrame中连续子序列的出现频率及对应ID
给定如下结构的DataFrame:
ID action 1 A 2 B 3 C 4 D 5 E 6 A 7 B 8 C ...
需求:查找子列表[A,B,C]的出现频率,同时获取每次匹配中各元素对应的ID,输出可采用字典或列表形式,参考输出如下:
出现频率为2: [A:1,B:2,C:3] [A:6,B:7,C:8]
解决方案(Python pandas实现)
import pandas as pd # 替换为你的实际DataFrame df = pd.DataFrame({ 'ID': [1,2,3,4,5,6,7,8], 'action': ['A','B','C','D','E','A','B','C'] }) target = ['A', 'B', 'C'] target_length = len(target) match_records = [] # 遍历所有可能的滑动窗口 for idx in range(len(df) - target_length + 1): # 取出当前窗口的action序列 current_window = df['action'].iloc[idx:idx+target_length].tolist() if current_window == target: # 生成当前匹配的action与ID映射 record = {df['action'].iloc[idx+i]: df['ID'].iloc[idx+i] for i in range(target_length)} match_records.append(record) # 输出结果 print(f"出现频率为{len(match_records)}:") for item in match_records: print([f"{k}:{v}" for k, v in item.items()])
说明
- 代码通过滑动窗口遍历DataFrame的
action列,逐一匹配目标子序列 - 匹配成功时,记录对应
action和ID的映射关系 - 最终输出匹配频率及每次匹配的详细信息,实际使用时只需替换示例中的DataFrame为你的数据即可
内容的提问来源于stack exchange,提问作者whoiswho
相关产品推荐
相关产品推荐

