Python如何从日志文件中查找符合指定规则的用户profile
按规则匹配日志中符合条件的用户Profile
匹配规则
- 同一profile在同一秒内严格按顺序触发三类操作,顺序为
user logged in(登录)→user changed password(改密)→user logged off(登出) - 上述三类操作必须连续出现,中间不能插入其他任何日志条目
- 日志以
|作为字段分隔符,每行第3个字段为profile名称
实现逻辑
不需要额外安装第三方库,基于已导入的标准库即可完成功能,核心步骤如下:
- 逐行读取日志,按
|分割每行内容,提取时间戳、profile名称、操作类型三个核心字段 - 用固定长度为3的双端队列做滑动窗口,缓存最近读取的3条日志
- 每读入一条新日志填充窗口后,校验窗口内3条日志是否满足:profile完全一致、时间戳完全一致、操作顺序和目标序列完全匹配
- 匹配成功的profile存入集合自动去重,最终统一输出
参考代码
import collections import time # 目标操作序列,顺序不能改 TARGET_OPS = ("user logged in", "user changed password", "user logged off") # 存符合条件的profile,自动去重 valid_profiles = set() # 长度为3的滑动窗口,存最近3条解析后的日志:(时间戳, profile, 操作) window = collections.deque(maxlen=3) # 第一次遍历日志做规则匹配 with open("logfiles.txt", "r", encoding="utf-8") as f: for raw_line in f: line = raw_line.strip() if not line: continue # 分割字段,自动去除字段前后多余空格 fields = [item.strip() for item in line.split("|")] # 字段索引按实际日志结构调整:索引从0开始计数,第三字段为profile对应索引2 # 示例日志行:2024-06-01 14:25:30 | 10.0.0.1 | asdf | user logged in log_time = fields[0] profile = fields[2] op = fields[-1] window.append( (log_time, profile, op) ) # 窗口凑够3条就做校验 if len(window) == 3: t1, p1, op1 = window[0] t2, p2, op2 = window[1] t3, p3, op3 = window[2] if p1 == p2 == p3 and t1 == t2 == t3 and (op1, op2, op3) == TARGET_OPS: valid_profiles.add(p1) # 保留原本实现的日志行频次统计功能 line_counter = collections.Counter() with open("logfiles.txt", "r", encoding="utf-8") as f: for raw_line in f: line = raw_line.strip() if line: line_counter[line] += 1 print("日志行出现频次统计:") for content, count in line_counter.items(): print(f"出现{count}次:{content}") # 输出符合条件的profile print("\n符合规则的profile列表:") print("、".join(valid_profiles)) # 保留原本的暂停逻辑 time.sleep(10)
调整说明
- 如果日志里时间戳、操作内容的字段位置和代码默认值不一致,直接修改
fields[x]的索引值即可 - 输出格式和期望的
asdf、klij、plnb、zzad格式完全一致 - 原有已实现的频次统计、程序暂停10秒的逻辑全部保留,不需要删除原有代码
内容的提问来源于stack exchange,提问作者anub13
相关产品推荐
相关产品推荐

