使用Bash提取并格式化auth.log中的授权数据
解决auth.log授权数据格式化报错问题
核心问题场景分析
格式化报错大多源于正则匹配分组缺失或日志时间解析逻辑缺陷,结合你已实现3小时数据提取、仅格式化环节报错的情况,以下是针对性修复方案:
1. 常见auth.log日志格式示例
Oct 12 14:32:01 server sshd[1234]: Accepted password for john from 192.168.1.100 port 22 ssh2 Oct 12 15:05:22 server sshd[5678]: Failed password for invalid user root from 10.0.0.5 port 22 ssh2 Oct 12 16:10:45 server sshd[9012]: Connection closed by 172.16.3.200 port 22 [preauth]
2. 典型错误与修复方案
常见错误信息(正则分组越界)
Traceback (most recent call last): File "parse_auth.py", line 23, in <module> user = match.group(2) IndexError: no such group
修正后的完整脚本
import re from datetime import datetime, timedelta def parse_auth_log(log_path): # 兼容多场景的正则表达式:覆盖登录成功/失败/连接关闭等日志 log_pattern = re.compile( r'^(\w{3}\s+\d{1,2}\s+\d{2}:\d{2}:\d{2})\s+\S+\s+sshd\[\d+\]:\s+' r'(Accepted password for (\w+)|Failed password for (invalid user )?(\w+)|Connection closed by (\S+))' r'(?: from (\S+))?' ) # 计算3小时前的时间阈值 time_threshold = datetime.now() - timedelta(hours=3) with open(log_path, 'r') as log_file: for line in log_file: line = line.strip() if not line: continue # 解析日志时间(auth.log默认不含年份,需补充当前年份) time_components = line.split(' ', 3)[:3] if len(time_components) < 3: continue try: log_datetime = datetime.strptime(f"{datetime.now().year} {' '.join(time_components)}", "%Y %b %d %H:%M:%S") except ValueError: continue # 过滤超出3小时的日志 if log_datetime < time_threshold: continue # 匹配日志内容 match_result = log_pattern.match(line) if not match_result: continue # 提取各字段,兼容无匹配的情况 date = log_datetime.strftime('%Y-%m-%d') time = log_datetime.strftime('%H:%M:%S') user = match_result.group(3) or match_result.group(5) or 'unknown' ip = match_result.group(7) or match_result.group(6) or 'unknown' # 判定操作类型 if 'Accepted' in line: action = 'Login Success' elif 'Failed' in line: action = 'Login Failed' elif 'Connection closed' in line: action = 'Connection Closed' else: action = 'Unknown' # 格式化输出(可替换为写入文件逻辑) print(f"{date}\t{time}\t{user}\t{action}\t{ip}") if __name__ == '__main__': parse_auth_log('/var/log/auth.log')
3. 关键修复点
- 正则分支覆盖:用多分支正则匹配不同类型的日志条目,避免因某类日志不匹配导致分组缺失
- 字段判空处理:用
or运算符给字段设置默认值,防止无匹配时抛出IndexError - 时间解析健壮性:补充日志缺失的年份,增加异常捕获跳过无效时间格式的日志行
- 无效行过滤:跳过空行和格式异常的日志,避免后续逻辑报错
内容的提问来源于stack exchange,提问作者phantaserrr
相关产品推荐
相关产品推荐

