如何遍历字符串列表提取指定正则匹配内容(日期、时间、主机名)
日志字符串提取指定字段的问题解决
需要遍历日志字符串列表,提取日期、时间、主机名(shost字段)到独立列表,用于后续构建DataFrame。已经写了正则表达式,但迭代处理时触发TypeError: expected string or bytes-like object错误,不知道如何修复。
示例日志列表
logs = [ "Feb 24 2023 20:37:42 somedomain.com Label=Risk_Level cs5=Low cs2Label=Policy act=Deny shost=VD-DONALD dntdom=disney\\ ", "Feb 24 2023 20:46:10 somedomain.com Label=Risk_Level cs5=High cs2Label=Policy act=Terminate shost=VD-MICKEY dntdom=disney\\ ", ]
原正则代码的问题
你之前写的正则直接传入了整个logs列表,而re.findall()需要接收字符串类型参数,这是报错的核心原因。错误的正则代码如下:
date = ''.join(re.findall('\w{3}\s\d{2}\s\d{4}', logs)) timestamp = ''.join(re.findall('\d{2}:\d{2}:\d{2}', logs)[0]) target_host = ''.join(re.findall('shost=([^\s]+)', logs))
你尝试的错误迭代代码
尝试1(错误遍历每个字符而非匹配字段)
date_list = [] for log in logs: for date in log: date_list.append(date) print(date_list)
尝试2(逻辑错误,未正确使用正则匹配)
for log in logs: for log_item in log: if date in log_item: print(date)
正确解决方案
正确的做法是遍历列表中的每个日志字符串,对单个字符串应用正则提取,再将结果存入对应列表:
import re logs = [ "Feb 24 2023 20:37:42 somedomain.com Label=Risk_Level cs5=Low cs2Label=Policy act=Deny shost=VD-DONALD dntdom=disney\\ ", "Feb 24 2023 20:46:10 somedomain.com Label=Risk_Level cs5=High cs2Label=Policy act=Terminate shost=VD-MICKEY dntdom=disney\\ ", ] dates_list = [] timestamp_list = [] host_list = [] # 遍历每个日志条目 for log in logs: # 提取日期 date_match = re.search(r'\w{3}\s\d{2}\s\d{4}', log) if date_match: dates_list.append(date_match.group()) # 提取时间戳 time_match = re.search(r'\d{2}:\d{2}:\d{2}', log) if time_match: timestamp_list.append(time_match.group()) # 提取shost host_match = re.search(r'shost=([^\s]+)', log) if host_match: host_list.append(host_match.group(1)) print(dates_list) print(timestamp_list) print(host_list)
预期结果
运行上述代码后将得到:
dates_list = ['Feb 24 2023', 'Feb 24 2023'] timestamp_list = ['20:37:42', '20:46:10'] host_list = ['VD-DONALD','VD-MICKEY']
内容的提问来源于stack exchange,提问作者CUI
相关产品推荐
相关产品推荐

