Python如何高效从日志字符串中提取末尾数组和可选type字段
日志字段提取解决方案
针对两类日志的统一提取需求,直接用正则表达式做一次匹配即可,不需要多套拆分逻辑,执行效率更高,也能完美适配空格分布不均、数组带空格的场景。
实现代码
import re # 预编译正则,适配可选type字段 + 数组的场景,一次编译可重复用于所有行匹配 pattern = re.compile(r'(?:for type "([^"]+)")?.*are (\[.*?\])') with open('abc.log', 'r') as f: for line in f: # 先过滤不含Binning的行,减少正则匹配开销 if 'Binning' not in line: continue res = pattern.search(line) if not res: continue type_val, array_str = res.groups() # 按需处理结果,type不存在时type_val为None print(f"type字段: {type_val if type_val else '未携带'}") print(f"数组内容: {array_str}") # 若需要把数组转成Python列表,可加以下代码: # import ast # array_list = ast.literal_eval(array_str)
效果说明
- 对Type1格式日志,输出
type字段: 未携带,数组内容为[6, 3, 0, 1, 0, 1] - 对Type2格式日志,输出
type字段: VTC,数组内容为[6, 3, 0, 1, 0, 1] - 数组会作为完整字符串返回,不受内部空格影响,无需额外拼接处理
内容的提问来源于stack exchange,提问作者Souradip Roy
相关产品推荐
相关产品推荐

