如何编写正则从stderr日志提取Hadoop MR计数器数据并匹配制表符行对?
Hadoop MR计数器提取与缩进行匹配正则解决方案
问题1:提取stderr日志中的Hadoop MR计数器数据
日志中计数器区域从Counters: \d+开始,到下一条非缩进日志结束。可以分两步提取:
步骤1:捕获完整计数器块
先把整个计数器区域从日志中分离出来,避免无关内容干扰:
import re # 匹配计数器块的正则 counter_block_pattern = re.compile(r'Counters: \d+\n(.*?)(?=\n\d{2}/\d{2}/\d{2}|\Z)', re.DOTALL) # 提取计数器块内容 counter_block = counter_block_pattern.search(your_log_text).group(1)
(?=\n\d{2}/\d{2}/\d{2}|\Z)是正向预查,确保只匹配到下一条日期格式的日志行或文本结尾。
步骤2:提取类别与对应计数器项
针对计数器块,匹配一级缩进(4个空格,对应你说的\t)的类别行,以及二级缩进(8个空格,对应\t\t)的所有项:
# 匹配类别和对应项的正则(日志用空格缩进) category_items_pattern = re.compile(r'^\s{4}(\S.*?)\n(^\s{8}.*?(?:\n\s{8}.*?)*)', re.MULTILINE | re.DOTALL) # 获取所有类别-项对 counter_pairs = category_items_pattern.findall(counter_block)
如果日志确实用制表符\t缩进,替换正则中的空格为制表符即可:
category_items_pattern = re.compile(r'^\t(\S.*?)\n(^\t\t.*?(?:\n\t\t.*?)*)', re.MULTILINE | re.DOTALL)
问题2:匹配\t开头行与对应\t\t开头行的正则修正
你之前的正则存在三个问题:
- 用
\n\t\w+作为结束边界,会漏掉最后一个类别(无后续\t行) .*?无限制匹配,可能捕获无关内容[a-zA-Z\s]+无法匹配带连字符的类别名(如Map-Reduce Framework)
修正后的正则
# 制表符缩进版本 corrected_pattern = re.compile(r'(\t[\w\s-]+)\n((?:\t\t.*?\n)*)', re.MULTILINE | re.DOTALL) # 空格缩进版本(对应日志实际格式) corrected_pattern = re.compile(r'(\s{4}[\w\s-]+)\n((?:\s{8}.*?\n)*)', re.MULTILINE | re.DOTALL)
[\w\s-]+覆盖所有类别名的字符(字母、空格、连字符)(?:\t\t.*?\n)*只匹配\t\t开头的行,非捕获组避免多余分组- 利用
re.MULTILINE确保每行独立匹配,re.DOTALL允许跨行匹配
完整运行示例
import re # 替换为你的日志文本 log_text = """ 23/01/16 14:26:13 INFO mortbay.log: Conf is not init. 23/01/16 14:26:14 INFO mapreduce.Job: Counters: 246 File System Counters FILE: Number of bytes read=104971581500 FILE: Number of bytes written=287906526786 FILE: Number of read operations=0 FILE: Number of large read operations=0 FILE: Number of write operations=0 HDFS: Number of bytes read=758223470025 HDFS: Number of bytes written=97994290043 HDFS: Number of read operations=24275 HDFS: Number of large read operations=0 HDFS: Number of write operations=2000 VIEWFS: Number of bytes read=0 VIEWFS: Number of bytes written=0 VIEWFS: Number of read operations=0 VIEWFS: Number of large read operations=0 VIEWFS: Number of write operations=0 Job Counters Killed map tasks=3 Killed reduce tasks=2 Launched map tasks=6427 Launched reduce tasks=1002 Other local map tasks=33 Data-local map tasks=3746 Rack-local map tasks=2648 Total time spent by all maps in occupied slots (ms)=358061940 Total time spent by all reduces in occupied slots (ms)=858021936 Total time spent by all map tasks (ms)=119353980 Total time spent by all reduce tasks (ms)=107252742 Total vcore-milliseconds taken by all map tasks=119353980 Total vcore-milliseconds taken by all reduce tasks=107252742 Total megabyte-milliseconds taken by all map tasks=305546188800 Total megabyte-milliseconds taken by all reduce tasks=878614462464 Map-Reduce Framework Map input records=30951997 Map output records=30951997 Shuffled Maps =6425000 Failed Shuffles=46 Merged Map outputs=6425000 File Input Format Counters Bytes Read=0 File Output Format Counters Bytes Written=0 23/01/16 14:26:14 INFO streaming.StreamJob: Output directory: + [[ 0 -ne 0 ]] + exit 0 """ # 提取计数器块 counter_block = re.search(r'Counters: \d+\n(.*?)(?=\n\d{2}/\d{2}/\d{2}|\Z)', log_text, re.DOTALL).group(1) # 提取类别与项 counter_pairs = re.findall(r'^\s{4}(\S.*?)\n(^\s{8}.*?(?:\n\s{8}.*?)*)', counter_block, re.MULTILINE | re.DOTALL) # 输出结果 for category, items in counter_pairs: print(f"【类别】: {category.strip()}") print("【计数器项】:") for item in items.strip().split('\n'): print(f" {item.strip()}") print("-" * 50)
内容的提问来源于stack exchange,提问作者lxneng
相关产品推荐
相关产品推荐

