如何用Python读取日志文件并移除行中存在的时间戳?
移除日志时间戳并提取统计数据为CSV的解决方案
核心思路
不用硬截取固定长度字符,而是用正则表达式精准匹配时间戳格式,只移除行开头符合DD/MM/YY HH:MM:SS:SSS 的内容,这样不管行有没有时间戳,都不会破坏原有内容。
修改后的完整代码
import re from argparse import ArgumentParser import csv if __name__ == '__main__': # 解析命令行参数:日志文件路径 parser = ArgumentParser() parser.add_argument("--logDestination", dest="logDest", help="Provide the directory of the log file") args = parser.parse_args() log_path = str(args.logDest).strip() # 匹配行开头的时间戳格式(注意末尾的空格) timestamp_regex = r'^\d{2}/\d{2}/\d{2} \d{2}:\d{2}:\d{2}:\d{3} ' # 配置CSV输出 csv_output_path = 'log_statistics.csv' csv_headers = ['Procedure', 'Count', 'Retry', 'Success', 'Failure'] with open(log_path, 'r') as log_file, open(csv_output_path, 'w', newline='') as csv_file: csv_writer = csv.writer(csv_file) csv_writer.writerow(csv_headers) is_in_stats_section = False for raw_line in log_file: # 移除时间戳并清理首尾空格 cleaned_line = re.sub(timestamp_regex, '', raw_line.strip()) if not cleaned_line: continue # 标记进入统计数据区域 if cleaned_line == 'EMM_PROCEDURE:': is_in_stats_section = True continue # 处理统计行数据 if is_in_stats_section: # 跳过统计表头行 if cleaned_line.startswith('[Procedure]'): continue # 按任意数量空格分割字段(适配日志中不一致的空格数) stats_fields = re.split(r'\s+', cleaned_line) # 校验字段数量后写入CSV if len(stats_fields) == len(csv_headers): csv_writer.writerow(stats_fields) # 如需查看所有清理后的行,取消注释 # print(cleaned_line)
关键细节说明
正则匹配逻辑:
^锚定行开头,确保只移除开头的时间戳,不会误删中间内容- 时间戳格式严格匹配示例中的
DD/MM/YY HH:MM:SS:SSS,如果日志年份是四位(如2026/10/22),只需把正则里的\d{2}改成\d{4}即可
CSV处理逻辑:
- 通过
is_in_stats_section标记是否进入统计数据区域,自动过滤无关行 - 用
re.split(r'\s+', ...)分割字段,适配日志中字段间数量不一的空格 - 自定义CSV表头,确保输出格式规范
- 通过
兼容性:
- 没有时间戳的行不会被修改,直接保留原内容
- 空行会被自动跳过,避免无效处理
内容的提问来源于stack exchange,提问作者CaffeineAndCode
相关产品推荐
相关产品推荐

