You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修正Python IIS日志解析器的字段问题并优化日期时间处理

IIS日志转CSV脚本问题解决方案

一、移除字段指示器与提取表头

IIS日志中以#Fields:开头的行是字段定义,其他以#开头的为注释行,处理时直接过滤即可:

  • 遍历日志行时,跳过所有以#开头的行,仅提取#Fields:后的内容作为CSV表头
  • 表头仅写入CSV一次,避免重复输出

示例代码片段:

import csv

with open('iis_log.txt', 'r') as log_f, open('output.csv', 'w', newline='') as csv_f:
    csv_writer = csv.writer(csv_f)
    header_written = False
    log_date = None

    for line in log_f:
        line = line.strip()
        if not line:
            continue
        # 提取日志日期
        if line.startswith('#Date:'):
            log_date = line.split(' ', 1)[1].strip()
            continue
        # 提取字段表头
        if line.startswith('#Fields:'):
            if not header_written:
                fields = line.replace('#Fields:', '').strip().split()
                # 将time字段替换为datetime,方便后续合并
                fields[fields.index('time')] = 'datetime'
                csv_writer.writerow(fields)
                header_written = True
            continue
        # 跳过其他注释行
        if line.startswith('#'):
            continue
        # 处理日志内容行...

二、修复字段向右偏移问题

偏移的核心原因是直接用split()分割时,无法处理带空格的字段(比如cs(User-Agent)包含空格)。改用csv.reader处理日志行,指定分隔符为空格并忽略开头空格,可正确识别带引号的字段:

修改日志行处理逻辑:

# 替换原有的行读取逻辑,用csv.reader处理日志行
log_reader = csv.reader(log_f, delimiter=' ', skipinitialspace=True)
for row in log_reader:
    if not row:
        continue
    # 处理#Date:、#Fields:等行的逻辑保持不变...
    # 此时row已经是正确分割的字段列表

三、合并日期与时间

IIS日志的日期单独存放在#Date:行,每行日志仅包含时间,只需将捕获到的日期与每行的时间字段合并即可:

  • 先从#Date:行提取日期字符串
  • 找到每行的时间字段(默认是第一个字段),将日期与时间拼接为YYYY-MM-DD HH:MM:SS格式的完整时间戳
  • 替换原时间字段为合并后的完整时间戳

完整示例代码:

import csv

def convert_iis_log_to_csv(log_path, csv_path):
    with open(log_path, 'r') as log_f, open(csv_path, 'w', newline='') as csv_f:
        log_reader = csv.reader(log_f, delimiter=' ', skipinitialspace=True)
        csv_writer = csv.writer(csv_f)
        header_written = False
        log_date = None
        time_idx = -1

        for row in log_reader:
            if not row:
                continue
            # 获取日志日期
            if row[0] == '#Date:':
                log_date = row[1]
                continue
            # 处理字段表头
            if row[0] == '#Fields:':
                if not header_written:
                    fields = row[1:]
                    time_idx = fields.index('time')
                    # 替换time字段为datetime
                    fields[time_idx] = 'datetime'
                    csv_writer.writerow(fields)
                    header_written = True
                continue
            # 跳过其他注释行
            if row[0].startswith('#'):
                continue
            # 合并日期与时间
            if log_date and time_idx != -1 and len(row) > time_idx:
                time_str = row[time_idx]
                row[time_idx] = f"{log_date} {time_str}"
            # 写入CSV
            csv_writer.writerow(row)

# 调用示例
convert_iis_log_to_csv('access.log', 'access.csv')

关键说明

  • 如果日志包含多天内容(存在多个#Date:行),每次遇到新的#Date:会自动更新日期,后续日志行都会使用最新日期
  • csv.reader的skipinitialspace=True参数会自动忽略字段前的空格,避免分割出空元素
  • 合并后的时间戳可直接用于Excel、Pandas等工具的日志分析

内容的提问来源于stack exchange,提问作者Kippers

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 21:32:31