You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在PowerShell中提取括号间的Loading Length数值并导出CSV

提取日志中"Loading Length"括号内数值并导出为CSV

需求说明

  • 从日志文件中提取日期、时间、完整消息、Loading Length后括号内的纯数值
  • 仅处理包含"Loading Length"的条目,忽略所有"RunOut"相关内容
  • 将结果导出为分栏清晰的CSV文件

日志示例

2024-05-20 14:30:15 INFO: Loading Length (123.45) - processing file
2024-05-20 14:30:20 WARN: RunOut Length (67.89) - low resources
2024-05-20 14:30:25 INFO: Loading Length (987.65) - completed task

当前代码问题

你提供的代码会提取所有含"Length"的条目,无法区分"Loading Length"和"RunOut Length",导致非目标数值被混入结果。示例代码如下:

import re
import csv

def process_log(log_path, output_csv):
    with open(log_path, 'r') as f, open(output_csv, 'w', newline='') as csv_f:
        writer = csv.writer(csv_f)
        writer.writerow(['Date', 'Time', 'Message', 'Length'])
        
        for line in f:
            dt_match = re.match(r'(\d{4}-\d{2}-\d{2}) (\d{2}:\d{2}:\d{2})', line)
            if not dt_match:
                continue
            date, time = dt_match.groups()
            
            length_match = re.search(r'Length \((.*?)\)', line)
            length = length_match.group(1) if length_match else ''
            
            message = line[20:].strip()
            
            writer.writerow([date, time, message, length])

process_log('app.log', 'output.csv')

解决方案:精准匹配修改

修改后的代码

import re
import csv

def process_log(log_path, output_csv):
    # 正则精准匹配含Loading Length的日志行,捕获日期、时间、消息前缀、数值、消息后缀
    log_regex = re.compile(
        r'(\d{4}-\d{2}-\d{2}) (\d{2}:\d{2}:\d{2}) (.*?): Loading Length \(([\d.]+)\)(.*)'
    )
    
    with open(log_path, 'r', encoding='utf-8') as log_file, open(output_csv, 'w', newline='', encoding='utf-8') as csv_file:
        csv_writer = csv.writer(csv_file)
        # 写入CSV表头
        csv_writer.writerow(['Date', 'Time', 'Message', 'Loading Length'])
        
        for line in log_file:
            match_result = log_regex.search(line)
            if match_result:
                date, time, log_level, length_value, message_detail = match_result.groups()
                # 拼接完整消息内容
                full_message = f"{log_level}: Loading Length ({length_value}){message_detail}".strip()
                # 写入CSV行
                csv_writer.writerow([date, time, full_message, length_value])

# 调用函数,替换为你的日志路径和输出路径
process_log('app.log', 'loading_length_output.csv')

关键修改点

  1. 精准正则匹配:使用log_regex只匹配包含"Loading Length"的行,自动过滤"RunOut"条目
  2. 数值提取:通过([\d.]+)专门捕获括号内的数字(支持整数和小数格式)
  3. 编码处理:添加encoding='utf-8'避免中文或特殊字符乱码
  4. 消息完整性:拼接日志级别、Loading Length内容和消息后缀,保留完整原始消息

期望CSV输出

Date,Time,Message,Loading Length
2024-05-20,14:30:15,INFO: Loading Length (123.45) - processing file,123.45
2024-05-20,14:30:25,INFO: Loading Length (987.65) - completed task,987.65

内容的提问来源于stack exchange,提问作者UselessCoder

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 11:23:35