You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何拆分并提取方括号间的文本?日志内容提取方案求解

提取日志中指定时间与消息大小的实现方案

嗨,我来帮你搞定这个日志提取的问题!你提到想用split方法,其实有两种实用思路——split配合字符串处理,或者用正则表达式(更适配这种结构化日志场景),我都给你详细讲讲:

首先先明确你的原始日志内容:

[2018-01-06 18:52:00,376] INFO [workflow@xxx] - [com.app.gateway] size of [1] message 
[2018-01-06 18:54:00,188] INFO [workflow@xxx] - [com.app.gateway] size of [3] message 
[2018-01-06 18:55:00,140] INFO [workflow@xxx] - [com.app.gateway] size of [5] message

方法一:使用split逐步拆分

既然你想尝试split,我们可以利用日志的固定结构,通过拆分字符来定位目标内容:

  1. 先把日志按换行拆分成单独的行
  2. 对每一行用]拆分,拿到所有被方括号包裹的片段
  3. 从第一个片段中提取不带毫秒的时间,从包含size of的片段中提取数字

代码示例(Python):

log_content = """[2018-01-06 18:52:00,376] INFO [workflow@xxx] - [com.app.gateway] size of [1] message 
[2018-01-06 18:54:00,188] INFO [workflow@xxx] - [com.app.gateway] size of [3] message 
[2018-01-06 18:55:00,140] INFO [workflow@xxx] - [com.app.gateway] size of [5] message"""

result = []
# 按行处理日志
for line in log_content.splitlines():
    # 用']'拆分当前行的所有片段
    parts = line.split(']')
    # 提取时间:去掉开头的'[',再按逗号拆分取前半部分(去掉毫秒)
    raw_time = parts[0].strip('[')
    clean_time = raw_time.split(',')[0]
    result.append(f"[{clean_time}]")
    
    # 提取size数字:筛选出包含'size of'的片段,拆分'['后取数字部分
    size_segment = [p for p in parts if 'size of' in p][0]
    size_num = size_segment.split('[')[-1].strip()
    result.append(f"[{size_num}]")

# 拼接成最终输出格式
final_output = ' '.join(result)
print(final_output)

方法二:正则表达式(更简洁高效)

因为你的日志格式非常固定,用正则表达式可以直接精准捕获我们需要的时间和数字,代码会更简洁:
我们的正则会匹配两个核心内容:

  • 不带毫秒的时间:(\d{4}-\d{2}-\d{2} \d{2}:\d{2}:\d{2})
  • 消息大小的数字:(\d+)

代码示例(Python):

import re

log_content = """[2018-01-06 18:52:00,376] INFO [workflow@xxx] - [com.app.gateway] size of [1] message 
[2018-01-06 18:54:00,188] INFO [workflow@xxx] - [com.app.gateway] size of [3] message 
[2018-01-06 18:55:00,140] INFO [workflow@xxx] - [com.app.gateway] size of [5] message"""

# 定义正则模式,捕获时间和size数字
pattern = re.compile(r'\[(\d{4}-\d{2}-\d{2} \d{2}:\d{2}:\d{2}),\d+\].*size of \[(\d+)\]')

result = []
# 遍历所有匹配结果
for match in pattern.finditer(log_content):
    time_str, size_str = match.groups()
    result.append(f"[{time_str}]")
    result.append(f"[{size_str}]")

# 生成最终输出
final_output = ' '.join(result)
print(final_output)

两种方法运行后都会输出你想要的结果:

[2018-01-06 18:52:00] [1] [2018-01-06 18:54:00] [3] [2018-01-06 18:55:00] [5]

内容的提问来源于stack exchange,提问作者causita

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:39:44