You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中提取文本特定内容?附日志文本处理需求

Extract "Notes" Content from Log Files with Python

Here's a practical approach to pull out the notes from your log entries and save them to a new file. I'll cover two methods—one using basic string manipulation (great if your logs have a rock-solid consistent format) and another with regex (more flexible for slight variations in log structure).

Method 1: Basic String Splitting (Consistent Log Structure)

If your logs always follow the exact pattern you shared (where notes end right before the username starting with u), this simple split method gets the job done quickly:

# Open input log file and output file in one go (cleaner resource handling)
with open('input_logs.txt', 'r') as input_file, open('extracted_notes.txt', 'w') as output_file:
    for line in input_file:
        # Skip empty lines to avoid clutter in output
        if not line.strip():
            continue
        # Check if the line has the "Notes: " section we need
        if "Notes: " in line:
            # Split the line to get everything after "Notes: "
            notes_section = line.split("Notes: ")[1]
            # Split again at the start of the username (assuming it starts with 'u')
            # This isolates the notes text from the trailing username/date/time
            notes_text = notes_section.split(" u")[0].strip()
            # Write the cleaned note to the output file
            output_file.write(f"{notes_text}\n")

Method 2: Regular Expression (Flexible Log Format)

If your logs might have minor variations (like different username prefixes or slight date formatting shifts), regex is more robust. This pattern targets the notes content specifically, ignoring the trailing metadata:

import re

# Regex pattern to capture notes between "Notes: " and the trailing username/date/time block
note_pattern = re.compile(r'Notes: (.*?)\s+\w+\s+\d{2}\s+\w{3}\s+\d{4}\s+\d{2}:\d{2}:\d{2}')

with open('input_logs.txt', 'r') as input_file, open('extracted_notes.txt', 'w') as output_file:
    for line in input_file:
        match = note_pattern.search(line)
        if match:
            # Grab the captured notes text and trim any extra whitespace
            notes_text = match.group(1).strip()
            output_file.write(f"{notes_text}\n")

Quick Breakdown of the Regex Pattern:

  • Notes: : Matches the exact phrase to kick off our capture
  • (.*?): Non-greedy capture of all characters until we hit the next metadata block
  • \s+\w+: Catches the username (a word with no spaces)
  • \d{2}\s+\w{3}\s+\d{4}: Matches dates like 21 Mar 2018
  • \d{2}:\d{2}:\d{2}: Matches timestamps like 10:38:46

Testing with Your Sample Logs:

For your example lines:

This offense was closed with reason: Non-Issue. Notes: There is no suspicious situation. u0t9231 21 Mar 2018 10:38:46
This offense was closed with reason: Non-Issue. Notes: Expected traffic. u0t9231 21 Mar 2018 14:20:11

Both methods will output these lines to extracted_notes.txt:

There is no suspicious situation.
Expected traffic.

Just swap out input_logs.txt with your actual log file path, and extracted_notes.txt with your preferred output file name.

内容的提问来源于stack exchange,提问作者Kirchoff

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 12:21:36