如何在Python中提取文本特定内容?附日志文本处理需求
Here's a practical approach to pull out the notes from your log entries and save them to a new file. I'll cover two methods—one using basic string manipulation (great if your logs have a rock-solid consistent format) and another with regex (more flexible for slight variations in log structure).
Method 1: Basic String Splitting (Consistent Log Structure)
If your logs always follow the exact pattern you shared (where notes end right before the username starting with u), this simple split method gets the job done quickly:
# Open input log file and output file in one go (cleaner resource handling) with open('input_logs.txt', 'r') as input_file, open('extracted_notes.txt', 'w') as output_file: for line in input_file: # Skip empty lines to avoid clutter in output if not line.strip(): continue # Check if the line has the "Notes: " section we need if "Notes: " in line: # Split the line to get everything after "Notes: " notes_section = line.split("Notes: ")[1] # Split again at the start of the username (assuming it starts with 'u') # This isolates the notes text from the trailing username/date/time notes_text = notes_section.split(" u")[0].strip() # Write the cleaned note to the output file output_file.write(f"{notes_text}\n")
Method 2: Regular Expression (Flexible Log Format)
If your logs might have minor variations (like different username prefixes or slight date formatting shifts), regex is more robust. This pattern targets the notes content specifically, ignoring the trailing metadata:
import re # Regex pattern to capture notes between "Notes: " and the trailing username/date/time block note_pattern = re.compile(r'Notes: (.*?)\s+\w+\s+\d{2}\s+\w{3}\s+\d{4}\s+\d{2}:\d{2}:\d{2}') with open('input_logs.txt', 'r') as input_file, open('extracted_notes.txt', 'w') as output_file: for line in input_file: match = note_pattern.search(line) if match: # Grab the captured notes text and trim any extra whitespace notes_text = match.group(1).strip() output_file.write(f"{notes_text}\n")
Quick Breakdown of the Regex Pattern:
Notes:: Matches the exact phrase to kick off our capture(.*?): Non-greedy capture of all characters until we hit the next metadata block\s+\w+: Catches the username (a word with no spaces)\d{2}\s+\w{3}\s+\d{4}: Matches dates like21 Mar 2018\d{2}:\d{2}:\d{2}: Matches timestamps like10:38:46
Testing with Your Sample Logs:
For your example lines:
This offense was closed with reason: Non-Issue. Notes: There is no suspicious situation. u0t9231 21 Mar 2018 10:38:46
This offense was closed with reason: Non-Issue. Notes: Expected traffic. u0t9231 21 Mar 2018 14:20:11
Both methods will output these lines to extracted_notes.txt:
There is no suspicious situation. Expected traffic.
Just swap out input_logs.txt with your actual log file path, and extracted_notes.txt with your preferred output file name.
内容的提问来源于stack exchange,提问作者Kirchoff

