如何去除Python日志冗余信息并解析为带指定标签的XML?
Got it, let's fix those messy logs and turn them into clean, tagged XML that only keeps the data you care about. Here's a straightforward Python solution tailored to your log format:
Step 1: Parse the Log Entries
First, we'll use regular expressions to yank out the four key fields (phone number, message text, user ID, event ID) from each log line. Your log lines follow a pattern where the critical data is in a tuple like ('29999999', 'text1', '666', '1'), so we can target that specifically.
Step 2: Generate Structured XML
Once we have the extracted fields, we'll build valid XML with dedicated tags for each piece of data. We'll use Python's built-in xml.etree.ElementTree module to ensure proper XML formatting (no manual string concatenation mess!).
Full Code Example
import re import xml.etree.ElementTree as ET from xml.dom import minidom def parse_log_to_xml(log_file_path, output_xml_path): # Regex pattern to match the tuple with our target fields log_pattern = re.compile(r"\('(\d+)', '([^']*)', '(\d+)', '(\d+)'\)") # Create root element for XML root = ET.Element("sms_logs") with open(log_file_path, 'r', encoding='utf-8') as log_file: for line_num, line in enumerate(log_file, 1): # Skip empty lines if not line.strip(): continue # Search for the pattern in the line match = log_pattern.search(line) if match: phone_num, message, user_id, event_id = match.groups() # Create a child element for each SMS event sms_event = ET.SubElement(root, "sms_event") ET.SubElement(sms_event, "phone_number").text = phone_num ET.SubElement(sms_event, "message_content").text = message ET.SubElement(sms_event, "user_id").text = user_id ET.SubElement(sms_event, "event_id").text = event_id else: # Optional: Log lines that don't match the pattern print(f"Skipping line {line_num}: No matching data found") # Pretty-print the XML for readability rough_string = ET.tostring(root, 'utf-8') reparsed = minidom.parseString(rough_string) pretty_xml = reparsed.toprettyxml(indent=" ") # Write the formatted XML to file with open(output_xml_path, 'w', encoding='utf-8') as xml_file: xml_file.write(pretty_xml) # Usage parse_log_to_xml("sms.log", "cleaned_sms_logs.xml")
What This Does:
- Regex Matching: The pattern
r"\('(\d+)', '([^']*)', '(\d+)', '(\d+)'\)"targets the tuple in your logs and captures each field into groups. - XML Structure: Creates a root
<sms_logs>element, with each valid log entry becoming a<sms_event>child containing your tagged fields. - Readability: Uses
minidomto pretty-print the XML so it's easy to read, instead of a single block of text. - Error Handling: Skips lines that don't match the expected format and prints a warning (you can adjust this to log errors to a file if needed).
Example Output XML:
<?xml version="1.0" ?> <sms_logs> <sms_event> <phone_number>29999999</phone_number> <message_content>text1</message_content> <user_id>666</user_id> <event_id>1</event_id> </sms_event> <!-- More <sms_event> entries here --> </sms_logs>
If your log lines have slight variations (like different quote styles or extra characters in the tuple), just tweak the regex pattern to match your actual log format. Let me know if you need help adjusting it!
内容的提问来源于stack exchange,提问作者NaturalHabitat

