Python3实现含格式错误的XML日志文件容错式格式化(Pretty Print)方案咨询
I get it—dealing with malformed XML logs that break your pretty-print workflow is super frustrating. Your current code fails because xml.dom.minidom is strict and will crash at the first invalid tag or unclosed element. Let's walk through two practical solutions to either skip corrupted lines or fix minor errors automatically.
Solution 1: Skip Corrupted Lines/Entries
If you want to stick with the standard library and just skip any unparseable parts, we can adjust your code to process content in chunks (like individual lines) and only keep the parts that parse successfully. This way, valid XML gets formatted, and broken parts get logged and ignored.
Here's the revised code:
import xml.dom.minidom def parse_and_prettify(xml_chunk): """Attempt to parse an XML chunk and return formatted XML, or empty string on failure.""" try: # Strip whitespace to avoid parsing issues with empty lines cleaned_chunk = xml_chunk.strip() if not cleaned_chunk: return "" dom = xml.dom.minidom.parseString(cleaned_chunk) return dom.toprettyxml(indent=" ") except Exception as e: # Log the issue without stopping the whole process print(f"Skipping malformed chunk: {cleaned_chunk[:60]}... (Error: {str(e)})") return "" def process_xml_file(input_path, output_path): formatted_content = [] with open(input_path, 'r') as infile: # Process line by line (adjust if your XML spans multiple lines) for line in infile: pretty_chunk = parse_and_prettify(line) if pretty_chunk: formatted_content.append(pretty_chunk) # Write only valid, formatted content to a new file (don't overwrite original!) with open(output_path, 'w') as outfile: outfile.writelines(formatted_content) if __name__ == "__main__": try: # Use a separate output file to preserve your original log process_xml_file(r'output\output_xml.txt', r'output\formatted_output.xml') print("XML processing complete—valid content formatted successfully!") except Exception as e: print(f"File operation error: {str(e)}")
Key Notes:
- We process each line individually, so corrupted lines are skipped without breaking the entire job.
- We write to a new file instead of overwriting the original—always a good practice when modifying logs!
- If your XML entries span multiple lines, you could adjust this to read chunks until you hit a closing root/entry tag, but line-by-line is a simple starting point.
Solution 2: Auto-Repair Minor Errors with lxml
If your XML issues are minor (like unclosed tags, duplicate attributes—common in messy logs), using the lxml library with recovery mode can fix these issues automatically instead of skipping them. This is often more useful than just discarding content.
First, install lxml if you haven't:
pip install lxml
Then use this code:
from lxml import etree def repair_and_prettify(xml_content): try: # Enable recovery mode to fix common XML errors parser = etree.XMLParser(recover=True, encoding='utf-8') # Parse the content—lxml will auto-repair unclosed tags, etc. tree = etree.fromstring(xml_content.encode('utf-8'), parser=parser) # Return formatted XML as a string return etree.tostring(tree, pretty_print=True, encoding='unicode') except Exception as e: print(f"Could not repair XML: {str(e)}") return "" def process_repairable_xml(input_path, output_path): with open(input_path, 'r') as infile: full_content = infile.read() formatted_content = repair_and_prettify(full_content) if formatted_content: with open(output_path, 'w') as outfile: outfile.write(formatted_content) print("XML repaired and formatted successfully!") else: print("Failed to repair or format the XML content.") if __name__ == "__main__": process_repairable_xml(r'output\output_xml.txt', r'output\repaired_formatted.xml')
How This Works:
lxml'srecover=Trueflag tells the parser to fix common issues like unclosed tags, mismatched tags, and duplicate attributes.- It will attempt to make the XML valid before formatting it, so you keep more of your log data instead of skipping it.
Which to Choose?
- Use Solution 1 if your logs have severe, unrecoverable corruption (like random non-XML lines mixed in).
- Use Solution 2 if your issues are typical XML syntax mistakes (unclosed tags, minor formatting errors)—it preserves more data.
内容的提问来源于stack exchange,提问作者ginger_beard

