You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python3实现含格式错误的XML日志文件容错式格式化(Pretty Print)方案咨询

Handling Malformed XML Logs for Pretty-Printing (Skipping Corrupted Entries)

I get it—dealing with malformed XML logs that break your pretty-print workflow is super frustrating. Your current code fails because xml.dom.minidom is strict and will crash at the first invalid tag or unclosed element. Let's walk through two practical solutions to either skip corrupted lines or fix minor errors automatically.

Solution 1: Skip Corrupted Lines/Entries

If you want to stick with the standard library and just skip any unparseable parts, we can adjust your code to process content in chunks (like individual lines) and only keep the parts that parse successfully. This way, valid XML gets formatted, and broken parts get logged and ignored.

Here's the revised code:

import xml.dom.minidom

def parse_and_prettify(xml_chunk):
    """Attempt to parse an XML chunk and return formatted XML, or empty string on failure."""
    try:
        # Strip whitespace to avoid parsing issues with empty lines
        cleaned_chunk = xml_chunk.strip()
        if not cleaned_chunk:
            return ""
        dom = xml.dom.minidom.parseString(cleaned_chunk)
        return dom.toprettyxml(indent="  ")
    except Exception as e:
        # Log the issue without stopping the whole process
        print(f"Skipping malformed chunk: {cleaned_chunk[:60]}... (Error: {str(e)})")
        return ""

def process_xml_file(input_path, output_path):
    formatted_content = []
    with open(input_path, 'r') as infile:
        # Process line by line (adjust if your XML spans multiple lines)
        for line in infile:
            pretty_chunk = parse_and_prettify(line)
            if pretty_chunk:
                formatted_content.append(pretty_chunk)
    
    # Write only valid, formatted content to a new file (don't overwrite original!)
    with open(output_path, 'w') as outfile:
        outfile.writelines(formatted_content)

if __name__ == "__main__":
    try:
        # Use a separate output file to preserve your original log
        process_xml_file(r'output\output_xml.txt', r'output\formatted_output.xml')
        print("XML processing complete—valid content formatted successfully!")
    except Exception as e:
        print(f"File operation error: {str(e)}")

Key Notes:

  • We process each line individually, so corrupted lines are skipped without breaking the entire job.
  • We write to a new file instead of overwriting the original—always a good practice when modifying logs!
  • If your XML entries span multiple lines, you could adjust this to read chunks until you hit a closing root/entry tag, but line-by-line is a simple starting point.

Solution 2: Auto-Repair Minor Errors with lxml

If your XML issues are minor (like unclosed tags, duplicate attributes—common in messy logs), using the lxml library with recovery mode can fix these issues automatically instead of skipping them. This is often more useful than just discarding content.

First, install lxml if you haven't:

pip install lxml

Then use this code:

from lxml import etree

def repair_and_prettify(xml_content):
    try:
        # Enable recovery mode to fix common XML errors
        parser = etree.XMLParser(recover=True, encoding='utf-8')
        # Parse the content—lxml will auto-repair unclosed tags, etc.
        tree = etree.fromstring(xml_content.encode('utf-8'), parser=parser)
        # Return formatted XML as a string
        return etree.tostring(tree, pretty_print=True, encoding='unicode')
    except Exception as e:
        print(f"Could not repair XML: {str(e)}")
        return ""

def process_repairable_xml(input_path, output_path):
    with open(input_path, 'r') as infile:
        full_content = infile.read()
    
    formatted_content = repair_and_prettify(full_content)
    if formatted_content:
        with open(output_path, 'w') as outfile:
            outfile.write(formatted_content)
        print("XML repaired and formatted successfully!")
    else:
        print("Failed to repair or format the XML content.")

if __name__ == "__main__":
    process_repairable_xml(r'output\output_xml.txt', r'output\repaired_formatted.xml')

How This Works:

  • lxml's recover=True flag tells the parser to fix common issues like unclosed tags, mismatched tags, and duplicate attributes.
  • It will attempt to make the XML valid before formatting it, so you keep more of your log data instead of skipping it.

Which to Choose?

  • Use Solution 1 if your logs have severe, unrecoverable corruption (like random non-XML lines mixed in).
  • Use Solution 2 if your issues are typical XML syntax mistakes (unclosed tags, minor formatting errors)—it preserves more data.

内容的提问来源于stack exchange,提问作者ginger_beard

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 03:32:34