You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效从TXT生成XML?替代awk的方案咨询

Hey there! Great call steering clear of reinventing the wheel—let's break down some better, more maintainable options for converting that TXT file to XML, instead of cobbling together a custom awk script.


1. Python + Standard Library (Most Flexible & Maintainable)

Python's built-in xml.etree.ElementTree requires no extra installs, is highly readable, and makes it easy to tweak XML structures or handle more complex TXT formatting later on.

Here's a quick example:

import xml.etree.ElementTree as ET

# Create root XML node
root = ET.Element("usage_categories")

# Parse the TXT file line by line
with open("input.txt", "r") as txt_file:
    for line in txt_file:
        line = line.strip()
        if not line:
            continue
        
        # Split line into components (handles spaces in category names)
        parts = line.split(maxsplit=2)
        prefix, category_id, category_name = parts
        
        # Build child node for each entry
        category = ET.SubElement(root, "category")
        category.set("prefix", prefix)
        category.set("id", category_id)
        category.text = category_name

# Format and write the final XML
tree = ET.ElementTree(root)
ET.indent(tree, space="  ", level=0)  # Prettify output (Python 3.9+)
tree.write("output.xml", encoding="utf-8", xml_declaration=True)
2. Template Engine (Jinja2 for Intuitive XML Structure)

If your TXT format is fixed, using Jinja2 lets you define your XML structure as a template, then plug in parsed TXT data. This is perfect if you need to customize the XML layout quickly.

First install Jinja2: pip install jinja2

Example code:

from jinja2 import Template

# Define your XML template (easy to modify later)
xml_template = """<?xml version="1.0" encoding="UTF-8"?>
<usage_categories>
{% for entry in entries %}
  <category prefix="{{ entry.prefix }}" id="{{ entry.id }}">{{ entry.name }}</category>
{% endfor %}
</usage_categories>
"""

# Parse TXT entries into a list of dictionaries
entries = []
with open("input.txt", "r") as txt_file:
    for line in txt_file:
        line = line.strip()
        if line:
            prefix, cat_id, cat_name = line.split(maxsplit=2)
            entries.append({"prefix": prefix, "id": cat_id, "name": cat_name})

# Render template and write XML output
template = Template(xml_template)
final_xml = template.render(entries=entries)

with open("output.xml", "w") as xml_file:
    xml_file.write(final_xml)
3. Optimized Awk Script (If You Prefer Command-Line Tools)

If you still want to use Awk, you don't have to call it "reinventing the wheel"—just write a clean, maintainable script. Here's a concise version:

BEGIN {
    print "<?xml version=\"1.0\" encoding=\"UTF-8\"?>"
    print "<usage_categories>"
}

# Process lines with at least 3 fields (ignores empty lines)
NF >= 3 {
    prefix = $1
    cat_id = $2
    # Combine all remaining fields into the category name
    cat_name = ""
    for (i=3; i<=NF; i++) {
        cat_name = cat_name " " $i
    }
    gsub(/^ /, "", cat_name)  # Trim leading space
    printf "  <category prefix=\"%s\" id=\"%s\">%s</category>\n", prefix, cat_id, cat_name
}

END {
    print "</usage_categories>"
}

Run it with: awk -f convert.awk input.txt > output.xml


Quick Recommendation
  • Go with Python + Standard Library if you need to extend the logic later (e.g., add validation, handle edge cases).
  • Use Jinja2 if you want to tweak the XML structure without touching parsing code.
  • Stick to the optimized Awk script if you need a lightweight, dependency-free command-line solution.

内容的提问来源于stack exchange,提问作者Ger Cas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:22:51