如何高效从TXT生成XML?替代awk的方案咨询
Hey there! Great call steering clear of reinventing the wheel—let's break down some better, more maintainable options for converting that TXT file to XML, instead of cobbling together a custom awk script.
Python's built-in xml.etree.ElementTree requires no extra installs, is highly readable, and makes it easy to tweak XML structures or handle more complex TXT formatting later on.
Here's a quick example:
import xml.etree.ElementTree as ET # Create root XML node root = ET.Element("usage_categories") # Parse the TXT file line by line with open("input.txt", "r") as txt_file: for line in txt_file: line = line.strip() if not line: continue # Split line into components (handles spaces in category names) parts = line.split(maxsplit=2) prefix, category_id, category_name = parts # Build child node for each entry category = ET.SubElement(root, "category") category.set("prefix", prefix) category.set("id", category_id) category.text = category_name # Format and write the final XML tree = ET.ElementTree(root) ET.indent(tree, space=" ", level=0) # Prettify output (Python 3.9+) tree.write("output.xml", encoding="utf-8", xml_declaration=True)
If your TXT format is fixed, using Jinja2 lets you define your XML structure as a template, then plug in parsed TXT data. This is perfect if you need to customize the XML layout quickly.
First install Jinja2: pip install jinja2
Example code:
from jinja2 import Template # Define your XML template (easy to modify later) xml_template = """<?xml version="1.0" encoding="UTF-8"?> <usage_categories> {% for entry in entries %} <category prefix="{{ entry.prefix }}" id="{{ entry.id }}">{{ entry.name }}</category> {% endfor %} </usage_categories> """ # Parse TXT entries into a list of dictionaries entries = [] with open("input.txt", "r") as txt_file: for line in txt_file: line = line.strip() if line: prefix, cat_id, cat_name = line.split(maxsplit=2) entries.append({"prefix": prefix, "id": cat_id, "name": cat_name}) # Render template and write XML output template = Template(xml_template) final_xml = template.render(entries=entries) with open("output.xml", "w") as xml_file: xml_file.write(final_xml)
If you still want to use Awk, you don't have to call it "reinventing the wheel"—just write a clean, maintainable script. Here's a concise version:
BEGIN { print "<?xml version=\"1.0\" encoding=\"UTF-8\"?>" print "<usage_categories>" } # Process lines with at least 3 fields (ignores empty lines) NF >= 3 { prefix = $1 cat_id = $2 # Combine all remaining fields into the category name cat_name = "" for (i=3; i<=NF; i++) { cat_name = cat_name " " $i } gsub(/^ /, "", cat_name) # Trim leading space printf " <category prefix=\"%s\" id=\"%s\">%s</category>\n", prefix, cat_id, cat_name } END { print "</usage_categories>" }
Run it with: awk -f convert.awk input.txt > output.xml
- Go with Python + Standard Library if you need to extend the logic later (e.g., add validation, handle edge cases).
- Use Jinja2 if you want to tweak the XML structure without touching parsing code.
- Stick to the optimized Awk script if you need a lightweight, dependency-free command-line solution.
内容的提问来源于stack exchange,提问作者Ger Cas

