You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从URL解析XML并保存指定标签至文本文件及解析故障排查

Hey there! Let's work through your XML parsing problems together. I’ll go over what was off in your code and show you working solutions for both the ElementTree and minidom methods you tried.

1. Fixing the xml.etree.ElementTree Approach

Your original code had two key issues here:

  • ET.parse() expects a file path or file-like object, but you passed raw bytes from wp.read(). You need to use ET.fromstring() instead to parse the byte data directly.
  • Your file path had unescaped backslashes (Windows requires double backslashes or raw strings) and was incomplete.

Here's the corrected code:

import urllib.request
import xml.etree.ElementTree as ET

# Fetch XML content from the URL (using 'with' to auto-manage the connection)
url = 'https://www.boardgamegeek.com/xmlapi2/thing?id=13&stats=1'
with urllib.request.urlopen(url) as wp:
    xml_bytes = wp.read()

# Parse the XML from the byte string
root = ET.fromstring(xml_bytes)

# Example: Extract a specific tag (primary game name)
# Adjust the XPath to target whatever tag/content you need
primary_name = root.find('.//name[@type="primary"]').attrib['value']

# Save to a text file (use a raw string for Windows path to avoid escape issues)
output_path = r'C:\Users\Excel\Desktop\boardgame_info.txt'
with open(output_path, 'w', encoding='utf-8') as f:
    f.write(f"Primary Game Name: {primary_name}\n")
    # Add more extracted content here (e.g., stats, description) as needed

Key Fixes:

  • Used ET.fromstring(xml_bytes) instead of ET.parse() to handle the raw XML bytes.
  • Added a complete, properly formatted Windows file path using a raw string (r'path').
  • Used a with block for urlopen to ensure the connection is closed automatically.
2. Fixing the xml.dom.minidom Approach

Similar to the ElementTree method, you need to parse the raw bytes directly instead of using a file-focused method. Here's the working version:

from xml.dom import minidom
import urllib.request

url = 'https://www.boardgamegeek.com/xmlapi2/thing?id=13&stats=1'
with urllib.request.urlopen(url) as wp:
    xml_bytes = wp.read()

# Parse the XML byte string with minidom
dom = minidom.parseString(xml_bytes)

# Example: Extract the primary game name
name_tags = dom.getElementsByTagName('name')
primary_name = None
for tag in name_tags:
    if tag.getAttribute('type') == 'primary':
        primary_name = tag.getAttribute('value')
        break

# Save to text file
output_path = r'C:\Users\Excel\Desktop\boardgame_info_minidom.txt'
with open(output_path, 'w', encoding='utf-8') as f:
    if primary_name:
        f.write(f"Primary Game Name (minidom): {primary_name}\n")
    else:
        f.write("Primary game name not found in XML.\n")

Key Fixes:

  • Used minidom.parseString(xml_bytes) to parse the raw byte data (instead of trying to use parse() which targets files).
  • Added proper iteration to filter the correct name tag by its type attribute.
  • Fixed the file path formatting for Windows.

Quick Notes for Customization

  • To extract other content (like stats, descriptions, or other tags), adjust the XPath (for ElementTree) or tag filtering logic (for minidom) to match the structure of the XML from BoardGameGeek.
  • Always use encoding='utf-8' when writing to text files to handle special characters in the XML content.

内容的提问来源于stack exchange,提问作者ASH

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 08:18:16