如何从URL解析XML并保存指定标签至文本文件及解析故障排查
Hey there! Let's work through your XML parsing problems together. I’ll go over what was off in your code and show you working solutions for both the ElementTree and minidom methods you tried.
1. Fixing the xml.etree.ElementTree Approach
Your original code had two key issues here:
ET.parse()expects a file path or file-like object, but you passed raw bytes fromwp.read(). You need to useET.fromstring()instead to parse the byte data directly.- Your file path had unescaped backslashes (Windows requires double backslashes or raw strings) and was incomplete.
Here's the corrected code:
import urllib.request import xml.etree.ElementTree as ET # Fetch XML content from the URL (using 'with' to auto-manage the connection) url = 'https://www.boardgamegeek.com/xmlapi2/thing?id=13&stats=1' with urllib.request.urlopen(url) as wp: xml_bytes = wp.read() # Parse the XML from the byte string root = ET.fromstring(xml_bytes) # Example: Extract a specific tag (primary game name) # Adjust the XPath to target whatever tag/content you need primary_name = root.find('.//name[@type="primary"]').attrib['value'] # Save to a text file (use a raw string for Windows path to avoid escape issues) output_path = r'C:\Users\Excel\Desktop\boardgame_info.txt' with open(output_path, 'w', encoding='utf-8') as f: f.write(f"Primary Game Name: {primary_name}\n") # Add more extracted content here (e.g., stats, description) as needed
Key Fixes:
- Used
ET.fromstring(xml_bytes)instead ofET.parse()to handle the raw XML bytes. - Added a complete, properly formatted Windows file path using a raw string (
r'path'). - Used a
withblock forurlopento ensure the connection is closed automatically.
2. Fixing the xml.dom.minidom Approach
Similar to the ElementTree method, you need to parse the raw bytes directly instead of using a file-focused method. Here's the working version:
from xml.dom import minidom import urllib.request url = 'https://www.boardgamegeek.com/xmlapi2/thing?id=13&stats=1' with urllib.request.urlopen(url) as wp: xml_bytes = wp.read() # Parse the XML byte string with minidom dom = minidom.parseString(xml_bytes) # Example: Extract the primary game name name_tags = dom.getElementsByTagName('name') primary_name = None for tag in name_tags: if tag.getAttribute('type') == 'primary': primary_name = tag.getAttribute('value') break # Save to text file output_path = r'C:\Users\Excel\Desktop\boardgame_info_minidom.txt' with open(output_path, 'w', encoding='utf-8') as f: if primary_name: f.write(f"Primary Game Name (minidom): {primary_name}\n") else: f.write("Primary game name not found in XML.\n")
Key Fixes:
- Used
minidom.parseString(xml_bytes)to parse the raw byte data (instead of trying to useparse()which targets files). - Added proper iteration to filter the correct
nametag by itstypeattribute. - Fixed the file path formatting for Windows.
Quick Notes for Customization
- To extract other content (like stats, descriptions, or other tags), adjust the XPath (for ElementTree) or tag filtering logic (for minidom) to match the structure of the XML from BoardGameGeek.
- Always use
encoding='utf-8'when writing to text files to handle special characters in the XML content.
内容的提问来源于stack exchange,提问作者ASH
相关产品推荐
相关产品推荐

