You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python根据type字段值提取XML文件指定内容

Extracting XML Entries with type='summary' Using Python

First, let's fix your approach—right now you're iterating over all <body> and <type> elements independently, which won't link each body to its corresponding type. Instead, we need to target the parent elements that contain all four fields (body, id, type, snippets) for each entry, then filter those where type equals summary.

Here's a revised version of your code that does this correctly:

from bs4 import BeautifulSoup

def load_extract(path):
    # Initialize a list to store structured extracted data
    summary_entries = []
    
    # Parse the XML file (using 'with' ensures proper file handling)
    with open(path, "r") as xml_file:
        soup = BeautifulSoup(xml_file, "html.parser")
    
    # Replace 'entry' with the actual parent tag name in your XML (e.g., 'item', 'record')
    # This tag should wrap each individual set of body/id/type/snippets fields
    for entry in soup.find_all('entry'):
        # Get the type value and check if it's 'summary'
        type_element = entry.find('type')
        if type_element and type_element.text.strip() == 'summary':
            # Extract each field with fallback for missing values
            entry_id = entry.find('id').text.strip() if entry.find('id') else 'N/A'
            body_content = entry.find('body').text.strip() if entry.find('body') else 'N/A'
            snippets_content = entry.find('snippets').text.strip() if entry.find('snippets') else 'N/A'
            
            # Add as a dictionary for easy access later
            summary_entries.append({
                'id': entry_id,
                'body': body_content,
                'type': 'summary',
                'snippets': snippets_content
            })
    
    return summary_entries

# Example usage: Replace with your XML file path
result = load_extract("your_data.xml")
for entry in result:
    print(f"ID: {entry['id']}\nBody: {entry['body']}\nSnippets: {entry['snippets']}\n---")

Key Tips:

  • Parent Tag Adjustment: Make sure to replace 'entry' with the actual name of the tag that wraps each full entry in your XML. For example, if your XML uses <item> for each record, use soup.find_all('item') instead.
  • Robustness: The code includes checks to handle missing fields gracefully (returning 'N/A' instead of crashing).
  • Readability: Storing entries as dictionaries makes it much easier to work with the extracted data compared to separate lists.

If you can share a small sample of your XML structure, I can tweak this code to match your exact file format even better!

内容的提问来源于stack exchange,提问作者user3568044

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:55:30