You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python解析多层XML文件并转换为DataFrame?

Fixing XML to DataFrame Conversion in Python

Hey there! Let's sort out that XML-to-DataFrame issue you're facing. Your nested loops aren't working because you're reusing the same variable name (element) and not properly navigating the nested sections (like Location, ListingDetails, and BasicDetails) under each Listing entry. Here's a straightforward solution using lxml (or the standard xml.etree library) and pandas:

Step-by-Step Solution

First, we'll parse the XML, extract all relevant data from each Listing into a dictionary, collect these dictionaries into a list, and then convert that list to a DataFrame.

1. Import Required Libraries

import pandas as pd
from lxml import etree  # Or use xml.etree.ElementTree as ET if you prefer the standard library

2. Parse the XML

If your XML is in a file:

tree = etree.parse("your_listings.xml")
root = tree.getroot()

If you're working with an XML string:

xml_str = """<Listings>
  <Listing>
    <Location>
      <City>Amagansett</City>
      <State>NY</State>
      <Zip>11930</Zip>
      <Latitude>4.12</Latitude>
      <Longitude>2.13</Longitude>
      <DisplayAddress>No</DisplayAddress>
    </Location>
    <ListingDetails>
      <Status>For Rent</Status>
      <Price>120000</Price>
      <ListingUrl>http://www.co.com/listing.aspx?Region=LI3&amp;ListingID=122</ListingUrl>
      <MlsId>122</MlsId>
      <DateListed>2011-06-10</DateListed>
      <NewDevelopment>N</NewDevelopment>
    </ListingDetails>
    <BasicDetails>
      <PropertyType>Other</PropertyType>
      <Description>Rental Registration #: the master suite has a lavish bath and its own terrace with small ocean views..</Description>
      <Bedrooms>5</Bedrooms>
      <Bathrooms>4</Bathrooms>
      <FullBathrooms>4</FullBathrooms>
      <HalfBathrooms>0</HalfBathrooms>
      <LivingArea>5775</LivingArea>
      <LotSize>0.8</LotSize>
    </BasicDetails>
  </Listing>
</Listings>"""

root = etree.fromstring(xml_str)

3. Extract Data and Convert to DataFrame

# Initialize an empty list to hold all listing data
listing_records = []

# Iterate over each <Listing> element in the XML
for listing in root.findall("Listing"):
    # Create a dictionary to store data for this listing
    listing_data = {}
    
    # Extract data from <Location> section
    location = listing.find("Location")
    for elem in location:
        listing_data[elem.tag] = elem.text
    
    # Extract data from <ListingDetails> section
    details = listing.find("ListingDetails")
    for elem in details:
        listing_data[elem.tag] = elem.text
    
    # Extract data from <BasicDetails> section
    basic_details = listing.find("BasicDetails")
    for elem in basic_details:
        listing_data[elem.tag] = elem.text
    
    # Add the listing's data to our records list
    listing_records.append(listing_data)

# Convert the list of dictionaries to a pandas DataFrame
df = pd.DataFrame(listing_records)

# View the result
print(df.head())

Why This Works

  • We explicitly target each nested section (Location, ListingDetails, BasicDetails) under every Listing, so we don't miss any fields.
  • Each element's tag (like City, Price) becomes a column name in the DataFrame, and the element's text becomes the corresponding value.
  • We avoid variable name conflicts (your original code reused element for different loops, which caused confusion).

If Using the Standard xml.etree Library

The code is nearly identical—just replace the import and parsing lines with:

import xml.etree.ElementTree as ET

# For a file:
tree = ET.parse("your_listings.xml")
root = tree.getroot()

# For a string:
root = ET.fromstring(xml_str)

内容的提问来源于stack exchange,提问作者Liu Yu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 10:18:47