如何使用Python解析多层XML文件并转换为DataFrame?
Hey there! Let's sort out that XML-to-DataFrame issue you're facing. Your nested loops aren't working because you're reusing the same variable name (element) and not properly navigating the nested sections (like Location, ListingDetails, and BasicDetails) under each Listing entry. Here's a straightforward solution using lxml (or the standard xml.etree library) and pandas:
Step-by-Step Solution
First, we'll parse the XML, extract all relevant data from each Listing into a dictionary, collect these dictionaries into a list, and then convert that list to a DataFrame.
1. Import Required Libraries
import pandas as pd from lxml import etree # Or use xml.etree.ElementTree as ET if you prefer the standard library
2. Parse the XML
If your XML is in a file:
tree = etree.parse("your_listings.xml") root = tree.getroot()
If you're working with an XML string:
xml_str = """<Listings> <Listing> <Location> <City>Amagansett</City> <State>NY</State> <Zip>11930</Zip> <Latitude>4.12</Latitude> <Longitude>2.13</Longitude> <DisplayAddress>No</DisplayAddress> </Location> <ListingDetails> <Status>For Rent</Status> <Price>120000</Price> <ListingUrl>http://www.co.com/listing.aspx?Region=LI3&ListingID=122</ListingUrl> <MlsId>122</MlsId> <DateListed>2011-06-10</DateListed> <NewDevelopment>N</NewDevelopment> </ListingDetails> <BasicDetails> <PropertyType>Other</PropertyType> <Description>Rental Registration #: the master suite has a lavish bath and its own terrace with small ocean views..</Description> <Bedrooms>5</Bedrooms> <Bathrooms>4</Bathrooms> <FullBathrooms>4</FullBathrooms> <HalfBathrooms>0</HalfBathrooms> <LivingArea>5775</LivingArea> <LotSize>0.8</LotSize> </BasicDetails> </Listing> </Listings>""" root = etree.fromstring(xml_str)
3. Extract Data and Convert to DataFrame
# Initialize an empty list to hold all listing data listing_records = [] # Iterate over each <Listing> element in the XML for listing in root.findall("Listing"): # Create a dictionary to store data for this listing listing_data = {} # Extract data from <Location> section location = listing.find("Location") for elem in location: listing_data[elem.tag] = elem.text # Extract data from <ListingDetails> section details = listing.find("ListingDetails") for elem in details: listing_data[elem.tag] = elem.text # Extract data from <BasicDetails> section basic_details = listing.find("BasicDetails") for elem in basic_details: listing_data[elem.tag] = elem.text # Add the listing's data to our records list listing_records.append(listing_data) # Convert the list of dictionaries to a pandas DataFrame df = pd.DataFrame(listing_records) # View the result print(df.head())
Why This Works
- We explicitly target each nested section (
Location,ListingDetails,BasicDetails) under everyListing, so we don't miss any fields. - Each element's tag (like
City,Price) becomes a column name in the DataFrame, and the element's text becomes the corresponding value. - We avoid variable name conflicts (your original code reused
elementfor different loops, which caused confusion).
If Using the Standard xml.etree Library
The code is nearly identical—just replace the import and parsing lines with:
import xml.etree.ElementTree as ET # For a file: tree = ET.parse("your_listings.xml") root = tree.getroot() # For a string: root = ET.fromstring(xml_str)
内容的提问来源于stack exchange,提问作者Liu Yu

