如何用R将复杂不动产登记XML转换为DataFrame?
Hey there! Let's work through your problem: converting that large (20k buildings) XML file into a structured table, while fixing that annoying XPath error you're hitting.
First, Why That XPath Error Happens
The error message points to your XPath expression //consulta_dnp[1]/ — the trailing slash is the culprit! XPath requires something after the slash (like a child node name or axis), so leaving it hanging makes the expression invalid. You just need to remove that final slash, or specify the child node you want to target.
Solution 1: Using the XML Package (Your Original Tool)
Let's adjust your code to properly parse the XML, handle both valid building nodes and error nodes, and extract all the fields you need.
# Load required package library(XML) # Read the XML file (useInternalNodes=TRUE is better for large files) doc <- xmlParse("your_real_estate.xml", useInternalNodes = TRUE) # Get all <consulta_dnp> nodes all_dnp_nodes <- getNodeSet(doc, "//consulta_dnp") # Define a function to process each node process_single_dnp <- function(node) { # Check if this is an error node (has <lerr> child) is_error <- length(getNodeSet(node, "./lerr")) > 0 if (is_error) { # Return a row marked as error (match the column count) return(data.frame( pc1 = NA, pc2 = NA, car = NA, cc1 = NA, cc2 = NA, np = NA, nm = NA, luso = NA, sfc = NA, cpt = NA, ant = "error", stringsAsFactors = FALSE )) } # Helper function to extract field values (returns NA if node doesn't exist) extract_field <- function(field_name) { field_value <- xpathSApply(node, paste0("./", field_name), xmlValue) if (length(field_value) == 0) NA else field_value } # Extract all required fields for valid nodes data.frame( pc1 = extract_field("pc1"), pc2 = extract_field("pc2"), car = extract_field("car"), cc1 = extract_field("cc1"), cc2 = extract_field("cc2"), np = extract_field("np"), nm = extract_field("nm"), luso = extract_field("luso"), sfc = extract_field("sfc"), cpt = extract_field("cpt"), ant = extract_field("ant"), stringsAsFactors = FALSE ) } # Process all nodes and combine into a single table final_table <- do.call(rbind, lapply(all_dnp_nodes, process_single_dnp)) # View the result head(final_table)
Solution 2: Using xml2 (Modern & More Intuitive)
If you're open to switching packages, xml2 is more user-friendly and handles large files smoothly. Here's how to do it:
# Load required packages library(xml2) library(dplyr) # Define a helper operator for handling NULLs (returns NA if node is missing) `%||%` <- function(x, y) if (is.null(x)) y else x # Read the XML file doc <- read_xml("your_real_estate.xml") # Get all <consulta_dnp> nodes all_dnp_nodes <- xml_find_all(doc, "//consulta_dnp") # Process each node process_node <- function(node) { # Check for error node if (length(xml_find_all(node, "./lerr")) > 0) { return(tibble( pc1 = NA, pc2 = NA, car = NA, cc1 = NA, cc2 = NA, np = NA, nm = NA, luso = NA, sfc = NA, cpt = NA, ant = "error" )) } # Extract fields (trim whitespace and handle missing nodes) tibble( pc1 = xml_text(xml_find_first(node, "./pc1"), trim = TRUE) %||% NA_character_, pc2 = xml_text(xml_find_first(node, "./pc2"), trim = TRUE) %||% NA_character_, car = xml_text(xml_find_first(node, "./car"), trim = TRUE) %||% NA_character_, cc1 = xml_text(xml_find_first(node, "./cc1"), trim = TRUE) %||% NA_character_, cc2 = xml_text(xml_find_first(node, "./cc2"), trim = TRUE) %||% NA_character_, np = xml_text(xml_find_first(node, "./np"), trim = TRUE) %||% NA_character_, nm = xml_text(xml_find_first(node, "./nm"), trim = TRUE) %||% NA_character_, luso = xml_text(xml_find_first(node, "./luso"), trim = TRUE) %||% NA_character_, sfc = xml_text(xml_find_first(node, "./sfc"), trim = TRUE) %||% NA_character_, cpt = xml_text(xml_find_first(node, "./cpt"), trim = TRUE) %||% NA_character_, ant = xml_text(xml_find_first(node, "./ant"), trim = TRUE) %||% NA_character_ ) } # Combine all processed nodes into a table final_table <- bind_rows(lapply(all_dnp_nodes, process_node))
Key Notes for Large Files
- Both methods should handle 20k nodes without issues, but if you run into memory constraints, consider processing nodes in smaller batches.
- The code handles missing fields by returning
NA, so you won't get errors if some buildings don't have all the required nodes. - Error nodes are marked with "error" in the
antcolumn (you can adjust this to a blank row or another marker if preferred).
内容的提问来源于stack exchange,提问作者Jammoyano

