You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R将复杂不动产登记XML转换为DataFrame?

Fixing XPath Error & Converting Large Real Estate XML to Table in R

Hey there! Let's work through your problem: converting that large (20k buildings) XML file into a structured table, while fixing that annoying XPath error you're hitting.

First, Why That XPath Error Happens

The error message points to your XPath expression //consulta_dnp[1]/ — the trailing slash is the culprit! XPath requires something after the slash (like a child node name or axis), so leaving it hanging makes the expression invalid. You just need to remove that final slash, or specify the child node you want to target.

Solution 1: Using the XML Package (Your Original Tool)

Let's adjust your code to properly parse the XML, handle both valid building nodes and error nodes, and extract all the fields you need.

# Load required package
library(XML)

# Read the XML file (useInternalNodes=TRUE is better for large files)
doc <- xmlParse("your_real_estate.xml", useInternalNodes = TRUE)

# Get all <consulta_dnp> nodes
all_dnp_nodes <- getNodeSet(doc, "//consulta_dnp")

# Define a function to process each node
process_single_dnp <- function(node) {
  # Check if this is an error node (has <lerr> child)
  is_error <- length(getNodeSet(node, "./lerr")) > 0
  
  if (is_error) {
    # Return a row marked as error (match the column count)
    return(data.frame(
      pc1 = NA, pc2 = NA, car = NA, cc1 = NA, cc2 = NA,
      np = NA, nm = NA, luso = NA, sfc = NA, cpt = NA,
      ant = "error",
      stringsAsFactors = FALSE
    ))
  }
  
  # Helper function to extract field values (returns NA if node doesn't exist)
  extract_field <- function(field_name) {
    field_value <- xpathSApply(node, paste0("./", field_name), xmlValue)
    if (length(field_value) == 0) NA else field_value
  }
  
  # Extract all required fields for valid nodes
  data.frame(
    pc1 = extract_field("pc1"),
    pc2 = extract_field("pc2"),
    car = extract_field("car"),
    cc1 = extract_field("cc1"),
    cc2 = extract_field("cc2"),
    np = extract_field("np"),
    nm = extract_field("nm"),
    luso = extract_field("luso"),
    sfc = extract_field("sfc"),
    cpt = extract_field("cpt"),
    ant = extract_field("ant"),
    stringsAsFactors = FALSE
  )
}

# Process all nodes and combine into a single table
final_table <- do.call(rbind, lapply(all_dnp_nodes, process_single_dnp))

# View the result
head(final_table)

Solution 2: Using xml2 (Modern & More Intuitive)

If you're open to switching packages, xml2 is more user-friendly and handles large files smoothly. Here's how to do it:

# Load required packages
library(xml2)
library(dplyr)

# Define a helper operator for handling NULLs (returns NA if node is missing)
`%||%` <- function(x, y) if (is.null(x)) y else x

# Read the XML file
doc <- read_xml("your_real_estate.xml")

# Get all <consulta_dnp> nodes
all_dnp_nodes <- xml_find_all(doc, "//consulta_dnp")

# Process each node
process_node <- function(node) {
  # Check for error node
  if (length(xml_find_all(node, "./lerr")) > 0) {
    return(tibble(
      pc1 = NA, pc2 = NA, car = NA, cc1 = NA, cc2 = NA,
      np = NA, nm = NA, luso = NA, sfc = NA, cpt = NA,
      ant = "error"
    ))
  }
  
  # Extract fields (trim whitespace and handle missing nodes)
  tibble(
    pc1 = xml_text(xml_find_first(node, "./pc1"), trim = TRUE) %||% NA_character_,
    pc2 = xml_text(xml_find_first(node, "./pc2"), trim = TRUE) %||% NA_character_,
    car = xml_text(xml_find_first(node, "./car"), trim = TRUE) %||% NA_character_,
    cc1 = xml_text(xml_find_first(node, "./cc1"), trim = TRUE) %||% NA_character_,
    cc2 = xml_text(xml_find_first(node, "./cc2"), trim = TRUE) %||% NA_character_,
    np = xml_text(xml_find_first(node, "./np"), trim = TRUE) %||% NA_character_,
    nm = xml_text(xml_find_first(node, "./nm"), trim = TRUE) %||% NA_character_,
    luso = xml_text(xml_find_first(node, "./luso"), trim = TRUE) %||% NA_character_,
    sfc = xml_text(xml_find_first(node, "./sfc"), trim = TRUE) %||% NA_character_,
    cpt = xml_text(xml_find_first(node, "./cpt"), trim = TRUE) %||% NA_character_,
    ant = xml_text(xml_find_first(node, "./ant"), trim = TRUE) %||% NA_character_
  )
}

# Combine all processed nodes into a table
final_table <- bind_rows(lapply(all_dnp_nodes, process_node))

Key Notes for Large Files

  • Both methods should handle 20k nodes without issues, but if you run into memory constraints, consider processing nodes in smaller batches.
  • The code handles missing fields by returning NA, so you won't get errors if some buildings don't have all the required nodes.
  • Error nodes are marked with "error" in the ant column (you can adjust this to a blank row or another marker if preferred).

内容的提问来源于stack exchange,提问作者Jammoyano

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:54:21