使用R语言读取指定XML文件失败及数据转换问题求助
Hey there! Let's work through this XML-to-DataFrame issue you're facing with that Stortinget dataset. I’ve dealt with similar XML parsing headaches in R before, so here’s a step-by-step solution using the xml2 package (since that’s the one that successfully read your data):
Step 1: Confirm you're reading the XML correctly
First, let’s formalize the working read step to make sure we’re starting on solid ground:
# Load required packages library(xml2) library(tidyverse) # For easy DataFrame manipulation # Read the XML document from the URL xml_doc <- read_xml("https://data.stortinget.no/eksport/sak?sakid=76122")
If this runs without errors, you’re all set—xml2 handled the XML structure that tripped up the older XML package.
Step 2: Parse XML nodes into a DataFrame
XML is hierarchical, so we can’t convert it to a DataFrame in one click. We need to target specific nodes and extract the data we want. Let’s start with core details from the main <sak> node and its direct children:
# Extract main <sak> node content and convert to a DataFrame sak_df <- xml_find_all(xml_doc, ".//sak") %>% purrr::map_df(function(node) { # Pull attributes and text from key child elements tibble( sak_id = xml_attr(node, "sakid"), sak_type = xml_attr(node, "saktype"), tittel = xml_find_first(node, ".//tittel") %>% xml_text(), status = xml_find_first(node, ".//status") %>% xml_text(), registreringsdato = xml_find_first(node, ".//registreringsdato") %>% xml_text() ) }) # View the resulting DataFrame head(sak_df)
This code uses xml_find_all to locate the main node, then map_df to iterate over each node and extract specific attributes/text into a tibble (which behaves exactly like a standard DataFrame).
Bonus: Handling nested child nodes
If your XML has repeated child nodes (like multiple <del> entries under <sak>), we can either nest those entries or unnest them for a flat table. Here’s how to handle nested <del> nodes:
# Extract main sak data plus nested del entries sak_with_dels <- xml_find_all(xml_doc, ".//sak") %>% purrr::map_df(function(sak_node) { # Extract all <del> nodes for this sak del_data <- xml_find_all(sak_node, ".//del") %>% purrr::map_df(function(del_node) { tibble( del_id = xml_attr(del_node, "delid"), del_tittel = xml_find_first(del_node, ".//tittel") %>% xml_text(), del_status = xml_find_first(del_node, ".//status") %>% xml_text() ) }) # Combine main sak data with nested del data tibble( sak_id = xml_attr(sak_node, "sakid"), sak_tittel = xml_find_first(sak_node, ".//tittel") %>% xml_text(), del_entries = list(del_data) # Store del data as a nested list column ) }) # Optional: Unnest to get a flat table structure sak_flat <- sak_with_dels %>% unnest(del_entries) head(sak_flat)
Why did the XML package throw an error?
The error XML content does not seem to be XML: '' typically pops up when the older XML package can’t parse the content returned by the URL. This is likely due to:
- The URL returning content with encoding that
XMLdoesn’t handle well - Minor structural quirks in the XML that the more modern
xml2parser is more tolerant of
Sticking with xml2 is the right call here—it’s actively maintained and handles modern XML structures far more reliably.
内容的提问来源于stack exchange,提问作者Giovanni Romeo

