R语言中将XML文件转换为DataFrame的问题求助
Hey there! Let's break down what's going wrong and get your bank fee XML data into a usable DataFrame.
First: Why xmlToDataFrame() Threw an Error
That error happens because you're mixing two different XML packages:
read_xml()comes from the xml2 package, which creates anxml_documentobject.xmlToDataFrame()is part of the older XML package, which expects objects likeXMLInternalDocument(fromxmlTreeParse()) instead.
You can't pass an xml2 object to an XML package function directly—hence the "no inherited method" error.
Fixing the Nested List Problem from Your Current Code
Your code using xmlTreeParse() and xmlSApply() is pulling all data into a nested list because you're extracting values at the root level instead of targeting the individual transaction/fee record nodes. Bank statements almost always have repeating nodes (like <feeRecord> or <transaction>) that hold each line item—you need to isolate those first.
Solution 1: Using the XML Package (Your Original Tool)
Assuming your XML has a structure like this (adjust node names to match your actual file):
<bankFeeStatement> <feeRecord> <transactionID>1001</transactionID> <feeAmount>2.50</feeAmount> <feeType>Monthly Maintenance</feeType> <chargeDate>2024-03-01</chargeDate> </feeRecord> <feeRecord> <transactionID>1002</transactionID> <feeAmount>1.75</feeAmount> <feeType>ATM Withdrawal</feeType> <chargeDate>2024-03-05</chargeDate> </feeRecord> </bankFeeStatement>
Here's how to extract and reshape the data correctly:
library(XML) # Read the XML file with XML package's parser xml_doc <- xmlTreeParse("file.xml", useInternalNodes = TRUE) # Use XPath to target all repeating record nodes (replace //feeRecord with your actual node path) record_nodes <- xpathSApply(xml_doc, "//feeRecord") # For each record, extract its child values and convert to a single-row data frame record_list <- lapply(record_nodes, function(node) { # Get values from all child elements of the record node_values <- xmlSApply(node, xmlValue) # Convert to a row (t() transposes the list into a row) as.data.frame(t(node_values), stringsAsFactors = FALSE) }) # Combine all rows into one DataFrame final_df <- do.call(rbind, record_list)
Solution 2: Using the xml2 Package (Modern & Recommended)
The xml2 package has a more intuitive API and plays nicely with tidyverse tools. Here's how to do the same conversion:
library(xml2) library(dplyr) library(purrr) # Read the XML file xml_doc <- read_xml("file.xml") # Target all record nodes with XPath (adjust the path to match your XML) records <- xml_find_all(xml_doc, "//feeRecord") # Extract each field from every record and bind into a DataFrame final_df <- records %>% map_dfr(function(record) { tibble( transactionID = xml_find_first(record, "./transactionID") %>% xml_text(), feeAmount = xml_find_first(record, "./feeAmount") %>% xml_text() %>% as.numeric(), feeType = xml_find_first(record, "./feeType") %>% xml_text(), chargeDate = xml_find_first(record, "./chargeDate") %>% xml_text() %>% as.Date() ) })
Key Notes for Your Specific XML File
- Replace
//feeRecordwith the actual XPath to your repeating line-item nodes (use tools like XML editors or your browser's dev tools to inspect the XML structure). - Adjust the field names (like
transactionID,feeAmount) to match the tags in your file. - If some fields are optional, add
%>% replace_na("")or similar to handle missing values.
Once you run either of these, you'll have a flat DataFrame where each row is a single bank fee record, and you can access columns normally!
内容的提问来源于stack exchange,提问作者ChrisTL

