从指定XML链接提取表格遇阻,求R语言优化方案
Hey there! Let's break down your problem and find better solutions for extracting that table data, plus clarify the differences between rvest's html_* and xml_* functions.
1. Why your initial rvest attempts failed
The URL you're targeting returns an XML document, not HTML. When you use read_html(), it parses the XML as if it were HTML, which misinterprets the structure—hence why html_nodes("table") and .report td didn't work as expected.
2. Key differences between html_* and xml_* functions
html_*functions: Built for HTML documents. They leverage CSS selectors (like.report td) and have helper functions likehtml_table()tailored to standard HTML table structures (with<thead>,<tr>,<td>tags following web conventions).xml_*functions: Designed for XML documents, which have more flexible (and sometimes unstructured) tag names. XML doesn't have a universal "table" standard like HTML, so there's noxml_table()function. Instead, you use XPath queries to target specific nodes and extract their content manually.
3. A simpler XML-based solution (no RSelenium needed)
Since the source is XML, use the xml2 package (which rvest relies on under the hood) to parse it directly into a dataframe. Here's how:
library(xml2) url <- "http://reports.ieso.ca/public/GenOutputbyFuelHourly/PUB_GenOutputbyFuelHourly.xml" xml_doc <- read_xml(url) # Extract all table rows rows <- xml_find_all(xml_doc, "//table/tr") # Get header text from the first row's <th> tags header <- xml_text(xml_find_all(rows[1], "./th")) # Extract data from remaining rows' <td> tags table_data <- lapply(rows[-1], function(row) { xml_text(xml_find_all(row, "./td")) }) # Convert to a dataframe and set column names df <- as.data.frame(do.call(rbind, table_data), stringsAsFactors = FALSE) colnames(df) <- header
4. Optimizing your RSelenium workflow
If you still prefer using RSelenium (maybe for more dynamic content), you can skip manually splitting the values vector by grabbing the table's HTML directly and parsing it with rvest:
library(RSelenium) library(rvest) remDr <- remoteDriver(port = 4567L, browserName = "phantomjs") remDr$open() remDr$navigate(url) # Grab the full HTML of the table element table_html <- remDr$findElement("css", ".report table")$getElementAttribute("outerHTML")[[1]] # Parse the HTML table directly into a dataframe df <- read_html(table_html) %>% html_table() %>% .[[1]] remDr$close()
This way, you avoid the tedious step of splitting the vector into rows—rvest handles the table structure for you once you have the HTML.
内容的提问来源于stack exchange,提问作者jsg51483

