You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从指定XML链接提取表格遇阻,求R语言优化方案

Hey there! Let's break down your problem and find better solutions for extracting that table data, plus clarify the differences between rvest's html_* and xml_* functions.

1. Why your initial rvest attempts failed

The URL you're targeting returns an XML document, not HTML. When you use read_html(), it parses the XML as if it were HTML, which misinterprets the structure—hence why html_nodes("table") and .report td didn't work as expected.

2. Key differences between html_* and xml_* functions

  • html_* functions: Built for HTML documents. They leverage CSS selectors (like .report td) and have helper functions like html_table() tailored to standard HTML table structures (with <thead>, <tr>, <td> tags following web conventions).
  • xml_* functions: Designed for XML documents, which have more flexible (and sometimes unstructured) tag names. XML doesn't have a universal "table" standard like HTML, so there's no xml_table() function. Instead, you use XPath queries to target specific nodes and extract their content manually.

3. A simpler XML-based solution (no RSelenium needed)

Since the source is XML, use the xml2 package (which rvest relies on under the hood) to parse it directly into a dataframe. Here's how:

library(xml2)

url <- "http://reports.ieso.ca/public/GenOutputbyFuelHourly/PUB_GenOutputbyFuelHourly.xml"
xml_doc <- read_xml(url)

# Extract all table rows
rows <- xml_find_all(xml_doc, "//table/tr")

# Get header text from the first row's <th> tags
header <- xml_text(xml_find_all(rows[1], "./th"))

# Extract data from remaining rows' <td> tags
table_data <- lapply(rows[-1], function(row) {
  xml_text(xml_find_all(row, "./td"))
})

# Convert to a dataframe and set column names
df <- as.data.frame(do.call(rbind, table_data), stringsAsFactors = FALSE)
colnames(df) <- header

4. Optimizing your RSelenium workflow

If you still prefer using RSelenium (maybe for more dynamic content), you can skip manually splitting the values vector by grabbing the table's HTML directly and parsing it with rvest:

library(RSelenium)
library(rvest)

remDr <- remoteDriver(port = 4567L, browserName = "phantomjs")
remDr$open()
remDr$navigate(url)

# Grab the full HTML of the table element
table_html <- remDr$findElement("css", ".report table")$getElementAttribute("outerHTML")[[1]]

# Parse the HTML table directly into a dataframe
df <- read_html(table_html) %>% html_table() %>% .[[1]]

remDr$close()

This way, you avoid the tedious step of splitting the vector into rows—rvest handles the table structure for you once you have the HTML.

内容的提问来源于stack exchange,提问作者jsg51483

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:34:22