You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从HTML文件提取DT R包生成的表格为tibble?

从DT生成的HTML中提取表格并转为命名tibble列表

解决方案代码

library(rvest)
library(jsonlite)
library(tibble)
library(purrr)

# 读取目标HTML文件
html_content <- read_html("demo.html")

# 定位到"Display data"章节区域
target_section <- html_element(html_content, "#display-data")

# 获取表格对应的名称(从二级标题提取)
table_titles <- target_section |> 
  html_elements("h2") |> 
  html_text()

# 提取所有DT表格对应的JSON脚本
dt_json_scripts <- target_section |> 
  html_elements('script[type="application/json"][data-for^="htmlwidget"]')

# 解析JSON并转换为tibble列表
dt_tibble_list <- map(dt_json_scripts, function(script_node) {
  # 提取脚本内的JSON文本
  raw_json <- html_text(script_node)
  # 解析JSON
  parsed_data <- fromJSON(raw_json)
  # 将表格数据转为tibble
  as_tibble(parsed_data$x$data)
})

# 为列表添加对应名称
named_dt_list <- set_names(dt_tibble_list, table_titles)

# 查看最终结果
named_dt_list

步骤说明

  1. 定位目标区域:通过#display-data选择器定位到包含所有表格的章节,避免处理页面其他无关内容。
  2. 提取表格名称:从章节内的<h2>标签提取文本,作为最终列表的命名依据,确保和表格一一对应。
  3. 筛选DT脚本:通过属性筛选,只提取DT表格对应的application/json类型脚本(这类脚本存储了DT的完整数据)。
  4. 解析JSON并转tibble:遍历每个脚本,提取JSON文本后解析,从x$data节点获取原始表格数据,转为tibble格式。
  5. 命名列表:将提取的标题和tibble列表绑定,得到符合需求的命名列表。

这个方法无需重新运行参数化RMarkdown,直接从已生成的HTML中提取数据,节省计算资源和时间。

内容的提问来源于stack exchange,提问作者Ashirwad

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 21:05:26