You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R脚本提取HTML文件中的指定元数据(如creation_date)

R脚本提取HTML中meta标签的creation_date字段

你当前的代码问题在于定位标签的方式不对——creation_date不是HTML标签名,而是meta标签的name属性值,所以得用属性选择器来定位对应的meta标签,并且meta标签的目标内容存在于content属性里,不是标签文本,不能用html_text2()提取。

正确提取单个creation_date字段的代码

# 先确保安装并加载rvest包
library(rvest)

# 读取本地HTML文件
html <- read_html("/Users/.../A1.html")

# 定位name为creation_date的meta标签,提取content属性值
creation_date <- html %>% 
  html_element('meta[name="creation_date"]') %>% 
  html_attr("content")

print(creation_date)

一次性提取所有需要的字段

如果要批量提取Record Type、Creator、Creation Date、Subject、To这些字段,可以用批量选择的方式直接生成数据框:

library(rvest)

html <- read_html("/Users/.../A1.html")

# 定义需要提取的字段名(对应meta的name属性)
target_fields <- c("record_type", "creator", "creation_date", "subject", "to")

# 批量提取每个字段的content属性值,同时去除前后空白
extracted_data <- sapply(target_fields, function(field) {
  html %>% 
    html_element(paste0('meta[name="', field, '"]')) %>% 
    html_attr("content") %>% 
    trimws()
})

# 转成数据框方便查看和处理
extracted_df <- as.data.frame(t(extracted_data))
print(extracted_df)

运行后就能得到干净的目标字段内容,比如creation_date会返回2000-11-22。

内容的提问来源于stack exchange,提问作者benjamin1989

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 07:22:57