使用R脚本提取HTML文件中的指定元数据(如creation_date)
R脚本提取HTML中meta标签的creation_date字段
你当前的代码问题在于定位标签的方式不对——creation_date不是HTML标签名,而是meta标签的name属性值,所以得用属性选择器来定位对应的meta标签,并且meta标签的目标内容存在于content属性里,不是标签文本,不能用html_text2()提取。
正确提取单个creation_date字段的代码
# 先确保安装并加载rvest包 library(rvest) # 读取本地HTML文件 html <- read_html("/Users/.../A1.html") # 定位name为creation_date的meta标签,提取content属性值 creation_date <- html %>% html_element('meta[name="creation_date"]') %>% html_attr("content") print(creation_date)
一次性提取所有需要的字段
如果要批量提取Record Type、Creator、Creation Date、Subject、To这些字段,可以用批量选择的方式直接生成数据框:
library(rvest) html <- read_html("/Users/.../A1.html") # 定义需要提取的字段名(对应meta的name属性) target_fields <- c("record_type", "creator", "creation_date", "subject", "to") # 批量提取每个字段的content属性值,同时去除前后空白 extracted_data <- sapply(target_fields, function(field) { html %>% html_element(paste0('meta[name="', field, '"]')) %>% html_attr("content") %>% trimws() }) # 转成数据框方便查看和处理 extracted_df <- as.data.frame(t(extracted_data)) print(extracted_df)
运行后就能得到干净的目标字段内容,比如creation_date会返回2000-11-22。
内容的提问来源于stack exchange,提问作者benjamin1989
相关产品推荐
相关产品推荐

