You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R rvest提取FAOLEX网页数据失败,求正确XPath方案

解决rvest提取FAOLEX网页信息返回空值的问题

核心问题分析

FAOLEX详情页的信息项采用固定的document-info-item类结构,之前的XPath要么依赖绝对路径(易因页面微小结构变动失效),要么未精准匹配标签与值的对应关系,导致提取不到有效内容。

精准XPath写法(针对目标字段)

以下是语言、日期、文本类型、关键词等字段的稳健XPath,通过标签文本匹配定位父容器,再提取对应值:

  • 语言://div[contains(@class, 'document-info-item') and .//label[text()='Language']]/div[@class='document-info-value']
  • 通过日期://div[contains(@class, 'document-info-item') and .//label[text()='Date of adoption']]/div[@class='document-info-value']
  • 文本类型://div[contains(@class, 'document-info-item') and .//label[text()='Type of text']]/div[@class='document-info-value']
  • 关键词://div[contains(@class, 'document-info-item') and .//label[text()='Keywords']]/div[@class='document-info-value']

可运行的R代码示例

library(rvest)
library(dplyr)

# 单页面信息提取函数
extract_faolex_info <- function(url) {
  page <- read_html(url)
  
  # 提取各字段,html_text2自动处理换行与多余空格
  language <- page %>% 
    html_node(xpath = "//div[contains(@class, 'document-info-item') and .//label[text()='Language']]/div[@class='document-info-value']") %>% 
    html_text2()
  
  adoption_date <- page %>% 
    html_node(xpath = "//div[contains(@class, 'document-info-item') and .//label[text()='Date of adoption']]/div[@class='document-info-value']") %>% 
    html_text2()
  
  text_type <- page %>% 
    html_node(xpath = "//div[contains(@class, 'document-info-item') and .//label[text()='Type of text']]/div[@class='document-info-value']") %>% 
    html_text2()
  
  keywords <- page %>% 
    html_node(xpath = "//div[contains(@class, 'document-info-item') and .//label[text()='Keywords']]/div[@class='document-info-value']") %>% 
    html_text2()
  
  # 返回结构化数据框
  tibble(
    url = url,
    language = language,
    adoption_date = adoption_date,
    text_type = text_type,
    keywords = keywords
  )
}

# 测试单页面提取
test_url <- "https://www.fao.org/faolex/results/details/en/c/LEX-FAOC213092"
extract_faolex_info(test_url)

批量迭代与追加保存

将需要处理的URL存入向量,用purrr批量提取后合并,再追加到已有数据文件:

library(purrr)

# 待处理的URL列表
faolex_urls <- c(
  "https://www.fao.org/faolex/results/details/en/c/LEX-FAOC213092",
  "https://www.fao.org/faolex/results/details/en/c/LEX-FAOC213093",
  "https://www.fao.org/faolex/results/details/en/c/LEX-FAOC213094"
)

# 批量提取所有页面数据
all_data <- map_df(faolex_urls, extract_faolex_info)

# 追加到已有CSV文件(文件不存在则新建)
if (file.exists("faolex_data.csv")) {
  write.table(all_data, "faolex_data.csv", sep = ",", append = TRUE, row.names = FALSE, col.names = FALSE)
} else {
  write.csv(all_data, "faolex_data.csv", row.names = FALSE)
}

额外提示

  1. 若遇到JS动态加载的内容,可改用RSelenium或rvest结合V8处理渲染后的HTML。
  2. 若页面标签文本有变体(如多语言页面),需调整XPath中的text()匹配内容。

内容的提问来源于stack exchange,提问作者NM_

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 14:12:42