使用R rvest提取FAOLEX网页数据失败,求正确XPath方案
解决rvest提取FAOLEX网页信息返回空值的问题
核心问题分析
FAOLEX详情页的信息项采用固定的document-info-item类结构,之前的XPath要么依赖绝对路径(易因页面微小结构变动失效),要么未精准匹配标签与值的对应关系,导致提取不到有效内容。
精准XPath写法(针对目标字段)
以下是语言、日期、文本类型、关键词等字段的稳健XPath,通过标签文本匹配定位父容器,再提取对应值:
- 语言:
//div[contains(@class, 'document-info-item') and .//label[text()='Language']]/div[@class='document-info-value'] - 通过日期:
//div[contains(@class, 'document-info-item') and .//label[text()='Date of adoption']]/div[@class='document-info-value'] - 文本类型:
//div[contains(@class, 'document-info-item') and .//label[text()='Type of text']]/div[@class='document-info-value'] - 关键词:
//div[contains(@class, 'document-info-item') and .//label[text()='Keywords']]/div[@class='document-info-value']
可运行的R代码示例
library(rvest) library(dplyr) # 单页面信息提取函数 extract_faolex_info <- function(url) { page <- read_html(url) # 提取各字段,html_text2自动处理换行与多余空格 language <- page %>% html_node(xpath = "//div[contains(@class, 'document-info-item') and .//label[text()='Language']]/div[@class='document-info-value']") %>% html_text2() adoption_date <- page %>% html_node(xpath = "//div[contains(@class, 'document-info-item') and .//label[text()='Date of adoption']]/div[@class='document-info-value']") %>% html_text2() text_type <- page %>% html_node(xpath = "//div[contains(@class, 'document-info-item') and .//label[text()='Type of text']]/div[@class='document-info-value']") %>% html_text2() keywords <- page %>% html_node(xpath = "//div[contains(@class, 'document-info-item') and .//label[text()='Keywords']]/div[@class='document-info-value']") %>% html_text2() # 返回结构化数据框 tibble( url = url, language = language, adoption_date = adoption_date, text_type = text_type, keywords = keywords ) } # 测试单页面提取 test_url <- "https://www.fao.org/faolex/results/details/en/c/LEX-FAOC213092" extract_faolex_info(test_url)
批量迭代与追加保存
将需要处理的URL存入向量,用purrr批量提取后合并,再追加到已有数据文件:
library(purrr) # 待处理的URL列表 faolex_urls <- c( "https://www.fao.org/faolex/results/details/en/c/LEX-FAOC213092", "https://www.fao.org/faolex/results/details/en/c/LEX-FAOC213093", "https://www.fao.org/faolex/results/details/en/c/LEX-FAOC213094" ) # 批量提取所有页面数据 all_data <- map_df(faolex_urls, extract_faolex_info) # 追加到已有CSV文件(文件不存在则新建) if (file.exists("faolex_data.csv")) { write.table(all_data, "faolex_data.csv", sep = ",", append = TRUE, row.names = FALSE, col.names = FALSE) } else { write.csv(all_data, "faolex_data.csv", row.names = FALSE) }
额外提示
- 若遇到JS动态加载的内容,可改用
RSelenium或rvest结合V8处理渲染后的HTML。 - 若页面标签文本有变体(如多语言页面),需调整XPath中的
text()匹配内容。
内容的提问来源于stack exchange,提问作者NM_
相关产品推荐
相关产品推荐

