You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

rvest::html_nodes()重复调用返回结果不一致的原因咨询

问题:相同解析操作两次执行结果不同的原因

我运行以下R代码:

# load required libraries
library(rvest)
library(tidyverse)
 
# provide sample url
url <- "https://www.thewholesaler.co.uk/cgi-bin/go.cgi?id=3074"
 
# read html
doc <- rvest::read_html(url)

# parse html 1st time
doc %>% rvest::html_nodes("meta")

# parse html 2nd time
doc %>% rvest::html_nodes("meta")

得到如下运行结果:

> # parse html 1st time
> doc %>% rvest::html_nodes("meta")
{xml_nodeset (2)}
[1] <meta name="robots" content="noindex">
[2] <meta http-equiv="REFRESH" content="0;URL=http://www.puckator-dropship.co.uk/gifts/	">
> 
> # parse html 2nd time
> doc %>% rvest::html_nodes("meta")
{xml_nodeset (3)}
[1] <meta http-equiv="Content-Type" content="text/html; charset=UTF-8">
[2] <meta name="robots" content="noindex">
[3] <meta http-equiv="REFRESH" content="0;URL=http://www.puckator-dropship.co.uk/gifts/	">

在代码完全相同、查询的HTML文档未更改的情况下,两次解析得到的结果却不相同?我使用的是rvest v1.0.3版本。


解答

这是因为rvest依赖的底层解析库libxml2有延迟补全HTML节点的特性:

  • 调用read_html()读取文档时,libxml2不会立刻完成完整的HTML结构补全,只会先加载原始文档内容。
  • 第一次调用html_nodes()时,libxml2启动初步解析,但还没补全HTML规范要求的所有必要节点(比如Content-Type这类meta标签,原文档可能缺失,但解析器会自动补全)。
  • 第二次调用html_nodes()时,libxml2已经完成全部结构补全操作,所以能返回包含补全节点的完整结果。

这是libxml2的标准行为,并非rvest的bug,在基于libxml2的rvest版本(包括v1.0.3)中都会出现这个现象。


内容的提问来源于stack exchange,提问作者Rez99

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 03:33:34