You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用RSelenium提取<script>标签内data数组的数值?

解决方法

你可以通过字符串提取+正则匹配的方式从script标签内容中提取目标数值数组,具体步骤如下:

步骤1:提取script标签内的纯代码文本

你已经用htmlParse解析了outerHTML,接下来先提取script标签里的实际代码内容:

# 从解析后的doc中获取script标签内的文本
script_content <- xpathSApply(doc, "//script", xmlValue)[[1]]

步骤2:用正则表达式匹配data后的数值部分

使用正则捕获data: [ ... ]中括号内的数值字符串:

# 匹配data: 后的数组内容(需先安装stringr包)
library(stringr)
data_match <- str_match(script_content, "data:\\s*\\[(.*?)\\]")[,2]

如果不想额外安装包,也可以用基础R实现:

# 基础R正则匹配方式
data_match <- regmatches(script_content, gregexpr("data:\\s*\\[(.*?)\\]", script_content))[[1]]
data_match <- sub("data:\\s*\\[(.*?)\\]", "\\1", data_match)

步骤3:转换为数值列表

把捕获到的字符串分割并转为数值型向量(即R中的列表形式):

# 分割字符串并转换为数值
num_list <- as.numeric(strsplit(trimws(data_match), "\\s*,\\s*")[[1]])
# 去除原数组末尾逗号导致的空值
num_list <- num_list[!is.na(num_list)]

完整示例代码

# 你的原有代码
textprog <- remDr$findElement(using="xpath","/html/body/div1/div1/div/div[2]/div/div/div[2]/div[4]/div[2]/div[2]/div[3]/div/div/script")$getElementAttribute('outerHTML')[1]
doc <- htmlParse(textprog)

# 提取script内容
script_content <- xpathSApply(doc, "//script", xmlValue)[[1]]

# 匹配并提取数值
library(stringr)
data_match <- str_match(script_content, "data:\\s*\\[(.*?)\\]")[,2]
num_list <- as.numeric(strsplit(trimws(data_match), "\\s*,\\s*")[[1]])
num_list <- num_list[!is.na(num_list)]

# 查看结果
print(num_list)

处理后num_list即为你需要的数值列表,输出结果为[1] 46 59 41 25 66 2 97 9。

内容的提问来源于stack exchange,提问作者WickSchozen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 18:30:13