如何修改R爬虫脚本以从必应图片搜索获取目标URL?
必应图片搜索爬虫代码修改方案
我正在构建用于从搜索引擎抓取产品图片的R爬虫脚本,目前已通过以下代码成功从谷歌图片搜索获取包含图片的URL:
google_urls <- GET("https://www.google.com/search?q=WWF%20CUB%20CLUB%20WWF16215003&tbm=isch", user_agent = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/93.0.4577.63 Safari/537.36") %>% read_html() %>% html_nodes(xpath = "//td/a") %>% html_attr("href") %>% `[`(str_detect(., "/url\\?")) %>% strsplit("=|\\&") %>% sapply(`[`, 2)
为扩展爬虫的搜索范围,我希望同样从必应图片搜索抓取类似URL,但复用以下代码后bing_urls为空,无法得到结果:
bing_urls <- GET("https://www.bing.com/images/search?q=WWF%20CUB%20CLUB%20WWF16215003", user_agent = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/93.0.4577.63 Safari/537.36") %>% read_html() %>% html_nodes(xpath = "//td/a") %>% html_attr("href") %>% `[`(str_detect(., "/url\\?")) %>% strsplit("=|\\&") %>% sapply(`[`, 2)
修改后的必应爬虫代码
library(httr) library(rvest) library(stringr) bing_urls <- GET("https://www.bing.com/images/search?q=WWF%20CUB%20CLUB%20WWF16215003", user_agent = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/93.0.4577.63 Safari/537.36") %>% read_html() %>% # 定位必应图片条目中的a标签 html_nodes(xpath = "//div[@class='iusc']/a") %>% html_attr("href") %>% # 提取href中的imgurl参数 str_extract("imgurl=(.*?)&") %>% # 清理参数前缀和后缀 str_remove("imgurl=") %>% str_remove("&") %>% # URL解码得到真实图片地址 URLdecode()
修改说明
必应图片搜索的HTML结构与谷歌完全不同:
- 原代码中的
//td/a选择器无法匹配必应的图片链接元素,必应的图片条目通常包裹在class为iusc的div容器内; - 必应的图片原始URL以
imgurl参数的形式编码在a标签的href属性中,需要通过字符串提取和解码才能得到真实地址; - 使用
URLdecode()可以还原被编码的特殊字符,确保URL可直接访问。
内容的提问来源于stack exchange,提问作者nba2020
相关产品推荐
相关产品推荐

