Python实现提取Yandex反向图片搜索结果中的目标URL
提取Yandex反向图片搜索的图片来源URL方法
直接通过字符串截取HTML标签里的URL容易出错,推荐用HTML解析库BeautifulSoup来处理,步骤如下:
1. 安装依赖库
执行命令安装所需工具:
pip install requests beautifulsoup4
2. 编写解析代码
通过解析HTML结构提取目标URL,同时处理Yandex的跳转链接(Yandex会把真实URL封装在跳转链接的参数里):
import requests from bs4 import BeautifulSoup from urllib.parse import unquote, urlparse, parse_qs # 替换为你的Yandex反向图片搜索页面地址 target_page_url = "https://yandex.ru/images/search?rpt=imageview&url=..." # 模拟浏览器请求头,避免被反爬 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } # 获取页面源码 response = requests.get(target_page_url, headers=headers) response.encoding = "utf-8" # 解析HTML soup = BeautifulSoup(response.text, "html.parser") # 定位图片来源的<a>标签(需根据当前Yandex页面结构调整class或属性,以下为常见示例) source_a_tags = soup.find_all("a", class_="Link Link_theme_normal") for tag in source_a_tags: raw_href = tag.get("href") if not raw_href: continue # 处理Yandex的跳转链接,提取真实目标URL if raw_href.startswith("https://yandex.ru/clck/jsredir"): parsed = urlparse(raw_href) # 从query参数中取出真实URL real_url = parse_qs(parsed.query).get("url", [""])[0] # 解码URL编码的内容 print(unquote(real_url)) else: # 直接输出非跳转链接 print(raw_href)
注意事项
- 页面结构适配:Yandex的页面class或标签结构可能会更新,需要打开浏览器开发者工具,查看图片来源链接对应的标签属性,调整
find_all的参数。 - 反爬处理:必须设置合法的
User-Agent,如果遇到请求被拦截,可以添加Cookie或使用代理。 - 动态内容处理:如果页面有滚动加载的内容,requests无法获取JS渲染的部分,这种情况需要改用Selenium等工具模拟浏览器操作。
内容的提问来源于stack exchange,提问作者Wiki
相关产品推荐
相关产品推荐

