如何在Python爬虫中下载Google图片的高清版本?
如何用Python爬取Google图片的高清原图
我用Python 3练习网页爬虫,尝试爬取Google图片,但下载的都是低分辨率的预览图(示例:
)。
现有代码如下:
import requests, webbrowser, bs4 # this code specifically looks for cats res = requests.get(f"https://www.google.com/search?q=cat&tbm=isch") res.raise_for_status() soup = bs4.BeautifulSoup(res.text, "html.parser") # gets the first 5 images from google images searchImgs = soup.select("img")[:6] if searchImgs == []: sys.exit("No images found. Exiting..") else: print("Images found. Opening browser..") # opens the browser with the images for i in range(len(searchImgs)): webbrowser.open(searchImgs[i].get("src"))
问题出在直接获取了页面上的预览图链接,要获取高清原图得拿到图片的原始链接,但我不知道怎么在Python里获取并下载这些链接,试过找方法但没成功。
问题根源
你当前代码选中的img标签的src属性是Google生成的低分辨率缩略图,高清原图的链接藏在页面的专属属性或脚本数据中。另外,直接用requests请求Google图片页面会触发反爬机制,需要模拟浏览器请求才能正常访问。
解决方案
- 添加请求头:模拟浏览器发送请求,规避Google的反爬拦截。
- 提取高清链接:Google图片的高清原图链接通常存储在
img标签的data-src属性中(针对当前页面结构),直接提取该属性即可获取原图地址。
以下是修改后的可运行代码,支持获取并下载高清原图:
import requests import bs4 import os import sys # 替换为你自己浏览器的User-Agent,可在开发者工具中查看 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } search_query = "cat" url = f"https://www.google.com/search?q={search_query}&tbm=isch" # 发送请求 res = requests.get(url, headers=headers) try: res.raise_for_status() except requests.exceptions.HTTPError as e: print(f"请求页面失败: {e}") sys.exit(1) soup = bs4.BeautifulSoup(res.text, "html.parser") # 提取带高清链接的img元素,取前5张 img_elements = soup.select("img[data-src]")[:5] if not img_elements: print("未找到高清图片链接") sys.exit(1) # 创建图片保存目录 save_dir = f"{search_query}_hd_images" os.makedirs(save_dir, exist_ok=True) # 下载并保存图片 for idx, elem in enumerate(img_elements): img_url = elem.get("data-src") if not img_url: print(f"第{idx+1}张图片无高清链接,跳过") continue try: img_res = requests.get(img_url, headers=headers) img_res.raise_for_status() # 保存图片,自动处理后缀 img_ext = img_url.split(".")[-1].split("?")[0] save_path = os.path.join(save_dir, f"cat_hd_{idx+1}.{img_ext}") with open(save_path, "wb") as f: f.write(img_res.content) print(f"已保存: {save_path}") except Exception as e: print(f"第{idx+1}张图片下载失败: {e}")
注意事项
- User-Agent替换:必须使用真实浏览器的User-Agent,否则会被Google拒绝请求。
- 页面结构变化:Google图片的页面结构可能随时调整,如果
data-src属性失效,可在浏览器开发者工具中重新查找高清链接所在的属性(比如srcset)。 - 反爬限制:不要频繁发送请求,必要时可添加
time.sleep(1)等间隔,避免IP被封禁。
内容的提问来源于stack exchange,提问作者xtourmaline
相关产品推荐
相关产品推荐

