You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python爬虫中下载Google图片的高清版本?

如何用Python爬取Google图片的高清原图

我用Python 3练习网页爬虫,尝试爬取Google图片,但下载的都是低分辨率的预览图(示例:低分辨率预览图)。

现有代码如下:

import requests, webbrowser, bs4

# this code specifically looks for cats
res = requests.get(f"https://www.google.com/search?q=cat&tbm=isch")
res.raise_for_status()

soup = bs4.BeautifulSoup(res.text, "html.parser")

# gets the first 5 images from google images
searchImgs = soup.select("img")[:6]         

if searchImgs == []:
    sys.exit("No images found. Exiting..")
else:
    print("Images found. Opening browser..")

    # opens the browser with the images 
    for i in range(len(searchImgs)):
        webbrowser.open(searchImgs[i].get("src"))

问题出在直接获取了页面上的预览图链接,要获取高清原图得拿到图片的原始链接,但我不知道怎么在Python里获取并下载这些链接,试过找方法但没成功。


问题根源

你当前代码选中的img标签的src属性是Google生成的低分辨率缩略图,高清原图的链接藏在页面的专属属性或脚本数据中。另外,直接用requests请求Google图片页面会触发反爬机制,需要模拟浏览器请求才能正常访问。

解决方案

  1. 添加请求头:模拟浏览器发送请求,规避Google的反爬拦截。
  2. 提取高清链接:Google图片的高清原图链接通常存储在img标签的data-src属性中(针对当前页面结构),直接提取该属性即可获取原图地址。

以下是修改后的可运行代码,支持获取并下载高清原图:

import requests
import bs4
import os
import sys

# 替换为你自己浏览器的User-Agent,可在开发者工具中查看
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

search_query = "cat"
url = f"https://www.google.com/search?q={search_query}&tbm=isch"

# 发送请求
res = requests.get(url, headers=headers)
try:
    res.raise_for_status()
except requests.exceptions.HTTPError as e:
    print(f"请求页面失败: {e}")
    sys.exit(1)

soup = bs4.BeautifulSoup(res.text, "html.parser")

# 提取带高清链接的img元素,取前5张
img_elements = soup.select("img[data-src]")[:5]

if not img_elements:
    print("未找到高清图片链接")
    sys.exit(1)

# 创建图片保存目录
save_dir = f"{search_query}_hd_images"
os.makedirs(save_dir, exist_ok=True)

# 下载并保存图片
for idx, elem in enumerate(img_elements):
    img_url = elem.get("data-src")
    if not img_url:
        print(f"第{idx+1}张图片无高清链接,跳过")
        continue
    
    try:
        img_res = requests.get(img_url, headers=headers)
        img_res.raise_for_status()
        
        # 保存图片,自动处理后缀
        img_ext = img_url.split(".")[-1].split("?")[0]
        save_path = os.path.join(save_dir, f"cat_hd_{idx+1}.{img_ext}")
        
        with open(save_path, "wb") as f:
            f.write(img_res.content)
        
        print(f"已保存: {save_path}")
    except Exception as e:
        print(f"第{idx+1}张图片下载失败: {e}")

注意事项

  • User-Agent替换:必须使用真实浏览器的User-Agent,否则会被Google拒绝请求。
  • 页面结构变化:Google图片的页面结构可能随时调整,如果data-src属性失效,可在浏览器开发者工具中重新查找高清链接所在的属性(比如srcset)。
  • 反爬限制:不要频繁发送请求,必要时可添加time.sleep(1)等间隔,避免IP被封禁。

内容的提问来源于stack exchange,提问作者xtourmaline

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 04:43:36