You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BeautifulSoup携带请求头爬取图片返回base64问题解决方法

问题描述

我有一段用于爬取图片的代码:

import requests, base64
from bs4 import BeautifulSoup


baseurl = "https://www.google.com/search?q=cat&sxsrf=APq-WBuyx07rsOeGlVQpTsxLt262WbhlfA:1650636332756&source=lnms&tbm=shop&sa=X&ved=2ahUKEwjQr5HC66f3AhXxxzgGHejKC9sQ_AUoAXoECAIQAw&biw=1920&bih=937&dpr=1"
headers = {"User-Agent" : "Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:99.0) Gecko/20100101 Firefox/99.0"}

r_images = requests.get(url=baseurl, headers=headers)


soup_for_image = BeautifulSoup(r_images.text, 'html.parser') 
# 查找商品图片
productimages = [] 
product_images = soup_for_image.findAll('img')
for item in product_images:
    # print(item['src'])
    if "data:image/svg+xml" not in item['src']:
        productimages.append(item.get('src'))
print(productimages)

不携带请求头时上述代码运行正常,但添加请求头后爬取到的图片均为base64格式,需要实现在携带请求头的前提下正常爬取目标图片。

原因说明

出现这个差异的核心原因是谷歌搜索的返回逻辑:

  • 不带合法浏览器UA的请求会被判定为简单爬虫请求,返回无懒加载的简化版页面,所有图片地址直接写在img标签的src属性中,直接读取就能拿到真实地址
  • 携带标准浏览器UA的请求会被判定为正常用户访问,返回带懒加载优化的正式页面:首屏src属性里放的都是低清base64占位图,真实图片地址存放在img标签的data-src或data-iurl自定义属性中,等用户滚动到对应位置时前端JS才会替换src触发真实图片加载,所以直接读src只能拿到base64占位内容。
解决方法

不需要修改请求逻辑,只要调整解析规则,优先读取存放真实地址的属性,同时过滤掉base64格式的占位图即可,修改后的核心解析代码如下:

productimages = [] 
product_images = soup_for_image.findAll('img')
for item in product_images:
    # 优先读取懒加载属性里的真实地址,兜底读src
    img_real_url = item.get('data-src') or item.get('data-iurl') or item.get('src')
    # 过滤svg占位图、base64占位图
    if img_real_url and "data:image/svg+xml" not in img_real_url and not img_real_url.startswith('data:image/'):
        productimages.append(img_real_url)
print(productimages)

如果需要获取更高清的原图、降低被反爬拦截的概率,可以在请求头里补充Accept、Accept-Language等字段,更贴近真实浏览器的请求特征。

爬取公开页面内容请遵守目标站点的相关规则与当地法律法规要求。

内容的提问来源于stack exchange,提问作者Awan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.03 07:27:36