BeautifulSoup携带请求头爬取图片返回base64问题解决方法
问题描述
我有一段用于爬取图片的代码:
import requests, base64 from bs4 import BeautifulSoup baseurl = "https://www.google.com/search?q=cat&sxsrf=APq-WBuyx07rsOeGlVQpTsxLt262WbhlfA:1650636332756&source=lnms&tbm=shop&sa=X&ved=2ahUKEwjQr5HC66f3AhXxxzgGHejKC9sQ_AUoAXoECAIQAw&biw=1920&bih=937&dpr=1" headers = {"User-Agent" : "Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:99.0) Gecko/20100101 Firefox/99.0"} r_images = requests.get(url=baseurl, headers=headers) soup_for_image = BeautifulSoup(r_images.text, 'html.parser') # 查找商品图片 productimages = [] product_images = soup_for_image.findAll('img') for item in product_images: # print(item['src']) if "data:image/svg+xml" not in item['src']: productimages.append(item.get('src')) print(productimages)
不携带请求头时上述代码运行正常,但添加请求头后爬取到的图片均为base64格式,需要实现在携带请求头的前提下正常爬取目标图片。
原因说明
出现这个差异的核心原因是谷歌搜索的返回逻辑:
- 不带合法浏览器UA的请求会被判定为简单爬虫请求,返回无懒加载的简化版页面,所有图片地址直接写在img标签的
src属性中,直接读取就能拿到真实地址 - 携带标准浏览器UA的请求会被判定为正常用户访问,返回带懒加载优化的正式页面:首屏
src属性里放的都是低清base64占位图,真实图片地址存放在img标签的data-src或data-iurl自定义属性中,等用户滚动到对应位置时前端JS才会替换src触发真实图片加载,所以直接读src只能拿到base64占位内容。
解决方法
不需要修改请求逻辑,只要调整解析规则,优先读取存放真实地址的属性,同时过滤掉base64格式的占位图即可,修改后的核心解析代码如下:
productimages = [] product_images = soup_for_image.findAll('img') for item in product_images: # 优先读取懒加载属性里的真实地址,兜底读src img_real_url = item.get('data-src') or item.get('data-iurl') or item.get('src') # 过滤svg占位图、base64占位图 if img_real_url and "data:image/svg+xml" not in img_real_url and not img_real_url.startswith('data:image/'): productimages.append(img_real_url) print(productimages)
如果需要获取更高清的原图、降低被反爬拦截的概率,可以在请求头里补充Accept、Accept-Language等字段,更贴近真实浏览器的请求特征。
爬取公开页面内容请遵守目标站点的相关规则与当地法律法规要求。
内容的提问来源于stack exchange,提问作者Awan
相关产品推荐
相关产品推荐

