基于Requests的图片URL校验过慢,求优化方案
我用Requests写了个简单的is_image_url函数,用来校验给定URL是否可达且返回内容是图片。函数能正常运行,但有时候速度特别慢。我经验不足不知道原因,想请教怎么优化这个函数,或者有没有更高效的实现方式。
原函数代码:
def is_image_url(url): try: response = requests.get(url) status = response.status_code print(status) if response.status_code == 200: content_type = response.headers.get('content-type') print(content_type) if content_type.startswith('image'): return True return False except Exception as e: print(e) return False
我本来想用这个函数过滤掉无法访问的图片(比如Instagram、Facebook这些平台的图),因为只需要校验10张来自Google自定义搜索API的图片,以为效果会不错。现在我已经把这个函数从应用里移除了,试着通过检查图片宽高来验证它是否存在且可达,但效果不理想。
调用函数的循环代码:
for image in img['items']: height = image['image']['height'] width = image['image']['width'] imgTitle = image['title'] imgHtmlTitle = image['htmlTitle'] imgContext = image['image']['contextLink'] imgLink = image['link'] workingImg = is_image_url(imgLink) print(f'alt {height}.... larg {width}') image_info = {'imgTitle' : imgTitle, 'imgHtmlTitle' : imgHtmlTitle, "imgContext" : imgContext, "imgLink" : imgLink} if workingImg: image_results.append(image_info)
你的函数慢的核心原因有两个:一是用requests.get会完整下载整个图片文件,哪怕是几MB的大图也会全量拉取;二是循环里串行请求,10个请求得挨个等前一个完成才会发下一个。下面给你几个实用的优化方向:
1. 用HEAD请求替代GET,只拿响应头
我们只需要状态码和Content-Type这两个信息,完全不需要下载图片内容。HTTP的HEAD请求只会返回响应头,不会返回响应体,速度能提升一大截。
修改后的函数:
import requests def is_image_url(url): try: # 发送HEAD请求,允许重定向(很多图片URL会跳转),设置超时避免挂死 response = requests.head(url, allow_redirects=True, timeout=5) if response.status_code == 200: content_type = response.headers.get('content-type') # 处理带参数的Content-Type,比如image/png; charset=utf-8 if content_type and content_type.split(';')[0].startswith('image'): return True return False except Exception as e: print(e) return False
2. 用GET+stream模式,不下载完整内容
如果有些服务器不支持HEAD请求(虽然很少见),可以用stream=True开启流式下载,只读取响应头后立即关闭连接,同样不用下载整个图片:
def is_image_url(url): try: response = requests.get(url, stream=True, timeout=5, allow_redirects=True) if response.status_code == 200: content_type = response.headers.get('content-type') if content_type and content_type.split(';')[0].startswith('image'): return True # 立即关闭连接,避免占用资源 response.close() return False except Exception as e: print(e) return False
3. 并行请求,缩短总耗时
10个请求串行跑的话,总时间是每个请求耗时的总和;如果用多线程并行跑,总时间差不多等于最慢的那个请求的耗时。用Python自带的concurrent.futures.ThreadPoolExecutor就能轻松实现:
先调整函数去掉打印(并行时打印会混乱),再改写循环逻辑:
import requests from concurrent.futures import ThreadPoolExecutor def is_image_url(url): try: response = requests.head(url, allow_redirects=True, timeout=5) if response.status_code == 200: content_type = response.headers.get('content-type') return content_type and content_type.split(';')[0].startswith('image') return False except Exception: return False def process_images(img_items): image_results = [] # 准备待校验的URL与对应图片信息 tasks = [(image['link'], image) for image in img_items] # 用线程池并行处理,最大线程数设为10刚好匹配你的需求 with ThreadPoolExecutor(max_workers=10) as executor: results = executor.map(lambda task: (is_image_url(task[0]), task[1]), tasks) # 过滤有效图片并整理信息 for is_valid, image in results: if is_valid: image_info = { 'imgTitle': image['title'], 'imgHtmlTitle': image['htmlTitle'], 'imgContext': image['image']['contextLink'], 'imgLink': image['link'] } image_results.append(image_info) return image_results # 调用示例 # image_results = process_images(img['items'])
补充说明
你之前试的检查宽高没用,是因为Google自定义搜索API返回的宽高是它爬取时的元数据,图片可能已经失效、被删除,但元数据还留在API结果里,所以必须校验真实URL的可达性和内容类型。
另外,有些平台(比如Instagram)会对非浏览器请求做拦截,遇到这种情况可以给请求加个模拟浏览器的User-Agent头:
headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'} response = requests.head(url, headers=headers, allow_redirects=True, timeout=5)
内容的提问来源于stack exchange,提问作者jklzz

