You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Requests的图片URL校验过慢,求优化方案

问题描述

我用Requests写了个简单的is_image_url函数,用来校验给定URL是否可达且返回内容是图片。函数能正常运行,但有时候速度特别慢。我经验不足不知道原因,想请教怎么优化这个函数,或者有没有更高效的实现方式。

原函数代码:

def is_image_url(url):
    try:
        response = requests.get(url)
        status = response.status_code
        print(status)
    
        if response.status_code == 200:
            content_type = response.headers.get('content-type')
            print(content_type)
            if content_type.startswith('image'):
                return True
        return False
    
    except Exception as e:
        print(e)
        return False

我本来想用这个函数过滤掉无法访问的图片(比如Instagram、Facebook这些平台的图),因为只需要校验10张来自Google自定义搜索API的图片,以为效果会不错。现在我已经把这个函数从应用里移除了,试着通过检查图片宽高来验证它是否存在且可达,但效果不理想。

调用函数的循环代码:

for image in img['items']:
    height = image['image']['height']
    width = image['image']['width']
    imgTitle = image['title']
    imgHtmlTitle = image['htmlTitle']
    imgContext = image['image']['contextLink']
    imgLink = image['link']
    workingImg = is_image_url(imgLink)
    print(f'alt {height}.... larg {width}')
    image_info = {'imgTitle' : imgTitle, 'imgHtmlTitle' : imgHtmlTitle, "imgContext" : imgContext, "imgLink" : imgLink}
    if workingImg:
       image_results.append(image_info)
优化方案

你的函数慢的核心原因有两个:一是用requests.get会完整下载整个图片文件,哪怕是几MB的大图也会全量拉取;二是循环里串行请求,10个请求得挨个等前一个完成才会发下一个。下面给你几个实用的优化方向:

1. 用HEAD请求替代GET,只拿响应头

我们只需要状态码和Content-Type这两个信息,完全不需要下载图片内容。HTTP的HEAD请求只会返回响应头,不会返回响应体,速度能提升一大截。

修改后的函数:

import requests

def is_image_url(url):
    try:
        # 发送HEAD请求,允许重定向(很多图片URL会跳转),设置超时避免挂死
        response = requests.head(url, allow_redirects=True, timeout=5)
        if response.status_code == 200:
            content_type = response.headers.get('content-type')
            # 处理带参数的Content-Type,比如image/png; charset=utf-8
            if content_type and content_type.split(';')[0].startswith('image'):
                return True
        return False
    except Exception as e:
        print(e)
        return False

2. 用GET+stream模式,不下载完整内容

如果有些服务器不支持HEAD请求(虽然很少见),可以用stream=True开启流式下载,只读取响应头后立即关闭连接,同样不用下载整个图片:

def is_image_url(url):
    try:
        response = requests.get(url, stream=True, timeout=5, allow_redirects=True)
        if response.status_code == 200:
            content_type = response.headers.get('content-type')
            if content_type and content_type.split(';')[0].startswith('image'):
                return True
        # 立即关闭连接,避免占用资源
        response.close()
        return False
    except Exception as e:
        print(e)
        return False

3. 并行请求,缩短总耗时

10个请求串行跑的话,总时间是每个请求耗时的总和;如果用多线程并行跑,总时间差不多等于最慢的那个请求的耗时。用Python自带的concurrent.futures.ThreadPoolExecutor就能轻松实现:

先调整函数去掉打印(并行时打印会混乱),再改写循环逻辑:

import requests
from concurrent.futures import ThreadPoolExecutor

def is_image_url(url):
    try:
        response = requests.head(url, allow_redirects=True, timeout=5)
        if response.status_code == 200:
            content_type = response.headers.get('content-type')
            return content_type and content_type.split(';')[0].startswith('image')
        return False
    except Exception:
        return False

def process_images(img_items):
    image_results = []
    # 准备待校验的URL与对应图片信息
    tasks = [(image['link'], image) for image in img_items]
    
    # 用线程池并行处理,最大线程数设为10刚好匹配你的需求
    with ThreadPoolExecutor(max_workers=10) as executor:
        results = executor.map(lambda task: (is_image_url(task[0]), task[1]), tasks)
    
    # 过滤有效图片并整理信息
    for is_valid, image in results:
        if is_valid:
            image_info = {
                'imgTitle': image['title'],
                'imgHtmlTitle': image['htmlTitle'],
                'imgContext': image['image']['contextLink'],
                'imgLink': image['link']
            }
            image_results.append(image_info)
    return image_results

# 调用示例
# image_results = process_images(img['items'])

补充说明

你之前试的检查宽高没用,是因为Google自定义搜索API返回的宽高是它爬取时的元数据,图片可能已经失效、被删除,但元数据还留在API结果里,所以必须校验真实URL的可达性和内容类型。

另外,有些平台(比如Instagram)会对非浏览器请求做拦截,遇到这种情况可以给请求加个模拟浏览器的User-Agent头:

headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'}
response = requests.head(url, headers=headers, allow_redirects=True, timeout=5)

内容的提问来源于stack exchange,提问作者jklzz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 15:17:03