You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python和Requests读取在线图片尺寸时完成时间不稳定问题

Fixing Unstable Load Times When Fetching Web Image Dimensions in Python 2.7

Hey there, let's break down why your image dimension fetching code has inconsistent read times and how to fix it.

First, let's recap your approach: using a Range header to only grab the first 1024 bytes of the image is a smart move to avoid downloading full files—but here are the key issues causing those unpredictable speeds:

  • Server compatibility with Range requests: Not all servers honor the Range header. If a server doesn't support partial content, it'll send the entire image instead, which drastically slows things down.
  • Missing timeout handling: Without a timeout, your request could hang indefinitely if the server is slow or unresponsive, leading to random long waits.
  • PIL's overhead for partial data: PIL isn't optimized for reading truncated image data in every scenario, which might add unexpected processing delays.

Here are actionable fixes you can implement right away:

1. Validate Partial Content Support First

Before sending a Range request, check if the server allows partial content with a HEAD request. This lets you fall back to fetching the full image only when necessary:

import requests
from io import BytesIO
from PIL import Image

def get_web_img_dimensions(url):
    """Returns image dimensions with optimized range handling"""
    headers = {'User-Agent': 'Mozilla/5.0 (Windows; U; Windows NT 5.1; en-US; rv:1.9.0.1) Gecko/2008070208 Firefox/3.0.1'}
    
    # Check if server supports partial content
    try:
        head_req = requests.head(url, headers=headers, timeout=5)
        head_req.raise_for_status()
        accept_ranges = head_req.headers.get('Accept-Ranges')
        
        # Use range request only if supported
        if accept_ranges == 'bytes':
            req = requests.get(url, headers={**headers, "Range": "bytes=0-1023"}, timeout=5)
        else:
            req = requests.get(url, headers=headers, timeout=5)
        req.raise_for_status()
        
        # Read dimensions
        img_data = BytesIO(req.content)
        with Image.open(img_data) as img:
            return img.size
    except requests.exceptions.RequestException as e:
        print(f"Request failed: {e}")
        return None

2. Add Timeout Parameters

Always include a timeout in your requests to prevent hanging. I added a 5-second timeout in the example above—adjust it based on your network's typical speed.

3. Use a Lighter-Weight Library for Dimension Reading

PIL is great for full image processing, but for just reading dimensions, libraries like imagesize are faster and don't require loading the entire image into memory. It works perfectly with Python 2.7:
First install it:

pip install imagesize

Then modify your code:

import requests
import imagesize
from io import BytesIO

def get_web_img_dimensions(url):
    """Returns image dimensions using a lightweight library"""
    headers = {'User-Agent': 'Mozilla/5.0 (Windows; U; Windows NT 5.1; en-US; rv:1.9.0.1) Gecko/2008070208 Firefox/3.0.1'}
    
    try:
        # Check partial content support first
        head_req = requests.head(url, headers=headers, timeout=5)
        head_req.raise_for_status()
        accept_ranges = head_req.headers.get('Accept-Ranges')
        
        if accept_ranges == 'bytes':
            req = requests.get(url, headers={**headers, "Range": "bytes=0-1023"}, timeout=5)
        else:
            req = requests.get(url, headers=headers, timeout=5)
        req.raise_for_status()
        
        # Read dimensions without full image parsing
        return imagesize.get(BytesIO(req.content))
    except requests.exceptions.RequestException as e:
        print(f"Request failed: {e}")
        return None
    except imagesize.UnknownImageFormatError:
        print("Unsupported image format")
        return None

4. Cache Results (If Reusing URLs)

If you're fetching dimensions for the same URL multiple times, cache the results to avoid redundant requests. A simple dictionary works for basic use cases:

dimensions_cache = {}

def get_web_img_dimensions(url):
    if url in dimensions_cache:
        return dimensions_cache[url]
    
    # ... rest of the code from option 2 ...
    
    dimensions = imagesize.get(BytesIO(req.content))
    dimensions_cache[url] = dimensions
    return dimensions

Testing with Your URLs

Both of your test URLs support partial content, so the Range request should work consistently once you add timeout handling and validate support.

内容的提问来源于stack exchange,提问作者Hassan Baig

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:01:34