使用Python和Requests读取在线图片尺寸时完成时间不稳定问题
Hey there, let's break down why your image dimension fetching code has inconsistent read times and how to fix it.
First, let's recap your approach: using a Range header to only grab the first 1024 bytes of the image is a smart move to avoid downloading full files—but here are the key issues causing those unpredictable speeds:
- Server compatibility with Range requests: Not all servers honor the
Rangeheader. If a server doesn't support partial content, it'll send the entire image instead, which drastically slows things down. - Missing timeout handling: Without a timeout, your request could hang indefinitely if the server is slow or unresponsive, leading to random long waits.
- PIL's overhead for partial data: PIL isn't optimized for reading truncated image data in every scenario, which might add unexpected processing delays.
Here are actionable fixes you can implement right away:
1. Validate Partial Content Support First
Before sending a Range request, check if the server allows partial content with a HEAD request. This lets you fall back to fetching the full image only when necessary:
import requests from io import BytesIO from PIL import Image def get_web_img_dimensions(url): """Returns image dimensions with optimized range handling""" headers = {'User-Agent': 'Mozilla/5.0 (Windows; U; Windows NT 5.1; en-US; rv:1.9.0.1) Gecko/2008070208 Firefox/3.0.1'} # Check if server supports partial content try: head_req = requests.head(url, headers=headers, timeout=5) head_req.raise_for_status() accept_ranges = head_req.headers.get('Accept-Ranges') # Use range request only if supported if accept_ranges == 'bytes': req = requests.get(url, headers={**headers, "Range": "bytes=0-1023"}, timeout=5) else: req = requests.get(url, headers=headers, timeout=5) req.raise_for_status() # Read dimensions img_data = BytesIO(req.content) with Image.open(img_data) as img: return img.size except requests.exceptions.RequestException as e: print(f"Request failed: {e}") return None
2. Add Timeout Parameters
Always include a timeout in your requests to prevent hanging. I added a 5-second timeout in the example above—adjust it based on your network's typical speed.
3. Use a Lighter-Weight Library for Dimension Reading
PIL is great for full image processing, but for just reading dimensions, libraries like imagesize are faster and don't require loading the entire image into memory. It works perfectly with Python 2.7:
First install it:
pip install imagesize
Then modify your code:
import requests import imagesize from io import BytesIO def get_web_img_dimensions(url): """Returns image dimensions using a lightweight library""" headers = {'User-Agent': 'Mozilla/5.0 (Windows; U; Windows NT 5.1; en-US; rv:1.9.0.1) Gecko/2008070208 Firefox/3.0.1'} try: # Check partial content support first head_req = requests.head(url, headers=headers, timeout=5) head_req.raise_for_status() accept_ranges = head_req.headers.get('Accept-Ranges') if accept_ranges == 'bytes': req = requests.get(url, headers={**headers, "Range": "bytes=0-1023"}, timeout=5) else: req = requests.get(url, headers=headers, timeout=5) req.raise_for_status() # Read dimensions without full image parsing return imagesize.get(BytesIO(req.content)) except requests.exceptions.RequestException as e: print(f"Request failed: {e}") return None except imagesize.UnknownImageFormatError: print("Unsupported image format") return None
4. Cache Results (If Reusing URLs)
If you're fetching dimensions for the same URL multiple times, cache the results to avoid redundant requests. A simple dictionary works for basic use cases:
dimensions_cache = {} def get_web_img_dimensions(url): if url in dimensions_cache: return dimensions_cache[url] # ... rest of the code from option 2 ... dimensions = imagesize.get(BytesIO(req.content)) dimensions_cache[url] = dimensions return dimensions
Testing with Your URLs
Both of your test URLs support partial content, so the Range request should work consistently once you add timeout handling and validate support.
内容的提问来源于stack exchange,提问作者Hassan Baig

